Architected a local-first RAG platform with FastAPI and Next.js, designed to run without external model APIs for privacy-sensitive document search.
Artha uses a local-first architecture with FastAPI handling the API layer and Next.js providing the frontend. The RAG pipeline consists of: 1. **Document Ingestion**: Documents are parsed, hierarchically chunked, embedded, and processed asynchronously using Redis and Celery to avoid blocking user workflows. 2. **Query Processing**: User queries pass through query rewriting, retrieval quality checks, and HyDE fallback paths when confidence is low. 3. **Retrieval**: Hybrid retrieval combines vector similarity, trigram search, Reciprocal Rank Fusion (RRF), and reranking to improve recall and precision on technical documents. 4. **Generation**: Context is passed to local Ollama models for grounded response generation, with GraphRAG-inspired relationship-aware retrieval patterns explored to improve context selection.