Back to Projects

Artha

Local-First RAG Platform

Architected a local-first RAG platform with FastAPI and Next.js, designed to run without external model APIs for privacy-sensitive document search.

Role: Full Stack Developer
Duration: 3 months
Team: Solo

Tech Stack

PythonFastAPILangGraphPostgreSQLRedisOllama

Key Highlights

Local-first RAG architecture for privacy-sensitive document search
Asynchronous ingestion pipeline with Redis and Celery
Hybrid retrieval with vector search, trigram matching, RRF, and reranking
LangGraph workflows with query rewriting, quality gates, and HyDE fallback

Architecture

Artha uses a local-first architecture with FastAPI handling the API layer and Next.js providing the frontend. The RAG pipeline consists of: 1. **Document Ingestion**: Documents are parsed, hierarchically chunked, embedded, and processed asynchronously using Redis and Celery to avoid blocking user workflows. 2. **Query Processing**: User queries pass through query rewriting, retrieval quality checks, and HyDE fallback paths when confidence is low. 3. **Retrieval**: Hybrid retrieval combines vector similarity, trigram search, Reciprocal Rank Fusion (RRF), and reranking to improve recall and precision on technical documents. 4. **Generation**: Context is passed to local Ollama models for grounded response generation, with GraphRAG-inspired relationship-aware retrieval patterns explored to improve context selection.

Challenges

  • 01Handling document versioning and updates without re-indexing entire corpora
  • 02Optimizing local model performance and retrieval quality on technical documents
  • 03Designing fault-tolerant LangGraph workflows with quality gates and fallbacks
  • 04Balancing retrieval recall, precision, and latency in a hybrid search pipeline

Key Learnings

  • Local-first systems require careful orchestration of compute, storage, and queueing
  • Hybrid retrieval significantly improves robustness over single-strategy search
  • LangGraph works well for resilient retrieval and decision workflows
  • Asynchronous ingestion is essential for scalable document processing

Future Work

  • →Multi-modal document support for images, tables, and diagrams
  • →Collaborative knowledge base features with fine-grained access control
  • →Advanced retrieval evaluation and observability dashboards
← Back to Projects