HemGPT: Enterprise RAG System
A production-grade RAG platform enabling secure multi-document reasoning over 10,000+ internal documents via hybrid vector search, Reciprocal Rank Fusion, and local Ollama inference.
Rationale
Why HemGPT? Enterprise data privacy and vendor cost management. Sending sensitive internal documents to third-party LLM APIs risks compliance violations and compounding per-token costs at scale. By deploying local models via Ollama alongside ChromaDB and BM25, we achieve zero data leakage while maintaining high precision.
The Hardest Challenge: Context window budgeting and chunking drift. I implemented a parent-child chunking hierarchy and sentence-level context compression. The retriever searches fine-grained child chunks for pinpoint accuracy, but inflates the surrounding parent context for the LLM generator.
Tech Stack

Core Engineering Challenge
The Bottleneck: Pure vector search (ChromaDB) missed exact keyword queries, while keyword search (BM25) missed semantic intent.
The Solution
Implemented a parallel hybrid retrieval pipeline executing dense and sparse queries simultaneously, fusing candidate scores with Reciprocal Rank Fusion (RRF) and filtering via Cross-Encoder reranking.
The Trade-off
Reranking introduces a ~150ms latency overhead, but boosts context precision by over 50%, ensuring higher accuracy before prompt delivery.
Key Highlights
- ▹Engineered a hybrid retrieval engine pairing ChromaDB vector embeddings with BM25 keyword search, fused via Reciprocal Rank Fusion.
- ▹Implemented cross-encoder reranking and context compression to eliminate low-relevance passages before LLM context injection.
- ▹Integrated local Ollama inference for query expansion and token streaming with zero external API data leakage.
- ▹Built async background ingestion via FastAPI and Celery workers, guarded by Redis rate limiters, caching, and circuit breakers.
- ▹Configured full observability with Prometheus metrics, Grafana dashboards, Jaeger distributed tracing, and Kubernetes HPA auto-scaling.
Architecture Details
The HemGPT RAG System decouples ingestion, hybrid retrieval, and generation to handle heavy concurrent query loads reliably.
1. Data Ingestion & Async Processing
- FastAPI handles file validation and queues ingestion jobs to Celery workers with Redis backends.
- Workers perform semantic chunking, parent-child splitting, and async vector index generation without blocking the web thread.
2. Hybrid Retrieval & RRF Fusion
- Queries undergo Query Expansion via Ollama to generate parallel search variations.
- Search queries execute concurrently against ChromaDB (dense semantic vector search) and BM25/Elasticsearch (sparse lexical search).
- Reciprocal Rank Fusion (RRF) merges candidate lists, followed by a Cross-Encoder Reranker that scores exact passage relevance.
3. Fault Tolerance & Observability
- Redis provides response caching, token-bucket rate limiting, and circuit breakers around inference calls.
- Prometheus and Jaeger capture query-level latency histograms and distributed traces across vector lookup and model generation steps.
Interactive System Design
FastAPI Rate-Limited Gateway
Engineering Insight
Validates incoming queries rapidly while checking strict sliding-window rate limits via Redis to prevent unauth LLM abuse.
Platform Showcase

Technical Walkthrough
Engineering Insights
The Reality of Production RAG
"An LLM-as-a-judge that agrees with itself is not an eval."
Level 1: Prompt Wrapper
Direct API calls without context grounding or local memory.
Level 2: Basic Vector Search
Naive vector retrieval stuck with single-embedding recall gaps.
Level 3: Hybrid Retrieval & Reranking
BM25 + Dense vector search fused with cross-encoder precision.
Level 4: Observability & Resilience
Circuit breakers, tracing, async ingestion, and containerized scale.
Moving from basic RAG to a production system requires rigorous infrastructure. Implementing hybrid retrieval (BM25 + Vector), Reciprocal Rank Fusion, cross-encoder reranking, and Jaeger tracing proved that query accuracy is an architectural problem, not just a prompt engineering exercise.