Skip to main content
Hemanth.
Back to Feed
2025AI & Data

HemGPT: Enterprise RAG System

A production-grade RAG platform enabling secure multi-document reasoning over 10,000+ internal documents via hybrid vector search, Reciprocal Rank Fusion, and local Ollama inference.

Rationale

Why HemGPT? Enterprise data privacy and vendor cost management. Sending sensitive internal documents to third-party LLM APIs risks compliance violations and compounding per-token costs at scale. By deploying local models via Ollama alongside ChromaDB and BM25, we achieve zero data leakage while maintaining high precision.

The Hardest Challenge: Context window budgeting and chunking drift. I implemented a parent-child chunking hierarchy and sentence-level context compression. The retriever searches fine-grained child chunks for pinpoint accuracy, but inflates the surrounding parent context for the LLM generator.

Tech Stack

PythonFastAPILangChainChromaDBBM25OllamaCross-EncoderCeleryRedisPrometheusJaegerDocker
System Architecture

Core Engineering Challenge

The Bottleneck: Pure vector search (ChromaDB) missed exact keyword queries, while keyword search (BM25) missed semantic intent.

The Solution

Implemented a parallel hybrid retrieval pipeline executing dense and sparse queries simultaneously, fusing candidate scores with Reciprocal Rank Fusion (RRF) and filtering via Cross-Encoder reranking.

The Trade-off

Reranking introduces a ~150ms latency overhead, but boosts context precision by over 50%, ensuring higher accuracy before prompt delivery.

Key Highlights

  • ▹Engineered a hybrid retrieval engine pairing ChromaDB vector embeddings with BM25 keyword search, fused via Reciprocal Rank Fusion.
  • ▹Implemented cross-encoder reranking and context compression to eliminate low-relevance passages before LLM context injection.
  • ▹Integrated local Ollama inference for query expansion and token streaming with zero external API data leakage.
  • ▹Built async background ingestion via FastAPI and Celery workers, guarded by Redis rate limiters, caching, and circuit breakers.
  • ▹Configured full observability with Prometheus metrics, Grafana dashboards, Jaeger distributed tracing, and Kubernetes HPA auto-scaling.

Architecture Details

The HemGPT RAG System decouples ingestion, hybrid retrieval, and generation to handle heavy concurrent query loads reliably.

1. Data Ingestion & Async Processing

  • FastAPI handles file validation and queues ingestion jobs to Celery workers with Redis backends.
  • Workers perform semantic chunking, parent-child splitting, and async vector index generation without blocking the web thread.

2. Hybrid Retrieval & RRF Fusion

  • Queries undergo Query Expansion via Ollama to generate parallel search variations.
  • Search queries execute concurrently against ChromaDB (dense semantic vector search) and BM25/Elasticsearch (sparse lexical search).
  • Reciprocal Rank Fusion (RRF) merges candidate lists, followed by a Cross-Encoder Reranker that scores exact passage relevance.

3. Fault Tolerance & Observability

  • Redis provides response caching, token-bucket rate limiting, and circuit breakers around inference calls.
  • Prometheus and Jaeger capture query-level latency histograms and distributed traces across vector lookup and model generation steps.

Interactive System Design

Click nodes to inspect engineering flow

FastAPI Rate-Limited Gateway

FastAPIRedisPydantic
Engineering Insight

Validates incoming queries rapidly while checking strict sliding-window rate limits via Redis to prevent unauth LLM abuse.

Platform Showcase

HemGPT: Enterprise RAG System Interface Banner

Technical Walkthrough

Engineering Insights

The Reality of Production RAG

"An LLM-as-a-judge that agrees with itself is not an eval."

Level 1: Prompt Wrapper

Direct API calls without context grounding or local memory.

Level 2: Basic Vector Search

Naive vector retrieval stuck with single-embedding recall gaps.

Level 3: Hybrid Retrieval & Reranking

BM25 + Dense vector search fused with cross-encoder precision.

Level 4: Observability & Resilience

Circuit breakers, tracing, async ingestion, and containerized scale.

Moving from basic RAG to a production system requires rigorous infrastructure. Implementing hybrid retrieval (BM25 + Vector), Reciprocal Rank Fusion, cross-encoder reranking, and Jaeger tracing proved that query accuracy is an architectural problem, not just a prompt engineering exercise.