A production Retrieval-Augmented Generation system that answers Indian legal queries with paragraph-level source citations. In sub-500ms retrieval, across 8 Indian languages, with near-zero hallucinations.
India's legal ecosystem runs on judgments. Every argument in court, every contract dispute, every advisory memo starts with a question: what did courts say about this before? Answering that question the traditional way means opening Indian Kanoon or Manupatra, running keyword searches, opening 20 tabs, reading through irrelevant hits, and finally landing on the two or three cases that matter. After 3 to 6 hours of research.
When we started building VakeelSaathi (Decipher's legal operations SaaS), the founding hypothesis was simple: a lawyer should be able to ask "what have courts held about arbitral awards being challenged under Section 34?" and get a cited answer in 30 seconds.
Building it was harder than the hypothesis suggested.
Ask ChatGPT the same question and it will confidently cite Union of India v. Some Case, (2018) 4 SCC 123. A citation it just made up. This is the hallucination problem. In legal work, a fabricated citation isn't a bug; it's malpractice. We needed a system where the model could only quote from real, verifiable judgments.
VakeelSaathi's legal Q&A runs on a RAG pipeline. Retrieval-Augmented Generation is straightforward in theory. Fetch relevant documents, give them to an LLM as context, let the LLM answer using only what you gave it. In practice, at 50,000+ documents and Indian legal complexity, every step required careful engineering.
Before locking in Claude, we ran evaluations on 500 curated Indian legal questions across GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, and Llama 3.1 70B (self-hosted). Grading criteria: factual accuracy, correct citation format, sensitivity to jurisdiction (SC vs High Court binding force), and refusal quality on out-of-scope queries.
Results:
| Model | Accuracy | Notes |
|---|---|---|
| Claude 3.5 Sonnet | 89% | Near-perfect citation format, best refusal behavior |
| GPT-4o | 84% | Occasional citation fabrications, slightly weaker on Indian statute interpretation |
| Gemini 1.5 Pro | 79% | Verbose answers, unreliable citation format |
| Llama 3.1 70B (self-hosted) | 71% | Cheapest to run, best for high-volume simple queries |
Production runs Claude for user-facing legal Q&A. Llama 3.1 handles internal batch processing (bulk case summarization, catchword extraction). This split cuts our monthly API bill by roughly 60% versus running Claude everywhere.
"Reduce hallucinations" is what every AI vendor claims. We measured ours. Here's what worked:
Most hallucinations trace back to bad retrieval. The model got weak context, so it filled the gap. Reranking with Cohere improved top-3 retrieval precision from ~62% to ~91% on our eval set. This single change removed the majority of hallucination sources.
The system prompt is 800+ words. Key rules:
Below a certain retrieval-similarity threshold, the system doesn't attempt an answer. It surfaces the top matches it found and lets the user decide. In legal, "I don't know" is a valid and expected answer. Pretending otherwise is dangerous.
500 curated questions run through the system every night. Any regression on accuracy or citation quality alerts our on-call. This is standard software practice applied to AI. A habit far too few teams building AI products actually maintain.
On our 500-question eval, hallucination rate dropped from an initial 14% (unstructured RAG) to under 1% (production system). The remaining ~1% is almost entirely "refusal when it should have answered". The safe failure mode.
Initial retrieval was around 1.8 seconds. Too slow for a conversational UI. Four optimizations got us to 480ms average:
ap-south-1 (Mumbai) since our users are in India. Network latency halved.Full user-facing latency (including LLM generation) is 2-4 seconds for a cited answer. We stream tokens so users see the answer forming immediately.
VakeelSaathi supports legal drafting in Hindi, Marathi, Tamil, Bengali, Telugu, Kannada, Malayalam, and Gujarati. English is the primary language for legal research (since Indian case law is in English); regional languages are used for drafting notices, agreements, and client-facing documents.
The setup:
For comparison: hiring a single mid-level lawyer to do the same volume of research work costs ~₹1.2 lakh/month fully loaded, and they can process a small fraction of the queries per day.
Two things, in hindsight:
Every RAG project we've shipped since VakeelSaathi uses the same fundamentals: chunking that preserves citation ability, hybrid retrieval + reranking, strict grounding prompts, evals from day one. The domain changes (contracts, product docs, medical protocols, insurance policies) but the pattern holds.
If you're evaluating a RAG chatbot for your business. Internal knowledge, customer support, document Q&A, or industry-specific research. Ask Deci or use our contact form. We'll show you what a working prototype on your data would look like.
10,000 SMS/second, 99.9% uptime. The messaging infrastructure playbook.
View case studies CS-03 Managed Ops · E-commerceMTTR from 4.2h to 22min for a D2C brand. The alerting redesign + AI anomaly detection.
Read case study CS-04 Compliance PlatformDPDP + GDPR + ISO 27001 + SOC 2 on one modular platform.
View case studiesReach us via the contact form. Tell us what documents you want your bot to answer from. We'll come back with the stack, timeline, and a fixed prototype price.