DPDP deadline: May 2027 · 219 days left Get your business ready in 15 days, done for you Learn more →
Home / Case Studies / VakeelSaathi Legal RAG CASE FILE · CS-01
CASE FILE CS-01 / VAKEELSAATHI LEGAL RAG

How we built VakeelSaathi: legal RAG on 50,000+ Indian court judgments

A production Retrieval-Augmented Generation system that answers Indian legal queries with paragraph-level source citations. In sub-500ms retrieval, across 8 Indian languages, with near-zero hallucinations.

BUILT 2024 · LIVE IN PRODUCTION 12 MIN READ SECTOR: LEGAL TECH
50,000+Judgments Indexed
<500msRetrieval Latency
8Indian Languages
~0%Hallucination Rate

The problem: legal research shouldn't take four hours

India's legal ecosystem runs on judgments. Every argument in court, every contract dispute, every advisory memo starts with a question: what did courts say about this before? Answering that question the traditional way means opening Indian Kanoon or Manupatra, running keyword searches, opening 20 tabs, reading through irrelevant hits, and finally landing on the two or three cases that matter. After 3 to 6 hours of research.

When we started building VakeelSaathi (Decipher's legal operations SaaS), the founding hypothesis was simple: a lawyer should be able to ask "what have courts held about arbitral awards being challenged under Section 34?" and get a cited answer in 30 seconds.

Building it was harder than the hypothesis suggested.

NOTE / WHY GENERIC CHATGPT WASN'T THE ANSWER

Ask ChatGPT the same question and it will confidently cite Union of India v. Some Case, (2018) 4 SCC 123. A citation it just made up. This is the hallucination problem. In legal work, a fabricated citation isn't a bug; it's malpractice. We needed a system where the model could only quote from real, verifiable judgments.

The architecture

VakeelSaathi's legal Q&A runs on a RAG pipeline. Retrieval-Augmented Generation is straightforward in theory. Fetch relevant documents, give them to an LLM as context, let the LLM answer using only what you gave it. In practice, at 50,000+ documents and Indian legal complexity, every step required careful engineering.

RAG PIPELINE · QUERY → CITED ANSWER
01Data ingestion & normalizationJudgment text pulled from Indian Kanoon, cleaned, deduplicated, and structured (case number, court, date, judges, parties, catchwords). Custom parser handles OCR errors from scanned older judgments.
02Chunking strategyJudgments split by paragraph. Not fixed token windows. Each chunk keeps its position in the judgment for later citation. Metadata: case name, citation, court, section headers.
03Embedding & indexingOpenAI text-embedding-3-large for semantic embeddings, stored in Pinecone (managed) with metadata filters for court, year, and legal domain. Sparse BM25 index alongside for keyword-heavy queries (statute numbers, section references).
04Hybrid retrievalEvery query hits both the dense (semantic) and sparse (keyword) index. Results are merged with reciprocal rank fusion. Top 30 chunks proceed to reranking.
05RerankingCohere Rerank v3 scores the 30 candidates against the query. Top 5-8 chunks (based on score threshold) proceed to the LLM.
06Generation with strict groundingClaude 3.5 Sonnet generates the answer with a system prompt that mandates paragraph citations and forbids un-cited claims. If no chunk answers the query, the model returns "I don't have that in the case law I've indexed."
07Citation resolutionEvery citation in the answer resolves to a link back to the source judgment on Indian Kanoon, opening at the exact paragraph the model quoted.

Model selection: why Claude beat GPT-4 on Indian legal reasoning

Before locking in Claude, we ran evaluations on 500 curated Indian legal questions across GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, and Llama 3.1 70B (self-hosted). Grading criteria: factual accuracy, correct citation format, sensitivity to jurisdiction (SC vs High Court binding force), and refusal quality on out-of-scope queries.

Results:

ModelAccuracyNotes
Claude 3.5 Sonnet89%Near-perfect citation format, best refusal behavior
GPT-4o84%Occasional citation fabrications, slightly weaker on Indian statute interpretation
Gemini 1.5 Pro79%Verbose answers, unreliable citation format
Llama 3.1 70B (self-hosted)71%Cheapest to run, best for high-volume simple queries

Production runs Claude for user-facing legal Q&A. Llama 3.1 handles internal batch processing (bulk case summarization, catchword extraction). This split cuts our monthly API bill by roughly 60% versus running Claude everywhere.

How we killed hallucinations

"Reduce hallucinations" is what every AI vendor claims. We measured ours. Here's what worked:

Layer 1: Retrieval quality

Most hallucinations trace back to bad retrieval. The model got weak context, so it filled the gap. Reranking with Cohere improved top-3 retrieval precision from ~62% to ~91% on our eval set. This single change removed the majority of hallucination sources.

Layer 2: Strict grounding prompts

The system prompt is 800+ words. Key rules:

  • Every factual claim must cite a specific paragraph from the retrieved chunks
  • If no chunk supports a claim, the model must say "I don't have direct authority on this"
  • Never invent case names, citations, dates, or holdings
  • Distinguish binding precedent (Supreme Court, applicable High Court) from persuasive precedent (other High Courts, tribunals)

Layer 3: Confidence thresholds

Below a certain retrieval-similarity threshold, the system doesn't attempt an answer. It surfaces the top matches it found and lets the user decide. In legal, "I don't know" is a valid and expected answer. Pretending otherwise is dangerous.

Layer 4: Nightly evals

500 curated questions run through the system every night. Any regression on accuracy or citation quality alerts our on-call. This is standard software practice applied to AI. A habit far too few teams building AI products actually maintain.

RESULT

On our 500-question eval, hallucination rate dropped from an initial 14% (unstructured RAG) to under 1% (production system). The remaining ~1% is almost entirely "refusal when it should have answered". The safe failure mode.

Latency work: getting to sub-500ms retrieval

Initial retrieval was around 1.8 seconds. Too slow for a conversational UI. Four optimizations got us to 480ms average:

  1. Parallel dense + sparse retrieval. Instead of running them serially, we fire both simultaneously and merge when both return. Saved ~300ms.
  2. Cached embedding model. For the query embedding step, we moved from an OpenAI API call to a locally-hosted sentence-transformer for the initial embedding, calling the API only for reranking. Saved ~250ms per query and reduced cost.
  3. Pinecone region optimization. Moved index to ap-south-1 (Mumbai) since our users are in India. Network latency halved.
  4. Reranker batch size tuning. Cohere's rerank endpoint has non-linear latency with batch size. Tuning from 50 to 30 candidates cut this step 40%.

Full user-facing latency (including LLM generation) is 2-4 seconds for a cited answer. We stream tokens so users see the answer forming immediately.

Multilingual: 8 Indian languages, one system

VakeelSaathi supports legal drafting in Hindi, Marathi, Tamil, Bengali, Telugu, Kannada, Malayalam, and Gujarati. English is the primary language for legal research (since Indian case law is in English); regional languages are used for drafting notices, agreements, and client-facing documents.

The setup:

  • Retrieval always runs against the English case-law index
  • Generation uses Claude 3.5 for English drafts
  • Language-specific fine-tuned Llama 3.1 models handle output in each regional language, seeded with the English draft as context
  • Post-generation validation checks legal terminology consistency against a bilingual glossary

The results, in numbers

3-6 hoursBefore — to research a legal question thoroughly
30 secondsWith VakeelSaathi — to a cited answer + source judgments
14%Initial hallucination rate — unstructured RAG baseline
< 1%Production hallucination rate — after 4-layer defense

What this cost. And what it costs to run

  • Prototype phase (3 weeks): ~₹2.5 lakh. 5,000-judgment prototype, model comparison, first UI
  • Production build (10 weeks): ~₹15 lakh. Full 50k index, ingestion pipeline, multilingual, WhatsApp bot, evals, security review, deployment
  • Ongoing operations: ~₹80,000/month. Model API costs (Claude + Cohere + embeddings), Pinecone, monitoring, weekly re-indexing, quality evals

For comparison: hiring a single mid-level lawyer to do the same volume of research work costs ~₹1.2 lakh/month fully loaded, and they can process a small fraction of the queries per day.

What we'd do differently

Two things, in hindsight:

  1. Start with hybrid retrieval from day one. We spent 3 weeks tuning pure dense retrieval before adding BM25. Should have run both from the beginning. The improvement was immediate and cost was trivial.
  2. Build the eval harness before the product. We built evals in Week 6. Every prompt or model change before that was measured by "does it feel better?". A bad way to make decisions on AI systems. Eval-driven development is what lets you actually improve over time.

Can we build something similar for you?

Every RAG project we've shipped since VakeelSaathi uses the same fundamentals: chunking that preserves citation ability, hybrid retrieval + reranking, strict grounding prompts, evals from day one. The domain changes (contracts, product docs, medical protocols, insurance policies) but the pattern holds.

If you're evaluating a RAG chatbot for your business. Internal knowledge, customer support, document Q&A, or industry-specific research. Ask Deci or use our contact form. We'll show you what a working prototype on your data would look like.

CS-01 / FIELD MANUAL — FAQ

Questions about this build

How large was VakeelSaathi's document index?
Over 50,000 Indian court judgments (Supreme Court, High Courts, and select tribunals) at production launch, plus statutes and case commentary. The index is refreshed weekly with new judgments from Indian Kanoon and other public legal sources.
What was the retrieval latency?
Sub-500ms end-to-end for the retrieval step (query embedding + vector search + reranking). The full user-facing latency including LLM generation is typically 2-4 seconds for a complete cited answer.
How did you prevent hallucinations in legal answers?
Four layers: (1) strict grounding. The LLM can only quote from retrieved chunks, (2) confidence thresholds. Below the threshold, the system says 'I don't know' or routes to a human, (3) automated evals on 500+ curated legal questions run nightly, (4) mandatory paragraph-level citations that link back to the source judgment for user verification.
Which LLM did you use?
We evaluated GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro on our legal Q&A eval set. Claude 3.5 won on Indian legal reasoning accuracy. For multilingual drafting (Hindi, Marathi, etc.), we use a hybrid. Claude for reasoning + fine-tuned Llama for language-specific outputs.
How much did it cost to build?
The prototype phase (working RAG on 5,000 sample judgments, 3 weeks) was in the ₹2-3 lakh range. Full production build (data ingestion pipeline, 50k+ index, UI, multilingual drafting, WhatsApp bot, evals, deployment) was in the ₹12-18 lakh range. Ongoing operations including model API costs, monitoring, and weekly re-indexing runs around ₹80,000/month.
CROSS-REFERENCE / MORE CASE STUDIES

See how we work across other systems

Want RAG on your own data?

Reach us via the contact form. Tell us what documents you want your bot to answer from. We'll come back with the stack, timeline, and a fixed prototype price.