✦ Methodology
How Courts & Cases works — transparently.
You should know exactly how a legal AI tool produces its output, where its data comes from,
and what it can and cannot do. Here is our methodology, our numbers, and our limitations.
Where the corpus comes from
Courts & Cases indexes judgment records published by Indian courts through their public case
information systems. Two layers of coverage exist:
- Case-registry records — Complete lists of disposed and pending records published by the
courts' public information systems, spanning the Supreme Court and High Courts. These provide the
case identifier, court, and decided-on date for every decision published that way.
- Full-text judgments — Scanned and OCR-processed judgment documents for a large subset of
the registry, which we ingest, process into structured sections (facts, issues, reasoning, decision),
and index for retrieval.
6.4M+judgment records indexed from the Supreme Court and all 25 High Courts
890K+full-text judgments with 1,000+ word processed text available for deep research
64,564full-text Supreme Court judgments
25 / 25High Courts represented in the corpus
Honest about OCR. A meaningful part of the full-text corpus was digitised from scans via OCR,
and OCR is imperfect. Typographical errors, garbled names, or misread numbers can appear in processed text.
Where a quoted passage matters — and it always does before filing — read the judgment and, where necessary,
verify against the official court record linked on every case page.
Research methodology
When you describe a matter, Courts & Cases:
- Extracts legal issues — AI identifies the legal questions in your facts.
- Retrieves before it generates — A fast full-text pre-filter finds candidate judgments,
then a semantic ranker re-orders them by factual alignment and legal-issue overlap against the retrieved
texts. The model is only ever given what was retrieved.
- Quotes verbatim — Quoted passages come from the processed judgment text itself, with the
paragraph they belong to. They are not paraphrased by the model.
- Returns real, openable judgments — Every citation links to the full judgment text in the
corpus. A judgment that is not in the corpus cannot be cited, because nothing is cited from memory.
Retrieval-first, by enforcement. The drafting and research pipeline refuses to answer from
the model's stored knowledge about Indian case law. Authorities must come from what retrieval found; the
model's role is to read, rank, summarise and quote those retrieved texts — not to recall cases.
Drafting methodology
When you draft a document:
- Clarifying questions — The AI asks targeted follow-up questions to fill gaps in the facts.
- Research-backed generation — The draft is grounded in the judgments and statutes retrieved
for your matter; citations are embedded only where the underlying text is in the corpus.
- Statute-aware — Drafts recognise Indian statutes and flag the governing provisions,
including which procedural code (pre-2024 or the BNS/BNSS) applies to the stage of the matter.
- Your review — The draft is a starting point. You review, edit, and finalize before filing.
Judge Intelligence methodology
Judge profiles are computed directly from the judgment database, never from AI recollection:
- Descriptive, not predictive — We show patterns from a judge's own published judgments:
subject-matter mix, how comparable matters were decided, and the reasoning relied on.
- DB-computed numbers — Every count, distribution and statistic is drawn from the indexed
judgment records. The AI writes only the narrative, grounded on verbatim excerpts from the same records.
- Calculated neutrality — We deliberately do not claim to predict how a judge will rule.
Profiles are for preparation and advocacy, not prediction.
Today the corpus yields reliable profiles for 5,000+ judges; the profile surface is limited to judges with
a meaningful number of published judgments so patterns are real and not noise.
Case Law Digest verification
Every Case Law Digest note is checked against the judgment text and the case record before it is
generated and again before it is published — case name, court, decision date, bench, citation, case
number, quotations, and precedent treatment all have to trace back to the judgment. A note that fails a
check on its case name, court, or date is not published at all; a note that fails a smaller check simply
drops that one detail rather than showing something unverified. See our
editorial policy for the full list of checks.
Statute indexing
Statutes and sections are extracted from the judgment text itself, grouped into canonical act families
(IPC/BNS, CrPC/BNSS, CPC, the Constitution, and so on), and exposed through statute hubs that list the
judgments citing each provision. Citation counts you see on those pages are computed from the corpus and
precomputed nightly, off-peak, rather than on every page load. A repealed predecessor Act with a
different year (for example the Motor Vehicles Act, 1939, or the Code of Criminal Procedure, 1898) is not
counted as a citation of the current Act with the same short name.
Keeping the Central Acts registry current
Alongside statutes found in judgment text, we maintain a registry of official Central Acts. A job
discovers newly listed Acts from PRS Legislative Research's public listings and fetches the official
text PDF for each one, honouring the site's crawl delay. A discovered Act is stored only if its extracted
text passes a set of quality checks: the PDF must have a genuine text layer (not a scan with no text),
its section numbers must be unique and run forward without gaps that suggest unrelated content (such as a
Schedule listing another Act's sections), and Section 1 must be kept. Amendment, appropriation and
finance Acts are recorded in the registry but never stored as the law itself. India Code and the
Legislative Department's own sites block automated access, so Acts without a reachable official copy
elsewhere wait for manual review rather than being guessed at. This keeps the registry growing without
ever publishing text we can't stand behind — we don't claim specific new Acts are live from this page;
check the statute hub for what's actually indexed.
What we never do
- We never invent citations — Every case citation links to a real judgment in our corpus.
- We never fabricate judgments — If a case isn't in our database, we don't pretend it is.
If the platform cannot find a solid authority, it says so rather than guessing.
- We never guarantee outcomes — Legal research informs strategy; it doesn't predict results.
- We never replace your judgment — Courts & Cases is a research tool. You are the advocate.
Limitations, stated plainly
- Corpus coverage — A 6.4M-record corpus is broad but not everything ever decided. Some
older or regional decisions may be absent, and official publication channels change over time.
- OCR quality — Scanned judgments carry OCR risk. We filter the full-text search surface
to judgments with substantial processed text, and we recommend verifying any high-stakes quote against the
official record before filing.
- Full text is not universal — Many registry records are metadata-rich but may not have
processed full text for every judgment. Those records still surface in browse and court pages, but deep
verbatim quotation requires the full-text layer.
- AI summaries — AI-generated summaries are starting points. Always read the full judgment
before relying on it.
- Official source — Courts & Cases is a research interface. Verify important judgments
against the official court source before filing.