Skip to content

Research the historical archive

ArchiveClient queries public Supreme Court and High Court judgment datasets using DuckDB. Install bharat-judgements[archive]; no AWS account or credentials are needed for the public buckets.

Search metadata

import asyncio
from bharat_judgements.archive import ArchiveClient


async def main():
    async with ArchiveClient() as client:
        records = await client.search(court="delhi", year=2020, party="tata", limit=10)
        for record in records:
            print(record.cnr, record.title, record.decision_date)
        print(client.cache_info())


asyncio.run(main())

Specify a court and year where possible to reduce the metadata scanned. A year can be a single integer or inclusive pair such as (2018, 2024). search returns a date-ordered Judgment list; missing dates sort last.

Match filters to the dataset

Filter Supreme Court High Courts
court, year Select source/year Select court/year
judge Judge substring Judge substring
party Petitioner, respondent or title substring Title substring
citation Citation substring Ignored: no citation column
cnr Exact identifier match Exact identifier match

This is metadata search, not a search of the body text of judgment PDFs. There is no district court judgment archive in this client. When court is omitted, list search splits its result budget between SCI and HC sources; one source is not guaranteed to fill unused capacity from the other.

Stream a larger result set

iter_judgments yields individual records with SQL paging. batch_size defaults to 500; max_results can cap the total. It streams SCI first and then HC, with each source internally ordered by decision date and CNR. It does not globally merge the sources by date.

import asyncio
from bharat_judgements.archive import ArchiveClient


async def main():
    async with ArchiveClient() as client:
        async for record in client.iter_judgments(
            court="delhi", year=2020, batch_size=100, max_results=200
        ):
            print(record.to_json(exclude_none=True))


asyncio.run(main())

Fetch a PDF

fetch_pdf(record) accepts a Judgment; fetch_pdf(cnr) resolves an archive record first. A recognised CNR prefix can route the lookup to the correct court. Unknown identifiers, missing paths and failed downloads can raise errors.

High Court PDFs are downloaded individually. Supreme Court PDFs are extracted from a cached yearly language tar: fetching the first PDF can download a large archive. Use prefetch_sci_year(year, language="english") deliberately before a batch, accounting for network and disk usage.

SCI language support comes from the stored archive language map and available files. Selecting a language does not translate a PDF. High Court retrieval uses the stored PDF path.

Counts and cache controls

count(court=..., year=...) returns source counts rather than a judgment list. cache_info() reports PDF/tar cache information. Metadata and PDF caches live under ~/.cache/bharat-judgements/archive/.

The metadata cache defaults to 30 days of freshness. PDF/tar storage defaults to a 5 GiB cap with eviction; metadata has separate storage. Configure these through the process environment as described in configuration.

bharat-judgements --json archive query --court delhi --year 2020 --party tata --limit 10
bharat-judgements archive cache

These commands need the cli extra. archive cache --clear deletes the archive cache directory; use it only when you intend to discard downloaded files. archive download --court sci --year 2020 prefetches the SCI yearly tar; it is not a general HC bulk-download command.

Freshness and attribution

The publishers list bi-monthly SCI updates and quarterly HC updates. Dataset publication and your metadata cache both affect freshness; no fixed maximum lag is guaranteed. A missing recent judgment may require a live source query with appropriate inputs.

The datasets list CC-BY-4.0. Keep provenance and attribution when reusing archive material. See data and licensing for publisher links and source distinctions.

Next: archive API, facade routing, or live text search.