How we verify
Where this text comes from
Every provision on Vidhara is traced to the government’s own published PDF. This page explains how, what we check, what we have found wrong and fixed — and, just as important, what we do not claim.
- Acts
- 36
- Sections
- 5,594
- Old⇄new mappings
- 1,271
The source is the government's PDF, not another website
Each act is taken from India Code or the Gazette of India — the official publications — and never copied from another legal site. For each one we record the exact file we used, its size in bytes and its SHA-256 hash, so the text on this page can be traced back to a specific government document rather than to “a PDF we found”.
Every section page carries that provenance at the bottom, including a link to the source document, so you can check any provision against the original yourself.
Extraction reads the page, not the copy-paste
Statute PDFs put marginal notes in one column and text in another, and print schedules as tables. Copying text out collapses those columns and silently interleaves them. We instead read the coordinates of every word on the page and rebuild the structure from its geometry, which is why a three-column schedule keeps each period beside the limb it belongs to.
What we check before publishing
Where the structure allows it, extraction is verified lossless: we compare every word of the parsed result against the source document and require the counts to match exactly. The Limitation Act’s Schedule, for instance, was published only once its 4,581 word tokens matched the PDF’s 4,581.
A scanner then runs over the whole corpus looking for known defect shapes — text that duplicates itself, provisions that ended up empty, headings that leaked into a body, numbering that jumps. It currently reports zero severity-1 defects across all 5,594 published sections.
Things we have found wrong — and fixed
This is the part that matters, because a corpus assembled in bulk will carry these and nobody will know. Each of these was a real defect in our own data, found by checking:
- One State’s amendment shown as national law. India Code prints State amendments immediately after the central section they modify, and our text had absorbed them. This was found in stages and the last of it was only cleared on 2 August 2026: 68 sections — including CrPC §438 (anticipatory bail), §125 (maintenance) and §154 (FIR) — were still carrying a State’s amending text inside the central provision, about 142,000 characters of it. A further 21 provisions were published as sections in their own right though no such section exists nationally: IPC 354E, 376F, 509A and 509B (Chhattisgarh), 379A and 379B (Gujarat), 382B–382F (Tripura), IEA 114B (Chhattisgarh), and Registration Act 80A–80G and 89C–89D (Bengal and Uttar Pradesh).
- The Constitution was missing Part II. Citizenship — Articles 5 to 11 — had no Part of its own, and those articles sat under “The Union and its Territory”.
- Schedule entries read as sections. The NDPS Act’s schedule of substances parsed as sections numbered past 110, none of which exist in the Act.
- Headings swallowed into the text. Chapter and Part headings appended to the end of the previous section’s body.
Every correction is recorded with the date, the cause and the fix, and lives in a versioned data bundle rather than being edited straight into the live database — so a republish can never quietly undo it.
State amendments are shown, and shown separately
Several States amend central Acts in their own application. Keeping that text out of the section is not enough on its own: a reader in Karnataka who sees only the central provision has been told, in effect, that nothing else applies. Silence is its own wrong answer.
So where the source records one, the amendment now appears in its own block beneath the section — labelled with the State, with the amending Act cited, and collapsed until you open it. There are currently 320 such amendments across the corpus. What is shown is the amending instruction as India Code prints it (“in section 17, after clause (b), insert…”), never a consolidated State version of the section — writing that ourselves would mean composing statute text, which we do not do.
What we do not claim
Being useful means being clear about the edges. All of the following are true today:
- 36 acts is narrow. Other apps carry hundreds. We would rather add them slowly and check each one than publish a corpus we cannot vouch for.
- Parsing is automated, with spot checks against the PDF. A full clause-by-clause human proofread of every section has not been done.
- Footnotes and amendment history are excluded. You get the provision as currently printed, not the record of how it changed.
- Most schedules are not ingested. The Limitation Act’s Schedule is; others are not yet. Where a schedule matters, read it from the source PDF.
- There is no case law here. Vidhara tells you what a provision says, not how courts have read it.
- State amendments are recorded, not consolidated. You get the amending instruction and its citation, not the section as it reads in that State — and only where the source prints one, which is not the same as everywhere one exists.
- An act is held back rather than published if we cannot vouch for it. Six acts’ PDFs print amendment footnotes in the same size as their text, which destroys sections outright — the Special Marriage Act lost the conditions for a valid marriage, and the SC/ST (Prevention of Atrocities) Act lost its central punishment provision. Each is published only because the repaired text matches that Act’s own arrangement of sections, section for section; the repair refuses to run otherwise. Nothing is currently withheld.
For anything you file or rely on professionally, check the bare act. This is a reading and reference tool, not a substitute for the official text.
If you find a mistake
Wrong text is the most serious kind of bug we can have, and it is treated that way — a report goes to the top of the queue ahead of any feature. Every section page has a report link scoped to that exact provision, or you can tell us here.
Counts on this page are read from the live database, so they stay accurate as the corpus grows. Last checked when this page was generated.