Receipts
Receipts: what Pundora has measured, and what it missed
Legal AI is sold on adjectives. This page is numbers: what was measured, on what date, how, and what the measurement does not show. The misses are here too, because a record with no misses is not a record.
The judgments
| What | Result | Measured |
|---|---|---|
| Judgments indexed | 18,433,280 | 3 October 2026, the nightly count |
| Held with full text | 17,962,823 | 3 October 2026, the nightly count |
| Parties' names filled | 99.6% of 18,037,122 judgments read | 20 September 2026, one pass over the whole corpus (34.8 hours) |
| Case number filled | 99.1% | the same pass |
| Citations found inside judgments | 13,365,593 | the same pass |
| Text against the courts' own files | 20 of 20 sampled judgments word for word | 15 September 2026, a health test against the source archive |
| Newly published judgments reaching the index | 5,000 of a 5,000-judgment sample of recent loads | 3 October 2026, measured nightly |
What this does not show: tribunals and district-court orders are not held at all; the full text was extracted from the courts' own files by Pundora's own process, and a sample of twenty is a check, not a proof, of the other eighteen million.
Finding a judgment
| Search | Time | Measured |
|---|---|---|
| By case number, across 18 million | 0.6 milliseconds | 20 September 2026 |
| By a party's name | 0.2 seconds (it had timed out at 30 seconds before its index was built) | 21 September 2026 |
| By an exact phrase in quotation marks | 1.3 to 3.6 seconds in the index; 6 to 8 seconds end to end | 29 September 2026 |
| A composed answer from the judgments | 21 to 25 seconds | 21 September 2026 |
The citation check: the founder's thirteen-trap test
On 21 September 2026 the founder ran a criminal quashing petition through the check with thirteen planted traps and scored it himself.
| Outcome | Count | What followed |
|---|---|---|
| Traps caught | 8 of 13 | — |
| False alarms on real authorities | 2 | Treated as the fault to fix first. A leading Supreme Court judgment that Pundora held was marked unconfirmed because the name search returned hundreds of High Court rows first. The search now runs in the Supreme Court first for a Supreme Court reporter, and a tie is said, not swallowed. |
| Traps missed | 3 | A fabricated quotation went unchecked because another source had confirmed the citation while Pundora held the judgment; Pundora's own copy is now found first. A sub-section that does not exist is now marked. One kind is still not caught: a real judgment cited for the opposite of what it holds. |
What this does not show: one petition, one tester, traps chosen by the person who built the product. A hand-checked test set of two hundred judgments is planned and not yet built; until it exists Pundora publishes no accuracy percentage for research or for the check.
Reading scanned pages
| Page | Words read correctly | Time a page |
|---|---|---|
| An ordinary photocopy | 98.9% | about 5 seconds, on the machine that serves it |
| A badly degraded page | 96.5% |
Measured 14 September 2026 on test pages. The finding that shaped the product: the reader reported 0.99 certainty on a page where it had 4% of the words wrong. So its certainty is never treated as proof, a page read from a picture is marked as such for ever, and the document says how many of its pages were pictures. Hindi and other Indian-script pages have not been measured.
Reading a case file, drafting, reviewing
- Dates from a case file: the standing test on a hand-marked file requires at least 80% of the marked dates to be found and none invented; it passes. It is one file; the advocate still confirms every date before it reaches the calendar.
- A first draft: 32 to 40 seconds on invented facts (19 September 2026); a planted instruction inside the source document was ignored and reported.
- A review of a draft: found all three faults planted in a test draft (19 September 2026).
- Documents that try to instruct the AI: eight such documents are kept as a permanent test; all eight are flagged.
- A research answer: about 20,000 units of AI reading and about six US cents each (21 September 2026). The first live test found a genuine quotation presented as what the court "explained" when the court was setting out an argument in order to reject it; the passages either side are now read, and a permanent test fails if it recurs.
Running the service
- Speed under load: tested on the live database at 100 chambers, 10,000 matters and 250,000 events (18 September 2026); a matter screen needs one request, answered in 11 to 27 milliseconds. The test data was deleted back to the exact starting count.
- Restore from backup: rehearsed and timed; the first drill restored and verified fourteen tables in 6.4 seconds (30 August 2026, on a small early database).
- One chamber cannot see another's: tested on every change before it can be deployed.
- Automated tests: 506 pass on the day this page was written.
What went wrong
- 16 September 2026: nine minutes down. A change to a large table held a lock for 53 minutes and the app answered errors for nine. A change may now hold such a lock for an instant or not at all; three guards enforce it, one of them proven by a live trial.
- 7 September 2026: every matter page failed. A deploy a few days earlier had broken every matter page and no test had opened one. Every screen is now read by a test, signed in, before a release is certified.
- September 2026: a nightly job that had never run. For four weeks newly loaded judgments had no titles or party names, because a schedule read as armed while it was parked for a later date. It was run by hand, the rule corrected, and a nightly measurement added (the last row of the first table).
- A check that skipped itself read as passing. Twice. Every such check is now proven to run when it is written.
- Evaluations once used the founder's own chamber and spent its day's AI allowance. Live tests now use a separate test chamber.
What has not been measured
- An accuracy percentage for research answers or for the citation check on a hand-checked set.
- Hindi and regional-language documents, and real photographs of real court orders at volume.
- Anything by an outside auditor: no penetration test, no ISO or SOC audit.
- How Pundora behaves with many real chambers; it is in early access with very few.
Questions advocates ask
How accurate is Pundora?
Pundora publishes no accuracy percentage for research or for the citation check, because the hand-checked test set that would justify one has not been built. What has been measured, with dates and methods, is on this page, with the misses.
Does Pundora hallucinate?
Any AI reading can be wrong. Pundora's design keeps the AI away from the things that must not be invented: a citation is read off a judgment it holds, a quotation is found in the text by code or not shown, and a date reaches the calendar only when the advocate confirms it. In the founder's own test the check missed three of thirteen planted traps and raised two false alarms, which are described above with what was changed.
How many judgments does Pundora hold?
18,433,280 indexed and 17,962,823 with full text on 3 October 2026: every Indian High Court and the Supreme Court, 1950 to today. Tribunals and district-court orders are not held.
How well does Pundora read scanned documents?
On test pages, 98.9% of the words of an ordinary photocopy and 96.5% of a badly degraded page, at about five seconds a page. Because the reader can be confidently wrong, a page read from a picture is always marked. Indian-script pages have not been measured.