Indic & Citizen Services
Devanagari OCR on Scanned Government Records
· 9 minute read
Devanagari OCR on a district record is not a demo on a printed book. Stamps, bleed-through, handwriting and a clerk's last photocopy will invent a survey number if you let the agent trust the Unicode.
The record room sent up a bundle that had been copied twice, stamped four times, and bound so the inner margin ate a digit. The vendor's OCR, trained on clean books, returned a survey number that did not exist in the village. The agent cited that number with confidence. The patwari, looking at the paper, saw a different digit under the staple rust.
Devanagari is not uniquely impossible. It is unforgiving on the documents government actually keeps: conjuncts that break when the photocopy smears, numerals that look like letters, rubber stamps over the operative line, and a century of mixed typewriters.
This guide is how to put OCR under an agent without letting the Unicode become the record. It is not a claim that any engine, including one we ship, 'solves Hindi OCR'.
The page is the record. The Unicode is an index.
A court, an audit party and a citizen will ask for the page. They will not ask for your confidence score. Design the store so every extracted field points at a crop. When the agent proposes a fact from a scan, the officer sees the crop beside the proposal.
If you only keep the OCR text, you have thrown away the evidence and kept the guess. That is the opposite of digitisation.
What breaks Devanagari on your shelf
Second-generation photocopies. Blue and red stamps across the body. Punch holes through a matra. Bleed-through from the reverse. Carbon copies. Typewritten Hindi with a worn ribbon. Mixed Devanagari and English on one line. Handwritten amounts in the margin. Bound volumes that cannot lie flat on a feeder.
Numerals deserve their own row. A ३ that becomes a २ after a stain is not a language problem. It is a rights problem if it is an amount or a survey number. Score numerals separately from running text.
Do not average those conditions into one percentage. Report by generation of copy, by printed versus handwritten, and by field type.
| Pack | What you pull | What you score |
|---|---|---|
| Clean print | First-generation laser Hindi circulars | This is the vendor's favourite. Do not let it be the only pack. |
| Stamped body | Pages where a seal sits on a name or number | Field under the stamp versus the rest of the line |
| Bound gutter | Inner-margin digits from registers | Whether a digit was invented to complete a token |
| Margin hand | Orders written in the side or foot | Recall of the handwritten operative text |
| Mixed line | English file numbers inside Devanagari notes | Identifiers left intact |
How the agent must behave on a low-confidence page
If the engine is unsure about a field that can change a right, the agent stops and shows the crop. It does not guess a survey number. It does not silently replace a name with a common one from the voter list.
Retrieval should be able to use both the OCR index and, where you have it, an image embedding — but the citation the officer sees is the page. Do not let a vector neighbour from a cleaner copy overwrite a worse original without saying so.
Human keying is not a failure of the programme. For registers that decide land, money or liberty, budget a two-pass: machine index plus targeted human read of operative fields.
- Never auto-correct a proper name against a modern spelling list without a flag.
- Never drop the image after 'successful' OCR.
- Never run a hosted enhance-and-forget step on a page that names a person unless the hop is written.
- Keep a hash of the image and of the OCR output so a later dispute can reconstruct what the agent saw.
This is not only a Hindi problem
The same discipline applies to other Indic scripts on poor scans. If your record room is Odia, Telugu or Modi-era Marathi, do not borrow a Devanagari book-scan number and relabel it. Build the packs from your shelf.
Where older records use a script the current staff cannot read, OCR will not create competence. Budget a human who can. A model that transliterates a script nobody in the office can check is a new way to be confidently wrong.
Objections you will hear — and what to do with them
These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.
We will only OCR clean [circulars](/blog/translating-circulars-without-losing-legal-meaning), not the record room.
Then write that limit. Do not let the same pipeline silently ingest the record room later because someone found a scanner.
A larger vision-language model will read the page end to end.
It might draft a useful description. It still needs the crop, the human stop on operative fields, and a location story. Do not confuse a demo caption with a certified extract.
Human two-pass doubles the cost.
Human two-pass on operative fields is cheaper than a land dispute. Cost the fields that matter. Let the machine index the rest.
Our vendor already quoted 98 percent.
Ask: on which pack, at what field type, with what human adjudication. A book-scan number is not your bundle.
Do not certify a number
An acceptance certificate that recites a single OCR percentage is a gift to the next dispute. Recite the packs, the field list, and the stop rule instead. If a vendor needs a number for their slide, they can print their own. It does not belong on your certificate.
The same discipline applies when a vision-language model captions the page. A caption is a draft. The crop is the extract. Do not let a fluent caption replace a numbered field.
Three weeks in the record room
Borrow a scanner and a sceptical patwari. Leave the brochure in the car.
- Week 1: photograph the five worst conditions on the shelf. Those are your packs.
- Week 1 end: pick twenty operative fields — names, amounts, survey numbers, dates, orders.
- Week 2: run the proposed engine on those pages. Do not let the vendor pick the pages.
- Week 2 end: sit an officer with the crops. Label invented digits and ignored hands.
- Week 3: write the agent rule — stop and show crop on those fields.
- Week 3: write the store rule — image and OCR hashed, image never deleted after 'success'.
- Week 3 end: decide which registers stay human-keyed. Put that in the go-live note.
How this shows up in the file
The note should say: OCR is an index. The scan is the evidence. Operative fields from poor Devanagari pages require a crop view and, where listed, a human pass. We have not accepted a single accuracy percentage as a substitute for the attached packs.
Attach the packs' hashes, the field list, and the stop rule.
This article is informational field guidance for Indian public institutions, not legal, procurement, security-accreditation, linguistics or engineering advice. Confirm against the current Gazette, Official Languages Act and Rules, state official-language law, GIGW, RPwD Act, DPDP text and Rules, CERT-In directions, departmental manual and your counsel before you file it.
How to test this with real speech, not staff English
“Devanagari OCR on Scanned Government Records” fails in the field if you only tested officers. A P4 Programme/Implementation should hear a first-generation student, a rural caller, or a Hinglish grievance before claiming “Devanagari OCR accuracy”.
Devanagari OCR on a district record is not a demo on a printed book. Stamps, bleed-through, handwriting and a clerk's last photocopy will invent a survey number if you let the agent trust the Unicode. Twenty-two scheduled languages is a Constitution fact, not a model fact. Script support is not language support. Official language rules may require bilingual output even when the model prefers one script.
- Name the languages and scripts in the eval set.
- Include code-mix and scheme-name tests.
- Measure comprehension, not BLEU alone.
- Design a human fallback when language fails.
Close this loop before the next CAB
Put “Devanagari OCR on Scanned Government Records” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P4 Programme/Implementation, not “the vendor.”
Revisit the item when the model, the GeM term, the region, or the SI changes. “Devanagari OCR accuracy” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.
Questions this usually raises
- Can we quote a single Devanagari OCR accuracy figure in the RFP?
- Not as a national fact. Accuracy on your scans depends on DPI, generation of photocopy, paper, stamp ink, handwriting and the font of the original. Demand a test on your volumes. Do not copy a vendor's book-scan number.
- Should the agent retrieve from OCR text or from the image?
- Keep the image as the evidence. Use OCR as an index. When a number or a name will change a right, show the officer the crop, not only the Unicode.
- Is a handwritten margin in scope?
- If officers wrote the real decision in the margin, yes. Say so. Printed-body OCR with ignored margins is how you lose the stay order.
- Does OCR output become personal data?
- If the page can identify a person, the image, the OCR text and the embedding are all in scope. Location, retention and access follow the record, not the file format.
- Can we clean scans in a hosted enhancement API?
- Only with a written hop. A 'denoise' service that leaves India is still a processor of the page, stamps and all.
Sources
- Official Languages Act, 1963 — Department of Official Language
- Constitution of India — Eighth Schedule (languages)
- Digital Personal Data Protection Act, 2023
- CERT-In Directions dated 28 April 2022 (180-day ICT logs)
- MeitY — India AI Governance Guidelines (5 November 2025, PIB PDF)
- Prcept AI — on-prem / air-gapped agents