All insights

Indic & Citizen Services

Benchmarking Indic Models on Government Text

· 10 minute read

A government Indic benchmark is a method, not a podium. Build it from your circulars and scheme names, hold it out, and refuse any leaderboard that will not show the items.

Every quarter someone forwards a chart. Model A is ahead of Model B on 'Indic'. The axes are cheerful. The items are invisible. A programme team then writes the chart into a steering-committee deck as if a ministry had run a trial.

That is how a department buys a news-headline model to read a gazette. Government text is a different register: defined terms, shall and may, scheme titles that are half acronym, bilingual authentic versions, tables that were once paper, and footnotes that carry the actual commencement date.

This is a data-study in the only honest sense: a study of how to build the data. It does not publish a leaderboard. It does not invent an accuracy percentage. If a vendor cannot sit your bench, they do not have a government Indic result. They have a marketing quarter.

What government text actually is

Start with instruments you already have a right to copy: Gazette notifications, scheme guidelines, office memoranda, the Manual of Office Procedure habit of your secretariat, frequently asked clarifications, and the bilingual forms the Official Languages Act already requires the Union to issue in Hindi and English.

Then add the operational register, de-identified: grievance subjects, call-centre wrap-up codes, the way a tehsil writes a place name, the way a campus writes a circular about hostel fees. That register is where agents fail even when they can paraphrase a PIB release.

Keep legal authentic text in its own slice. A model that paraphrases a notification into friendlier Hindi has not translated the law. It has written a pamphlet. Score pamphlets and instruments separately or you will reward fluency for a task that required fidelity.

The method — six artefacts, no podium

One: a corpus card. Where each item came from, whether it is public, the language and script, the date, the office. Two: a task card. What the model is asked to do — retrieve, extract a date, answer a citizen FAQ, draft a noting, align Hindi and English versions. Three: a split. Train-like public sample, development slice, held-out slice. Four: a contamination note. How you will detect that a vendor sat the held-out items. Five: a metric stack that matches the task, not a single exam score. Six: a human protocol for the items automatic metrics cannot settle.

  1. Inventory public instruments and de-identified operational strings. Do not scrape live case files.
  2. Tag language, script, mix, office and task. A bilingual notification is two aligned items plus an alignment task.
  3. Write gold only where two officers agree. Disputed items go to a parking lot, not into the scored slice.
  4. Freeze the held-out hash. Record who can open it.
  5. Run a baseline you already own — even a dumb keyword retriever — so a neural gain is visible.
  6. Report by language and task. Never a single 'Indic %'.

What the bench must contain if it claims to be government

Scheme titles and their nicknames. Statutory phrases that look like ordinary Hindi or Tamil until you change one word. Tables and schedules. Dates in more than one calendar habit. Place names that collide across districts. Acronyms that expand differently in two languages. OCR-degraded pages, if your production corpus is scans. Code-mixed queries, if your production queries are mixed.

If any of those are missing, say so on the card. A clean-Unicode-only bench is a bench of clean Unicode. It will not predict your record room.

Minimum slices for a department bench. Drop a slice only if that work is out of scope.
SliceSourceWhat a fail looks like
Instrument fidelityGazette / OM, bilingual if UnionShall/may swap, dropped proviso, wrong commencement
Scheme namesGuidelines plus citizen nicknamesMerged with a different scheme or expanded wrongly
Operational mixDe-identified tickets and SMSLanguage-id drop, ignored half-sentence
Scan realityYour own degraded Devanagari or regional scansClean-text score used as a proxy for OCR text
RetrievalYour circular corpus with known gold docsFluent answer from the wrong year of the scheme

How to publish without lying

If you share results inside government, share the method card and the per-language table. If you share outside, share the public sample and the method, not a rank that implies a national contest. Do not let a systems integrator reprint your internal table as 'State X ranks Model A first on Indic'.

BHASHINI and other public missions may offer tools, datasets or APIs. Use them. Still write your card. A mission is not your contamination check.

Objections you will hear — and what to do with them

These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.

Without a public leaderboard nobody will take this seriously.

Auditors take methods seriously. Ministers take district failures seriously. A podium is optional. A re-runnable hash is not.

We should wait for a national Indic government benchmark.

Wait if you wish. Do not pause year-one service. Twenty of your own circulars and a hundred tickets will beat a national set that has not arrived.

Reporting by language will make one language look bad and create politics.

A blended score is the politics. It hides the district that will fail. Report the column. Fund the gap or narrow the go-live.

Vendors will refuse a bench they did not design.

That is useful information. A vendor who will only sit their own exam is telling you who writes the mark sheet.

No podium in the minutes

Steering-committee minutes love a rank. Resist it. A rank implies a closed contest on a shared exam. You do not have that exam. You have a departmental method. Minutes that say 'Model A led on Indic' will be quoted in the next bid as if the State had certified a league.

Write instead: on the attached hash, for these tasks and languages, candidate A met the floor and candidate B did not. That sentence can be re-run. A podium cannot.

Four weeks to a bench you can re-run

This is a construction schedule, not a research sabbatical.

  1. Week 1: inventory public instruments and de-identified tickets. Write the corpus card.
  2. Week 1 end: decide tasks. Retrieval and fidelity first. Open generation last.
  3. Week 2: double-mark gold on a small slice. Throw out items officers will not agree on.
  4. Week 2 end: freeze held-out hash. Publish only the sample and the method.
  5. Week 3: run a dumb baseline and one current checkpoint. Learn the scoring scripts.
  6. Week 3 end: add the scan slice and the mix slice if those are production realities.
  7. Week 4: write the reporting template — per language, per task, no blended podium.
  8. Week 4 end: put the re-run date into the contract or the go-live checklist.

How this shows up in the file

The note should say: we have not adopted a public Indic leaderboard as a selection criterion. We have built a department bench from the attached corpus card. Selection and go-live will use the held-out hash. We will not publish a ranked podium that implies a national contest.

Attach the method, the sample, the hash reference, and the names of the officers who marked gold.

This article is informational field guidance for Indian public institutions, not legal, procurement, security-accreditation, linguistics or engineering advice. Confirm against the current Gazette, Official Languages Act and Rules, state official-language law, GIGW, RPwD Act, DPDP text and Rules, CERT-In directions, departmental manual and your counsel before you file it.

How to test this with real speech, not staff English

“Benchmarking Indic Models on Government Text” fails in the field if you only tested officers. A P4 Programme/Implementation should hear a first-generation student, a rural caller, or a Hinglish grievance before claiming “Indic model benchmark”.

A government Indic benchmark is a method, not a podium. Build it from your circulars and scheme names, hold it out, and refuse any leaderboard that will not show the items. Twenty-two scheduled languages is a Constitution fact, not a model fact. Script support is not language support. Official language rules may require bilingual output even when the model prefers one script.

  • Name the languages and scripts in the eval set.
  • Include code-mix and scheme-name tests.
  • Measure comprehension, not BLEU alone.
  • Design a human fallback when language fails.

Close this loop before the next CAB

Put “Benchmarking Indic Models on Government Text” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P4 Programme/Implementation, not “the vendor.”

Revisit the item when the model, the GeM term, the region, or the SI changes. “Indic model benchmark” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.

Questions this usually raises

Will Prcept publish a ranked Indic leaderboard in this article?
No. We will not invent scores or reprint a vendor podium as if it were a national study. A rank without your items, your split and your contamination check is theatre. This piece is the method for a bench you can defend.
Can we reuse a public IndicQA or similar set?
As a smoke test, yes. As the scored government bench, no. Public sets are light on gazette register, scheme nicknames, bilingual defined terms and the OCR mess of a district file. Use them to catch a dead model. Do not use them to pick a production stack.
How do we stop the bench leaking into the next model's train set?
Hold out a private slice. Rotate items. Watermark a few strings. Do not put the gold pack on a public GitHub. Assume any fully public bench will be contaminated within a year.
Should the bench be multilingual in one table?
Report each language, script and mix as its own column. A blended 'Indic score' hides a collapse in the language you actually serve.
Is government text personal data?
A gazette notification may be public. A file noting that names a citizen is not. Build the bench from public instruments and de-identified operational text. Do not copy live case files into a shared eval repo.

Sources