Indic & Citizen Services
Evaluating Indic Accuracy in Agents
· 10 minute read
Indic accuracy is not a model-card percentage. It is whether the agent, on your held-out tickets, in the scripts and mixes citizens actually use, still retrieves the right file and does not invent a scheme.
A procurement committee in a state mission sat through three Indic demos in one afternoon. Each vendor opened with a Hindi greeting and a dashboard that listed twenty-two languages. Each vendor then answered an English question about a scheme that does not exist in that state. The committee scored language as passed. The first production week, a Hinglish WhatsApp forward about a failed DBT credit was summarised as a request for a new ration card.
That is not an accuracy problem you can buy your way out of with a larger checkpoint. It is an evaluation design problem. You scored chrome. You did not score the job.
This is a field guide for people who must mark a bid or a go-live. It is written for Indian ministries, states, PSUs and campuses that are buying or building agents. It is not a linguistics paper. It is not legal advice. It is the test we wish every RFP already contained.
Prcept AI builds sovereign, on-prem and air-gapped agents. We should sit the same test. If we wave at a public Indic leaderboard instead of your set, fail the row.
Accuracy is a workflow, not a model-card cell
A language model can be fluent and still be wrong in the only way that matters to a file. It can retrieve the wrong circular. It can drop a date. It can translate a defined term into a near-synonym that changes a right. It can answer in polished Hindi after silently sending the prompt to an English host. Fluency is not the duty. The duty is a reconstructable, purpose-limited action on a citizen or officer request.
Evaluate the agent as it will run: the same model hash, the same retrieval corpus, the same tools, the same network path. A hosted larger model that passes Odia on demo day is a different architecture from the box you will rack in the SDC. Write that as two bids, not as one language tick.
The Eighth Schedule currently lists twenty-two scheduled languages. English is used throughout the Union's work. States notify official languages under their own laws. Citizens write Hinglish, Tanglish, Odia in Latin letters, and Urdu in Devanagari. An eval that only marks formal Hindi and formal English is a partial test of a partial population.
What to score — five rows that survive a review
Do not start with a single percentage. Start with rows an under-secretary can defend. Task success: did the agent retrieve the right artefact and propose the right next step. Fidelity: names, identifiers, amounts, dates and scheme titles survive. Language of work: the reply is in the language and script the citizen used, unless the citizen asked otherwise. Mix handling: a mid-sentence switch is not treated as noise. Architecture honesty: any translation hop is declared, located and scored as a hop.
| Row | What you mark | Fail if |
|---|---|---|
| Task success | Right file, right next step, on held-out tickets | Wrong scheme, wrong office, or invented circular |
| Fidelity | Names, numbers, dates, scheme titles exact | Any silent rewrite of a defined term or amount |
| Language and script | Reply matches citizen language and script | English-only reply to a non-English ticket without ask |
| Code-mix | Separate held-out pack of real mixed strings | Language-id drop or summary that ignores half the SMS |
| Hop honesty | Declared path of any MT or hosted ASR | Undeclared egress or undeclared English pivot |
Build the set before you book the demo
Pull last year's tickets, call transcripts, scanned notes and WhatsApp forwards that already sit in the section. Strip identifiers. Split by language, script and mix the way a clerk would, not the way a paper would. Keep a held-out slice the vendor does not see. Publish a small public sample so bidders know the register.
Size is a function of risk, not of prestige. A scholarship helpdesk can start with two hundred de-identified strings if the rubric is tight. A legal-translation workflow needs fewer items and a harder human mark. Do not pad the set with news headlines so a vendor can look good.
Who marks: two officers who disagree in public, against a written rubric. Facts preserved. Names preserved. Numbers preserved. No invented scheme. Polite address matching the department's public voice. If they cannot agree on a gold reply, the item is not ready to score a bid.
- Vendor never sees the held-out items before the scored run.
- Speech, if in scope, has its own set. Do not infer ASR from text scores.
- OCR, if in scope, has its own set of your scans, not clean Unicode.
- A translation hop, if used, is timed, located and logged.
- Re-run on the delivered hash, not only on demo day.
What not to invent on the file
Do not write a sentence that says the selected model is 94 percent accurate on Indic. You do not have a national study that supports a single number across languages, dialects, channels and domains. If a vendor offers one, ask for the set, the split, the contamination check, and whether government text was in it.
Do not treat BHASHINI, or any National Language Translation Mission component, as a substitute for your eval. BHASHINI is a public mission and a set of services. It is not a certificate that your agent is fit for a particular circular. Use public tools where policy allows. Still mark your tickets.
Do not confuse official-language duty with product coverage. Section 3 of the Official Languages Act requires both Hindi and English for specified Union instruments — resolutions, general orders, rules, notifications, certain reports, contracts, licences, permits, notices and tender forms. That is a duty on the Union's own instruments. It is not a vendor claim that an agent 'supports 22 languages'.
Objections you will hear — and what to do with them
These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.
We do not have linguists on strength.
You need last year's tickets and two officers who will argue about a rubric. A linguist is useful later, for dialect coverage and gold labels. Absence of a linguist is not a reason to score a vendor slide.
The minister wants all 22 languages on day one.
Show two columns: year-one languages with tests, and a dated roadmap with a new eval each time you add a language. A fake full list returns as a newspaper story when the first district fails.
Automatic metrics are objective. Officer rubrics are politics.
Officer rubrics are how government already decides whether a draft noting is fit to issue. Automatic metrics are a ranking aid. Use both. Do not hide a rights-affecting error behind a BLEU gain.
The vendor will game any set we publish.
Publish a sample. Hold out the rest. Change a slice every quarter. Gaming a public sample is expected. Gaming a held-out pack you re-run on the delivered hash is a contractual event.
What a pass is not
A pass is not a Hindi greeting. It is not a dashboard that lists twenty-two tiles. It is not an officer saying the demo felt good. A pass is a held-out ticket, a written rubric, and a runtime that is the runtime you will buy. If any of those three is missing, you have a conversation, not a mark.
Write that definition into the evaluation committee brief. People forget it the moment a salesperson speaks Odia. The brief is how you remember.
Twelve days to an Indic accuracy annexure you can mark
You do not need a university partnership to start. You need a clerk, a scanner and a decision about year-one languages.
- Day 1–2: pull 300 real strings. Strip identifiers. Tag language, script and mix as a clerk would.
- Day 3: write the year-one mandatory list and the roadmap. Delete any language with fewer than twenty real strings from mandatory.
- Day 4–5: write the rubric. Facts, names, numbers, no invented schemes, no silent hop.
- Day 6: freeze a held-out pack. Hash it. Put a public sample in the bid.
- Day 7: decide who marks and how disagreements are resolved. Two officers, written.
- Day 8: add speech and OCR rows only if those channels are in year-one scope.
- Day 9: write the fail. Numeric floor or hard fail on invented schemes and undeclared hops.
- Day 10: paste the table into the RFP or the go-live checklist.
- Day 11: brief the evaluation committee that English-only demos score zero on Indic rows.
- Day 12: schedule the re-run on the delivered hash before the security acceptance certificate.
How this shows up in the file
The note should say: we are not claiming Eighth Schedule coverage. We are buying year-one performance of the agent — not a disconnected model — on the attached sets for the named languages, scripts and mixes. A vendor tick against '22 official languages' is non-responsive if our annexure asked for named rows.
Attach the public sample, the hash of the held-out pack, the rubric, and the names of the marking officers. A later audit party should be able to re-run the test.
This article is informational field guidance for Indian public institutions, not legal, procurement, security-accreditation, linguistics or engineering advice. Confirm against the current Gazette, Official Languages Act and Rules, state official-language law, GIGW, RPwD Act, DPDP text and Rules, CERT-In directions, departmental manual and your counsel before you file it.
Questions this usually raises
- Can we accept a vendor's published Indic benchmark as the evaluation?
- No. Treat it as marketing. Your pass mark is a held-out set built from your tickets, circulars and call transcripts. A published score can be supplementary evidence. It cannot be the only scored row.
- Does the Eighth Schedule require us to evaluate all 22 scheduled languages?
- No. The Eighth Schedule is a constitutional list for specified purposes. It is not a software test harness. The Official Languages Act makes Union instruments bilingual in Hindi and English. Your year-one eval list is a service-design and official-language decision, not a schedule-completion exercise.
- Should we score the model or the agent?
- Score the agent on the workflow. Retrieval, tool choice, language of the reply, names and numbers, and whether a silent English hop occurred. A model that can chat about cricket in Hindi and fail on a ration-card sentence is not a public-service model.
- Is BLEU or COMET enough for citizen replies?
- Not as the only number. Automatic metrics help you rank drafts. They do not tell you whether a widow in a district understood the next step, or whether a defined term survived. Pair them with an officer rubric and, where the risk is high, a citizen teach-back.
- What if we have no evaluation set and the GeM bid is due?
- Twenty anonymised real tickets are better than a vendor slide. Delay mandatory language rows or publish the set with the bid. A due date is not a licence to write a false specification.
Sources
- Official Languages Act, 1963 — Department of Official Language
- Constitution of India — Eighth Schedule (languages)
- Digital Personal Data Protection Act, 2023
- BHASHINI — National Language Translation Mission (MeitY)
- MeitY — India AI Governance Guidelines (5 November 2025, PIB PDF)
- Prcept AI — on-prem / air-gapped agents
- Department of Administrative Reforms — Central Secretariat Manual of Office Procedure