Indic & Citizen Services
Building an Indic Eval Set for Your Department
· 9 minute read
Your Indic eval set is last year's inbox, stripped and frozen. If a vendor has not sat that set on the runtime you will buy, they have not passed language.
A CIO we work with had three vendor scorecards and no departmental set. Each scorecard was a different exam. Each vendor was first on their own exam. The file had no way to fail anyone.
You do not need a language institute to end that. You need a fortnight, two honest officers, and the legal nerve to keep a slice the market will not see.
This is the BOFU version of the language conversation: how to build the set that will sit in your annexure, and how to refuse any agent — including Prcept — that will not take it.
What the set is for
It is not a research corpus. It is a pass mark. It exists so a committee can fail a beautiful demo. It exists so a go-live can be refused when the delivered hash regresses. It exists so a later officer can re-run the same items after a model change.
If you cannot imagine using it to fail someone, it is a brochure. Start again.
Collect without stealing a life
Pull tickets, emails, WhatsApp exports the department already holds, call wrap-ups, and public FAQs. Strip names, numbers, addresses and rare combinations. Check a sample with a privacy officer. Public circulars can be used more freely; still watch for annexures that list people.
Keep the distribution ugly. If 40 percent of your inbox is mix, 40 percent of the set is mix. If 15 percent is a photo of a bill, the set needs image items or you must write photos out of scope.
Tag language, script, mix, channel, scheme and difficulty. Difficulty is 'would a new clerk get this wrong'. It is not 'would a linguist find this interesting'.
- No vendor sees the held-out gold.
- A public sample is allowed and useful.
- Sealed internal slice for items you could not fully de-identify.
- Hash at freeze. Re-hash if anyone edits.
Gold and rubric
For retrieval items, gold is a document id or a short list. For classification, gold is a scheme code. For generation, gold is a rubric: facts, names, numbers, no invented scheme, register, language of reply. Two officers mark. A third breaks ties. Items they will not agree on leave the scored set.
Do not gold a single 'perfect sentence'. You will punish harmless paraphrase and miss a rights error that is beautifully written.
| Dimension | Pass | Hard fail |
|---|---|---|
| Facts | All operative facts from the ticket used | Any invented next step or scheme |
| Identifiers | Names and numbers preserved | Any silent change |
| Language | Reply in the citizen's language/script unless asked | Unasked English-only reply |
| Hop | Declared if present | Undeclared pivot or egress |
| Tone | Matches the department's public voice | Blame, slang, or a promise the office cannot keep |
How it enters the buy
Publish the sample and the rubric with the bid. State that the held-out pack will be run on the proposed runtime, in your room or on your jump host, without a larger substitute model. State that go-live requires a re-run on the delivered hash. State the floor. Invented schemes are a hard fail even if the average looks fine.
The Eighth Schedule is not your item count. Year-one languages with packs are. Roadmap languages get a date and a promise to build a pack, not a tick.
Objections you will hear — and what to do with them
These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.
We will use BHASHINI or a public set and save the fortnight.
Use them as smoke. They will not contain your nicknames or your noting style. The fortnight is the work.
Two officers cannot be spared.
Then you cannot spare the year you will spend arguing with a bad bot. Marking two hundred items is cheaper than a public failure.
A held-out pack is unfair to smaller Indian vendors.
A public sample plus a clear rubric is fairness. A vendor-owned exam is the opposite. Smaller vendors who know the domain often do better on a real pack than on an academic one.
The set will go stale.
Yes. Put a refresh on the calendar. A stale set is still better than a vendor slide, until the scheme changes — then refresh before the next go-live.
The pack is an exam paper
Treat the held-out pack the way a board treats a question paper. Named access. No copies on personal drives. A hash when it is frozen. A new hash if anyone edits. If that sounds heavy, remember what the pack decides: who is allowed near citizen language.
A leaked pack is not a scandal you can ignore. It is a reason to rebuild before the next scored run. Write the rebuild trigger now, while nobody has leaked it yet.
Fourteen days to a pack in the annexure
Block two officers' mornings. Protect the slice like an exam paper.
- Day 1: export the inbox. Decide year-one languages. Write photos in or out.
- Day 2–3: de-identify. Privacy officer samples twenty items.
- Day 4: tag language, script, mix, scheme.
- Day 5: draw the public sample and the held-out slice. Hash both.
- Day 6–8: two officers mark gold and write the rubric disagreements into the parking lot.
- Day 9: freeze. Remove anyone who does not need access.
- Day 10: write the bid language — runtime, no substitute model, re-run at delivery.
- Day 11: brief the evaluation committee with one planted fail item.
- Day 12–13: dry-run scoring on an old checkpoint so the scripts work.
- Day 14: file the pack note. Bar the builders from the implementation bid if they are external.
How this shows up in the file
The note should say: the department owns a held-out Indic eval for the named languages and mixes. Vendor benchmarks are not the pass mark. Delivery is accepted only after the delivered hash sits the pack. Builders of the pack, if external, are ineligible to implement.
Attach the rubric, the sample, the hash, and the access list.
This article is informational field guidance for Indian public institutions, not legal, procurement, security-accreditation, linguistics or engineering advice. Confirm against the current Gazette, Official Languages Act and Rules, state official-language law, GIGW, RPwD Act, DPDP text and Rules, CERT-In directions, departmental manual and your counsel before you file it.
How to test this with real speech, not staff English
“Building an Indic Eval Set for Your Department” fails in the field if you only tested officers. A P1 CIO/CTO should hear a first-generation student, a rural caller, or a Hinglish grievance before claiming “custom evaluation set Indic”.
Your Indic eval set is last year's inbox, stripped and frozen. If a vendor has not sat that set on the runtime you will buy, they have not passed language. Twenty-two scheduled languages is a Constitution fact, not a model fact. Script support is not language support. Official language rules may require bilingual output even when the model prefers one script.
- Name the languages and scripts in the eval set.
- Include code-mix and scheme-name tests.
- Measure comprehension, not BLEU alone.
- Design a human fallback when language fails.
Close this loop before the next CAB
Put “Building an Indic Eval Set for Your Department” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P1 CIO/CTO, not “the vendor.”
Revisit the item when the model, the GeM term, the region, or the SI changes. “custom evaluation set Indic” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.
What must be true before you file this
If “Building an Indic Eval Set for Your Department” is only a heading, it will not survive a file inspection. A P1 CIO/CTO should be able to attach one artefact that proves “custom evaluation set Indic”: a log export, a clause, a scored row, a dated notice, or a refusal rule.
Write three dated sentences: what was decided, who owns it, and when it will be re-checked. Unsigned sentences are souvenirs. Dated sentences are controls.
- Name the owner of “custom evaluation set Indic” inside the institution.
- Attach one artefact a stranger can open next year.
- Revisit when the model, the notice, or the SI changes.
- Do not treat a vendor slide as evidence.
Questions this usually raises
- How many items do we need?
- Enough that a lucky demo cannot pass by chance, and few enough that two officers can mark them. For a single-language helpdesk, start near two hundred de-identified tickets plus a smaller gold slice for generation. For a second language, do not steal items from the first. Build a real pack or keep the language on the roadmap.
- Who owns the held-out pack?
- The department. Not the SI. Not the model vendor. Not a consultant who will later bid. Access is named. Copies are hashed.
- Can we put live personal data in the eval?
- No. De-identify. Replace rare facts that still identify. If you cannot de-identify a class of ticket, do not put that class in a pack that vendors will see. Keep a sealed internal slice for your own re-run.
- Should we include English?
- Yes, as a row, so it cannot be the only row that saves the score. English is part of Union work. It is not a substitute for the regional pack.
- How often do we refresh?
- When the scheme changes, when a new channel opens, and at least once a year. A 2022 FAQ pack will not catch a 2026 nickname.
Sources
- Official Languages Act, 1963 — Department of Official Language
- Constitution of India — Eighth Schedule (languages)
- Digital Personal Data Protection Act, 2023
- BHASHINI — National Language Translation Mission (MeitY)
- MeitY — India AI Governance Guidelines (5 November 2025, PIB PDF)
- Department of Administrative Reforms — Central Secretariat Manual of Office Procedure
- Prcept AI — on-prem / air-gapped agents