All insights

Indic & Citizen Services

Domain Glossaries: The Cheapest Accuracy Win

· 10 minute read

Before you rent a larger checkpoint, lock the fifty phrases that already decide your files. A glossary is cheaper than a model, and you can audit it.

A programme team had budgeted a larger multilingual model because 'Hindi quality is poor'. We sat with last month's failures. Forty of sixty were scheme titles, office names, form numbers and shall/may. Sixteen were mix. Four were genuine fluency. They were about to buy parameters to fix a list.

The list is the cheapest accuracy win in this cluster. It is also the one you can show an audit party without a research story.

This guide is how to build, lock, version and evaluate a domain glossary so the model is not allowed to be creative where the file is not.

What goes in — and what stays out

In: scheme titles and nicknames, form numbers, office and designation names, statute short titles, the dangerous legal phrases, place names that collide, product names of your own portals, bilingual authentic pairs you have already issued.

Out: ordinary adjectives, slogans, and phrases officers like to rewrite every week. If everything is locked, nothing is. The glossary is a spine, not a style guide for the whole language.

Fields that make a row usable

A usable row is not 'Hindi = x, English = y'. It is a small record.

Minimum fields per glossary entry.
FieldWhy it exists
Canonical idWhat retrieval and eval key off, not a display string
Official formsEach authentic language, plus script
AliasesNicknames and common misspellings you will accept
Banned formsPoetic MT output you have already seen
StatusLive, closed, merged-into
CiteDocument id the agent must retrieve
Owner / versionWho may edit, and the date of the last OM

How the agent uses it

Detect aliases on the way in. Retrieve by canonical id. Mask official forms during any translation hop. Constrain the decoder or run a post-check that fails the turn if a banned form appears. Show the officer the glossary hit when the draft is for a Class A instrument.

If the glossary and the model disagree, the glossary wins and the disagreement is logged. A model that 'knows better' is how a closed scheme returns from the dead.

Governance so the list does not rot

One owner in the programme cell. A trigger on every OM that creates, renames or closes a scheme. A quarterly pass on nicknames from the inbox. An export test: if the vendor left tomorrow, could you load the same file into another runtime.

Do not let the glossary become a personal spreadsheet on a deputy's laptop. That is how versions fork and two authentic Hindi titles appear.

Put the glossary version in the agent's response metadata, at least for officer-facing drafts. When a later correction is needed, you must know which list the model saw. A nameless list is how two sections argue about which Hindi title was official last Tuesday.

Objections you will hear — and what to do with them

These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.

This is just terminology work. It is not AI.

Correct. That is why it is cheap and why it works. AI that skips terminology work is a cost centre.

Our language is too fluid for a list.

The gazette is not fluid on the title of a scheme. Lock the gazette. Leave the rest.

Vendors will not accept a hard fail on banned forms.

Then they are asking to invent your schemes. Write the fail. Watch who stays.

We already have a Rajbhasha glossary.

Good. Import it. Add nicknames, status and cite. Official-language glossaries rarely include the WhatsApp name of a scheme.

Export or it is not yours

A glossary that lives only inside a vendor prompt will die when the vendor does. The export test is not bureaucracy. It is how you keep the cheapest accuracy win after the contract ends.

If the vendor says the terms are their IP, take the official titles out. Those were never theirs. Nicknames from your inbox were never theirs. What remains is their style, which you did not need to lock.

Fourteen days to a spine

Start with money-moving schemes and the ten legal phrases that have already caused a correction.

  1. Day 1–2: export titles from the website and the last year of OMs.
  2. Day 3: add nicknames from tickets.
  3. Day 4: import any Rajbhasha or official bilingual pairs.
  4. Day 5: fill status and cite.
  5. Day 6: list banned forms from last month's bad drafts.
  6. Day 7: load the file into the runtime. Implement inbound aliasing.
  7. Day 8: implement output check and mask-on-translate.
  8. Day 9: put glossary hits into the officer view.
  9. Day 10: write the owner and the OM trigger.
  10. Day 11–12: sit a fifty-item name pack. Compare to last month.
  11. Day 13: export test — can another runtime load the file.
  12. Day 14: file the version. Postpone the larger-model indenture if the pack moved enough.

How this shows up in the file

The note should say: we will not buy parameters to fix unlocked names. The attached glossary is authoritative on titles, offices and listed legal phrases. The agent must fail a turn that emits a banned form. The owner and the OM trigger are named.

Attach the file version and the export test.

This article is informational field guidance for Indian public institutions, not legal, procurement, security-accreditation, linguistics or engineering advice. Confirm against the current Gazette, Official Languages Act and Rules, state official-language law, GIGW, RPwD Act, DPDP text and Rules, CERT-In directions, departmental manual and your counsel before you file it.

How to test this with real speech, not staff English

“Domain Glossaries: The Cheapest Accuracy Win” fails in the field if you only tested officers. A P4 Programme/Implementation should hear a first-generation student, a rural caller, or a Hinglish grievance before claiming “domain glossary AI accuracy”.

Before you rent a larger checkpoint, lock the fifty phrases that already decide your files. A glossary is cheaper than a model, and you can audit it. Twenty-two scheduled languages is a Constitution fact, not a model fact. Script support is not language support. Official language rules may require bilingual output even when the model prefers one script.

  • Name the languages and scripts in the eval set.
  • Include code-mix and scheme-name tests.
  • Measure comprehension, not BLEU alone.
  • Design a human fallback when language fails.

Close this loop before the next CAB

Put “Domain Glossaries: The Cheapest Accuracy Win” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P4 Programme/Implementation, not “the vendor.”

Revisit the item when the model, the GeM term, the region, or the SI changes. “domain glossary AI accuracy” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.

What must be true before you file this

If “Domain Glossaries: The Cheapest Accuracy Win” is only a heading, it will not survive a file inspection. A P4 Programme/Implementation should be able to attach one artefact that proves “domain glossary AI accuracy”: a log export, a clause, a scored row, a dated notice, or a refusal rule.

Write three dated sentences: what was decided, who owns it, and when it will be re-checked. Unsigned sentences are souvenirs. Dated sentences are controls.

  • Name the owner of “domain glossary AI accuracy” inside the institution.
  • Attach one artefact a stranger can open next year.
  • Revisit when the model, the notice, or the SI changes.
  • Do not treat a vendor slide as evidence.

One more artefact before you close the file

Add a one-page owner map: who runs this after the vendor leaves, who can stop it, and where the logs live. If those three names are missing, the project is still a demo.

Date the page. File it next to the contract. That is the difference between a blog you read and a control you can audit.

What the next file must contain

“Domain Glossaries: The Cheapest Accuracy Win” earns a line in the noting only if a P4 Programme/Implementation can attach proof of “domain glossary AI accuracy.” A heading is not proof. A vendor slide is not proof. A workshop photograph is not proof.

Write three dated sentences: what was decided, who owns it after the next posting order, and when it will be re-checked. If you cannot write the three sentences, you are not ready to buy, to sell, or to go live.

Leave unsourced percentages out of the note. DPDP is not a blanket localisation statute. The November 2025 AI governance text is guidance, not an Act. CERT-In’s 28 April 2022 directions still set specified incident and log clocks. A PAC, when lawful, lives in GFR Rule 166.

  • Name the designation that owns “domain glossary AI accuracy.”
  • Attach one artefact a stranger can open next year.
  • Record the instrument you are actually using.
  • Revisit when the model, the SI, the notice or the posting changes.

Questions this usually raises

Is a glossary just a word list?
No. Each entry has official forms in each authentic language, banned equivalents, aliases, a status, and a document id. A list without those fields will be ignored by the model and by the staff.
Will a glossary make the agent sound stiff?
It will make the agent sound like the department on the words that must not drift. You can still write a warm sentence around a locked title.
Can we buy a national Indic glossary and skip this?
You can reuse public term banks as a start. Your nicknames, your office names and your live-versus-closed status will not be in them. Do the fortnight.
Where does the glossary live?
In version control the department owns, loaded at runtime, cited in the eval. Not in a slide. Not only inside a vendor prompt you cannot export.
Does a glossary replace fine-tuning?
It replaces a surprising amount of fine-tuning for names and legal phrases. Fine-tune later for register, if you must, on text that already respects the glossary.

Sources