Indic & Citizen Services
Script Support vs Language Support: The Gap
· 9 minute read
A portal that paints Tamil chrome and reasons in English has script support. It does not have language support. Score them as different rows or you will buy a font.
The demonstration opened in Tamil. The menu was Tamil. The vendor said the product was Tamil-first. The evaluator typed a citizen sentence about a night-soil worker's death benefit. The retrieval set that appeared in the debug pane was an English FAQ from another state. The reply was fluent Tamil that described a different scheme.
The committee had a row called language support. They had already given it full marks when the login page rendered. They had scored a script.
This explainer is the gap we keep finding in RFPs: script, encoding, input method, rendering, and language competence are five different purchases. If you fold them into one tick, you will receive a font pack and a press note.
Five layers that get one word in a brochure
Encoding: can the store hold the code points without mojibake. Rendering: can the officer's machine draw the conjuncts. Input: can a citizen type with a phonetic keyboard or a native one. Script competence: can the model read and write that script, including mixed lines. Language competence: can it do the job in that language — retrieve, keep terms, match the register.
A product can pass the first three and fail the last two. That product will look finished in a screenshot.
| Layer | What you test | Typical false pass |
|---|---|---|
| Encoding | Round-trip a nasty string through the MIS | UTF-8 in the app, legacy encoding in the dump |
| Rendering | Officer laptop and citizen phone, default fonts | Vendor laptop with a paid font pre-installed |
| Input | Phonetic and native keyboards on a cheap Android | Only English keyboard with paste from a translator |
| Script competence | Held-out strings in that script, including mix | English prompt behind a translated chrome |
| Language competence | Task success on your tickets in that language | A greeting and a proverb |
The English brain behind a regional face
Many stacks still translate inbound text to English, retrieve in English, and translate the reply. That can be an honest architecture if you declare it, locate the hop, and accept the damage to terms. It is not 'native language support'. Write it as 'English pivot, declared'.
The Official Languages Act's bilingual duty on Union instruments does not mean an English brain is always acceptable. It means the Union already issues many things in two authentic languages. Your agent should be able to retrieve both, not flatten both into one pivot.
One script is not one language
Devanagari carries Hindi, Marathi, Konkani, Sanskrit, Nepali and more. Perso-Arabic in India may be Urdu or Kashmiri. Latin carries English and a large share of romanised regional messages. If you write 'Devanagari supported', you have specified a script. You have not specified Marathi competence on a Maharashtra scheme.
The reverse is also true. A model that knows spoken Tamil from speech data may still fail written Tamil with the right conjuncts. Speech, script and language are different rows.
How to write the RFP so a font cannot win
Give script and rendering a small, fail-able platform score. Give language competence the marks, on a held-out set, per language you will actually serve. Forbid demo chrome from counting toward language marks. Ask for the reasoning language and the index language in the architecture note.
If a vendor wants marks for twenty-two scripts, ask for twenty-two rendering tests and still refuse twenty-two language ticks without twenty-two evals.
Write the year-one script list separately from the year-one language list. A state that must render Urdu, Hindi and English on officer machines is not thereby buying Urdu, Hindi and English agents. The first list is a desktop standard. The second is an eval. Mixing them is how a font pack wins a language row.
Objections you will hear — and what to do with them
These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.
Users only look at the screen. Chrome is what they mean by language.
Users look at the answer. A correctly drawn wrong scheme is worse than an ugly right one. Score chrome. Do not stop there.
We can add languages by adding fonts.
You can add rendering that way. You cannot add competence that way. Put the sentence in the minutes so the next indenture cannot pretend otherwise.
A multilingual embedding space means the model understands every script we throw at it.
Embeddings can retrieve neighbours. They do not certify that a defined term survived, or that the reply is in the citizen's language. Sit the eval.
Splitting rows will make the technical mark sheet too long.
Five rows is shorter than a year of a portal that only looks local. Long mark sheets are how you stop buying adjectives.
Screenshots are not evals
A screenshot of Tamil chrome will circulate in the WhatsApp group of the committee the night before marking. Prepare a counter-screenshot: Tamil chrome, English retrieval debug, wrong scheme. Practice the sentence: this is script support. Language is the next row.
If the vendor will not show the retrieval language, score competence as not demonstrated. Silence is a result.
A week to split the tick
Take the current draft RFP and a red pen.
- Day 1: highlight every use of 'language support'. Replace each with script, input, rendering or competence.
- Day 2: add a platform test for encoding and fonts on an officer laptop that is not the vendor's.
- Day 3: write the architecture question: what language does the index live in, and is there a pivot.
- Day 4: attach a ten-item competence set per year-one language.
- Day 5: brief the committee with two screenshots — Tamil chrome, English retrieval. Practice failing it.
How this shows up in the file
The note should say: we distinguish script, encoding, input, rendering and language competence. Marks for language require the attached eval. A translated user interface is scored under platform, not under language.
Attach the five-layer table and the architecture question the vendor must answer in one paragraph.
This article is informational field guidance for Indian public institutions, not legal, procurement, security-accreditation, linguistics or engineering advice. Confirm against the current Gazette, Official Languages Act and Rules, state official-language law, GIGW, RPwD Act, DPDP text and Rules, CERT-In directions, departmental manual and your counsel before you file it.
How to test this with real speech, not staff English
“Script Support vs Language Support: The Gap” fails in the field if you only tested officers. A P2 Procurement should hear a first-generation student, a rural caller, or a Hinglish grievance before claiming “script vs language support AI”.
A portal that paints Tamil chrome and reasons in English has script support. It does not have language support. Score them as different rows or you will buy a font. Twenty-two scheduled languages is a Constitution fact, not a model fact. Script support is not language support. Official language rules may require bilingual output even when the model prefers one script.
- Name the languages and scripts in the eval set.
- Include code-mix and scheme-name tests.
- Measure comprehension, not BLEU alone.
- Design a human fallback when language fails.
Close this loop before the next CAB
Put “Script Support vs Language Support: The Gap” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P2 Procurement, not “the vendor.”
Revisit the item when the model, the GeM term, the region, or the SI changes. “script vs language support AI” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.
What must be true before you file this
If “Script Support vs Language Support: The Gap” is only a heading, it will not survive a file inspection. A P2 Procurement should be able to attach one artefact that proves “script vs language support AI”: a log export, a clause, a scored row, a dated notice, or a refusal rule.
Write three dated sentences: what was decided, who owns it, and when it will be re-checked. Unsigned sentences are souvenirs. Dated sentences are controls.
- Name the owner of “script vs language support AI” inside the institution.
- Attach one artefact a stranger can open next year.
- Revisit when the model, the notice, or the SI changes.
- Do not treat a vendor slide as evidence.
Questions this usually raises
- If the UI is in Devanagari, does the agent support Hindi?
- Not necessarily. The chrome can be translated while the model, the corpus and the tools stay English. Ask where reasoning happens and on which index.
- Is Unicode compliance the same as Indic support?
- Unicode is the floor for storage and rendering. It does not mean the model can read a mixed ticket, keep a scheme name, or draft a noting. Put Unicode in the platform row. Put language in the eval row.
- Does GIGW require language support or script support?
- GIGW is about quality, accessibility and lifecycle of government websites and apps, including multilingual publishing expectations. It does not certify that your agent understands a language. Meet GIGW. Still sit a language eval.
- Can a Latin-script model 'support' Hindi?
- It can support romanised Hindi if you test romanised Hindi. That is not Devanagari Hindi. Write the script in the row. Citizens will use both.
- Should we score fonts in the AI tender?
- Score fonts in the client row if officers cannot otherwise read the output. Do not let a font pack steal marks from the language eval.
Sources
- Official Languages Act, 1963 — Department of Official Language
- Constitution of India — Eighth Schedule (languages)
- Guidelines for Indian Government Websites and Apps (GIGW 3.0)
- BHASHINI — National Language Translation Mission (MeitY)
- MeitY — India AI Governance Guidelines (5 November 2025, PIB PDF)
- Prcept AI — on-prem / air-gapped agents