All insights

Indic & Citizen Services

Northeast Languages: The Coverage Gap

· 10 minute read

A vendor slide that ticked Assamese because the Eighth Schedule exists is not coverage. Many languages of the Northeast are missing from commercial models, from public eval sets, and from your last RFP. Name the gap. Do not invent a list of who 'supports' what.

A mission-mode RFP for a 'national' citizen agent listed twenty-two official languages and a footnote that other languages would be added in phase two. The demonstration used Assamese headlines from a news crawl. The first real ticket from a district in the hills was a Roman-script mix the crawl had never seen. The helpdesk closed it as 'language not supported' and suggested English. The citizen already knew English was the gate. That was the complaint.

This is a data-study in the only honest sense: a study of what you do not know. We will not publish a table of which commercial models 'support' Khasi, Mizo, Bodo, Garo, Nagamese, Meitei, or the many languages the Census records at small counts. Those tables rot. Vendors change checkpoints. A blog that pretends to have the current matrix is a lie with footnotes.

What is stable is the gap. Public pre-training mass, open evaluation sets, keyboard and font infrastructure, and government digitised corpora are uneven. They are thinner for several Northeast languages and for Roman-script practices that never look like the literary standard. If your service area includes those speakers, the gap is an operational fact, not a diversity slide.

Not legal advice, not a linguistic survey, not a claim that Prcept or anyone else covers the region. It is a way to write the file so a later officer is not surprised.

What the famous lists do not do

The Eighth Schedule of the Constitution lists languages for specified constitutional purposes. It includes some languages widely used in the Northeast and excludes others that people actually speak in government offices, markets and churches. It is not a software inventory. It is not a ranking of dignity. It is not a procurement checklist.

The Census language tables are a different list again. They are valuable and they are not a real-time map of what a helpdesk will hear in 2026. People report a language for identity and use another for the form. Roman script is common in several States for languages that also have another script. Meitei Mayek has been revived in official use in ways a 2011 crawl will miss.

Official-language laws of the States and of the Union add a third list. An agent that satisfies a Schedule citation and fails the language of the counter has satisfied a citation.

The gap is structural, not a missing tick

Commercial multilingual models are trained where text is cheap. Web text, news, and parallel religious or literary corpora are not evenly distributed. Several Northeast languages have strong oral use, active local media, and thin machine-readable government corpora. Code-mix with English and with a regional lingua franca is the working register. If your eval is literary monolingual text, you are measuring a language the counter does not speak.

Script choice is a second gap. Some communities prefer Roman. Some official uses prefer a scheduled script. An agent that only emits one is not 'the language'. It is one orthography. Specify both if both appear in last year's tickets.

Speech is a third gap. Do not infer speech coverage from text. Do not infer text from the existence of a radio station. If you did not sit a local eval, you do not have a number. You have a hope.

A honesty table for the file. Fill from your tickets. Do not fill from a vendor brochure.
What you might assumeWhat you must checkWhat you must not write
Assamese is 'done' because it is scheduledYour tickets: script, code-mix, districtsModel X supports Assamese at production quality
Manipuri / Meitei is one rowMeitei Mayek vs Bengali script vs Roman in your channelsA single tick for 'Manipuri'
Nagamese will be covered by AssameseWhether your speakers actually write that wayCreole coverage inferred from a scheduled language
Khasi / Garo / Mizo can be phase twoShare of last year's tickets, not the minister's hometownA date without an eval set
English is a fair fallbackWhether the service is then gated on EnglishNational coverage with an English back door

How to measure without faking a leaderboard

Build a local set from de-identified tickets, school-office notes, and scheme forms. Pay speakers from the community to mark a rubric: facts, names, numbers, polite address, no invented scheme. Hold out a slice. That is the only leaderboard that matters.

If you cannot find twenty real strings for a language, you do not have a mandatory row. You have a research wish. Write the wish as a funded corpus project, not as a go-live claim.

Publish the method. You can keep the held-out items closed. The method — who marked, what the rubric was, which districts — should be on the file. A secret method is how vendor benches sneak back in.

Procurement honesty is the only coverage that scales

Year-one languages with tests. Roadmap languages with dates and new tests. Human desks for the rest, in the languages those desks actually speak. That triangle is more honourable than a map of India shaded in one vendor colour.

Fund the corpus. A department that wants a language in 2027 should be paying for annotated tickets in 2026. Models follow data. Speeches do not produce data.

Two ways to miss the hills

Objections you will hear — and what to do with them

These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.

Give us the list of models that support each Northeast language.

We will not. Any list we publish on 17 August 2026 will be wrong on 18 August. Sit your own eval. Ask the vendor to sit it. File the date.

A national mission will fix this next year.

Missions can fund corpora and compute. They do not relieve you of the year-one map. Write what you will do on Monday.

These languages have too few speakers to [justify the cost](/blog/cost-of-adding-one-more-language-modelled).

Then say so in the file, name the human desk, and do not print a coverage map. Cost is a real argument. Erasure by brochure is not.

Roman script is not the real language.

If citizens write it on the form, it is in scope. Literary committees can argue standards. The queue cannot wait for the argument.

Four weeks to an honest Northeast annexure

Do this with the State language directorate and two district offices, not only with a Delhi consultant.

  1. Week 1: count last year's tickets by district and by the language/script the clerk actually saw. No Schedule names until the count exists.
  2. Week 2: pick year-one rows that have enough strings to test. Everything else is roadmap or human desk.
  3. Week 3: pay community markers to build a small held-out pack. Write the rubric. Forbid vendor-supplied translations as the only source.
  4. Week 4: put the triangle in the RFP — tested languages, dated roadmap, staffed desks. Delete any sentence that says the product covers the Northeast.

File note you can paste

Subject: Language coverage for districts in the Northeast — honest map.

This department will not claim coverage of 'Northeast languages' as a class. Year-one agent languages are those listed in Annexure A with evaluation sets. Other languages will be served by named human desks whose staffing is attached. Roadmap languages have dates and a corpus plan, not a vendor tick.

Eighth Schedule citation is not evidence of model quality. No model-support matrix is adopted in this file because such matrices go stale. This note is not a linguistic survey and not legal advice.

What we will not shade on a map

Prcept AI will not publish a tricolour map of the Northeast with our logo in every district. We will sit the evals you fund. We will refuse a 22-language tick. We will tell you when a Roman-script mix is the real workload.

If a competitor shades the map, ask for the held-out items from your hills, not from a crawl. The crawl does not stand in the queue.

What the next file must contain

“Northeast Languages: The Coverage Gap” earns a line in the noting only if a P1 CIO/CTO can attach proof of “northeast India language AI.” A heading is not proof. A vendor slide is not proof. A workshop photograph is not proof.

Write three dated sentences: what was decided, who owns it after the next posting order, and when it will be re-checked. If you cannot write the three sentences, you are not ready to buy, to sell, or to go live.

Leave unsourced percentages out of the note. DPDP is not a blanket localisation statute. The November 2025 AI governance text is guidance, not an Act. CERT-In’s 28 April 2022 directions still set specified incident and log clocks. A PAC, when lawful, lives in GFR Rule 166.

  • Name the designation that owns “northeast India language AI.”
  • Attach one artefact a stranger can open next year.
  • Record the instrument you are actually using.
  • Revisit when the model, the SI, the notice or the posting changes.

This article is informational field guidance for Indian public institutions, not legal, language-policy, procurement, finance or engineering advice. Confirm against the live Gazette, Official Languages Act and Rules, your State's official-language law, MeitY / IndiaAI notices, GFR, GeM terms, DPDP text, departmental manuals and your counsel before you file it.

Questions this usually raises

Which Northeast languages does Prcept support today?
We will not give you a brochure list in this article. Ask for a dated eval on your tickets. If we have not sat that eval, we do not support the language for your purpose, whatever a slide once said.
Does the Eighth Schedule require us to offer every listed language in the region?
No. The Schedule is not a software mandate. Your official-language law and your service design decide the offer. Honesty about desks and tests is the duty you can actually perform.
Is Nagamese a dialect of Assamese for procurement?
Do not decide that in an RFP. If people write it, test it as its own row or staff a desk. Inferring coverage across languages is how queues get English gates.
Can we use machine translation from English as coverage?
Only if you declare the hop, test it on local names and scheme words, and accept the failure modes. Hidden translation is not coverage.
Where do we find speakers to mark an eval?
District offices, universities, language directorates, and paid community annotators. Do not use a single Delhi-based speaker of a language as the gold standard for a hill district.
Is this article a dataset?
No. It is a method. We have not enumerated speaker counts as if they were 2026 operational facts. Use Census publications and your own tickets, dated.

Sources