All insights

Indic & Citizen Services

When to Use a Small Indic Model Instead

· 9 minute read

A small Indic model you can rack, time and eval will beat a giant general model you cannot isolate — on the tasks that are most of a government inbox.

A district NIC room cannot rack a frontier cluster. A state SDC can rack some cards but not an ever-growing family of giant checkpoints. A ministry can rent a hosted giant and pretend the residency slide is the architecture. Those three facts, not a Twitter argument about parameters, should decide the model.

Most government agent tasks are classification, retrieval, glossary-locked answers, and short drafts. Those are places a focused Indic model, or a small general model adapted on your tickets, often wins on the only scores that matter: task success, latency, and whether the path is isolatable.

This is a decision guide. It is not a claim that Prcept only ships small models, or that large models are unethical. It is when the smaller box should win the file.

Where small wins

Classification and routing with a short taxonomy. Status answers against a structured MIS plus a glossary. Language-id and mix flags. Masked translation of a known pair. OCR-index cleanup with a human on operative fields. Any turn that must finish before a rural caller speaks into the hole.

Small also wins when the alternative is a hop you cannot write. If the giant only exists as a foreign API and the RFP said personal data stays, the small local model is not a quality compromise. It is the only compliant candidate.

Where small loses — and what to do instead

Long, messy drafting across many domains with no retrieval. Open legal research without a corpus. A first draft of a speech in a language you have not adapted. If you need that, either isolate a larger checkpoint you can actually run, or narrow the task until retrieval plus a small model is enough.

Do not use a small model as a fig leaf for an undeclared giant hop on the hard turns. That architecture is 'small until it fails', which is the opposite of a decision.

A decision table you can paste into a note.
If this is truePreferDo not
Task is classify / retrieve / locked answerSmall Indic or adapted small general, localRent a giant for the adjective
Rural voice latency budget is tightSmall on the path you timedA multi-hop giant 'just for NLU'
Personal data must not leave the rackWhatever you can isolate, often smallA larger hosted model 'only for Indic'
Task is open drafting, no corpusNarrow the task, or isolate a larger boxPretend a 7B will write the budget speech
Year-one language has no evalBuild the eval firstAssume either size covers it

Hardware honesty

A small model you cannot operate — no one to patch, no one to load the glossary, no one to re-run the eval — is not small. It is abandoned. Count operators, not only VRAM.

A large model you can time-slice on an existing SDC queue may be more honest than twenty abandoned mini-boxes in twenty districts. Read the companion problem of scale in the on-prem cluster of this site. Size of model and shape of estate are separate choices.

Indic is not a size

People say 'we need a big model for Indic'. Sometimes a larger multilingual checkpoint is better on a language you cannot adapt. Sometimes a smaller model trained or adapted on your language and your schemes wins on your tickets. The eval set in this cluster is how you know. Size is a hypothesis.

BHASHINI and other public components may give you a small, localisable piece for a slice of the job. Use them as pieces. They do not answer the size question for your agent as a whole.

Objections you will hear — and what to do with them

These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.

Leadership wants the most advanced model.

Give them the most advanced architecture: isolated, evaluated, replaceable. Advanced is not a parameter count on a slide.

Small models are a dead end.

Small models are a replaceable component. Binding yourself to one giant API is the dead end.

We will use small for Hindi and a giant API for the long tail of languages.

Write the hop. Limit the data class on the tail. Do not let the tail become the path for the languages you claimed were local.

Our SI only supports one hosted giant.

That is an SI constraint, not a physics constraint. Record it. It may be a reason to change SI.

Replaceable is the requirement

Whatever size you choose, the checkpoint must be a component. Tools, glossary, eval and logs stay. If a vendor's 'platform' will not load a second checkpoint, you have not bought a decision. You have bought a lease.

Write replaceability as a scored row. Small versus large can then be an engineering conversation instead of a hostage situation.

Ten days to a size decision on one workflow

Pick one workflow. Do not hold a theological seminar about all AI.

  1. Day 1: write the task in one sentence and the data class.
  2. Day 2: write the isolation rule from the RFP or the architecture note.
  3. Day 3: write the latency budget if a human is waiting.
  4. Day 4: sit a small local candidate on the eval pack.
  5. Day 5: sit the giant you were about to rent, on the same pack, on the real path.
  6. Day 6: compare task success, hop honesty, and latency. Not a news benchmark.
  7. Day 7: decide. If neither passes isolation, neither wins.
  8. Day 8: write the replaceability rule — checkpoint is a component.
  9. Day 9: cost the operators, not only the licence.
  10. Day 10: file the table. Refuse a parameter argument that ignores it.

How this shows up in the file

The note should say: for this workflow we chose the candidate that passed isolation and the departmental eval inside the latency budget. Parameter count was not a scored row. The checkpoint is replaceable without rewriting tools or the glossary.

Attach the two-candidate table.

This article is informational field guidance for Indian public institutions, not legal, procurement, security-accreditation, linguistics or engineering advice. Confirm against the current Gazette, Official Languages Act and Rules, state official-language law, GIGW, RPwD Act, DPDP text and Rules, CERT-In directions, departmental manual and your counsel before you file it.

How to test this with real speech, not staff English

“When to Use a Small Indic Model Instead” fails in the field if you only tested officers. A P4 Programme/Implementation should hear a first-generation student, a rural caller, or a Hinglish grievance before claiming “small language model Indic”.

A small Indic model you can rack, time and eval will beat a giant general model you cannot isolate — on the tasks that are most of a government inbox. Twenty-two scheduled languages is a Constitution fact, not a model fact. Script support is not language support. Official language rules may require bilingual output even when the model prefers one script.

  • Name the languages and scripts in the eval set.
  • Include code-mix and scheme-name tests.
  • Measure comprehension, not BLEU alone.
  • Design a human fallback when language fails.

Close this loop before the next CAB

Put “When to Use a Small Indic Model Instead” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P4 Programme/Implementation, not “the vendor.”

Revisit the item when the model, the GeM term, the region, or the SI changes. “small language model Indic” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.

What must be true before you file this

If “When to Use a Small Indic Model Instead” is only a heading, it will not survive a file inspection. A P4 Programme/Implementation should be able to attach one artefact that proves “small language model Indic”: a log export, a clause, a scored row, a dated notice, or a refusal rule.

Write three dated sentences: what was decided, who owns it, and when it will be re-checked. Unsigned sentences are souvenirs. Dated sentences are controls.

  • Name the owner of “small language model Indic” inside the institution.
  • Attach one artefact a stranger can open next year.
  • Revisit when the model, the notice, or the SI changes.
  • Do not treat a vendor slide as evidence.

Questions this usually raises

Is smaller always more sovereign?
No. A small hosted model is still a hop. A large model in your rack can be isolated. Sovereignty is the path and the control, not the parameter count.
Will a small model handle all 22 scheduled languages?
Unlikely as a serious year-one claim, and you should not buy that claim in a large model either without evals. Pick year-one languages. Measure.
Can we start small and swap up later?
Yes if the interface is the agent — tools, glossary, eval, logs — and the checkpoint is a replaceable component. If you bind the business logic to one vendor's giant API, you cannot swap.
Do small models need less evaluation?
They need the same eval, plus a latency and isolation eval the giant API often skips. Small is not a waiver.
What is 'small' in this article?
A checkpoint you can run on the hardware you already have or can honestly buy this quarter, with a latency budget a rural turn can survive. We are not printing a parameter religion.

Sources