All insights

Indic & Citizen Services

Measuring Comprehension, Not Just Translation

· 10 minute read

The scored event is whether the citizen can do the next step. A high translation score on a sentence they cannot use is a failed public service.

A portal team celebrated a jump in Hindi automatic metrics. In a district office, we asked five people who had just used the agent what they would do tomorrow. Three pointed at the wrong window. One had understood a date that was a week earlier than the circular. One had not heard the part about documents because the reply was a single spoken paragraph.

The metric had measured resemblance to a reference sentence. The citizen had to measure a building.

This guide is how to score comprehension for Indic agents: teach-back, task completion, and the humility to shorten a correct sentence until it can be used.

The scored event

After the reply, can the person say the next step, the place or number to use, the date, and the document to carry — in their own words. Then, if you can, can they actually start the step. Those are the marks. Style is a distant third.

If the service is voice, the teach-back is spoken. If the service is WhatsApp, the teach-back can be a reply. If the person cannot read the script you sent, you have already failed, however perfect the translation memory.

Design the test without theatre

Recruit people who look like the catchment, including low-literacy users if they are in the catchment. Pay them. Do not use only staff. Staff already know the window.

Give the same prompt the public will give. Do not coach. After the agent's reply, ask for the next step. Record whether the official fact survived. Then offer a human explanation and see whether the design, not the citizen, was the problem.

Split by language and by channel. A sentence that works as text may fail as speech. A sentence that works in formal Hindi may fail in the mix the citizen actually uses.

A minimum comprehension scorecard.
ItemPassFail
Next actionUser names the action the office intendedWrong window, wrong scheme, or no action
Date / amountMatches the instrumentAny drift
Document listUser can list what to carryInvented document or a missed mandatory one
Channel fitUser can consume the reply on that channelUnread script, unheard paragraph, clipped audio
Correctness firstThe understood action is the lawful oneFluent wrong instruction

What to change when they fail

Shorten. One idea a turn. Put the next step first. Keep the scheme title locked. Offer a number key or a missed-call path if the SMS script is unreadable. Do not 'add more explanation' to a sentence that already failed. Length is often the defect.

If one language cell fails and another passes, you have a language problem, not a general UX problem. Fund that cell or stop advertising it.

Ethics of the test

Do not record people without a purpose and a retention clock. Do not publish their faces as a success story. Do not use the session to collect extra personal data for a model. Comprehension testing is service design, not a hidden training harvest.

Objections you will hear — and what to do with them

These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.

This is qualitative. We cannot put it in a tender.

You can require a teach-back protocol and a minimum pass on the next-action item. You already accept officer rubrics. This is the citizen version.

Users will be shy and we will fail good text.

Shy users are users. If they cannot say the step in a paid test, they will not find the window in a crowded office.

We will do this after go-live with analytics.

Analytics will show drop-off, not the wrong window. A week of teach-back before go-live is cheaper than a newspaper photograph of the wrong queue.

Plain language will water down the legal meaning.

Then send two layers: a one-line next step the user proved they understood, and a link to the authentic text. Do not make the authentic paragraph the only SMS.

Windows beat metrics

If three paid users walk to the wrong window, the automatic score is a side issue. The window is the service. Committees that have never watched a teach-back will over-weight a BLEU table because it looks like engineering. Invite them to the room.

A one-page note after the session — next-action pass by language and channel, no names — is enough to stop a go-live. You do not need a journal paper.

Twelve days to a teach-back you can file

Pick one task. Five to fifteen people per year-one language is enough to start.

  1. Day 1: pick the task and the lawful next step.
  2. Day 2: write the scorecard. Correctness first.
  3. Day 3–4: recruit and pay. Include low-literacy users if they are in scope.
  4. Day 5: run text channel.
  5. Day 6: run voice or WhatsApp if those are in scope.
  6. Day 7: tabulate next-action passes. Do not average away a language cell.
  7. Day 8–9: shorten the failing replies. Do not add a paragraph.
  8. Day 10: retest the failures.
  9. Day 11: write the rule — no go-live if next-action pass is below the floor you set.
  10. Day 12: file the protocol and the anonymised counts.

How this shows up in the file

The note should say: we will not go live on a language-channel pair that failed teach-back on the next action, even if automatic translation metrics improved. Unofficial explainers remain labelled. Authentic text remains linked.

Attach the scorecard and the anonymised cell counts. Do not attach names or faces.

This article is informational field guidance for Indian public institutions, not legal, procurement, security-accreditation, linguistics or engineering advice. Confirm against the current Gazette, Official Languages Act and Rules, state official-language law, GIGW, RPwD Act, DPDP text and Rules, CERT-In directions, departmental manual and your counsel before you file it.

How to test this with real speech, not staff English

“Measuring Comprehension, Not Just Translation” fails in the field if you only tested officers. A P1 CIO/CTO should hear a first-generation student, a rural caller, or a Hinglish grievance before claiming “comprehension testing citizen services”.

The scored event is whether the citizen can do the next step. A high translation score on a sentence they cannot use is a failed public service. Twenty-two scheduled languages is a Constitution fact, not a model fact. Script support is not language support. Official language rules may require bilingual output even when the model prefers one script.

  • Name the languages and scripts in the eval set.
  • Include code-mix and scheme-name tests.
  • Measure comprehension, not BLEU alone.
  • Design a human fallback when language fails.

Close this loop before the next CAB

Put “Measuring Comprehension, Not Just Translation” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P1 CIO/CTO, not “the vendor.”

Revisit the item when the model, the GeM term, the region, or the SI changes. “comprehension testing citizen services” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.

What must be true before you file this

If “Measuring Comprehension, Not Just Translation” is only a heading, it will not survive a file inspection. A P1 CIO/CTO should be able to attach one artefact that proves “comprehension testing citizen services”: a log export, a clause, a scored row, a dated notice, or a refusal rule.

Write three dated sentences: what was decided, who owns it, and when it will be re-checked. Unsigned sentences are souvenirs. Dated sentences are controls.

  • Name the owner of “comprehension testing citizen services” inside the institution.
  • Attach one artefact a stranger can open next year.
  • Revisit when the model, the notice, or the SI changes.
  • Do not treat a vendor slide as evidence.

What the next file must contain

“Measuring Comprehension, Not Just Translation” earns a line in the noting only if a P1 CIO/CTO can attach proof of “comprehension testing citizen services.” A heading is not proof. A vendor slide is not proof. A workshop photograph is not proof.

Write three dated sentences: what was decided, who owns it after the next posting order, and when it will be re-checked. If you cannot write the three sentences, you are not ready to buy, to sell, or to go live.

Leave unsourced percentages out of the note. DPDP is not a blanket localisation statute. The November 2025 AI governance text is guidance, not an Act. CERT-In’s 28 April 2022 directions still set specified incident and log clocks. A PAC, when lawful, lives in GFR Rule 166.

  • Name the designation that owns “comprehension testing citizen services.”
  • Attach one artefact a stranger can open next year.
  • Record the instrument you are actually using.
  • Revisit when the model, the SI, the notice or the posting changes.

Questions this usually raises

Is comprehension testing a scientific literacy study?
It is a service test. You ask a small, paid group of real users to do the next step after the agent's reply. You are not publishing a national literacy paper. You are deciding whether to send the reply.
Can we use BLEU or COMET as a proxy?
As a ranking aid for drafts, yes. As the go-live number, no. Those metrics do not know whether a widow can find the window.
How small can the user group be?
Small enough to run this month, large enough that a lucky five cannot hide a fail. Stratify by language and literacy, even if each cell is modest. Empty cells stay empty.
Does GIGW require this?
GIGW requires usable, accessible government properties. Comprehension testing is how you evidence usability for generated Indic text. It is good practice even where a clause does not use the word teach-back.
What if users 'comprehend' a wrong fact because the agent was fluent?
That is a hard fail. Comprehension of a wrong next step is worse than confusion. Score correctness first, then whether the correct step was understood.

Sources