Indic & Citizen Services
Testing With Real Citizens, Not Just Staff
· 9 minute read
A section officer who types in both scripts is not your user. The user is the person who travelled forty kilometres with a PDF on a cheap phone and one working language. If you have not watched that person fail, you have not tested the service.
The UAT sign-off sat in a conference room with good air-conditioning and six people who already knew the scheme. They typed complete sentences. They used the words the circular used. They smiled when the agent used their name. The minutes said user testing complete. The first Tuesday of production, a woman in the queue asked the kiosk, in one breath, whether the widow pension and the house-tax rebate were the same form. The agent offered her a link to the home page. She left. The clerk later filed her case as 'citizen did not cooperate'.
Staff testing is necessary. It is not sufficient. Officers are paid to be patient with software. They have staff IDs, bilingual keyboards, and a colleague to ask. Citizens have a bus to catch, a language the RFP forgot, and a phone that still has 11 percent battery.
This guide is for CIOs who need a field test that will survive a newspaper story and a DPDP question. It is not a market-research playbook from a consumer app. It is how you watch real people use a public agent without turning them into unpaid labour or an unlawful dataset.
Not legal advice. Ethics board, DPO and the scheme owner still decide who you may approach and what you may record.
Who is not a user, however senior
The vendor's bilingual demo person is not a user. The NIC intern is not a user. The secretary who tries the English path on a laptop is not a user. The CSMMC operator who has seen the script twenty times is not a user. They can break builds. They cannot certify the service.
A real user, for a citizen-facing agent, is someone who needs the thing the agent claims to do, in the language and channel you claim to support, without having been briefed on the ontology. If you only have staff, you have a dress rehearsal.
- At least two districts, not both headquarters.
- At least one person who does not read the UI language well.
- At least one feature phone or low-end Android if you claim those channels.
- At least one code-mix speaker if your tickets show code-mix.
- At least one person who arrived with a photo of a document, not a PDF.
Lawful, kind, and not a hidden training run
Field tests process personal data. Under the Digital Personal Data Protection Act, 2023 you still need a purpose, a lawful basis story your DPO will sign, and a plan for what happens to the recordings. 'We are improving the model' is a different purpose from 'we are watching whether the kiosk is usable'. Do not smuggle training into a usability session.
Consent for a recorded session is not a 14-page notice on a sunlit screen. Say who you are, what you will record, how long you will keep it, that refusal does not affect their real application, and how they can leave. Then mean it. If the person still has to finish the real form, pay for their time or run the test on a dummy case. Unpaid queues are not research. They are extraction.
Do not test eligibility decisions on live rights without a human who can complete the job the same day. A failed prototype is allowed to confuse a volunteer. It is not allowed to strand a widow.
What you are measuring, if you are honest
Not satisfaction stars. Not BLEU. You are measuring whether a person who needed a result got a next step they could perform, in a language they could use, without being blamed. Time to first useful answer. Number of times they switched to the clerk. Number of times the agent asked them to retype a number. Number of times the failure message was in the wrong language.
Watch hands, not only transcripts. People cover the camera. They tilt the Aadhaar card. They ask a neighbour. They accept the first button because it is green. A transcript that looks clean can still be a failure if the neighbour did the work.
Score language separately from task success. An agent that completes the English task and abandons the Hindi speaker is not 50 percent successful. It is a gated service.
| Signal | How you capture it | Fail if |
|---|---|---|
| Task completion without a clerk | Observer tick, not self-report | Majority need the clerk for the happy path |
| Language match | Citizen's first utterance vs agent's replies | Agent switches to English and stays there |
| Numeral survival | Compare spoken/written amount to store | Any silent change to a money or date field |
| Failure path | Force one outage in a dry run | Message is English-only or has no next step |
| Dignity | Observer note, optional short ask | Citizen is blamed, mocked, or asked to 'speak properly' |
Staff still matter — just not as fake citizens
Use staff to break the admin path, the override, the audit export, and the evil inputs. Use them to check that the clerk can take over without retyping. Do not dress them as villagers in the UAT minutes.
Field staff are a different population again. They will tell you what they will actually do at 4:40 pm. That is the next article. Do not mix their workshop with citizen sessions. The power in the room is different.
Two tests, only one counted
Objections you will hear — and what to do with them
These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.
We do not have time before the launch date.
Then launch to staff only, or launch a channel you have tested. A date in a speech is not a reason to use untested citizens as crash-test dummies.
Real citizens will leak the scheme and game it.
Usability tests use dummy cases or already-decided cases. You are not previewing eligibility. You are watching whether people can ask and understand.
A panel vendor can recruit for us.
They can recruit. They cannot replace your observer, your language mix, or your DPO's purpose note. A consumer panel in a metro mall is not a tehsil queue.
[Accessibility](/blog/accessibility-standards-for-indic-ai-interfaces) is a later phase.
Then do not claim the kiosk is for everyone. Write the excluded population on the file. Later-phase is honest. Silent exclusion is not.
A four-week citizen test that fits a government calendar
Start from last month's tickets, not from a persona workshop.
- Week 1: DPO writes the purpose note. Scheme owner writes which tasks are dummy-only. Pull the language mix from real tickets.
- Week 2: recruit through field offices, not through the vendor. Pay travel or run at a place people already are, after they finish the real work.
- Week 3: ten to twenty sessions. Two observers. One is silent. Record only what the note allows. Stop a session that becomes a live rights decision.
- Week 4: file the score table, the clips you are allowed to keep, and the change tickets. UAT sign-off that ignores this file is incomplete.
File note you can paste
Subject: Citizen testing of the proposed agent — conditions before public go-live.
Staff UAT is recorded separately and is not treated as citizen evidence. A field series of at least ten sessions, covering two locations and the languages we claim, will be completed under a DPO purpose note. Recordings will not be used to train a model unless a separate, lawful basis is written. Refusal will not affect any real application.
Go-live on a citizen channel requires the attached score table to pass. This note is not legal advice.
What we will sit through
Prcept AI will sit at the back of a tehsil test without a sales script. We will take the fail. We will not ask you to sign UAT on our bilingual staff. If we cannot watch a stranger use the agent in the language on the brochure, we will tell you the channel is not ready.
A sovereign stack that only officers can use is a secretariat toy. Build for the queue.
This article is informational field guidance for Indian public institutions, not legal, language-policy, procurement, finance or engineering advice. Confirm against the live Gazette, Official Languages Act and Rules, your State's official-language law, MeitY / IndiaAI notices, GFR, GeM terms, DPDP text, departmental manuals and your counsel before you file it.
Questions this usually raises
- How many citizen sessions are enough?
- Enough to cover each claimed language and channel at least twice, in more than one location, with people who were not briefed. Ten honest sessions beat a hundred staff scripts. There is no magic n that replaces a missing language.
- Can we use the sessions to fine-tune the model?
- Only if that purpose was written, explained, and lawful. Default is no. Usability and training are different purposes. Do not bury training in a clipboard form.
- Do we need an ethics board?
- Many departments will not have one. You still need a DPO note, a scheme-owner note, and a way to refuse. Universities and health departments may have additional review. Ask before you record faces.
- What if citizens only speak a language we did not build?
- That is a coverage finding, not a user failure. Write it. Offer the human desk. Do not force them through a language they do not have so the completion rate looks better.
- Should the vendor run the sessions?
- The vendor may log issues. The department owns the protocol, the recruiting, and the pass mark. A vendor-run test will find vendor-shaped problems.
- Is a phone survey after the visit enough?
- It is a supplement. People say the service was fine because they finally got the certificate from a clerk. Watch the hands.