All insights

Indic & Citizen Services

Handling Transliterated Names in Records

· 9 minute read

A transliterated name is not a typo. It is how the same person entered three systems. If the agent auto-merges on a fuzzy score, you will either deny a benefit or invent a double draw.

The scholarship file had Mohd. Irfan, Mohammed Irfan, and इरफान. The bank file had IRFAN MD. The agent, proud of its fuzzy matcher, merged them with a man in the next block whose father was different and whose year of birth was a decade off. Two credits went one way. A genuine student was flagged as a duplicate.

Indian public records are a museum of romanisation. There was no single official Latin spelling for most names when the first computers arrived. There still is not, outside a few documents. An agent that treats spelling as identity will punish the citizen who moved between a Hindi register, an English bank, and a regional-language school.

This guide is the matching policy we ask departments to write before they turn on auto-merge. It is not a new national ID design. It is not legal advice on Aadhaar use.

Variants are expected. Silent merges are not.

You will see vowel length ignored, honorifics moved, Mohammed shortened, Singh and Sinh, Chandra and Chander, Tamil names with and without the father's initial, Northeast names respelled by a clerk who did not speak the language, and women whose surname changed in one system and not another.

You will also see two different people who share a common name and a village. Fuzzy match without a second key will join them. That error is worse than a false miss, because it is hard to see until money moves.

Write the policy as rows, not as a library name

A product that says 'we handle Indic names' is not a policy. A policy says: which keys may propose a match, which keys may confirm a match, who may accept, what is logged, and how a citizen unpicks a wrong join.

A starter matrix. Replace the examples with your schemes.
ActionAllowed keysWho acceptsLog
Propose matchNormalised name plus one more (village / parent / year)Agent may propose onlyBoth original strings, scripts, scores
Confirm match for inquiryTwo keys plus a unique file numberSystem, if policy says soThe keys used, not a raw Aadhaar dump
Merge for paymentStatutory identifier your scheme already uses, or officerNamed officer for exceptionsBefore/after records, officer id
Refuse mergeAny conflict on sex, decade of birth, or parent nameSystem must refuseThe conflicting fields

Keep every surface form the citizen used

Store the original string, the script, and any working key. When you write to a citizen, prefer the form they last used on that channel. When you write to another system, use the form that system already has, and say you are doing so.

Do not overwrite a Devanagari school name with a Latin 'standard' because the matcher prefers ASCII. The transfer certificate will not match your standard.

Transliteration direction matters. Going from Indic script to Latin loses vowels. Going from Latin to Indic invents vowels. Do not round-trip a name through a model and save the result as truth.

What the agent may say

The agent may say: we found more than one record that might be you, here are the non-sensitive differences, please confirm. It may not say: we have updated your identity. It may not silently pick the spelling that looks most 'official'.

If the scheme already has an identity rule — a tokenised Aadhaar match, a student unique id — the agent follows that rule. It does not invent a parallel fuzzy democracy.

Objections you will hear — and what to do with them

These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.

If we do not auto-merge, the duplicate-benefit dashboard will look bad.

A dishonest merge makes the dashboard look good and the audit look terrible. Report proposed duplicates and confirmed duplicates as different numbers.

Soundex / Metaphone will handle Indian names.

Those algorithms were built for other languages. If you use a phonetic key, test it on your list and still require a second key. Do not import a 1918 American census trick as identity policy.

The vendor's entity-resolution graph is state of the art.

Then it can emit proposals with reasons. State of the art is not a delegation of the merge.

Citizens should spell their names consistently.

They should be able to. They often cannot, because the systems asked them differently. Fix the systems. Do not punish the variants.

Aliases are not merges

An alias says: this string has been used by this person. A merge says: these two records are one legal subject. The first is almost always needed. The second is almost never a job for a default library. Keep the words apart in the MIS and in the minutes.

If a dashboard cannot show proposed aliases without collapsing them, fix the dashboard. Do not fix the citizen.

Twelve days to a name-matching note

You need last year's collision list more than you need a new library.

  1. Day 1–2: pull known collisions and known false merges from the last payment cycle.
  2. Day 3: list every surface form those people used, with script.
  3. Day 4–5: write propose versus confirm versus refuse rules with two officers from the scheme.
  4. Day 6: ban silent overwrite of the original string.
  5. Day 7: decide which unique identifiers you already have a legal basis to use.
  6. Day 8: put a citizen correction path on the form — 'this is not me'.
  7. Day 9: log both strings and the keys used, in India, with a retention clock.
  8. Day 10: trial the agent in propose-only mode.
  9. Day 11: count false proposes and missed true pairs. Adjust keys, not the marketing.
  10. Day 12: write the note. Auto-merge for payment stays off until the numbers are boring.

How this shows up in the file

The note should say: transliteration variants are expected. The agent may propose matches. It may not merge records that move money or status without the keys listed in the matrix or an officer acceptance. Original surface forms are retained.

Attach the matrix, the collision sample, and the correction path.

This article is informational field guidance for Indian public institutions, not legal, procurement, security-accreditation, linguistics or engineering advice. Confirm against the current Gazette, Official Languages Act and Rules, state official-language law, GIGW, RPwD Act, DPDP text and Rules, CERT-In directions, departmental manual and your counsel before you file it.

How to test this with real speech, not staff English

“Handling Transliterated Names in Records” fails in the field if you only tested officers. A P4 Programme/Implementation should hear a first-generation student, a rural caller, or a Hinglish grievance before claiming “name matching transliteration India”.

A transliterated name is not a typo. It is how the same person entered three systems. If the agent auto-merges on a fuzzy score, you will either deny a benefit or invent a double draw. Twenty-two scheduled languages is a Constitution fact, not a model fact. Script support is not language support. Official language rules may require bilingual output even when the model prefers one script.

  • Name the languages and scripts in the eval set.
  • Include code-mix and scheme-name tests.
  • Measure comprehension, not BLEU alone.
  • Design a human fallback when language fails.

Close this loop before the next CAB

Put “Handling Transliterated Names in Records” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P4 Programme/Implementation, not “the vendor.”

Revisit the item when the model, the GeM term, the region, or the SI changes. “name matching transliteration India” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.

What must be true before you file this

If “Handling Transliterated Names in Records” is only a heading, it will not survive a file inspection. A P4 Programme/Implementation should be able to attach one artefact that proves “name matching transliteration India”: a log export, a clause, a scored row, a dated notice, or a refusal rule.

Write three dated sentences: what was decided, who owns it, and when it will be re-checked. Unsigned sentences are souvenirs. Dated sentences are controls.

  • Name the owner of “name matching transliteration India” inside the institution.
  • Attach one artefact a stranger can open next year.
  • Revisit when the model, the notice, or the SI changes.
  • Do not treat a vendor slide as evidence.

Questions this usually raises

Should the agent treat close transliterations as the same person?
Not automatically. Treat them as a proposed match with a reason. A human or a second identifier — a parent name plus village plus year, an Aadhaar token your law allows you to use, a file number — must close the merge when a right is at stake.
Is this a language problem or an identity problem?
Both. Script and spelling create the variants. Identity policy decides whether two strings may become one person. Do not let a library's default edit distance become your identity policy.
Can we normalise every name to English?
You can store a working Latin key if you also keep the original script form the citizen used. Normalising away the original is how you lose the only spelling the school register will accept.
What about caste or community honorifics attached to names?
Do not strip tokens you do not understand. A dropped suffix can be a different person or a different family. Put honorific handling in a written list, not in a silent cleaner.
Does DPDP change name matching?
Matching that relates records to an identifiable person is processing of personal data. Purpose, minimisation and a correction path belong on the file. A 'data quality' merge is still processing.

Sources