Indic & Citizen Services
Urdu and Right-to-Left in Government Systems
· 10 minute read
Urdu is not Hindi in another hat, and it is not a font file. Perso-Arabic script runs right to left, mixes left-to-right numbers and English, and breaks every government PDF pipeline that assumed Devanagari or Latin. Budget the engineering.
A State minority-welfare portal announced Urdu support on a Friday. The UI string was a PNG of Nastaliq. The form fields still ran left to right. The agent replied in Urdu that a layout engine then drew left to right, so every sentence started at the wrong end. The PDF certificate put the citizen's name in isolated letter forms that even the family could not read. The press note still said Urdu enabled.
This is a system-integrator field guide. Urdu in the Perso-Arabic script is a right-to-left language with a mature Unicode story, a well-known bidirectional algorithm, and a long list of ways Indian government stacks get it wrong. If your year-one scope includes Urdu, you are buying engineering, fonts, test users, and a PDF pipeline — not a language pack.
We are not declaring where Urdu 'must' be offered. That is official-language politics, State law, and service design. We are declaring that if you offer it, the offer is fake until RTL works in the store, the screen, the agent, and the paper.
Not legal advice and not a calligraphy brief. Confirm fonts, keyboard standards and any State Urdu academy instructions with the people who already run your Urdu desk.
Right-to-left is a stack, not a typeface
The Unicode Bidirectional Algorithm (UAX #9) decides how a mixed string of Urdu letters, ASCII digits, the rupee amount, and an English scheme name should display. Your browser, your Android WebView, your old Java report engine, and your laser printer do not all implement that story the same way. If you only set font-family to a Nastaliq face, you have decorated a broken bidi context.
HTML needs a dir attribute on the right elements, not a CSS trick on the whole portal. A portal that is English chrome with one Urdu paragraph needs that paragraph marked. A form that is Urdu-first needs the form marked, and then the numbers and Latin file IDs marked the other way. W3C has had this guidance for years. Government templates still ignore it.
Plain-text SMS and some IVRS transcripts have no markup. The logical order in the store must still be correct. Display hacks that reverse the stored string will destroy search, audit, and the next channel you add.
Joining, fonts, Nastaliq, and the PDF that no one can read
Arabic-script letters change shape by position. If your PDF library prints isolated forms, you have not printed Urdu. You have printed a puzzle. Test on the printer the tehsil actually has, not on the designer's Mac.
Nastaliq is the expected book and newspaper face for much Urdu public text in India. Naskh is easier for some UI and some OCR. Neither is optional decoration if your users cannot read the one you picked. Ask the existing Urdu cell which face they already use on notices. Then licence a font that is legally usable on government machines, embeds in PDFs, and supports the digits you will emit.
Do not screenshot text. Screen-reader users, search, and later agents cannot read a PNG. If your only 'Urdu' is an image asset, you do not have Urdu. You have a banner.
| Layer | What breaks | What 'done' looks like |
|---|---|---|
| Store | Reversed strings, mojibake, NFC/NFD fights | Logical-order Unicode, normalised, original preserved |
| Input | No keyboard, Inpage leftovers, Latin-only IME | A supported IME on kiosk and officer machines |
| UI | dir missing, English layout crushes the field | Correct bidi on forms, labels, errors |
| Agent | Model emits Urdu, renderer draws LTR | Tokens display as speakers would read them |
| PDF / certificate | Isolated letters, missing font embed | Print sample signed off by the Urdu desk |
| OCR / upload | Nastaliq scans become noise | A tested path or an honest out-of-scope |
The mixed string is the real workload
Citizens do not type literary Urdu. They type Urdu with an English scheme name, a Hindi office name, an ASCII mobile number, and a date in digits. Officers reply the same way. If your eval is a clean paragraph from a newspaper, you will pass and then fail.
Numbers are the usual injury. Digits in an RTL paragraph can jump to the wrong side of a rupee amount. A file number like A/12/2026 can split. Test those strings on every renderer you ship: web, Android, PDF, thermal receipt if you have one.
Urdu written in Devanagari, and Hindi written in Perso-Arabic, appear in some districts. If they appear in your tickets, they are rows. If they do not, do not invent them for a slide.
Models, OCR and the temptation to hop to English
Many stacks still translate Urdu to English, reason, and translate back. That hop must be declared. It will mangle honorifics, place names, and legal terms. It may also be a transfer if the translator is hosted. A beautiful Urdu reply that was reasoned in English on another continent is not a sovereignty story.
OCR on Nastaliq remains harder than OCR on Naskh or Devanagari in a lot of production pipelines. Do not assume the same confidence you have on printed Hindi. If officers upload Urdu affidavits, sit an eval on those scans or keep the field human.
Two deliveries that were not Urdu
Objections you will hear — and what to do with them
These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.
We will add the font in phase two.
Without a font and a bidi context there is no phase-one Urdu. Put RTL engineering on the critical path or remove the language from year one.
Unicode already solved this.
Unicode specifies the characters and the bidi algorithm. It does not configure your report engine, licence your font, or train your officers' keyboards. Those are project tasks.
Our model card lists Urdu.
A model that emits Urdu tokens into a broken renderer is a broken service. Score the full path, including paper.
Only a small percentage of users need this.
Then either staff an Urdu desk and say so, or build it properly for that percentage. A broken path is worse than an honest English-plus-desk path, because it produces unreadable legal paper.
Four weeks to know whether you should claim Urdu
Bring the existing Urdu translator or academy officer into the first meeting. They have already met your PDF library's victims.
- Week 1: inventory channels — web, app, SMS, PDF, agent, kiosk. Print one mixed string on each. Photograph the failures.
- Week 2: decide year-one scope. If PDF certificates are in scope, the font and library are in scope. If not, write the exclusion.
- Week 3: build a mixed-string eval: names, amounts, file IDs, addresses, one legal paragraph. Include one Nastaliq scan if OCR is claimed.
- Week 4: sit the SI and the model vendor on the same eval. Fail either row independently. A good model in a bad PDF is a fail.
File note you can paste
Subject: Urdu / Perso-Arabic right-to-left engineering — go-live conditions.
Urdu support, if claimed, includes input method, bidirectional display, a licensed embeddable font, PDF joining, and an agent path tested on mixed strings. Logical order will be stored. Visual-order hacks are not permitted in the system of record.
A model card that lists Urdu is not acceptance. The Urdu desk will sign the print samples. This note is not legal advice.
What we will build, and what we will not fake
Prcept AI will treat RTL as a product row: store, UI, agent, paper. We will not ship a banner and call it Urdu. We will declare any translation hop. We will sit your mixed-string pack on your renderer.
If your year-one volume does not justify the engineering, we will say so and help you staff the desk instead. A dignified desk beats an unreadable certificate.
What the next file must contain
“Urdu and Right-to-Left in Government Systems” earns a line in the noting only if a P4 System Integrator can attach proof of “Urdu RTL government systems.” A heading is not proof. A vendor slide is not proof. A workshop photograph is not proof.
Write three dated sentences: what was decided, who owns it after the next posting order, and when it will be re-checked. If you cannot write the three sentences, you are not ready to buy, to sell, or to go live.
Leave unsourced percentages out of the note. DPDP is not a blanket localisation statute. The November 2025 AI governance text is guidance, not an Act. CERT-In’s 28 April 2022 directions still set specified incident and log clocks. A PAC, when lawful, lives in GFR Rule 166.
- Name the designation that owns “Urdu RTL government systems.”
- Attach one artefact a stranger can open next year.
- Record the instrument you are actually using.
- Revisit when the model, the SI, the notice or the posting changes.
This article is informational field guidance for Indian public institutions, not legal, language-policy, procurement, finance or engineering advice. Confirm against the live Gazette, Official Languages Act and Rules, your State's official-language law, MeitY / IndiaAI notices, GFR, GeM terms, DPDP text, departmental manuals and your counsel before you file it.
Questions this usually raises
- Is Urdu just Hindi with a different script?
- No. Shared vocabulary does not make a shared stack. Script, direction, fonts, and much of the public register differ. Procurement that treats them as skins will fail both communities.
- Do we need Nastaliq in the UI?
- Ask the people who already publish your Urdu notices. Many public bodies expect Nastaliq on paper and will accept a clearer UI face on screens. Write the decision. Do not let a designer pick a face nobody reads.
- Can we store Urdu in Latin transliteration?
- Not as the only store if you claim Urdu service. Transliteration can be an input aid. The record the citizen receives should be in the script you promised.
- Why do numbers jump in Urdu sentences?
- Because digits are typically left-to-right runs inside a right-to-left paragraph. Without correct bidi markup or isolates, renderers reorder them. This is expected engineering, not a model bug.
- Is RTL only an Urdu problem in India?
- Perso-Arabic Urdu is the usual government case. Other RTL scripts may appear in specific communities and foreign-language desks. Do not generalise; specify the script you will ship.
- Does a larger model fix joining in PDFs?
- No. Joining is a font and shaping problem. Buy a library that knows Arabic script. Do not hope the checkpoint will draw glyphs.