Compute & Cost
Inference Cost per Citizen Interaction, Modelled
· 9 minute read
A citizen turn is not one forward pass. It is retrieval, tools, retries, a failure pack, and sometimes a human. Model the unit cost that way. Do not publish a fake rupee-per-chat for India.
A dashboard said each chat cost ₹1.20. The arithmetic was a monthly cloud bill divided by sessions that included accidental opens and staff tests. It ignored the clerk who finished half the cases, the retrieval cluster, and the Indic token tax. The minister quoted ₹1.20 in a speech. The true factory cost was a different animal. Nobody had modelled the animal.
This is a unit-cost method. We will not crown a national rupee-per-interaction. Workloads, languages, and whether a human finishes the job move the number by more than any GPU discount. If you need a rupee, fill the sheet from your meters and your roster.
The method is useful even without rupees. It tells you which lever to pull: shorter retrieval, fewer retries, a better failure pack, a smaller model for FAQ, a desk for the tail. Levers are why you model.
Not a tariff. Not a GeM rate. A sheet.
Define the unit or you will lie
Pick one: a citizen session that reaches an outcome; a single turn; a completed grievance; a completed IVRS call. Do not mix. Staff tests do not count. Accidental opens do not count. If you must report sessions, also report completed outcomes.
An outcome can be 'answered from FAQ', 'handed to window 2', or 'filed a ticket'. 'Bot spoke' is not an outcome.
The stack inside a turn
Retrieval tokens and embedding lookups. Generation tokens, including the system prompt and tool traces the citizen never sees. Retries when the tool fails. The pre-written failure pack (cheap) or a model-generated apology (not cheap and unsafe). The human fallback when language or rights require it — usually the dominant rupee if it happens.
On-prem, convert GPU-minutes plus a share of people and power into the unit. On a billed API, use their counter on your Indic pack, plus the same people share. Do not compare an API sticker to an on-prem sticker without people.
| Component | How you measure | Often small / often large |
|---|---|---|
| Retrieval + embeddings | Calls × tokens × your rate or GPU-min | Large on document-heavy turns |
| Generation | In/out tokens at production mix | Grows with Indic tax and long answers |
| Retries / tool loops | Extra generations per outcome | Large when MIS is flaky |
| Failure pack | Almost free if pre-written | Becomes large if the model narrates |
| Human fallback share | Desk minutes × loaded wage × rate of handoff | Dominates whenever language fails |
| Allocated operate (power, AMC, owners) | Monthly factory / completed outcomes | Dominates if volume is tiny |
A worked shape, not a price
Imagine 10,000 completed citizen outcomes a month. 80 percent die on a short FAQ with modest retrieval. 15 percent need a tool and a retry. 5 percent go to a human for eight minutes. The GPU line on the 80 percent may look like paise-to-a-few-rupees on many stacks; we will not pin it. The 5 percent desk line will look like tens of rupees per those outcomes, and a few rupees when spread across all 10,000. The story is the mix, not the paise.
If you improve language and the handoff rate falls from 5 to 3 percent, the unit cost moves more than if you shave 10 percent off the GPU sticker. That is why the language cluster sits next to this compute cluster.
If volume is 400 outcomes a month, allocated operate dominates and unit cost looks 'bad'. That is not a reason to switch models. It is a reason to ask whether the service should exist at that volume, or should share a pair.
What not to present to a minister
A single rupee without the mix. A comparison to a private chatbot's blog. A number that excludes the desk. A number computed in English tokens. A number from a week that included a load test.
Do present the mix, the levers, and a band. Ministers can hear 'mostly paise of silicon and a few rupees of desk, unless language fails'. They cannot hear a false precision.
Two unit costs
Objections you will hear — and what to do with them
These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.
Just give us the industry benchmark.
There is no honest Indian public-service benchmark we will cite. Anyone selling one minted it.
We should ignore the desk; it existed before.
If the agent changes the desk load, the change is in the unit. If it does not change the desk, say so and do not claim savings.
Token billing will give us the real number automatically.
It will give you the silicon slice if you are on an API. It will not give you retries done off-meter, or the desk.
Per-interaction cost will be used to kill the project.
Then show the mix and the alternative — full desk, or no service. A killed project that was a vanity queue is not a tragedy.
Three weeks to a unit band
Logs first, rupees second.
- Week 1: define the outcome. Pull mix: FAQ / tool / handoff. Measure tokens on the Indic pack.
- Week 2: attach GPU-min or live rate card. Attach loaded desk wage. Allocate operate honestly.
- Week 3: publish a band and two levers. Do not publish a single rupee unless finance demands it, and then show the mix beside it.
File note you can paste
Subject: Unit cost model for [service] — method.
The unit is [completed outcome]. Mix, tokens, desk minutes and allocated operate are attached. No national benchmark is adopted. Silicon rates come from [live card or on-prem minutes], dated. This is not a tariff order.
Sensitivity without theatre
Move one lever at a time. Handoff rate plus or minus two points. Token tax plus 30 percent for a fatter Indic circular. Volume halved. Publish those three bands. A tornado chart with twelve coloured bars is how you hide that you do not know the mix.
If the unit only looks acceptable when handoff is magically 1 percent, you do not have a unit cost. You have a hope. Fund the language row or fund the desk. Do not fund a slide.
What we will not average
Prcept AI will fill this sheet from your logs. We will not give a reporter a single rupee with our name on it. We will show you if language, not silicon, is the expensive atom.
A sovereign service that cannot state its mix is not cheap. It is unmeasured.
This article is informational field guidance for Indian public institutions, not legal, language-policy, procurement, finance or engineering advice. Confirm against the live Gazette, Official Languages Act and Rules, your State's official-language law, MeitY / IndiaAI notices, GFR, GeM terms, DPDP text, departmental manuals and your counsel before you file it.
How to put this in the finance note
A P1 CIO/CTO searching “cost per interaction AI” needs a number a CFO can defend, not a GPU brand. “Inference Cost per Citizen Interaction, Modelled” belongs in a cost model with people, power, idle time, AMC and the cost of a failed pilot.
A citizen turn is not one forward pass. It is retrieval, tools, retries, a failure pack, and sometimes a human. Model the unit cost that way. Do not publish a fake rupee-per-chat for India. IndiaAI subsidy, if you use it, is a live notice — not a permanent discount. On-prem TCO includes ops headcount. Do not invent Rs/hour. Cite the source of every rupee.
- Separate capex, opex, and one-time cleanup.
- Show utilisation, not just peak GPUs.
- Price the human fallback, not only inference.
- Date every tariff and subsidy assumption.
Close this loop before the next CAB
Put “Inference Cost per Citizen Interaction, Modelled” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P1 CIO/CTO, not “the vendor.”
Revisit the item when the model, the GeM term, the region, or the SI changes. “cost per interaction AI” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.
What must be true before you file this
If “Inference Cost per Citizen Interaction, Modelled” is only a heading, it will not survive a file inspection. A P1 CIO/CTO should be able to attach one artefact that proves “cost per interaction AI”: a log export, a clause, a scored row, a dated notice, or a refusal rule.
Write three dated sentences: what was decided, who owns it, and when it will be re-checked. Unsigned sentences are souvenirs. Dated sentences are controls.
- Name the owner of “cost per interaction AI” inside the institution.
- Attach one artefact a stranger can open next year.
- Revisit when the model, the notice, or the SI changes.
- Do not treat a vendor slide as evidence.
Questions this usually raises
- Should we cost per turn or per session?
- Cost the outcome the citizen wanted. Turns are an engineering metric. Sessions without outcomes are a vanity metric.
- How do we treat partial automation?
- Split the mix. A turn that always needs a clerk is a desk product with a draft. Cost it that way.
- Do we include citizen travel if the kiosk fails?
- Not in the silicon unit. You may show it as a dignity cost in the same note. Do not pretend it is zero.
- Will a smaller model cut unit cost in half?
- Only if tokens and retries fall and quality does not push people to the desk. Bench the mix, not the parameter count.
- Can we recover this as a user charge?
- That is a political and legal question we will not answer. Most citizen services should not try to recover GPU paise at the window.
- How often do we recompute?
- When the language mix, the model hash, or the handoff rate changes, and quarterly otherwise.