All insights

Compute & Cost

Token Economics for Document-Heavy Workloads

· 9 minute read

The expensive atom in a secretariat agent is not the citizen's sentence. It is the circular you retrieve, re-embed and stuff into context every time. Treat documents as a store with a budget, not as a pile you pour into the window.

An officer asked a one-line question about a house-tax rebate. The agent retrieved four gazettes, two annexures, and a 90-page FAQ, all in Hindi, all freshly embedded the night before because the job 'refresh corpus' had no increment. The window filled. The answer cited the wrong year. The GPU had worked hard to be wrong.

Document-heavy work is the default in government. Circulars, affidavits, judgments, marksheets, cadastral extracts, college rules. Token economics here is not a chat problem. It is a library problem: what you keep on the desk, what you keep on the shelf, and what you refuse to photocopy every morning.

This guide is for integrators who inherited a RAG diagram and a cron job. It sits on the tokenisation teardown and the unit-cost sheet. It will not give you a rupee per page. Pages are not tokens. Indic pages are not English pages. Measure.

Not a records-retention order. Your manual still says what you must keep. This says what the model may see.

Where the tokens actually go

Ingestion. OCR text, sometimes larger than the page's truth because the engine stutters. Embedding tokens, paid whenever you re-embed. A nightly full re-embed of a gazette that did not change is a scheduled waste.

Retrieval. You pay to search, then you pay again to stuff chunks into the prompt. Stuffing four 90-page documents is not retrieval. It is panic.

Generation. Officers ask for 'quote the para'. Quoting a para is cheaper than stuffing the file. Summarising a 90-page FAQ into the prompt 'just in case' is how you buy the wrong year.

Retries. A tool that fetches the wrong PDF will fetch another. Each fetch is tokens plus dignity if the officer waits.

Library rules for a government corpus

Version the circular. The unit of retrieval is a dated instrument, not a blob called latest.pdf. When two years are retrieved, the prompt must see the dates.

Chunk on official structure — para, clause, schedule — not on 512 English tokens. Indic conjuncts and table cells die in naive chunks.

Cache embeddings of unchanged hashes. The cron should ask 'what changed' before it spends the night.

Cap the stuff. A hard token budget per turn, with a visible 'I did not read page 70–90' rather than a silent drop. Silent drop is how the wrong year wins.

Do not embed what the model must not see. Legal holds, exam papers, and personal annexures do not belong in a general officer index. Purpose tags are cheaper than a scandal.

Spend controls an SI can implement in a sprint.
ControlWhat it stopsHow you know it works
Hash-based incremental embedNightly full re-embedJob logs show skipped hashes
Dated instrument IDsLatest.pdf confusionCitations carry a date
Structure-aware chunkingSplit paras / split tablesEval on table questions
Per-turn stuff budgetPanic stuffingp95 prompt tokens fall
Declare unread pagesSilent dropOfficer sees the gap

Paper is a channel, not a PDF worship

A photographed affidavit is not a circular. It is a one-off with personal data. Do not add it to the standing index. Process, extract, forget or retain under the file's schedule — do not let it become everyone else's retrieval noise.

Nastaliq, Odia conjuncts, and stamps will OCR badly. Budget a human pass for rights-bearing paper. Token spend on garbage OCR is how you pay twice: once for the model, once for the correction.

Print the retrieved set into the packet. Document-heavy agents that cannot show which para they used are not cheaper. They are indefensible.

When to refuse the pile

If the officer wants 'read the whole file' and the file is 400 pages of mixed scans, the honest product is a human with a marker, or a structured extraction job that runs overnight — not a chat that pretends. Write the refusal. Sell the overnight job.

If the corpus is the entire State gazette from 1956 without dates, stop. Build the dated store first. Tokens cannot organise a dump that records staff would not recognise.

Two libraries

Objections you will hear — and what to do with them

These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.

Long-context models make chunking obsolete.

They make larger stuff possible and slower. They do not date your circulars. They do not stop personal annexures from leaking into an index. Chunking is still library science.

We must retrieve everything for completeness.

Completeness is the right para plus a declaration of what was not read. Everything is panic.

Re-embedding nightly is safer.

Safer for the engineer who does not want to think about hashes. More dangerous for quality, power and traces. Incremental with an audit log is safer.

Give us tokens per typical Indian gazette page.

No. Script, tables and OCR quality move it. Run your tokeniser on ten pages you actually have.

Four weeks to a cheaper, safer library

Start with the cron. It is usually the villain.

  1. Week 1: inventory corpuses. Kill any personal one-offs in the standing index. Date the instruments.
  2. Week 2: turn on hash-based incremental embed. Log skips.
  3. Week 3: structure-aware chunking for the top ten circulars. Sit a table-question eval.
  4. Week 4: set a per-turn stuff budget and an unread declaration. Re-measure p95 prompt tokens and the unit-cost mix.

File note you can paste

Subject: Token and corpus controls for document-heavy agents.

Standing indexes will hold dated instruments only. Embeddings refresh on hash change. Per-turn context has a budget and an unread declaration. Personal paper is not indexed for general retrieval. Token counts will be measured on our tokeniser, not on an English rule of thumb. This is not a retention schedule.

What we will not pour into the window

Prcept AI will treat your circulars as a library: versions, hashes, budgets, packets. We will refuse a demo that stuffs 400 pages on a hosted giant and calls it a plan. We will show unread pages rather than silently drop them.

Sovereignty includes knowing which para was in the room. A cheap token that cannot be cited is still expensive when the court asks.

This article is informational field guidance for Indian public institutions, not legal, language-policy, procurement, finance or engineering advice. Confirm against the live Gazette, Official Languages Act and Rules, your State's official-language law, MeitY / IndiaAI notices, GFR, GeM terms, DPDP text, departmental manuals and your counsel before you file it.

How to put this in the finance note

A P4 System Integrator searching “token cost document processing” needs a number a CFO can defend, not a GPU brand. “Token Economics for Document-Heavy Workloads” belongs in a cost model with people, power, idle time, AMC and the cost of a failed pilot.

The expensive atom in a secretariat agent is not the citizen's sentence. It is the circular you retrieve, re-embed and stuff into context every time. Treat documents as a store with a budget, not as a pile you pour into the window. IndiaAI subsidy, if you use it, is a live notice — not a permanent discount. On-prem TCO includes ops headcount. Do not invent Rs/hour. Cite the source of every rupee.

  • Separate capex, opex, and one-time cleanup.
  • Show utilisation, not just peak GPUs.
  • Price the human fallback, not only inference.
  • Date every tariff and subsidy assumption.

Close this loop before the next CAB

Put “Token Economics for Document-Heavy Workloads” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P4 System Integrator, not “the vendor.”

Revisit the item when the model, the GeM term, the region, or the SI changes. “token cost document processing” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.

What must be true before you file this

If “Token Economics for Document-Heavy Workloads” is only a heading, it will not survive a file inspection. A P4 System Integrator should be able to attach one artefact that proves “token cost document processing”: a log export, a clause, a scored row, a dated notice, or a refusal rule.

Write three dated sentences: what was decided, who owns it, and when it will be re-checked. Unsigned sentences are souvenirs. Dated sentences are controls.

  • Name the owner of “token cost document processing” inside the institution.
  • Attach one artefact a stranger can open next year.
  • Revisit when the model, the notice, or the SI changes.
  • Do not treat a vendor slide as evidence.

Questions this usually raises

Should we embed scanned PDFs or wait for clean text?
If you embed garbage OCR you will retrieve garbage. For rights-bearing paper, budget a human or a better OCR path. For circulars you control, keep clean text.
How big should a chunk be?
One official unit — a para or a schedule row — plus enough overlap to keep a table together. Do not start from an English blog's 512.
Do we pay tokens to cite a para the officer already has?
You pay to retrieve and to generate the citation. You should not pay to stuff the whole file so the model can retype a para the officer could have been pointed to.
Is re-ranking worth it?
Often, if it stops panic stuffing. Bench it as p95 prompt tokens and as citation accuracy, not as a fashion.
Where do CERT-In logs meet the corpus?
Who retrieved what, when, for which purpose. A library without a reader list is not auditable.
Can IndiaAI hours absorb a bad cron?
They can hide it for a month. Then the letter expires or the queue grows. Fix the cron.

Sources