Compute & Cost
Forecasting Usage Before You Have Users
· 10 minute read
You already have a forecast: last year's tickets, forms and peak weeks. Convert those. Do not buy an eight-GPU node because a slide showed exponential adoption.
Every bad GPU purchase begins with a sentence about users we do not have yet. The users exist. They are the officers who already open the grievance MIS, the students who already file the form, the citizens who already stand in the queue. Their last twelve months are a better forecast than a startup's adoption curve.
This guide converts an existing workload into a first-year inference and support forecast. It will be wrong, like every forecast. It will be wrong in a direction you can explain, with a sensitivity, rather than wrong because you believed a hockey stick. That is the standard.
It is not a scientific demand model. It is a file method. If your MIS cannot emit last year's counts, fix that before you size a cluster. A department that cannot count tickets cannot count tokens.
Start from the MIS, not from the model
Pull twelve months of the triggering events: tickets created, forms submitted, files opened, calls logged, exam scripts received. Keep the weekly peak, not only the mean. Government work is seasonal — budget close, admissions, disaster weeks, board exams. Size to a percentile you can name, not to a mean that makes the node look busy in a slide and idle in June.
Convert events to model calls. A grievance draft may be one call. A retrieval-heavy file may be several. An officer who pastes a 40-page PDF is a different animal from a schema-bound extraction. Write a calls-per-event assumption and a tokens-per-call or milliseconds-per-call assumption. Those two lines are the whole forecast. Everything else is arithmetic.
Apply an adoption fraction. Not everyone will use the agent on day one. A first-year fraction of 20–50 percent of eligible events is a common planning band in the files we see — and it is a band, not a measurement. Put 20, 40 and 70 percent on the sheet. If the hardware decision flips inside the band, you are buying a hope. Buy the smaller node and a reserved burst instead.
Peaks, queues and the myth of 8760
Most citizen-facing cells work a shift, not a year. A node powered 8760 hours for a 10-to-5 desk is a power forecast, not a usage forecast. Separate busy hours from powered hours. The first sizes the GPU. The second sizes the electricity line.
Queue time is part of the forecast. If the 95th percentile event cluster arrives at 11:15, concurrency — not monthly totals — decides whether officers wait. A small model that serves twelve concurrent threads can beat a large model that serves two, at the same monthly volume. Forecast concurrency, or you have forecast a blog post.
Training and eval are not usage. They are projects. Put them on a calendar: a night window, a reserved IndiaAI week, a quiet month. Do not add them into the inference mean and then buy a permanently larger card.
| Input | Source | Planning note |
|---|---|---|
| Events / week, with peak week | MIS, last 12 months | Keep the peak week; do not average it away |
| Calls per event | Design of the workflow | Draft = 1; RAG-heavy file = several; PDF-paste is a defect |
| Tokens or ms per call | A ten-case pilot on the actual model | Measure once; do not import a US token blog |
| Adoption fraction | Judgement, with a band | Show 20 / 40 / 70 before you buy silicon |
| Concurrency at peak hour | Peak events × calls / working minutes | This sizes the card; monthly totals do not |
| Training / eval hours | A calendar, not a rate | Buy a reservation or a night window |
What not to import from a vendor deck
Do not import users grow 15 percent month-on-month. Public services do not acquire users like a consumer app. They acquire gazetted mandates and, sometimes, a circular that makes the tool compulsory. If a circular is coming, write the date. If it is not, do not model it.
Do not import tokens-per-user from a coding-assistant case study. Officers are not completing functions. They are opening files. Measure ten of your files. Do not import GPU counts from another department unless you have their event volumes and their model class. They bought eight is not a forecast. It is peer pressure.
Re-forecast is part of the design
Write the first re-forecast date into the sanction. Day 60 is a good default. If actual concurrency is 2× the base case, you trigger the reserved extra card or the smaller second node you already priced. If actual concurrency is 0.3×, you power-cap and tell finance before they notice the idle on their own.
A forecast that cannot be wrong in public is not a forecast. It is a commitment dressed as arithmetic. The file should say what would make you smaller as clearly as what would make you larger.
Objections you will hear — and what to do with them
These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.
We have no MIS extracts
Then you are not ready to size a cluster. Count two weeks by hand at the counter if you must. A clipboard is a forecast. A hockey stick is not.
The vendor's calculator is industry standard
Ask for the inputs. If the inputs are not your events, the standard is a brochure.
Better to overbuy than to queue
Overbuy a reservation or a smaller extra card you can turn off. Do not overbuy a permanently powered 8-GPU node for a maybe.
Adoption will explode once officers see it
Sometimes. Model it as a sensitivity, not as the base case. Year-two budget cuts start here.
A fifteen-day forecast
- Days 1–5: extract 12 months of events, weekly, from every relevant MIS queue.
- Days 6–8: run ten live cases on the candidate model. Measure calls and tokens or latency.
- Days 9–12: build the skeleton at 20 / 40 / 70 percent adoption. Size concurrency at the peak week.
- Days 13–15: choose hardware or hours that survive the band without a new node. Write what would trigger a second card.
How this shows up in the file
Subject: Usage forecast for [workflow] — first year.
Events from [MIS] for [period]. Peak week [date, count]. Calls per event [n], measured on [sample]. Adoption band 20/40/70 percent. Concurrency at peak [x]. Hardware / hour recommendation survives the band: [yes/no]. Training and eval are calendared separately. No external hockey stick was used.
This forecast is a planning aid, not a measurement and not legal advice. Re-forecast after sixty live days.
This article is informational field guidance for Indian public institutions, not legal, procurement, tax, accounting, tariff or engineering advice. Confirm against the current Gazette, GFR, GeM term, SERC tariff order, IndiaAI portal rule, CAG mandate, DPDP text, departmental finance manual and your counsel before you file it. Figures are methods and order-of-magnitude illustrations, not a dataset of real deployments and not a substitute for a live quote.
How to put this in the finance note
A P1 CIO/CTO searching “AI usage forecasting” needs a number a CFO can defend, not a GPU brand. “Forecasting Usage Before You Have Users” belongs in a cost model with people, power, idle time, AMC and the cost of a failed pilot.
You already have a forecast: last year's tickets, forms and peak weeks. Convert those. Do not buy an eight-GPU node because a slide showed exponential adoption. IndiaAI subsidy, if you use it, is a live notice — not a permanent discount. On-prem TCO includes ops headcount. Do not invent Rs/hour. Cite the source of every rupee.
- Separate capex, opex, and one-time cleanup.
- Show utilisation, not just peak GPUs.
- Price the human fallback, not only inference.
- Date every tariff and subsidy assumption.
Close this loop before the next CAB
Put “Forecasting Usage Before You Have Users” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P1 CIO/CTO, not “the vendor.”
Revisit the item when the model, the GeM term, the region, or the SI changes. “AI usage forecasting” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.
What the next file must contain
“Forecasting Usage Before You Have Users” earns a line in the noting only if a P1 CIO/CTO can attach proof of “AI usage forecasting.” A heading is not proof. A vendor slide is not proof. A workshop photograph is not proof.
Write three dated sentences: what was decided, who owns it after the next posting order, and when it will be re-checked. If you cannot write the three sentences, you are not ready to buy, to sell, or to go live.
Leave unsourced percentages out of the note. DPDP is not a blanket localisation statute. The November 2025 AI governance text is guidance, not an Act. CERT-In’s 28 April 2022 directions still set specified incident and log clocks. A PAC, when lawful, lives in GFR Rule 166.
- Name the designation that owns “AI usage forecasting.”
- Attach one artefact a stranger can open next year.
- Record the instrument you are actually using.
- Revisit when the model, the SI, the notice or the posting changes.
Questions this usually raises
- How accurate can a pre-user forecast be?
- Good enough to choose a size class and a burst plan. Not good enough to claim three significant figures. Re-forecast after 60 live days.
- Should we forecast tokens or GPU-hours?
- Whichever you pay for. On-prem, forecast concurrency and powered hours. On a meter, forecast tokens or GPU-seconds and put a ceiling in the sanction.
- What if the circular making use compulsory arrives mid-year?
- Write that as a dated scenario. Do not put it in the base case until the circular exists.
- Can IndiaAI absorb the uncertainty?
- For burst training and early evals, often yes if you are eligible. For standing citizen-facing inference, only if reservation and residency fit. Uncertainty is not a reason to skip those tests.
- How does Prcept size first racks?
- From your MIS peaks and a measured sample, toward the smallest card that holds concurrency in the adoption band, with a written trigger for the next card.
- Is a hockey-stick user curve acceptable in a DPR?
- Not as the base case. Public services acquire gazetted mandates, not viral loops. Model explosion only as a sensitivity with a date.