Compute & Cost
How Many GPUs Does a Department Need?
· 9 minute read
The number of GPUs is not a status signal. It is concurrency × tokens × latency, plus a spare, plus a place to eval. Most departments need a boring pair and a queue. A few need a small cluster. Almost none need the keynote.
A directorate asked for eight flagship GPUs because a neighbouring State had eight on a tour. Their workload was a helpdesk that peaked at forty concurrent chats and a nightly batch of PDFs. Two cards, a spare, and a queue would have been an engineering answer. Eight cards would have been a heating answer. They bought eight. They now have a room that is warm and a utilisation graph that is not.
This guide is the counting method. It sits next to the existing sizing piece and the tokenisation teardown. It will not name a SKU as the government standard. SKUs move. The method does not.
IndiaAI hours can cover bursts you should not buy silicon for. They cannot replace the count for a steady, on-prem, citizen-facing SLO. Do not mix the two without writing the split.
Not a vendor quote. Measure, then buy.
The inputs that actually set the count
Concurrency. How many in-flight generations at the p95 minute of the p95 day. Not registered users. Not 'the whole State'. Chats, IVRS slots, and officer copilots are different queues. Do not add them as if they shared a context.
Tokens. Prompt plus retrieval plus generation, at p95, in the production language mix. Indic token tax is real until measured — see that teardown. English demo tokens will under-count.
Latency. A kiosk that must speak in two seconds is a different machine from a nightly summary. Tight SLOs need headroom. Batch can wait and pack.
Model footprint. Weights plus KV cache at your context. A 70B-class chat at long context is a different animal from a 7B classifier. Do not size from the name of the family. Size from the hash you will run.
Eval and staging. If you only have production cards, you will eval on production. That is how you drop the SLO to run a language pack. Buy a place to eval, even if it is off-peak on the same pair with a written freeze.
A worked count you can imitate, not copy
Suppose a service forecasts 30 concurrent chats at p95, 2,000 tokens in and 400 out after retrieval, and a 7B-to-small-30B-class model that your bench says does 40 of those completions per card per minute at that context. You need less than one card for the mean and perhaps two for the p95 plus hiccups. A pair plus a cold spare is a department-shaped answer. We used rounded, illustrative throughputs. Your bench replaces them.
Suppose instead you want a large long-context model for legal research at 8 concurrent officers and 32k context. The KV cache may force multiple cards per replica. Now you are in small-cluster territory. The jump is the model and the context, not the dignity of the department.
Suppose you process 50,000 pages a night. That is a batch window. Pack the GPU until morning. Do not buy the night's work as if it were noon concurrency.
| Question | Where you get it | What a fake answer looks like |
|---|---|---|
| p95 in-flight generations | Existing IVRS/chat logs | 'One per citizen in the State' |
| p95 tokens in/out | Tokeniser on your corpus | English demo traces |
| Completions per card per minute | Bench on the intended hash | A keynote tokens/s number |
| Batch pages per night | Last month's scans | A vendor 'unlimited documents' |
| Eval hours per week | Language pack plan | Zero, we will think later |
Spares, queues, and the card you should not buy
One card is not a service. Failure is a day. Two cards plus a queue is a service. Three if you cannot tolerate a failure during peak and you have no cloud burst that is allowed.
A queue is cheaper than a card. Citizens will wait five seconds. They will not wait for a procurement that bought hope instead of a queue.
Do not buy training-shaped clusters for inference-shaped work because a tour showed you NVLink fabric. If you will not train, do not heat the room for training.
Do not buy the next SKU because the current one is embarrassing in a brochure. Embarrassing and sufficient is a professional choice.
When the count is 'fewer, plus a burst door'
If eval and occasional fine-tunes are the only reason you are about to double the cluster, look at live IndiaAI or GeM hours for that burst, if the data may go. Keep inference local. Write the door. Close it when the burst ends.
If the data may not go, buy the eval time as off-peak on the pair, or buy one extra card that sleeps. Sleeping is cheaper than an unused eight-pack.
Two counts
Objections you will hear — and what to do with them
These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.
We might train a foundation model later.
Then write a later paper. Do not heat a room for a later paper. Foundation training is not a departmental surprise.
The vendor says the model needs eight-way tensor parallel.
Some hashes do. Many that a department should run do not. Bench the hash you need, not the hash that flatters the vendor's cluster.
GPUs will be scarce, buy now.
Scarcity is a lead-time input, not a licence to oversize by 4×. Buy the pair and a written option, or a GeM rate contract, rather than a museum.
How many do other secretariats have?
Irrelevant unless they published their concurrency and tokens. Ask for their method, not their invoice.
Three weeks to a number you can defend
If you cannot fill the starred cells, you are not ready to raise a purchase proposal.
- Week 1: pull logs. Forecast p95 concurrency. Run the tokeniser pack.
- Week 2: bench two candidate hashes on one rented or borrowed card if you do not own one. Record completions per minute at your context.
- Week 3: compute online cards, batch window, spare, eval. Write the burst door. Circulate the sheet. Only then pick a SKU from a live quote.
File note you can paste
Subject: GPU count for [service] — method and recommendation.
The attached sheet uses measured concurrency, tokens and a bench of hash [insert]. Recommended online cards: [n], spare: [n], batch policy: [queue]. Burst / IndiaAI: [yes with letter / no]. A photograph of another State's rack is not part of this method. This is not a commercial offer.
What we will count
Prcept AI will bench on your mix and recommend a boring number. We will not upsell an eight-pack to match a tour. If you need one card and a queue, we will say one card and a queue — and then remind you that one card is not a service, so it is two.
Sovereignty is not measured in board count. It is measured in whether the queue still moves when a card dies.
This article is informational field guidance for Indian public institutions, not legal, language-policy, procurement, finance or engineering advice. Confirm against the live Gazette, Official Languages Act and Rules, your State's official-language law, MeitY / IndiaAI notices, GFR, GeM terms, DPDP text, departmental manuals and your counsel before you file it.
How to put this in the finance note
A P1 CIO/CTO searching “GPU requirements government AI” needs a number a CFO can defend, not a GPU brand. “How Many GPUs Does a Department Need?” belongs in a cost model with people, power, idle time, AMC and the cost of a failed pilot.
The number of GPUs is not a status signal. It is concurrency × tokens × latency, plus a spare, plus a place to eval. Most departments need a boring pair and a queue. A few need a small cluster. Almost none need the keynote. IndiaAI subsidy, if you use it, is a live notice — not a permanent discount. On-prem TCO includes ops headcount. Do not invent Rs/hour. Cite the source of every rupee.
- Separate capex, opex, and one-time cleanup.
- Show utilisation, not just peak GPUs.
- Price the human fallback, not only inference.
- Date every tariff and subsidy assumption.
Close this loop before the next CAB
Put “How Many GPUs Does a Department Need?” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P1 CIO/CTO, not “the vendor.”
Revisit the item when the model, the GeM term, the region, or the SI changes. “GPU requirements government AI” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.
Questions this usually raises
- Is there a MeitY-mandated GPU count per department?
- Not that we will invent. Size from workload. Ignore folklore. Put that in the file next to “GPU requirements government AI” so a stranger can reconstruct it. A one-line yes/no under “How Many GPUs Does a Department Need?” is not an answer a secretary can defend. Confirm against the live Gazette, circular or GeM term; this is not legal advice.
- Can CPU inference replace a GPU for us?
- For some small classifiers and some offline jobs, yes. For interactive generation at department scale, bench it. Do not assume either way.
- Should training and inference share cards?
- Only with a written freeze. Training jobs will eat the SLO. Prefer off-peak or a burst door.
- How many GPUs for a university versus a district?
- The method is the same. The inputs differ. A university that trains may need a small cluster. A district FAQ may need a pair. Names of institutions do not set counts.
- Do we need the same count in DR?
- You need a written RPO/RTO. That may be a cold spare in another hall, not a hot clone of an oversized cluster.
- Where do we rent a card to bench?
- Live IndiaAI or GeM lists, or a lab that will let you bring the hash. Date-stamp the rate. Do not steal a neighbour's node without a letter.