All insights

Compute & Cost

On-Prem vs Subsidised Cloud: 3-Year TCO

· 10 minute read

A three-year TCO that only compares GPU stickers is a children's drawing. Add people, power, SI, AMC and air-gap operations. Use live IndiaAI or GeM rates. We will give you honest bands, not a fabricated national study.

Two numbers arrived in the same file. The cloud column was a subsidised hourly rate multiplied by a hopeful utilisation. The on-prem column was a GPU quote from a distributor, divided by three years, with no power and no people. Cloud won. Then the department hired an SI, posted two operators, built a small air-gap for the land module, and paid a diesel bill the model never saw. Cloud had not won. Arithmetic had been refused.

This is a costing study in the honest sense: a method, worked ranges, and a refusal to invent a national average. We have not run your meters. We have seen enough files to know which lines dominate and which lines get deleted to make a slide.

IndiaAI hours, if you are eligible and if the data class may leave the SDC, belong in the cloud column at the live rate after subsidy — the rate on the printout dated today, not a percentage from memory. See the starred subsidy guide. This piece will not restate commercials.

Not financial advice. Your finance controller, GFR and any State IT policy still own the heads.

The six lines that must appear

People. Product owner, DPO time, language markers, gate officers, operators, and the desk that takes fallback. In almost every honest file we see, people are not a rounding error. Over three years they often rival or beat the GPU invoice. If your TCO has no salaries, stop.

GPU. Purchase, lease, or hourly. On-prem is capex plus spares. Cloud is hours plus the risk that the rate card or the subsidy letter moves in year two. Do not assume year-three hours cost year-one hours.

Power. Cards, cooling, UPS, and the generator culture of your SDC. Idle cards still draw. Cloud hides this in the hour. On-prem puts it on the electricity file. Ignoring it is how on-prem looks falsely expensive or falsely cheap, depending on who is lying.

SI. Integration to MIS, identity, PDF, logging, and the first year of 'please make it work with our 2014 stack'. Rarely small. Often forgotten in the cloud column because 'it is just an API'.

AMC and refresh. Support, spare, software subscriptions, and the mid-life GPU that will not be the SKU you bought. Cloud's AMC is the vendor's margin inside the hour — until you need a dedicated term that is not on the portal.

Air-gap operations. If any workload cannot go, you are not choosing cloud or on-prem. You are choosing cloud-plus-remainder or on-prem-for-all. The remainder has people, power and AMC too. A TCO that deletes the remainder is a different project.

Honest bands, labelled as bands

These are planning bands from files, not a sample survey and not your quote. Adjust or discard.

A small departmental inference pair — two modern datacentre GPUs, not a research cluster — is often a mid-to-high seven-figure rupee capital line before GST, before SI, before power. Treat that as an order of magnitude to check against today's distributor quote, not as a number to paste.

Power plus cooling for that pair, in a typical SDC with Indian tariffs and mediocre PUE, is often a low-to-mid six-figure rupee annual line if the cards spend much of the day awake. If they sleep, the line shrinks but does not die — see the idle-cluster piece.

SI for a first citizen agent wired into one MIS is often comparable to the first GPU invoice, sometimes larger, especially when PDFs, Indic fonts and identity are in scope. A 'just a chatbot' SI is cheaper and usually incomplete.

People for operate — even a thin roster — over three years regularly exceed the hardware. A TCO that shows hardware as 80 percent of cost is usually missing the roster.

IndiaAI-subsidised hours can undercut on-prem GPU stickers on eligible, delay-tolerant, non-sensitive training or eval bursts. They rarely undercut a full operate roster. They do not delete the air-gap remainder.

A TCO skeleton. Fill rupees from quotes and live cards. The skeleton is the deliverable.
LineOn-prem (3 yr)Subsidised cloud (3 yr)Hybrid remainder
People (named roster)AlwaysAlwaysAlways if remainder exists
GPU / hoursQuote + sparesLive card × hours × yearsQuote for the remainder
Power / cooling / UPS shareMeter or SDC tariffInside hour (ask)Remainder only
SI / integrationUsually year 1–2 heavyUsually still existsOften duplicated
AMC / support / refreshContractedTerms + lock-in riskOn the remainder
Air-gap ops / transfer mediaIf you are air-gappedUsually N/AThe hidden cloud cost
Wait / eligibility riskLead time to buyCommittee + RFE movementBoth

How to run the three years without fiction

Year one is SI and mistakes. Year two is operate. Year three is refresh pressure and contract fatigue. A model that spreads SI evenly across thirty-six months is smoothing to win.

Utilisation is a claim. If you cannot show a queue, do not cost cloud at 80 percent utilisation and on-prem at 20 percent to make a point. Use the same workload forecast in both columns, then apply the idle-cluster math to on-prem and the reserved-versus-spot story to cloud — if the programme even offers one.

Exit. On-prem exit is a used GPU and a data store you still hold. Cloud exit is a trace story, a fine-tune story, and a re-procurement. Put one line in each column for exit. Committees forget. Auditors do not.

A decision rule that is not a slogan

If more than one critical workflow cannot leave the perimeter, cost hybrid or cost on-prem-for-all. Do not cost pure cloud.

If the only honest utilisation is bursty eval and training, and the data may go, live subsidised hours can win the GPU line. Still add people.

If citizen inference is steady all day, on-prem utilisation can look rational even before sovereignty. Run the numbers. Do not assume.

If IndiaAI eligibility is not yet granted, the cloud column is a hope. Hope is not a third-year actual.

Two TCOs

Objections you will hear — and what to do with them

These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.

Give us a single rupee outcome for a typical department.

There is no typical department. Workload, tariff, roster and data class move the result. We will not mint a fake study.

Cloud has no power cost.

It has a power cost inside the hour. Ask. On-prem has a power cost on your bill. Both exist.

People exist whether we buy GPUs or not.

Some do. Some are incremental: operators, markers, an SI who would not be on the file. Incremental is the TCO test. Write the names.

Three years is too long; our budget is annual.

Then show year one clearly and still compute three. Annual blindness is how refresh and AMC become scandals.

Four weeks to a TCO a controller will initial

Start from named people, not from a GPU SKU.

  1. Week 1: name the roster and the SI scope. Split data classes. Decide if hybrid is mandatory.
  2. Week 2: collect today's GPU quote, SDC tariff, and the live IndiaAI or GeM printout. No memory rates.
  3. Week 3: fill the six lines for on-prem, cloud, and hybrid. Use the same workload forecast.
  4. Week 4: sensitivity — subsidy withdrawn in year two, utilisation halved, SI +30 percent. Present the bands, not a single hero number.

File note you can paste

Subject: Three-year TCO — on-prem, subsidised cloud, hybrid — [project].

Six lines are costed: people, GPU/hours, power, SI, AMC/refresh, air-gap ops. Unit GPU rates and any subsidy are taken from annexures dated [today], not restated. Bands, not a single figure, are presented. This is not a market study and not financial advice.

What we will put in the columns

Prcept AI will not delete people to make on-prem win. We will not delete the air-gap remainder to make cloud win. We will sit with your controller and fill the skeleton from your quotes.

If hybrid loses on money and wins on law, say both sentences. A TCO that hides the law is as bad as a legal note that hides the money.

What the next file must contain

“On-Prem vs Subsidised Cloud: 3-Year TCO” earns a line in the noting only if a P1 CIO/CTO can attach proof of “on-prem AI TCO India.” A heading is not proof. A vendor slide is not proof. A workshop photograph is not proof.

Write three dated sentences: what was decided, who owns it after the next posting order, and when it will be re-checked. If you cannot write the three sentences, you are not ready to buy, to sell, or to go live.

Leave unsourced percentages out of the note. DPDP is not a blanket localisation statute. The November 2025 AI governance text is guidance, not an Act. CERT-In’s 28 April 2022 directions still set specified incident and log clocks. A PAC, when lawful, lives in GFR Rule 166.

  • Name the designation that owns “on-prem AI TCO India.”
  • Attach one artefact a stranger can open next year.
  • Record the instrument you are actually using.
  • Revisit when the model, the SI, the notice or the posting changes.

This article is informational field guidance for Indian public institutions, not legal, language-policy, procurement, finance or engineering advice. Confirm against the live Gazette, Official Languages Act and Rules, your State's official-language law, MeitY / IndiaAI notices, GFR, GeM terms, DPDP text, departmental manuals and your counsel before you file it.

Questions this usually raises

Should we include IndiaAI subsidy in the base case?
Only if the letter exists or the live rule makes you clearly eligible and you can bear the wait. Otherwise put it in a sensitivity, not in the base.
Is diesel part of power?
If your SDC burns it when the grid dies, yes. AI nodes that cannot die at 11 am are diesel customers.
Where does language cost sit?
In people and in SI, and in the yearly refresh. It is not a seventh mystery line unless you want to show it. It is not zero.
Do we discount on-prem because we already own a rack?
You may treat sunk capital as sunk, but still cost power, AMC, operators and refresh. Free GPUs are never free to run.
How often do we refresh the TCO?
When the rate card, the roster, or the data-class split changes, and at least at every RE.
Can we use this as a bid score?
You can require bidders to fill the skeleton. You cannot require them to match a fake national ratio.

Sources