All insights

Air-Gapped & On-Prem

Cooling and Power for a Departmental GPU Rack

· 10 minute read

A GPU rack is a room heater with a purchase order. Use order-of-magnitude kilowatts, insist on a site survey, and do not let a brochure pick your breaker.

Every bad GPU purchase we have sat in had a sentence like 'power will be arranged'. Arranged by whom, from which panel, into which row, with what leftover cooling, at what outlet type, with what UPS autonomy, was never written. The crate then taught everyone physics.

This guide gives order-of-magnitude thinking you can take to a survey. It will not give you a fake exact kilowatt for a named SKU, and it will not give you a rupee figure. Those numbers change by generation, by PSU efficiency, by utilisation and by how optimistic the datasheet was. Anyone who quotes them as constants is selling.

ASHRAE's datacom guidance exists because heat, humidity and air volume are engineering. Indian halls and campus server rooms vary wildly. The only honest last step is a walk-through with a meter and a floor plan.

Order of magnitude, not a datasheet

Think in layers. The card has a thermal design power in the hundreds of watts for modern data-centre GPUs — sometimes toward the upper hundreds. The server adds CPUs, memory, NICs, disks and fans. Power-supply inefficiency adds another slice. A four-GPU inference node often behaves like a multi-kilowatt space heater. An eight-GPU node can challenge older racks that were planned around 4–8 kW total.

Utilisation matters. A RAG agent that bursts and then idles is not a training run that sits at high power for days. Still size the breaker and the cooling for the high point plus margin. Electrical systems are not impressed that your average was lower.

These are planning hints for a conversation with an engineer. They are not sanctionable figures.
ObjectOrder-of-magnitude thinkingWhat you must still do
One modern DC GPUHundreds of watts at the card, SKU-specificRead that SKU's board spec
1–2 GPU workstationOften survivable in a well-designed roomCheck noise, dust, AC, circuit
4 GPU 2U/4U serverSeveral kW class is a fair planning hintSurvey row kW and cooling
8 GPU nodeCan exceed older rack budgetsMay need a dedicated row or cage
HeatRoughly 1 kW electrical ≈ 3,400 BTU/hHVAC must remove it continuously
UPSMinutes, not hours, unless you bought hoursWrite the drain-and-shutdown plan

What the survey must answer

  1. Which panel and breaker, and what else already hangs there?
  2. What is the spare capacity at the UPS and at the row PDU, not only at the building incomer?
  3. What is the cooling left in that aisle at the worst recent season — not in January in Delhi?
  4. Floor loading and ramp: can the crate actually enter?
  5. Fire suppression compatibility. Some rooms were designed around a different thermal load.
  6. Noise and people: a departmental 'server room' that is also an office will lose the argument with humans.
  7. Redundancy: if one CRAC or one feeder dies, do you shed the GPU or the treasury VM?

Write unknown when you do not know. Unknown is a reason to measure. Invented precision is a reason to overheat.

India-specific annoyances

Summer ambient in many cities leaves little margin for comfort ACs asked to do datacom work. Dust and construction next door will coat filters. Diesel rotation during feeder work will look different if the GPU is the largest load on the UPS. Coastal humidity is a condensation risk if someone sets the AC like a hotel room. None of this is exotic. All of it is why a Bengaluru hall survey is not a copy-paste for a district building.

Design choices that change the heat

  • Inference versus training: do not buy a training box for a RAG desk unless you have a written fine-tune case.
  • Fewer, better-utilised GPUs beat a half-empty dense node you cannot cool.
  • CPU offload and smaller models sometimes remove a GPU entirely. That is a valid cooling strategy.
  • Schedule batch embedding at night if the hall's cooling is happier then — only if the electrical side still has margin.
  • Do not forget the admit host, the registry and the vector store. They are smaller, but they sit in the same story.

Objections you will hear — and what to do with them

The vendor quoted 2 kW typical.

Ask for idle, typical and peak, at the wall, for the whole node, with the PSU type. Then add margin. Typical is a marketing ambient. Breakers see peaks.

We will use existing UPS headroom.

Show the last load reading, not a nameplate from 2016. Headroom that exists on a spreadsheet may already have been promised to another project.

Liquid cooling will save us.

Liquid cooling moves the problem to a CDU and a facility that can take it. It is not a sticker you put on a 2014 CRAC. If the hall cannot do liquid, do not write liquid.

A survey-first playbook

  1. Before any PO: shortlist SKU classes and their board-level power ranges from datasheets, as ranges.
  2. Walk the candidate rooms with electrical and HVAC. Photograph panels. Note other loads.
  3. Pick a home: existing row, new row, or another hall. Write the spare kW as a range or as unknown.
  4. Only then issue the indent. Include PDU type, inlet spec, and a commissioning test that measures wall power at a synthetic high load.
  5. File the survey next to the sanction. A later outage will start with that page.

How this shows up in the file

The concurrence should say 'subject to site survey' until the survey exists, and then it should quote the survey's spare capacity in cautious language. Never let a GPU line item travel alone. Power, cooling and a shutdown plan travel with it, or the crate will teach the committee.

Commissioning tests worth writing

A node that has not been loaded is a rumour about kilowatts. During commissioning, run a synthetic high-load job the vendor provides — not a training run on personal data — and measure wall power, inlet temperature and neighbour-rack temperature. Write the numbers as a range observed on that day, with ambient noted. Then compare to the survey's spare capacity.

Also test shutdown: pull the UPS as far as the hall allows in a planned window, or simulate it. Does the agent drain connections and stop cleanly, or does it corrupt an index? Power is not only about staying up. It is about going down without creating a second incident.

  • Photograph the PDU mapping after install. Paper labels fall off.
  • Record idle, moderate and high load at the wall if you can do so safely.
  • Confirm the CRAC set-point was not quietly lowered to hide a margin problem.
  • Put a temperature alert into the same SIEM as the agent, not only into the hall BMS that nobody in the department sees.

If commissioning exceeds the survey's cautious range, stop and re-home the node. Do not 'see how summer goes'. Summer will answer.

Electrical and HVAC decisions need licensed site professionals. Treat every figure here as orientation. This article is a field guide, not legal, procurement, electrical or engineering advice. Confirm numbers, duties and designs against the current Gazette, CERT-In directions, your SDC / NIC / campus standards, a site survey and your counsel before you file them.

How to prove this on a rack, not on a slide

“Cooling and Power for a Departmental GPU Rack” only matters if a CISO can fail it. A P1 CIO/CTO should be able to point at a cable, a registry, a licence file, a PDU reading or a SIEM index and say: this is the control. If the only evidence is a brochure that mentions “GPU rack power cooling India”, you do not have the control.

A GPU rack is a room heater with a purchase order. Use order-of-magnitude kilowatts, insist on a site survey, and do not let a brochure pick your breaker. Air-gap and on-prem programmes die in the second month, when the first update, the first crash, or the first GPU lead-time slip arrives. Budget the boring path — media, offline licence, local registry, local traces — in the same note as the model name.

On-prem is not air-gapped. An India region is not either. Write the forbidden path (outbound HTTPS, licence phone-home, crash reporter, hidden model API) as a numbered list and test it with the internet off. Whatever still dies was a dependency you did not draw.

  1. Draw the data path for one user-visible answer under “GPU rack power cooling India”.
  2. Disable outbound internet on staging and run the demo script.
  3. List every remaining hop: update, licence, registry, NTP, DNS, SIEM.
  4. Give each hop an owner inside the department, not only the SI.
  5. Minute the restore or the media-transfer once before go-live.

Close this loop before the next CAB

Put “Cooling and Power for a Departmental GPU Rack” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P1 CIO/CTO, not “the vendor.”

Revisit the item when the model, the GeM term, the region, or the SI changes. “GPU rack power cooling India” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.

Questions this usually raises

Can you tell us the exact kilowatts for an H100 or similar node?
No, not as a number you should put in a sanction without the datasheet and a survey. As an order of magnitude, a modern data-centre GPU is often hundreds of watts at the card; a 4–8 GPU server plus CPU, fans and PSU loss commonly lands in the several-kilowatt range. Read the specific SKU. Measure the row.
Is a normal departmental split AC enough?
Sometimes for a single low-power workstation GPU in a well-ventilated room. Often not for a rack server. Heat is continuous. A comfort AC sized for people will short-cycle or freeze, and humidity control is usually missing. Get an electrical and HVAC opinion on site.
Do we need liquid cooling for a departmental inference box?
Usually not for a small inference footprint. Liquid appears when density or the hall's air margin forces it. Do not buy liquid because a keynote did. Do not refuse a hall that only has air if the survey says air is enough.
Who signs the survey?
The hall or campus electrical engineer, the HVAC owner, the fire officer if the rules say so, and the application owner. A vendor slide is not a survey.

Sources