Air-Gapped & On-Prem
On-Prem AI on Existing State Cloud: Feasible?
· 10 minute read
A state cloud is on-prem to the state, not automatically on-prem to your duty. Feasible if you can isolate, place GPUs, and kill phone-homes. Otherwise it is a shared VM with a slogan.
Every state now has something it calls a cloud: an SDC IaaS, a NIC-hosted estate, a MeghRaj tenant, or a private OpenStack that a 2018 SI left behind. Application owners hear 'put the AI there' because that is where the other VMs went. Sometimes that is the right sentence. Sometimes it is how a GPU ends up on a noisy neighbour VLAN with a default route to the internet.
Feasibility is not a yes or no to 'state cloud'. It is a yes or no to four properties: placement (will they take the box), isolation (is your data path yours), identity (do officers log in as themselves), and egress (can you default-deny). If three of four fail, you are not deploying on-prem AI. You are deploying a chatbot on a shared hypervisor.
This explainer is how a CIO asks those four without starting a fight with the SDC director. The director is often your best ally. They already know which catalogue lines are theatre.
What a state cloud usually is
It is typically IaaS or a small set of PaaS services — VMs, block storage, a load balancer, maybe Kubernetes — operated by the SDC, NIC, or an MSP under a state contract. It is rarely a GPU supermarket. It is rarely an air-gapped product. It is a way to get a VM without buying a blade for every website.
| Property | Often true today | What AI needs |
|---|---|---|
| Tenancy | Shared hypervisor, shared storage fabric | Clear isolation for the data class |
| Images | Golden images plus internet repos | Private registry, pinned digests |
| GPU | Absent or a pilot node | Named placement, power-checked |
| Network | Default route, central proxy | Default deny, documented exceptions |
| Identity | AD / LDAP already there | Reuse it; no local vendor users |
| Ops | VM tickets, business hours | A named path for the bag and the GPU |
| Backup | VM-centric | Application-aware for indexes |
Three feasible patterns
Pattern A — Dedicated host in the state cloud hall
The hall hosts your GPU server as a dedicated host or a dedicated rack, on a VLAN they agree to isolate. You reuse physical security, UPS and visitors' protocol. You do not reuse a shared Kubernetes that other departments can namespace into. This is usually the most honest 'on-prem on state cloud'.
Pattern B — Shared IaaS for non-sensitive planes
Build room, static websites, public-circular RAG, and CI sit on the shared IaaS. The run room for personal data sits on a dedicated island. The bag moves between them. This is hybrid inside the state, and it is allowed if you draw it.
Pattern C — Thin tenants, thick hall service
The SDC offers a managed inference service with your index in your project. Feasible only if the SDC will sign processor-like duties, refuse training, and export logs to you. Most halls are not ready to be a model vendor. Do not force them.
How to ask without a fight
- Start with what you will reuse: identity, backup, visitor protocol, SOC. Praise the hall for those.
- Ask for a closed project: no default egress, no shared filesystem, no standing MSP id.
- Ask where a GPU would actually sit and whether a dedicated host is a catalogue item.
- Ask who patches the hypervisor and who can snapshot your disks.
- If the answers fail your data class, ask the director to help you pick Pattern A or another hall. Bring them the problem.
When the honest answer is no
No if the only GPU is a shared card on a research VLAN. No if images can only be pulled from the internet. No if the MSP refuses to drop standing access. No if the hall cannot say where logs live. No is not a feud. No is how you avoid a later incident that will be blamed on the SDC anyway.
Objections you will hear — and what to do with them
Policy says all new workloads go on the state cloud.
Then take Pattern A and make the dedicated host a state-cloud offering. Policy wants consolidation of halls, not the destruction of isolation.
If we leave the state cloud we lose support.
If support equals standing admin on a medical-board VM, you did not have support. You had a shared key. Buy support that can live with break-glass.
We will start shared and isolate later.
You will not. Snapshots, golden images and helper jobs will grow roots. Isolate first or write the data class down as public.
A twelve-day feasibility playbook
- Days 1–3: read the state cloud service catalogue and the MSP contract clauses on access.
- Days 4–7: workshop with SDC director — four properties, three patterns.
- Days 8–10: if Pattern A or B is viable, write the project spec (VLAN, egress, GPU, backup). If not, write the no.
- Days 11–12: file the feasibility note. Do not raise a silent VM ticket in parallel.
How this shows up in the file
The feasibility note names the pattern, the four properties, and the residual MSP access. A later auditor should see why this workload is on a dedicated host while the department website is not. That distinction is the entire intelligence of the file.
Questions for the MSP contract, not the portal
Portals show flavours. Contracts show who can snapshot you. Before you call a state-cloud AI deployment feasible, read the MSP or NICSI-style operations clause with the CISO. Standing access, offshore support, and 'we may copy configurations for quality' are hops.
- Is support allowed to copy a disk out of the hall?
- Are operations staff in India, and is that contractual or anecdotal?
- Can you refuse a hypervisor-level snapshot?
- Where are the management logs, and can they land in your SIEM?
- What happens to your project on contract exit — export, or archaeology?
If the contract cannot bend and the data class cannot live with the default, Pattern A (dedicated host) or a no is the answer. Do not try to defeat a contract with a security group. Security groups do not bind MSPs.
State clouds remain a public good for the workloads they were built for. Stretching them silently around a medical board is how good programmes get a bad incident attached to their name.
This article is a field guide, not legal, procurement, electrical or engineering advice. Confirm numbers, duties and designs against the current Gazette, CERT-In directions, your SDC / NIC / campus standards, a site survey and your counsel before you file them.
How to prove this on a rack, not on a slide
“On-Prem AI on Existing State Cloud: Feasible?” only matters if a CISO can fail it. A P1 CIO/CTO should be able to point at a cable, a registry, a licence file, a PDU reading or a SIEM index and say: this is the control. If the only evidence is a brochure that mentions “state cloud AI deployment”, you do not have the control.
A state cloud is on-prem to the state, not automatically on-prem to your duty. Feasible if you can isolate, place GPUs, and kill phone-homes. Otherwise it is a shared VM with a slogan. Air-gap and on-prem programmes die in the second month, when the first update, the first crash, or the first GPU lead-time slip arrives. Budget the boring path — media, offline licence, local registry, local traces — in the same note as the model name.
On-prem is not air-gapped. An India region is not either. Write the forbidden path (outbound HTTPS, licence phone-home, crash reporter, hidden model API) as a numbered list and test it with the internet off. Whatever still dies was a dependency you did not draw.
- Draw the data path for one user-visible answer under “state cloud AI deployment”.
- Disable outbound internet on staging and run the demo script.
- List every remaining hop: update, licence, registry, NTP, DNS, SIEM.
- Give each hop an owner inside the department, not only the SI.
- Minute the restore or the media-transfer once before go-live.
Close this loop before the next CAB
Put “On-Prem AI on Existing State Cloud: Feasible?” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P1 CIO/CTO, not “the vendor.”
Revisit the item when the model, the GeM term, the region, or the SI changes. “state cloud AI deployment” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.
Questions this usually raises
- If the state cloud is in our SDC, is that on-prem AI?
- It is on-prem to the state. Whether it meets your air-gap, tenancy and admin-path duties depends on the offering. Ask who can see the hypervisor, where images come from, and whether your VLAN is default-deny outbound.
- Can we just raise a VM ticket for a GPU?
- Sometimes, if the catalogue has one and the hall has power. Often the ticket will return 'no SKU'. That is a planning input, not a personal insult. Forecast or bring a node the hall will host as a dedicated host.
- Is state cloud cheaper than our own rack?
- It can be, because you reuse shift staff, UPS and identity. It can be more expensive if you force the hall into a one-off GPU cage and still hire an SI. Compare the whole duty map, not the VM rate card — and do not invent a GeM GMV or a fake unit price here.
- Does using the state cloud satisfy DPDP by itself?
- No. You still have purpose, processor terms with whoever operates the cloud, and a map of hops. Shared operations staff are people. Write them into the story.