All insights

Air-Gapped & On-Prem

Migrating a Cloud Pilot to On-Prem Production

· 11 minute read

The pilot was easy because the cloud hid identity, logs and model hosting. Production on a government rack will surface every one of those. Migrate the workflow, not the tenant.

The most common Indian AI story in 2026 is not a greenfield air gap. It is a cloud pilot that worked in front of a secretary, followed by a sentence in the minutes: now put it on the SDC. The pilot used a hosted model, a hosted vector store, the vendor's identity, and a Slack channel for support. None of that exists on the government rack. The migration is not a lift of a resource group. It is a rebuild of every assumption the pilot was allowed to skip.

Departments lose six months here because they try to clone the tenant. They export a snapshot, argue about who pays egress, discover the embeddings were in a global index, and then discover the eval set still holds last month's scholarship names in the vendor's cloud. The honest object of migration is the workflow: which officer, which system of record, which tools, which model class, which approval. The tenant is a sketch. Treat it as a sketch.

This playbook is what we run with programme owners who have a working hosted pilot and a political deadline to be on-prem before the next assembly session. It assumes you already decided that production personal data should not stay with the hosted vendor. If you have not decided that, stop and write the decision. Migration without a decision is just a second bill.

What you are actually moving

You are moving a purpose, a role map, a tool list and a quality bar. You are not moving a virtual machine. If the pilot's only artefact is a URL and a login, you do not have something to migrate. You have a demonstration to rebuild.

List the objects before you book the SDC cage. Prompts and system instructions. Tool contracts and the systems they call. Retrieval corpora and whether they contain personal data. Eval questions and gold answers. Identity groups. Audit events you promised the DPO. Model family and context length. Latency the officer will accept. None of those objects are the cloud account.

Then list the objects you must not move. Live personal data that landed in the vendor's default logs. Embeddings of citizen files that were never supposed to leave. Support recordings. The vendor's improvement set. If those exist, your first migration task is a deletion and a certificate, not a copy. Copying a dirty pilot into a clean rack launders the dirt.

Migrate instructions and tests. Rebuild data, identity, logs and inference.
Pilot objectBring to on-prem?How
System prompts and tool specsYesExport as versioned files the department owns
Synthetic / redacted eval setYesRe-key and store in the department Git or artefact repo
Live tickets, names, Aadhaar, marksNoLeave them. Rebuild retrieval from the system of record inside the perimeter
Vendor-hosted embeddingsNoRebuild the index on-prem from authorised corpora
Identity (vendor IdP)ReplaceNIC / department AD or the campus IdP. No leftover vendor SSO.
ObservabilityReplaceSDC SIEM or campus syslog. CERT-In 180-day copy in India.
Model endpointReplaceLocal runtime on GPU or CPU, or a government-hosted model you control
Support channelReplaceTicketed jump host. No Slack with production snippets.

The cutover that does not lie

Run both paths on the same eval set before any officer sees the on-prem agent. Score groundedness, refusal, latency and tool-error handling. If on-prem is worse, say so in the note and either fix the model class or reduce the workflow. A silent quality drop is how collectors lose trust and go back to WhatsApp.

Do not dual-write live personal data to the cloud tenant just in case. Dual-write is how you fail the reason you left. If you need a rollback, rollback to the pre-pilot process: the existing MIS, the existing desk. That is an operational rollback, not a cloud rollback.

Freeze the hosted tenant the day production data is pointed at the rack. Change the password, revoke tokens, export the configuration, then request deletion of personal stores and ask for a dated certificate. If the vendor says deletion is on the next cycle, write that cycle into the file and keep chasing. A forgotten pilot tenant is a second processing activity you are still paying for.

What breaks first on the rack

Identity. The pilot used the vendor's users. Production needs NIC or campus accounts, groups that match actual desks, and a joiner-mover-leaver process. Budget a week for this even if the SI says it is just SAML. Retrieval quality. Cloud pilots often used a hosted embedding API and a large context window. On-prem you may have a smaller model and a local embedder. Rebuild chunking. Do not assume the same PDFs will behave. Tool reachability. The cloud agent called public APIs. The on-prem agent must call the district MIS on an internal name, with a service account, with timeouts that match a 2014 application. Write those timeouts down. People. The cloud vendor watched the dashboard at 11pm. On-prem, someone in the SDC or the NIC room has to. Name them before cutover. An unmanned rack is not production.

  1. Decide the production perimeter and the data classes that may never re-enter a hosted tenant.
  2. Export only prompts, tool specs and a redacted eval set.
  3. Stand identity, logs, registry and model runtime on the rack.
  4. Rebuild retrieval from the system of record. Do not import vendor embeddings of live files.
  5. Score on-prem against the eval set. Publish the delta in the note.
  6. Cut officers over. Freeze the hosted tenant. Demand deletion certificates.
  7. Run a 14-day hypercare on the rack with a named bridge, not a vendor Slack.

Objections you will hear — and what to do with them

These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.

We will lose the model's memory of our last three months.

If that memory is unique personal data in a hosted index, losing it is the point. If it is style and workflow, you already have it in prompts and evals. Rebuild the index from systems you own.

The SDC cannot match cloud latency.

Then shrink the workflow or buy the hardware the workflow needs. Do not keep a hosted path for the slow queries. That split becomes the real architecture, and the slow queries are usually the ones with the thickest files.

The vendor will migrate us as a professional-services SKU.

They can rebuild with you. They cannot migrate the tenant without moving stores you should be deleting. Buy rebuild hours against a checklist you wrote. Do not buy a black-box migration.

A ninety-day rebuild, not a lift-and-shift

Treat the hosted pilot as a laboratory notebook. The laboratory stays behind. The notebook comes with you.

  1. Days 1–10: write the production purpose, the data classes that stay inside, and the deletion request to the hosted vendor.
  2. Days 11–30: rack identity, logs, registry, model runtime. Prove a no-outbound or named-egress staging path.
  3. Days 31–55: rebuild retrieval and tools against internal systems. Port prompts. Build the eval harness from redacted items only.
  4. Days 56–75: score, fix, and tabletop a rollback to the pre-pilot desk process.
  5. Days 76–90: cut over, freeze the tenant, file the deletion chase, name the hypercare roster.

How this shows up in the file

File the decision to leave the hosted tenant, the list of objects brought across, the list of objects deleted, the eval delta, and the name of the on-prem owner for nights and Sundays. A later audit should not have to reconstruct this from Slack.

If the hosted tenant still has a live password after cutover, the migration is not finished. Say that in the weekly note until the certificate arrives.

What to do with the eval set you already love

The only artefact from a hosted pilot that is almost always worth keeping is the list of questions officers actually asked and the answers they marked as good or dangerous. That list is institutional knowledge. It is also often full of names. Redact before it crosses into the SDC. Replace unique identifiers with tokens you can map in a locked file the DPO holds. If you cannot redact, transcribe the pattern and throw away the original ticket.

Do not import the vendor's automatic eval tenant. That tenant is their product-improvement store wearing your department's history. Ask what it holds. Ask for deletion. Do not copy it so we do not lose quality. Quality is the questions, not their vectors of last year's children.

Identity cutover without a weekend outage

Create the on-prem groups two weeks before cutover. Add the same officers who already use the hosted URL. Let them log into staging with the new identity while the hosted path still runs on synthetic data. On cutover day you are changing a bookmark and a DNS name, not inventing accounts at 6pm. Leave the hosted tenant frozen, not dual-live.

Joiner-mover-leaver must be written. A deputation that ends in March should end the agent's rights in March. Hosted pilots skip this because the vendor's CSM adds emails by hand. Production that still works that way will have a former intern in the privileged group.

Rollback is the old desk, not the old tenant

Write a rollback that a collector can understand: if the on-prem agent is dark, officers use the existing MIS and the existing phone roster, and a public notice says so. Do not write a rollback that re-opens a hosted tenant you just asked to delete. That rollback is a transfer with a panic button.

Tabletop the rollback once. Time how long it takes to post the notice and to brief the desk. If it takes a day because nobody owns the notice, you do not have a rollback. You have a hope.

This article is informational field guidance for Indian public institutions, not legal, procurement, security-accreditation or engineering advice. Confirm against the current Gazette, CERT-In direction, GFR, GeM term, SDC policy and your counsel before you file it.

Questions this usually raises

Can we keep the cloud pilot as a fallback after on-prem go-live?
Only if it holds no production personal data and the file says so. A fallback that still sees live tickets is a second production, not a safety net.
Who should own the migration inside the department?
A programme owner who can command identity, the system of record and the SDC cage — not only the AI enthusiast who ran the pilot. The pilot owner can keep quality. They cannot usually command AD and the firewall.
What if the vendor says the prompts are their IP?
Then you paid for a demonstration, not a transferable system. Negotiate export of the instructions you jointly wrote. If they refuse, rewrite. Do not steal. Do not stay hostage either.
Does moving to on-prem finish the DPDP analysis?
No. You still need roles, basis, purpose, processor terms and erasure design. Location is one control. It is not the Act.

Sources