Sovereignty & Data Residency
Telemetry Is the Sovereignty Loophole
· 10 minute read
The model can sit in your rack and still ship the prompt. Telemetry is how sovereignty dies after the architecture slide is approved.
The CIO of a power-sector PSU had already won the argument. Inference would run on two servers in the plant's DMZ-adjacent VLAN. No public model API. The board note said sovereign. Three weeks after go-live, the network team dumped outbound TLS from the new VLAN and found a steady trickle to an American observability host. The bodies were compressed traces. Inside the traces were prompt prefixes, officer usernames and the account numbers the agent had been asked to reconcile.
Nobody had lied in the architecture review. The model was local. The product's default installer had also enabled a hosted APM because that is how the vendor debugs customers. The checkbox was buried under Advanced. The MSA said data remains in India. The MSA did not define data to include span payloads.
This teardown is the list we use when a CIO says the rack is ours, so we are done. You are not done until every packet that leaves the VLAN has an owner, a purpose and a redaction rule. Sovereignty that ignores telemetry is a press note.
The eight pipes that ignore your architecture slide
Vendors do not call these transfers. They call them health, product improvement, crash reporting, customer success. The packet does not care what they are called.
| Pipe | What usually rides in it | What to demand |
|---|---|---|
| APM / tracing agent | Spans with URLs, headers, sometimes full request bodies | Local collector only; payload redaction on; foreign host denied at the proxy |
| Crash / exception reporter | Stack traces plus the request that died, often with PII | Offline crash folder; manual triage; no automatic upload |
| Product analytics | Clicks, session replay, sometimes typed text | Disable or self-host; no session replay on citizen workflows |
| Feature flags / remote config | Environment names, user keys, rule evaluations | Local flags; no user keys that are officer emails |
| Licence / entitlement check | Host IDs, versions, sometimes hardware and user counts | Published payload; offline licence file; no sample data |
| Model or guardrail fallback | The full prompt when the local path fails | Fail closed; no silent hosted fallback |
| Support tunnel / remote assist | Whatever the engineer can see, plus copied logs | Break-glass with your jump box; recorded sessions; Indian staff if the file requires it |
| Container and model pulls | Registry metadata; sometimes build secrets in layers | Internal registry; signed artefacts; no live pull from a public host |
Two of these pipes are cultural, not technical. Product analytics exists because SaaS companies live on funnel metrics. Support tunnels exist because the alternative is teaching your team to debug. Both are rational for a vendor. Neither is rational as a default inside a citizen workflow.
How to tear the installer down before it lands
Do not start with the brochure. Start with a staging VLAN that has no default route except through a logging proxy. Install the product the way the vendor wants it installed. Then read what tried to leave.
- Capture DNS and TLS destinations for seventy-two hours of soak, including a simulated crash and a licence renewal window.
- Open the config tree and search for ingest, sentry, segment, datadog, honeycomb, launchdarkly, amplitude, pendo, intercom, and the vendor's own telemetry hostname.
- Trigger one failure that includes a dummy PAN or Aadhaar-shaped string. See whether that string appears in any outbound body.
- Ask for the data map of every outbound host: fields, retention, country, subprocessors.
- Refuse any host that cannot be switched off without breaking inference.
If the vendor says the product will not run without their cloud collector, you do not have an on-prem product. You have a thin client with a local GPU. Write that sentence in the evaluation note. Committees understand it.
Objections from vendors and from your own ops
We cannot support you if we cannot see crashes. Then you will support through a folder of local crash dumps that we upload after review. You will not have a standing pipe into production prompts. If that costs more professional-services days, price it. Do not smuggle it through telemetry.
The data is aggregated. Aggregation is a process, not a vibe. Show the transform. If the transform happens after the foreign collector receives the raw event, the transfer already occurred.
CERT-In requires us to monitor. Yes. It requires specified logs retained in India and incidents reported fast. It does not require you to give a SaaS vendor a copy of the prompt. Local SIEM is how you satisfy the direction without opening the loophole.
Our own SRE team wants Datadog. Then run Datadog, or any other tool, as an in-country collector with a written transfer decision if anything still leaves. Do not confuse a familiar dashboard with a lawful path.
Feature flags as a remote control
A feature-flag service can change the product under you. A flag that enables a hosted critic, a new OCR vendor, or a session-replay sample is a change of means. If that service lives abroad and reads user keys, it is also a standing transfer of staff identity.
Demand local flags, or a frozen flag file you promote like a model. If the vendor says flags must be live so they can disable a bad feature, give them a ticket path and an officer who can flip a local flag. Do not give them a remote hand on a citizen workflow.
Read the default sample rates. A hundred-percent trace sample with bodies on is not observability. It is a second case file in someone else’s cloud. Turn it to skeletons, locally, before the first officer logs in.
A two-week telemetry hunt, then a 90-day close
- Week 1: staging VLAN, logging proxy, full installer, crash test, licence window. Produce the destination list.
- Week 2: kill or replace every destination that carries content. Write the allow-list into the proxy and into the contract.
- Days 15–45: repeat the hunt after the first patch. Chart updates are how pipes return.
- Days 46–75: point remaining observability at the in-VLAN collector and the SIEM. Prove CERT-In retention locally.
- Days 76–90: tabletop a leak. If a prompt fragment left, treat it as a processor incident and practise the notice path.
What goes in the file
The file is a packet story, not a slogan. Keep the proxy logs from the soak, the destination table, the redaction config, the MSA clause that forbids undeclared collectors, and the date of the last hunt after a patch. If those artefacts are missing, the sovereignty claim is a hope.
Prcept AI is built to run without a vendor collector. If we ever need a support artefact, it should leave as a file you chose to send, not as a daemon you forgot was there.
How to defend this in the file
A P1 CIO/CTO will be asked to explain “Telemetry Is the Sovereignty Loophole” to a secretary who has ten minutes. Do not start with the model. Start with the store, the hop, the clause, or the residual risk. “AI telemetry data leakage” is a search phrase. The file needs a decision.
The model can sit in your rack and still ship the prompt. Telemetry is how sovereignty dies after the architecture slide is approved. DPDP does not define sovereign AI. Transfers can be lawful and still be a bad idea. Sector circulars can be stricter than DPDP. Write which instrument you are using.
If you cannot name the Data Fiduciary, the processor, the location of traces, and the erasure method, you are not ready for production personal data — whatever the architecture PDF says.
- One sentence on lawful basis or the procurement rule you are invoking.
- One sentence on where prompts, embeddings and logs live.
- One sentence on who can compel the operator.
- One artefact: packet capture, DPA schedule, or deletion certificate template.
Close this loop before the next CAB
Put “Telemetry Is the Sovereignty Loophole” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P1 CIO/CTO, not “the vendor.”
Revisit the item when the model, the GeM term, the region, or the SI changes. “AI telemetry data leakage” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.
Questions this usually raises
- If inference is on-prem, can telemetry still leave India?
- Yes. Crash reporters, APM agents, licence servers, feature-flag services and support tunnels are independent egress paths. A local GPU does not close them.
- Is usage analytics personal data?
- It can be. An officer identifier, a prompt fragment, an IP, a case ID in a span tag, or a screenshot in an error report can identify a person. Treat the analytics pipe as a processor until you prove it only carries counters.
- Do we have to keep logs in India?
- CERT-In's 28 April 2022 directions require specified logs to be retained in India for 180 days and incidents to be reported on a six-hour clock. Shipping those logs only to a foreign SaaS can fail the retention duty even if you also keep a copy.
- Can we allow a vendor to see stack traces but not prompts?
- Only if the product actually separates them. Many traces include the request body by default. Demand a config that redacts payloads, then verify with a packet capture, not a slide.
- Is a licence ping a sovereignty issue?
- A ping that sends a node ID and a version may be acceptable if documented. A ping that sends environment variables, hostnames that encode department names, or sample requests is a transfer. Read the payload.