All insights

Sovereignty & Data Residency

Telemetry Is the Sovereignty Loophole

· 10 minute read

The model can sit in your rack and still ship the prompt. Telemetry is how sovereignty dies after the architecture slide is approved.

The CIO of a power-sector PSU had already won the argument. Inference would run on two servers in the plant's DMZ-adjacent VLAN. No public model API. The board note said sovereign. Three weeks after go-live, the network team dumped outbound TLS from the new VLAN and found a steady trickle to an American observability host. The bodies were compressed traces. Inside the traces were prompt prefixes, officer usernames and the account numbers the agent had been asked to reconcile.

Nobody had lied in the architecture review. The model was local. The product's default installer had also enabled a hosted APM because that is how the vendor debugs customers. The checkbox was buried under Advanced. The MSA said data remains in India. The MSA did not define data to include span payloads.

This teardown is the list we use when a CIO says the rack is ours, so we are done. You are not done until every packet that leaves the VLAN has an owner, a purpose and a redaction rule. Sovereignty that ignores telemetry is a press note.

The eight pipes that ignore your architecture slide

Vendors do not call these transfers. They call them health, product improvement, crash reporting, customer success. The packet does not care what they are called.

If a pipe is not on this table, add the row. Do not assume it is innocent because it is small.
PipeWhat usually rides in itWhat to demand
APM / tracing agentSpans with URLs, headers, sometimes full request bodiesLocal collector only; payload redaction on; foreign host denied at the proxy
Crash / exception reporterStack traces plus the request that died, often with PIIOffline crash folder; manual triage; no automatic upload
Product analyticsClicks, session replay, sometimes typed textDisable or self-host; no session replay on citizen workflows
Feature flags / remote configEnvironment names, user keys, rule evaluationsLocal flags; no user keys that are officer emails
Licence / entitlement checkHost IDs, versions, sometimes hardware and user countsPublished payload; offline licence file; no sample data
Model or guardrail fallbackThe full prompt when the local path failsFail closed; no silent hosted fallback
Support tunnel / remote assistWhatever the engineer can see, plus copied logsBreak-glass with your jump box; recorded sessions; Indian staff if the file requires it
Container and model pullsRegistry metadata; sometimes build secrets in layersInternal registry; signed artefacts; no live pull from a public host

Two of these pipes are cultural, not technical. Product analytics exists because SaaS companies live on funnel metrics. Support tunnels exist because the alternative is teaching your team to debug. Both are rational for a vendor. Neither is rational as a default inside a citizen workflow.

How to tear the installer down before it lands

Do not start with the brochure. Start with a staging VLAN that has no default route except through a logging proxy. Install the product the way the vendor wants it installed. Then read what tried to leave.

  1. Capture DNS and TLS destinations for seventy-two hours of soak, including a simulated crash and a licence renewal window.
  2. Open the config tree and search for ingest, sentry, segment, datadog, honeycomb, launchdarkly, amplitude, pendo, intercom, and the vendor's own telemetry hostname.
  3. Trigger one failure that includes a dummy PAN or Aadhaar-shaped string. See whether that string appears in any outbound body.
  4. Ask for the data map of every outbound host: fields, retention, country, subprocessors.
  5. Refuse any host that cannot be switched off without breaking inference.

If the vendor says the product will not run without their cloud collector, you do not have an on-prem product. You have a thin client with a local GPU. Write that sentence in the evaluation note. Committees understand it.

Objections from vendors and from your own ops

We cannot support you if we cannot see crashes. Then you will support through a folder of local crash dumps that we upload after review. You will not have a standing pipe into production prompts. If that costs more professional-services days, price it. Do not smuggle it through telemetry.

The data is aggregated. Aggregation is a process, not a vibe. Show the transform. If the transform happens after the foreign collector receives the raw event, the transfer already occurred.

CERT-In requires us to monitor. Yes. It requires specified logs retained in India and incidents reported fast. It does not require you to give a SaaS vendor a copy of the prompt. Local SIEM is how you satisfy the direction without opening the loophole.

Our own SRE team wants Datadog. Then run Datadog, or any other tool, as an in-country collector with a written transfer decision if anything still leaves. Do not confuse a familiar dashboard with a lawful path.

Feature flags as a remote control

A feature-flag service can change the product under you. A flag that enables a hosted critic, a new OCR vendor, or a session-replay sample is a change of means. If that service lives abroad and reads user keys, it is also a standing transfer of staff identity.

Demand local flags, or a frozen flag file you promote like a model. If the vendor says flags must be live so they can disable a bad feature, give them a ticket path and an officer who can flip a local flag. Do not give them a remote hand on a citizen workflow.

Read the default sample rates. A hundred-percent trace sample with bodies on is not observability. It is a second case file in someone else’s cloud. Turn it to skeletons, locally, before the first officer logs in.

A two-week telemetry hunt, then a 90-day close

  • Week 1: staging VLAN, logging proxy, full installer, crash test, licence window. Produce the destination list.
  • Week 2: kill or replace every destination that carries content. Write the allow-list into the proxy and into the contract.
  • Days 15–45: repeat the hunt after the first patch. Chart updates are how pipes return.
  • Days 46–75: point remaining observability at the in-VLAN collector and the SIEM. Prove CERT-In retention locally.
  • Days 76–90: tabletop a leak. If a prompt fragment left, treat it as a processor incident and practise the notice path.

What goes in the file

The file is a packet story, not a slogan. Keep the proxy logs from the soak, the destination table, the redaction config, the MSA clause that forbids undeclared collectors, and the date of the last hunt after a patch. If those artefacts are missing, the sovereignty claim is a hope.

Prcept AI is built to run without a vendor collector. If we ever need a support artefact, it should leave as a file you chose to send, not as a daemon you forgot was there.

How to defend this in the file

A P1 CIO/CTO will be asked to explain “Telemetry Is the Sovereignty Loophole” to a secretary who has ten minutes. Do not start with the model. Start with the store, the hop, the clause, or the residual risk. “AI telemetry data leakage” is a search phrase. The file needs a decision.

The model can sit in your rack and still ship the prompt. Telemetry is how sovereignty dies after the architecture slide is approved. DPDP does not define sovereign AI. Transfers can be lawful and still be a bad idea. Sector circulars can be stricter than DPDP. Write which instrument you are using.

If you cannot name the Data Fiduciary, the processor, the location of traces, and the erasure method, you are not ready for production personal data — whatever the architecture PDF says.

  • One sentence on lawful basis or the procurement rule you are invoking.
  • One sentence on where prompts, embeddings and logs live.
  • One sentence on who can compel the operator.
  • One artefact: packet capture, DPA schedule, or deletion certificate template.

Close this loop before the next CAB

Put “Telemetry Is the Sovereignty Loophole” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P1 CIO/CTO, not “the vendor.”

Revisit the item when the model, the GeM term, the region, or the SI changes. “AI telemetry data leakage” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.

Questions this usually raises

If inference is on-prem, can telemetry still leave India?
Yes. Crash reporters, APM agents, licence servers, feature-flag services and support tunnels are independent egress paths. A local GPU does not close them.
Is usage analytics personal data?
It can be. An officer identifier, a prompt fragment, an IP, a case ID in a span tag, or a screenshot in an error report can identify a person. Treat the analytics pipe as a processor until you prove it only carries counters.
Do we have to keep logs in India?
CERT-In's 28 April 2022 directions require specified logs to be retained in India for 180 days and incidents to be reported on a six-hour clock. Shipping those logs only to a foreign SaaS can fail the retention duty even if you also keep a copy.
Can we allow a vendor to see stack traces but not prompts?
Only if the product actually separates them. Many traces include the request body by default. Demand a config that redacts payloads, then verify with a packet capture, not a slide.
Is a licence ping a sovereignty issue?
A ping that sends a node ID and a version may be acceptable if documented. A ping that sends environment variables, hostnames that encode department names, or sample requests is a transfer. Read the payload.

Sources