All insights

Sovereignty & Data Residency

Where Sovereign AI Claims Break Under Audit

· 9 minute read

Auditors do not grade adjectives. They grade paths. These are the paths that keep failing in real evaluations.

The claim usually dies in the second hour, not the first slide. Someone asks to see the admin path. Someone asks where last week's traces went. Someone asks whether the model vendor is a sub-processor. The room goes quiet. That silence is the audit.

This article is a script for that second hour. It is written for DPOs, internal audit and the technical member of a bid committee. It assumes the vendor has already said the word sovereign. Your job is to make them show a path.

The seven breaks we see repeatedly

BreakWhat the brochure saidWhat the log showed
Control plane abroadIndia regionCluster API and billing in another country
Hidden model APIOn-prem inferenceFallback calls to a public endpoint on timeout
TelemetryNo data leavesCrash dumps and prompts in a foreign APM
Support break-glassYour admins onlyVendor engineers with standing access
Sub-processorsSingle vendorUnlisted label, OCR or eval vendors
Training leakageWe never train on youFine-tune job in a shared tenancy
Unowned artefactsYou own outputsAdapters stuck in the vendor registry

Every row is testable. If you accept a verbal answer on any row, you have not audited. You have been briefed.

How to test each break

Control plane

Ask where the Kubernetes API, the feature-flag service, the marketplace billing and the identity provider live. Then ask for a traceroute or an equivalent from the staging cluster. If the answer is our India region handles everything and they cannot show the control-plane endpoints, mark the row amber at best.

Hidden model API

Disable outbound internet on staging. Run the demo script. Whatever still needs the internet was not on-prem. Search the orchestrator config for base URLs. Search source control for common hostnames. Ask for the timeout-and-retry policy in writing. Fallbacks are where honest local stacks become transfers.

[Telemetry](/blog/telemetry-is-the-sovereignty-loophole)

Ask for last week's traces opened in your SIEM, not in their portal. Ask whether crash dumps include prompts. Ask where support session recordings sit. CERT-In already wants specified logs in India for 180 days. A foreign APM that holds prompts is both a transfer and a log-retention failure.

Support access

Export privileged users. Look for vendor domains. Ask for the last five support tickets and whether access was time-bound. A standing admin user named support is a foreign control path, even if nobody used it this month.

Sub-processors

Demand a schedule, not a URL. URLs change after signature. Match the schedule to the endpoints you already found. Hosted OCR, speech, guardrails and eval labs are processors even if the salesperson does not think of them as AI.

Training and artefacts

Read the DPA and the acceptable-use policy. Improvement, safety and service quality are the usual hiding places for training rights. Then ask for an export of any adapter trained on your data, loaded onto a second model you already run. If they cannot export, you do not own it.

What good looks like in a room

  • An egress policy with a named exception process and a packet capture from staging.
  • A processor and sub-processor schedule with countries, purposes and a change-control clause.
  • A key ceremony the institution can repeat without the salesperson.
  • A restore of a fine-tune onto a second model the institution already runs.
  • An audit-log export the DPO can open in their own SIEM.
  • A deletion-certificate template that lists backups and eval sets by name.

If the vendor needs a week to produce these, they do not have them. Architecture that exists only after you ask is not architecture. It is a custom promise.

How to write the audit into the RFP

Put the seven rows in the bid as eligibility or as a high-weight quality block. Require artefacts with the technical bid, not after L1 is declared. Require a staging review before financial opening if the department's rules allow it. Price on a false stack is not economy. It is a future file.

The second-hour script, minute by minute

Minutes 0–15: egress. Show the proxy log from staging with outbound internet disabled, then with it enabled. Compare. Any new destination is a finding.

Minutes 15–30: identity. Export privileged users. Read domains. Ask for the last support ticket and the recording. If there is no recording, that is a finding even if the access was legitimate.

Minutes 30–45: model path. Name default and fallback. Kill the network. Run the demo script. Whatever errors name a host is a finding.

Minutes 45–60: artefacts. Demand the adapter export and a deletion-certificate template. If they need to go back to legal, the artefact does not exist.

Minutes 60–90: DPA. Read the improvement clause aloud. Ask the salesperson to explain it without the word safety. If they cannot, counsel marks the clause.

End with a written list of findings the same day. Findings that wait a week become memories. Memories become polite emails. Polite emails do not change architectures.

After the audit

Classify findings as fail, condition, or observation. Fails stop the bid or the go-live. Conditions become clauses with dates. Observations go to the risk register. Then assign an owner inside the department, not only inside the vendor. Vendor-owned findings die with the account manager.

Re-audit after the conditions are claimed closed. Closing on email is how the timeout fallback returns under a new name.

Build an internal audit programme, not a one-off raid

One heroic second hour will catch the first lie. The second lie arrives in a release note. Schedule a short re-audit every quarter and after every material change: new model, new region, new SI, new tool. Ninety minutes. Same script. Same rows. Same-day findings.

Staff it like any other IT audit. Internal audit chairs. CISO owns control-plane tests. DPO owns DPA and purpose. A CERT-In empanelled auditor is useful once a year for logs and storage, not for every quarterly pass.

Track findings in the same register you use for other systems. AI findings that live in a side spreadsheet will not get owners or dates. Owners and dates are how a timeout fallback dies and stays dead.

  • Fail, condition, observation — never 'discussed'.
  • Department owner on every fail, not only vendor owner.
  • Re-test before closure. Email is not closure.
  • Publish a one-page score to the CIO after each pass.

After two quarters the vendor will stop treating the audit as a surprise attack. That is the goal. Predictable scrutiny changes architecture. Surprise scrutiny changes slides.

Objections you will hear — and what to do with them

The vendor will say this audit is unfair because other departments accepted the same stack. Other departments' files are not a control. Ask for the artefact anyway. If they have it, the extra hour is cheap. If they do not, you have just been told you would have inherited someone else's miss.

They will say a red-team report or a SOC 2 closes the seven breaks. Those reports close the systems they name. Ask which tenant, which region, which fallback, which annotation vendor. If the report is silent, it is background reading. Background reading does not close a finding.

They will say showing egress logs would reveal proprietary architecture. Proprietary architecture that cannot be shown to a buyer who will run it on public records is not an architecture you can defend in a PAC. Offer an NDA. If they still refuse, mark the row fail.

Internal officers will say we do not have time for a second hour. The second hour is cheaper than the year you will spend after a journalist finds the fallback. Put the script in the RFP as a mandatory staging review. Time then becomes a bidder problem, not a committee problem.

After you start finding things, someone will ask whether the programme should be paused. Pause production personal data. Do not pause the audit. An audit that stops when it becomes inconvenient is a workshop. Workshops do not change stacks.

  • Unfair compared with other departments — still need the artefact.
  • SOC 2 / ISO as a substitute — only if it names this deployment.
  • Proprietary secrecy — NDA first, fail if still hidden.
  • No time — make the script a bid deliverable.
  • Bad findings — pause data, not the scrutiny.

How this shows up in the file

After the second hour, the file should contain a same-day findings list classified as fail, condition or observation, each with a department owner. Vendor-owned findings die with the account manager. Department-owned findings get dates.

Re-test before you close anything. Email is not closure. The timeout fallback that left in a release note will return under a new name if you close on courtesy.

Put the ninety-minute script in the next RFP as a mandatory staging review. Time then becomes a bidder problem. That is the only way the second hour happens on every buy, not only on the ones where someone was brave.

Questions this usually raises

Who should run this audit?
Internal audit plus the DPO, with the CISO on the control-plane questions. External CERT-In empanelled auditors are useful for logs and storage. Counsel should review the compulsion and DPA questions.
How long does a serious review take?
Two hours if artefacts exist. Two weeks if you have to discover the architecture from a demo. Do not let a vendor compress it into a lunch presentation.
What is an automatic fail?
An undeclared outbound model or telemetry path when the RFP forbade egress. A training right on customer data. A refusal to list subprocessors. Everything else can be a condition. Those three are eligibility.

Sources