All insights

Governance & Audit

Proving an Agent Followed Policy, Not Vibes

· 10 minute read

A vendor who says the model is aligned has given you a vibe. A reconstructable case file — prompt hash, retrieval set, tool calls, officer action — is evidence. Collect the second thing.

The integrator's slide said the agent was constitutionally aligned and policy-aware. The demonstration rejected a fabricated ineligible file with a fluent paragraph about fairness. The committee applauded. Two months later an RTI applicant asked why her identical neighbour had been coached toward a scheme she had been denied. The department produced the slide. The applicant produced a printout of her chat. The two documents did not know each other. There was no prompt version on the file, no retrieval list, and no hash of the circular the agent claimed to have read. Alignment had been a feeling in a conference room.

This is a teardown of that feeling. It is written for DPOs, internal auditors and the officer who will have to explain a single case. The claim is simple. You cannot prove an agent followed policy unless you can bind four artefacts to the case: which policy text it was allowed to see, which prompt and model it ran, which tools it called, and which human action closed the loop. Everything else is branding.

It is not legal advice. CAG has not issued a departmental circular that lists those four artefacts as a new legal form. RTI will still turn on your existing exemptions. DPDP will still turn on purpose and security. The four artefacts are how you remain able to tell the truth under all three.

What vibes look like in a bid and in a go-live note

Vibes have a vocabulary. Aligned. Guardrailed. Constitutional. Safety-tuned. RLHF'd. Policy-aware. Grounded. Responsible by design. None of those words name a file you can retrieve for case 2026/GRV/1184 on a Tuesday when the applicant is sitting outside.

Vibes also have props. A model card. A red-team slideshow. A certificate from a private lab. A quote from the India AI Governance Guidelines. Those props can be useful annexures. They do not bind a particular output to a particular policy version.

If the go-live note uses the vocabulary and does not attach the bindings, the note is a vibe that someone will later be asked to defend as if it were a control.

The four bindings that turn a chat into a case

Policy binding. The corpus identifiers and document versions the agent was allowed to retrieve for that workflow. Not 'the department's circulars'. The number, the date, the hash or the record-room identifier. If the agent can see a superseded circular, the binding must show that it did or did not.

Runtime binding. Model hash, prompt hash, tool-allow-list hash, temperature or decoding settings if they can change the output. A prompt that was edited on Friday afternoon without a hash change is how two identical inputs diverge and nobody can say why.

Action binding. Every tool call, with arguments and results, and every refused tool call. Retrieval without a tool log is a ghost citation. A send without a tool log is an unowned speech act.

Closure binding. What the officer did: accepted, edited, rejected, or never saw. If the officer never saw it and the citizen did, you do not have a policy-following agent. You have an unofficial spokesperson.

Ask for these objects in acceptance testing. A demo that cannot export them has not followed policy. It has performed policy.
BindingMinimum objectFail if
PolicyDocument ids + versions + hash or record number for the retrieval setAnswer cites 'rules' with no identifier, or the set cannot be replayed
RuntimeModel hash, prompt id/hash, tool-list hashProduction prompt can change without a ticket, or hashes are 'available on request'
ActionOrdered tool calls, including denialsOnly the final paragraph is stored
ClosureOfficer identity, decision, timestamp, optional override noteThumbs-up in a private chat, or auto-send with no identity

Replay is the test

Evidence that cannot support a replay is a story. Take a closed case. Reload the same corpus versions, the same prompt hash, the same tools. You will not always get the same tokens. That is a known property of these systems. You should get the same retrieval set, the same permitted tools, and the same closure rule. If even those drift, you cannot say the agent followed policy. You can only say it followed a mood.

Where bit-identical replay is impossible, record the sampling settings and keep the actual output artefact. Courts and auditors reconstruct what happened, not what would happen in a perfect laboratory.

Do not accept 'the model is deterministic in this mode' as a substitute for storing the output. Store the output anyway.

What law asks without naming hashes

A CAG party reconstructing a rejected scholarship will not ask for your alignment story. They will ask who approved the system, what it retrieved, what it proposed, and what the officer signed. That is the same four bindings in administrative language. There is no need to invent a CAG AI circular to justify collecting them.

RTI applicants will ask for the information that led to a decision. If your only record is a vanished consumer-chat log, you will discover the cost of shadow AI and of vibe-based official AI in the same reply.

CERT-In already wants specified ICT logs kept 180 days in India. DPDP will want you to know what you processed and why. Hashes are how engineers meet those older duties for a new kind of actor. They are not a new religion.

Teardown of common substitutes

The model card. Useful for intended use and limits. Does not bind case 1184.

The red-team report. Useful for known abuse classes. Dated the day before go-live, it does not describe Tuesday's retrieval set.

The officer's memory. Perishable. Transfer-prone. Not a record.

The vendor's 'constitutional classifier'. A component. If it is in the path, it needs a version and a log like everything else. Its marketing name is not a binding.

A screenshot of the chat. Better than nothing, worse than a stored artefact with hashes. Screenshots miss tool calls and are easy to crop.

Objections you will hear — and what to do with them

These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.

Storing all that will create a personal-data problem.

You already have a personal-data problem if the case exists. Bindings are how you apply retention and erasure with precision. A soup of unindexed chats is worse. Design retention with the DPO. Do not use privacy as an excuse for amnesia.

Our model vendor will not give us hashes.

Then you cannot prove what ran. That is a procurement fact, not a philosophical one. Score it. Prefer a stack that can show a hash on the rack you control.

Non-determinism makes replay pointless.

Replay the bindings, not the poetry. Retrieval set, tools, closure rule, stored output. That is enough to say whether policy was in the path.

The guidelines already require responsible AI, so we are covered.

Guidelines are not evidence. A minute that says we are responsible is a vibe. Bind the case.

A 20-day evidence drill

Pick one live or staging workflow that already drafts. Do not start with a theoretical agent.

  1. Days 1–4: export one closed case. List what you can already prove against the four bindings. The gaps are the backlog.
  2. Days 5–10: instrument hashes and retrieval ids in staging. Refuse to accept a build that cannot export them.
  3. Days 11–15: replay three cases. Record what matched and what only 'sounded similar'.
  4. Days 16–20: write the acceptance row into the contract or the go-live note. No export, no production personal data.

How this shows up in the file

Attach three exported cases, the hash scheme, and the sentence that production will not run if bindings are missing. That packet is the opposite of a slide titled aligned.

When someone asks whether the agent followed policy, hand them a case, not an adjective.

What the next file must contain

“Proving an Agent Followed Policy, Not Vibes” earns a line in the noting only if a P6 Compliance/DPO can attach proof of “AI policy compliance evidence.” A heading is not proof. A vendor slide is not proof. A workshop photograph is not proof.

Write three dated sentences: what was decided, who owns it after the next posting order, and when it will be re-checked. If you cannot write the three sentences, you are not ready to buy, to sell, or to go live.

Leave unsourced percentages out of the note. DPDP is not a blanket localisation statute. The November 2025 AI governance text is guidance, not an Act. CERT-In’s 28 April 2022 directions still set specified incident and log clocks. A PAC, when lawful, lives in GFR Rule 166.

  • Name the designation that owns “AI policy compliance evidence.”
  • Attach one artefact a stranger can open next year.
  • Record the instrument you are actually using.
  • Revisit when the model, the SI, the notice or the posting changes.

This article is informational field guidance for Indian public institutions, not legal, procurement, security-accreditation, academic-regulation or engineering advice. Confirm against the current Gazette, DPDP text and Rules, CERT-In direction, India AI Governance Guidelines, UGC/AICTE/NAAC notices, NEP documents, GFR, departmental manual and your counsel before you file it. Guidelines are not statute. Circulars move.

Questions this usually raises

Is a model card enough to prove policy compliance?
No. A model card describes intended use. Compliance evidence is a bound case: what was retrieved, what ran, what was called, what the officer did.
Do we need bit-identical replay of the model's words?
No. You need replay of the bindings and storage of the actual output. Token-identical regeneration is a laboratory luxury, not the administrative test.
How long should we keep the bindings?
At least as long as the underlying file must be kept, and not shorter than CERT-In's 180-day ICT log floor where those directions apply. Align with your record-retention schedule and the DPO. Do not invent a special AI number.
What if the officer edited the draft heavily?
Store both artefacts and the fact of the edit. An edited draft is still evidence of what the agent proposed. Hiding the proposal makes the officer the sole author of a machine-shaped text.
Does DPDP require prompt hashes?
The Act does not use that phrase. It does require you to know what you processed and to secure it. Hashes are an engineering method. They are not a statutory form.

Sources