Air-Gapped & On-Prem
SIEM Integration for Disconnected AI
· 9 minute read
An agent that cannot speak to the SIEM is an unmonitored processor. An agent that speaks through a cloud connector is a leak. Build a one-way, content-careful feed.
The CISO of a state SDC asked a reasonable question two weeks after the assistant went live. Why does every other application show up in the SIEM except the one that now reads beneficiary files? The SI said the platform had its own cloud security console. The CISO said that console was in another country and that the CERT-In clock did not pause for a vendor UI.
Disconnected AI is not an excuse to go unmonitored. It is a reason to be picky about the pipe. The SIEM needs events that help you notice abuse, outage and exfiltration. It does not need the novel of every prompt. The pipe must not become the fourth leak path from the on-prem article.
This guide is the integration we want in the file before a go-live certificate is signed. It assumes you already have, or are willing to run, an Indian-retained SIEM. It does not require you to buy a new logo.
Events worth shipping
| Event class | Why the SIEM cares | Payload discipline |
|---|---|---|
| Authentication and role change | Who touched the agent | User id, source host, success/fail; no prompt |
| Admin and promotion | Who changed the artefact or the flags | Hash, version, approver |
| Tool graph anomalies | Unexpected tool or burst of export-like tools | Tool names, counts, request id |
| Data-export actions | Bulk retrieve, bulk download, unusual attachment size | Counts and destinations; not the file |
| Serve death and queue collapse | Availability is a security story when people then use phones | Metrics already, plus an alert |
| Policy denials | Guardrail or DLP blocks | Rule id, not the blocked text, unless an incident extract |
The pipe that must not become a path
On a fail-closed on-prem estate, syslog or an HTTP ingest to the SDC SIEM on the management network is usually enough. Lock that route so it cannot be reused as a general proxy. On an air-gapped estate, prefer a one-way copy: write events to a store on the dark side, carry or diode them to the SIEM, never let the SIEM console administer the agent.
Do not install a SIEM agent that also does remote shell, automatic update, or cloud enrichment. Enrichment that calls a public threat-intel API from the agent host is a path. Enrich on the SIEM side, on a network that is allowed to do that, with events that already contain no prompt.
Retention: specified logs in India for 180 days is the CERT-In floor for the types the directions cover. Confirm with your security wing which of your agent events fall in. Do not keep content-bearing extracts for 180 days just because the security log next to them must live that long. Split the indexes.
Parsers, time and the clock that lies
If the agent host and the SIEM disagree about time, your detections will fire in the wrong hour and your six-hour narrative will look sloppy. Air-gapped estates that lost public NTP sometimes drift. Pin both sides to the same internal time source and alert on skew. This is unglamorous and it decides whether a timeline survives an incident review.
Write a parser that maps vendor fields onto names your SOC already uses: user, src, action, outcome, object. Leave a raw field for the rest. A unique schema that only the integrator can read will be an unindexed pile by the third month.
Test the pipe by generating a canary event — a fake bulk retrieve with a known request id — and watching it land. If the canary does not appear, you do not have integration. You have a configuration file. Repeat the canary after every agent upgrade, because shipper configs get reset the same way exporters do.
Objections
The SOC cannot parse a new schema. Then emit ECS-like or CEF-like fields and sit with them for a day. A unique JSON that only the vendor loves is how events rot in a dead index.
We already have the vendor’s cloud console. That console is their telemetry. It may even be useful to them. It is not your CERT-In store and not your correlation with the rest of the SDC.
Six hours is impossible for AI incidents. Six hours is the reporting clock after you notice a specified incident type. If your argument is that you will never notice, that is the integration problem, not a reason to skip the SIEM.
Who owns the six-hour clock
CERT-In’s six-hour incident report is not an AI feature. It is an organisational clock. Name the person who decides that an agent event is a specified incident, the person who drafts the report, and the person who files it. If those names are the same exhausted CISO, write a deputy.
Practise with a tabletop that starts from a SIEM alert, not from a phone call. The alert says bulk retrieve at 01:10. What do you pull from the observability store? Who talks to the officer whose account was used? When do you pull the network? When do you call CERT-In? Write the times in the tabletop. You will find that content extracts, if needed, take longer than the clock if you never practised the break-glass.
DPDP breach duties, once operational provisions apply, are a second clock. Do not assume they are the same as CERT-In’s. Counsel should say which events are which. The SIEM is how you notice. It is not how you finish the legal analysis.
A five-week SIEM hookup
- Week 1: pick event classes from the table. Kill prompt fields in the shipper config.
- Week 2: land events in a dedicated index with a 180-day (or agreed) retention and a shorter index for any content extracts.
- Week 3: write three detections — failed admin, bulk retrieve, off-hours promotion.
- Week 4: tabletop a specified incident, including who calls CERT-In and who holds the prompt extract.
- Week 5: remove any temporary path used during integration. Repeat a soak capture.
What goes in the file
- The event catalogue and the shipper config that proves prompts are absent.
- The network drawing of the pipe, including any diode or carry step.
- The three detections and the on-call roster.
- The tabletop record for the six-hour clock.
- The index retention settings, split between security logs and content extracts.
When those five sit in the file, the CISO’s question has an answer. The assistant is in the SIEM. The prompt is not in another country. That is a go-live condition, not an enhancement.
Prcept should emit structured security events you can ship to the SIEM you already run. We should not require our own cloud console as the only place those events exist.
How to prove this on a rack, not on a slide
“SIEM Integration for Disconnected AI” only matters if a CISO can fail it. A P1 CIO/CTO should be able to point at a cable, a registry, a licence file, a PDU reading or a SIEM index and say: this is the control. If the only evidence is a brochure that mentions “on-prem AI SIEM integration”, you do not have the control.
An agent that cannot speak to the SIEM is an unmonitored processor. An agent that speaks through a cloud connector is a leak. Build a one-way, content-careful feed. Air-gap and on-prem programmes die in the second month, when the first update, the first crash, or the first GPU lead-time slip arrives. Budget the boring path — media, offline licence, local registry, local traces — in the same note as the model name.
On-prem is not air-gapped. An India region is not either. Write the forbidden path (outbound HTTPS, licence phone-home, crash reporter, hidden model API) as a numbered list and test it with the internet off. Whatever still dies was a dependency you did not draw.
- Draw the data path for one user-visible answer under “on-prem AI SIEM integration”.
- Disable outbound internet on staging and run the demo script.
- List every remaining hop: update, licence, registry, NTP, DNS, SIEM.
- Give each hop an owner inside the department, not only the SI.
- Minute the restore or the media-transfer once before go-live.
Close this loop before the next CAB
Put “SIEM Integration for Disconnected AI” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P1 CIO/CTO, not “the vendor.”
Revisit the item when the model, the GeM term, the region, or the SI changes. “on-prem AI SIEM integration” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.
Questions this usually raises
- Must AI logs go to the same SIEM as the rest of the SDC?
- Preferably yes, so correlation works. They must not go only to a vendor cloud. If the SIEM is connected and the agent is air-gapped, copy events one-way without a return admin path.
- What is the six-hour clock?
- CERT-In’s 28 April 2022 directions require reporting of specified cyber incidents within six hours of noticing. Your SIEM and on-call must be able to notice. A dashboard nobody looks at does not start the clock in a useful way.
- Do we send prompts to the SIEM?
- Default no. Send security-relevant events: auth failures, privilege use, unusual tool graphs, data-export tools, admin actions. Content stays in the tighter observability store unless an incident requires a controlled extract.
- Is a cloud SIEM with an India region acceptable?
- Only with the same residency and telemetry reading you would give any processor: keys, subprocessors, no training, Indian retention. Many security teams accept it for connected estates. It is a poor fit for an air-gap.
- Which SIEM product must we use?
- Whichever you already operate well in India, if it can ingest the events and retain them. This article will not rank products or invent market shares.