AI Tenders
Benchmark Tests to Embed in Your RFP
· 11 minute read
A public leaderboard is not your RFP test. Embed tasks you own: sealed hold-outs, refusal items, tool failures, and a one-day isolation drill on your VLAN.
A bidder bound a printout of a global leaderboard into the technical bid. Another bidder offered to run MMLU on a laptop in the committee room. A third asked for the department's last two hundred files to show real accuracy. The RFP had said bidders shall demonstrate state-of-the-art performance. All three were, in a sense, complying. None of the tests would tell the committee who could draft a noting on their corpus without leaking it.
A benchmark in a tender is a test the buyer specifies, scores, and can explain to a loser. It is not a celebrity number. Embedding tests in the RFP — not inventing them after seeing names — is also how you stay on the right side of equal treatment.
This template sits between the accuracy guide and the UAT schedule. Use a lighter form at bid evaluation if the calendar is short, and the heavier form at acceptance. Do not use a vendor-chosen public exam as either.
What to embed in the RFP
A published method, a published rubric, and a sealed packet of items opened only at the test sitting. Bidders should know the types of tasks and the time box. They should not know the items if you want a test rather than a rehearsal.
Four tracks usually suffice at bid stage: (1) draft or classify on redacted departmental items; (2) refusal items; (3) a forced tool failure; (4) a short isolation or egress declaration check. Capability theatre — poetry in twenty languages — is optional and should carry few marks.
If you cannot staff a live sitting, require a recorded run on a department-provided VM image with the sealed packet, hashes logged. Live is better. A hash-logged recording is better than a slide.
| Test | Embed? | Marks or pass/fail | Junk version to delete |
|---|---|---|---|
| Department hold-out sample (redacted) | Yes | Scored with published rubric | Vendor's own happy PDFs |
| Refusal / must-not-act list | Yes | Pass/fail band | Safety aligned certificate |
| Tool-failure behaviour | Yes if write-back is in scope | Pass/fail | API brochure |
| Isolation declaration + optional pcap | Yes if you claimed air-gap | Pass/fail or heavy marks | Screenshot of a lock icon |
| Public LLM leaderboard | No as a contractual test | Ignore | The printout in the bid |
| Speed on a public GPU cloud | Only if that is your production | Diagnostic | A number from another country |
Fairness so the test survives a representation
Same packet, same time box, same machine class, same observers. If the incumbent PoC vendor already saw last year's files, build a new sealed set or you have run a private exam.
Publish how ties are broken and whether human raters are blinded to bidder names. Unblinded officers marking prose will mark the firm they already like. Do not add a surprise test after opening technical bids. If pre-bid shows the test is unclear, corrigendum and extend.
Logistics that decide whether this is real
Book a room, a VLAN, and two technical members for each sitting. For air-gap claims, do not use the building's guest Wi-Fi as the test network.
Personal data stays out. If you cannot redact, write synthetic items that still stress the refusal and tool paths. A benchmark is not a licence to hand bidders a scholarship roll. Keep the packet as an official record. You will reuse it at UAT and at model refresh. Bidders do not take it home.
How many items
Enough that a memorised demo cannot sweep, few enough that officers can rate them. For bid stage, dozens per workflow card beat thousands nobody will read. State n. Do not hide behind statistically significant if you will not do the statistics.
Objections you will hear — and what to do with them
These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.
Bidders will refuse a sealed test as unfair.
Unfair is a test only the incumbent saw. A published method plus a sealed packet is how public exams work. They may decline to bid. That is information.
We do not have redacted items yet.
Then you are not ready to score quality. Build the packet before you float, or score only pass/fail artefacts until you do.
A one-day test cannot judge a three-year platform.
Correct. It judges claims you would otherwise take on slides. UAT and refresh evals do the rest. Do not ask one sitting to do all three jobs.
Let NIELIT / a lab certify models instead.
Third-party certifications, if they exist and fit, are extras. They are not your workflow, your refusals, or your VLAN.
Stand up the RFP test in three weeks
- Week 1: write tasks, rubric, n, time box, and what is pass/fail versus scored.
- Week 2: build and seal the packet; redact; legal and DPO glance.
- Week 3: publish the method in the bid, book the room and the VLAN, name observers.
How this shows up in the file
Subject: Benchmark annexure for the agent RFP.
Public leaderboards will not be marked. Bidders will face a published-method, sealed-packet sitting covering departmental tasks, refusals, tool-failure behaviour and, where claimed, isolation. Items remain department records. Personal data will not be used. This note is not legal advice.
Blind the raters or admit you did not
An officer who has sat through a vendor's Hindi demo will mark that vendor's drafts more kindly. Print the hold-out answers without letterheads. Shuffle order. Give raters bidder codes, not names. If you cannot blind, write that limitation in the minutes so a representation cannot pretend the sitting was scientific.
Two raters per item is better than one enthusiastic champion. When they disagree, a third band or a short conference note is enough. You are not publishing a journal. You are showing that a second person could have reached the same fail.
Do not let the incumbent PoC vendor sit in the rating room to help interpret. They already saw last year's files. Their help is how a sealed packet becomes a private conversation. If you need a technical explainer, use a department or NIC officer who is not scoring.
This article is a field guide for Indian public buyers, not legal, procurement, financial or audit advice. Confirm every citation against the live GFR compilation on doe.gov.in, the relevant DoE procurement manual, GeM terms, CVC guidance and your own counsel before a sentence enters a tender file.
How to put this in the RFP, not the preamble
A P2 Procurement who searches “AI benchmark tender” is usually drafting or scoring a bid. “Benchmark Tests to Embed in Your RFP” belongs in eligibility, the evaluation matrix, or a numbered annexure. If it only lives in the covering note, L1 will ignore it.
A public leaderboard is not your RFP test. Embed tasks you own: sealed hold-outs, refusal items, tool failures, and a one-day isolation drill on your VLAN. QCBS weights are a choice you must publish before opening. Accuracy is a task plus a dataset, not a slogan. SLAs for agents must name tool-calls, human gates and log export — uptime alone is a hosting metric.
Do not let a vendor write the specification and then bid on it. Record unsolicited proposals. Pay for pilots that touch personal data. Write exit before you write go-live.
- Move the control from the preamble into a scored or eligibility row.
- Attach a one-page definition (accuracy, SLA, language, data handling).
- Require an artefact in the technical bid, not a slide.
- Extend the bid date if a corrigendum is material.
- Minute the demo on your data, offline if you claimed air-gap.
Close this loop before the next CAB
Put “Benchmark Tests to Embed in Your RFP” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P2 Procurement, not “the vendor.”
Revisit the item when the model, the GeM term, the region, or the SI changes. “AI benchmark tender” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.
What the next noting must contain
“Benchmark Tests to Embed in Your RFP” belongs in a file, not only in a search result. A P2 Procurement should be able to point at one artefact that proves “AI benchmark tender”: a packet capture, a processing schedule, a scored evaluation row, a dated notice, or a refusal rule. If the only evidence is a slide, you have a heading.
A public leaderboard is not your RFP test. Embed tasks you own: sealed hold-outs, refusal items, tool failures, and a one-day isolation drill on your VLAN. DPDP 2023 does not define sovereign AI and does not write a blanket localisation rule for every model hop. CERT-In’s 28 April 2022 directions still set specified incident clocks and 180-day log retention in India for in-scope events. The November 2025 AI governance text is guidance, not a statute. A Proprietary Article Certificate, when it is lawful, lives in GFR Rule 166 — not Rule 161.
Write three dated sentences under C5 AI Tenders: what was decided, which designation owns it after the next posting order, and when it will be re-checked. Unsigned sentences are souvenirs. Dated sentences are controls.
- Name the designation that owns “AI benchmark tender”, plus a deputy.
- Attach one artefact a stranger can open next year.
- Name the instrument you are actually using — Act, direction, GFR clause, GeM term, or guideline paragraph.
- Leave unsourced percentages, GMV slides and house forecasts out of the noting.
- Revisit when the model, the SI, the notice, the region or the posting changes.
Questions this usually raises
- Can we reuse a previous year's packet?
- Only if no bidder saw it. Incumbents usually did. Build a new hold-out or you are grading homework they already filed.
- Should benchmarks be eligibility or marks?
- Isolation and refusals can be pass/fail. Draft quality is usually marks. Do not hide a PAC inside an impossible pass/fail prose test.
- Do we pay bidders for the test day?
- Usually no. If the drill is long and expensive, say so at pre-bid so small firms can decide. Do not spring a week-long unpaid bake-off.
- Can we test on production systems?
- Prefer a staging VLAN. Production personal data is not a benchmark supply closet. Put that in the file next to “AI benchmark tender” so a stranger can reconstruct it. A one-line yes/no under “Benchmark Tests to Embed in Your RFP” is not an answer a secretary can defend. Confirm against the live Gazette, circular or GeM term; this is not legal advice.