AI Tenders
Defining "Accuracy" in a Tender Document
· 9 minute read
Accuracy without a task is a slogan. Write the job, the hold-out items, the metric, and who reviews a sample. Leave the magic 95 percent out of the bid.
The draft bid said the solution shall achieve at least 95 percent accuracy. Finance liked the number because it looked like a specification. The technical member asked 95 percent of what. Silence. A vendor later offered a certificate that their model scored 95.2 on an English benchmark that had nothing to do with Hindi notings. Both sides felt scientific. Neither could have marked a fail.
Accuracy is not a property of a platform. It is a score of a system on a task, on a dataset, under a metric, with a review rule. Change any one of those and the number is a different object. Tenders that nail a percentage and leave the rest to industry standards are buying a press note.
This guide is for officers writing evaluation and acceptance schedules. It will not give you a mandated percentage. There isn't one. It will give you a way to write accuracy so a committee can fail it without pretending to be machine-learning researchers.
Four nouns that must sit next to the word accuracy
Task. One verb from the workflow card: classify this grievance into a published scheme list; draft a noting that cites only retrieved paragraphs; extract four fields from an application. Answer officer queries is not a task.
Dataset. Items the department owns, drawn from real work, redacted if they contain personal data, frozen before the bidder tunes. If the bidder helped build the set, say so and keep a sealed hold-out they never saw.
Metric. Exact match on a label; citation-supportedness judged by a rubric; field-level precision and recall; refusal correctness. Do not say BLEU unless you know why. Do not say semantic similarity without a judge.
Human-review sample. A published n, a sampling rule, and who may overrule the metric. Agents produce fluent wrongness that metrics miss. A person has to look.
| Bad sentence in a bid | Why it cannot be marked | Replace with |
|---|---|---|
| 95 percent accuracy | No task, no set, no metric | On task T, set S, metric M, score at acceptance ≥ the published bar |
| State-of-the-art performance | No artefact | Compare two systems on S with the same rubric |
| High-quality Hindi | Taste | Officer rubric on a sample of n drafts, published bands |
| Hallucination-free | Absolute claims fail honestly | Citation-supportedness; unsourced claims are errors |
| Vendor's published MMLU / IndicQA score | Wrong task | Allowed as background, never as acceptance |
How to set a bar without a magic number
Measure the current human process on the same set if you can. If officers already classify grievances at a known error rate, the agent does not need a fantasy percentage. It needs to beat or match a band you can live with, with better logs.
If you have no baseline, run a small departmental scoring of a dumb rule or a previous tool. Write a provisional bar and a review after ninety days of shadow mode. Do not invent 95 because it sounds strict. Unreachable bars empty the bid or force everyone to lie.
Separate must-refuse items from must-answer items. A system that answers 95 percent of questions by guessing on the other 5 percent of dangerous ones is not accurate. It is reckless. Score refusals on their own line.
Where a number belongs — and where it does not
A number belongs in the acceptance schedule and in a periodic evaluation duty. It does not belong as a monthly liquidated-damage trigger unless you enjoy vendors who refuse hard cases to protect the average.
A number does not belong in eligibility (only firms whose model is 95 percent accurate may bid). That sentence is unverifiable at bid time and is a gift to whoever prints a certificate.
Public leaderboards do not belong as contractual bars. They are not your files. They leak nothing about your MIS tools.
Personal data in the evaluation set
Gold answers built from live files can be personal data. DPDP does not vanish because you said eval. Redact. Limit access. Do not let the vendor keep the set as a marketing benchmark. Assign the set to the institution. That doctrine is the same as the ownership note on fine-tunes: the set is often the more valuable artefact.
Objections you will hear — and what to do with them
These are the lines that stall the file. Answer them in the room, then put the answer in the note. A spoken answer without paper will be forgotten by the next officer.
Leadership wants a single number for the press note.
Give them the method in the file and a cautious band after acceptance. A press-note percentage is how you get a representation from a loser who also has a certificate.
We are not statisticians.
You do not need to be. You need a task, a folder of items, a rubric, and a sample. That is how you already check a junior's drafts.
The vendor will overfit if they see the set.
That is why you seal a hold-out. Share a development sample if you must. Never share the acceptance items.
Indic language quality cannot be metricated.
Then use a published officer rubric and a sample, not a fake percentage. Rubrics are allowed. Mysticism is not required.
Replace 95 percent in one week
- Day 1: delete every bare percentage from the draft bid.
- Day 2–3: for each year-one card, write task, metric, and who reviews.
- Day 4–5: assemble a redacted sample and a sealed hold-out. Decide n.
- Day 6–7: write the acceptance bar as a band or a comparison to the current process. Brief the chair so they do not put 95 back in.
How this shows up in the file
Subject: Evaluation method in lieu of a blanket accuracy percentage.
The phrase 95 percent accuracy has been removed. Each year-one workflow states a task, a department-owned dataset (including a sealed hold-out), a metric, a human-review sample, and a separate refusal score. Public model leaderboards are not acceptance tests. The evaluation set is an official artefact and will not be released to bidders as a marketing corpus.
This note is not legal advice.
A rubric an officer can mark in twenty minutes
You do not need a research lab. For a noting draft, four bands are enough: cites only retrieved paragraphs; cites a mix of retrieved and invented law; fluent but unsourced; refuses when it should have drafted. The officer ticks a band and writes one sentence. That sheet, attached to the hold-out IDs, is a metric. It is ugly. It is markable.
Publish the bands in the bid so bidders know they will be judged on citation honesty, not on how ministerial the prose sounds. A vendor who wants extra marks for eloquence is asking you to score the brochure. Eloquence without a citation is the failure mode that survives a 95 percent clause.
Keep the completed rubrics with the UAT pack. When the model refreshes in year two, the same officer class can mark the same items. Lineage of scores is how you notice that a silent upgrade broke refusals. A percentage with no sheet has no lineage.
This article is a field guide for Indian public buyers, not legal, procurement, financial or audit advice. Confirm every citation against the live GFR compilation on doe.gov.in, the relevant DoE procurement manual, GeM terms, CVC guidance, the Copyright Act, DPDP text and your own counsel before a sentence enters a tender file.
How to put this in the RFP, not the preamble
A P2 Procurement who searches “AI accuracy tender specification” is usually drafting or scoring a bid. “Defining "Accuracy" in a Tender Document” belongs in eligibility, the evaluation matrix, or a numbered annexure. If it only lives in the covering note, L1 will ignore it.
Accuracy without a task is a slogan. Write the job, the hold-out items, the metric, and who reviews a sample. Leave the magic 95 percent out of the bid. QCBS weights are a choice you must publish before opening. Accuracy is a task plus a dataset, not a slogan. SLAs for agents must name tool-calls, human gates and log export — uptime alone is a hosting metric.
Do not let a vendor write the specification and then bid on it. Record unsolicited proposals. Pay for pilots that touch personal data. Write exit before you write go-live.
- Move the control from the preamble into a scored or eligibility row.
- Attach a one-page definition (accuracy, SLA, language, data handling).
- Require an artefact in the technical bid, not a slide.
- Extend the bid date if a corrigendum is material.
- Minute the demo on your data, offline if you claimed air-gap.
Close this loop before the next CAB
Put “Defining "Accuracy" in a Tender Document” on the next change-advisory or bid-opening agenda as a single line item with an owner. If it cannot earn a line item, it will not earn a control. The owner should be a P2 Procurement, not “the vendor.”
Revisit the item when the model, the GeM term, the region, or the SI changes. “AI accuracy tender specification” is not a one-time workshop. It is a watch item. Date the last check. Unsigned watch items are souvenirs.
Questions this usually raises
- Is there a government-mandated accuracy percentage for AI?
- No. Anyone who tells you 95 percent is required is selling a number, not a method.
- Can we use the vendor's benchmark certificate?
- As background, perhaps. As acceptance, no. Certificates are not your files. Put that in the file next to “AI accuracy tender specification” so a stranger can reconstruct it. A one-line yes/no under “Defining "Accuracy" in a Tender Document” is not an answer a secretary can defend. Confirm against the live Gazette, circular or GeM term; this is not legal advice.
- What if two metrics disagree?
- Publish which one governs acceptance and which is diagnostic. A file that says both will accept everything.
- Do we publish the hold-out questions in the bid?
- Publish the method and the types of items. Keep the actual hold-out sealed or you have published the exam.