Red-Teaming Your Own AI-Assisted Spec Before Sign-Off
Three probes on your draft — not the model's wording, and not a detector score.
#ArtificialIntelligence #RequirementsEngineering #BusinessAnalysis #SpecReview #SoftwareEngineering

For a while I treated an AI-assisted requirements draft the way I once treated a tidy spreadsheet: if the sections lined up and the language sounded like a specification, the hard thinking was done. The model had filled the gaps. The peer who skimmed it said it looked complete. Someone asked whether we could sign.
It took a stock-feed go-live that stumbled on a late file — a failure the draft never named — to understand the mature move. Fluency is not readiness. Sign-off is the moment you put your name on what the system will do under ordinary stress, and the dangerous gaps are usually yours: assumptions you accepted because the draft filled them smoothly.
Sign-Off Is Not a Grammar Check
In practice, checking that a specification is well formed is already a different job from agreeing that construction may proceed. Verification asks whether the document meets quality standards and is usable for further work. Approval asks whether accountable people understand and accept the requirements so delivery can continue.
An AI draft is very good at looking verified. It produces headings, acceptance criteria, and SLA language that pass a writing check. What it does not do — unless you force the question — is make you accept the commitments underneath those sentences.
The room often confuses the two. People read for tone. They ask whether it sounds like a BRD. They ask, sometimes, whether a detector thinks a human wrote it. Detectors answer a different question entirely, and even their vendors have admitted the weakness: one widely used classifier was withdrawn after it caught only about a quarter of AI text on a challenge set while falsely flagging human text nearly one time in ten. At sign-off, authorship score is theatre. The live question is whether you are prepared to own the behaviour the draft describes when Monday morning is messy.
Fluency is not readiness. It is a risk factor.
AI risk frameworks name the same human failure under another label. People unjustifiably perceive generated content as higher quality and defer to it — automation bias — which then worsens confabulation and homogenisation risks. In requirements work, the bias looks ordinary: you stop hunting for the late file, the shared override account, and the business rule that moved three months ago, because the paragraphs already feel finished.
Three Probes — Ordinary Failure, Ownership Bend, Reality Drift
You do not need a new red-team role. You need an hour on the draft you are about to sign, using the same offensive-defence habit you should already apply to designs: attack it with desk-scale threats you have actually seen. The hour is visible cost. The false certainty of signing a fluent draft is invisible until the late file, the shared override, or the drifted rule shows up in production — and that cost is rarely bounded.
The three probes match that habit. Each one targets a different silence in a fluent document.
The model wrote the sentences. You still own the silences.
Ordinary Failure
Ask what happens when the happy path is merely late, duplicated, or half-complete — not attacked by an elite adversary.
Picture the retail stock feed that lands in the warehouse receiving system every night. The AI draft specifies mapping, retries, and an SLA. Ordinary failure asks: the file arrives at 05:40 instead of 02:00, the warehouse is already mid-shift, and the previous night's hold still sits in a staging table — who decides, and what does the system do without a heroic phone call?
If the draft only narrates success, it has not been attacked. Write the late-file path into the document, or refuse to sign until someone who owns operations can answer it in one sentence.
Ownership Bend
Ownership bend starts elsewhere. Ask who can quietly change outcomes without the change showing up as a named decision.
On the same feed, the draft may invent a clean role model. Ownership bend looks for the shared service account used "just for overrides," the export nobody logs when stock counts disagree, and the person who can force a receipt past validation because the truck is waiting. Those are not nation-state plots. They are the permissions that accumulate after the third reorg.
If the document cannot name a single accountable human for each override path, the fluency is covering a bend. Name the owner, or cut the path.
Reality Drift
Reality drift is the quieter probe. Ask what changed in the business after the brief was frozen — and whether that change ever entered the prompt context.
Here the AI is not the villain so much as the amplifier. Models resolve underspecified context by inserting plausible assumptions so the prose stays coherent. Practitioners call the silent version of that move assumption injection: unverified context slips in so the artifact can look complete. If that term is new, the one-sentence version worth keeping is: the draft fills a gap you never agreed to fill.
Industry voices on AI for requirements keep reporting the same pattern — specific stakeholder needs flattened into generic specifications that ignore organisational constraint. Reality drift is how that lands on your desk. Three months ago operations moved from "hold overnight" to "auto-cancel at 06:00." The brief never caught up. The model, given last quarter's notes, wrote a confident hold policy that no longer exists.
Open the live process, not the last prompt. Where the draft and the floor disagree, the floor wins — and the disagreement becomes an explicit gap until a human promotes a fix.
Worked Walkthrough — Retail Stock Feed to Warehouse Receiving
Take one AI-assisted integration draft as a running example: nightly stock positions from the retail ERP into warehouse receiving.
What the draft said, in polished form: file format, field mapping, retry three times, alert the integration mailbox, SLA of 04:00 completion, override available to "warehouse supervisors."
Probe 1 found the late-file silence. The warehouse starts receiving at 05:00. A 05:40 arrival with a partial file was undefined. We added a hold-and-escalate rule with a named duty manager — not another paragraph of "system will retry."
Probe 2 found the ownership bend. "Warehouse supervisors" mapped, in practice, to a shared AD group that included temporary contractors and an old service identity. We replaced the group with a named role, logged overrides, and removed the service identity from the happy-path section.
Probe 3 found reality drift. Auto-cancel at 06:00 had been live for a quarter; the draft still described overnight holds. We struck the hold language and wrote the cancel rule with the operations owner who had actually changed it.
None of that required detecting whether a model wrote the sentences. It required treating the document as a commitment under ordinary stress.
The detectors would have graded the wrong thing.
Teams that get value from AI in analysis tend to invest in context first, then generate — and they still watch for verbose, over-eager drafts that invent scope. The probes are how you spend that attention in the hour before sign-off.
After the Probes — What Ready Means
Ready means the three silences have answers on the page, or an explicit gap owned by a named person until they do. It does not mean the prose is elegant, a detector is quiet, or a peer said "looks fine."
Evaluative judgment stays human for the same reason approval was never a spellcheck: someone has to understand and accept what construction will build. The probes are how that understanding is forced when the draft arrives already fluent.
If you only have time for one pass, run ordinary failure first. Most go-lives that embarrass a tidy AI spec fail on the late file, not on sophisticated misuse.
A signed AI-assisted spec does not need prettier sentences. It needs an ordinary-failure path someone can run, a named owner for every override, and one check against the process that actually runs on the floor.
More in People
Protecting Deep Work When the Calendar Owns Your Week
Analyst focus blocks on hot programmes — treated as delivery infrastructure, not a productivity hobby.
6 min · July 25, 2026
Insider Signals in Handover Week — What Delivery People Notice
The clearest insider-risk signals on a delivery team are not nation-state theatre. They are ordinary exports, shared logins, and quiet workarounds that pile up when someone is leaving.
6 min · July 18, 2026
One Live Coordination Exercise — an Interview Loop Without Trivia
Most teams know how to stack interview stages. Far fewer design one shared problem that grades how a candidate holds ambiguity with you in the room.
8 min · July 16, 2026