Consider a safety-critical healthcare product where a surgical team opens each procedure with a short voice briefing. An LLM processes the briefing and auto-fills a safety checklist: patient identity, procedure, allergies, equipment, side. The surgeon says “no known penicillin allergy as far as the chart goes.” The LLM transcribes “no known penicillin allergy.” The checklist item for allergies auto-checks as satisfied. The patient’s chart in fact says penicillin-anaphylactic. The anaesthetist is one drug selection away from a case that will end up in a quality review.
Nothing I have described is unusual. LLMs make this class of error. Negation parsing, hedged speech, overlapping voices, accent, vocabulary that is not in the training distribution. These are the known failure modes of voice-to-structured-data pipelines, and no amount of fine-tuning drives them to zero.
But the mistake in the scenario above is not the LLM’s. The mistake is that a signal from a model was allowed to write a safety claim into the state of the system.
That design choice is more common than it should be, and in domains where the cost of being wrong is bounded, it is usually a reasonable trade. In a safety-critical context, it is the single most dangerous pattern a team can ship.
The uncomfortable principle
The design rule I would carry into any team shipping AI into a safety-critical workflow is this:
In safety-critical AI, confident automation is the dangerous kind. The default state of a safety claim must be unchecked, and only an explicit human act is allowed to promote it to checked. Confidence is not evidence. Signal is not confirmation.
That principle is easy to state and uncomfortable to implement, because it costs real time at the point of use. In an operating theatre, every item on the pre-incision checklist that cannot be auto-filled is three to five seconds of human attention. Across a shift, across a hospital, across a vendor’s install base, those seconds add up. The pressure to let the model carry the write is constant.
The principle holds anyway. The seconds are the price of the pattern. The alternative is a system whose correctness depends on the model being right on the days it is wrong.
The first-order mistake teams make
The pattern I see most often in practice is confidence-gated automation.
The product team instruments the LLM to return a confidence score alongside its proposed output. The backend checks the score against a threshold. Above the threshold, the item auto-fills as satisfied. Below the threshold, it is flagged for human review. The team ships, feels good about the safety angle, and moves to the next feature.
That design is better than nothing. It is not safe.
Four problems with it, in roughly increasing severity.
One. The confidence calibration of modern LLMs is not epistemically reliable in the sense that would warrant trust. Models can be confidently wrong, especially on inputs that look superficially similar to things they have seen. A hallucinated confidence of 0.95 reads to the system identically to a well-grounded confidence of 0.95.
Two. Negation handling is a known weak spot. “No known allergy”, “no known penicillin allergy as far as the chart goes”, “the chart does not currently list a penicillin reaction” are nuanced enough that a model can collapse them into “no allergy” with high confidence. The confidence is for the token sequence, not for the medical fact.
Three. “Flag for review” degrades to click-through fatigue the moment the system becomes trusted. Clinicians who see ten flagged items a day and nine of them are correct will start to approve the tenth by habit. The review step was supposed to be a defence; it becomes a formality.
Four. The system still writes “done” into state based on model output, in the auto-fill case. The confidence threshold moved the line between model-write and human-write, but the line still exists on the model-write side for every input above the threshold.
The result is a system that is safer than one with no gating, and still has model output in the critical path for every input the model is confident about. Confident inputs are precisely the inputs where a confident wrong answer is most dangerous, because the human has the least reason to look twice.
The design rule I would keep
The rule I would carry forward in any safety-critical product I shipped:
The model proposes. The human confirms. The system only writes after confirmation.
Three primitives make it enforceable.
- The schema makes the acknowledgement structurally mandatory. For items that carry a safety claim, the type system refuses a record that does not carry an acknowledgement event.
- The write path refuses to persist without the acknowledgement event. The discipline sits at the data layer, not the UI. A direct API call, a data migration, or an internal script cannot bypass it.
- The UI shows “proposed” and “confirmed” as two visually distinct states. The operator sees at a glance which items the model has suggested and which items a human has stood behind.
The confidence score is not thrown away. It is used for prioritisation (surface the uncertain items first), for logging (so quality teams can analyse drift), and for product analytics (so leadership can see how often the model is well-grounded). It is not used to decide whether the system can skip the human step.
Snippet one: the output schema
A reasonable TypeScript shape for a safety checklist item, with the fields that matter named explicitly:
type SafetyChecklistItem = {
id: string;
label: string; // "No known allergies"
proposed_state: "satisfied" | "not_satisfied" | "unknown";
model_confidence: number; // informational only
evidence_quote: string | null; // what the model heard
requires_human_ack: true; // not negotiable for safety items
ack: null | {
user_id: string;
role: "circulating_nurse" | "anaesthetist" | "surgeon";
timestamp: string;
method: "tap" | "voice_confirm";
};
};Two choices in that type are deliberate.
The field requires_human_ack is typed as the literal true, not as a boolean. A developer who tries to add a new safety item with requires_human_ack: false is caught by the compiler before the code reaches review. That constraint lives in the type system, which is the cheapest place to enforce invariants.
The field ack carries a role. The write path can require that the acknowledgement comes from a specific role for specific items (for example, allergies must be acknowledged by the anaesthetist, not any available clinician). This is where the pattern composes with the organisation’s clinical protocol instead of fighting it.
Snippet two: the guard on the write path
A thin wrapper around the persistence call that refuses to write unacknowledged safety items:
class UnacknowledgedSafetyItemError extends Error {
constructor(public readonly context: {
item_id: string;
proposed_state: string;
model_confidence: number;
}) {
super(`Safety item ${context.item_id} cannot be persisted without ack`);
}
}
async function persistChecklistItem(item: SafetyChecklistItem) {
if (item.requires_human_ack && item.ack === null) {
throw new UnacknowledgedSafetyItemError({
item_id: item.id,
proposed_state: item.proposed_state,
model_confidence: item.model_confidence,
});
}
await db.safetyItems.put(item);
await auditLog.record({
event: "safety_item_confirmed",
item_id: item.id,
ack: item.ack,
});
}The guard is intentionally boring. It is not a middleware, not a decorator, not a rule in a policy engine. It is a conditional at the single function that writes safety state. The simpler the guard, the harder it is to accidentally remove.
The audit log entry gives the compliance team the evidence they need months later when a case review asks “who confirmed this item, and when?” That audit trail is itself a product requirement in safety-critical domains. It is also what makes the pattern defensible to regulators.
Why “flag on doubt” is not enough
I owe this section a direct comparison, because the “flag on doubt” pattern is widespread and well-intentioned.
The pattern in short: when the model is uncertain, require explicit human review. When the model is confident, auto-fill.
The problem is that model confidence is not a trustworthy input to a safety decision. The pattern reads safe. In practice it ships every confident wrong answer directly into state. A confident wrong answer is also the kind of error the human reviewer is least likely to catch, because by construction the model’s own self-report says there is nothing to check.
The pattern I argued above requires the human acknowledgement regardless of confidence. That costs the team roughly two to three seconds per item at the point of use. In exchange, the system eliminates the entire class of “the model was confident and wrong” incidents, which is the class most likely to hurt someone.
Teams I have worked with that already implement “doubt-to-tap” are halfway there. They have accepted that the model’s uncertainty should block automation. The step they usually have not taken is to accept that the model’s confidence should also block automation. The design rule I kept forces that step.
For a related failure mode on the engineering side, I wrote Stop Delegating Your Thinking to AI. The debugging-trap section there is the same shape: a confident wrong answer applied without the two-layers-below check. On the architecture-review side, the paper response Reading “The End of Software Engineering” Against My 2026 Shipping Log names verification fidelity as the industry’s weakest link. The pattern in this piece is one concrete answer to that weakness in the specific domain where getting it wrong hurts the most.
Generalising the rule
The rule generalises cleanly past healthcare, with the same shape and the same trade-off.
Fintech anti-fraud. A model that marks a transaction as “not fraud” should not be allowed to auto-approve the transaction on confidence alone. The write of the “not fraud” claim needs an explicit decision event, whether that event is a reviewer click or a product-level policy stamp with a named owner.
Legal review. A model that marks a contract clause as “standard, no deviation” should not auto-greenlight the clause past review. The default state of the review item is unchecked, and only a lawyer’s acknowledgement moves it.
HR screening. A model that marks a candidate as “passes screen” should not auto-advance them to the next stage. The write of the advancement is a human decision, with the model’s output as evidence.
Industrial IoT. A model that marks a sensor reading as “nominal, no alert” should not auto-dismiss the alert path. The dismissal is a human act with role and timestamp.
The common thread: the write of a safety claim must be gated on a human act, not a model signal. The specific domain decides what counts as the human act and who is allowed to perform it. The design pattern is the same.
Closing: the two questions
The two questions I would ask every team shipping AI into a safety-critical context.
First question. What is the default state of a safety claim when the model is uncertain? If the answer is “we flag it for review”, you have the first-order fix. Keep going.
Second question. What is the default state when the model is confident? If the answer is “we auto-write it as done”, you are shipping confident automation. In the scenario this article opened with, that is one wrong auto-check away from a patient event. In softer contexts, it is one wrong confident call away from a product incident. In both cases the fix is the same.
In safety-critical AI, the model proposes, the human confirms, and the system only writes after. Everything else is confident automation. Confident automation is the dangerous kind.