Decision safety keeps AI-assisted decisions within acceptable risk by combining calibrated confidence thresholds, the ability to abstain, surfacing disagreement between sources, hard safety envelopes, preference for reversible actions, and escalation to accountable humans when any of these limits are reached.
Why decision safety is different from model safety
Model safety asks whether a model’s output is accurate or harmful. Decision safety asks whether the action taken on that output is acceptable in this context, under this authority, with this reversibility. A 95%-accurate model can still drive an unsafe decision if the 5% lands on an irreversible, high-consequence action. Gartner expects that by 2027, 25% of ungoverned LLM-based decisions will cause financial or reputational loss.
Seven controls
| Control | What it does | Example |
|---|---|---|
| Confidence thresholds | Below a calibrated threshold, the system asks or escalates. | Under 70% confidence, route claim to a nurse reviewer. |
| Abstention | Out-of-domain or novel situations are declined, not guessed. | New robot model, no training history: abstain. |
| Disagreement | Conflicting sources are shown, not averaged away. | Sensor says tank 70% full; manual dip says 82%. |
| Safety envelope | Hard physical or regulatory limits no recommendation can cross. | Robot speed never exceeds 1.2 m/s near people. |
| Reversibility | Prefer actions that can be undone; require more authority for those that can’t. | Hold a payment (reversible) vs send it (irreversible). |
| Rate & blast-radius limits | Cap how many actions run per hour and how much value they touch. | Max 20 reroutes per hour without supervisor sign-off. |
| Escalation | A named human gets the decision with context and the reason it escalated. | Procurement lead paged for > $100K reallocation. |
Safety envelopes for physical AI
When decisions move physical things — robots, vehicles, valves, aircraft — the safety envelope is not a policy document but a hard boundary enforced below the AI layer. Decision intelligence plans inside the envelope and escalates anything that would approach it. See decision intelligence for robotics and physical AI.
Consequence × reversibility
| Instantly reversible | Reversible at cost | Irreversible | |
|---|---|---|---|
| Low consequence | Automated | Automated / on-the-loop | Augmented |
| Medium | On-the-loop | Augmented | Augmented + second approver |
| High | Augmented | Augmented + second approver | Human only |
A simple starting policy for setting autonomy and approval requirements per decision type.
A decision safety checklist
- Classify consequenceLow, medium, high, catastrophic — per decision type.
- Classify reversibilityInstantly reversible, reversible at cost, irreversible.
- Set confidence thresholdsCalibrate them on historical outcomes, not intuition.
- Define abstain conditionsNovelty, missing context, conflicting evidence.
- Set envelopes and rate limitsWhat must never happen, and how fast things may happen.
- Name the escalation ownerWho receives the decision, within what time.
- Decision safety governs actions, not just model outputs.
- Low confidence, novelty and disagreement must change behavior.
- Irreversible, high-consequence actions need more authority.
- Physical AI needs hard envelopes enforced below the AI layer.