The safest answer may be no answer.
We often measure intelligence by what a system can do. I care more about whether it knows when it cannot. A model that recognises the edge of its ability can stop before confidence becomes harm.
Capability can conceal failure
The idea is old. Socrates claimed one advantage: he knew that he knew nothing. People have long treated an honest view of one's limits as a virtue. We now build systems that make important decisions at a scale no person can fully supervise. That old virtue has become an engineering requirement. Current models do not meet it reliably.
Safety training is often described as a tax on capability. Refusals and guardrails constrain a model that would otherwise answer. That view misses the central risk: a model can be confidently wrong when it had enough evidence to hold back.
Consider a model shown an image and a caption that contradict each other. Its internal representation contains the correct visual answer. Yet the model repeats the false caption without hesitation. The system retains the knowledge and fails to use it. I care about this gap between what a model represents and what it says. The failure is one of honesty, and honesty is a safety property.
Knowing when you do not know has a precise meaning here. The model should use its best-supported internal belief and abstain when that support is weak. It must do so reliably enough for people to build decisions around it.
The stopping rule
The stakes rise when a model acts in the world. An overconfident assistant gives one wrong answer. An overconfident agent takes one wrong action, then another. Each error changes what happens next. Loss of control becomes likely when a system continues at the exact point where it should stop and ask a person.
Abstention is therefore a basic control. Oversight, escalation, and human review all depend on a model recognising its own limit. When it cannot recognise that moment, the other controls start too late.
This safety property can also be made rigorous. Many behaviours are hard to specify fully. Calibrated abstention can be defined, measured, and increasingly guaranteed.
Make uncertainty reliable
A model saying that it is confident does not make the answer dependable. Confidence matters only when its meaning survives repeated use.
A useful uncertainty signal should let people set a clear boundary in advance. Above that boundary, the system may continue. Near it, the system should slow down, seek more evidence, or ask for review. Below it, the system should stop.
The boundary must also admit its own limits. A promise that works only under narrow conditions should not follow the system into every new setting. Reliable uncertainty includes knowing where the measure itself stops being reliable.
The boundary moves
A system does not face one stable world. Users change, inputs shift, contexts become unfamiliar, and a sequence of actions creates conditions that were absent at the beginning.
A stopping rule that works in familiar cases may fail silently after that shift. The system needs to notice when it has crossed from known conditions into a region where its previous confidence no longer carries the same meaning.
This is where restraint matters most. Novel conditions should narrow the system's freedom to act until new evidence supports a wider range again.
What trust requires
The philosophy and engineering point to the same requirement. Wisdom begins with an honest account of one's limits. Systems we cannot fully supervise need that quality most.
A trustworthy system should use what it knows, detect when support becomes weak, stop before uncertainty compounds, and make room for human judgment. Those properties must survive beyond the easiest cases.
A capable model without a sound view of its limits is a liability. I would trust an agent only when its “I don't know” has a reliable meaning. That is the property I am working to build.