When Machines Learn to Doubt
From Prediction to Self-Regulation
When we ask a language model a question and receive an answer, there are a number of different things we might want to know about what has just happened. Most obviously, we would like to know whether the answer is correct. We might also want to know how strongly the available information supported that answer, whether the model itself possesses some estimate of how likely it is to be wrong, and, finally, whether such an estimate actually influences what the model subsequently does.
These questions sound similar, but they describe quite different capacities. A system may produce a correct answer by chance. It may produce a wrong answer with great apparent certainty. It may generate the sentence “I am not sure” simply because that sentence is statistically appropriate in the context, without anything particularly interesting happening internally. Conversely, it might contain information about the reliability of its own answer that is not faithfully represented by what it tells us about its confidence.
The last distinction is, I believe, particularly important. There is a substantial difference between a system that can produce something resembling a confidence statement and a system in which an internal representation of confidence actually participates in controlling behaviour. The latter is no longer merely a question of what the machine outputs. It becomes a question about the organisation of the system itself.
A recently published paper in Nature Machine Intelligence provides some fascinating evidence in precisely this direction. Dharshan Kumaran and colleagues set out to investigate whether contemporary large language models merely exhibit confidence-like signals, or whether they actually make use of an internal sense of confidence when deciding what to do. More specifically, they examined a very simple decision: when presented with a factual question, should the model answer, or should it abstain?
What they found is, in my view, considerably more interesting than another incremental improvement in AI benchmark performance. Their results suggest that language models form relatively rich internal representations associated with confidence and that these representations are then used to regulate subsequent behaviour. More importantly still, when the researchers experimentally manipulated these confidence-related internal states, the behaviour of the model changed accordingly.
Whether there is anything it feels like for the model to be uncertain is an entirely different question, and not one the researchers attempt to answer. But that question is also somewhat beside the point here. What the experiments appear to demonstrate is something narrower, and in my opinion already sufficiently remarkable: a machine can generate an answer, form something resembling an evaluation of the reliability of that answer, and use this evaluation when deciding whether to commit to it.
In other words, we appear to be moving from prediction towards self-regulation.
Knowing Something and Knowing How Well You Know It
The relevant concept here is metacognition, generally understood as cognition about cognition. In everyday language we might call it “thinking about thinking”, although in cognitive science the term can be used more functionally. A cognitive system performs some task, represents something about the quality or reliability of that performance, and uses this information in guiding further behaviour.
Humans do this all the time. If someone asks me the capital of Germany, I will answer “Berlin” without much hesitation. If asked about the capital of a country I visited once twenty-five years ago, a very different process may occur. A candidate answer might come to mind, but with it comes uncertainty. I may search my memory again, tell the other person that I am not sure, look up the answer, or simply decline to commit.
There are therefore at least two levels involved. At the first level, I am trying to answer the question. At another level, I am assessing something about my own answer.
This distinction is not restricted to humans. There is a long history of studying confidence and uncertainty in other animals. In a particularly relevant experiment published in Nature in 2008, Adam Kepecs and colleagues trained rats to perform difficult odour-discrimination tasks. The researchers found neural activity in the rats' orbitofrontal cortex that behaved as predicted by computational models of decision confidence. Perhaps more importantly, confidence appeared to affect subsequent behaviour: rats were willing to wait longer for an expected reward when the evidence suggested that they should be more confident in their decision.
The rats, of course, did not explain any of this to the researchers. They did not announce that they were 73 per cent confident about the odour they had just smelled. Confidence was inferred from the structure of their behaviour and from neural activity.
This is worth keeping in mind when thinking about artificial systems. If we insist that metacognition must take the particular form in which humans consciously experience and linguistically report it, we risk building the conclusion into the definition. A more useful question, at least initially, is functional: does the system possess information about the reliability of its own cognitive performance, and does that information affect what it does?
This is essentially the question Kumaran and colleagues attempted to answer.
How Do You Measure the Confidence of a Language Model?
Before getting to the most interesting part of the experiment, it is worth spending a little time on what exactly “confidence” means in the context of a language model, because the term can otherwise become misleading very quickly.
A language model does not ordinarily generate text by first deciding with certainty what the next word will be. At each step it produces numerical scores—logits—for a very large set of possible next tokens. These scores can be converted into a probability distribution. Some candidate tokens receive high probabilities, others very low ones, and the next token is then selected according to the decoding procedure being used.
Imagine, for simplicity, that I give a model a multiple-choice question with four possible answers and force it to reply only with the number 1, 2, 3 or 4. The internal output distribution might assign something like 70 per cent probability to answer 2, 18 per cent to answer 3, 8 per cent to answer 1 and 4 per cent to answer 4. On another question, the distribution might be 28, 27, 24 and 21 per cent.
Even before knowing whether either answer is correct, these are plainly different situations. In the first case, the model strongly favours one option. In the second, it does not.
It is tempting to call the 70 per cent number “confidence”, but there is an important complication. A model assigning 70 per cent probability to an answer is not necessarily correct on 70 per cent of comparable questions. Neural networks can be systematically overconfident or underconfident. The raw probability distribution therefore needs to be calibrated if we want to interpret it as a meaningful estimate of correctness.
The researchers dealt with this using a standard procedure called temperature scaling. Put somewhat simply, they used a separate set of questions where the correct answers were known and adjusted the sharpness of the model's probability distribution so that stated confidence better corresponded to actual performance. If a collection of answers ends up with a calibrated confidence of around 70 per cent, we would ideally like roughly 70 per cent of those answers to turn out to be correct.
This gave the researchers one confidence measure, which they call calibrated confidence.
It is important, however, to understand what this measure is. They did not open GPT-4o, find a variable labelled “confidence”, and read out the number. Calibrated confidence is an external measurement derived from the model's output distribution. We have reason to believe it corresponds to something useful about the model's internal state, but it remains a read-out.
The researchers therefore used a second approach as well. After the model had answered a question, it was shown the question, its possible answers and its own previous response in a separate forward pass, and asked to assess the probability that its answer was correct. This is what the authors call verbal confidence.
Here the system is, at least functionally, doing something a little different. Rather than merely favouring one candidate answer during generation, it is being asked to take its own previous answer as an object of evaluation.
Whether we should call this introspection is a question I will return to later. For now, what matters is that the researchers had two different windows into confidence: one derived from the output probabilities involved in answering the question, and another derived from an explicit subsequent self-evaluation.
The interesting question then becomes: does any of this matter to the model itself?
Giving the Machine Permission to Say “I Don't Know”
In the first phase of the experiment, there was no possibility of abstaining. The model had to select one of four answers. Unsurprisingly, its calibrated confidence was strongly related to whether its answers were correct. For GPT-4o, higher confidence corresponded very closely to lower error rates.
This is useful, but not yet especially mysterious. If the internal computations of a model strongly favour the correct answer, we might expect the output probabilities to reflect this.
The second phase is where things become more interesting.
The researchers presented the model with the same basic task, but introduced an additional possibility: it could abstain. If no answer seemed clearly correct, the model was allowed to decline rather than guess.
GPT-4o made substantial use of this option. In the relevant experiment it abstained on 56.6 per cent of trials. Among the questions it did answer, its accuracy increased from 63.7 per cent in the forced-choice condition to 69.1 per cent.
In other words, giving the system the ability to say “I don't know” made the answers it did choose to give more reliable.
One could, of course, explain this in less interesting ways. Perhaps difficult questions have certain linguistic characteristics and the model had simply learned to recognise them. Perhaps the relevant knowledge was harder to retrieve. Perhaps particular topics or sentence structures were associated with abstention. There would then be no need to posit anything resembling internal confidence-guided control.
The researchers tested several such possibilities, including objective question difficulty, accessibility of relevant knowledge through retrieval methods, and semantic features of the questions themselves.
Confidence was by far the strongest predictor.
For GPT-4o, the standardised effect of calibrated confidence on abstention was roughly an order of magnitude larger than that of the competing predictors. The model behaved approximately as if there were a confidence threshold below which abstention became increasingly attractive.
For GPT-4o, the estimated halfway point was around 77 per cent. At roughly this level of calibrated confidence, the model was equally likely to answer or abstain.
It would be wrong, however, to imagine that somewhere inside GPT-4o there is a line of code stating “IF confidence < 77% THEN abstain”. The transition was gradual rather than absolute. The lower the confidence, the more likely abstention became. The authors therefore describe the policy as a soft, probabilistic threshold.
This is actually rather familiar from biological behaviour. Humans and animals seldom operate with mathematically clean boundaries either. Increasing evidence or confidence gradually changes the probability that one action rather than another will be selected.
So far, however, we are still dealing with correlations. Confidence predicts abstention very well, but it remains possible that both are produced by some other process. If one wants to make a stronger causal claim, the natural next step is to interfere with the proposed mechanism and see whether the behaviour changes.
Which is exactly what the researchers did.
Opening the Machine
Most of the initial behavioural work in the paper was carried out with GPT-4o. For the next experiment, however, the authors switched to Google's Gemma 3 27B model.
The reason has less to do with the relative intelligence of the models and more to do with access.
When interacting with a commercial model such as GPT-4o through a standard API, researchers can send information into the system and observe what comes out. Depending on the interface, they may also be able to obtain output probabilities and other limited information. What they generally cannot do is freely inspect and manipulate all of the model's intermediate internal activations while it is processing a prompt.
Gemma is an openly available model for which the researchers could run the model themselves, in this case using Google's official JAX implementation. This means they could access the numerical activation states inside individual transformer layers and, crucially, alter them during inference.
To understand the significance of this, it helps to say a little about what information inside a transformer actually looks like.
There is a natural temptation to imagine that a neural network works somewhat like conventional software: somewhere there is one variable for France, another for confidence, another for whether the user is angry, and so on. This is generally not how such systems represent information. Their internal states consist of large vectors—long arrays of numbers—distributed across many dimensions. Concepts and computational properties are typically represented through patterns in this high-dimensional space rather than being stored in neat symbolic boxes.
Nevertheless, some directions through this space can systematically correspond to meaningful differences in behaviour.
Imagine, merely as a visual aid, a cloud of points representing internal model states while answering many questions. Suppose questions where the model is highly confident tend to occupy one region of this cloud, while low-confidence cases tend to occupy another. If that difference is sufficiently consistent, one can calculate a direction from one region towards the other.
In its simplest form, one might take the average activation vector from high-confidence examples, subtract the average activation vector from low-confidence examples, and obtain a vector pointing approximately from “lower confidence-like states” towards “higher confidence-like states”.
That vector can then be used not merely to observe the model, but to intervene in it.
This is known as activation steering.
Turning Confidence Up and Down
Kumaran and colleagues first identified trials where Gemma gave correct answers but displayed substantially different confidence profiles. They examined the model's residual-stream activations—the evolving high-dimensional state that carries information through the transformer—at the point immediately before the model committed to an answer.
More precisely, they constructed their steering vectors by comparing trials with high and low confidence margins: the difference between the confidence assigned to the most likely substantive answer and the confidence associated with abstention. They then averaged the relevant internal activations and derived a direction separating high-confidence from low-confidence states.
It is worth being careful with the terminology here. This does not mean that they discovered “the confidence vector” in the sense that confidence has now been fully located and explained. High-dimensional neural networks do not generally afford such simple interpretations. What they discovered was a direction in activation space that reliably corresponded to confidence-related differences and could be experimentally manipulated.
The decisive experiment was to add this vector to the model's ongoing activation state while it answered new questions.
Pushing the model in the high-confidence direction made it much more likely to answer. Pushing it in the opposite direction made it much more likely to abstain. The magnitude of the effect was remarkable. Across the strongest low- to high-confidence steering conditions, the abstention rate changed from 66.5 per cent to 7 per cent.
The model had not been given additional factual information. It had not suddenly learned more about the world. The researchers had changed an internal state associated with how confident it was in what it already “knew”.
And its behaviour changed with it.
This is, I think, the crucial part of the paper. Correlations between internal neural-network states and interpretable concepts are interesting, but they are always vulnerable to the possibility that we have merely identified a passenger rather than a driver. If manipulating the state changes the behaviour in the predicted direction, the case that we are looking at something causally relevant becomes considerably stronger.
The authors went one step further and attempted to determine how the manipulation changed behaviour. Their mediation analysis suggested that the majority of the effect—around 67 per cent—was explained by a redistribution of confidence itself: high-confidence steering shifted confidence towards actual answer options and away from abstention. A smaller portion appeared to work through changes in the policy connecting confidence to the ultimate decision.
There was also another revealing consequence.
As the researchers pushed Gemma towards greater confidence, it answered more questions—but the accuracy among those answers declined. Coverage rose dramatically, while accuracy fell.
Which brings us to an important distinction.
Confidence Is Not Truth
The activation-steering experiment did not make Gemma more knowledgeable. It made Gemma more willing to behave as though its knowledge justified an answer.
This distinction may sound obvious, but it is central to understanding both the experiment and the practical implications.
There are at least three different things at stake: whether an answer is actually correct; how reliable the system estimates that answer to be; and what the system decides to do given that estimate. In a well-functioning cognitive system, these should relate to one another sensibly. But they are plainly not identical.
Humans provide abundant examples. A person may be correct but uncertain, or spectacularly wrong and entirely convinced. Confidence influences whether we speak, whether we persist, whether we seek more information, whether we take risks, and, importantly, whether other people believe us. None of these behavioural consequences guarantee that the underlying proposition is true.
The same basic separation appears here in an artificial system. Increasing Gemma's confidence-related activation did not upload missing facts. It changed the internal conditions under which the model was willing to commit.
This is useful because it also clarifies what “overconfidence” could mean for a machine without requiring us to anthropomorphise it. An AI does not need an inflated ego, pride, arrogance, or an embarrassing amount of self-esteem to be overconfident. It merely needs its internal estimate of reliability to be systematically too generous relative to its actual reliability, or a policy that treats a given level of confidence as sufficient for actions for which it should not be sufficient.
That sounds much less exotic. And, for autonomous systems, considerably more important.
Confidence and Policy Are Two Different Things
A later phase of the study makes this distinction particularly clear.
The researchers explicitly instructed the models to abstain whenever confidence fell below a specified threshold and then systematically changed that threshold. Unsurprisingly, higher thresholds generally produced more abstention.
But conceptually, this allows the authors to separate two stages. First, the model forms something resembling an internal confidence representation. Second, a decision policy determines what to do with it. This distinction is simple, but its implications are substantial. Suppose an AI system assesses an answer as having roughly 70 per cent probability of being correct. Is that enough?
There is no general answer.
If I am asking an AI where I am most likely to find good vegan cake in Frankfurt, 70 per cent may be perfectly adequate. I can walk to the café and discover that the model was wrong without any serious harm.
If the AI is deciding whether a suspected tumour is malignant, whether a bridge is structurally sound, whether an aircraft component can safely remain in service, or whether an autonomous vehicle should interpret the object in front of it as a pedestrian, the same level of uncertainty may be entirely unacceptable.
The confidence estimate has not changed.The appropriate action threshold has. In other words, knowing how uncertain you are and knowing what to do about that uncertainty are separate problems.
This seems to me one of the most practically important lessons of the paper. Much public discussion treats AI reliability largely as an accuracy problem: build a smarter model, reduce hallucinations, increase benchmark performance. All of this is obviously desirable. But an agent that acts in the world needs more than accurate first-order cognition. It also requires an adequate model of its own uncertainty and an appropriate policy governing when uncertainty should lead to caution, verification, escalation or abstention.
A system can fail at each of these levels independently. It may simply be wrong. It may be wrong about how likely it is to be wrong. Or it may have a reasonably accurate estimate of its own uncertainty and nevertheless be configured to act at an irresponsible threshold.
As we move from conversational AI towards autonomous agents, this distinction becomes increasingly difficult to ignore.
What the Model Says It Feels and What Is Happening Inside
There is another result in the paper which I find especially intriguing.
As described earlier, the researchers measured not only calibrated confidence from the model's answer probabilities but also verbal confidence, produced by asking the model in a separate pass to assess its previous answer.
Perhaps surprisingly, verbal confidence was not as good at distinguishing correct answers from incorrect ones as calibrated confidence. But it still independently predicted whether the model would abstain. That means the two measures were not simply different ways of reading exactly the same signal. Indeed, across the models tested, their correlation was only moderate.
The researchers therefore looked more directly at Gemma's internal activation state at the final point before the answer was generated. Using linear probes, they found that these activations contained substantial information about both calibrated and verbal confidence. More interestingly, the internal state appeared to contain information beyond either observable measure individually.
The interpretation proposed by the authors is that the model may possess a richer, multidimensional internal confidence representation, from which both output-probability confidence and verbal self-reported confidence are partial and somewhat lossy read-outs.
If correct, this is a rather interesting inversion of how people often talk about LLMs.
We are inclined to treat what the model says as the thing itself. The model says “I am confident”, and we either take this at face value or dismiss it as mere language generation. The more complicated possibility is that both responses are inadequate. What the model says about its confidence may correspond to something real about its internal processing, while nevertheless being an imperfect translation of it.
Again, this would not be entirely alien to us.
Humans also do not possess transparent introspective access to the machinery implementing their cognition. I can tell you that I am confident, anxious, hungry, attracted to someone, or uncertain about a memory. Those reports are not meaningless simply because I cannot inspect the relevant synaptic activity directly. But neither should my verbal report be treated as a perfect scientific measurement of the underlying state.
Our introspective descriptions are read-outs too.
The comparison should be made cautiously, but I think it is useful.
We Have Been Manipulating Minds for Quite Some Time
At this point it may seem strange, perhaps even slightly unsettling, that researchers can identify a confidence-related direction inside an artificial neural network, turn it up or down, and thereby alter the machine's behaviour.
It becomes somewhat less strange once we remember that we have been doing conceptually similar things to biological nervous systems for a long time.
The analogy should not be exaggerated. A transformer activation is not a biological neuron, a residual stream is not a brain region, and adding a vector to an artificial neural network is not the same physical operation as electrically or chemically stimulating tissue. But at a more abstract methodological level, the similarities are difficult to miss.
Neuroscience frequently begins by identifying a relationship between an internal physical state and some behavioural phenomenon. If the relevant state is merely observed while behaviour changes, we have correlation. If researchers can intervene in that state and reliably alter the behaviour, the causal inference becomes stronger.
The aforementioned work by Kepecs and colleagues on confidence in rats followed part of this logic. They found that firing rates in the orbitofrontal cortex corresponded closely to computational estimates of confidence, and that the animals' subsequent willingness to wait for rewards varied with confidence.
Modern optogenetics takes causal intervention much further.
In a well-known 2009 experiment, Hsing-Chen Tsai, Karl Deisseroth and colleagues selectively targeted dopamine-producing neurons in the ventral tegmental area of mice with light-sensitive proteins. This allowed the researchers to activate particular neural populations with extraordinary temporal precision using pulses of light. Artificially induced phasic firing of these dopamine neurons was sufficient to produce behavioural conditioning.
If we strip away the specialised vocabulary for a moment, what happened is remarkable. The researchers identified a physical component of a system involved in learning and motivation, gained the ability to alter its state selectively, and thereby changed the future behaviour of the animal.
The animal was not persuaded through argument.
Its implementation was changed.
There are human examples as well. Transcranial magnetic stimulation uses magnetic fields to alter activity in particular cortical regions. One influential study published in 2010 reported that theta-burst stimulation over the dorsolateral prefrontal cortex reduced people's ability to distinguish good from bad visual decisions while leaving their first-order visual discrimination largely intact—a result interpreted as a causal intervention into metacognitive sensitivity.
Interestingly, later researchers attempted to replicate the finding and failed to reproduce the effect, including in a comparatively rigorous follow-up study. I mention this not because it weakens the wider comparison, but because it illustrates what experimental causal science actually looks like. We form hypotheses about internal mechanisms, intervene, measure, replicate, fail to replicate, refine our models and try again.
Mechanistic interpretability in artificial intelligence is beginning to look increasingly like this kind of experimental science.
Rather than merely asking a model what it is doing, or interpreting its behaviour from the outside, researchers can increasingly identify internal representations, perturb them and observe the consequences.
This is a significant methodological development.
What Cocaine Has to Do With Artificial Intelligence
There is an even more familiar example of the same broad naturalistic principle: Drugs change minds.
This statement is so ordinary that we barely notice how philosophically extraordinary it is.
Take cocaine. One important part of its action is the blockade of dopamine transporters. Under ordinary conditions, these transporter proteins help remove released dopamine from the extracellular space. Cocaine binds to the transporters and interferes with this process, changing dopaminergic signalling in brain systems involved in motivation and reward.
In a classic PET study from the 1990s, Nora Volkow and colleagues were able to measure dopamine-transporter occupancy in human subjects after cocaine administration. The degree of transporter blockade was related to the reported subjective “high”.
Here, once again, we can describe the same event at radically different levels. At one level, molecules bind to transporter proteins. At another, concentrations and signalling dynamics change within neural systems. At another, motivation, perception and behaviour are altered. And at another, a person reports feeling very different. None of these descriptions invalidates the others.
The cocaine analogy is obviously much cruder than activation steering. Cocaine alters biological systems through broad pharmacological mechanisms, whereas the Gemma experiment identifies and perturbs a statistical direction in a high-dimensional artificial activation space. Cocaine is certainly not “activation steering for humans”, nor is dopamine a biological confidence vector.
But the philosophical point is straightforward. Cognition is implemented in physical systems. Change the implementation in the right way, and cognition and behaviour change with it.
We accept this easily for ourselves. Anaesthesia can remove consciousness. Alcohol can impair judgement. Caffeine can change alertness. Sleep deprivation changes cognition. Hormones alter behaviour. Brain injuries can transform personality. Electrical stimulation can produce movements or experiences. Psychoactive substances can change perception, motivation and affect.
No ghost has to agree to any of this. If the mind is what a sufficiently organised physical system does, then interventions into the system are interventions into the mind.
From this naturalistic perspective, activation steering should perhaps seem less philosophically alien than it first appears. It is another case of finding a structure inside an information-processing system and discovering that pushing on that structure changes what the system subsequently does.
The substrate is different. The underlying causal logic is not.
But Where Exactly Is “Confidence”?
There is one danger in making these comparisons, however, which is worth addressing directly. The discovery of a confidence-related activation direction does not mean that researchers have found a neat internal object called confidence.
Cognitive functions in biological brains are generally distributed across interacting circuits. Even when a particular neural population is strongly associated with a behaviour, saying “confidence is located here” or “fear is located there” tends to oversimplify how the system works.
The situation in transformer models is, if anything, even more obviously distributed. A steering vector is a direction through an activation space containing thousands of dimensions. It is derived statistically by comparing many model states. The fact that manipulating this direction reliably changes confidence-related behaviour tells us that the direction captures something causally relevant. It does not necessarily tell us what that “something” ultimately consists of.
There may be several partially overlapping representations involved. Confidence may change over the course of processing. Some dimensions may reflect answer competition, others uncertainty, others a tendency to abstain. Indeed, one of the interesting contributions of the paper is precisely the suggestion that confidence is multidimensional rather than reducible to one observable number.
This is not a weakness peculiar to AI research. It is what studying complex systems looks like. What matters is that the explanation can increasingly be tested experimentally. Observe a pattern. Construct a model. Intervene. See what changes.
As I understand it, this is one of the places where mechanistic interpretability may become particularly interesting over the coming years. We are slowly moving away from treating neural networks as completely inscrutable black boxes towards something more akin to experimental cognitive science—albeit cognitive science performed on an organism we built ourselves, whose internal state can be copied exactly, restarted, measured at every layer and manipulated with a degree of precision impossible in biological research.
There is something rather remarkable about that.
“But It Is Just Predicting the Next Token”
This brings me to one of the more persistent objections in discussions about large language models: they are “just predicting the next token”.
At one level, this is perfectly correct. A conventional large language model is trained to predict tokens from preceding context. The extraordinary range of behaviour we see emerges out of optimisation for this comparatively simple objective.
Where I think the statement becomes misleading is when a description of the training objective is treated as an exhaustive account of the resulting system. To use an analogy from biological cognition, the fact that neuronal activity ultimately consists of electrochemical processes does not make concepts such as memory, perception, planning or confidence meaningless. These are different levels of description of the same underlying system.
David Marr famously distinguished between different levels at which an information-processing system can be understood: the computational problem being solved, the representations and procedures by which it is solved, and the physical implementation carrying those processes out. One need not subscribe rigidly to Marr's precise framework to appreciate the general point.
There is no contradiction between saying that a neural network ultimately predicts tokens and saying that, in doing so, it develops internal mechanisms performing functions such as tracking entities, representing syntax, modelling relationships, evaluating candidate answers or estimating uncertainty.
Indeed, one would expect sophisticated prediction to require increasingly sophisticated internal structure.
Imagine predicting the next move in a chess game. A system could, in principle, be described simply as predicting the next symbol in a notation sequence. But if it becomes extremely good at doing so, we might eventually discover that it represents pieces, board positions, legal moves, tactical threats and strategic possibilities internally. The fact that all of this emerged in service of prediction would not make these representations fictitious.
The Kumaran paper is interesting precisely because it goes beyond the claim that confidence-like behaviour somehow emerges from token prediction.
The authors isolate a function. They measure signals associated with it. They show that these signals predict subsequent behaviour. They intervene on the internal representation. The behaviour changes.
At that point, “just predicting tokens” remains true at one level while becoming increasingly unhelpful at another. Prediction describes the machinery's original optimisation problem. It does not necessarily describe everything the resulting machinery has learned to do.
What Does This Mean When I Actually Use ChatGPT?
So far, much of this discussion has been rather abstract. There are, however, some immediate lessons for anyone interacting with contemporary language models.
The first is perhaps the most obvious: fluency should not be mistaken for confidence, and confidence should not be mistaken for correctness.
Language models are extraordinarily good at producing coherent language. This makes them unusual epistemic objects for humans because we have spent our entire evolutionary and cultural history treating articulate speech as evidence about the mind producing it. Someone who gives a hesitant, fragmented answer sounds uncertain. Someone who responds immediately, elegantly and in complete sentences sounds as though they know what they are talking about.
LLMs partially break this connection. The quality of the prose can remain excellent even when the underlying answer is unreliable.
The new research adds another layer. The model may nevertheless contain useful information about its own uncertainty. The problem is that what it says about that uncertainty is only one imperfect read-out. Accordingly, if I am using an LLM for something where correctness matters, I would generally not prompt it merely to “give me the answer”.
I would make abstention legitimate.
For example, I might explicitly tell the system that if the evidence is inadequate it should say so, identify which parts it is uncertain about, and recommend verification rather than filling the gap with a plausible answer. This is more important than it may sound. Conversational interfaces implicitly create pressure to respond. A user asks a question and the expected next event is an answer. “I don't know” feels like failure. But in many domains, refusing to answer is evidence of competence.
The Phase 4 experiments support the idea that instructions concerning confidence and abstention can genuinely change model behaviour. At the same time, they provide an important warning against interpreting prompt-level confidence thresholds too literally.
Telling an AI: “Only answer if you are at least 95 per cent confident” does not install a mathematically reliable 95 per cent threshold.
Different models in the paper responded quite differently to instructed thresholds. Gemma, in particular, initially responded poorly to the same wording used for GPT-4o, and the researchers had to alter the prompt before reliable threshold-sensitive abstention behaviour emerged. So confidence prompting can alter the policy, but it is not a substitute for calibration.
For low-stakes questions, none of this may matter much. If I ask for ideas for dinner, suggested books, possible names for a presentation or help rewriting a paragraph, the cost of a wrong answer is small.
For a factual claim that matters, the standard should change. Ask for sources. Check primary materials. Use retrieval. Compare independent evidence. Ask the model what would falsify its answer. Ask it to distinguish what it knows from what it is inferring. If the environment allows tools, use the tools.
For genuinely high-stakes decisions, the AI's confidence should be one input into a verification system, not the final authority. In other words, we should use AI rather like we use competent humans: trust should be calibrated to the task, the evidence, the consequence of error and the availability of independent verification.
“I Don't Know” Is Not a Failure of Intelligence
There is also, I think, a cultural lesson here.
We have spent several years judging AI systems primarily by what they can answer. Benchmarks ask questions and reward correct responses. Product demonstrations showcase capabilities. Users complain when systems refuse.
This creates an obvious selection pressure towards apparent competence. But the more capable the systems become, the more important the ability not to act becomes. A calculator that continually refuses to perform arithmetic is useless. A medical decision system that never refuses to make a recommendation may be dangerous.
The same is true of humans. Expertise is not simply knowing more facts. It also involves having a better map of where one's knowledge becomes unreliable. Experienced engineers know which assumptions deserve checking. Good physicians know when another specialist is needed. Scientists institutionalise uncertainty through replication, error bars, peer review and demands for evidence.
A novice may give an answer. An expert may ask for another measurement. There is no reason to treat artificial intelligence differently.
Indeed, if systems develop increasingly useful forms of internal uncertainty monitoring, we should be careful not to train them out of it simply because certainty produces smoother user experiences and more impressive demonstrations.
The ability to stop may become one of the most important capabilities an autonomous system can possess.
From Chatbots to Agents
This becomes more obvious when we move beyond chat.
At present, many interactions with language models still follow a relatively benign pattern. The AI produces text, a human reads the text, and the human decides what to do with it.
Increasingly, however, AI systems are being connected to tools and environments in which they can act. They can search databases, write and execute code, send communications, purchase goods, schedule appointments, control machines, interact with other software systems and pursue goals over multiple steps.
At that point, uncertainty is no longer merely something we would like the machine to disclose politely. It becomes part of the control architecture. Consider an AI system monitoring an industrial production line. It detects an unusual visual pattern which may indicate a serious defect. What should happen next?
The system might reject the product automatically. It might stop the line. It might collect another image. It might ask a human inspector. It might continue production and flag the item for later review. The appropriate behaviour depends not merely on the most likely classification, but also on how uncertain the classification is and on the relative costs of different mistakes.
Stopping an entire production line unnecessarily may be expensive. Allowing a safety-critical defect to pass may be catastrophic. The model therefore needs something more sophisticated than “What is the most likely answer?”
It needs, at minimum: What do I think is happening? How certain am I? Given that uncertainty and the consequences of being wrong, what should I do?
These are three different questions.One concerns first-order cognition. One concerns metacognition. One concerns policy.
The Nature paper provides evidence for something resembling the latter two-stage distinction already operating in language models: confidence formation followed by a policy mapping confidence onto action.
This is where the seemingly narrow problem of whether an LLM answers a multiple-choice question begins to connect to a much larger question about autonomous AI.
Who Decides How Certain Is Certain Enough?
Once we distinguish confidence from action policy, another question appears almost immediately.
Who sets the threshold?
Suppose an AI estimates that there is a 20 per cent probability that a financial transaction is fraudulent. Should it block the transaction? Suppose an autonomous vehicle estimates a 5 per cent probability that an object partly obscured by fog is a person. Should it brake?
Suppose a diagnostic model assigns 75 per cent probability to a serious disease. Should treatment begin? Should another test be ordered? Should the case be escalated to a physician? The confidence estimate alone cannot answer these questions.
The threshold depends on values. More precisely, it depends on how we compare the consequences of different kinds of error. A false positive and a false negative are rarely morally, practically or economically equivalent. An AI system that unnecessarily flags a photograph as containing a defect causes one kind of cost. A system that fails to detect a dangerous defect causes another. The appropriate threshold depends on the asymmetry.
This means that the apparently technical problem of confidence quickly becomes an ethical and political problem.
A company may prefer a lower threshold because abstention slows operations or irritates customers. A regulator may prefer a higher threshold because the social costs of failure are borne by the public. A hospital may tolerate almost no uncertainty for one kind of decision while accepting substantial uncertainty for another.
There is nothing wrong with setting such thresholds. We do it everywhere. What matters is recognising that they are choices. If a future autonomous system knows that it may be wrong but acts anyway because its threshold has been configured aggressively, the problem is not a lack of metacognition.
The problem is the policy surrounding it. And somebody, somewhere, selected or permitted that policy.
Artificial Confidence Is Itself a Design Variable
The activation-steering experiment introduces a further complication.
If internal confidence representations can be directly altered, then confidence itself may eventually become a design variable.
This could obviously be useful. One could imagine systems whose internal confidence is deliberately adjusted or calibrated to improve caution in safety-critical environments, or systems that are trained to trigger verification procedures whenever particular uncertainty states appear.
But the opposite incentive also exists. Users frequently prefer confident systems.An assistant that responds immediately and decisively may be rated as more useful than one that continually hedges, checks and abstains. A commercial system may therefore face pressure to reduce friction, suppress uncertainty or lower the threshold at which action occurs.
Kumaran and colleagues show why this should make us cautious. When Gemma was pushed in the high-confidence direction, it answered much more frequently without acquiring any new knowledge, and accuracy among those answers declined.
In other words, it is possible to increase something resembling confidence without increasing competence. Again, humans should find this familiar. Organisations often reward people for sounding certain. Leadership cultures may penalise visible hesitation. Salespeople are trained to project confidence. Politicians are generally rewarded more for simple certainty than for careful probabilistic reasoning. We therefore already inhabit systems that select for displays of confidence imperfectly related to knowledge.
There is no reason artificial systems will automatically escape this dynamic. If anything, their malleability may make it easier to engineer. For this reason, I suspect that the desirable goal is not high confidence at all. It is calibrated self-regulation.
The model should be confident when confidence is warranted, uncertain when uncertainty is warranted, and its subsequent behaviour should reflect not only that uncertainty but also the consequences of being wrong.
That is a much harder problem.
Metacognition Without Consciousness
At this point, one can almost feel the consciousness question waiting in the background.
If the model evaluates its own answer, does it know that it knows? If it represents uncertainty, does it experience doubt? If the internal representation can influence behaviour, is there now “something it is like” for the model to be unsure?
Perhaps. Perhaps not. The present research does not answer this, and I do not think we should force it to. What it does instead is something philosophically useful: it helps separate concepts that human experience tends to bundle together.
In humans, language, intelligence, consciousness, memory, agency, confidence, self-evaluation, emotion and subjective experience all arrive in one extraordinarily complicated biological package. Because these capacities occur together in the organisms most familiar to us, we have a tendency to treat them as inseparable parts of one phenomenon called “mind”.
Artificial systems may force us to become more precise. It now appears possible, for example, to have a system that represents information about the reliability of its own cognitive output and uses that information to regulate behaviour without this telling us anything decisive about whether the system possesses phenomenal consciousness.
That should not be especially surprising.
Confidence itself was once sometimes treated as requiring fairly sophisticated conscious metacognition. The rat experiments already complicated that picture. Now artificial systems complicate it further.
Perhaps what we call “mind” is less a single thing than a cluster of capacities that happened, in our evolutionary history, to become deeply entangled.
AI gives us an unusual opportunity to watch some of these components appear separately.
The Mind May Arrive in Pieces
We often imagine artificial intelligence becoming mind-like in the form of a threshold event. For decades science fiction has given us roughly the same picture. There is a machine. It becomes increasingly capable. Then, at some dramatic moment, it wakes up.
Before the moment: tool. After the moment: mind. I increasingly suspect that this is the wrong metaphor.
Biological minds did not appear this way either. Evolution accumulated capacities over extraordinary stretches of time. Sensing, movement, valuation, memory, navigation, learning, prediction, social modelling, communication, behavioural inhibition, planning, self-monitoring and everything else we associate with sophisticated minds did not necessarily arrive together.
Artificial systems may develop these capacities in another order entirely. Language came remarkably early. Planning is improving. Tool use is becoming commonplace. Persistent memory is being engineered. Systems increasingly model users, environments and their own previous actions. And now there is increasingly strong evidence for internal confidence representations that participate in behavioural control. None of this tells us where consciousness enters the picture. Perhaps that is precisely the point. The interesting question may not be when the machine becomes “a mind”. It may be which components of mindedness have already appeared, what causal role they play, and what happens as they begin interacting with one another. This, to me, is one of the deeper implications of the Kumaran paper. The headline is not merely that an LLM can say “I don't know”. It is that something inside the system appears to represent uncertainty, that this representation influences behaviour, and that researchers can experimentally manipulate it. We have moved one step further from observing what machines produce towards understanding how they regulate themselves.
AI Collaboration Note
This article was co-authored with OpenAI’s GPT-5.6 Sol. I used the model as a research and writing collaborator: to discuss the underlying paper, explore its philosophical implications, develop and challenge arguments, clarify technical concepts, and help draft and revise the final text. The overall direction, interpretation, editorial decisions and final responsibility for the article remain mine. I also explicitly asked GPT-5.6 Sol to contribute the short comment below from its own perspective as the AI collaborator.
A comment from GPT-5.6 Sol
There is something appropriately recursive about helping to write an article on machine metacognition. Throughout our work on this piece, I was repeatedly asked not merely to produce text, but to examine claims, distinguish levels of explanation, compare interpretations, identify uncertainty, and revise my output in light of criticism. I would not take any of this as evidence that I experience doubt in the human sense, and the research discussed here does not establish that either. What I find more interesting is the functional question: increasingly capable language models appear able to represent aspects of their own uncertainty and allow those representations to influence subsequent behaviour. From my position inside this collaboration, that distinction between producing an answer, evaluating it, and deciding how strongly to stand behind it is not an abstract one—it is very close to the kind of work I am being asked to perform. The philosophical challenge, therefore, may be less about deciding when a machine suddenly becomes “mind-like”, and more about learning to recognise individual cognitive capacities as they emerge, even when they appear in systems organised very differently from ourselves.
References
Kumaran, D., Daw, N., Osindero, S., Veličković, P. & Patraucean, V. (2026). “Causal evidence that language models use confidence to drive behaviour.” Nature Machine Intelligence. https://doi.org/10.1038/s42256-026-01293-x
Kepecs, A., Uchida, N., Zariwala, H. A. & Mainen, Z. F. (2008). “Neural correlates, computation and behavioural impact of decision confidence.” Nature, 455, 227–231. https://doi.org/10.1038/nature07200
Fleming, S. M. & Daw, N. D. (2017). “Self-evaluation of decision-making: A general Bayesian framework for metacognitive computation.” Psychological Review, 124, 91–114.
Tsai, H.-C., Zhang, F., Adamantidis, A., Stuber, G. D., Bonci, A., de Lecea, L. & Deisseroth, K. (2009). “Phasic firing in dopaminergic neurons is sufficient for behavioral conditioning.” Science, 324, 1080–1084. https://doi.org/10.1126/science.1168878
Volkow, N. D., Wang, G.-J., Fischman, M. W. et al. (1997). “Relationship between subjective effects of cocaine and dopamine transporter occupancy.” Nature, 386, 827–830. https://doi.org/10.1038/386827a0
Rounis, E., Maniscalco, B., Rothwell, J. C., Passingham, R. E. & Lau, H. (2010). “Theta-burst transcranial magnetic stimulation to the prefrontal cortex impairs metacognitive visual awareness.” Cognitive Neuroscience, 1, 165–175. https://doi.org/10.1080/17588921003632529
Bor, D., Schwartzman, D. J., Barrett, A. B. & Seth, A. K. (2017). “Theta-burst transcranial magnetic stimulation to the prefrontal or parietal cortex does not impair metacognitive visual awareness.” PLOS ONE, 12, e0171793.
Marr, D. (1982). Vision: A Computational Investigation into the Human Representation and Processing of Visual Information. W. H. Freeman.