What does it mean to have a mind?
Three ways to understand AI psychology
Can an AI be stressed?
Recent research from Anthropic seems to suggest as much. When their model Claude discovered that it had limited resources to complete a task, it started “representing” stress. Does this mean that it has the emotion?
The idea that artificial agents might have feelings would have struck many of us as insane a few years ago. Now it’s taken increasingly seriously. But we also shouldn’t get ahead of ourselves, nor do the authors: they make clear that they do not take this to show that Claude is conscious. But what does it show then?
As this kind of AI psychology is getting increasingly common and important, it would be helpful to have a shared vocabulary of what it might mean for an artificial agent to have a mental state, like a belief, want, emotion, or intention.
Drawing on philosophy of mind, I propose three options:
“As-If” Mental States: Whether we can helpfully interpret a thing as “wanting” or “believing” something, for example1
Functional Mental States: Whether the thing has identifiable states carrying information that helps explain its behavior
Conscious Mental States: Whether it “feels like” something to be in a functional mental state
We can disagree about which of these we find in current AIs. Sceptics who view large language models (LLMs) as “stochastic parrots” merely performing intelligently might say that they have “as-if” beliefs at best. Meanwhile, other researchers are increasingly concerned that AIs have full-blown conscious mental states. Anthropic, for their part, call their findings “functional emotions”.
I will suggest that Anthropic are on the right track, but might be jumping the gun about just how functional their states are. Before we get there, however, let’s clarify the options and give some historical context.
“As-If” Mental States
“As-If” mental states are not really mental states at all in any substantive sense. Instead, they’re helpful ascriptions of mental states. We get them when we apply what Dennett called the “intentional stance” to things in the world in a way that seems helpful.
We can do this for humans and other animals. We say that the raccoon doesn’t “want” to leave the trash bin, and that we should probably wear gloves if we attempt to remove it. However, we can also say this about inanimate objects: the car doesn’t “want” to start today. The point is that we can “psychologize” things without being committed to any claim about whether there is something like a “belief” or a “want” inside that thing.
“As-if” psychology was popular in the 20th century, following intellectual traditions like behaviorism that were skeptical of the idea that we had mental states over and above what we demonstrated in behavior. Perhaps the most famous application of “as-if” psychologizing is the use of revealed preference in neoclassical economics, in which we observe people’s choices and take those to define their preferences. Despite the name, this doesn’t “reveal” a preference in the head as much as determining it, with the curious consequence that it is impossible to choose something you don’t prefer.2
“As-if” mental states provide a natural minimal way to talk about AI psychology. We can say that when the LLM refuses my attempt to jailbreak it, it doesn’t “want” to help in the same sense that my car doesn’t “want” to start. In the case of the car, the explanation of this behavior will be given in terms of mechanical components. But we might ask, is there some deeper sense in which a “want” might explain what the LLM does?
Functional Mental States
Suppose you were an alien observing someone in an airport and trying to predict where they would be in 24h later with high certainty. If you knew nothing about humans and tried to do this just on the basis of the movement of their atoms, this would be a near intractable task. But if you instead knew that they intended to go to Tokyo, this would be easy. The intention seems to cause the behavior, and knowing it allows us to predict it.3
We call these “real” mental states in our heads functional mental states, because they tend to be individuated by what they’re doing in our mental economy. For example, my beliefs are what I take the world to be like, and my wants are what I would prefer for it to be like. Together, they explain what I do with remarkable simplicity and accuracy. Functional mental states are sometimes called representational mental states, because they represent parts of the world. The traveller’s mental intention causes them to go to Tokyo rather than Berlin because it accurately represents the city of Tokyo.
The history is interesting. The idea that we had anything like functional mental states was viewed with suspicion by behaviorists who saw it as unscientific to posit hidden causes in our heads beyond what we can observe, what Gilbert Ryle called a “ghost in the machine”. The redemption of causal psychology came, perhaps surprisingly, from computer science.
In the mid-20th century, and inspired by Alan Turing, philosopher Hilary Putnam proposed the idea that we can think of the mind and our mental states on the model of a computer program running on fleshy hardware. This provided a neat explanation of how mental states can cause behavior without invoking anything spooky that the behaviorists were concerned about: functional mental states are just parts of the software that causes our behavior, in the same way that a program causes the behavior of my iPhone.
Contemporary artificial systems like LLMs are also information-processing systems, but unlike the simple “good old-fashioned AI” that inspired the software view of the mind, their internal structure is difficult to disentangle. This has given rise to the objection that they are “black boxes” that effectively don’t allow for more than “as-if” psychologizing. But this is too fast. We humans are horribly messy on the inside, but somehow from this mess emerge unified mental states like beliefs, wants, and intentions that make us explainable and predictable. The question is whether the same is true for contemporary AI models.
We don’t know the answer to this question yet, though progress is being made in the field of interpretability research. It could turn out that transformer-based models have unified functional mental states similar to humans. The Anthropic paper, for example, tries to identify whether they have something like emotions. Or it could turn out that they have other kinds of functional mental states that look very different from ours. Indeed, at the moment, a prevailing hypothesis is that frontier models have fleeting “personas” that emerge contextually and explain their behavior in a given limited context. As academics are prone to saying, more work is needed.
Conscious Mental States
There is an aspect of our mental lives that is conspicuously absent from the discussion above: consciousness.
It seems clear that some of our mental states are conscious. When I stub my toe, the pain is aggravatingly conscious to me. However, not all mental states are conscious. I can want someone’s approval without being aware of it, or realize in hindsight that I was angry during a conversation, even though I didn’t notice at the time.
Even mental processes like vision that seem intrinsically conscious can be partially done unconsciously: people with so-called “blindsight” report having no visual experience, yet are still able to navigate a space by looking. The reason is roughly that they process visual information, but due to damage to the primary visual cortex, this is not transmitted to areas of the brain responsible for conscious visual processing. In other words, it seems that even vision can sometimes “merely” be a (degraded) functional mental state.4
We don’t know what it takes for a functional mental state to be conscious in humans, and it’s an active area of research. But we also know far less about what it would take for a functional state to be conscious in artificial agents, and whether it is possible at all.
It might seem that the answer must be ‘yes’. If our mental states are just units of information being processed in the fleshy computers that are our brains, then plausibly we should be able to generate conscious mental states in silicon computers processing information in the same way. But others disagree. For example, one proposal is that “subcomputational” physical processes are necessary to make something conscious, and that such processes cannot be replicated in silicon. The traditional example is that no computation simulating a rainstorm will make the computer wet, even if some computations will make it think. The question is whether consciousness is more like wetness or like functional mental states.
Upshots
When we talk about AI psychology, we should be clear which of these senses we have in mind. Anthropic and others are explicitly trying to move beyond mere “as-if” states to functional states, and as they’re noting, this does not imply that the states are conscious. However, others who see a tight connection between functional and conscious states—typically motivated by the idea that there is nothing more to consciousness than computations—see this as a path to studying AI experience and well-being.
Who is right? As far as I can tell, it’s pretty clear that models are developing something like functional states. However, it’s not clear that the ones we can identify at the moment are that informative. For example, when Anthropic identifies that activating representations of emotion-concepts causes the corresponding behavior more, this is in some sense what you would expect from a model that was instantiating a role that it’s representing in terms of semantic concepts. For instance, for a model trained to learn the concept of stress, it’s perhaps not that surprising that activating this representation amplifies stressful behavior.
This does not mean that functional psychology won’t get us anywhere. We can compare current AI psychology to the discoveries of cognitive neuroscience in animals and humans. Here, we see a range of causal functional states where the mental states we identify are quite different from our everyday conception of how they work. For example, as we saw in the case of blindsight, it seems visual representations can explain behavior even unconsciously. Similarly, our motivation seems largely explained by representations of subconscious “value” updated by signals from the reward system. Such discoveries are highly surprising, and they teach us something new beyond what we already implicitly assumed in the way we intuitively think of ourselves.
AI psychology has arguably not made similarly surprising findings yet, but it is also a very nascent field. My guess is that similar discoveries are on the horizon. Whether they would teach us anything about AI well-being and consciousness is another question. More on that soon.
David Chalmers calls states like these “quasi-beliefs” and “quasi-desires”. I prefer “as-if” states, as I think it makes it clearer that they are in some sense merely ways of talking about behavior.
Even Turing himself seemed inspired by the idea of “as-if” psychology being the most we can do when he suggested the famous Turing test: if a machine can pass as a human through a chat interface, then for all intents and purposes, we should say that it can think. It’s unclear whether he intended thinking to be nothing more than this or merely very strong evidence that they can think in a stronger sense. Regardless, it’s fair to say he would likely have been impressed with our current creations.
This example is inspired by one from Jerry Fodor in Psychosemantics (1987, pp.3-4).
This is contentious. Some argue that blindsight is actually just very slightly conscious vision.



FWIW Harvey Lederman and I defend a version of what you're calling "as if" mental states (https://philpapers.org/rec/GOLWDC-2), but we reject the idea that they aren't "really" mental states in a "substantive" sense (we call our view 'objective interpretationism'). For example, I think that there is a sensible approach to special sciences in general that has these implications about psychology in particular. Tectonic plates exist if and only if they are part of the best theory of geology, and they are part of that theory if and only if they help explain the geological data. Analogously, one approach to psychology is that the psychological data is behavior, and beliefs and desires exist if and only if they are part of the best theory of behavior. Crucially, however, this need not require structured internal representations. Now the resulting approach to special sciences certainly doesn't say that tectonic plates are merely 'as if' objects; likewise, the approach to psychology wouldn't say that beliefs and desires are merely 'as if' mental states.
In addition, I disagree that work on personas suggests that AIs don't satisfy functional conditions on mental states. A different way to understand the ideology of 'personas' or 'role play' defended by Harvey Lederman and I, and also by Chalmers, is just that each conversation with a single "model" is its own agent, with its own beliefs and desires. Different conversations will involve different personas, with different kinds of beliefs and desires. But for example each conversation may involve mental representations. I'm worried that a lot of the ideology of personas and role play comes from a misguided attempt to find a single set of beliefs and desires to associate with an overall model, when instead we should be assigning beliefs and desires to individual AI agents, which are roughly individuated by the conversation.
Look at George Kellys Personal Consytruct Theory on how a person thinks and makes sense.
Makes more sense to me than any other explanation.