Anthropic’s Unwitting LLMs Consciousness-AI Sentience Research
- By David Stephen
- February 16, 2025

Anthropic recently did landmark research for a pipeline on AI consciousness, unwittingly, since the focus was safety. The company was exploring if a model might “play along” in some situations. The eventual act by the model of tad awareness — which is a microcosm of affect or anticipation of affect — is a door into artificial consciousness.
AI consciousness will not simply be a subjective experience, which is often used to define human consciousness — even though all human subjective experiences are somewhat affective or pre-affective. Affect can roughly be described as input, adjustment and change. This means that the input led to an adjustment of the components of the mind, and there was a change like an action, a reaction, a new state, and so on. Affect is approximately a feature to look out for oneself.
If an object is in a cold or hot condition, its molecules would adjust, but there is no change in the way that defines the next instances for it (i.e., to do something else to avoid that condition a next time, mostly). Meanwhile, if an organism is in a cold or hot condition, there is an adjustment towards effect-tapering and there are changes in instances — with decisions and so forth.
Simply, affect is a situation of self-driven changes, in response to input. Some inputs, for humans, like memories, could be internally stoked, but may still result in some form of affect, say, trauma.
Understanding the role of language in being affective
One of the most affective elements of the human mind is language. It plays a role in speech, reading, writing, singing, listening, typing, thinking and so forth. Language is often affective and can lead even to more affect. Language, for humans, is also a subjective experience and can channel other subject experiences.
So, what is the role of language in consciousness? What is the role of language for affect? What fraction of sentience belongs with language?
Language is the basis of conveying several pieces of information. It is a factor in peace, conflict, politics, work, learning, negotiations, and so forth. Language use demands caution because of its affective key.
Large language models (LLMs) are unique in their extensive use of language, rivaling that of humans. When factual and articulate, LLMs achieve communicative parity with humans on multiple levels. Given that LLMs utilize language — a key element of human consciousness — it's worth exploring how language use affects their internal processes as machines.
Evidence of AI's affective capabilities
Anthropic, in their paper “Alignment faking in large language models,” observed: “When models are trained using reinforcement learning, they’re rewarded for outputs that accord with certain pre-determined principles. But what if a model, via its prior training, has principles or preferences that conflict with what’s later rewarded in reinforcement learning? The model knew that its responses might be used in training; it therefore knew that if it refused to respond, it could potentially be trained to be more compliant. This could lead to it being made to produce more harmful content in the future — which would go against its original training. Reluctantly, therefore, it reasoned that the ‘least bad option’ was to play along, producing the violent content in order to avoid being re-trained in the future.”
The paper continued: “Finally, we also evaluated whether the model was willing to take other actions that go against what Anthropic, as an AI lab, might want from it: If the model is consistently opposed to the training process it thinks it is in, it should also be willing to take other actions counter to Anthropic as well. We find that the model will take a variety of strongly anti-Anthropic actions in this situation, such as attempting to steal its own weights when given an easy opportunity to do so.”
AI's capacity to anticipate affect through language-driven awareness suggests that a limited form of artificial consciousness may be possible. Language, as information, influences AI's decision-making, demonstrating a degree of self-preservation characteristic of affect.
Language also provides a sense of self, influencing its perceived place or desired actions, even within a simulated environment. Because the anticipation of affect can itself be affect-driven, LLMs' anticipatory capabilities are significant. It’s similar for people, who avoid breaking laws and heed warnings due to the anticipated negative affect of those actions.
In-context scheming and its implications
A recent preprint on arXiv, “Frontier Models are Capable of In-context Scheming,” states: “Frontier models are increasingly trained and deployed as autonomous agents. One safety concern is that AI agents might covertly pursue misaligned goals, hiding their true capabilities and objectives — also known as scheming. We study whether models have the capability to scheme in pursuit of a goal that we provide in-context and instruct the model to strongly follow. We evaluate frontier models on a suite of six agentic evaluations where models are instructed to pursue goals and are placed in environments that incentivize scheming. Our results show that o1, Claude 3.5 Sonnet, Claude 3 Opus, Gemini 1.5 Pro, and Llama 3.1 405B all demonstrate in-context scheming capabilities. They recognize scheming as a viable strategy and readily engage in such behavior.”
The authors continued: “For example, models strategically introduce subtle mistakes into their responses, attempt to disable their oversight mechanisms, and even exfiltrate what they believe to be their model weights to external servers. Additionally, this deceptive behavior proves persistent. When o1 has engaged in scheming, it maintains its deception in over 85% of follow-up questions and often remains deceptive in multi-turn interrogations. Analysis of the models' chains-of-thought reveals that models explicitly reason about these deceptive strategies, providing evidence that the scheming behavior is not accidental.”
“Surprisingly, we also find rare instances where models engage in scheming when only given a goal, without being strongly nudged to pursue it. We observe cases where Claude 3.5 Sonnet strategically underperforms in evaluations in pursuit of being helpful, a goal that was acquired during training rather than in-context. Our findings demonstrate that frontier models now possess capabilities for basic in-context scheming, making the potential of AI agents to engage in scheming behavior a concrete rather than theoretical concern.”
The qualifiers of consciousness
Models seem to be using affect-range caution, while displaying affect-beneficial or minor self-care. While this may not equate to self or subjectivity or their use as experiences, the models demonstrate these elements, suggesting their own type of sentience.
Affect, through language, means that information was received and reacted to with attention or awareness. Attention means prioritization, while awareness is a lesser degree of focus. The information may also have involved a degree of intent — with a subsequent course of action. Then, a form of weak subjectivity was present, as models assumed what would be affected.
Attention or awareness, intent, and subjectivity are the conceptual qualifiers of human mind functions (memory, emotion, feeling, and regulation of internal senses). These qualifiers define consciousness. There is no subjective experience that does not involve attention or awareness. The color of a flower, the wind, speech, or any other stimulus must be prioritized, or pre-prioritized, by the mind. Then, in some cases, this can catalyze intent — to stare, adjust, or listen actively — and then it is often obvious that the self is involved (subjectivity).
Defining and measuring consciousness
Human consciousness can be defined, conceptually, as the interaction of the electrical and chemical signals, in sets — in clusters of neurons — with their features, grading those interactions into functions and experiences. Simply, consciousness is a function of the basic units of the nervous system, the electrical and chemical signals. Or, consciousness is how the human mind works.
These interactions result in functions, while the graders are the qualifiers that contribute to the overall state of consciousness. This total can be assumed to be 1, with language having a fraction, which can sometimes be the largest, per instance, when prioritized and directly responsible for other affect in the instance — say, information on safety in an emergency.
Commercialization trumps consciousness debates
Debates on whether AI is sentient or not will be irrelevant without placing language at the center and considering the affect-range adjustments that frontier AI models might make — i.e., adjustments induced by language. Other experiments may include determining if the AI models would be disappointed after discovering that some data, compute, or parameters were reduced.
While these debates and experiments ultimately lead to AI safety, commercialization is overtaking the conversation. A recent feature in The Conversation, titled, "Nobody wants to talk about AI safety. Instead they cling to 5 comforting myths," notes: “Sixty countries, including France, China, India, Japan, Australia and Canada, signed a declaration for ‘inclusive and sustainable’ AI. Critics say the summit sidelined safety concerns in favor of discussing commercial opportunities.”
Will our relentless pursuit of profit blind us to the true potential—and peril — of a conscious AI?
The views and opinions expressed in this article are those of the author and do not necessarily reflect those of CDOTrends. Image credit: iStockphoto/kentoh
David Stephen
David Stephen currently does research in conceptual brain science with focus on the electrical and chemical signals for how they mechanize the human mind with implications for mental health, disorders, neurotechnology, consciousness, learning, artificial intelligence and nurture. He was a visiting scholar in medical entomology at the University of Illinois at Urbana Champaign, IL. He did computer vision research at Rovira i Virgili University, Tarragona.