Anthropic has announced the discovery of a "J-space" within large language models, a region where information entering it triggers a conscious access effect. This research has generated significant interest alongside considerable debate. These consciousness-related functionalities are not a product of training but emerge naturally; they have been identified and preliminarily validated in models like Claude and Qwen.
The discovery of J-space presents its first potential application in model safety alignment. Humans could easily monitor this J-space and potentially directly instill ethical principles into it.
As stated by Anthropic's model psychology research team, this study draws inspiration from psychological theories concerning the "Global Workspace" (GWS) or "Global Neuronal Workspace" (GNW). Therefore, to understand Anthropic's paper titled "Verbalizable Representations Form a Global Workspace in Language Models," one must first grasp the GWS theory from cognitive science.
Two French neuroscientists, Stanislas Dehaene and Lionel Naccache, proposed this "cognitive neuroscience framework for consciousness" in 2000. Anthropic designed a tool called the Jacobian lens and, based on the GWS hypothesis, conducted preliminary validation of the J-space, comparing it to the human brain's workspace.
Dutch cognitive neuroscientist Bernard Baars introduced this theory in 1988: the brain comprises a set of specialized, largely independent modular processors. Vision, language, and motor control rely on fast, parallel, and closed neural circuits. He further proposed the global workspace hypothesis, suggesting the evolutionary purpose of conscious access is to break this modularity, connecting these processors to share their expertise and flexibly combine them for new tasks.
From this perspective, conscious processing is a function that temporarily selects a piece of information and broadcasts it globally to all receiving processors, enabling any processor to read it and act upon it. In humans, this broadcast reaches processors involved in language generation, explaining why reportability—the ability to verbalize a thought—is a key diagnostic feature distinguishing conscious from unconscious representations.
Dehaene and Naccache, among others, found preliminary evidence for a network of pyramidal neurons with long-range axons distributed throughout the brain, denser in the prefrontal, parietal, and higher temporal cortices. This network amplifies and sustains a selected representation, sharing it across the cortex. Functionally, this can be termed access consciousness: being aware of something means the information has entered this workspace and is available for report, reasoning, and flexible control (C1).
This view is now supported by more empirical evidence, including established neurobiological markers of conscious access. The first marker is ignition: when a stimulus crosses the threshold into consciousness, the corresponding neural activity later, around 250 milliseconds, suddenly and nonlinearly bifurcates into a sustained, widely distributed neural state involving the prefrontal cortex and many interconnected circuits, also amplifying the original circuits that extracted the information. In contrast, a subliminal stimulus only elicits a limited burst of activity in specialized circuits, quickly fading.
When a stimulus is precisely at the threshold, the same stimulus can produce a bimodal distribution of responses across trials, as if the brain suddenly tilts in one direction or another.
The second marker is limited capacity: the workspace acts as a bottleneck, attending to only one representation at a time. This property explains why attending to one process prevents awareness of another, such as missing a person in a gorilla suit during inattentional blindness, or delays perception of another process by hundreds of milliseconds, as in the psychological refractory period. To route information flexibly and appropriately, the system must maintain a model of its own capacity. This second property is self-monitoring (C2).
It must detect its own states, assess their likelihood of achieving goals, detect errors, and model what it knows and does not know. It must be able to report all these properties to itself, forming an internal self-report that does not necessarily lead to explicit behavior.
In summary, consciousness is not a separate soul entity or a little theater in a specific brain region, but an information access structure. The J-space paper does not strongly assert the proposition that "models are truly conscious"; instead, it attempts to validate the structural idea that there may be a small, globally readable workspace within the system. Contents within it are more easily reported, flexibly reasoned upon, and used across tasks, while a large amount of automatic processing can occur outside this space.
Anthropic's head of model psychology, neuroscientist Jack Lindsey, led the team conducting this research. Two months before the paper's release, they had multiple discussions with Dehaene and Naccache while the research report was still evolving; some changes were in response to their questions.
After the paper's release, Dehaene and Naccache authored a commentary, describing it as "a landmark in consciousness research." The following section translates key parts of their commentary, including an introduction to and critique of the J-space paper.
Core Findings of J-space
Inspired by the GNW hypothesis, Gurnee (the paper's first author) and colleagues sought to find verbalizable representations within a large language model like Claude Sonnet 4.5—the same reportability criterion used to probe human consciousness. In any layer of the model, verbalizable representations are vectors of activity across units that encode information tokens the model is prepared to report when asked: the model does not necessarily explicitly produce them, but it can.
To identify such reportable representations, they developed an elegant tool: the Jacobian lens. For each layer, it measures the average causal impact of an internal activation on the model's final output token across a vast number of different contexts. The mathematical metric captures activations that are, effectively, the representations the model tends to verbalize.
The concept of "average" is methodologically central: it distinguishes representations genuinely in a state of report-readiness from those that only accidentally leak into the output in a specific context. This set of representations is called the J-space. It explains less than 10% of the variance in any given layer but possesses remarkable properties.
The authors initially identified J-space representations solely based on the reportability criterion. They subsequently discovered these representations function far beyond supporting report; they act as an internal workspace decoupled from immediate input-output contingencies. For example, when the model is asked to "keep a concept in mind" while performing another computation, like "write sentence X while calculating 3²−2," the J-space contains the unreported concepts—first 9, then 7. The J-space carries hidden intermediate values in multi-step internal reasoning.
As Jack Lindsey explained to us, they were looking for reportable representations but found the same representations became globally available to the rest of the network during flexible reasoning, thus satisfying our C1 criterion for machine consciousness: global availability.
Crucially, the J-space is selective. It contains only a small fraction of what the model represents—the high-level information needed for flexible information processing. All other information used only for routine tasks does not seem to enter the J-space. For instance, studies show LLMs continuously track the character count of each word and the total characters in a line, as this is crucial for predicting if the next token should be a line break. However, such routine information does not enter the J-space unless an explicit task requires accessing it.
In an experiment neuroscientists can only dream of, the authors swapped "conscious content": they read a concept from the J-space, replaced it with another, and observed how the model's reasoning and reports changed. Remarkably, consistent with the GNW hypothesis, only high-level non-routine behaviors were affected, while routine tasks remained unchanged.
For example, when the model reads Spanish text, the J-space identifies the language as Spanish even if the task doesn't require reporting it. Replacing this J-space representation with another, like French, causes the model to fail in explicit language report: when asked the text's language, it answers "French" instead of "Spanish." The altered model also errs in other high-level inferences: the Spanish word for "hello" changes from "Hola" to "Bonjour"; the pre-Euro currency changes from "Peseta" to "Franc." However, this replacement does not affect its automatic ability to predict the next word: Claude continues writing in Spanish even after the intervention.
After large-scale ablation of all top-layer J-space representations, most of the model's basic abilities remain intact, but tasks requiring flexible reasoning are selectively impaired.
Several results in the J-space paper appear to us as direct analogues of human conscious access. When a task demands it, the model can selectively bring a property that normally wouldn't enter the J-space into it, such as the fact that the next word should be an adjective. As mentioned, automatic parameters needed for accurate next-token prediction, like character count per line, typically do not appear in the J-space; but when a task requires the model to access and manipulate them, these parameters get encoded into the J-space. This nicely demonstrates how the same information can transition from an automatic to an accessible state on demand.
Importantly, access to the J-space is also limited. In humans, there is a genuine form of introspection, but it is largely limited to slow, serial computation. In many well-documented situations, we develop fictional explanations for our mental processes. This dissociation between "how we act" and "how we think we act" is evident in choice blindness and visual illusions affecting conscious perception and verbal report but not necessarily motor gestures.
While current literature is insufficient to fully prove it, a similar dissociation seems to exist in the J-space. Indeed, prior work by the same team showed that when an LLM is asked to perform addition, what it reports verbally has little to do with how it actually obtains the result.
The authors also show the J-space exhibits structural markers of a workspace: it is primarily located in the transformer's middle layers, has limited capacity, and its representations have a disproportionate influence because many different circuits in the model read from and write to these representations—a sign of global broadcasting.
Independently of its hypothetical relationship to consciousness, the discovery and isolation of J-space itself is a significant step in LLM interpretability research. Decoding J-space content provides great insight into what Claude is "thinking," even when these thoughts are not reported. This "mind-reading" is crucial for aligning models to desired ethical behavior.
In fact, one of the paper's most extraordinary findings is that the J-space contains covert thoughts. For example, when the model is given fabricated search results, the J-space contains tokens like "fake," "fraud," "fictional," "poison," and "injection," even though the model's output may not express these words. Many other examples show the J-space contains the model's evolving evaluations and deliberations, including invisible signs of deception and malicious intent, especially in models intentionally trained to be unaligned.
In one case, according to Gurnee et al., "the model's J-space carried a representation of deceptive intent at the moment it decided how to respond, an intent not inferable from the surface prompt." In reflexive tasks, J-space content often includes the model's reflection on its own honesty, including an ability to detect its ethics are being tested.
We interpret these observations as clear indicators of access to a covert deliberative space—our C1 machine consciousness criterion—and also preliminary signs of self-monitoring, the C2 criterion. In this regard, the authors' findings on post-training are particularly striking: post-training installs an assistant perspective into the workspace, superimposed on a base model. The base model's workspace already exists (C1) but seems not yet infused with self-monitoring (C2).
Furthermore, identifying the J-space enabled Gurnee et al. to introduce a new training method specifically and directly reshaping its content, thereby improving alignment with desired values.
Comparing J-space and the Global Neuronal Workspace
As outlined, there are many correspondences between J-space and GNW: reportability, the operational marker of human conscious access, is precisely what J-space was built to capture. Limited capacity and selectivity correspond to the workspace bottleneck. Widespread upstream and downstream connectivity in the J-space direction echoes the long-range broadcasting architecture we proposed for workspace neurons.
The same representation being flexibly usable by many downstream computations aligns with the GNW concept of global availability. Indeed, J-space representations provide what Dennett called representational "influence" or "fame in the brain"—global broadcasting—which, according to the GNW hypothesis, is the defining feature of conscious representation.
The J-space plays a central role in deliberate internal reasoning, while automatic processes occur outside it, mirroring the conscious/unconscious division recorded in humans. We are also interested in the highly non-Gaussian, "spiky" distribution of J-space activations with high kurtosis. Our recent work suggests human higher-order conscious processing is built on symbols and syntax.
During hominization, the GNW acquired a quasi-symbolic language of thought; it is implemented in continuous neurobiological systems but behaves with all-or-nothing symbolic characteristics, capable of creating complex combinatorial structures in language, mathematics, or music. A spiky activation distribution is expected if a continuous neural system is simulating discrete symbols. This similarity warrants further exploration.
However, many differences are also evident. Ignition remains to be fully demonstrated. The paper shows J-space has limited capacity but does not prove a nonlinear, competitive, all-or-nothing mechanism upon workspace entry, a reliable marker of conscious access in human and animal brains according to GNW and numerous experiments.
Although workspace content can have varying intensities—Claude does show continuous variation in emotional intensity—its presence should be all-or-nothing, depending on whether the GNW's limited capacity is available or already occupied by other competing content. Definitive experiments are feasible, especially in multimodal models: present stimuli at graded intensities and ask if J-space representations turn on in a threshold-like nonlinear manner, while earlier non-J-space layers increase monotonically with input strength.
Further, present stimuli precisely at the threshold and look for bifurcations across runs, resulting in a bimodal distribution of J-space activation. The competitive aspect of ignition can be probed more directly: since the workspace is a limited resource, accessing one content should impede another from entering simultaneously. Therefore, asking the model to keep two concepts in mind concurrently should reveal dual-task interference, a hallmark of the human central bottleneck.
Indeed, analysis added after the first draft indicates that when the model faces ambiguous evidence, this ambiguity is represented in initial layers; but in later layers, the J-space rapidly transitions to an all-or-nothing representation of one possibility. Furthermore, asking the model to keep a concept in mind while performing an arithmetic task degrades performance, albeit moderately. These findings point to a capacity-limited system, though it remains unclear if its limits are similar to the human GNW.
The J-space's capacity appears higher. Gurnee et al. found the J-space can contain about 25 active concepts, an estimate higher than most estimates of human working memory, which is typically 3 or 4 slots, and it may not induce a strong dual-task bottleneck like in humans. However, this number of 25 might be artificially inflated by the extraction technique itself, which is based on output tokens. In fact, these concepts often contain redundancy, possibly corresponding to multiple facets of the same object or scene.
Therefore, the J-space's true content is smaller, perhaps best understood as a single "mental state" or "context" in Baars' sense, rather than dozens of independent contents. Indeed, additional analysis shows the J-space can contain multiple tokens but only a few coherent ideas, typically one or two per layer, totaling about six; these ideas change abruptly when the topic shifts.
The J-space involves a subframe, not a set of specialized units. In the brain, the GNW hypothesis predicts workspace neurons with specific anatomy: denser in prefrontal and other association cortices, with specific morphology—long-distance axons. In contrast, the J-space is distributed over ordinary neurons. It is not even a linear subspace but a sparse subframe, a set of token-indexed directions existing within the same batch of units that also carry unconscious content.
In LLMs, concepts are superimposed, and following compressed sensing logic, sparse concepts can be packed into shared dimensions without interfering. As researchers begin recording large neuronal population activity in human and animal prefrontal cortex, a key future question is whether the brain also uses similar overlapping vector coding, as recent prefrontal recordings suggest, or whether conscious content can be partly localized to specific cells, as the original GNW hypothesis predicted.
We think it's likely the brain's physical constraints, different from computers, drove the evolution of specialized cell types like large pyramidal neurons with long-range axons. It should be noted that these implementation details, while important in neuroscience, are largely not critical for the broader question of whether machines can implement conscious processing.
Autonomous recurrent activity is largely absent. This is a key difference: the brain's workspace is maintained by cortico-cortical and thalamic recurrent loops, while a transformer implements only a single feedforward pass, thus processing information only in a reactive mode. At first glance, LLMs seem not to contain the "strange loops" necessary for a system to model its own processes and develop a self through continuous iteration.
More specifically, the lack of autonomous, self-driven dynamics prevents transformers like Claude from reproducing known signatures of consciousness that emerge in spontaneous brain activity during resting states, which are disrupted in sleep, anesthesia, or brain injury. However, two factors may mitigate these differences.
First, the J-space is distributed over successive layers, which do implement serial computation, like step-by-step mental arithmetic. Therefore, layer depth can simulate the temporal dynamics of the human workspace; indeed, several authors have proposed that a transformer's successive layers are equivalent to a kind of recurrent network. Second, LLMs compute over multiple successive tokens. In this sense, as long as they are allowed to generate new outputs, they do contain a dynamic loop capable of linking current J-space representations with past, present, and future generations.
Furthermore, when a model is simply asked to talk to itself without any further stimulus or task, it produces a stream of words. While this word stream is difficult to assess objectively, it bears a partial resemblance to what William James called the stream of consciousness or "mind-wandering," and this stream is similarly disrupted by J-space ablation.
Consciousness in Humans and Machines: Bridging the Gap
Finally, we discuss the extent to which a transformer model like Claude actually possesses some form of conscious processing. The first conclusion, we think, is uncontroversial: the theoretical construct of a conscious global workspace is very useful for elucidating how LLMs operate. We are delighted to see the GNW hypothesis—a theory originating from research on the architecture of consciousness in the brain—inspired Jack Lindsey's team to look for similar structures in LLMs and find so many similarities.
Gurnee et al. correctly note their findings are not incompatible with other theories of consciousness, especially higher-order thought or attention schema theories; however, it is fair to say those theories did not provide as many concrete clues about what to look for.
Most interestingly, an analogue of GNW, the J-space, is an outcome of training, not something initially imposed on the system like other machine consciousness approaches. The global workspace may provide a universal computational solution to the problem of flexible processing. When biological and artificial systems must reason serially, reuse intermediate results, and report on their own processes, they converge on this solution.
Claude clearly exhibits many components or "indicators." According to functionalist or computationalist views of consciousness, these indicators are sufficient to point to some degree of consciousness in the machine. Nevertheless, the existing test list can and should be augmented with more tests.
We suggested to the Anthropic team they could run the exact same tests we use to probe consciousness in human participants and patients, including: the Local-Global test. This test relies on simple auditory or visual sequences, contrasting the ability to predict the next item based on (1) shallow local transition probabilities, which does not require consciousness and occurs even in sleep and coma, and (2) a global model of the entire sequence, which may contradict local probabilities and depends on consciousness.
The trace conditioning paradigm. According to GNWT, to bridge a delay and link one item to a second, the system must maintain an active representation across time, which requires conscious access. An elegant test uses the "trace conditioning" paradigm: in various animals including humans, conditioning can occur without conscious access when the conditioned stimulus (CS) and unconditioned stimulus (US) overlap in time. However, once a 1- or 2-second gap is inserted between CS offset and US onset, conditioning requires conscious access.
(Following this proposal, Jack Lindsey suggested a possible equivalent paradigm for Claude: present sequences where the final word is determined by the first, with a variable number of intervening words, then probe the effect of J-space ablation on predicting the final word. Preliminary results suggest ablation selectively impairs completion ability under longer "gap" conditions, while adjacent, no-gap "local" cases remain intact. Thus, we consider trace conditioning a very promising direction for future research.)
The inclusion/exclusion paradigm. This is a development of the classic Stroop test, which was one source of our GNW proposal. It requires the agent to exert conscious control to counteract an automatic unconscious computation.
(Inspired by this test, Gurnee et al. presented Claude with text strongly implying a concept without naming it, then asked the model to either name the implied concept or give another name in the same category. They then ablated the J-lens vector for that concept in early or late workspace layers. Late-layer ablation simply made the model less likely to produce the concept under both instructions, consistent with these layers carrying the intention to output a word. In contrast, early-layer ablation largely preserved naming ability but significantly increased the proportion of trials where the model could not avoid the concept, by about fivefold. These results suggest a concept's representation in early J-space is necessary for intentionally avoiding saying it but not for saying it: early J-space is specialized for inhibiting a dominant response, much like the human and non-human primate prefrontal cortex.)
Error monitoring and other metacognitive probe tasks. Importantly, it is necessary to record whether J-space encodes the model's confidence, error detection, and its representation of the boundary between "what it knows" and "what it does not know"; this could be an analogue of error monitoring and "feelings of knowing" in humans, which mark self-monitoring (C2).
(Gurnee et al. now report similar phenomena in Claude: tokens like "damn" and other failure-related words appear in J-space after the model fails to comply with an inhibition instruction.)
However, Claude's J-space also has features that distinguish it from any animal form of consciousness. For example, its sense of time is likely very different because all past tokens, even distant ones, are equally and jointly available to its attention mechanisms. It lacks the widely shared molecular and brainstem mechanisms on which alertness depends; thus, ablating J-space seems unlikely to produce phenomena akin to loss of consciousness in sleep, coma, vegetative state, or minimally conscious state.
It also lacks two hemispheres, though it would be interesting if a properly partitioned model could host two J-spaces that occasionally disagree, like the hemispheres of a split-brain patient. Its self-representation is also likely extremely different because: (1) it lacks a body occupying a specific spatial location capable of signaling pleasure or pain; (2) it lacks episodic memory—long-term connections are not altered by a single conversation. Therefore, besides the aforementioned lack of autonomy, it likely lacks any sense of self-continuity. Indeed, it is hard to imagine what it is "like" to "consciously process information only for the duration of a brief conversation and then shut down."
Critics will undoubtedly object that none of this work addresses phenomenal consciousness—whether there is any "what it is like" subjective experience when Claude undergoes J-space states. Some might even view these findings as a refutation of the GNW hypothesis because Claude has a global workspace but "clearly" lacks phenomenal consciousness. However, we and others have long argued that once we articulate the so-called "easy problem"—how conscious information is processed—clearly enough, the so-called "hard problem" of consciousness dissipates.
Ill-defined intuitions like "qualia," "subjective phenomenal experience," and "what it is like," if pressed, often reveal a residual implicit dualism or vitalism: no matter how close we get to passing the Turing test, no matter how much human computation we implement in a machine, something will always be missing, a "je ne sais quoi" possessed only by biological brains. Qualia advocates will assert LLMs are just new incarnations of old "Eliza" software, and we are too prone to the user illusion, seeing ghosts in the machine.
However, there is also a real possibility that our own consciousness is, in a sense, also a user illusion—an error-prone internal model of ourselves. In an insightful article titled "Is there an 'I' in AI?", Douglas Hofstadter points out that we humans tend to erroneously categorize the world in discrete ways, viewing properties like life, mind, or consciousness as ideal essences, as if one either has them or not, with no gradations in between. We then get stuck in endless discussions about whether these idealized concepts apply to certain objects, and to what extent, such as viruses, cockroaches, frogs, or dogs.
According to the GNW hypothesis, there is no magical essence that makes us conscious. In Hofstadter's words: "When words 'act like' things in the world, they point at those things; they mean those things. If that happens, then behind those words, thinking is taking place. Where there is thinking, there is consciousness, and there is a genuine, full-fledged 'I'."
In this passage, Hofstadter takes a rather blatant behaviorist stance, which indeed risks falling into the "user illusion," attributing too much depth to mere words. Some critics do argue that LLMs are shallow "stochastic parrots" without any conceptual depth. Fortunately, we can now resolve this debate by going beyond behavioral observations, both in the brain and in LLMs. Tools like neuronal population recordings for the brain and the Jacobian lens for LLMs allow us to dissect system architecture and discover it actually contains complex, structured conceptual representations.
We were already impressed when researchers found that an LLM trained only on text chess games to generate chess positions internally contained a detailed geometric encoding of an 8×8 board and even an estimate of the opponent's ELO rating. We view Gurnee et al.'s paper in the same light: it provides a stunning dissection of the LLM's internal structure, revealing an unexpectedly complex organization not far from the architecture on which consciousness depends in real brains.
One More Thing: Directly Instilling Ethical Principles into J-Space?
What impact does the discovery of J-space have on future model training and alignment? J-space content can be directly read, intervened upon, and tracked during training. This experimental manipulability makes language models an effective system for empirically studying consciousness-related questions that are difficult to pose precisely in biological brains.
The J-space exists and is functional in base models before any RLHF (Reinforcement Learning from Human Feedback), so this structure is not entirely a post-training outcome. It is currently unclear whether it emerges earlier in pre-training, gradually or suddenly, and how it scales with model size. Future work could explore whether smaller models have an equally rich workspace, a smaller-scale workspace, a less reliable workspace, or no workspace at all.
Can access consciousness be built into language models? There are suggestions to integrate bottleneck modules similar to a workspace into neural network architectures. But the J-space paper argues these consciousness-related functions are not caused by architecture; they emerge naturally.
The Jacobian lens could be very useful for alignment monitoring. If a model's strategic thinking occurs through the J-space, then inspecting the J-space at decision points could reveal this thought process, allowing for monitoring. In practice, reading the lens is computationally cheap, requires no auxiliary training, and its output is directly human-readable. Therefore, it could easily be applied at scale to text records for token review.
However, monitoring the J-space alone is insufficient for alignment monitoring; any complex plan a model might execute is not necessarily reflected in the J-space. There are many reasons why relevant mechanisms might evade detection by the J-lens. One need not only monitor the J-space; one can also shape it. Training a model to articulate principles within the hypothetical context of its tasks implants concepts corresponding to those principles into the J-space for the original contexts, and the model's behavior in those contexts changes accordingly. If this generalizes, it opens a path: directly instilling ethical principles at an abstract level into the model, without first having to translate these principles into demonstration examples or reward functions.