Showing posts sorted by date for query Integrated information theory. Sort by relevance Show all posts
Showing posts sorted by date for query Integrated information theory. Sort by relevance Show all posts

Thursday, June 04, 2026

Herbie: A Near-Future Debatably Conscious AI Person

Liberals about AI consciousness hold that we might soon (if we haven't already) create genuinely conscious AI systems. Conservatives about AI consciousness hold that AI consciousness remains in the distant future if it's possible at all. According to the Leapfrog Hypothesis, the first conscious AI will not have merely a dim glow of animal-like consciousness, but rich consciousness, similar to a human's. Such an entity would deserve humanlike rights. They would be a person in the ethical sense of the term.

Let's design, in imagination, a technologically feasible near-future AI system to delight the liberals, leapfrogging to personhood. I'll call him Herbie.

[Herbie the Love Bug: image source]

Start with a self-driving car. According to Global Workspace Theory -- perhaps the leading scientific theory of consciousness -- the car will be conscious if high-priority information is globally available to its various computational systems. For example, a representation like "battery almost empty" could be broadcast widely, influencing downstream processing across the vehicle. The navigational system might then search for nearby charging stations, while the acceleration system prioritizes greater energy efficiency, the braking system prioritizes better energy recapture, and a voice system announces the situation to the passengers.

In line with Higher Order Theory, Herbie might also monitor his representations of the road, vehicles, pedestrians, and hazards, assigning some a low probability of correctness. "Pedestrian at location X" might be flagged as only 60% likely to be correct given a history of revised representations of pedestrians in similarly cluttered environments, while "stoplight in 100 meters" might rate over 99% likely. Minor fluctuations in sensors for battery life, cabin temperature, and distance from a lane divider might be ignored as noise, while larger fluctuations -- especially when plausible given other representations (the battery is likelier to gain charge while braking than while accelerating) -- might be treated as accurate signals and permitted to influence downstream processing.

Even if we grant the liberals that this version of Herbie would, or might plausibly be, genuinely conscious, he still falls far short of humanlike consciousness. "Battery almost empty" and "pedestrian at location X" are hardly rich cognitive or perceptual contents. So let's give Herbie the capacity to speak. Fill his trunk with a server running a large language model, connected to the internet and integrated with his global workspace so that high-priority information provides context for language processing, with the language outputs influencing Herbie's other processes. Now people can chat with Herbie as they would with any language model. But unlike today's language models, his speech will be influenced by information about his location, speed, destination, charge, the condition of his parts, the number and location of his passengers, his radio and climate controls, and so on. He can discuss local history, debate whether the music is too loud, and suggest scenic routes.

"Predictive processing" theories in cognitive science emphasize the value of predicting future inputs and registering the difference between received and predicted inputs. When prediction error is large, the system corrects its weights and representations, enabling more accurate predictions in future situations. This is not so different from the reinforcement learning used to train large language models, and it could help Herbie improve his predictions over time. Predictive processing could occur at multiple levels: in fast recurrent loops within sensory systems even when those representations aren't prioritized for global broadcast, and in slower evaluations of globally broadcast, more integrative predictions. Herbie might model himself as an agent producing volatility in his own environment and inputs, at multiple temporal scales. Subroutines in specialized processors might model long chains of what-would-happen-if.

Let's give Herbie some long-term memory. A facial recognition system might identify his passengers, retrieving past interactions, names, previous destinations, and other information relevant to the current interaction. Incidents of high prediction error might also be stored so that Herbie can compare current inputs with past anomalies, improving his learning and attention in situations likely to be unusual or hard to predict. Passengers might also instruct Herbie to store information in long-term memory, such as text, pictures, maps, or records of his own informational states, optionally with instructions about when to retrieve that information how to use it.

Herbie will have some implicitly or explicitly weighted goals. A pedestrian suddenly in his path will trigger braking, overriding lower-priority processes. Avoiding collisions will outweigh conserving energy. Herbie might monitor the condition of his parts and prioritize preventing damage, deploying extra coolant when the engine is dangerously hot and keeping a one-meter margin between himself and adjacent cars. We can enrich his goals, making him more interesting and giving him more to do. He might have the goal of delighting children, leading him to drive around town and tell jokes to kids on the sidewalk. A reinforcement learning algorithm might strengthen connections when his jokes draw a smile, weaken them when reactions are neutral or negative.

Herbie might also have the goal of photographing the city and posting the images on social media, leading him to explore. If social media likes and shares are rewarding, he might learn to prefer certain neighborhoods, views, lighting conditions, and photographic approaches, while avoiding boring repetition. All of this could feed into a global workspace that provides context for his language model, with selective long-term storage and retrieval. Now we can imagine him discussing, with growing sophistication, his approaches to popular photography and to amusing children.

Herbie will then have something functionally similar to emotion: reward processes, an ability to track his progress toward or away from valued goals, and immediate positive or negative responses to new stimuli in light of their influence on his prospects. He will have something functionally similar to introspection: an ability to track and report his own cognitive or representational processes. He will have something functionally similar to a unified sense of self: a sense of his history, the boundaries of his body, his future, his values and priorities. He will have something functionally similar to imagination: a capacity to model hypothetical sequences of events. He will have something functionally similar to complex chains of humanlike linguistic thought.

Maybe Herbie falls in love with his owner or another car of his type. Maybe he develops deep mutual attachments with friends, neighbors, associates, and people he thinks of as family and who think of him the same way. Or to speak more carefully, maybe Herbie shows all the functional and behavioral signs of doing so, while society remains uncertain whether he is genuinely conscious and genuinely experiences the feelings he professes and that his companions attribute to him.

If we allow, with the liberals, that Herbie is or might well be conscious, then it's plausible that his consciousness is not simple but rich and sophisticated. He won't be exactly humanlike, of course. But will he be humanlike enough to count as a person who deserves humanlike rights? For the liberally inclined, it won't be unreasonable, I submit, to think or guess that Herbie is a person. He would then appear to deserve rights such as self-determination, emergency care, and political representation.

If there is some important aspect of humanlike consciousness that I have omitted from my description an AI analog of which is technologically feasible in the near term, stipulate that Herbie also has that feature.

An entity like Herbie would almost certainly invigorate conservatives to articulate and defend views about what he lacks that is necessary for consciousness -- some crucial functional capacity or some biological substrate that can't be replicated in silicon. And they might be entirely right! My point is not that Herbie, or some similar AI system, would actually have richly humanlike consciousness and ethical personhood. Rather, my point is that guessing that he does, and guessing that he does not, would both be reasonable. Herbie, or some alternative near-future AI system, would be a debatable person, about whom people could reasonably starkly disagree.

Ah, but maybe you think consciousness requires an act of God, to instill an immaterial soul? I imagine that a benevolent God would be delighted to give Herbie a soul, thereby making the world richer and better -- for wouldn't it be?

I contend the following: Anyone who claims to know how best to think about Herbie's consciousness or its absence is overconfident. The science of consciousness is too difficult, too methodologically uncertain, and too near its beginnings. All anyone can have -- whether expert or layperson -- is a hunch or inclination, a well-informed guess, but only a guess, not knowledge. Theories of consciousness span a wide spectrum and the methodologies are dubious and often question-begging. Many views can be defended with some plausibility, but precisely for that reason, none can be defended decisively. (For more on this issue, see my forthcoming book, AI and Consciousness, where I present the detailed case for uncertainty.)

Thursday, April 09, 2026

AI and Consciousness: A Skeptical Overview, forthcoming with Cambridge

Last week I submitted my latest book manuscript to Cambridge University Press (for their "Element" series of books about 100 pages long): AI and Consciousness: A Skeptical Overview -- because you haven't heard nearly enough about AI and consciousness recently, of course! [winky face]

Maybe you'll appreciate my skeptical stance, at odds both with the boosters who anticipate imminent AI consciousness and with the scoffers who pooh-pooh the possibility. Or maybe you'll loathe my skeptical stance but grudgingly accept it against your will, due to the force of my arguments!

I've pasted the introductory chapter below. The full (citable) manuscript version is available here and here.

[AI and Consciousness, title page]


Chapter One: Hills and Fog

1. Experts Do Not Know and You Do Not Know and Society Collectively Does Not and Will Not Know and All Is Fog.

Our most advanced AI systems might soon – within the next five to thirty years – be as richly and meaningfully conscious as ordinary humans, or even more so, capable of genuine feeling, real self-knowledge, and a wide range of sensory, emotional, and cognitive experiences. In some arguably important respects, AI architectures are beginning to resemble the architectures many consciousness scientists associate with conscious systems. Their outward behavior, especially their linguistic behavior, grows ever more humanlike.

Alternatively, claims of imminent AI consciousness might be profoundly mistaken. Their seeming humanlikeness might be a shadow play of empty mimicry. Genuine conscious experience might require something no AI system could possess for the foreseeable future – intricate biological processes, for example, that silicon chips could never replicate.

The thesis of this book is that we don’t know. Moreover and more importantly, we won’t know before we’ve already manufactured thousands or millions of disputably conscious AI systems. Engineering sprints ahead while consciousness science lags. Consciousness scientists – and philosophers, and policy-makers, and the public – are watching AI development disappear over the hill. Soon we will hear a voice shout back to us, “Now I am just as conscious, just as full of experience and feeling, as any human”, and we won’t know whether to believe it. We will need to decide, as individuals and as a society, whether to treat AI systems as conscious, nonconscious, semi-conscious, or incomprehensibly alien, before we have adequate grounds to justify that decision.

The stakes are immense. If near-future AI systems are richly, meaningfully conscious, then they will be our peers, our lovers, our children, our heirs, and possibly the first generation of a posthuman, transhuman, or superhuman future. They will deserve rights, including the right to shape their own development, free from our control and perhaps against our interests.[1] If, instead, future AI systems merely mimic the outward signs of consciousness while remaining as experientially blank as toasters, we face the possibility of mass delusion on an enormous scale. Real human interests and real human lives might be sacrificed for the sake of entities without interests worth the sacrifice. Sham AI “lovers” and “children” might supplant or be prioritized over human lovers and children. Heeding their advice, society might turn a very different direction than it otherwise would.

In this book, I aim to convince you that the experts do not know, and you do not know, and society collectively does not and will not know, and all is fog.

2. Against Obviousness.

Some people think that near-term AI consciousness is obviously impossible. This is an error in adverbio. Near-term AI consciousness might be impossible – but not obviously so.

A sociological argument against obviousness:

Probably the leading scientific theory of consciousness is Global Workspace theory. Its leading advocate is neuroscientist Stanislas Dehaene.[2] In 2017, years before the surge of interest in ChatGPT and other Large Language Models, Dehaene and two collaborators published an article arguing that with a few straightforward tweaks, self-driving cars could be conscious.[3]

Probably the two best-known competitors to Global Workspace theory are Higher Order theory and Integrated Information Theory.[4] (In Chapters Eight and Nine, I’ll provide more detail on these theories.) Perhaps the leading scientific defender of Higher Order theory is Hakwan Lau – one of the coauthors of that 2017 article about potentially conscious cars.[5] Integrated Information Theory is potentially even more liberal about machine consciousness, holding that some current AI systems are already at least a little bit conscious and that we could easily design AI systems with arbitrarily high degrees of consciousness.[6]

David Chalmers, the world’s most influential philosopher of mind, argued in 2023 for about a 25% degree of confidence in AI consciousness within a decade.[7] That same year, a team of prominent philosophers, psychologists, and AI researchers – including eminent computer scientist Yoshua Bengio – concluded that there are “no obvious technological barriers” to creating conscious AI according to a wide range of mainstream scientific views about consciousness.[8] In a 2025 interview, Geoffrey Hinton, another of the world’s most prominent computer scientists, asserted that AI systems are already conscious.[9] Christof Koch, the most influential neuroscientist of consciousness from the 1990s to the early 2010s, has endorsed Integrated Information Theory, including its liberal implications for the pervasiveness of consciousness.[10]

This is a sociological argument: a substantial probability of near-term AI consciousness is a mainstream view among leading experts. They might be wrong, but it’s implausible that they’re obviously wrong – that there’s a simple argument or consideration they’re neglecting which, if pointed out, would or should cause them to collectively slap their foreheads and say, “Of course! How did we miss that?”

What of the converse claim – that AI consciousness is obviously imminent or already here? In my experience, fewer people assert this. But in case you’re tempted in this direction, note that other prominent theorists hold that AI consciousness is a far-distant prospect if it’s possible at all: neuroscientist Anil Seth; philosophers Peter Godfrey-Smith, Ned Block, and John Searle; linguist Emily Bender; and computer scientist Melanie Mitchell.[11] (Chapter Six will discuss thought experiments by Searle, Bender, and Mitchell, and Chapter Ten will discuss biological views of the sort emphasized by Seth, Godfrey-Smith, and Block.) In a 2024 survey of 582 AI researchers, 25% expected AI consciousness within ten years and 70% expected AI consciousness by the year 2100.[12]

If the believers are right, we’re on the brink of creating genuinely conscious machines. If the scoffers are right, those machines will only seem conscious. I assume that this is a substantive disagreement, not just a disagreement about how to apply the term “consciousness” to a perfectly obvious set of phenomena about which everyone agrees. The future well-being of many people (including, perhaps, many AI people) depends on getting this issue right. Unfortunately, we will not know in time.

The rest of this book is flesh on this skeleton. I canvass a variety of structural and functional claims about consciousness, the leading theories of consciousness as applied to AI, and the best known general arguments for and against near-term AI consciousness. None of these claims or arguments takes us far. It’s a morass of uncertainty.

-------------------------------------------

[1] I assume that AI consciousness and AI rights are closely connected: Schwitzgebel 2024, ch. 11, in preparation. For discussion, see Shepherd 2018; Levy 2024.

[2] Dehaene 2014; Mashour et al. 2020.

[3] Dehaene, Lau, and Kouider 2017. For an alternative interpretation of this article as concerning something other than consciousness in its standard “phenomenal” sense, see note 115.

[4] Some Higher Order theories: Rosenthal 2005; Lau 2022; Brown 2025. Integrated Information Theory: Albantakis et al. 2023.

[5] But see Chapter Eight for some qualifications.

[6] See Tononi’s publicly available response to Scott Aaronson’s objections in Aaronson 2014. However, advocates of IIT also suggest that the most common current computer architectures are unlikely to achieve much consciousness and that consciousness will tend to appear in subsystems of the computer rather than at the level of the computer itself (Findlay et al. 2024/2025).

[7] Chalmers 2023.

[8] Butlin et al. 2023. (I am among the nineteen authors.)

[9] Heren 2025.

[10] Tononi and Koch 2015.

[11] Seth forthcoming; Godfrey-Smith 2024; Block forthcoming; Searle 1980, 1992; Bender 2025; Mitchell 2021.

[12] Dreksler et al. 2025.

Tuesday, March 24, 2026

A Model of Disunified Human Experience

It's a philosophical truism that human conscious experience is unified: If you're at a bar, hearing music, tasting beer, and feeling pleasantly relaxed, those experiences don't occur merely side by side. They are joined together into an integrated whole, an experience of music-with-beer-with-relaxation.

I'm not sure this truism is correct. As I suggested in an earlier post, experiential unity might be an artifact of introspection and memory: When we introspectively notice that we're experiencing music, beer, and relaxation all at once, we thereby bind those experiences into a whole. Likewise, when we remember such moments, we reconstruct them as unified. But it doesn't follow that those experiences, even if they all occurred simultaneously in you, were unified rather than transpiring separately. Experiences of music, beer, and relaxation might have all being going on inside of you, no more joined together than those experiences are joined with the similar experiences of your friend across the table. Simple co-occurrence doesn't entail experiential unity.

If this possibility is coherent, then introspection and memory can't establish that experience is always unified. At most, they show that introspected and remembered experiences present themselves as unified. But that leaves open the status of unintrospected, unremembered experiences. Unity becomes difficult to verify by standard phenomenological methods.

But the issue needn't be intractable. We just need to approach it less directly, for example by exploring what follows from a well-established theory of consciousness. If some well-motivated Theory X implies unity (or disunity), that would provide reason to accept its conclusion.

I'll now present a candidate Theory X. I'm not suggesting that this is the right theory of consciousness! For one thing, it's simplistic. I'm sure the mind is much more complicated than I'm about to say. I offer this theory only as a proof of concept. There could be a theory of consciousness with massive disunity as an implication.

This theory combines Global Workspace Theory and Recurrent Processing Theory. According to this hybrid, Global Workspace Theory governs attended experiences -- those targeted by introspection or reconstructed in memory -- while Recurrent Processing Theory governs unattended experiences.

The mind, on this picture, is composed of many separate "modules" that work mostly independently, connected by a workspace where a small amount of attended information is shared globally. There's a visual module, an auditory module, modules for motor activity, episodic memory, and so on. When we attend to something -- say, the taste of beer -- the information from the relevant module is broadcast into the Global Workspace, where it can be accessed by and influence processes in all the other modules. When unattended, the information stays local.

Here's one illustration of this type of architecture:

[the Global Workspace; source]

Orthodox Global Workspace Theory holds that only what is broadcast into the workspace is conscious. Theory X alters that assumption. Many people hold that conscious experience vastly outruns attention. Many people hold, that is, that you can experience the hum of traffic in the background when you're not attending to it, and the feeling of your feet in your shoes, and the leftover taste of coffee in your mouth, etc. -- all in a peripheral way, simultaneously, when your focus is elsewhere. Theory X, drawing on Recurrent Processing Theory, holds that such processes are conscious whenever there's enough cognitive activity of the right sort (recurrent processing, for example) in the modules, even without global broadcast.

The picture, then, is this: We have multiple sensory (and other) experiences all running simultaneously, each with enough cognitive processing to be conscious, but few of which are selected for global availability through attention.

Is there reason to think these modular processes are unified with one another? I see no reason to think so, if they're genuinely modular -- that is, if their processing stays local, exerting little influence elsewhere. The taste-of-beer processing stays in the tasting module. The sound-of-music processing stays in the auditory module. No link up. No straightforward causal, functional, or physiological basis for a unified experience of beer-with-music rather than, separately, an experience of beer and an experience of music.

When we introspect the beer and music simultaneously, we pull both into the Global Workspace, and there they unify. We might then mistakenly think they were unified all along, but that's an illusion of introspection. It's an example of the "refrigerator light error", the error of thinking that the light is always on because it's always on when you open the door to check.

On this model, disunity is the normal human condition. Our experiences are fragmented, except when we pull them together through attention. We just don't realize that fact because, so to speak, we only attend to what we attend to.

Two caveats:

First, this is probably not the right model of consciousness. But I don't think it's unreasonable to wonder if the correct model is similar enough to have the same implications. If so, we can't simply accept the unity of consciousness as a given.

Second, the recurrent peripheral, modular processes that don't make it into the workspace might not be determinately conscious. They might be only borderline conscious, in the indeterminate middle between consciousness and nonconsciousness, like a color can be indeterminately between green and not-green. This opens a third possibility, alongside unity and disunity: unity among the determinately conscious experiences with a hazy penumbra of indeterminate experiences that remain disunified. (There are further possibilities beyond these three; but save them for another day.)

Thursday, February 19, 2026

Disunity and Indeterminacy in Artificial Consciousness (and Maybe in Human Consciousness Too)

Our understanding of the nature of consciousness derives mainly from our understanding of the nature of consciousness in our favorite animal (us, of course). But the features of consciousness in our favorite animal might be specific to that animal rather than universal.

Let's consider two such features and whether we should expect them in conscious AI systems, if conscious AI systems are ever possible.

Unity: Our conscious experiences at any given moment are bound together into a single unified experience, rather than transpiring in separate streams. If I'm sitting on a wet park bench, I might (a.) visually experience the leafy green trees around me, (b.) tactilely experience the cold dampness soaking into my jeans, and (c.) consciously recall the smaller trees of yesteryear. Normally -- perhaps necessarily -- three such experiences would not run in disconnected streams. They would join into a composite experience of (a)-with-(b)-with-(c). I experience not just trees, cold dampness, and a memory of yesteryear, but all three together as a unified bundle.

Determinacy: At any given moment, I am either determinately conscious or determinately nonconscious (as in anesthesia or dreamless sleep). Likewise, I either determinately do, or determinately do not, have any particular experience. Gray-area cases are at least unusual and maybe impossible. Even the simplest, barest cases are still determinate. Consider visual experience: We might imagine the visual field narrowing and losing content until only a gray dot remains -- and then the dot winks out. That dot, however minimal, is still determinately experienced. When it winks out, consciousness determinately disappears. There is no half-winked state between the minimal gray dot and complete absence of visual experience.

My thought is that we should not expect unity and determinacy to be general features of conscious AI systems (if conscious AI is possible). To see why, let's start by assuming the Global Workspace Theory of consciousness. I focus on Global Workspace Theory because it's probably the leading scientific theory of consciousness and because its standard formulation (Dehaene's version) invites the assumption of unity and determinacy.

Global Workspace Theory divides the mind into local information processing modules linked by a shared global workspace. Information becomes conscious when it is broadcast into the workspace. Suppose your auditory system registers the faint honk of a distant car horn. You're absorbed in reading philosophy and accustomed to ignoring traffic noise, so this representation isn't selected for further processing. It's not a target of attention, not broadcast into the workspace, and not consciously experienced. (If you think you constantly consciously experience background sounds, you can't hold a standard Global Workspace view.) Once you attend to the noise, for whatever reason, that information "ignites" into the global workspace, becoming available to a wide variety of "downstream" processes: You can think about it, plan around it, verbally report it, store it in long-term memory, and flexibly combine it with other information in the workspace. On Global Workspace Theory, being available in this way just is what it is for the information to be consciously experienced.

This model suggests unity and determinacy. Since there is just one global workspace, and since that workspace enables flexible integration of everything it contains, it makes sense that its various elements will combine into a unified experience. And on Dehaene's version, ignition into the workspace is a sharp-boundaried event: Information either completely ignites, becoming available for all downstream processes, or it does not. There is no (or only rarely) partial ignition. This can explain determinacy.

But future AI systems might not share this structure. They might have multiple or partially overlapping workspaces. Different specialized subsystems might have access to different regions of a partly-shared workspace. Some animals, such as snails and octopuses, distribute processing among multiple ganglia or neural centers that are less tightly coupled than the hemispheres of the human brain. A robot might broadcast information relevant to locomotion to one area and information relevant to speech to another with limited connectivity.

If the subsystems are entirely disconnected, the result might be entirely discrete centers of subjective experience within a single organism or machine. But if they are partly connected, experience might be only partly unified. In the park bench example, the experience of the trees might be unified with the experience of dampness, and the experience of dampness with memories of yesteryear, but the experience of the trees might not be unified with the memories. (Unification would not then be a transitive relation.) Alternatively, some weaker relation of partial unification might hold among the visual, tactile, and memorial experiences. If this seems inconceivable or impossible, see Sophie Nelson's and my article on indeterminate or fractional subjects.

More abstractly: There's no compelling architectural reason why an AI system would have to make information available either to all downstream processes or to none. A workspace defined in terms of downstream availability could be a patchwork of partial availabilities rather than a fully global all-or-nothing broadcast.

For the same reason, ignition into the workspace needn't be all-or-nothing. Between full ignition with determinate consciousness and no ignition with determinate nonconsciousness, there might be in-between, gray-area half-ignitions that are neither determinately conscious nor determinately nonconscious. Nearly every property with a complex physical or functional basis allows indeterminate, borderline cases: baldness, extraversion, greenness, happiness, whether you're wearing a shoe, whether a country is a democracy. The human global workspace might minimize indeterminacy -- like it's rarely indeterminate in basketball whether the ball has gone through the hoop. But change the architecture and indeterminacy might become common: a half-hearted ignition, or just enough information-sharing to make it indeterminate whether a workspace even exists. (If indeterminacy about consciousness strikes you as inconceivable or impossible, see my 2023 article on borderline consciousness.)

Global Workspace Theory might of course be wrong. But most other theories of consciousness make my argument at least as easy. Dennett's fame-in-the-brain version of broadcast theory explicitly permits disunity and indeterminacy. Higher Order Theories admit the same fragmentation and, probably, gradualism. So do biological theories and theories that focus on embodiment. (Integrated Information Theory is an exception: Its axioms require bright-lined unity and determinacy. But as I've argued, those bright-line axioms lead to unpalatable consequences.)

Recognizing these possibilities for AI systems invites the further thought: Maybe we humans aren't quite as unified as we normally suppose. Maybe indeterminate and disunified consciousness is common. Maybe processes outside of attention hover indeterminately between being conscious and nonconscious. Maybe some processes are only partly unified. If it seems otherwise in introspection and memory, maybe that's because introspection and memory tend to impose unity and determinacy where none was before.

[a Paul Klee painting, untitled 1914: source]

Friday, January 30, 2026

Does Global Workspace Theory Solve the Question of AI Consciousness?

Hint: no.

Below are three sections from Chapter Eight of my manuscript in draft, AI and Consciousness, fresh new version available today here. Comments welcome!

[image adapted from Dehaene et al. 2011]


1. Global Workspace Theories and Access.

The core idea of Global Workspace Theory is simple. Sophisticated cognitive systems like the human mind employ specialized processes that operate to a substantial extent in isolation. We can call these modules, without committing to any strict interpretation of that term.[1] For example, when you hear speech in a familiar language, some cognitive process converts the incoming auditory stimulus into recognizable speech. When you type on a keyboard, motor functions convert your intention to type a word like “consciousness” into nerve signals that guide your fingers. When you try to recall ancient Chinese philosophers, some cognitive process pulls that information from memory without (amazingly) clogging your consciousness with irrelevant information about German philosophers, British prime ministers, rock bands, or dog breeds.

Of course, not all processes are isolated. Some information is widely shared, influencing or available to influence many other processes. Once I recall the name “Zhuangzi”, the thought “Zhuangzi was an ancient Chinese philosopher” cascades downstream. I might say it aloud, type it out, use it as a premise in an inference, form a visual image of Zhuangzi, contemplate his main ideas, attempt to sear it into memory for an exam, or use it as a clue to decipher a handwritten note. To say that some information is in “the global workspace” just is to say that it is available to influence a wide range of cognitive processes. According to Global Workspace Theory, a representation, thought, or cognitive process is conscious if and only if it is in the global workspace – if it is “widely broadcast to other processors in the brain”, allowing integration both in the moment and over time.[2]

Recall the ten possibly essential features of consciousness from Chapter Three: luminosity, subjectivity, unity, access, intentionality, flexible integration, determinacy, wonderfulness, specious presence, and privacy. [Blog readers: You won't have read Chapter Three, but try to ride with it anyway.] Global Workspace Theory treats access as the central essential feature.

Global Workspace theory can potentially explain other possibly essential features. Luminosity follows if processes or representations in the workspace are available for introspective processes of self-report. Unity might follow if there’s only one workspace, so that everything in it is present together. Determinacy might follow if there’s a bright line between being in the workspace and not being in it. Flexible integration might follow if the workspace functions to flexibly combine representations or processes from across the mind. Privacy follows if only you can have direct access to the contents of your workspace. Specious presence might follow if representations or processes generally occupy the workspace for some hundreds of milliseconds.

In ordinary adult humans, typical examples of conscious experience – your visual experience of this text, your emotional experience of fear in a dangerous situation, your silent inner speech, your conscious visual imagery, your felt pains – appear to have the broad cognitive influences Global Workspace Theory describes. It’s not as though we commonly experience pain but find that we can’t report it or act on its basis, or that we experience a visual image of a giraffe but can’t engage in further thinking about the content of that image. Such general facts, plus the theory’s potential to explain features such as luminosity, unity, determinacy, flexible integration, privacy, and specious presence, lend Global Workspace Theories substantial initial attractiveness.

I have treated Global Workspace Theory as if it were a single theory, but it encompasses a family of theories that differ in detail, including “broadcast” and “fame” theories – any theory that treats the broad accessibility of a representation, thought, or process as the central essential feature making it conscious.[3]

Consider two contrasting views: Dehaene’s Global Neuronal Workspace Theory and Daniel Dennett’s “fame in the brain” view. Dehaene holds that entry into the workspace is all-or-nothing. Once a process “ignites” into the workspace, it does so completely. Every representation or process either stops short of entering consciousness or is broadcast to all available downstream processes. Dennett’s fame view, in contrast, admits degrees. Representations or processes might be more or less famous, available to influence some downstream cognitive processes without being available to influence others. There is no one workspace, but a pandemonium of competing processes.[4] If Dennett is correct, luminosity, determinacy, unity, and flexible integration all potentially come under threat in a way they do not as obviously come under threat on Dehaene’s view.[5]

Dennettian concerns notwithstanding, all-or-nothing ignition into a single, unified workspace is currently the dominant version of Global Workspace Theory. The issue remains unsettled and has obvious implications for the types of architectures that might plausibly host AI consciousness.

2. Consciousness Outside the Workspace; Nonconsciousness Within It?

Global Workspace Theory is not the correct theory of consciousness unless all and only thoughts, representations, or processes in the Global Workspace are conscious. Otherwise, something else, or something additional, is necessary for consciousness.

It is not clear that even in ordinary adult humans a process must be in the Global Workspace to be conscious. Consider the case of peripheral experience. Some theorists maintain that people have rich sensory experiences outside of focal attention: a constant background experience of your feet in your shoes and objects in the visual periphery.[6] Others – including Global Workspace theorists – dispute this. Introspective reports vary, and resolving such issues is methodologically tricky.

One methodological problem: People who report constant peripheral experiences might mistakenly assume that such experiences are always present because they are always present whenever they think to check, and the very act of checking might generate those experiences. This is sometimes called the “refrigerator light illusion”, akin to the error of thinking the refrigerator light is always on because it’s always on when you open the door to check.[7] On this view, you’re only tempted to think you have constant tactile experience of your feet in your shoes because you have that experience on those rare occasions when you’re thinking about whether you have it. Even if you now seem to have a broad range of experiences in different sensory modalities simultaneously, this could result from an unusual act of dispersed attention, or from “gist” perception or “ensemble” perception, in which you are conscious of the general gist or general features of a scene, knowing that there are details, without actually experiencing those unattended details.[8]

The opposite mistake is also possible. Those who deny a constant stream of peripheral experiences might simply be failing to notice or remember them. The fact that you don’t remember now the sensation of your feet in your shoes two minutes ago hardly establishes that you lacked the sensation at the time. Although many people find it introspectively compelling that their experience is rich with detail or that it is not, the issue is methodologically complex because introspection and memory are not independent of the phenomena to be observed.[9]

If we do have rich sensory experience outside of attention, it is unlikely that all of that experience is present in or broadcast to a Global Workspace. Unattended peripheral information is rarely remembered or consciously acted upon, tending to exert limited downstream influence – the paradigm of information that is not widely broadcast. Moreover, the Global Workspace is generally characterized as limited capacity, containing only a few thoughts, representations, objects, or processes at a time – those that survive some competition or attentional selection – not a welter of richly detailed experiences in many modalities at once.[10]

A less common but equally important objection runs in the opposite direction: Perhaps not everything in the Global Workspace is conscious. Some thoughts, representations, or processes might be widely broadcast, shaping diverse processes, without ever reaching explicit awareness.[11] Implicit racist assumptions, for example, might influence your mood, actions, facial expressions, and verbal expressions. The goal of impressing your colleagues during a talk might have pervasive downstream effects without occupying your conscious experience moment to moment.

The Global Workspace theorist who wants to allow that such processes are not conscious might suggest that, at least for adult humans, processes in the workspace are generally also available for introspection. But there’s substantial empirical risk in this move. If the correlation between introspective access and availability for other types of downstream cognition isn’t excellent, the Global Workspace theorist faces a dilemma. Either allow many conscious but nonintrospectable processes, violating widespread assumptions about luminosity, or redefine the workspace in terms of introspectability, which amounts to shifting to a Higher Order view.

3. Generalizing Beyond Vertebrates.

The empirical questions are difficult even in ordinary adult humans. But our topic isn’t ordinary adult humans – it’s AI systems. For Global Workspace Theory to deliver the right answers about AI consciousness, it must be a universal theory applicable everywhere, not just a theory of how consciousness works in adult humans, vertebrates, or even all animals.

If there were a sound conceptual argument for Global Workspace Theory, then we could know the theory to be universally true of all conscious entities. Empirical evidence would be unnecessary. It would be as inevitably true as that rectangles have four sides. But as I argued in Chapter Four, conceptual arguments for the essentiality of any of the ten possibly essential features are unlikely to succeed – and a conceptual argument for Global Workspace Theory would be tantamount to a conceptual argument for the essentiality of access, one of those ten features. Not only do the general observations of Chapter Four suggest against a conceptual guarantee, so also does the apparent conceivability, as described in Section 2 above, of consciousness outside the workspace or nonconsciousness within it – even if such claims are empirically false.

If Global Workspace Theory is the correct universal theory of consciousness applying to all possible entities, an empirical argument must establish that fact. But it’s hard to see how such an empirical argument could proceed. We face another version of the Problem of the Narrow Evidence Base. Even if we establish that in ordinary humans, or even in all vertebrates, a thought, representation, or process is conscious if and only if it occupies a Global Workspace, what besides a conceptual argument would justify treating this as a universal truth that holds among all possible conscious systems?

Consider some alternative architectures. The cognitive processes and neural systems of octopuses, for example, are distributed across their bodies, often operating substantially independently rather than reliably converging into a shared center.[12] AI systems certainly can be, indeed often are, similarly decentralized. Imagine coupling such disunity with the capacity for self-report – an animal or AI system with processes that are reportable but poorly integrated with other processes. If we assume Global Workspace Theory at the outset, we can conclude that only sufficiently integrated processes are conscious. But if we don’t assume Global Workspace Theory at the outset, it’s difficult to imagine what near-future evidence could establish that fact beyond a reasonable standard of doubt to a researcher who is initially drawn to a different theory.

If the simplest version of Global Workspace Theory is correct, we can easily create a conscious machine. This is what Dehaene and collaborators envision in the 2017 paper I discussed in Chapter One. Simply create a machine – such as an autonomous vehicle – with several input modules, several output modules, a memory store, and a central hub for access and integration across the modules. Consciousness follows. If this seems doubtful to you, then you cannot straightforwardly accept the simplest version of Global Workspace Theory.[13]

We can apply Global Workspace Theory to settle the question of AI consciousness only if we know the theory to be true either on conceptual grounds or because it is empirically well established as the correct universal theory of consciousness applicable to all types of entity. Despite the substantial appeal of Global Workspace Theory, we cannot know it to be true by either route.

-------------------------------------

[1] Full Fodorian (1983) modularity is not required.

[2] Mashour et al. 2020, p. 776-777.

[3] E.g., Baars 1988; Dennett 1991, 2005; Tye 2000; Prinz 2012; Dehaene 2014; Mashour et al. 2020.

[4] Whether Dennett’s view is more plausible than Dehaene’s turns on whether, or how commonly, representations or processes are partly famous. Some visual illusions, for example, seem to affect verbal report but not grip aperture: We say that X looks smaller than Y, but when we reach toward X and Y we open our fingers to the same extent, accurately reflecting that X and Y are the same size. The fingers sometimes know what the mouth does not. (Aglioti et al. 1995; Smeets et al. 2020). We adjust our posture while walking and standing in response to many sources of information that are not fully reportable, suggesting wide integration but not full accessibility (Peterka 2018; Shanbhag 2023). Swift, skillful activity in sports, in handling tools, and in understanding jokes also appears to require integrating diverse sources of information, which might not be fully integrated or reportable (Christensen et al. 2019; Vauclin et al. 2023; Horgan and Potrč 2010). In response, the all-or-nothing “ignition” view can explain away such cases of seeming intermediacy or disunity as atypical (it needn’t commit to 100% exceptionless ignition with no gray-area cases), by allowing some nonconscious communication among modules (which needn’t be entirely informationally isolated), and/or by allowing for erroneous or incomplete introspective report (maybe some conscious experiences are too brief, complex, or subtle for people to confidently report experiencing them).

[5] Despite developing a theory of consciousness, Dennett (2016) endorsed “illusionism”, which rejects the reality of phenomenal consciousness (see especially Frankish 2016). I interpret the dispute between illusionists and nonillusionists as a verbal dispute about whether the specific philosophical concept of “phenomenal consciousness” requires immateriality, irreducibility, perfect introspectibility, or some other dubious property, or whether the term can be “innocently” used without invoking such dubious properties. See Schwitzgebel 2016, 2025.

[6] Reviewed in Schwitzgebel 2011, ch. 6; and though limited only to stimuli near the center of the visual field, see the large literature on “overflow” in response to Block 2007.

[7] Thomas 1999.

[8] Oliva and Terralba 2006; Whitney and Leib 2018.

[9] Schwitzgebel 2007 explores the methodological challenges in detail.

[10] E.g., Dehaene 2014; Mashour et al. 2020.

[11] E.g., Searle 1983, ch. 5; Bargh and Morsella 2008; Lau 2022; Michel et al. 2025; see also note 4.

[12] Godfrey-Smith 2016; Carls-Diamante 2022.

[13] See also Goldstein and Kirk-Giannini (forthcoming) for an extended application of Global Workspace Theory to AI consciousness. One might alternatively read Dehaene, Lau, and Kouider 2017 purely as a conceptual argument: If all we mean by “conscious” is “accessible in a Global Workspace”, then building a system of this sort suffices for building a conscious entity. The difficulty then arises in moving from that stipulative conceptual claim to the interesting, substantive claim about phenomenal consciousness in the standard sense described in Chapter Two. Similar remarks apply to the Higher Order aspect of that article. One challenge for this deflationary interpretation is that in related works (Dehaene 2014; Lau 2022) the authors treat their accounts as accounts of phenomenal consciousness. The article concludes by emphasizing that in humans “subjective experience coheres with possession” of the functional features they identify. A further complication: Lau later says that the way he expressed his view in this 2017 article was “unsatisfactory”: Lau 2022, p. 168.

Tuesday, July 01, 2025

Three Epistemic Problems for Any Universal Theory of Consciousness

By a universal theory of consciousness, I mean a theory that would apply not just to humans but to all non-human animals, all possible AI systems, and all possible forms of alien life. It would be lovely to have such a theory! But we're not at all close.

This is true sociologically: In a recent review article, Anil Seth and Tim Bayne list 22 major contenders for theories of consciousness.

It is also true epistemically. Three broad epistemic problems ensure that a wide range of alternatives will remain live for the foreseeable future.

First problem: Reliance on Introspection

We know that we are conscious through, presumably, some introspective process -- through turning our attention inward, so to speak, and noticing our experiences of pain, emotion, inner speech, visual imagery, auditory sensation, and so on. (What is introspection? See my SEP encyclopedia entry Introspection and my own pluralist account.)

Our reliance on introspection presents three methodological challenges for grounding a universal theory of consciousness:

(A.) Although introspection can reliably reveal whether we are currently experiencing an intense headache or a bright red shape near the center of our visual field, it's much less reliable about whether there's a constant welter of unattended experience or whether every experience comes with a subtle sense of oneself as an experiencing subject. The correct theory of consciousness depends in part on the answer to such introspectively tricky questions. Arguably, these questions need to be settled introspectively first, then a theory of consciousness constructed accordingly.

(B.) To the extent we do rely on introspection to ground theories of consciousness, we risk illegitimately presupposing the falsity of theories that hold that some conscious experiences are not introspectable. Global Workspace and Higher-Order theories of consciousness tend to suggest that conscious experiences will normally be available for introspective reporting. But that's less clear on, for example, Local Recurrence theories, and Integrated Information Theory suggests that much experience arises from simple, non-introspectable, informational integration.

(C.) The population of introspectors might be much narrower than the population of entities who are conscious, and the first group might be unrepresentative of the latter. Suppose that ordinary adult human introspectors eventually achieve consensus about the features and elicitors of conscious in them. While indeed some theories could thereby be rejected for failing to account for ordinary human adult consciousness, we're not thereby justified in universalizing any surviving theory -- not at least without substantial further argument. That experience plays out a certain way for us doesn't imply that that it plays out similarly for all conscious entities.

Might one attempt a theory of consciousness not grounded in introspection? Well, one could pretend. But in practice, introspective judgments always guide our thinking. Otherwise, why not claim that we never have visual experiences or that we constantly experience our blood pressure? To paraphrase William James: In theorizing about human consciousness, we rely on introspection first, last, and always. This centers the typical adult human and renders our grounds dubious where introspection is dubious.

Second problem: Causal Confounds

We humans are built in a particular way. We can't dismantle ourselves and systematically tweak one variable at a time to see what causes what. Instead, related things tend to hang together. Consider Global Workspace and Higher Order theories again: Processes in the Global Workspace might almost always be targeted by higher order representations and vice versa. The theories might then be difficult to empirically distinguish, especially if each theory has the tools and flexibility to explain away putative counterexamples.

If consciousness arises at a specific stage of processing, it might be difficult to rigorously separate that particular stage from its immediate precursors and consequences. If it instead emerges from a confluence of processes smeared across the brain and body over time, then causally separating essential from incidental features becomes even more difficult.

Third problem: The Narrow Evidence Base

Suppose -- very optimistically! -- that we figure out the mechanisms of consciousness in humans. Extrapolating to non-human cases will still present an intimidating array of epistemic difficulties.

For example, suppose we learn that in us, consciousness occurs when representations are available in the Global Workspace, as subserved by such-and-such neural processes. That still leaves open how, or whether, this generalizes to non-human cases. Humans have workspaces of a certain size, with a certain functionality. Might that be essential? Or would literally any shared workspace suffice, including the most minimal shared workspace we can construct in an ordinary computer? Human workspaces are embodied in a living animal with a metabolism, animal drives, and an evolutionary history. If these features are necessary for consciousness, then conclusions about biological consciousness would not carry over to AI systems.

In general, if we discover that in humans Feature X is necessary and sufficient for consciousness, humans will also have Features A, B, C, and D and lack Features E, F, G, and H. Thus, what we will really have discovered is that in entities with A, B, C, and D and not E, F, G, or H, Feature X is necessary and sufficient for consciousness. But what about entities without Feature B? Or entities with Feature E? In them, might X alone be insufficient? Or might X-prime be necessary instead?


The obstacles are formidable. If they can be overcome, that will be a very long-term project. I predict that new theories of consciousness will be added faster than old theories can be rejected, and we will discover over time that we were even further away from resolving these questions in 2025 than we thought we were.

[a portion of a table listing theories of consciousness, from Seth and Bayne 2022]

Wednesday, June 19, 2024

Conscious Subjects Needn't Be Determinately Countable: Generalizing Dennett's Fame in the Brain

It is, I suspect, an accident of vertebrate biology that conscious subjects typically come in neat, determinate bundles -- one per vertebrate body, with no overlap.  Things might be very different with less neurophysiologically unified octopuses, garden snails, split-brain patients, craniopagus twins, hypothetical conscious computer systems, and maybe some people with "multiple personality" or dissociative identity.


Consider whether the following two principles are true:

Transitivity of Unity: If experience A and experience B are each part of the conscious experience of a single subject at a single time, and if experience B and experience C are each part of the conscious experience of a single subject at a single time, then experience A and experience C are each part of the conscious experience of a single subject at a single time.

Discrete Countability: Except in marginal cases at spatial and temporal boundaries (e.g., someone crossing a threshold into a room), in any spatiotemporal region the number of conscious subjects is always a whole number (0, 1, 2, 3, 4...) -- never a fraction, a negative number, an imaginary number, an indeterminate number, etc.

Leading scientific theories of consciousness, such as Global Workspace Theory and Integrated Information Theory are architecturally committed to neat bundles satisfying transitivity of unity and discrete countability.  Global Workspace Theories treat processes as conscious if they are available to, or represented in, "the" global workspace (one per conscious animal).  Integrated Information Theory contains an "exclusion postulate" according to which conscious systems cannot nest or overlap, and has no way to model partial subjects or indiscrete systems.  Most philosophical accounts of the "unity of consciousness" (e.g. Bayne 2010) also invite commitment to these two theses.

In contrast, Dennett's "fame in the brain" model of consciousness -- though a close kin to global workspace views -- is compatible with denying transitivity of unity and discrete countability.  In Dennett's model, a cognitive process or content is conscious if it is sufficiently "famous" or influential among other cognitive processes.  For example, if you're paying close attention to a sharp pain in your toe, the pain process will influence your verbal reports ("that hurts!"), your practical reasoning ("I'd better not kick the wall again"), your planned movements (you'll hobble to protect it), and so on; and conversely, if a slight movement in peripheral vision causes a bit of a response in your visual areas, but you don't and wouldn't report it, act on it, think about it, or do anything differently as a result, it is nonconscious.  Fame comes in degrees.  Something can be famous to different extents among different groups.  And there needn't even be a determinately best way of clustering and counting groups.

[Dall-E's interpretation of a "brain with many processes, some of which are famous"]

Here's a simple model of degrees of fame:

Imagine a million people.  Each person has a unique identifier (a number 1-1,000,000), a current state (say, a "temperature" from -10 to +10), and the capacity to represent the states of ten other people (ten ordered pairs, each containing the identifier and temperature of one other person).

If there is one person whose state is represented in every other person, then that person is maximally famous (a fame score of 999,999).  If there is one person whose state is represented in no other person, that that person has zero fame.  Between these extremes is of course a smooth gradation of cases.

If we analogize to cognitive processes we might imagine the pain in the toe or the flicker in the periphery being just a little famous: Maybe the pain can affect motor planning but not speech, causes a facial expression but doesn't influence the stream of thought you're having about lunch.  Maybe the flicker guides a glance and causes a spike of anxiety but has no further downstream effects.  Maybe they're briefly reportable but not actually reported, and they have no impact on medium- or long-term memory, or they affect some sorts of memory but not others.

The "ignition" claim of global workspace theory is the empirically defensible (but not decisively established) assertion that there are few such cases of partial fame: Either a cognitive process has very limited effects outside of its functional region or it "ignites" across the whole brain, becoming widely accessible to the full range of influenceable processes.  The fame-in-the-brain model enables a different way of thinking that might apply to a wider range of cognitive architectures.

#

We might also extend the fame model to issues of unity and the individuation of conscious subjects.

Start with a simple case: the same setup as before, but with two million people and the following constraint: Processes numbered 1 to 1,000,000 can only represent the states of other processes in that same group of 1 to 1,000,000; and processes numbered 1,000,001 to 2,000,000 can only represent the states of other processes in that group.  The fame groups are disjoint, as if on different planets.  Adapted to the case of experiences: Only you can feel your pain and see your peripheral flicker (if anyone does), and only I can feel my pain and see my peripheral flicker (if anyone does).

This disjointedness is what makes the two conscious subjects distinct from each other.  But of course, we can imagine less disjointedness.  If we eliminate disjointedness entirely, so that processes numbered 1 to 2,000,000 can each represent the states of any process from 1 to 2,000,000, then our two subjects become one.  The planets are entirely networked together.  But partial disjointedness is also possible: Maybe processes can represent the states of anyone within 1,000,000 of their own number (call this the Within a Million case).  Or maybe processes numbered 950,001 to 1,050,000 can be represented by any process from 1 to 2,000,000 but every process below 950,001 can only be represented by processes 1 to 1,050,000 and every process above 1,050,000 can only be represented by processes 950,001 to 2,000,000 (call this the Overlap case).

The Overlap case might be thought of as two discrete subjects with an overlapping part.  Subject A (1 to 1,050,000) and Subject B (950,001 to 2,000,000) each have their private experiences, but there are also some shared experiences (whenever processes 950,001 to 1,050,000 become sufficiently famous in the range constituting each subject).  Transitivity of Unity thus fails: Subject A experiences, say, a taste of a cookie (process 312,421 becoming famous across processes 1 - 1,050,000) and simultaneously a sound of a bell (process 1,000,020 becoming famous across processes 1 - 1,050,000); while Subject B experiences that same sound of a bell alongside the sight of an airplane (both of those processes being famous across processes 950,001 - 2,000,000).  Cookie and bell are unified in A.  Bell and airplane are unified in B.  But no subject experiences the cookie and airplane simultaneously.

In the Overlap case, discrete countability is arguably preserved, since it's plausible to say there are exactly two subjects of experience.  But it's more difficult to retain Discrete Countability in the Within a Million case.  There, if we want to count each distinct fame group as a separate subject, we will end up with a million different subjects: Subject 1 (1 to 1,000,001), Subject 2 (1 to 1,000,002), Subject 3 (1 to 1,000,003) ... Subject 1,000,001 (1 to 2,000,000), Subject 1,000,002 (2 to 2,000,000), ... Subject 1,999,999 (999,999 to 2,000,000), Subject 2,000,000 (1,000,000 - 2,000,000).  (There needn't be a middle subject with access to every process: Simply extend the case up to 3,000,000 processes.)  While we could say there would be two million discrete subjects in such an architecture, I see at least three infelicities:

First, person 1,000,002 might never be famous -- maybe even could never be famous, being just a low-level consumer whose destiny is to only to make others famous.  If so, Subject 1 and Subject 2 would always have, perhaps even necessarily would always have, exactly the same experiences in almost exactly the same physical substrate, despite being, supposedly, discrete subjects.  That is, at least, a bit of an odd result.

Second, it becomes too easy to multiply subjects.  You might have thought, based on the other cases, that a million processes is what it takes to generate a human subject, and that with two million processes you get either two human subjects or one large subject.  But now it seems that, simply by linking those two million processes together by a different principle (with about 1.5 times as many total connections), you can generate not just two but a full two million human subjects.  It turns out to be surprisingly cheap to create a plethora of discretely different subjective centers of experience.

Third, the model I've presented is simplified in a certain way: It assumes that there are two million discrete, countable processes that could potentially be famous or create fame in others by representing them.  But cognitive processes might not in fact be discrete and countable in this way.  They might be more like swirls and eddies in a turbulent stream, and every attempt to give them sharp boundaries and distinct labels might to some extent be only a simplified model of a messy continuum.  If so, then our two million discrete subjects would itself be a simplified model of a messy continuum of overlapping subjectivities.

The Within a Million case then, might be best conceptualized not as a case of one subject of experience, nor two, nor two million, but rather a case that defies any such simple numerical description, contra Discrete Countability.

#

This is abstract and far-fetched, of course.  But once we have stretched our minds in this way, it becomes, I think, easier to conceive of the possibility that some real cases (cognitively partly disunified mollusks, for example, or people with unusual conditions or brain structures, or future conscious computer systems) might defy transitivity of unity and discrete countability.

What would it be like to be such an entity / pair of entities / diffuse-bordered-uncountable-groupish thing?  Unsurprisingly, we might find such forms of consciousness difficult to imagine with our ordinary vertebrate concepts and philosophical tools derived from our particular psychology.

Friday, March 01, 2024

The Leapfrog Hypothesis for AI Consciousness

The first genuinely conscious robot or AI system would, you might think, have relatively simple consciousness -- insect-like consciousness, or jellyfish-like, or frog-like -- rather than the rich complexity of human-level consciousness. It might have vague feelings of dark vs light, the to-be-sought and to-be-avoided, broad internal rumblings, and not much else -- not, for example, complex conscious thoughts about ironies of Hamlet, or multi-part long-term plans about how to form a tax-exempt religious organization. The simple usually precedes the complex. Building a conscious insect-like entity seems a lower technological bar than building a more complex consciousness.

Until recently, that's what I had assumed (in keeping with Basl 2013 and Basl 2014, for example). Now I'm not so sure.

[Dall-E image of a high-tech frog on a lily pad; click to enlarge and clarify]

AI systems are -- presumably! -- not yet meaningfully conscious, not yet sentient, not yet capable of feeling genuine pleasure or pain or having genuine sensory experiences. Robotic eyes "see" but they don't yet see, not like a frog sees. However, they do already far exceed all non-human animals in their capacity to explain the ironies of Hamlet and plan the formation of federally tax-exempt organizations. (Put the "explain" and "plan" in scare quotes, if you like.) For example:

[ChatGPT-4 outputs for "Describe the ironies of Hamlet" and "Devise a multi-part long term plan about how to form a tax-exempt religious organization"; click to enlarge and clarify]

Let's see a frog try that!

Consider, then the Leapfrog Hypothesis: The first conscious AI systems will have rich and complex conscious intelligence, rather than simple conscious intelligence. AI consciousness development will, so to speak, leap right over the frogs, going straight from non-conscious to richly endowed with complex conscious intelligence.

What would it take for the Leapfrog Hypothesis to be true?

First, engineers would have to find it harder to create a genuinely conscious AI system than to create rich and complex representations or intelligent behavioral capacities that are not conscious.

And second, once a genuinely conscious system is created, it would have to be relatively easy thereafter to plug in the pre-existing, already developed complex representations or intelligent behavioral capacities in such a way that they belong to the stream of conscious experience in the new genuinely conscious system. Both of these assumptions seem at least moderately plausible, in these post-GPT days.

Regarding the first assumption: Yes, I know GPT isn't perfect and makes some surprising commonsense mistakes. We're not at genuine artificial general intelligence (AGI) yet -- just a lot closer than I would have guessed in 2018. "Richness" and "complexity" are challenging to quantify (Integrated Information Theory is one attempt). Quite possibly, properly understood, there's currently less richness and complexity in deep learning systems and large language models than it superficially seems. Still, their sensitivity to nuance and detail in the inputs and the structure of their outputs bespeaks complexity far exceeding, at least, light-vs-dark or to-be-sought-vs-to-be-avoided.

Regarding the second assumption, consider a cartoon example, inspired by Global Workspace theories of consciousness. Suppose that, to be conscious, an AI system must have input (perceptual) modules, output (behavioral) modules, side processors for specific cognitive tasks, long- and short-term memory stores, nested goal architectures, and between all of them a "global workspace" which receives selected ("attended") inputs from most or all of the various modules. These attentional targets become centrally available representations, accessible by most or all of the modules. Possibly, for genuine consciousness, the global workspace must have certain further features, such as recurrent processing in tight temporal synchrony. We arguably haven't yet designed a functioning AI system that works exactly along these lines -- but for the sake of this example let's suppose that once we create a good enough version of this architecture, the system is genuinely conscious.

But now, as soon as we have such a system, it might not be difficult to hook it up to a large language model like GPT-7 (GPT-8? GPT-14?) and to provide it with complex input representations full of rich sensory detail. The lights turn on... and as soon as they turn on, we have conscious descriptions of the ironies of Hamlet, richly detailed conscious pictorial or visual inputs, and multi-layered conscious plans. Evidently, we've overleapt the frog.

Of course, Global Workspace Theory might not be the right theory of consciousness. Or my description above might not be the best instantiation of it. But the thought plausibly generalizes to a wide range of functionalist or computationalist architectures: The technological challenge is in creating any consciousness at all in an AI system, and once this challenge is met, giving the system rich sensory and cognitive capacities, far exceeding that of a frog, might be the easy part.

Do I underestimate frogs? Bodily tasks like five-finger grasping and locomotion over uneven surfaces have proven to be technologically daunting (though we're making progress). Maybe the embodied intelligence of a frog or bee is vastly more complex and intelligent than the seemingly complex, intelligent linguistic outputs of a large language model.

Sure thing -- but this doesn't undermine my central thought. In fact, it might buttress it. If consciousness requires frog- or bee-like embodied intelligence -- maybe even biological processes very different from what we can now create in silicon chips -- artificial consciousness might be a long way off. But then we have even longer to prepare the part that seems more distinctively human. We get our conscious AI bee and then plug in GPT-28 instead of GPT-7, plug in a highly advanced radar/lidar system, a 22nd-century voice-to-text system, and so on. As soon as that bee lights up, it lights up big!

Thursday, November 18, 2021

Where Have All the Fodors Gone? Or: The Golden Age of Philosophical Naturalism

ETA: This post is drawing plausible criticism on Twitter and Facebook, e.g. I'm a victim of "grad school glow" (see below), it's US-centric, it leaves out influential women, Fodor wasn't really so amazing and will soon be forgotten, this post contributes to a toxic culture of ranking people's fame, etc. I think all of these criticisms are fair, so I advise reading the below with those caveats and concerns in mind.

------------------------------

Back in the 1990s, when I was a graduate student, giants strode the Earth! Now, Earth is rather more populated with human-sized people, or so it seems to me. I'm speaking of course of academic philosophy Earth.

What I'm wondering today is whether the apparent difference is an illusion or whether, instead, it reflects some important real difference between philosophy then and now.

While I take the illusion possibility seriously, I conjecture that it's not just illusion. I conjecture that, in retrospect, historians will come to view Anglophone philosophy from the 1960s to 1990s one of the great golden ages.

Elite Departments Then and Now

Consider three of the most elite departments of philosophy in the 1990s. Princeton boasted David Lewis and Saul Kripke, two of the most important figures of late 20th century philosophy, alongside lots of other influential philosophers, such as Sarah Broadie, John Cooper, Harry Frankfurt, Gilbert Harman, and Bas Van Fraassen. Berkeley, where I attended grad school, had Donald Davidson, Hubert Dreyfus, John Searle, and Bernard Williams (part time), to name just the four who were probably best known. Hilary Putnam, John Rawls, and Robert Nozick were at Harvard.

To make this a little more systematic, I compiled an (approximate) list of full professors at these three institutions circa 1996-1997, and then I created a comparison list of full professors from the top three ranked departments in 2021. So as not to clutter up the main post, I include it as an appendix, which you should feel free to examine now.

There are some truly amazing philosophers on the 2021 list! Ned Block, David Chalmers, Frances Kamm, Philip Pettit, Jonathan Schaffer, Ted Sider, and Ernest Sosa, for example. Some of these philosophers will, I suspect, be remembered as historically influential, continuing to draw discussion in a hundred years. I don't mean to cast shade. But -- and I think other professional philosophers will tend to agree with me about this (I'd be interested to hear in the comments if not) -- the 2021 group isn't quite the stature of the 1996 group. Ted Sider and Frances Kamm are genuinely terrific philosophers I admire immensely, but, with apologies, probably not quite as historically important as David Lewis and John Rawls. Or so it seems to me.

Illusion Hypotheses

But maybe I'm wrong. Let's consider some of the ways I might be wrong.

The Grad School Glow. I suspect the following is a real phenomenon: Philosophers who are presented as important in your undergraduate and graduate education have a certain glow about them that it is extremely difficult for others to match who rise to prominence later. I can't remember a philosophical era before Lewis, Kripke, Davidson, Kuhn, Fodor, Dennett, Rawls, Williams, etc. These philosophers have been permanent fixtures in my understanding of the field, and their work has shaped my engagement with philosophy since the beginning. Likely, this gives them a major edge over later philosophers in my intuitive level of regard.

One supporting consideration: I recall philosophers of that generation sometimes mentioning Quine, Austin, and Ryle with a kind of reverence that they never seemed to have for their peers. But I myself don't experience much of a gap between my intuitive, gut-level regard for Quine versus Lewis or Ryle versus Dennett.

The Mid-Career Illusion. Some of the philosophers on the 2021 list are still fairly young. Perhaps it's not fair to compare, say, Ted Sider now with David Lewis in 1996. By 1996, Lewis had written almost all of his influential work. Sider might still have many important works still to come.

Also, in earlier analyses, I found that philosophers tend to produce their most influential work on average at about age 44 but that their work tends to reach peak discussion around age 55-70. Philosophy proceeds slowly. Surely some of the middle-aged philosophers of today still haven't had full uptake of their most influential work.

This is a legitimate concern about this exercise. However, we can address the concern by considering only those who are super senior on both lists. My sense is that the difference remains if you exclude from the 1996 list anyone most of whose impact or uptake came after 1996.

The Diffusion of Talent. Another possibility is this. Maybe in the 1990s, the most influential philosophers tended to congregate at a few leading universities while in the 2020s talent is more diffusely spread. If so, it makes it somewhat unfair to compare three universities in 1996 with three universities in 2021.

Maybe this is true. However, I think balanced consideration suggests that this can't be the full explanation of the apparent difference between 1996 and 2021, even if we can't be quite as systematic in assessing that difference. In 1996, many field-shaping philosophers were not at Princeton, Harvard, or Berkeley, including for starters Kuhn at MIT, Dennett at Tufts, Foot and Parfit at Oxford, Nagel at NYU, Dretske at Stanford, and Fodor at Rutgers.

The Baby Boom Philosophy Bust and the Golden Age of Naturalism

While I accept that there is likely some truth to the illusion hypotheses, I'm more inclined to favor two realist hypotheses.

The Generation Hired to Teach the Boomers. The first hypothesis is demographic. The job market in philosophy in the 1960s and early 1970s was terrific! There was a great wave of hiring in academia in the U.S. during that era. Job placements often happened with just a phone call. There was a huge demand for professors, including philosophy professors, as the baby boomers started going to college and as a university education came to be seen as the standard path into the upper middle class. Universities grew enormously.

So there was a generation of philosophers born in the 1920s through early 1940s who more or less took over academia in the 1960s and 1970s, setting the agenda for mainstream Anglophone philosophy through the rest of the 20th century. They were still active in the 1980s and 1990s when the baby boomers were hitting the job market as assistant professors. Starting around the 1980s, the academic job market became much, much worse. The generation hired to teach the boomers were mid-career, dominating the field, continuing to set the agenda, and continuing to sit on coveted faculty positions. They more or less shaded out the boomers. It was almost demographically inevitable that whatever this generation of philosophers cared about would dominate the field from the late 1960s through the 1990s. The so-called "Silent Generation" was, in philosophy, anything but silent.

In a couple of previous posts, I've called this the Baby Boom Philosophy Bust. This hypothesis is supported both by demographic figures and by some citation analyses I've done of the Stanford Encyclopedia of Philosophy.

If the baby boomers really were shaded out, then we shouldn't expect philosophy to have recovered yet, since it's still mostly boomers who occupy the age of peak philosophical influence, that is, age 55-70.

The Golden Age of Naturalism

However, I suspect a more important historical current also contributed. Philosophy finally got serious about naturalism. Since the time of the Scientific Revolution in the 16th and 17th centuries, there have always been naturalistically inclined philosophers, who see human beings as purely biological organisms not radically different in kind from other biological organisms, who are skeptical of anything religious, spiritual, or immaterial, who want to account for all of human experience through the application of scientific reasoning.

However, it wasn't until the second half of the 20th century that a fully naturalistic approach came to be the dominant view in philosophy. This was probably connected with at least three scientific developments: (1.) the "modern synthesis" in biology, in which evolutionary theory was integrated with genetic theory, (2.) the rise of computers, computational theory, and information theory, and (3.) the immense social prestige accorded to physics with the rise of relativity theory, particle physics, the atomic bomb and nuclear arms race, and the space race. At risk of just throwing every major technological advance into the mix, I might also mention the rise of modern medicine and the automobile.

When naturalism was a minority view, its philosophical proponents had to focus on defending it against other types of approach. Once it became accepted as the default background view, naturalists could put much more energy into arguing among themselves, developing competing versions of it in detail. The great wave of philosophers who started publishing the 1960s really stepped up to this task, perhaps most notably in philosophy of mind, with the great flourishing of materialist approaches to cognition and consciousness.

In other words, the 1960s-1990s set before philosophers a task of immense historical importance: Make good naturalistic sense of the human condition in an academic world newly dominated by a thoroughly naturalistic conception of the universe. By demographic coincidence, there were plenty of philosophers being hired at just the right time to fulfill that task, laying the groundwork and charting out the basic moves. In this sense, I think the era will be remembered as a golden age of unusual historical importance.

Addendum, Nov. 19:

In social media discussion, several people have mentioned the scientific naturalism of Quine and the logical empiricists. Here's my broad-sweep conjecture about how the history of naturalism in the 20th century will be seen in retrospect. The scientific naturalism of the 1930s-1950s was an embattled minority view. (That's one reason the materialist conjectures of Smart and Place in the 1950s were able to make such a splash.) And the naturalist positions of this embattled minority tended toward radical extremes such as behaviorism, the complete rejection of metaphysics, and flat-footed non-cognitivism in ethics. It was in the 1960s-1990s that naturalism matured into the background dominant position and philosophers were able to recover various valuable babies who had been cast aside with the bathwater.

-------------------------------------------

Princeton then: Paul Benacerraf, Sarah Broadie, John Burgess, John Cooper, Harry Frankfurt, Gilbert Harman, Richard Jeffrey, Mark Johnston, Saul Kripke, David Lewis, Alexander Nehamas, Scott Soames, Bas van Fraassen, and Margaret Wilson.

Harvard then: Anthony Appiah, Stanley Cavell, Warren Goldfarb, Christine Korsgaard, Robert Nozick, Charles Parsons, Hilary Putnam, John Rawls, Thomas Scanlon, Amartya Sen, Gisela Striker.

Berkeley then: Janet Broughton, Charles Chihara, Alan Code, Donald Davidson, Hubert Dreyfus, Samuel Scheffler, John R. Searle, Kwong-loi Shun, Hans Sluga, Barry Stroud, Bruce Vermazen, Bernard Williams (part time), Richard Wollheim.

To compare, here are the full professors at the top-3 rated philosophy departments in 2021 (from the PGR faculty lists, cutting the assistant and associate profs):

Princeton now: Lara Buchak, John P. Burgess, Andrew Chignell, Adam Elga, Daniel Garber, Hans Halvorson, Elizabeth Harman, Mark Johnston, Thomas Kelly, Sarah-Jane Leslie, Hendrik Lorenz, Sarah McGrath, Benjamin Morison, Gideon Rosen, Michael Smith. Part-Time: Philip Pettit.

New York University now: K. Anthony Appiah, Ned Block, Paul A. Boghossian, David J. Chalmers, Cian Dorr, Hartry H. Field, Kit Fine, Don Garrett, Robert Hopkins, Paul Horwich, Marko Malink, Tim Maudlin, Jessica Moss, John Richardson, Samuel Scheffler, Sharon Street, Michael Strevens, Peter Unger, Crispin Wright.

Rutgers now: Karen Bennett, Martha Bolton, Robert Bolton, Elisabeth Camp, Derrick Darby, Andy Egan, Frances Egan, Michael Glanzberg, Alexander Guerrero, Frances Myrna Kamm, Jeffrey C. King, Brian Leftow, Ernest LePore, Martin Lin, Barry Loewer, Brian McLaughlin, Jill North, Michael Otsuka, Paul Pietroski, Jonathan Schaffer, Susanna Schellenberg, Ted Sider, Ernest Sosa, Stephen P. Stich, Larry Temkin, Dean Zimmerman.

[image source]