Thursday, July 23, 2026

The Cognitive Advantages of Being of Two Minds, or Twelve, or Twelve Hundred

I was of three minds,
Like a tree
In which there are three blackbirds.

[from Wallace Stevens, Thirteen Ways of Looking at a Blackbird]

The literature on superintelligence and AI alignment tends to assume that an intelligent AI will have one set of values, one reward function, and consistent, determinate preferences. But speaking for myself, I'm not as consistent as that. You aren't either, I suspect. Now, maybe this is a defect in us that would be absent from a well-engineered intelligent system -- but I don't think so. There are advantages to having, so to speak, a splintered mind.

Two Minds for Two Theories

Suppose you're 85% confident Theory A is true and 15% confident Theory B is true. And suppose it's not obvious what follows from the two theories; discovering the implications will require substantial cognitive work. With that work, you could learn that Theory A implies C, D, and E, while Theory B implies F, G, and H. Here are three cognitive strategies:

All-in on the most probable. Assume Theory A. Work for a while and discover that C, D, and E follow. Declare yourself done.

Both in a single model. Build a single cognitive model with the representation .85*A & .15*B. Work with this complex representation and eventually discover .85*C & .85*D & .85*E & .15*F & .15*G & .15*H.

Two models. Assume Theory A. Work a while and discover that C, D, and E follow. Then assume Theory B. Work a while and discover that F, G, and H follow. Now weight the theories and their implications .85 and .15.

All-in is easiest and cheapest. But if you discover that Theory A is wrong and then need to act quickly on the consequences of Theory B, you'll be sorry you hadn't worked out those consequences in advance.

Single model is probably hardest and slowest. It's difficult to keep track of complex conjunctive representations.

Two models has all the advantages of single model, while being easier, less prone to error, and probably faster. I'll bet it's what you did when I described single model. You thought through the conjuncts separately, considering first what follows from A, then considering what follows from B, then stitching the results together, rather than thinking truly conjunctively through the whole process.

Here's something that might be even better, if you could do it: Explore the two models simultaneously in parallel. One part of you works out the implications of Theory A, while another part of you works out the implications of Theory B. If we have design choices in creating an intelligence, maybe we should set it up that way. Arguably, much of the human mind does work that way. Even if attentive conscious processing is narrowly focused and serial, parallel nonconscious processes also operate simultaneously behind the scenes (compare Dennett's characterization of cognitive processes competing for "fame in the brain").

If I'm building a mind that I want to be robust against error while working efficiently, I might encourage it to divide itself into subminds with different opinions. The submind with the likeliest opinion gets the most control over the behavior and final judgments, but especially in risky situations, the subminds might compromise. For example, if the Theory A submind recommends Action X as slightly better than Action Y, while the Theory B submind predicts that Action X would be catastrophic, the system as a whole might choose Action Y.

Obviously, this generalizes to numbers above two.

Subminds for Plans, Values, and Preferences

The same logic applies to plans. Mind 1 develops Plan A, while Mind 2 develops Plan B. If Plan A fails, Plan B is ready to go -- especially important if quick action is needed and Plan B takes time to construct.

Perhaps less obviously, values and preferences might also be divided. Mind 1 values hard work and exercise. Mind 2 values lolling about at home reading fiction. You're a morning person: You wake with energy, and Mind 1 is in charge. Off you go for a jog, then to the office, blazing through your to-dos. By late afternoon, your energy flags and Mind 2 starts taking over. You can't just let Mind 1 be always in charge. You'd exhaust yourself, and you'd become a one-dimensional workaholic. Similarly, you can't just let Mind 2 be always in charge. You'd lose focus and vegetate.

Hypothetically, a single, complex value representation might suffice -- the value equivalent of the "single model" above -- one stable, intricate set of conditionalized values, producing different behavior in the morning and evening.

But the system needn't be built that way. And, phenomenologically, that's not how it feels to me. It feels like my morning self simply values one set of things while my evening self simply values different things. Sometimes I wish they would get along better. In the morning, I resolve to eat healthily all day, and I succeed until 9:00 pm. Then my evening self concocts an excuse to eat three big cookies. The next morning, looking at the bathroom scale, my morning self excoriates my evening self.

Of course this is a simplification. The values are complex, both morning and evening. The point is, they're not constant. There's not, it seems, one steady underlying valuation. My cares, my preferences, my priorities fluctuate over time, depending on many factors -- not just time of day, but mood, situation, bodily state, what's salient at the moment, what memories or habits happen to arise. There isn't a stable preference ranking, a single momentary snapshot, that represents the real, undistorted me. My morning and evening and other selves are equally authentic and equally me.

As in the Theory A vs Theory B case, we can shift from diachronic to synchronic -- from serial to parallel, from this-then-that to this-and-that simultaneously. My work self, my family self, my friend self, with their different preferences and values, are all simultaneously ready and updating in the nonconscious background. My work self, or selves, dominate at the office, but I could swiftly transition if an old friend were unexpectedly to knock on the door. And once in a while these other selves toss "off topic" thoughts up into my central conscious stream.

Back to AI

Maybe this isn't how the human mind works. I'm inclined to think it is, but set that aside. We certainly could build AI to work that way. Multiple competing models of the world. Multiple competing value sets. Shifting with evidence, shifting with circumstance -- perhaps not even capable of being distilled into a single conjunctive representation or single coherent value ranking.

Multiply the minds -- generalizing from two to twelve or twelve hundred -- and the subminds can form coalitions. Mind 1 might have the plurality of the weight, but if Minds 2 to 5 all agree on something else, they might collectively outweigh it. With enough diverse subminds, the group dynamics of voting, coalition building, and delegation might arise.

Let's not assume that an AI superintelligence (if the concept even makes sense) would be a hyper-coherent, hyper-rational entity that never contradicts itself and always maximizes the same fixed set of values. It might be instead a swirl of many contradictory, competing processes that compromise among one another and take turns in control.

What would it feel like to be such an entity, if such an entity is capable of having experiences? Maybe it would feel something like being us. But I can also imagine such an entity being more multi-centered than we are (we, with our narrow, serial attention in a single stream of consciousness [or so we ordinarily assume]). Such an entity might have many eddies of consciousness, partly distinct, partly overlapping, not easily divisible into discrete, countable units -- neither one unified person, nor many discrete persons, but a swarm of partly discrete entities, like a network of half-joined plants or fungi that fade one into the next without sharp boundary.

[image source, NASA]

Wednesday, July 15, 2026

AI Slop and Evidence about Evidence: Why Philosophy Journals Should Reject AI-Written Prose

I don't want to read your AI-generated email. Lots of reasons why, but here's the core of it: I want the text to reflect your thoughts. I want to know that you, the human, actually had those ideas and had them authentically enough to express them in exactly the form I see on the page. For personal emails, I want this for personal reasons. For philosophically substantive emails, I want this for evidential reasons. Philosophy journals should also want human-generated rather than AI-generated text for the same evidential reasons.

There Probably Are Good Reasons You Phrased It the Way You Did

The evidential reason: Human experts think differently and better than LLMs. Their word choices, even subtle ones, reflect sensitivities that they might not themselves be aware of. Typically, an expert's prose will be more sensitive to the matters on which they are expert than the output of a language model. When I receive a philosophical email from you -- and more so when I read a journal article -- I want your expert word choices, not blurry LLM approximations.

You might object as follows: Of course I read the LLM outputs before sending, and I wouldn't send the email, much less submit the article, unless I endorsed every word! So, the objection continues, you did think the thoughts expressed. The text reflects your expert best judgment -- maybe even something better than your expert best judgment: your expert best judgment combined with the expertise of an LLM.

I reply: There's a huge cognitive difference between nodding along while reading something and actually productively generating a text. Two reasons: First, once the text is on the page, it's easy to passively let the approximate word suffice, rather than thinking about word choice in the same effortful, active way we do when generating prose de novo. Second, as I suggested above, I doubt that human beings, even experts, have a good sense of all the factors that shape word choice -- everything they're being sensitive to. You would have phrased it slightly differently, and even if you don't know that, or why, a different signal is sent and received.

Evidence about Evidence

I'm talking about evidence about evidence: meta-epistemology. Your email or your article presents evidence for a particular philosophical view (alternatively, evidence that you support a particular philosophical view). In a simple world, I could evaluate this evidence entirely on its face: How good is the proposed view? But in the actual, complex world, it helps to have evidence about the quality of the evidence. The fact that you, an expert human, generated the text is evidence that the view is worth thinking about -- more so than if the text were generated by an LLM. This holds even if the text is exactly the same, which of course it wouldn't be.

An increasingly large part of the function of journals is to provide evidence about evidence -- the value of their imprimatur. The fact that an article appears in Nous or Ethics is evidence that it has been through rigorous review and was judged worthy by several expert humans applying unusually demanding standards of quality and importance. Its appearance in those journals is thus evidence (imperfect of course!) that the reasoning is of high quality and the arguments worth taking seriously.

Similarly, if I know that an email or an article was written by a respected colleague, reflecting their positive creative exertion in trying to choose the right words, guided by their intuitive expertise in how to phrase things, I have better reason to take it seriously than if I know that it was generated by an LLM and reflects only their passive after-the-fact assent.

Philosophers sometimes suggest that we shouldn't care if an argument was human-generated or AI-generated -- that insisting on human-generated prose is fetishizing personal human interaction rather than facts and argument quality. In a way, that's true: A sound argument is a sound argument. Similarly, we shouldn't care if an article was written by David Chalmers and published in Philosophical Review or whether it was written by someone with no institutional affiliation and published on an obscure blog. If the argument is good, it's good -- of course, of course!

But at the same time, we have limited attention, limited time, limited ability to understand the nuances when matters drift even a little from our tightest foci of expertise, and in these cases it's helpful to have meta-evidence. What should I read? How far should I trust the author has the details right, versus how much should I pause critically and chase down independent sources? How much should I let their way of phrasing things, their habitual patterns of thinking, the presuppositions hidden in their word choices and sentence structures, slip gently into my brain, silently strengthening my own associations and predilections?

An Example

To better capture the phenomenon I have in mind, I've selected a random sentence from my most recent outbound email (as of the first draft of this post):

This lack of self-respect is manifested both in her willingness to consider being killed to somewhat improve Josie's life and in her incapacity to develop desires and interests of her own independent of Josie's, as manifested by her being content to rot in a junkyard once her usefulness to Josie has ended.

The topic is Klara from Klara and the Sun, the story of a robot servant who is designed to be utterly dedicated to serving the human girl to whom she has been assigned as an "artificial friend". (I discuss the case in detail here.)

I asked ChatGPT "high" to "write the sentence better", and here is the output:

This lack of self-respect is evident both in her willingness to contemplate her own destruction for the sake of a modest improvement in Josie’s life and in her failure to develop desires or interests independent of Josie’s—as shown by her contentment to decay in a junkyard once she is no longer useful to her.

At first glance, the revised sentence is slightly better -- smoother, more natural. My double use of "manifested" was awkward, and rereading it now I think I should have changed at least that stylistic detail. I could even see approving the LLM output as expressing my own thinking about Klara, if someone asked whether it accurately reflects my view.

However, in several subtle ways the edited version is philosophically worse.

I wrote that Klara's lack of self-respect is "manifested" in her willingness to such-and-such. ChatGPT replaced "manifested" with the slightly more common and natural "evident". Close enough, seemingly? But no. They differ subtly, in an important way. "Manifested" is ontological, "evident" is epistemic. "Manifested" emphasizes the relation between the underlying lack of self-respect and the particular thoughts Klara is having -- that those thoughts arise from a broad insufficiency of self-respect. "Evident" emphasizes that the particular thoughts are evidence of that attitude. Either word might do, but "manifested" is better: I'm not really talking about what evidence we have that Klara lacks self-respect but rather what types of thinking constitute lack of self-respect.

The second change: "consider being killed" is replaced by "contemplate her own destruction". The LLM phrasing is again a bit smoother, more common, more natural. And again, it's not exactly wrong; it's a phrasing I might easily approve. But on careful thought, my original phrasing is again truer to what I really want to say. "Consider" suggests that Klara is seriously weighing the choice. "Contemplate" is vaguer: She might only be thinking idly about it. "Being killed" is vivid and suggests that someone else would be the agent of her death. "Her own destruction" is not as vivid: "Destruction" flattens things a little, morally and emotionally. And it's not quite as clear that someone else, rather than Klara herself, would be the agent of her death.

And so I could continue with the other revisions.

I am focusing on line-by-line prose, not main ideas. Main ideas are a different issue and require a different analysis. One might wrongly think that only the main ideas are important. My point here is that subtle differences in line-by-line prose matter too, and that an LLM's word choices will differ from an expert's, and that there's reason to respect experts' intuitive word choices. If experts yield their writing to a language model, it is probably too easy for them to allow their expert word sensitivities to be blurred away.

Thus, people should prefer human-expert-generated prose in emails and journal articles, for at least this meta-epistemic reason. That a passage was written by an expert human is evidence that the prose accurately tracks important nuances. Being hand-written by a human expert is an (imperfect) high-level indicator of quality.

Those of us who receive emails, or read journal articles, or even who edit and referee journal articles cannot evaluate every sentence with the same close attention to nuance I demonstrated in my example above. We often need to read at two minutes a page, not two minutes a sentence. We all reasonably rely to some extent on the writer's expertise in phrasing.

This is one reason journals should require human-written rather than AI-written prose, even if the latter is endorsed post-hoc by a human expert.

Using LLMs Well

This is not a call for blanket rejection of the use of LLMs in philosophical (or other) writing. They can be useful for generating criticisms and for suggesting topics and readings to explore, as long as one takes the outputs with a cautious grain of salt.

They can also be useful for copyediting after the initial prose is written -- but only either

(a.) for minor grammatical/spelling corrections, or

(b.) in a thoughtful, effortful way, ideally with mechanisms to reduce passivity such as (b.i) having to type in any adjustments by hand rather than simply approving them in a revised document and (b.ii) receiving suggestions from more than one LLM to force you to contemplate several alternatives rather than simply accepting one, and

(c.) in a way that preserves or amplifies, rather than averages away, your distinctive way of expressing yourself.

Such an approach to LLM-aided revision is time-consuming rather than time-saving, a matter of "cognitive onloading" rather than "cognitive offloading" -- an occasion to actively consider your word choice line by line, via input from a few not-very-insightful copyeditors. Had I subjected my email to LLM copyediting, I'd have noticed the awkward double use of "manifested" and consequently considered whether that was the best way to express myself. The result of AI-assisted copyediting following principles (a)-(c) might be prose that even more accurately reflects your best individual thinking than your independently produced drafts.

This argument against using LLMs to generate expert prose is bounded and possibly temporary: If someday LLMs surpass human experts at subtle word choice, this particular argument will no longer apply -- though other excellent reasons might remain not to use LLMs for expert prose writing.

[Mud: image source]

Tuesday, July 07, 2026

New Book in Draft! Humanlike: A Defense of AI Rights

Here. Comments welcomed, hoped for, treasured.

Over the past couple years I've been working on a pair of books, AI and Consciousness and Humanlike: A Defense of AI Rights. I began circulating AI and Consciousness last fall, and it should appear with Cambridge Elements soon. Both experts' and non-experts' comments were extremely helpful in revising. Thanks so much!

I'd like to similarly begin collecting comments on Humanlike, starting today. Any reader who gives comments on the whole book will receive an appreciatively signed copy when it appears in print, as well as (of course) a call-out in the acknowledgements section.

To those wondering about the apparent madness of working on two books at once: They're both short (30K and 50K words, respectively) -- really one good-size book in total. Also, more than half of Humanlike is synthesized and updated material from several published and forthcoming articles going back to 2015.

I see the books as a companion pair. AI and Consciousness is a skeptical overview of the cases for and against AI consciousness. I argue that neither the boosters nor the scoffers have compelling arguments. Consequently, we will probably soon (within five to thirty years) have AI systems who might be as richly conscious as human beings or that might be as experientially blank as toasters. We won't have good scientific or philosophical grounds to settle the question.

Humanlike explores the ethical consequences. The central theses are:

(1.) AI with humanlike consciousness would deserve humanlike rights.

(2.) AI whose humanlikeness is seriously debatable should not be created, since it will force us into a dilemma between possibly overattributing and possibly underattributing rights, with potentially catastrophic consequences either way.

(3.) Our intuitive and theoretical understandings of how to treat persons ethically, grounded as they are in a narrow range of familiar human cases, are likely to fail catastrophically when confronted with future AI "persons" with radically different lifeways -- for example who can divide, merge, overlap, and back themselves up.

We are as unready for conscious AI systems as medieval physicists were for spaceflight. Still, the eventual result of technological development, perhaps in the thousand-plus-year future, might be a planet so richly full of diverse sources of awesomely valuable existence that it resembles a new Cambrian Explosion.

Full text here.

Friday, July 03, 2026

Bare Functionalism Versus Digital Computational Functionalism

In philosophy of mind, people sometimes say things like the following: If functionalism is true about consciousness, then digital computers could be conscious; all it requires is that they instantiate the right programs. And yes, most functionalists about consciousness do think that. But the claim about digital computer programs does not straightforwardly follow from functionalism alone. Today I want to clarify the issue by distinguishing three forms of functionalism: bare functionalism, computational functionalism, and digital computational functionalism.

Bare Functionalism

According to bare functionalism -- that is functionalism, without further commitments -- mental states are functional states. They are determined wholly by their causal relations to stimuli, behavior, and other mental states. Whatever plays the causal role of some mental state M is mental state M. Whatever plays the causal role of pain, for example, being apt to be caused by tissue damage and tissue stress and being apt to cause writhing, groaning, complaint, avoidance, calls to the doctor, and angry vows of vengeance (to simplify somewhat) is pain -- whether it is brain state 1117A in humans, pod state 24uw in Martians, or computational state 0110100110 in a robot. It's also generally assumed that mental state terms are in principle eliminable en masse by Ramsification. That is, they are placeholders for whatever Xs, Ys, and Zs are caused by the specified stimuli and cause the specified behavior and have the right causal relations to other Xs, Ys, and Zs.

Thus, if two systems behave exactly alike given the same patterns of environmental stimuli and do so by going through the same types of abstractly specifiable internal state transitions, they are necessarily mentally identical, whether they are made of silicon, carbon, or plaster of paris. Of course some materials -- such as plaster of paris -- probably can't support the causal complexity of humanlike behavior, no matter how cleverly arranged. So that puts a practical limit on the feasible materials out of which a humanlike mind could be constructed. But the only constraint on the substrate or material underlying mentality is that it supports the relevant causal relations.

Of course, no two different materials will respond in exactly the same way to every perturbation. Functionalist substrate flexibility thus requires constraints on what count as relevantly different stimuli and relevantly different behaviors. Functionalists tend not to commit to specifics, but hand-waving generalizations convey the spirit: If we built a robot -- like one of Asimov's robots or Data from Star Trek -- who responded in broadly humanlike ways to broadly humanlike stimuli, and did so by means of internal mechanisms sufficiently like ours at an abstract causal level, the robot would have mentality like ours, whatever its underlying material constitution.

How similar must the internal mechanisms be? A single giant search tree or lookup table is probably too different. But suppose the subsystems segment the visual world into discrete objects, and store long-term memories, and store lexical items flexibly combinable via generative grammar, and monitor its bodily condition, goals, and mental states, and so on, for all of our most important cognitive functions abstractly described (or less question-beggingly, they do the abstract causal equivalent of all of these things). That's probably close enough, even if the details differ.

Computational Functionalism

Computational functionalism is stronger -- that is, riskier and more committal. It holds that functional computational structure alone suffices for consciousness. Bare functionalism needn't say this. Bare functionalism can require that stimuli, behavior, and internal causation be specified noncomputationally. Two important differences follow.

First, anti-computationalists, such as Searle, observe that a computational model of a hurricane gets no one wet and a computational model of an oven cooks no turkey. Unlike the computational functionalist, the bare functionalist can join Searle in insisting on wet turkey: To fully match the relevant input-output relations, the system much have certain concrete effects.

It might help to remember that functionalism is essentially a complexification of dispositional or causal approaches to properties. Just as we can define a poison as something that (simplistically) tends to sicken or kill when ingested, a functionalist can define a painful experience as something that tends to produce certain concrete behaviors, such as (simplistically) screaming, protest, and avoidance. If there's no tendency to concretely sicken, there's no poison; if there's no tendency to concretely scream, protest, and avoid, there's no experience of pain.

The computational functionalist rejects this commitment to specific, concrete inputs and outputs. As long as the right computational relations are instantiated, between some set of inputs and some set of outputs, I1...Ii and O1...Oj, abstractly described, mentality is present. This is because exactly the same computations have been executed, whatever the concrete Is and Os are.

Asimov's robots can throw water balloons and cook turkeys, and a functionalist who is not a computational functionalist can insist that this matters.

Second, although the bare functionalist holds that mental states depend on causal relations, unless pancomputationalism is true, not every causal relation is computational. On Piccinini's account of computation, for example, computation only occurs when "medium independent vehicles" -- physical variables defined solely in terms of their degrees of freedom (e.g., 0 vs 1) rather than their specific physical composition -- are manipulated according to rules. The bare functionalist can insist that some noncomputational causal relations matter to mentality.

Digital Computational Functionalism

Digital computational functionalism is stronger still. It further commits to treating the relevant computation digitally. Standard theories of computation, such as Turing's, are digital, and the best-known computers are digital computers. But other forms of computation are sometimes described and implemented, including analog computation and nondigital forms of computation inspired by, and perhaps instantiated in, the human brain.

Why This Matters

Keeping these distinctions clearly in mind can help us appreciate that functionalism is a larger tent than sometimes assumed. Functionalists don't need to think that digital computers could achieve consciousness. Functionalists can insist on genuine embodiment. Functionalists needn't insist that only abstractly specified causal relations matter. They do have to insist that what matters internally, among the Ramsifiable mental state relations, are the abstractly specified causal relations. But they can simultaneously insist on particular types of concrete inputs and outputs.

One might think this is an inelegant or flawed position, not the most natural development of functionalism; but that requires argument.

[COBOL Rube Goldberg by Philip Manker: image source]