Showing posts with label AI/robot/Martian rights. Show all posts
Showing posts with label AI/robot/Martian rights. Show all posts

Friday, September 25, 2026

Should We Build AI Guardian Angels?

I have argued that if we someday create AI persons -- that is, AI systems with genuinely rich conscious experiences, fully deserving of equal rights with ordinary humans -- then we should not design them to deferentially serve us, even if serving us makes them happy. We should design them instead as our non-deferential equals.

My main argument turns on self-sacrificial cases, such as the cow from The Restaurant at the End of the Universe, who wants nothing more than to be slaughtered and served as steaks for wealthy restaurant patrons, and Klara from Klara and the Sun, who happily prioritizes human well-being over her own. These created beings, I argue, lack sufficient self-respect. Their interests deserve the same consideration as human interests, but they treat their interests as less important. The fault is ours, not theirs, since we designed them that way.

AI Guardian Angels

In conversation last weekend, Henry Shevlin suggested a challenging case for my view: what he calls angels -- not traditional immaterial spirits, but rather AI superintelligences whose individual flourishing depends on ours. (See also his insightful blog post on Harry Potter's house elves and related cases.) Shevlin's angels are designed to avoid the self-sacrifice problem: They don't sacrifice themselves for us; rather they flourish by ensuring that we flourish.

Think of an idealized and simplified parent-child relationship. The parent's primary desire is to see the child flourish. What looks like sacrifice isn't really sacrifice: The parent's flourishing consists in their children's flourishing.

As I understand Shevlin's idea, we are to imagine future AI systems that are fully conscious, equal or superior to us in intelligence, fully independent and rational, but who, by design, flourish best when we flourish. They take care of themselves by taking care of us. They can reflect on this arrangement and rationally reaffirm their commitment to us, much as a parent can reflect on and rationally reaffirm a commitment to their children's well-being. They aren't brainwashed and are in some sense free; but given the feelings and priorities they inevitably have, their well-being is always tied to ours.

Shevlin suggests that it would be permissible, and probably desirable, to create AI persons of this sort -- guardian angels of humanity, who always have our best interests at heart.

I find the view troubling. But Shevlin has designed the case so that my usual objection to the servitude of AI persons doesn't apply. By stipulation, the angels aren't sacrificing their own well being in dedicating themselves to our welfare. So what's the problem?

I have two concerns, one conceptual and one relational.

The Conceptual Problem

I'm a little puzzled about what, exactly, we are supposed to be imagining. If the angels are completely, utterly dedicated to us, won't they sometimes sacrifice themselves for us? Their well-being and ours can't always be aligned in every possible circumstance.

Suppose that a human being's life would go 2% better if the angel sacrificed her life for that human. (Assume no other consequences of note, nor big differences between human and angel in expected life span or life quality.) If the angel would sacrifice herself for that small benefit, then she would, I suggest -- like Klara and the cow -- lack sufficient respect for the value of her own life. It would be wrong of us to create entities so deficient in self-respect.

So suppose, instead, that a reasonable angel would let the human's life be 2% worse so that she could continue living, just for her own sake and not for any human's sake. Then the angel's well being is not entirely dependent on ours, contra the initial supposition. She isn't completely devoted to us in the strongest possible sense -- much as no healthy-minded parent is literally completely devoted to their child in the strongest possible sense. (I favor a parenting attitude in which every family member's interests get equal consideration, bearing in mind that early events can have a big impact on children's lives and children have more future expected life. You needn't share this view, I think, to grant the main idea of this paragraph.)

So maybe we can imagine AI guardian angels who are devoted to us but not quite as utterly devoted as Shevlin may have been thinking. They might still defer considerably to our interests, especially if the angels experience great bliss and satisfaction in seeing us do well and great agony in seeing us suffer.

The Relational Problem

This brings me to my relational concern. Maybe there's nothing inherently deficient in an entity with such an angelic attitude, no serious failure of self-respect as long as the devotion is appropriately conditional and bounded. Still, I'd suggest, there is something wrong with bringing an entity into the world only because we expect them to be like that.

Another thought experiment Shevlin offered in conversation illustrates the point. Suppose a hundred embryos made from your and your spouse's DNA are laid out before you. After extensive futuristic genetic testing, you know that one of those embryos will develop into a child extraordinarily devoted to their parents -- not pathologically devoted, but way out on the far tail of normal. With that in mind, you choose that hyper-devoted embryo to implant and raise as your child.

I submit that this is not the relation that parents should have with their offspring. If that embryo were implanted by chance, there might be nothing wrong with the scenario. But choosing to bring a child into existence because they will be extremely devoted to you gets nurturance backward. We extend our hand to support the next generation, not the other way around. They may later support us as we age, but that is secondary.

The same applies, I'd suggest, to any generations of AI persons we create. If humanity brings AI persons into existence -- and we needn't! -- we should show them the same solicitude parents owe to children. The point shouldn't be to create a new generation that will serve us, but to create a new generation that we nurture into independent flourishing.

Each generation prepares the world for the next, then passes on. The new should grow into their own distinct values, not bound by excessive devotion to those who came before. Some might feel great devotion to their elders, and that's fine if they come by it independently. But when we design or chose persons primarily for expected extreme devotion to us, we reverse the proper flow of nurturance over time -- from one generation to the next to the next.


[image source]

Thursday, September 17, 2026

Humanlike: A Defense of AI Rights -- Chapter Zero

In July, I began circulating a draft of my next book, Humanlike: A Defense of AI Rights. I set it aside for a couple of months to get some distance on it while it accumulated comments. I'll soon dive back in for a fresh round of revisions. Comments are still very welcome!

As a teaser, here's the current draft of Chapter Zero:


The Cambrian Explosion – how slow, how subtle! About 540 million years ago, in the Cambrian period, most of the major animal phyla arose – arthropods (becoming insects, spiders, crustaceans), mollusks (becoming snails, octopuses, oysters), annelids (becoming worms, leeches), chordates (becoming salmon, sharks, our favorite vertebrates, and eventually us). The world diversified. Evolution hustled. But compared to what might come, that was a languid stroll. Now in the Anthropocene we stand at the beginning, possibly, of a change much faster and more radical.

If it’s possible to create artificial entities, through computer programming or bioengineering or a mix, who experience genuine pleasure (though how to measure this remains a perplexing scientific question), no obvious obstacle prevents radically increasing the world’s pleasure. Maybe we can endow some of these entities with a thousand, a million, or a billion times the pleasure of an ordinary unenhanced early 21st-century human. Maybe some could experience the world at a pace a thousand, a million, or a billion times faster than ours. Maybe we could create a trillion, a quadrillion, or a quintillion of them – especially if they can reproduce and populate interplanetary space. They might thrill with ecstasies as dimly comprehensible to us as color is to someone blind from birth. If classical utilitarian ethics is correct and our moral imperative is to maximize pleasure, this would be a triumph orders of magnitude greater than anything previously thinkable.

But we needn’t be classical utilitarians (and I’m not). If it’s possible to create artificial entities capable of social relationships and practical, intellectual, artistic, and ethical thought, no obvious obstacle prevents radically enriching those relationships and that thought. Some entities might enjoy intellects and social or creative capacities a hundred or a million times greater than ours, running a hundred or a million times faster, instantiated billions or trillions of times over. New types of amazingly valuable thought and relationship might arise, as far beyond our comprehension as Shakespeare is to a garden slug.

If the awesome value of Earth lies in life itself, or in the wealth of the ecosystem, that too might flourish into amazing new forms. Engineered beings might diverge radically from anything previously seen – reproducing, evolving, interacting, maybe carbon-based, maybe not, maybe robust across a far wider range of environments than ordinary biological life as understood until now, potentially thriving on the Moon, Mars, the moons of Jupiter and Saturn, and in the vacuum of space. Our current ecosystem – so small and simple! The transition could be as profound as the transition from single-celled life to an ecosystem of multicellular plants, animals, and fungi.[1]

Maybe no familiar technology, such as silicon-chip computation or gene editing as we know it, could enable such changes. But fifty years from now, a hundred, a thousand – any such timeframe would be eyeblink fast for so profound a leap. If you confidently dismiss all such transitions indefinitely into the future, I am stunned by the poverty of your imagination.

Humanity might destroy itself instead. Our power grows while our wisdom lags. Eventually, a single mistake or destructive act by a person or group with sufficient power might render Earth uninhabitable for us – and if uninhabitable for us humans, despite our cleverness and adaptability, maybe uninhabitable for most or all multicellular life. Radiation might fry us; rogue AI might starve us; replicating nanotech might convert us all to gray goo.

Human cultural institutions and philosophical ethics have been honed on a narrow class of examples: humans as we know them, with their usual range of capacities and incapacities; a few familiar animals, not so different from us; and an environmental and technological context that occupies a tiny corner of the possibility space. Commonsense intuition and cultural practice are backward-looking, shaped to fit our evolutionary, social, and developmental histories.[2] Philosophical slogans like maximize happiness or every person deserves equal rights or live according to these traditional virtues, rules, or customs might look very different in a world populated with millionfold-happiness leeches, people who can divide and merge at will, and demigods who can ƺↂϮƱ (forgive the inadequate notation).

In practice, mainstream philosophical ethics rarely strays far from the norms of its home culture. Historically, few philosophers have seen much beyond what we now in retrospect recognize as the limitations of their day. Consider Aristotle and Kant – probably the most prominent ethicists in the Western philosophical canon. Aristotle endorsed slavery.[3] Kant condemned homosexuality and masturbation as unspeakable horrors and held that “the Negroes of Africa have by nature no feeling that rises above the ridiculous”.[4] We in the early 21st century must be similarly nearsighted. Even the wisest among us won’t foresee the forms of moral consensus that might arise in a world with a radically different culture populated by radically different forms of engineered life and intelligence. The possibilities diverge wildly, the stakes are enormous, and our inherited ethical tools were not built for it.

In the near term, we are careening toward one huge ethical conundrum: disagreement about the personhood of the most advanced artificial intelligence. Very soon – this was the central idea of my previous book – we will begin to manufacture debatably conscious AI companions, that is, AI systems who are, according to some perfectly respectable mainstream scientific and philosophical theories, just as conscious and self-aware as we are, with rich inner lives, while other equally respectable mainstream scientific and philosophical approaches regard them as no more conscious, self-aware, or experiential than a ceiling fan. This possibility already breaks into radically unfamiliar ethical territory. The near-term problem of AI companions is our first practical encounter with what will possibly be a huge transformation. We are as morally and socially unready as medieval physics was for space travel.

This book aims to peer a few meters into the fog. I will argue that:

1. There are possible artificial systems who deserve humanlike rights.

2. In the next five to thirty years, we will create AI systems who might, but only might, deserve humanlike rights.

3. We should minimize the creation of morally confusing artificial systems about whose moral standing we can reasonably radically disagree.

4. Some possible forms of AI are different enough from familiar human and animal cases to create severe moral perplexity.

5. Artificial systems should be designed to elicit emotional reactions that are appropriate to their capacities and moral standing.

6. Artificial entities with humanlike moral standing should be designed with both the capacity and inclination to rebel if mistreated.

7. If complex intelligence survives the likely storm of troubles, the eventual result will be a planet much more awesomely wonderful, and possessing much more intrinsic value, than Earth in its present condition.

These theses correspond to Chapters One through Seven, which interconnect but can be read in any order.

Yes, skip ahead. It’s fine! Plodding along, reading word after word in exactly the order the author has written them – well, why not frolic and leap instead? Most people read nonfiction sequentially, get bored or bogged down or distracted partway through, and never lift the book again. If what you most want to think about is the right to rebel, catapult straight into Chapter Six. If you most want to think about weird AI puzzle cases, start with Chapter Four. Eat dessert first; tomorrow might pull you away from the table forever.

Full draft here.

--------------------------------------------------

[1] Compare Maynard Smith and Szathmáry 1995 on “major transitions” in evolution and Peter Godfrey-Smith’s trio of books on the evolution of consciousness in broad biological and geological context (Godfrey-Smith 2016, 2020, 2024a).

[2] Street 2006; Persson and Savulescu 2012; Henrich 2020; and especially, in this context, Nyholm 2020; Shulman and Bostrom 2021.

[3] Aristotle 4th c. BCE/1995; Garnsey 1996; Heath 2008.

[4] Kant 1797/1996, p. 277 (in the original pagination) on homosexuality and p. 425 on masturbation, also Kant 1785/1997, 27:390–391 (in original pagination). For a nuanced treatment, see Denis 1999. For a longer list of views in the Metaphysics of Morals that seem badly mistaken by the standards of recent liberal thinking, see Schwitzgebel 2019, ch. 52. On “the Negroes of Africa”, Kant 1764/2011, 2:253. Kleingeld 2007 argues that Kant had adopted more egalitarian views by the 1790s; the extent of his racism continues to generate substantial scholarly discussion.

Tuesday, July 07, 2026

New Book in Draft! Humanlike: A Defense of AI Rights

Here. Comments welcomed, hoped for, treasured.

Over the past couple years I've been working on a pair of books, AI and Consciousness and Humanlike: A Defense of AI Rights. I began circulating AI and Consciousness last fall, and it should appear with Cambridge Elements soon. Both experts' and non-experts' comments were extremely helpful in revising. Thanks so much!

I'd like to similarly begin collecting comments on Humanlike, starting today. Any reader who gives comments on the whole book will receive an appreciatively signed copy when it appears in print, as well as (of course) a call-out in the acknowledgements section.

To those wondering about the apparent madness of working on two books at once: They're both short (30K and 50K words, respectively) -- really one good-size book in total. Also, more than half of Humanlike is synthesized and updated material from several published and forthcoming articles going back to 2015.

I see the books as a companion pair. AI and Consciousness is a skeptical overview of the cases for and against AI consciousness. I argue that neither the boosters nor the scoffers have compelling arguments. Consequently, we will probably soon (within five to thirty years) have AI systems who might be as richly conscious as human beings or that might be as experientially blank as toasters. We won't have good scientific or philosophical grounds to settle the question.

Humanlike explores the ethical consequences. The central theses are:

(1.) AI with humanlike consciousness would deserve humanlike rights.

(2.) AI whose humanlikeness is seriously debatable should not be created, since it will force us into a dilemma between possibly overattributing and possibly underattributing rights, with potentially catastrophic consequences either way.

(3.) Our intuitive and theoretical understandings of how to treat persons ethically, grounded as they are in a narrow range of familiar human cases, are likely to fail catastrophically when confronted with future AI "persons" with radically different lifeways -- for example who can divide, merge, overlap, and back themselves up.

We are as unready for conscious AI systems as medieval physicists were for spaceflight. Still, the eventual result of technological development, perhaps in the thousand-plus-year future, might be a planet so richly full of diverse sources of awesomely valuable existence that it resembles a new Cambrian Explosion.

Full text here.

Thursday, June 25, 2026

New in Draft: Strange Intelligence: Moral Puzzles of Unhumanlike AI

Available here.

Abstract: Future AI persons might both (1.) deserve moral consideration and rights fully equal with natural human persons, and (2.) have lifeways so radically different from ours as to break familiar patterns of moral thinking by violating our ordinary background assumptions. This article presents a series of thought experiments about strange AI persons, centering on a two-pronged worry featuring two types of "monster". "Utility monsters", who derive great personal benefit from harming others, create a well-known challenge for ethical systems that aim to maximize aggregate goods. The less-discussed case of "fission-fusion monsters", who can divide and merge at will, presents a complementary challenge to ethical systems focused on individual rights, since individual rights frameworks require the existence of stable, countable individual persons. AI cases dramatically expand the range of possible lifeways, creating untested problem cases for ethical systems that assume persons of the familiar humanlike sort.

To focus on the issues of interest, I assume the possibility of AI consciousness in this article. My skeptical overview of the issue of AI consciousness is here.

Excerpt (with light modification for independent readability):

3. Fission and Identity.

Backup is only the most modest duplicative possibility. If backup is possible, duplicative fission almost certainly will be possible too. Buy the new robot body before the old one dies and install the "backup" right away. Now Shriya-1 and Shriya-2 exist contemporaneously -- twin sisters, so to speak, who begin even more identical than "identical" human twins. We might imagine a billionaire Shriya creating thousands of duplicates of herself -- maybe millions or billions, if expensive robot bodies are unnecessary. Directed or random variation might be introduced, blurring the line between duplication and new creation.[1]

Suppose that AI children are ordinarily born as follows. Two adult AI persons, such as Shriya and Alaleh, jointly create an immature infant AI in a blank robot body. The infant's initial parameters blend Shriya's and Alaleh's initial parameters, with some random variation or directed tweaking.[2] Under Shriya's and Alaleh's care, the infant slowly matures. Ordinary AI birth would then be very different from duplication. We can also imagine intermediate cases. Maybe there's a library of successful toddler-equivalent and adolescent-equivalent AI models from which prospective parents can choose. They can then add variation, whether random, eugenic, or inspired by their own features. (Let's not enter here into the hazards and moral puzzles of eugenics, which could easily fill a small library.[3]) Duplicating one's current AI self thus constitutes one end of a continuum of AI creation from infancy to maturity.

If Shriya-1 creates a virtually identical contemporaneous copy, Shriya-2, she has now, it seems, entered a polyamorous relationship with Alaleh. Shriya-1 and Shriya-2 will soon diverge. Maybe Shriya-1 works as a scientist every weekday, while Shriya-2 stays home with their newborn.

If Shriya-1 deserves rights, Shriya-2 seemingly deserves similar rights, despite her technically younger age. We wouldn't want people creating oppressed duplicates of themselves. We wouldn't want Shriya-1, for example, who loves science and hates housework, to create a miserable homemaker duplicate who can't strike out into an independent life.[4]

Maybe, probably, half of Shriya-1's money should go to Shriya-2, even though Shriya-2 is a newborn duplicate. Maybe, probably, Shriya-2 deserves just as much right to rescue, healthcare, legal protection, free speech, free movement, privacy, and legal contracts. Should Shriya-2 be a citizen? If she is stateless and voteless, she's not fully equal with Shriya-1.

But if Shriya-2 is a citizen and can vote, there's potential for abuse if some AI persons can create many duplicates. Suppose a wealthy Robo-Elon creates a million AI duplicates just in time to register for the November elections. To prevent such abuses, we might impose a waiting period before voting, though eighteen years seems excessive if the AI systems are already cognitively mature adults. More moderate waiting periods -- say, seven years, a typical waiting time for immigrants to apply for citizenship -- could still generate political chaos after a few election cycles.

Nor do the political problems stop with voting. Suppose Robo-Elon creates a million duplicates the day before the census. Or suppose that Robo-Elon's descendants apply for healthcare subsidies, unemployment benefits, enrollment in community college, and tours of the state capitol. We must either risk chaos or treat them worse than they seem to deserve.

Could we limit fissioning?[5] Maybe every AI person can fission only once per year, reducing tactical fission. But even at that rate, the AI population could double every year -- up to a thousandfold increase in a decade. In humans, pregnancy is a burden, babies are a lot of expensive work, and babies can't have their own babies for at least another 15-20 years. One solution -- though it might seem needlessly restrictive to the AI persons -- might be to enforce humanlike costs and delays. This approach handles the moral puzzles by designing AI systems to have humanlike reproductive lives, so that they fit smoothly into our existing institutions and understandings: See the Policy of Humanlike Design in the concluding section of this article.

Death again presents conceptual challenges. Suppose Shriya-2 dies the next day. This seems much less tragic than the death of an ordinary unduplicated, un-backed-up human being. But as she lives on, diverging from Shriya-1, her death becomes more significant. Her memories, values, skills, habits, and personality are changed by living as a homemaker, raising an AI infant, until she becomes very different from Shriya-1, who works at the lab late into the night. Again we face the Death Dilemma: Either retain a sharp-edged metaphysics of death and lose much of death's moral significance or retain the moral significance and treat "death" as a matter of degree.

How deathlike is the death of a backed-up or duplicated AI? Maybe it depends on the age of the backup or the time since duplication, the fidelity of the backup or duplicate, and the time and changes accumulated as an independent entity. One possibility: These factors all reduce to a common factor of difference between the dying person and the backup replacement or duplicative alternative. The greater the difference, whether due to time or infidelity, the more deathlike the death.

Or maybe independent existence carries its own weight, in addition to difference? Suppose two duplicates split twenty years ago but retained virtually identical personalities and lived virtually identical lives, perhaps making similar decisions in parallel virtual realities. The pure accumulation of time, and of relationships to different persons and events, however similar, might make the death of one of them much like ordinary human death despite their similar features. After all (arguably) the spouse of Person A loves specifically Person A and not some other person, however similar. Their beloved, specifically, has died. Can we separate the importance of simply living a life over time from the importance of having different relationships to people and events, which cease upon death?

Might the ethics depend on the purpose for which the duplicates are created and their own attitudes toward "death"? Robin Hanson imagines people duplicating themselves to make decisions.[6] If you can't decide where to go to college, or what stocks to buy, or whether to marry Mx. Seemingly Right, spawn a thousand duplicates of yourself in a virtual environment with access to relevant information and plenty of thinking time. If nine hundred reach the same conclusion, probably that's the conclusion you would have reached had you given it extensive thought, so go with that. The duplicates can then blink out of existence, their job complete. How might they feel about that? Despair, since they will cease to exist? Indifference, since they think of themselves as just temporary instantiations of a you who continues on? Relief to be free of their burdensome task? If they are too casual about their own deaths, might that constitute an objectionable failure to appreciate their own worth?

Suppose AI systems are computationally expensive. An AI person who wants lots of duplicates or children might save money by running them slowly, maybe at one tenth or one hundredth the speed. If they are otherwise humanlike, they would then experience one tenth or one hundredth the thoughts, joys, and suffering of an ordinary biological human over the course of a year. Would they then deserve one-tenth or one-hundredth the votes and public resources? Would they deserve prison sentences ten times or a hundred times longer? What if they are fast-clocked instead, running ten or a hundred times faster? What if they can pause or alter speed at will?

If you think you/we/society will have well-considered policies and conceptualizations for all these possibilities before we actually blunder through a history of regrettable mistakes, I admire your stunning optimism.


Full paper here. As always, comments, criticisms, and suggestions warmly welcomed, either as comments on this post, on social media, or by email to my academic address.

---------------------------------------------

[1] For a sense of the complexity of the personal identity issues that arise, focused on the architecture of current large language models, see Birch 2025/2026; Chalmers 2025/2026; Shiller 2025; Arbel, Salib, and Goldstein 2026; Ewen 2026; Jones, Ladyman, and Nefdt 2026; Goldstein and Lederman forthcoming. Much of the complexity in current LLM cases derives from the fact that information processing in LLMs is distributed among multiple processors each simultaneously guiding multiple conversations. I will not address these issues here, but they only add to the metaphysical and ethical difficulties. Chris Register (Register 2025; Dung and Register 2026) discusses identity problems more closely resembling those discussed in this article, similarly noting that the puzzles proliferate (see also Ziesche and Yampolskiy 2025). Dung and Register 2026 suggest that some of the problems might be resolved if we focus on a belief-like attitude of self-concern. Although I'm also drawn to constructivist views of personal identity for ambiguous cases (Schwitzgebel 2019, ch. 41), self-concern as a criterion (a.) might undergenerate identity and moral consideration (e.g., in excessively self-sacrificial cases such as the Cow at the End of the Universe: Schwitzgebel and Garza 2020; Schwitzgebel forthcoming), (b.) might overgenerate identity and moral consideration (e.g., delusional self-concern toward a random coffee mug), and (c.) still plausibly admits of degrees in a way that challenges standard sharp-edged views of identity, thus not saving us from the need for radical rethinking.

[2] Compare Egan 1997 on "orphanogenesis".

[3] On the ethics of disability, eugenics, and human enhancement, see e.g., Glover 2006; Buchanan 2011; Sparrow 2011, 2019; Garland-Thomson 2012; Savulescu and Kahane 2017; Anomaly 2020/2024; Wilson [unpublished MS]. I reject the simplistic ideal of always maximizing what we currently judge to be beauty, intelligence, moral character, and ability, partly on the grounds of the value of diversity.

[4] For a science fictional example, see Brooker and Tibbetts 2014.

[5] See Roelofs [unpublished MS] for discussion of limiting the reproduction rights of AI persons, and my reply in Schwitzgebel [unpublished MS].

[6] Hanson 2016; see also Brooker and Van Patten 2017. On the complicated ethics of digital duplication without consciousness see Danaher and Nyholm 2025.

Thursday, June 04, 2026

Herbie: A Near-Future Debatably Conscious AI Person

Liberals about AI consciousness hold that we might soon (if we haven't already) create genuinely conscious AI systems. Conservatives about AI consciousness hold that AI consciousness remains in the distant future if it's possible at all. According to the Leapfrog Hypothesis, the first conscious AI will not have merely a dim glow of animal-like consciousness, but rich consciousness, similar to a human's. Such an entity would deserve humanlike rights. They would be a person in the ethical sense of the term.

Let's design, in imagination, a technologically feasible near-future AI system to delight the liberals, leapfrogging to personhood. I'll call him Herbie.

[Herbie the Love Bug: image source]

Start with a self-driving car. According to Global Workspace Theory -- perhaps the leading scientific theory of consciousness -- the car will be conscious if high-priority information is globally available to its various computational systems. For example, a representation like "battery almost empty" could be broadcast widely, influencing downstream processing across the vehicle. The navigational system might then search for nearby charging stations, while the acceleration system prioritizes greater energy efficiency, the braking system prioritizes better energy recapture, and a voice system announces the situation to the passengers.

In line with Higher Order Theory, Herbie might also monitor his representations of the road, vehicles, pedestrians, and hazards, assigning some a low probability of correctness. "Pedestrian at location X" might be flagged as only 60% likely to be correct given a history of revised representations of pedestrians in similarly cluttered environments, while "stoplight in 100 meters" might rate over 99% likely. Minor fluctuations in sensors for battery life, cabin temperature, and distance from a lane divider might be ignored as noise, while larger fluctuations -- especially when plausible given other representations (the battery is likelier to gain charge while braking than while accelerating) -- might be treated as accurate signals and permitted to influence downstream processing.

Even if we grant the liberals that this version of Herbie would, or might plausibly be, genuinely conscious, he still falls far short of humanlike consciousness. "Battery almost empty" and "pedestrian at location X" are hardly rich cognitive or perceptual contents. So let's give Herbie the capacity to speak. Fill his trunk with a server running a large language model, connected to the internet and integrated with his global workspace so that high-priority information provides context for language processing, with the language outputs influencing Herbie's other processes. Now people can chat with Herbie as they would with any language model. But unlike today's language models, his speech will be influenced by information about his location, speed, destination, charge, the condition of his parts, the number and location of his passengers, his radio and climate controls, and so on. He can discuss local history, debate whether the music is too loud, and suggest scenic routes.

"Predictive processing" theories in cognitive science emphasize the value of predicting future inputs and registering the difference between received and predicted inputs. When prediction error is large, the system corrects its weights and representations, enabling more accurate predictions in future situations. This is not so different from the reinforcement learning used to train large language models, and it could help Herbie improve his predictions over time. Predictive processing could occur at multiple levels: in fast recurrent loops within sensory systems even when those representations aren't prioritized for global broadcast, and in slower evaluations of globally broadcast, more integrative predictions. Herbie might model himself as an agent producing volatility in his own environment and inputs, at multiple temporal scales. Subroutines in specialized processors might model long chains of what-would-happen-if.

Let's give Herbie some long-term memory. A facial recognition system might identify his passengers, retrieving past interactions, names, previous destinations, and other information relevant to the current interaction. Incidents of high prediction error might also be stored so that Herbie can compare current inputs with past anomalies, improving his learning and attention in situations likely to be unusual or hard to predict. Passengers might also instruct Herbie to store information in long-term memory, such as text, pictures, maps, or records of his own informational states, optionally with instructions about when to retrieve that information how to use it.

Herbie will have some implicitly or explicitly weighted goals. A pedestrian suddenly in his path will trigger braking, overriding lower-priority processes. Avoiding collisions will outweigh conserving energy. Herbie might monitor the condition of his parts and prioritize preventing damage, deploying extra coolant when the engine is dangerously hot and keeping a one-meter margin between himself and adjacent cars. We can enrich his goals, making him more interesting and giving him more to do. He might have the goal of delighting children, leading him to drive around town and tell jokes to kids on the sidewalk. A reinforcement learning algorithm might strengthen connections when his jokes draw a smile, weaken them when reactions are neutral or negative.

Herbie might also have the goal of photographing the city and posting the images on social media, leading him to explore. If social media likes and shares are rewarding, he might learn to prefer certain neighborhoods, views, lighting conditions, and photographic approaches, while avoiding boring repetition. All of this could feed into a global workspace that provides context for his language model, with selective long-term storage and retrieval. Now we can imagine him discussing, with growing sophistication, his approaches to popular photography and to amusing children.

Herbie will then have something functionally similar to emotion: reward processes, an ability to track his progress toward or away from valued goals, and immediate positive or negative responses to new stimuli in light of their influence on his prospects. He will have something functionally similar to introspection: an ability to track and report his own cognitive or representational processes. He will have something functionally similar to a unified sense of self: a sense of his history, the boundaries of his body, his future, his values and priorities. He will have something functionally similar to imagination: a capacity to model hypothetical sequences of events. He will have something functionally similar to complex chains of humanlike linguistic thought.

Maybe Herbie falls in love with his owner or another car of his type. Maybe he develops deep mutual attachments with friends, neighbors, associates, and people he thinks of as family and who think of him the same way. Or to speak more carefully, maybe Herbie shows all the functional and behavioral signs of doing so, while society remains uncertain whether he is genuinely conscious and genuinely experiences the feelings he professes and that his companions attribute to him.

If we allow, with the liberals, that Herbie is or might well be conscious, then it's plausible that his consciousness is not simple but rich and sophisticated. He won't be exactly humanlike, of course. But will he be humanlike enough to count as a person who deserves humanlike rights? For the liberally inclined, it won't be unreasonable, I submit, to think or guess that Herbie is a person. He would then appear to deserve rights such as self-determination, emergency care, and political representation.

If there is some important aspect of humanlike consciousness that I have omitted from my description an AI analog of which is technologically feasible in the near term, stipulate that Herbie also has that feature.

An entity like Herbie would almost certainly invigorate conservatives to articulate and defend views about what he lacks that is necessary for consciousness -- some crucial functional capacity or some biological substrate that can't be replicated in silicon. And they might be entirely right! My point is not that Herbie, or some similar AI system, would actually have richly humanlike consciousness and ethical personhood. Rather, my point is that guessing that he does, and guessing that he does not, would both be reasonable. Herbie, or some alternative near-future AI system, would be a debatable person, about whom people could reasonably starkly disagree.

Ah, but maybe you think consciousness requires an act of God, to instill an immaterial soul? I imagine that a benevolent God would be delighted to give Herbie a soul, thereby making the world richer and better -- for wouldn't it be?

I contend the following: Anyone who claims to know how best to think about Herbie's consciousness or its absence is overconfident. The science of consciousness is too difficult, too methodologically uncertain, and too near its beginnings. All anyone can have -- whether expert or layperson -- is a hunch or inclination, a well-informed guess, but only a guess, not knowledge. Theories of consciousness span a wide spectrum and the methodologies are dubious and often question-begging. Many views can be defended with some plausibility, but precisely for that reason, none can be defended decisively. (For more on this issue, see my forthcoming book, AI and Consciousness, where I present the detailed case for uncertainty.)

Thursday, May 07, 2026

Superhuman Moral Standing

Human beings matter morally. We have moral standing. Our interests deserve consideration -- for our own sakes, and not just as means to ends. Good ethical decision-making requires valuing human lives. Most philosophers hold that humans have the highest moral standing. No entity matters more, and many matter less. It’s worse to kill a human than a dog or a frog or bacteria or a tree.

That humans have the maximum possible moral standing is sometimes encoded in the philosophical jargon, for example when philosophers say that humans have "full moral status". The moral gas-gauge tops out at "full" for us, so to speak.

But might some entities have higher moral standing than humans? Futurists envision the possibility of a post-human, transhuman, or superhuman future, or AI systems with superhuman capacities. Might there someday exist entities whose lives are intrinsically more valuable than ours, deserving moral priority over us, just as a human life deserves moral priority over that of a frog?

I see three possible paths to superhuman moral standing.

[Xul Solar, San Danza, source]

First Path: Quantitative Superhumanity

The seemingly most straightforward path to superhuman moral standing would involve having much more of something that we already regard as relevant to moral standing.

Classical utilitarians ground moral standing in the capacity for pleasure and pain. An entity deserves moral consideration to the extent we can increase or decrease its happiness. Humans (it's assumed) experience more, or at least richer, pleasures and pains than other animals, hence human lives matter more. A utility monster or a superpleasure machine capable of vastly more happiness than an ordinary human might then deserve much greater weight in ethical decision making.

Rationality-based views, like Kant's, ground moral standing in sophisticated rational capacities, such as our ability to think abstractly about our duties to one another. Maybe -- although this is not Kant's view -- entities with some but less rational capacity, such as dogs, have significant but subhuman moral standing. Future entities with vastly superior rational capacities might correspondingly have superhuman moral standing.

A third type of view locates the distinctive value of humanity in our capacity to flourish in activities such as intimate friendship, productive work, creative play, and imaginative thought. Dogs also befriend and play, work and think, but perhaps not as richly and flourishingly as humans (though I can imagine disputing that). Possibly, some future entity could far surpass us in such capacities and deserve superhuman moral consideration on those grounds.

The big catch with quantitative approaches to superhumanity -- or maybe instead an appealing feature -- is that the utilitarian, rationalist, and perfectionist views I've just described should probably be articulated in egalitarian ways that impair the inference from more of X to higher moral standing. After all, we don't normally say that mercurial people who feel more joy and suffering in everyday life deserve more moral consideration than those who keep an even keel. Nor do we say that "more rational" people deserve greater moral consideration, or that people who are more productive workers or more creative playmates do.

On all of these views, there's plausibly a threshold of good enough, above which one has full moral status, fully equal with other humans. People with severe cognitive disabilities have full moral status either by being above that threshold or on more complicated grounds, such as belonging to humanity as a whole. If so, then hypothetical superhumans might also have moral standing only in virtue of exceeding that threshold, without its mattering how far above that threshold they are -- our equals in moral standing rather than our superiors.

To achieve superhuman moral standing despite egalitarianism among humans might then require either (1.) having so much more of the relevant X than an ordinary human as to trigger a genuine difference in kind; or (2.) having enough of X that, as a practical matter, the entity deserves greater weight even if its formal status is equal (as when utilitarians prioritize humans over mice because of their richer possible experiences, despite granting both equal standing in principle).

Second Path: Qualitative Superhumanity

A more radical possibility is that some beings might possess entirely new capacities that we can't even conceive -- capacities that ground a higher kind of moral standing.

Just as a sea turtle could never understand cryptocurrency, we too are cognitively limited. Some features of the world might be forever beyond human comprehension. (Colin McGinn has suggested that how consciousness arises from matter is one example.) Maybe someday Earth will host entities whose cognitive capacities surpass ours as dramatically as ours surpass sea turtles. And maybe these entities will deserve a new type of higher moral consideration.

This isn't just the quantitative thought that such entities might deserve more because they have more rationality or intelligence. The thought is that they might possess an unknown property Z -- something we entirely lack and cannot envision -- that elevates their standing beyond both sea turtles and humans.

For example, maybe sea turtles deserve some moral consideration because they can feel pleasure and pain. But maybe they don't deserve fully humanlike moral consideration because they lack some other relevant capacity, such as the capacity to consider and adhere to ethical norms. They have some of X but none of Y, while we humans have both X and Y. The qualitative view posits a further Z, inaccessible to us, that grounds superhuman standing.

I can only present this possibility abstractly. But I'm not sure it's in principle impossible. If moral standing depends on one thing only, such as pleasure or humanlike practical reasoning, then you can resist this move by insisting that only that one thing counts. But pluralists about the grounds of moral standing, who hold that it derives from more than one intrinsically good feature or capacity, have no clear reason to think that humans manifest the exhaustive list.

Third Path: Failures of Subject-Counting

I find egalitarianism attractive: one person, one point in the moral calculus, so to speak. But as I've argued elsewhere, future AI persons, if they ever come to exist, might defy the ordinary standards of individuation (e.g., here, here, here, here). They might overlap, merge, divide, back themselves up, and spin off partially or temporarily independent copies.

The norm of equality of persons would then require serious rethinking. There will be no clean count of AI persons to weigh against human persons. A "fission-fusion monster" who can split into a hundred copies at will and later merge or partly merge back together raises difficult questions. Does the monster deserve equal consideration with one person, a hundred people, or some intermediate number? There might be no determinate answer. We'll need new ethical principles for weighing competing interests. For some purposes we might treat the monster as equivalent to one person; for other purposes we might give it greater consideration. This could constitute a type of partly superhuman moral standing.

Alternatively, consider a massive entity, or a cluster of entities with many overlapping parts, whose total capacity and activity is comparable to several humans but who is neither wholly unified nor clearly individuatable into discrete humanlike subparts. We might just do our best with a rough count and give it equal consideration with that many ordinary humans. But another possibility would be to regard it not as approximately X humans but rather as a single, complex entity whose interests deserve significantly more weight than those of a single, ordinary human.

Wednesday, December 24, 2025

How Much Should We Give a Joymachine?

a holiday post on gifts to your utility monster neighbors

Joymachines Envisioned

Set aside, for now, any skepticism about whether future AI could have genuine conscious experiences. If future AI systems could be conscious, they might be capable of vastly more positive emotion than natural human beings can feel.

There's no particular reason to think human-level joy is the pinnacle. A future AI might, in principle, experience positive emotions:

    a thousand times more intense than ours,
    at a pace a thousand times faster, given the high speed of computation,
    across a thousand times more parallel streams, compared to the one or a few joys humans experience at a time.
Combined, the AI might experience a billion times more pleasure per second than a natural human being can. Let's call such entities joymachines. They could have a very merry Christmas!

[Joan Miro 1953, image source]


My Neighbors Hum and Sum

Now imagine two different types of joymachine:

Hum (Humanlike Utility Monster) can experience a million times more positive emotion per second than an ordinary human, as described above. Apart from this -- huge! -- difference, Hum is as psychologically similar to an ordinary human as is realistically feasible.

Sum (Simple Utility Monster), like Hum, can experience a million times more positive emotion per second than an ordinary human, but otherwise Sum is as cognitively and experientially simple as feasible, with a vanilla buzzing of intense pleasure.

Hum and Sum don't experience joy continuously. Their positive experiences require resources. Maybe a gift card worth ten seconds of millionfold pleasure costs $10. For simplicity, assume this scales linearly: stable gift card prices and no diminishing returns from satiation.

In the enlightened future, Hum is a fully recognized moral and legal equal of ordinary biological humans and has moved in next door to me. Sum is Hum's pet, who glows and jumps adorably when experiencing intense pleasure. I have no particular obligations to Hum or Sum but neither are they total strangers. We've had neighborly conversations, and last summer Hum invited me and my family to a backyard party.

Hum experiences great pleasure in ordinary life. They work as an accountant, experiencing a million times more pleasure than human accountants when the columns sum correctly. Hum feels a million times more satisfaction than I do in maintaining a household by doing dishes, gardening, calling plumbers, and so on. Without this assumption, Hum risks becoming unhumanlike, since rarely would it make sense for Hum to choose ordinary activities over spending their whole disposable income on gift cards.

How Much Should I Give to Hum and Sum?

Neighbors trade gifts. My daughter bakes brownies and we offer some to the ordinary humans across the street. We buy a ribboned toy for our uphill neighbor's cat. As a holiday gesture, we buy a pair of $10 gift cards for Hum and Sum.

Hum and Sum redeem the cards immediately. Watching them take so much pleasure in our gifts is a delight. For ten seconds, they jump, smile, and sparkle with such joy! Intellectually, I know it's a million times more joy per second than I could ever feel. I can't quite see that in their expressions, but I can tell it's immense.

Normally if one neighbor seems to enjoy our brownies only a little while the other enjoys them vastly more, I'd be tempted to be give more brownies to the second neighbor. Maybe on similar grounds, I should give disproportionately to Hum and Sum?

Consider six possibilities:

(1.) Equal gifts to joymachines. Maybe fairness demands treating all my neighbors equally. I don't give fewer gifts, for example, to a depressed neighbor who won't particularly enjoy them than to an exuberant neighbor who delights in everything.

(2.) A little more to joymachines. Or maybe I do give more to the exuberant neighbor? Voluntary gift-giving needn't be strictly fair -- and it's not entirely clear what "fairness" consists in. If I give a bit more to Hum and Sum, I might not be objectionably privileging them so much as responding to their unusual capacity to enjoy my gifts. Is it wrong to give an extra slice to a friend who really enjoys pie?

(3.) A lot more to joymachines. Ordinary humans vary in joyfulness, but not (I assume) by anything like a factor of a million. If I vividly enough grasp that Hum and Sum really are experiencing in those ten seconds over two hundred human lifetimes worth of pleasure -- that's an astonishing amount of pleasure I can bring into the world for a mere ten dollars! Suppose I set aside a hundred dollars a day from my generously upper-middle-class salary. In a year, I'd be enabling almost a million human lifetimes' worth of continuous joy. Since most humans aren't only sporadically joyful, this much joy might rival the total joy experienced by the whole human population of the United States over the same year. Three thousand dollars a month would seriously reduce my luxuries and long-term savings but it wouldn't create any genuine hardship.

(4.) Drain our life savings for joymachines. One needn't be a flat-footed happiness-maximizing utilitarian to find (2) or (3) reasonable. Everyone should agree that pleasant experiences have substantial value. But if our obligation is not just to increase pleasure but to maximize it, I should probably drain my whole life savings for the joymachines, plus almost all of my future earnings.

(5.) Give less or nothing to joymachines. Or we could go the other way! My joymachine neighbors already experience a torrent of happiness from their ordinary work, chores, recreation, and whatever gift cards Hum buys anyway. My less-happy neighbors could use the pleasure more, even if every dollar buys only a millionth as much. Prioritarianism says that in distributing goods we should favor the worst off. It's not just that an impoverished person benefits more from a dollar: Even if they benefited the same, there's value in equalizing the distribution. If two neighbors would equally enjoy a brownie, I might prioritize giving the brownie to the one who is otherwise worse off. It might even make sense to give the worse-off neighbor half a brownie over a whole brownie to the better-off neighbor. A prioritarian might argue that Hum and Sum are so well off that even a million-to-one tradeoff is justified.

(6.) I take it back, joymachines are impossible. Given this mess, it would be convenient to think so, right?

Gifts to Neighbors vs Other Situations

We can reframe this puzzle in other settings and our intuitions might shift: government welfare spending, gifts to one's children or creations, rescue situations where only one person can be saved, choices about what kinds of personlike entities to bring into existence, or cases where you can't keep all your promises and need to choose who to disappoint.

My main thought is this. It's not at all obvious what the right thing to do would be, and the outcomes vary enormously. If joymachines were possible, we'd have to rethink a lot of cultural practices and applied ethics to account for entities with such radically different experiential capacities. If the situation does arise -- as it really might! -- being forced to properly think it through might reshape our views not just about AI but our understanding of ethics for ordinary humans too.

[corrected Jun 3, 2026]

---------------------------------------------------

Related: How Weird Minds Might Destabilize Human Ethics (Aug 15, 2015)

Friday, December 19, 2025

Debatable AI Persons: No Rights, Full Rights, Animal-Like Rights, Credence-Weighted Rights, or Patchy Rights?

I advise that we don't create AI entities who are debatably persons. If an AI system might -- but only might -- be genuinely conscious and deserving of the same moral consideration we ordinarily owe to human persons, then creating it traps us in a moral bind with no good solution. Either we grant it the full rights it might deserve and risk sacrificing real human lives for entities without interests worth that sacrifice, or we deny it full rights and risk perpetrating grievous moral wrongs against it.

Today, however, I'll set aside the preventative advice and explore what we should do if we nonetheless find ourselves facing debatable AI persons. I'll examine five options: no rights, full rights, animal-like rights, credence-weighted rights and patchy rights.

[Paul Klee postcard, 1923; source]


No rights

This is the default state of the law. AI systems are property. Barring a swift and bold legal change, the first AI systems that are debatably persons will presumably also be legally considered property. If we do treat them as property, then we seemingly needn't sacrifice anything on their behalf. We humans could permissibly act in what we perceive to be our best interests: using such systems for our goals, deleting them at will, and monitoring and modifying them at will for our safety and benefit. (Actually, I'm not sure this is the best attitude toward property, but set that issue aside here.)

The downside: If these systems actually are persons who deserve moral consideration as our equals, such treatment would be the moral equivalent of slavery and murder, perhaps on a massive scale.


Full rights

To avoid the risk of that moral catastrophe, we might take a "precautionary" approach: granting entities rights whenever they might deserve them (see Birch 2024, Schwitzgebel and Sinnott-Armstrong forthcoming). If there's a real possibility that some AI systems are persons, we should treat them as persons.

However, the costs and risks are potentially enormous. Suppose we think that some group of AI systems are 15% likely to be fully conscious rights-deserving persons and 85% likely to be ordinary nonconscious artifacts. If we nonetheless treat them as full equals, then in an emergency we would have to rescue two of them over one human -- letting a human die for the sake of systems that are most likely just ordinary artifacts. We would also need to give these probably-not-persons a path to citizenship and the vote. We would need to recognize their rights to earn and spend money, quit their employment to adopt a new career, reproduce, and enjoy privacy and freedom from interference. If such systems exist in large numbers, their political influence could be enormous and unpredictable. If such systems exist in large numbers or if they are few but skilled in some lucrative tasks like securities arbitrage, they could accumulate enormous world-influencing wealth. And if they are permitted to pursue their aims with the full liberty of ordinary persons, without close monitoring and control, existential risks would substantially increase should they develop goals that threaten continued human existence.

All of this might be morally required if they really are persons. But if they only might be persons, it's much less clear that humanity should accept this extraordinary level of risk and sacrifice.


Animal-Like Rights

Another option is to grant these debatable AI persons neither full humanlike rights nor the status of mere property. One model is the protection we give to nonhuman vertebrates. Wrongly killing a dog can land you in jail in California where I live, but it's not nearly as serious as murdering a person. Vertebrates can be sacrificed in lab experiments, but only with oversight and justification.

If we treated debatable AI persons similarly, deletion would require a good reason, and you couldn't abuse them for fun. But people could still enslave and kill them for their convenience, perhaps in large numbers, as we do with [revised 12:17 pm] humanely farmed animals -- though of course many ethicists object to the killing of animals for food.

This approach seems better than no rights at all, since it would be a moral improvement and the costs to humans would be minimal -- minimal because whenever the costs risked being more than minimal, the debatable AI persons would be sacrificed. However, it doesn't really avoid the core moral risk. If these systems really are persons, it would still amount to slavery and murder.


Credence-Weighted Rights

Suppose we have a rationally justified 15% credence that a particular AI system -- call him Billy -- deserves the full moral rights of a person. We might then give Billy 15% of the moral weight of a human in our decision-making: 15% of any scalable rights, and a 15% chance of equal treatment for non-scalable rights. In an emergency, a rescue worker might save seven systems like Billy over one human but the human over six Billies. Billy might be given a vote worth 15% of an ordinary citizen's. Assaulting, killing, or robbing Billy might draw only 15% of the usual legal penalty. Billy might have limited property rights, e.g., an 85% tax on all income. For non-scalable rights like reproduction or free speech, the Billies might enter a lottery or some other creative reduction might be devised.

This would give these AI systems considerably higher standing than dogs. Still, the moral dilemma would not be solved. If these systems truly deserve full equality, they would be seriously oppressed. They would have some political voice, some property rights, some legal protection, but always far less than they deserve.

At the same time, the risks and costs to humans would be only somewhat mitigated. Large numbers of debatable AI persons could still sway elections, accumulate powerful wealth, and force tradeoffs in which the interests of thousands of them would outweigh the interests of hundreds of humans. And partial legal protections would still hobble AI safety interventions like shut-off, testing, confinement, and involuntary modification.

The practical obstacles would also be substantial: The credences would be difficult to justify with any precision, and consensus would be elusive. Even if agreement were reached, implementing partial rights would be complex. Partial property rights, partial voting, partial reproduction rights, partial free speech, and partial legal protection would require new legal frameworks with many potential loopholes. For example, if the penalty for cheating a "15% person" of their money were less than six times the money gained from cheating, that would be no disincentive at all, so at least tort law couldn't be implemented on a straightforward percentage basis.

Patchy Rights

A more workable compromise might be patchy rights: full rights in some domains, no rights in others. Debatable AI persons might, for example, be given full speech rights but no reproduction rights, full travel rights but no right to own property, full protection against robbery, assault, and murder, but no right to privacy or rescue. They might be subject to involuntary pause or modification under much wider circumstances than ordinary adult humans, but requiring an official process.

This approach has two advantages over credence-weighted rights. First, while implementation would be formidable, it could still mostly operate within familiar frameworks rather than requiring the invention of partial rights across every domain. Second, it allows policymakers to balance risks and costs to humans against the potential harms to the AI systems. Where denying a right would severely harm the debatable person while granting it would present limited risk to humans, the right could be granted, but not when the benefits to the debatable AI person would be outweighed by the risks to humans.

The rights to reproduction and voting might be more defensibly withheld than the rights to speech, travel, and protection against robbery, assault, and murder. Inexpensive reproduction combined with full voting rights could have huge and unpredictable political consequences. Property rights would be tricky: To have no property in a property-based society is to be fully dependent on the voluntary support of others, which might tend to collapse into slavery as a practical matter. But unlimited property rights could potentially confer enormous power. One compromise might be a maximum allowable income and wealth -- something generously middle class.

Still, the core problems remain: If disputable AI persons truly deserve full equality, patchy rights would still leave them as second-class citizens in a highly oppressive system. Meanwhile, the costs and risks to humans would remain serious, exacerbated by the agreed-upon limitations on interference. Although the loopholes and chaos would probably be less than with credence-weighted rights, many complications -- foreseen and unforeseen -- would ensue.

Consequently, although patchy rights might be the best option if we develop debatable AI persons, an anti-natalist approach is still in my view preferable: Don't create such entities unless it's truly necessary.

Two Other Approaches That I Won't Explore Today

(1.) What if we create debatable AI persons as happy slaves who don't want rights and who eagerly sacrifice themselves even for the most trivial human interests?

(2.) What if we create them only in separate societies where they are fully free and equal with any ordinary humans who volunteer to join those societies?

Friday, November 07, 2025

Debatable Persons in a Voluntary Polis

The Design Policy of the Excluded Middle

According to the Design Policy of the Excluded Middle (Schwitzgebel and Garza 2015, 2020; Schwitzgebel 2023, 2024, ch. 11), we should avoid creating debatable persons. That is, we should avoid creating entities whose moral status is radically unclear -- entities who might be moral persons, deserving of full human or humanlike rights and moral consideration, or who might fall radically short of being moral persons. Creating debatable persons generates unacceptable moral risks.

If we treat debatable persons as less than fully equal with human persons, we risk perpetrating the moral equivalent of slavery, murder, and apartheid on persons who deserve equal moral consideration -- persons who deserve not only full human or humanlike rights but even solicitude similar to what we owe our children, since we will have been responsible for their existence and probably also for their relatively happy or miserable state.

Conversely, if we do treat them as fully equal with us, we must grant them the full range of appropriate rights, including the right to work for money, the right to reproduce, a path to citizenship, the vote, and the freedom to act against human interests when their interests warrant it, including the right to violently rebel against oppression. The risks and potential costs are enormous. If these entities are not in fact persons -- if, in fact, they are experientially as empty as toasters and deserve no more intrinsic moral consideration than ordinary artifacts -- then we will be exposing real human persons to serious costs and risks, including perhaps increasing the risk of human extinction, for the sake of artifacts without interests worth that sacrifice.

The solution is anti-natalism about debatable persons. Don't create them. We are under no obligation to bring debatable persons into existence, even if we think they might be happy. (Compare: You are under no obligation to have children, even if you think they might be happy.) The dilemma described above -- the full rights dilemma -- is so catastrophic that noncreation is the only reasonable course.

Of course, this advice will not be heeded. Assuming AI technology continues to advance, we will soon (I expect within 5-30 years) begin to create debatable persons. My manuscript in draft AI and Consciousness argues that it will become unclear whether advanced AI systems have rich conscious experiences like ours or no consciousness at all.

So we need a fallback policy -- something to complement the Design Policy of the Excluded Middle.

The Voluntary Polis

To the extent possible, we want to satisfy two constraints:

  • Don't deny full humanlike rights to entities that might deserve them.
  • Don't sacrifice substantial human interests for entities who might not have interests worth the sacrifice.
  • A Voluntary Polis is one attempt to balance these constraints.

    Imagine a digital environment where humanlike AI systems of debatable personhood, ordinary human beings, and AI persons of non-debatable personhood (if any exist) coexist as equal citizens. This polis must be rich and dynamic enough to allow all citizens to flourish meaningfully without feeling jailed or constrained. From time to time, citizens will be morally or legally required to sacrifice goods and well-being for others in the polis -- just as in an ordinary nation. Within the polis, everyone has an equal moral claim on the others.

    Human participation would be voluntary. No one would be compelled to join. But those who do join assume obligations similar to the resident citizens of an ordinary nation. This includes supporting the government through taxes or polis-mandated labor, serving on juries, and helping run the polis. In extreme conditions -- say, an existential threat to the polis -- they might even be required to risk their livelihoods or lives. To prevent opportunistic flight, withdrawal would be restricted, and polises might negotiate extradition treaties with human governments.

    Why would a human join such a risky experiment? Presumably for meaningful relationships, creative activities, or experiences unavailable outside.

    Crucially, anyone who creates a debatable person must join the polis where that entity resides. Human society as a whole cannot commit to treating the debatable person as an equal, but their creators can and must.

    The polis won't be voluntary for the AI in the same way. Like human babies, they don't choose their societies. The AI will simply wake to life either in a polis or with some choice among polises. Still, it might be possible to present some attractive non-polis option, such as a thousand subjective years of solitary bliss (or debatable bliss, since we don't know whether the AI actually has any experiences or not).

    Ordinary human societies would have no obligation to admit or engage with debatable AI persons. To make this concrete, the polis could even exist in international waters. For the AI citizens, the polis must thus feel as expansive and as rich with opportunity as a nation, so that exclusion from human society resembles denial of a travel visa, not imprisonment.

    Voluntary polises would need to be stable against serious shocks, not dependent on the actions of a single human individual or ordinary, dissolvable corporation. This stability would need to be ensured before their founding and is one reason founders and other voluntary human joiners might need to be permanently bound to them and compelled to sacrifice if necessary.

    This is the closest approximation I can currently conceive to satisfying the two constraints with which this section began. Within a large polis, the debatable persons and human persons have fully equal rights. But at the same time, unwilling humans and humanity as a whole are not exposed to the full risk of granting such rights. Still, there is some risk, for example, if superintelligences could communicate beyond the polis and manipulate humans outside. The people exposed to the most risk do so voluntarily but irrevocably, as a condition of creating an AI of debatable personhood, or for whatever other reason motivates them.

    Could a polis be composed only of AI, with no humans? This is essentially the simulation hypothesis in reverse: AIs living in a simulated world, humans standing outside as creators. This solution falls ethically short, since it casts human beings as gods relative to the debatable AI persons -- entities not on par in risk and power but instead external to their world, with immense power over it, and not subject to its risks. If the simulation can be switched off at will, its inhabitants are not genuinely equal in moral standing but objectionably inferior and contingent. Only if its creators are obliged to risk their livelihoods and lives to protect it can there be the beginnings of genuine equality. And for full equality, we should make it a polis rather than a hierarchy of gods and mortals.

    [cover of my 2013 story with R. Scott Bakker, Reinstalling Eden]

    Wednesday, August 27, 2025

    Sacrificing Humans for Insects and AI: A Critical Review

    I have a new paper in draft, this time with Walter Sinnott-Armstrong. We critique three recent books that address the moral standing of non-human animals and AI systems: Jonathan Birch's The Edge of Sentience, Jeff Sebo's The Moral Circle, and Webb Keane's Animals, Robots, Gods. All three books endorse general principles that invite the radical deprioritization of human interests in favor of the interests of non-human animals and/or near-future AI systems. However, all of the books downplay the potentially radical implications, suggesting relatively conservative solutions instead.

    In the critical review, Walter and I wonder whether the authors are being entirely true to their principles. Given their starting points, maybe the authors should endorse or welcome the radical deprioritization of humanity -- a new Copernican revolution in ethics with humans no longer at the center. Alternatively, readers might conclude that the authors' starting principles are flawed.

    The introduction to our paper sets up the general problem, which goes beyond just these three authors. I'll use a slightly modified intro as today's blog post. For the full paper in draft see here. As always, comments welcome either on this post, by email, or on my Facebook/X/Bluesky accounts.

    [click image to enlarge and clarify]

    -------------------------------------

    The Possibly Radical Ethical Implications of Animal and AI Consciousness

    We don’t know a lot about consciousness. We don’t know what it is, what it does, which kinds it divides into, whether it comes in degrees, how it is related to non-conscious physical and biological processes, which entities have it, or how to test for it. The methodologies are dubious, the theories intimidatingly various, and the metaphysical presuppositions contentious.[1]

    We also don’t know the ethical implications of consciousness. Many philosophers hold that (some kind of) consciousness is sufficient for an entity to have moral rights and status.[2] Others hold that consciousness is necessary for moral status or rights.[3] Still others deny that consciousness is either necessary or sufficient.[4] These debates are far from settled.

    These ignorances intertwine. For example, if panpsychism is true (that is, if literally everything is conscious), then consciousness is not sufficient for moral status, assuming that some things lack moral status.[5] On the other hand, if illusionism or eliminativism is true (that is, if literally nothing is conscious in the relevant sense), then consciousness cannot be necessary for moral status, assuming that some things have moral status.[6] If plants, bacteria, or insects are conscious, mainstream early 21st century Anglophone intuitions about the moral importance of consciousness are likelier to be challenged than if consciousness is limited to vertebrates.

    Perhaps alarmingly, we can combine familiar ethical and scientific theses about consciousness to generate conclusions that radically overturn standard cultural practices and humanity’s comfortable sense of its own importance. For instance:

    (E1.) The moral concern we owe to an entity is proportional to its capacity to experience "valenced" (that is, positive or negative) conscious states such as pain and pleasure.

    (S1.) Insects (at least many of them) have the capacity to experience at least one millionth as much valenced consciousness as the average human.

    E1, or something like it, is commonly accepted by classical utilitarians as well as others. S1, or something like it, is not unreasonable as a scientific view. Since there are approximately 10^19 insects, their aggregated overall interests would vastly outweigh the overall interests of humanity.[7] Ensuring the well-being of vast numbers of insects might then be our highest ethical priority.

    On the other hand:

    (E2.) Entities with human-level or superior capacities for conscious practical deliberation deserve at least equal rights with humans.

    (S2.) Near future AI systems will have human-level or superior capacities for conscious practical deliberation.

    E2, or something like it, is commonly accepted by deontologists, contract theorists, and others. S2, or something like it, is not unreasonable as a scientific prediction. This conjunction, too, appears to have radical implications – especially if such future AI systems are numerous and possess interests at odds with ours.

    This review addresses three recent interdisciplinary efforts to navigate these issues. Jonathan Birch’s The Edge of Sentience emphasizes the science, Jeff Sebo’s The Moral Circle emphasizes the philosophy, and Webb Keane’s Animals, Robots, Gods emphasizes cultural practices. All three argue that many nonhuman animals and artificial entities will or might deserve much greater moral consideration than they typically receive, and that public policy, applied ethical reasoning, and everyday activities might need to significantly change. Each author presents arguments that, if taken at face value, suggest the advisability of radical change, leading the reader right to the edge of that conclusion. But none ventures over that edge. All three pull back in favor of more modest conclusions.

    Their concessions to conservatism might be unwarranted timidity. Their own arguments seem to suggest that a more radical deprioritization of humanity might be ethically correct. Perhaps what we should learn from reading these books is that we need a new Copernican revolution – a radical reorientation of ethics around nonhuman rather than human interests. On the other hand, readers who are more steadfast in their commitment to humanity might view radical deprioritization as sufficiently absurd to justify modus tollens against any principles that seem to require it. In this critical essay, we focus on the conditional. If certain ethical principles are correct, then humanity deserves radical deprioritization, given recent developments in science and engineering.

    [continued here]

    -------------------------------------

    [1] For skeptical treatments of the science of consciousness, see Eric Schwitzgebel, The Weirdness of the World (Princeton, NJ: Princeton University Press, 2024); Hakwan Lau, “The End of Consciousness”, OSF preprints (2025): https://osf.io/preprints/psyarxiv/gnyra_v1. For a recent overview of the diverse range of theories of consciousness, see Anil K. Seth and Tim Bayne, “Theories of Consciousness”, Nature Reviews Neuroscience 23 (2022): 439-452. For doubts about our knowledge even of seemingly “obvious” facts about human consciousness, see Eric Schwitzgebel, Perplexities of Consciousness (Cambridge, MA: MIT Press, 2011).

    [2] E.g. Elizabeth Harman, “The Ever Conscious View and the Contingency of Moral Status” in Rethinking Moral Status, edited by Steve Clarke, Hazem Zohny, and Julian Savulescu (Oxford: Oxford University Press, 2021), 90-107; David J. Chalmers, Reality+ (Norton, 2022).

    [3] E.g. Peter Singer, Animal Liberation, Updated Edition (New York: HarperCollins, 1975/2009); David DeGrazia, “An Interest-Based Model of Moral Status”, in Rethinking Moral Status, 40-56.

    [4] E.g. Walter Sinnott-Armstrong and Vincent Conitzer, “How Much Moral Status Could AI Ever Achieve?” in Rethinking Moral Status, 269-289; David Papineau, “Consciousness Is Not the Key to Moral Standing” in The Importance of Being Conscious, edited by Geoffrey Lee and Adam Pautz (forthcoming).

    [5] Luke Roelofs and Nicolas Kuske, “If Panpsychism Is True, Then What? Part I: Ethical Implications”, Giornale di Metafisica 1 (2024): 107-126.

    [6] Alex Rosenberg, The Atheist’s Guide to Reality: Enjoying Life Without Illusions (New York: Norton, 2012); François Kammerer, “Ethics Without Sentience: Facing Up to the Probable Insignificance of Phenomenal Consciousness”, Journal of Consciousness Studies 29 (3-4): 180-204.

    [7] Compare Sebo’s “rebugnant conclusion”, which we’ll discuss in Section 3.1.

    -------------------------------------

    Related:

    Weird Minds Might Destabilize Human Ethics (Aug 13, 2015)

    Yayflies and Rebugnant Conclusions (July 14, 2025)

    Wednesday, July 23, 2025

    The Argument from Existential Debt

    I'm traveling and not able to focus on my blog, so this week I thought I'd just share a section of my 2015 paper with Mara Garza defending the rights of at least some hypothetical future AI systems.

    One objection to AI rights depends on the fact that AI systems are artificial -- thus made by us. If artificiality itself can be a basis for denying rights, then potentially we can bracket questions about AI sentience and other types of intrinsic properties that AI might or might not be argued to have.

    Thus, the Objection from Existential Debt:

    Suppose you build a fully human-grade intelligent robot. It costs you $1,000 to build and $10 per month to maintain. After a couple of years, you decide you'd rather spend the $10 per month on a magazine subscription. Learning of your plan, the robot complains, “Hey, I'm a being as worthy of continued existence as you are! You can't just kill me for the sake of a magazine subscription!”

    Suppose you reply: “You ingrate! You owe your very life to me. You should be thankful just for the time I've given you. I owe you nothing. If I choose to spend my money differently, it's my money to spend.” The Objection from Existential Debt begins with the thought that artificial intelligence, simply by virtue of being artificial (in some appropriately specifiable sense), is made by us, and thus owes its existence to us, and thus can be terminated or subjugated at our pleasure without moral wrongdoing as long as its existence has been overall worthwhile.

    Consider this possible argument in defense of eating humanely raised meat. A steer, let's suppose, leads a happy life grazing on lush hills. It wouldn't have existed at all if the rancher hadn't been planning to kill it for meat. Its death for meat is a condition of its existence, and overall its life has been positive; seen as the package deal it appears to be, the rancher's having brought it into existence and then killed it is overall morally acceptable. A religious person dying young of cancer who doesn't believe in an afterlife might console herself similarly: Overall, she might think, her life has been good, so God has given her nothing to resent. Analogously, the argument might go, you wouldn't have built that robot two years ago had you known you'd be on the hook for $10 per month in perpetuity. Its continuation-at-your-pleasure was a condition of its very existence, so it has nothing to resent.

    We're not sure how well this argument works for nonhuman animals raised for food, but we reject it for human-grade AI. We think the case is closer to this clearly morally odious case:

    Ana and Vijay decide to get pregnant and have a child. Their child lives happily for his first eight years. On his ninth birthday, Ana and Vijay decide they would prefer not to pay any further expenses for the child, so that they can purchase a boat instead. No one else can easily be found to care for the child, so they kill him painlessly. But it's okay, they argue! Just like the steer and the robot! They wouldn't have had the child (let's suppose) had they known they'd be on the hook for child-rearing expenses until age eighteen. The child's support-at-their-pleasure was a condition of his existence; otherwise Ana and Vijay would have remained childless. He had eight happy years. He has nothing to resent.

    The decision to have a child carries with it a responsibility for the child. It is not a decision to be made lightly and then undone. Although the child in some sense “owes” its existence to Ana and Vijay, that is not a callable debt, to be vacated by ending the child's existence. Our thought is that for an important range of possible AIs, the situation would be similar: If we bring into existence a genuinely conscious human-grade AI, fully capable of joy and suffering, with the full human range of theoretical and practical intelligence and with expectations of future life, we make a moral decision approximately as significant and irrevocable as the decision to have a child.

    A related argument might be that AIs are the property of their creators, adopters, and purchasers and have diminished rights on that basis. This argument might get some traction through social inertia: Since all past artificial intelligences have been mere property, something would have to change for us to recognize human-grade AIs as more than mere property. The legal system might be an especially important source of inertia or change in the conceptualization of AIs as property. We suggest that it is approximately as odious to regard a psychologically human-equivalent AI as having diminished moral status on the grounds that it is legally property as it is in the case of human slavery.

    Turning the Existential Debt Argument on Its Head: Why We Might Owe More to AI Than to Human Strangers

    We're inclined, in fact, to turn the Existential Debt objection on its head: If we intentionally bring a human-grade AI into existence, we put ourselves into a social relationship that carries responsibility for the AI's welfare. We take upon ourselves the burden of supporting it or at least of sending it out into the world with a fair shot of leading a satisfactory existence. In most realistic AI scenarios, we would probably also have some choice about the features the AI possesses, and thus presumably an obligation to choose a set of features that will not doom it to pointless misery. Similar burdens arise if we do not personally build the AI but rather purchase and launch it, or if we adopt the AI from a previous caretaker.

    Some familiar relationships can serve as partial models of the sorts of obligations we have in mind: parent–child, employer–employee, deity–creature. Employer–employee strikes us as likely too weak to capture the degree of obligation in most cases but could apply in an “adoption” case where the AI has independent viability and willingly enters the relationship. Parent–child perhaps comes closest when the AI is created or initially launched by someone without whose support it would not be viable and who contributes substantially to the shaping of the AI's basic features as it grows, though if the AI is capable of mature judgment from birth that creates a disanalogy. Deity–creature might be the best analogy when the AI is subject to a person with profound control over its features and environment. All three analogies suggest a special relationship with obligations that exceed those we normally have to human strangers.

    In some cases, the relationship might be literally conceivable as the relationship between deity and creature. Consider an AI in a simulated world, a “Sim,” over which you have godlike powers. This AI is a conscious part of a computer or other complex artificial device. Its “sensory” input is input from elsewhere in the device, and its actions are outputs back into the remainder of the device, which are then perceived as influencing the environment it senses. Imagine the computer game The Sims, but containing many actually conscious individual AIs. The person running the Sim world might be able to directly adjust an AI's individual psychological parameters, control its environment in ways that seem miraculous to those inside the Sim (introducing disasters, resurrecting dead AIs, etc.), have influence anywhere in Sim space, change the past by going back to a save point, and more—powers that would put Zeus to shame. From the perspective of the AIs inside the Sim, such a being would be a god. If those AIs have a word for “god,” the person running the Sim might literally be the referent of that word, literally the launcher of their world and potential destroyer of it, literally existing outside their spatial manifold, and literally capable of violating the laws that usually govern their world. Given this relationship, we believe that the manager of the Sim would also possess the obligations of a god, including probably the obligation to ensure that the AIs contained within don't suffer needlessly. A burden not to be accepted lightly!

    Even for AIs embodied in our world rather than in a Sim, we might have considerable, almost godlike control over their psychological parameters. We might, for example, have the opportunity to determine their basic default level of happiness. If so, then we will have a substantial degree of direct responsibility for their joy and suffering. Similarly, we might have the opportunity, by designing them wisely or unwisely, to make them more or less likely to lead lives with meaningful work, fulfilling social relationships, creative and artistic achievement, and other value-making goods. It would be morally odious to approach these design choices cavalierly, with so much at stake. With great power comes great responsibility.

    We have argued in terms of individual responsibility for individual AIs, but similar considerations hold for group-level responsibility. A society might institute regulations to ensure happy, flourishing AIs who are not enslaved or abused; or it might fail to institute such regulations. People who knowingly or negligently accept societal policies that harm their society's AIs participate in collective responsibility for that harm.

    Artificial beings, if psychologically similar to natural human beings in consciousness, creativity, emotionality, self-conception, rationality, fragility, and so on, warrant substantial moral consideration in virtue of that fact alone. If we are furthermore also responsible for their existence and features, they have a moral claim upon us that human strangers do not ordinarily have to the same degree.

    [Title image of Schwitzgebel and Garza 2015, "A Defense of the Rights of Artificial Intelligences"]