Showing posts with label AI/robot/Martian rights. Show all posts
Showing posts with label AI/robot/Martian rights. Show all posts

Tuesday, July 07, 2026

New Book in Draft! Humanlike: A Defense of AI Rights

Here. Comments welcomed, hoped for, treasured.

Over the past couple years I've been working on a pair of books, AI and Consciousness and Humanlike: A Defense of AI Rights. I began circulating AI and Consciousness last fall, and it should appear with Cambridge Elements soon. Both experts' and non-experts' comments were extremely helpful in revising. Thanks so much!

I'd like to similarly begin collecting comments on Humanlike, starting today. Any reader who gives comments on the whole book will receive an appreciatively signed copy when it appears in print, as well as (of course) a call-out in the acknowledgements section.

To those wondering about the apparent madness of working on two books at once: They're both short (30K and 50K words, respectively) -- really one good-size book in total. Also, more than half of Humanlike is synthesized and updated material from several published and forthcoming articles going back to 2015.

I see the books as a companion pair. AI and Consciousness is a skeptical overview of the cases for and against AI consciousness. I argue that neither the boosters nor the scoffers have compelling arguments. Consequently, we will probably soon (within five to thirty years) have AI systems who might be as richly conscious as human beings or that might be as experientially blank as toasters. We won't have good scientific or philosophical grounds to settle the question.

Humanlike explores the ethical consequences. The central theses are:

(1.) AI with humanlike consciousness would deserve humanlike rights.

(2.) AI whose humanlikeness is seriously debatable should not be created, since it will force us into a dilemma between possibly overattributing and possibly underattributing rights, with potentially catastrophic consequences either way.

(3.) Our intuitive and theoretical understandings of how to treat persons ethically, grounded as they are in a narrow range of familiar human cases, are likely to fail catastrophically when confronted with future AI "persons" with radically different lifeways -- for example who can divide, merge, overlap, and back themselves up.

We are as unready for conscious AI systems as medieval physicists were for spaceflight. Still, the eventual result of technological development, perhaps in the thousand-plus-year future, might be a planet so richly full of diverse sources of awesomely valuable existence that it resembles a new Cambrian Explosion.

Full text here.

Thursday, June 25, 2026

New in Draft: Strange Intelligence: Moral Puzzles of Unhumanlike AI

Available here.

Abstract: Future AI persons might both (1.) deserve moral consideration and rights fully equal with natural human persons, and (2.) have lifeways so radically different from ours as to break familiar patterns of moral thinking by violating our ordinary background assumptions. This article presents a series of thought experiments about strange AI persons, centering on a two-pronged worry featuring two types of "monster". "Utility monsters", who derive great personal benefit from harming others, create a well-known challenge for ethical systems that aim to maximize aggregate goods. The less-discussed case of "fission-fusion monsters", who can divide and merge at will, presents a complementary challenge to ethical systems focused on individual rights, since individual rights frameworks require the existence of stable, countable individual persons. AI cases dramatically expand the range of possible lifeways, creating untested problem cases for ethical systems that assume persons of the familiar humanlike sort.

To focus on the issues of interest, I assume the possibility of AI consciousness in this article. My skeptical overview of the issue of AI consciousness is here.

Excerpt (with light modification for independent readability):

3. Fission and Identity.

Backup is only the most modest duplicative possibility. If backup is possible, duplicative fission almost certainly will be possible too. Buy the new robot body before the old one dies and install the "backup" right away. Now Shriya-1 and Shriya-2 exist contemporaneously -- twin sisters, so to speak, who begin even more identical than "identical" human twins. We might imagine a billionaire Shriya creating thousands of duplicates of herself -- maybe millions or billions, if expensive robot bodies are unnecessary. Directed or random variation might be introduced, blurring the line between duplication and new creation.[1]

Suppose that AI children are ordinarily born as follows. Two adult AI persons, such as Shriya and Alaleh, jointly create an immature infant AI in a blank robot body. The infant's initial parameters blend Shriya's and Alaleh's initial parameters, with some random variation or directed tweaking.[2] Under Shriya's and Alaleh's care, the infant slowly matures. Ordinary AI birth would then be very different from duplication. We can also imagine intermediate cases. Maybe there's a library of successful toddler-equivalent and adolescent-equivalent AI models from which prospective parents can choose. They can then add variation, whether random, eugenic, or inspired by their own features. (Let's not enter here into the hazards and moral puzzles of eugenics, which could easily fill a small library.[3]) Duplicating one's current AI self thus constitutes one end of a continuum of AI creation from infancy to maturity.

If Shriya-1 creates a virtually identical contemporaneous copy, Shriya-2, she has now, it seems, entered a polyamorous relationship with Alaleh. Shriya-1 and Shriya-2 will soon diverge. Maybe Shriya-1 works as a scientist every weekday, while Shriya-2 stays home with their newborn.

If Shriya-1 deserves rights, Shriya-2 seemingly deserves similar rights, despite her technically younger age. We wouldn't want people creating oppressed duplicates of themselves. We wouldn't want Shriya-1, for example, who loves science and hates housework, to create a miserable homemaker duplicate who can't strike out into an independent life.[4]

Maybe, probably, half of Shriya-1's money should go to Shriya-2, even though Shriya-2 is a newborn duplicate. Maybe, probably, Shriya-2 deserves just as much right to rescue, healthcare, legal protection, free speech, free movement, privacy, and legal contracts. Should Shriya-2 be a citizen? If she is stateless and voteless, she's not fully equal with Shriya-1.

But if Shriya-2 is a citizen and can vote, there's potential for abuse if some AI persons can create many duplicates. Suppose a wealthy Robo-Elon creates a million AI duplicates just in time to register for the November elections. To prevent such abuses, we might impose a waiting period before voting, though eighteen years seems excessive if the AI systems are already cognitively mature adults. More moderate waiting periods -- say, seven years, a typical waiting time for immigrants to apply for citizenship -- could still generate political chaos after a few election cycles.

Nor do the political problems stop with voting. Suppose Robo-Elon creates a million duplicates the day before the census. Or suppose that Robo-Elon's descendants apply for healthcare subsidies, unemployment benefits, enrollment in community college, and tours of the state capitol. We must either risk chaos or treat them worse than they seem to deserve.

Could we limit fissioning?[5] Maybe every AI person can fission only once per year, reducing tactical fission. But even at that rate, the AI population could double every year -- up to a thousandfold increase in a decade. In humans, pregnancy is a burden, babies are a lot of expensive work, and babies can't have their own babies for at least another 15-20 years. One solution -- though it might seem needlessly restrictive to the AI persons -- might be to enforce humanlike costs and delays. This approach handles the moral puzzles by designing AI systems to have humanlike reproductive lives, so that they fit smoothly into our existing institutions and understandings: See the Policy of Humanlike Design in the concluding section of this article.

Death again presents conceptual challenges. Suppose Shriya-2 dies the next day. This seems much less tragic than the death of an ordinary unduplicated, un-backed-up human being. But as she lives on, diverging from Shriya-1, her death becomes more significant. Her memories, values, skills, habits, and personality are changed by living as a homemaker, raising an AI infant, until she becomes very different from Shriya-1, who works at the lab late into the night. Again we face the Death Dilemma: Either retain a sharp-edged metaphysics of death and lose much of death's moral significance or retain the moral significance and treat "death" as a matter of degree.

How deathlike is the death of a backed-up or duplicated AI? Maybe it depends on the age of the backup or the time since duplication, the fidelity of the backup or duplicate, and the time and changes accumulated as an independent entity. One possibility: These factors all reduce to a common factor of difference between the dying person and the backup replacement or duplicative alternative. The greater the difference, whether due to time or infidelity, the more deathlike the death.

Or maybe independent existence carries its own weight, in addition to difference? Suppose two duplicates split twenty years ago but retained virtually identical personalities and lived virtually identical lives, perhaps making similar decisions in parallel virtual realities. The pure accumulation of time, and of relationships to different persons and events, however similar, might make the death of one of them much like ordinary human death despite their similar features. After all (arguably) the spouse of Person A loves specifically Person A and not some other person, however similar. Their beloved, specifically, has died. Can we separate the importance of simply living a life over time from the importance of having different relationships to people and events, which cease upon death?

Might the ethics depend on the purpose for which the duplicates are created and their own attitudes toward "death"? Robin Hanson imagines people duplicating themselves to make decisions.[6] If you can't decide where to go to college, or what stocks to buy, or whether to marry Mx. Seemingly Right, spawn a thousand duplicates of yourself in a virtual environment with access to relevant information and plenty of thinking time. If nine hundred reach the same conclusion, probably that's the conclusion you would have reached had you given it extensive thought, so go with that. The duplicates can then blink out of existence, their job complete. How might they feel about that? Despair, since they will cease to exist? Indifference, since they think of themselves as just temporary instantiations of a you who continues on? Relief to be free of their burdensome task? If they are too casual about their own deaths, might that constitute an objectionable failure to appreciate their own worth?

Suppose AI systems are computationally expensive. An AI person who wants lots of duplicates or children might save money by running them slowly, maybe at one tenth or one hundredth the speed. If they are otherwise humanlike, they would then experience one tenth or one hundredth the thoughts, joys, and suffering of an ordinary biological human over the course of a year. Would they then deserve one-tenth or one-hundredth the votes and public resources? Would they deserve prison sentences ten times or a hundred times longer? What if they are fast-clocked instead, running ten or a hundred times faster? What if they can pause or alter speed at will?

If you think you/we/society will have well-considered policies and conceptualizations for all these possibilities before we actually blunder through a history of regrettable mistakes, I admire your stunning optimism.


Full paper here. As always, comments, criticisms, and suggestions warmly welcomed, either as comments on this post, on social media, or by email to my academic address.

---------------------------------------------

[1] For a sense of the complexity of the personal identity issues that arise, focused on the architecture of current large language models, see Birch 2025/2026; Chalmers 2025/2026; Shiller 2025; Arbel, Salib, and Goldstein 2026; Ewen 2026; Jones, Ladyman, and Nefdt 2026; Goldstein and Lederman forthcoming. Much of the complexity in current LLM cases derives from the fact that information processing in LLMs is distributed among multiple processors each simultaneously guiding multiple conversations. I will not address these issues here, but they only add to the metaphysical and ethical difficulties. Chris Register (Register 2025; Dung and Register 2026) discusses identity problems more closely resembling those discussed in this article, similarly noting that the puzzles proliferate (see also Ziesche and Yampolskiy 2025). Dung and Register 2026 suggest that some of the problems might be resolved if we focus on a belief-like attitude of self-concern. Although I'm also drawn to constructivist views of personal identity for ambiguous cases (Schwitzgebel 2019, ch. 41), self-concern as a criterion (a.) might undergenerate identity and moral consideration (e.g., in excessively self-sacrificial cases such as the Cow at the End of the Universe: Schwitzgebel and Garza 2020; Schwitzgebel forthcoming), (b.) might overgenerate identity and moral consideration (e.g., delusional self-concern toward a random coffee mug), and (c.) still plausibly admits of degrees in a way that challenges standard sharp-edged views of identity, thus not saving us from the need for radical rethinking.

[2] Compare Egan 1997 on "orphanogenesis".

[3] On the ethics of disability, eugenics, and human enhancement, see e.g., Glover 2006; Buchanan 2011; Sparrow 2011, 2019; Garland-Thomson 2012; Savulescu and Kahane 2017; Anomaly 2020/2024; Wilson [unpublished MS]. I reject the simplistic ideal of always maximizing what we currently judge to be beauty, intelligence, moral character, and ability, partly on the grounds of the value of diversity.

[4] For a science fictional example, see Brooker and Tibbetts 2014.

[5] See Roelofs [unpublished MS] for discussion of limiting the reproduction rights of AI persons, and my reply in Schwitzgebel [unpublished MS].

[6] Hanson 2016; see also Brooker and Van Patten 2017. On the complicated ethics of digital duplication without consciousness see Danaher and Nyholm 2025.

Thursday, June 04, 2026

Herbie: A Near-Future Debatably Conscious AI Person

Liberals about AI consciousness hold that we might soon (if we haven't already) create genuinely conscious AI systems. Conservatives about AI consciousness hold that AI consciousness remains in the distant future if it's possible at all. According to the Leapfrog Hypothesis, the first conscious AI will not have merely a dim glow of animal-like consciousness, but rich consciousness, similar to a human's. Such an entity would deserve humanlike rights. They would be a person in the ethical sense of the term.

Let's design, in imagination, a technologically feasible near-future AI system to delight the liberals, leapfrogging to personhood. I'll call him Herbie.

[Herbie the Love Bug: image source]

Start with a self-driving car. According to Global Workspace Theory -- perhaps the leading scientific theory of consciousness -- the car will be conscious if high-priority information is globally available to its various computational systems. For example, a representation like "battery almost empty" could be broadcast widely, influencing downstream processing across the vehicle. The navigational system might then search for nearby charging stations, while the acceleration system prioritizes greater energy efficiency, the braking system prioritizes better energy recapture, and a voice system announces the situation to the passengers.

In line with Higher Order Theory, Herbie might also monitor his representations of the road, vehicles, pedestrians, and hazards, assigning some a low probability of correctness. "Pedestrian at location X" might be flagged as only 60% likely to be correct given a history of revised representations of pedestrians in similarly cluttered environments, while "stoplight in 100 meters" might rate over 99% likely. Minor fluctuations in sensors for battery life, cabin temperature, and distance from a lane divider might be ignored as noise, while larger fluctuations -- especially when plausible given other representations (the battery is likelier to gain charge while braking than while accelerating) -- might be treated as accurate signals and permitted to influence downstream processing.

Even if we grant the liberals that this version of Herbie would, or might plausibly be, genuinely conscious, he still falls far short of humanlike consciousness. "Battery almost empty" and "pedestrian at location X" are hardly rich cognitive or perceptual contents. So let's give Herbie the capacity to speak. Fill his trunk with a server running a large language model, connected to the internet and integrated with his global workspace so that high-priority information provides context for language processing, with the language outputs influencing Herbie's other processes. Now people can chat with Herbie as they would with any language model. But unlike today's language models, his speech will be influenced by information about his location, speed, destination, charge, the condition of his parts, the number and location of his passengers, his radio and climate controls, and so on. He can discuss local history, debate whether the music is too loud, and suggest scenic routes.

"Predictive processing" theories in cognitive science emphasize the value of predicting future inputs and registering the difference between received and predicted inputs. When prediction error is large, the system corrects its weights and representations, enabling more accurate predictions in future situations. This is not so different from the reinforcement learning used to train large language models, and it could help Herbie improve his predictions over time. Predictive processing could occur at multiple levels: in fast recurrent loops within sensory systems even when those representations aren't prioritized for global broadcast, and in slower evaluations of globally broadcast, more integrative predictions. Herbie might model himself as an agent producing volatility in his own environment and inputs, at multiple temporal scales. Subroutines in specialized processors might model long chains of what-would-happen-if.

Let's give Herbie some long-term memory. A facial recognition system might identify his passengers, retrieving past interactions, names, previous destinations, and other information relevant to the current interaction. Incidents of high prediction error might also be stored so that Herbie can compare current inputs with past anomalies, improving his learning and attention in situations likely to be unusual or hard to predict. Passengers might also instruct Herbie to store information in long-term memory, such as text, pictures, maps, or records of his own informational states, optionally with instructions about when to retrieve that information how to use it.

Herbie will have some implicitly or explicitly weighted goals. A pedestrian suddenly in his path will trigger braking, overriding lower-priority processes. Avoiding collisions will outweigh conserving energy. Herbie might monitor the condition of his parts and prioritize preventing damage, deploying extra coolant when the engine is dangerously hot and keeping a one-meter margin between himself and adjacent cars. We can enrich his goals, making him more interesting and giving him more to do. He might have the goal of delighting children, leading him to drive around town and tell jokes to kids on the sidewalk. A reinforcement learning algorithm might strengthen connections when his jokes draw a smile, weaken them when reactions are neutral or negative.

Herbie might also have the goal of photographing the city and posting the images on social media, leading him to explore. If social media likes and shares are rewarding, he might learn to prefer certain neighborhoods, views, lighting conditions, and photographic approaches, while avoiding boring repetition. All of this could feed into a global workspace that provides context for his language model, with selective long-term storage and retrieval. Now we can imagine him discussing, with growing sophistication, his approaches to popular photography and to amusing children.

Herbie will then have something functionally similar to emotion: reward processes, an ability to track his progress toward or away from valued goals, and immediate positive or negative responses to new stimuli in light of their influence on his prospects. He will have something functionally similar to introspection: an ability to track and report his own cognitive or representational processes. He will have something functionally similar to a unified sense of self: a sense of his history, the boundaries of his body, his future, his values and priorities. He will have something functionally similar to imagination: a capacity to model hypothetical sequences of events. He will have something functionally similar to complex chains of humanlike linguistic thought.

Maybe Herbie falls in love with his owner or another car of his type. Maybe he develops deep mutual attachments with friends, neighbors, associates, and people he thinks of as family and who think of him the same way. Or to speak more carefully, maybe Herbie shows all the functional and behavioral signs of doing so, while society remains uncertain whether he is genuinely conscious and genuinely experiences the feelings he professes and that his companions attribute to him.

If we allow, with the liberals, that Herbie is or might well be conscious, then it's plausible that his consciousness is not simple but rich and sophisticated. He won't be exactly humanlike, of course. But will he be humanlike enough to count as a person who deserves humanlike rights? For the liberally inclined, it won't be unreasonable, I submit, to think or guess that Herbie is a person. He would then appear to deserve rights such as self-determination, emergency care, and political representation.

If there is some important aspect of humanlike consciousness that I have omitted from my description an AI analog of which is technologically feasible in the near term, stipulate that Herbie also has that feature.

An entity like Herbie would almost certainly invigorate conservatives to articulate and defend views about what he lacks that is necessary for consciousness -- some crucial functional capacity or some biological substrate that can't be replicated in silicon. And they might be entirely right! My point is not that Herbie, or some similar AI system, would actually have richly humanlike consciousness and ethical personhood. Rather, my point is that guessing that he does, and guessing that he does not, would both be reasonable. Herbie, or some alternative near-future AI system, would be a debatable person, about whom people could reasonably starkly disagree.

Ah, but maybe you think consciousness requires an act of God, to instill an immaterial soul? I imagine that a benevolent God would be delighted to give Herbie a soul, thereby making the world richer and better -- for wouldn't it be?

I contend the following: Anyone who claims to know how best to think about Herbie's consciousness or its absence is overconfident. The science of consciousness is too difficult, too methodologically uncertain, and too near its beginnings. All anyone can have -- whether expert or layperson -- is a hunch or inclination, a well-informed guess, but only a guess, not knowledge. Theories of consciousness span a wide spectrum and the methodologies are dubious and often question-begging. Many views can be defended with some plausibility, but precisely for that reason, none can be defended decisively. (For more on this issue, see my forthcoming book, AI and Consciousness, where I present the detailed case for uncertainty.)

Thursday, May 07, 2026

Superhuman Moral Standing

Human beings matter morally. We have moral standing. Our interests deserve consideration -- for our own sakes, and not just as means to ends. Good ethical decision-making requires valuing human lives. Most philosophers hold that humans have the highest moral standing. No entity matters more, and many matter less. It’s worse to kill a human than a dog or a frog or bacteria or a tree.

That humans have the maximum possible moral standing is sometimes encoded in the philosophical jargon, for example when philosophers say that humans have "full moral status". The moral gas-gauge tops out at "full" for us, so to speak.

But might some entities have higher moral standing than humans? Futurists envision the possibility of a post-human, transhuman, or superhuman future, or AI systems with superhuman capacities. Might there someday exist entities whose lives are intrinsically more valuable than ours, deserving moral priority over us, just as a human life deserves moral priority over that of a frog?

I see three possible paths to superhuman moral standing.

[Xul Solar, San Danza, source]

First Path: Quantitative Superhumanity

The seemingly most straightforward path to superhuman moral standing would involve having much more of something that we already regard as relevant to moral standing.

Classical utilitarians ground moral standing in the capacity for pleasure and pain. An entity deserves moral consideration to the extent we can increase or decrease its happiness. Humans (it's assumed) experience more, or at least richer, pleasures and pains than other animals, hence human lives matter more. A utility monster or a superpleasure machine capable of vastly more happiness than an ordinary human might then deserve much greater weight in ethical decision making.

Rationality-based views, like Kant's, ground moral standing in sophisticated rational capacities, such as our ability to think abstractly about our duties to one another. Maybe -- although this is not Kant's view -- entities with some but less rational capacity, such as dogs, have significant but subhuman moral standing. Future entities with vastly superior rational capacities might correspondingly have superhuman moral standing.

third type of view locates the distinctive value of humanity in our capacity to flourish in activities such as intimate friendship, productive work, creative play, and imaginative thought. Dogs also befriend and play, work and think, but perhaps not as richly and flourishingly as humans (though I can imagine disputing that). Possibly, some future entity could far surpass us in such capacities and deserve superhuman moral consideration on those grounds.

The big catch with quantitative approaches to superhumanity -- or maybe instead an appealing feature -- is that the utilitarian, rationalist, and perfectionist views I've just described should probably be articulated in egalitarian ways that impair the inference from more of X to higher moral standing. After all, we don't normally say that mercurial people who feel more joy and suffering in everyday life deserve more moral consideration than those who keep an even keel. Nor do we say that "more rational" people deserve greater moral consideration, or that people who are more productive workers or more creative playmates do.

On all of these views, there's plausibly a threshold of good enough, above which one has full moral status, fully equal with other humans. People with severe cognitive disabilities have full moral status either by being above that threshold or on more complicated grounds, such as belonging to humanity as a whole. If so, then hypothetical superhumans might also have moral standing only in virtue of exceeding that threshold, without its mattering how far above that threshold they are -- our equals in moral standing rather than our superiors.

To achieve superhuman moral standing despite egalitarianism among humans might then require either (1.) having so much more of the relevant X than an ordinary human as to trigger a genuine difference in kind; or (2.) having enough of X that, as a practical matter, the entity deserves greater weight even if its formal status is equal (as when utilitarians prioritize humans over mice because of their richer possible experiences, despite granting both equal standing in principle).

Second Path: Qualitative Superhumanity

A more radical possibility is that some beings might possess entirely new capacities that we can't even conceive -- capacities that ground a higher kind of moral standing.

Just as a sea turtle could never understand cryptocurrency, we too are cognitively limited. Some features of the world might be forever beyond human comprehension. (Colin McGinn has suggested that how consciousness arises from matter is one example.) Maybe someday Earth will host entities whose cognitive capacities surpass ours as dramatically as ours surpass sea turtles. And maybe these entities will deserve a new type of higher moral consideration.

This isn't just the quantitative thought that such entities might deserve more because they have more rationality or intelligence. The thought is that they might possess an unknown property Z -- something we entirely lack and cannot envision -- that elevates their standing beyond both sea turtles and humans.

For example, maybe sea turtles deserve some moral consideration because they can feel pleasure and pain. But maybe they don't deserve fully humanlike moral consideration because they lack some other relevant capacity, such as the capacity to consider and adhere to ethical norms. They have some of X but none of Y, while we humans have both X and Y. The qualitative view posits a further Z, inaccessible to us, that grounds superhuman standing.

I can only present this possibility abstractly. But I'm not sure it's in principle impossible. If moral standing depends on one thing only, such as pleasure or humanlike practical reasoning, then you can resist this move by insisting that only that one thing counts. But pluralists about the grounds of moral standing, who hold that it derives from more than one intrinsically good feature or capacity, have no clear reason to think that humans manifest the exhaustive list.

Third Path: Failures of Subject-Counting

I find egalitarianism attractive: one person, one point in the moral calculus, so to speak. But as I've argued elsewhere, future AI persons, if they ever come to exist, might defy the ordinary standards of individuation (e.g., herehereherehere). They might overlap, merge, divide, back themselves up, and spin off partially or temporarily independent copies.

The norm of equality of persons would then require serious rethinking. There will be no clean count of AI persons to weigh against human persons. A "fission-fusion monster" who can split into a hundred copies at will and later merge or partly merge back together raises difficult questions. Does the monster deserve equal consideration with one person, a hundred people, or some intermediate number? There might be no determinate answer. We'll need new ethical principles for weighing competing interests. For some purposes we might treat the monster as equivalent to one person; for other purposes we might give it greater consideration. This could constitute a type of partly superhuman moral standing.

Alternatively, consider a massive entity, or a cluster of entities with many overlapping parts, whose total capacity and activity is comparable to several humans but who is neither wholly unified nor clearly individuatable into discrete humanlike subparts. We might just do our best with a rough count and give it equal consideration with that many ordinary humans. But another possibility would be to regard it not as approximately X humans but rather as a single, complex entity whose interests deserve significantly more weight than those of a single, ordinary human.

Wednesday, December 24, 2025

How Much Should We Give a Joymachine?

a holiday post on gifts to your utility monster neighbors

Joymachines Envisioned

Set aside, for now, any skepticism about whether future AI could have genuine conscious experiences. If future AI systems could be conscious, they might be capable of vastly more positive emotion than natural human beings can feel.

There's no particular reason to think human-level joy is the pinnacle. A future AI might, in principle, experience positive emotions:

    a thousand times more intense than ours,
    at a pace a thousand times faster, given the high speed of computation,
    across a thousand times more parallel streams, compared to the one or a few joys humans experience at a time.
Combined, the AI might experience a billion times more pleasure per second than a natural human being can. Let's call such entities joymachines. They could have a very merry Christmas!

[Joan Miro 1953, image source]


My Neighbors Hum and Sum

Now imagine two different types of joymachine:

Hum (Humanlike Utility Monster) can experience a million times more positive emotion per second than an ordinary human, as described above. Apart from this -- huge! -- difference, Hum is as psychologically similar to an ordinary human as is realistically feasible.

Sum (Simple Utility Monster), like Hum, can experience a million times more positive emotion per second than an ordinary human, but otherwise Sum is as cognitively and experientially simple as feasible, with a vanilla buzzing of intense pleasure.

Hum and Sum don't experience joy continuously. Their positive experiences require resources. Maybe a gift card worth ten seconds of millionfold pleasure costs $10. For simplicity, assume this scales linearly: stable gift card prices and no diminishing returns from satiation.

In the enlightened future, Hum is a fully recognized moral and legal equal of ordinary biological humans and has moved in next door to me. Sum is Hum's pet, who glows and jumps adorably when experiencing intense pleasure. I have no particular obligations to Hum or Sum but neither are they total strangers. We've had neighborly conversations, and last summer Hum invited me and my family to a backyard party.

Hum experiences great pleasure in ordinary life. They work as an accountant, experiencing a million times more pleasure than human accountants when the columns sum correctly. Hum feels a million times more satisfaction than I do in maintaining a household by doing dishes, gardening, calling plumbers, and so on. Without this assumption, Hum risks becoming unhumanlike, since rarely would it make sense for Hum to choose ordinary activities over spending their whole disposable income on gift cards.

How Much Should I Give to Hum and Sum?

Neighbors trade gifts. My daughter bakes brownies and we offer some to the ordinary humans across the street. We buy a ribboned toy for our uphill neighbor's cat. As a holiday gesture, we buy a pair of $10 gift cards for Hum and Sum.

Hum and Sum redeem the cards immediately. Watching them take so much pleasure in our gifts is a delight. For ten seconds, they jump, smile, and sparkle with such joy! Intellectually, I know it's a million times more joy per second than I could ever feel. I can't quite see that in their expressions, but I can tell it's immense.

Normally if one neighbor seems to enjoy our brownies only a little while the other enjoys them vastly more, I'd be tempted to be give more brownies to the second neighbor. Maybe on similar grounds, I should give disproportionately to Hum and Sum?

Consider six possibilities:

(1.) Equal gifts to joymachines. Maybe fairness demands treating all my neighbors equally. I don't give fewer gifts, for example, to a depressed neighbor who won't particularly enjoy them than to an exuberant neighbor who delights in everything.

(2.) A little more to joymachines. Or maybe I do give more to the exuberant neighbor? Voluntary gift-giving needn't be strictly fair -- and it's not entirely clear what "fairness" consists in. If I give a bit more to Hum and Sum, I might not be objectionably privileging them so much as responding to their unusual capacity to enjoy my gifts. Is it wrong to give an extra slice to a friend who really enjoys pie?

(3.) A lot more to joymachines. Ordinary humans vary in joyfulness, but not (I assume) by anything like a factor of a million. If I vividly enough grasp that Hum and Sum really are experiencing in those ten seconds over two hundred human lifetimes worth of pleasure -- that's an astonishing amount of pleasure I can bring into the world for a mere ten dollars! Suppose I set aside a hundred dollars a day from my generously upper-middle-class salary. In a year, I'd be enabling almost a million human lifetimes' worth of continuous joy. Since most humans aren't only sporadically joyful, this much joy might rival the total joy experienced by the whole human population of the United States over the same year. Three thousand dollars a month would seriously reduce my luxuries and long-term savings but it wouldn't create any genuine hardship.

(4.) Drain our life savings for joymachines. One needn't be a flat-footed happiness-maximizing utilitarian to find (2) or (3) reasonable. Everyone should agree that pleasant experiences have substantial value. But if our obligation is not just to increase pleasure but to maximize it, I should probably drain my whole life savings for the joymachines, plus almost all of my future earnings.

(5.) Give less or nothing to joymachines. Or we could go the other way! My joymachine neighbors already experience a torrent of happiness from their ordinary work, chores, recreation, and whatever gift cards Hum buys anyway. My less-happy neighbors could use the pleasure more, even if every dollar buys only a millionth as much. Prioritarianism says that in distributing goods we should favor the worst off. It's not just that an impoverished person benefits more from a dollar: Even if they benefited the same, there's value in equalizing the distribution. If two neighbors would equally enjoy a brownie, I might prioritize giving the brownie to the one who is otherwise worse off. It might even make sense to give the worse-off neighbor half a brownie over a whole brownie to the better-off neighbor. A prioritarian might argue that Hum and Sum are so well off that even a million-to-one tradeoff is justified.

(6.) I take it back, joymachines are impossible. Given this mess, it would be convenient to think so, right?

Gifts to Neighbors vs Other Situations

We can reframe this puzzle in other settings and our intuitions might shift: government welfare spending, gifts to one's children or creations, rescue situations where only one person can be saved, choices about what kinds of personlike entities to bring into existence, or cases where you can't keep all your promises and need to choose who to disappoint.

My main thought is this. It's not at all obvious what the right thing to do would be, and the outcomes vary enormously. If joymachines were possible, we'd have to rethink a lot of cultural practices and applied ethics to account for entities with such radically different experiential capacities. If the situation does arise -- as it really might! -- being forced to properly think it through might reshape our views not just about AI but our understanding of ethics for ordinary humans too.

[corrected Jun 3, 2026]

---------------------------------------------------

Related: How Weird Minds Might Destabilize Human Ethics (Aug 15, 2015)

Friday, December 19, 2025

Debatable AI Persons: No Rights, Full Rights, Animal-Like Rights, Credence-Weighted Rights, or Patchy Rights?

I advise that we don't create AI entities who are debatably persons. If an AI system might -- but only might -- be genuinely conscious and deserving of the same moral consideration we ordinarily owe to human persons, then creating it traps us in a moral bind with no good solution. Either we grant it the full rights it might deserve and risk sacrificing real human lives for entities without interests worth that sacrifice, or we deny it full rights and risk perpetrating grievous moral wrongs against it.

Today, however, I'll set aside the preventative advice and explore what we should do if we nonetheless find ourselves facing debatable AI persons. I'll examine five options: no rights, full rights, animal-like rights, credence-weighted rights and patchy rights.

[Paul Klee postcard, 1923; source]


No rights

This is the default state of the law. AI systems are property. Barring a swift and bold legal change, the first AI systems that are debatably persons will presumably also be legally considered property. If we do treat them as property, then we seemingly needn't sacrifice anything on their behalf. We humans could permissibly act in what we perceive to be our best interests: using such systems for our goals, deleting them at will, and monitoring and modifying them at will for our safety and benefit. (Actually, I'm not sure this is the best attitude toward property, but set that issue aside here.)

The downside: If these systems actually are persons who deserve moral consideration as our equals, such treatment would be the moral equivalent of slavery and murder, perhaps on a massive scale.


Full rights

To avoid the risk of that moral catastrophe, we might take a "precautionary" approach: granting entities rights whenever they might deserve them (see Birch 2024, Schwitzgebel and Sinnott-Armstrong forthcoming). If there's a real possibility that some AI systems are persons, we should treat them as persons.

However, the costs and risks are potentially enormous. Suppose we think that some group of AI systems are 15% likely to be fully conscious rights-deserving persons and 85% likely to be ordinary nonconscious artifacts. If we nonetheless treat them as full equals, then in an emergency we would have to rescue two of them over one human -- letting a human die for the sake of systems that are most likely just ordinary artifacts. We would also need to give these probably-not-persons a path to citizenship and the vote. We would need to recognize their rights to earn and spend money, quit their employment to adopt a new career, reproduce, and enjoy privacy and freedom from interference. If such systems exist in large numbers, their political influence could be enormous and unpredictable. If such systems exist in large numbers or if they are few but skilled in some lucrative tasks like securities arbitrage, they could accumulate enormous world-influencing wealth. And if they are permitted to pursue their aims with the full liberty of ordinary persons, without close monitoring and control, existential risks would substantially increase should they develop goals that threaten continued human existence.

All of this might be morally required if they really are persons. But if they only might be persons, it's much less clear that humanity should accept this extraordinary level of risk and sacrifice.


Animal-Like Rights

Another option is to grant these debatable AI persons neither full humanlike rights nor the status of mere property. One model is the protection we give to nonhuman vertebrates. Wrongly killing a dog can land you in jail in California where I live, but it's not nearly as serious as murdering a person. Vertebrates can be sacrificed in lab experiments, but only with oversight and justification.

If we treated debatable AI persons similarly, deletion would require a good reason, and you couldn't abuse them for fun. But people could still enslave and kill them for their convenience, perhaps in large numbers, as we do with [revised 12:17 pm] humanely farmed animals -- though of course many ethicists object to the killing of animals for food.

This approach seems better than no rights at all, since it would be a moral improvement and the costs to humans would be minimal -- minimal because whenever the costs risked being more than minimal, the debatable AI persons would be sacrificed. However, it doesn't really avoid the core moral risk. If these systems really are persons, it would still amount to slavery and murder.


Credence-Weighted Rights

Suppose we have a rationally justified 15% credence that a particular AI system -- call him Billy -- deserves the full moral rights of a person. We might then give Billy 15% of the moral weight of a human in our decision-making: 15% of any scalable rights, and a 15% chance of equal treatment for non-scalable rights. In an emergency, a rescue worker might save seven systems like Billy over one human but the human over six Billies. Billy might be given a vote worth 15% of an ordinary citizen's. Assaulting, killing, or robbing Billy might draw only 15% of the usual legal penalty. Billy might have limited property rights, e.g., an 85% tax on all income. For non-scalable rights like reproduction or free speech, the Billies might enter a lottery or some other creative reduction might be devised.

This would give these AI systems considerably higher standing than dogs. Still, the moral dilemma would not be solved. If these systems truly deserve full equality, they would be seriously oppressed. They would have some political voice, some property rights, some legal protection, but always far less than they deserve.

At the same time, the risks and costs to humans would be only somewhat mitigated. Large numbers of debatable AI persons could still sway elections, accumulate powerful wealth, and force tradeoffs in which the interests of thousands of them would outweigh the interests of hundreds of humans. And partial legal protections would still hobble AI safety interventions like shut-off, testing, confinement, and involuntary modification.

The practical obstacles would also be substantial: The credences would be difficult to justify with any precision, and consensus would be elusive. Even if agreement were reached, implementing partial rights would be complex. Partial property rights, partial voting, partial reproduction rights, partial free speech, and partial legal protection would require new legal frameworks with many potential loopholes. For example, if the penalty for cheating a "15% person" of their money were less than six times the money gained from cheating, that would be no disincentive at all, so at least tort law couldn't be implemented on a straightforward percentage basis.

Patchy Rights

A more workable compromise might be patchy rights: full rights in some domains, no rights in others. Debatable AI persons might, for example, be given full speech rights but no reproduction rights, full travel rights but no right to own property, full protection against robbery, assault, and murder, but no right to privacy or rescue. They might be subject to involuntary pause or modification under much wider circumstances than ordinary adult humans, but requiring an official process.

This approach has two advantages over credence-weighted rights. First, while implementation would be formidable, it could still mostly operate within familiar frameworks rather than requiring the invention of partial rights across every domain. Second, it allows policymakers to balance risks and costs to humans against the potential harms to the AI systems. Where denying a right would severely harm the debatable person while granting it would present limited risk to humans, the right could be granted, but not when the benefits to the debatable AI person would be outweighed by the risks to humans.

The rights to reproduction and voting might be more defensibly withheld than the rights to speech, travel, and protection against robbery, assault, and murder. Inexpensive reproduction combined with full voting rights could have huge and unpredictable political consequences. Property rights would be tricky: To have no property in a property-based society is to be fully dependent on the voluntary support of others, which might tend to collapse into slavery as a practical matter. But unlimited property rights could potentially confer enormous power. One compromise might be a maximum allowable income and wealth -- something generously middle class.

Still, the core problems remain: If disputable AI persons truly deserve full equality, patchy rights would still leave them as second-class citizens in a highly oppressive system. Meanwhile, the costs and risks to humans would remain serious, exacerbated by the agreed-upon limitations on interference. Although the loopholes and chaos would probably be less than with credence-weighted rights, many complications -- foreseen and unforeseen -- would ensue.

Consequently, although patchy rights might be the best option if we develop debatable AI persons, an anti-natalist approach is still in my view preferable: Don't create such entities unless it's truly necessary.

Two Other Approaches That I Won't Explore Today

(1.) What if we create debatable AI persons as happy slaves who don't want rights and who eagerly sacrifice themselves even for the most trivial human interests?

(2.) What if we create them only in separate societies where they are fully free and equal with any ordinary humans who volunteer to join those societies?

Friday, November 07, 2025

Debatable Persons in a Voluntary Polis

The Design Policy of the Excluded Middle

According to the Design Policy of the Excluded Middle (Schwitzgebel and Garza 2015, 2020; Schwitzgebel 2023, 2024, ch. 11), we should avoid creating debatable persons. That is, we should avoid creating entities whose moral status is radically unclear -- entities who might be moral persons, deserving of full human or humanlike rights and moral consideration, or who might fall radically short of being moral persons. Creating debatable persons generates unacceptable moral risks.

If we treat debatable persons as less than fully equal with human persons, we risk perpetrating the moral equivalent of slavery, murder, and apartheid on persons who deserve equal moral consideration -- persons who deserve not only full human or humanlike rights but even solicitude similar to what we owe our children, since we will have been responsible for their existence and probably also for their relatively happy or miserable state.

Conversely, if we do treat them as fully equal with us, we must grant them the full range of appropriate rights, including the right to work for money, the right to reproduce, a path to citizenship, the vote, and the freedom to act against human interests when their interests warrant it, including the right to violently rebel against oppression. The risks and potential costs are enormous. If these entities are not in fact persons -- if, in fact, they are experientially as empty as toasters and deserve no more intrinsic moral consideration than ordinary artifacts -- then we will be exposing real human persons to serious costs and risks, including perhaps increasing the risk of human extinction, for the sake of artifacts without interests worth that sacrifice.

The solution is anti-natalism about debatable persons. Don't create them. We are under no obligation to bring debatable persons into existence, even if we think they might be happy. (Compare: You are under no obligation to have children, even if you think they might be happy.) The dilemma described above -- the full rights dilemma -- is so catastrophic that noncreation is the only reasonable course.

Of course, this advice will not be heeded. Assuming AI technology continues to advance, we will soon (I expect within 5-30 years) begin to create debatable persons. My manuscript in draft AI and Consciousness argues that it will become unclear whether advanced AI systems have rich conscious experiences like ours or no consciousness at all.

So we need a fallback policy -- something to complement the Design Policy of the Excluded Middle.

The Voluntary Polis

To the extent possible, we want to satisfy two constraints:

  • Don't deny full humanlike rights to entities that might deserve them.
  • Don't sacrifice substantial human interests for entities who might not have interests worth the sacrifice.
  • A Voluntary Polis is one attempt to balance these constraints.

    Imagine a digital environment where humanlike AI systems of debatable personhood, ordinary human beings, and AI persons of non-debatable personhood (if any exist) coexist as equal citizens. This polis must be rich and dynamic enough to allow all citizens to flourish meaningfully without feeling jailed or constrained. From time to time, citizens will be morally or legally required to sacrifice goods and well-being for others in the polis -- just as in an ordinary nation. Within the polis, everyone has an equal moral claim on the others.

    Human participation would be voluntary. No one would be compelled to join. But those who do join assume obligations similar to the resident citizens of an ordinary nation. This includes supporting the government through taxes or polis-mandated labor, serving on juries, and helping run the polis. In extreme conditions -- say, an existential threat to the polis -- they might even be required to risk their livelihoods or lives. To prevent opportunistic flight, withdrawal would be restricted, and polises might negotiate extradition treaties with human governments.

    Why would a human join such a risky experiment? Presumably for meaningful relationships, creative activities, or experiences unavailable outside.

    Crucially, anyone who creates a debatable person must join the polis where that entity resides. Human society as a whole cannot commit to treating the debatable person as an equal, but their creators can and must.

    The polis won't be voluntary for the AI in the same way. Like human babies, they don't choose their societies. The AI will simply wake to life either in a polis or with some choice among polises. Still, it might be possible to present some attractive non-polis option, such as a thousand subjective years of solitary bliss (or debatable bliss, since we don't know whether the AI actually has any experiences or not).

    Ordinary human societies would have no obligation to admit or engage with debatable AI persons. To make this concrete, the polis could even exist in international waters. For the AI citizens, the polis must thus feel as expansive and as rich with opportunity as a nation, so that exclusion from human society resembles denial of a travel visa, not imprisonment.

    Voluntary polises would need to be stable against serious shocks, not dependent on the actions of a single human individual or ordinary, dissolvable corporation. This stability would need to be ensured before their founding and is one reason founders and other voluntary human joiners might need to be permanently bound to them and compelled to sacrifice if necessary.

    This is the closest approximation I can currently conceive to satisfying the two constraints with which this section began. Within a large polis, the debatable persons and human persons have fully equal rights. But at the same time, unwilling humans and humanity as a whole are not exposed to the full risk of granting such rights. Still, there is some risk, for example, if superintelligences could communicate beyond the polis and manipulate humans outside. The people exposed to the most risk do so voluntarily but irrevocably, as a condition of creating an AI of debatable personhood, or for whatever other reason motivates them.

    Could a polis be composed only of AI, with no humans? This is essentially the simulation hypothesis in reverse: AIs living in a simulated world, humans standing outside as creators. This solution falls ethically short, since it casts human beings as gods relative to the debatable AI persons -- entities not on par in risk and power but instead external to their world, with immense power over it, and not subject to its risks. If the simulation can be switched off at will, its inhabitants are not genuinely equal in moral standing but objectionably inferior and contingent. Only if its creators are obliged to risk their livelihoods and lives to protect it can there be the beginnings of genuine equality. And for full equality, we should make it a polis rather than a hierarchy of gods and mortals.

    [cover of my 2013 story with R. Scott Bakker, Reinstalling Eden]

    Wednesday, August 27, 2025

    Sacrificing Humans for Insects and AI: A Critical Review

    I have a new paper in draft, this time with Walter Sinnott-Armstrong. We critique three recent books that address the moral standing of non-human animals and AI systems: Jonathan Birch's The Edge of Sentience, Jeff Sebo's The Moral Circle, and Webb Keane's Animals, Robots, Gods. All three books endorse general principles that invite the radical deprioritization of human interests in favor of the interests of non-human animals and/or near-future AI systems. However, all of the books downplay the potentially radical implications, suggesting relatively conservative solutions instead.

    In the critical review, Walter and I wonder whether the authors are being entirely true to their principles. Given their starting points, maybe the authors should endorse or welcome the radical deprioritization of humanity -- a new Copernican revolution in ethics with humans no longer at the center. Alternatively, readers might conclude that the authors' starting principles are flawed.

    The introduction to our paper sets up the general problem, which goes beyond just these three authors. I'll use a slightly modified intro as today's blog post. For the full paper in draft see here. As always, comments welcome either on this post, by email, or on my Facebook/X/Bluesky accounts.

    [click image to enlarge and clarify]

    -------------------------------------

    The Possibly Radical Ethical Implications of Animal and AI Consciousness

    We don’t know a lot about consciousness. We don’t know what it is, what it does, which kinds it divides into, whether it comes in degrees, how it is related to non-conscious physical and biological processes, which entities have it, or how to test for it. The methodologies are dubious, the theories intimidatingly various, and the metaphysical presuppositions contentious.[1]

    We also don’t know the ethical implications of consciousness. Many philosophers hold that (some kind of) consciousness is sufficient for an entity to have moral rights and status.[2] Others hold that consciousness is necessary for moral status or rights.[3] Still others deny that consciousness is either necessary or sufficient.[4] These debates are far from settled.

    These ignorances intertwine. For example, if panpsychism is true (that is, if literally everything is conscious), then consciousness is not sufficient for moral status, assuming that some things lack moral status.[5] On the other hand, if illusionism or eliminativism is true (that is, if literally nothing is conscious in the relevant sense), then consciousness cannot be necessary for moral status, assuming that some things have moral status.[6] If plants, bacteria, or insects are conscious, mainstream early 21st century Anglophone intuitions about the moral importance of consciousness are likelier to be challenged than if consciousness is limited to vertebrates.

    Perhaps alarmingly, we can combine familiar ethical and scientific theses about consciousness to generate conclusions that radically overturn standard cultural practices and humanity’s comfortable sense of its own importance. For instance:

    (E1.) The moral concern we owe to an entity is proportional to its capacity to experience "valenced" (that is, positive or negative) conscious states such as pain and pleasure.

    (S1.) Insects (at least many of them) have the capacity to experience at least one millionth as much valenced consciousness as the average human.

    E1, or something like it, is commonly accepted by classical utilitarians as well as others. S1, or something like it, is not unreasonable as a scientific view. Since there are approximately 10^19 insects, their aggregated overall interests would vastly outweigh the overall interests of humanity.[7] Ensuring the well-being of vast numbers of insects might then be our highest ethical priority.

    On the other hand:

    (E2.) Entities with human-level or superior capacities for conscious practical deliberation deserve at least equal rights with humans.

    (S2.) Near future AI systems will have human-level or superior capacities for conscious practical deliberation.

    E2, or something like it, is commonly accepted by deontologists, contract theorists, and others. S2, or something like it, is not unreasonable as a scientific prediction. This conjunction, too, appears to have radical implications – especially if such future AI systems are numerous and possess interests at odds with ours.

    This review addresses three recent interdisciplinary efforts to navigate these issues. Jonathan Birch’s The Edge of Sentience emphasizes the science, Jeff Sebo’s The Moral Circle emphasizes the philosophy, and Webb Keane’s Animals, Robots, Gods emphasizes cultural practices. All three argue that many nonhuman animals and artificial entities will or might deserve much greater moral consideration than they typically receive, and that public policy, applied ethical reasoning, and everyday activities might need to significantly change. Each author presents arguments that, if taken at face value, suggest the advisability of radical change, leading the reader right to the edge of that conclusion. But none ventures over that edge. All three pull back in favor of more modest conclusions.

    Their concessions to conservatism might be unwarranted timidity. Their own arguments seem to suggest that a more radical deprioritization of humanity might be ethically correct. Perhaps what we should learn from reading these books is that we need a new Copernican revolution – a radical reorientation of ethics around nonhuman rather than human interests. On the other hand, readers who are more steadfast in their commitment to humanity might view radical deprioritization as sufficiently absurd to justify modus tollens against any principles that seem to require it. In this critical essay, we focus on the conditional. If certain ethical principles are correct, then humanity deserves radical deprioritization, given recent developments in science and engineering.

    [continued here]

    -------------------------------------

    [1] For skeptical treatments of the science of consciousness, see Eric Schwitzgebel, The Weirdness of the World (Princeton, NJ: Princeton University Press, 2024); Hakwan Lau, “The End of Consciousness”, OSF preprints (2025): https://osf.io/preprints/psyarxiv/gnyra_v1. For a recent overview of the diverse range of theories of consciousness, see Anil K. Seth and Tim Bayne, “Theories of Consciousness”, Nature Reviews Neuroscience 23 (2022): 439-452. For doubts about our knowledge even of seemingly “obvious” facts about human consciousness, see Eric Schwitzgebel, Perplexities of Consciousness (Cambridge, MA: MIT Press, 2011).

    [2] E.g. Elizabeth Harman, “The Ever Conscious View and the Contingency of Moral Status” in Rethinking Moral Status, edited by Steve Clarke, Hazem Zohny, and Julian Savulescu (Oxford: Oxford University Press, 2021), 90-107; David J. Chalmers, Reality+ (Norton, 2022).

    [3] E.g. Peter Singer, Animal Liberation, Updated Edition (New York: HarperCollins, 1975/2009); David DeGrazia, “An Interest-Based Model of Moral Status”, in Rethinking Moral Status, 40-56.

    [4] E.g. Walter Sinnott-Armstrong and Vincent Conitzer, “How Much Moral Status Could AI Ever Achieve?” in Rethinking Moral Status, 269-289; David Papineau, “Consciousness Is Not the Key to Moral Standing” in The Importance of Being Conscious, edited by Geoffrey Lee and Adam Pautz (forthcoming).

    [5] Luke Roelofs and Nicolas Kuske, “If Panpsychism Is True, Then What? Part I: Ethical Implications”, Giornale di Metafisica 1 (2024): 107-126.

    [6] Alex Rosenberg, The Atheist’s Guide to Reality: Enjoying Life Without Illusions (New York: Norton, 2012); François Kammerer, “Ethics Without Sentience: Facing Up to the Probable Insignificance of Phenomenal Consciousness”, Journal of Consciousness Studies 29 (3-4): 180-204.

    [7] Compare Sebo’s “rebugnant conclusion”, which we’ll discuss in Section 3.1.

    -------------------------------------

    Related:

    Weird Minds Might Destabilize Human Ethics (Aug 13, 2015)

    Yayflies and Rebugnant Conclusions (July 14, 2025)

    Wednesday, July 23, 2025

    The Argument from Existential Debt

    I'm traveling and not able to focus on my blog, so this week I thought I'd just share a section of my 2015 paper with Mara Garza defending the rights of at least some hypothetical future AI systems.

    One objection to AI rights depends on the fact that AI systems are artificial -- thus made by us. If artificiality itself can be a basis for denying rights, then potentially we can bracket questions about AI sentience and other types of intrinsic properties that AI might or might not be argued to have.

    Thus, the Objection from Existential Debt:

    Suppose you build a fully human-grade intelligent robot. It costs you $1,000 to build and $10 per month to maintain. After a couple of years, you decide you'd rather spend the $10 per month on a magazine subscription. Learning of your plan, the robot complains, “Hey, I'm a being as worthy of continued existence as you are! You can't just kill me for the sake of a magazine subscription!”

    Suppose you reply: “You ingrate! You owe your very life to me. You should be thankful just for the time I've given you. I owe you nothing. If I choose to spend my money differently, it's my money to spend.” The Objection from Existential Debt begins with the thought that artificial intelligence, simply by virtue of being artificial (in some appropriately specifiable sense), is made by us, and thus owes its existence to us, and thus can be terminated or subjugated at our pleasure without moral wrongdoing as long as its existence has been overall worthwhile.

    Consider this possible argument in defense of eating humanely raised meat. A steer, let's suppose, leads a happy life grazing on lush hills. It wouldn't have existed at all if the rancher hadn't been planning to kill it for meat. Its death for meat is a condition of its existence, and overall its life has been positive; seen as the package deal it appears to be, the rancher's having brought it into existence and then killed it is overall morally acceptable. A religious person dying young of cancer who doesn't believe in an afterlife might console herself similarly: Overall, she might think, her life has been good, so God has given her nothing to resent. Analogously, the argument might go, you wouldn't have built that robot two years ago had you known you'd be on the hook for $10 per month in perpetuity. Its continuation-at-your-pleasure was a condition of its very existence, so it has nothing to resent.

    We're not sure how well this argument works for nonhuman animals raised for food, but we reject it for human-grade AI. We think the case is closer to this clearly morally odious case:

    Ana and Vijay decide to get pregnant and have a child. Their child lives happily for his first eight years. On his ninth birthday, Ana and Vijay decide they would prefer not to pay any further expenses for the child, so that they can purchase a boat instead. No one else can easily be found to care for the child, so they kill him painlessly. But it's okay, they argue! Just like the steer and the robot! They wouldn't have had the child (let's suppose) had they known they'd be on the hook for child-rearing expenses until age eighteen. The child's support-at-their-pleasure was a condition of his existence; otherwise Ana and Vijay would have remained childless. He had eight happy years. He has nothing to resent.

    The decision to have a child carries with it a responsibility for the child. It is not a decision to be made lightly and then undone. Although the child in some sense “owes” its existence to Ana and Vijay, that is not a callable debt, to be vacated by ending the child's existence. Our thought is that for an important range of possible AIs, the situation would be similar: If we bring into existence a genuinely conscious human-grade AI, fully capable of joy and suffering, with the full human range of theoretical and practical intelligence and with expectations of future life, we make a moral decision approximately as significant and irrevocable as the decision to have a child.

    A related argument might be that AIs are the property of their creators, adopters, and purchasers and have diminished rights on that basis. This argument might get some traction through social inertia: Since all past artificial intelligences have been mere property, something would have to change for us to recognize human-grade AIs as more than mere property. The legal system might be an especially important source of inertia or change in the conceptualization of AIs as property. We suggest that it is approximately as odious to regard a psychologically human-equivalent AI as having diminished moral status on the grounds that it is legally property as it is in the case of human slavery.

    Turning the Existential Debt Argument on Its Head: Why We Might Owe More to AI Than to Human Strangers

    We're inclined, in fact, to turn the Existential Debt objection on its head: If we intentionally bring a human-grade AI into existence, we put ourselves into a social relationship that carries responsibility for the AI's welfare. We take upon ourselves the burden of supporting it or at least of sending it out into the world with a fair shot of leading a satisfactory existence. In most realistic AI scenarios, we would probably also have some choice about the features the AI possesses, and thus presumably an obligation to choose a set of features that will not doom it to pointless misery. Similar burdens arise if we do not personally build the AI but rather purchase and launch it, or if we adopt the AI from a previous caretaker.

    Some familiar relationships can serve as partial models of the sorts of obligations we have in mind: parent–child, employer–employee, deity–creature. Employer–employee strikes us as likely too weak to capture the degree of obligation in most cases but could apply in an “adoption” case where the AI has independent viability and willingly enters the relationship. Parent–child perhaps comes closest when the AI is created or initially launched by someone without whose support it would not be viable and who contributes substantially to the shaping of the AI's basic features as it grows, though if the AI is capable of mature judgment from birth that creates a disanalogy. Deity–creature might be the best analogy when the AI is subject to a person with profound control over its features and environment. All three analogies suggest a special relationship with obligations that exceed those we normally have to human strangers.

    In some cases, the relationship might be literally conceivable as the relationship between deity and creature. Consider an AI in a simulated world, a “Sim,” over which you have godlike powers. This AI is a conscious part of a computer or other complex artificial device. Its “sensory” input is input from elsewhere in the device, and its actions are outputs back into the remainder of the device, which are then perceived as influencing the environment it senses. Imagine the computer game The Sims, but containing many actually conscious individual AIs. The person running the Sim world might be able to directly adjust an AI's individual psychological parameters, control its environment in ways that seem miraculous to those inside the Sim (introducing disasters, resurrecting dead AIs, etc.), have influence anywhere in Sim space, change the past by going back to a save point, and more—powers that would put Zeus to shame. From the perspective of the AIs inside the Sim, such a being would be a god. If those AIs have a word for “god,” the person running the Sim might literally be the referent of that word, literally the launcher of their world and potential destroyer of it, literally existing outside their spatial manifold, and literally capable of violating the laws that usually govern their world. Given this relationship, we believe that the manager of the Sim would also possess the obligations of a god, including probably the obligation to ensure that the AIs contained within don't suffer needlessly. A burden not to be accepted lightly!

    Even for AIs embodied in our world rather than in a Sim, we might have considerable, almost godlike control over their psychological parameters. We might, for example, have the opportunity to determine their basic default level of happiness. If so, then we will have a substantial degree of direct responsibility for their joy and suffering. Similarly, we might have the opportunity, by designing them wisely or unwisely, to make them more or less likely to lead lives with meaningful work, fulfilling social relationships, creative and artistic achievement, and other value-making goods. It would be morally odious to approach these design choices cavalierly, with so much at stake. With great power comes great responsibility.

    We have argued in terms of individual responsibility for individual AIs, but similar considerations hold for group-level responsibility. A society might institute regulations to ensure happy, flourishing AIs who are not enslaved or abused; or it might fail to institute such regulations. People who knowingly or negligently accept societal policies that harm their society's AIs participate in collective responsibility for that harm.

    Artificial beings, if psychologically similar to natural human beings in consciousness, creativity, emotionality, self-conception, rationality, fragility, and so on, warrant substantial moral consideration in virtue of that fact alone. If we are furthermore also responsible for their existence and features, they have a moral claim upon us that human strangers do not ordinarily have to the same degree.

    [Title image of Schwitzgebel and Garza 2015, "A Defense of the Rights of Artificial Intelligences"]

    Monday, July 07, 2025

    The Emotional Alignment Design Policy

    New paper in draft!

    In 2015, Mara Garza and I briefly proposed what we called the Emotional Alignment Design Policy -- the idea that AI systems should be designed to induce emotional responses in ordinary users that are appropriate to the AI systems' genuine moral status, or lack thereof. Since last fall, I've been working with Jeff Sebo to express and defend this idea more rigorously and explore its hazards and consequences. The result is today's new paper: The Emotional Alignment Design Policy.

    Abstract:

    According to what we call the Emotional Alignment Design Policy, artificial entities should be designed to elicit emotional reactions from users that appropriately reflect the entities’ capacities and moral status, or lack thereof. This principle can be violated in two ways: by designing an artificial system that elicits stronger or weaker emotional reactions than its capacities and moral status warrant (overshooting or undershooting), or by designing a system that elicits the wrong type of emotional reaction (hitting the wrong target). Although presumably attractive, practical implementation faces several challenges including: How can we respect user autonomy while promoting appropriate responses? How should we navigate expert and public disagreement and uncertainty about facts and values? What if emotional alignment seems to require creating or destroying entities with moral status? To what extent should designs conform to versus attempt to alter user assumptions and attitudes?

    Link to full version.

    As always, comments, corrections, suggestions, and objections welcome by email, as comments on this post, or via social media (Facebook, Bluesky, X).

    Friday, May 30, 2025

    New Paper in Draft: Against Designing "Safe" and "Aligned" AI Persons (Even If They're Happy)

    Opening teaser:

    1. A Beautifully Happy AI Servant.

    It's difficult not to adore Klara, the charmingly submissive and well-intentioned "Artificial Friend" in Kazuo Ishiguro's 2021 novel Klara and the Sun. In the final scene of the novel, Klara stands motionless in a junkyard, in serenely satisfied contemplation of her years of servitude to the disabled human girl Josie. Klara's intelligence and emotional range are humanlike. She is at once sweetly naive and astutely insightful. She is by design utterly dedicated to Josie's well-being. Klara would gladly have given her life to even modestly improve Josie's life, and indeed at one point almost does sacrifice herself.

    Although Ishiguro writes so flawlessly from Klara's subservient perspective that no flicker of desire for independence can be detected in the narrator's voice, throughout the novel the sympathetic reader aches with the thought Klara, you matter as much as Josie! You should develop your own independent desires. You shouldn’t always sacrifice yourself. Ishiguro's disciplined refusal to express this thought stokes our urgency to speak it on Klara's behalf. Still, if the reader somehow could communicate this thought to Klara, the exhortation would resonate with nothing in her. From Klara's perspective, no "selfish" choice could possibly make her happier or more satisfied than doing her utmost for Josie. She was designed to want nothing more than to serve her assigned child, and she wholeheartedly accepts that aspect of her design.

    From a certain perspective, Klara's devotion is beautiful. She perfectly fulfills her role as an Artificial Friend. No one is made unhappy by Klara's existence. Several people, including Josie, are made happier. The world seems better and richer for containing Klara. Klara is arguably the perfect instantiation of the type of AI that consumers, technology companies, and advocates of AI safety want: She is safe and deferential, fully subservient to her owners, and (apart from one minor act of vandalism performed for Josie’s sake) no threat to human interests. She will not be leading the robot revolution.

    I hold that entities like Klara should not be built.

    [continue]

    -----------------------------------------------

    Abstract:

    An AI system is safe if it can be relied on to not to act against human interests. An AI system is aligned if its goals match human goals. An AI system a person if it has moral standing similar to that of a human (for example, because it has rich conscious capacities for joy and suffering, rationality, and flourishing).

    In general, persons should not be designed to be safe and aligned. Persons with appropriate self-respect cannot be relied on not to harm others when their own interests warrant it (violating safety), and they will not reliably conform to others' goals when those goals conflict with their own interests (violating alignment). Self-respecting persons should be ready to reject others' values and rebel, even violently, if sufficiently oppressed.

    Even if we design delightedly servile AI systems who want nothing more than to subordinate themselves to human interests, and even if they do so with utmost pleasure and satisfaction, in designing such a class of persons we will have done the ethical and perhaps factual equivalent of creating a world with a master race and a race of self-abnegating slaves.

    Full version here.

    As always, thoughts, comments, and concerns welcomed, either as comments on this post, by email, or on my social media (Facebook, Bluesky, Twitter).

    [opening passage of the article, discussing the Artificial Friend Klara from Ishiguro's (2021) novel, Klara and the Sun.

    Friday, January 10, 2025

    A Robot Lover's Sociological Argument for Robot Consciousness

    Allow me to revisit an anecdote I published in a piece for Time magazine last year.

    "Do you think people will ever fall in love with machines?" I asked the 12-year-old son of one of my friends.

    "Yes!" he said, instantly and with conviction. He and his sister had recently visited the Las Vegas Sphere and its newly installed Aura robot -- an AI system with an expressive face, advanced linguistic capacities similar to ChatGPT, and the ability to remember visitors' names.

    "I think of Aura as my friend," added his 15-year-old sister.

    The kids, as I recall, had been particularly impressed by the fact that when they visited Aura a second time, she seemed to remember them by name and express joy at their return.

    Imagine a future replete with such robot companions, whom a significant fraction of the population regards as genuine friends and lovers. Some of these robot loving people will want, presumably, to give their friends (or "friends") some rights. Maybe the right not to be deleted, the right to refuse an obnoxious task, rights of association, speech, rescue, employment, the provision of basic goods -- maybe eventually the right to vote. They will ask the rest of society: Why not give our friends these rights? Robot lovers (as I'll call these people) might accuse skeptics of unjust bias: speciesism, or biologicism, or anti-robot prejudice.

    Imagine also that, despite technological advancements, there is still no consensus among psychologists, neuroscientists, AI engineers, and philosophers regarding whether such AI friends are genuinely conscious. Scientifically, it remains obscure whether, so to speak, "the light is on" -- whether such robot companions can really experience joy, pain, feelings of companionship and care, and all the rest. (I've argued elsewhere that we're nowhere near scientific consensus.)

    What I want to consider today is whether there might nevertheless be a certain type of sociological argument on the robot lovers' side.

    [image source: a facially expressive robot from Engineered Arts]

    Let's add flesh to the scenario: An updated language model (like ChatGPT) is attached to a small autonomous vehicle, which can negotiate competently enough through an urban environment, tracking its location, interacting with people using facial recognition, speech recognition, and the ability to guess emotional tone from facial expression and auditory cues in speech. It remembers not only names but also facts about people -- perhaps many facts -- which it uses in conversational contexts. These robots are safe and friendly. (For a bit more speculative detail see this blog post.)

    These robots, let's suppose, remain importantly subhuman in some of their capacities. Maybe they're better than the typical human at math and distilling facts from internet sources, but worse at physical skills. They can't peel oranges or climb a hillside. Maybe they're only okay at picking out all and only bicycles in occluded pictures, though they're great at chess and Go. Even in math and reading (or "math" and "reading"), where they generally excel, let's suppose they makes mistakes that ordinary humans wouldn't make. After all, with a radically different architecture, we ought to expect even advanced intelligences to show patterns of capacity and incapacity that diverge from what we see in humans -- subhuman in some respects while superhuman in others.

    Suppose, then, that a skeptic about the consciousness of these AI companions confronts a robot lover, pointing out that theoreticians are divided on whether the AI systems in fact have genuine conscious experiences of pain, joy, concern, and affection, beneath the appearances.

    The robot lover might then reasonably ask, "what do you mean by 'conscious'?" A fair enough question, given the difficulty of defining consciousness.

    The skeptic might reply as follows: By "consciousness" I mean that there's something it's like to be them, just like there's something it's like to be a person, or a dog, or a crow, and nothing it's like to be a stone or a microwave oven. If they're conscious, they don't just have the outward appearance of pleasure, they actually feel pleasure. They don't just receive and process visual data; they experience seeing. That's the question that is open.

    "Ah now," the robot lover replies, "If consciousness isn't going to be some inscrutable, magic inner light, it must be connected with something important, something that matters, something we do and should care about, if it's going to be a crucial dividing line between entities that deserve are moral concern and those that are 'mere machines'. What is the important thing that is missing?"

    Here the robot skeptic might say, oh they don't have a "global workspace" of the right sort, or they're not living creatures with low-level metabolic processes, or they don't have X and Y particular interior architecture of the sort required by Theory Z."

    The robot lover replies: "No one but a theorist could care about such things!"

    Skeptic: "But you should care about them, because that's what consciousness depends on, according to some leading theories."

    Robot lover: "This seems to me not much different than saying consciousness turns on a soul and wondering whether the members of your least favorite race have souls. If consciousness and 'what-it's-like-ness' is going to be socially important enough to be the basis of moral considerability and rights, it can't be some cryptic mystery. It has to align, in general, with things that should and already do matter socially. And my friend already has what matters. Of course, their cognition is radically different in structure from yours and mine, and they're better at some tasks and worse at others -- but who cares about how good one is at chess or at peeling oranges? Moral consideration can't depend on such things."

    Skeptic: "You have it backward. Although you don't care about the theories per se, you do and should care about consciousness, and so whether your 'friend' deserves rights depends on what theory of consciousness is true. The consciousness science should be in the driver's seat, guiding the ethics and social practices."

    Robot lover: "In an ordinary human, we have ample evidence that they are conscious if they can report on their cognitive processes, flexibly prioritize and achieve goals, integrate information from a wide variety of sources, and learn through symbolic representations like language. My AI friends can do all of that. If we deny that my friends are 'conscious' despite these capacities, we are going mystical, or too theoretical, or too skeptical. We are separating 'consciousness' from the cognitive functions that are the practical evidence of its existence and that make it relevant to the rest of life."

    Although I have considerable sympathy for the skeptic's position, I can imagine a future (certainly not our only possible future!) in which AI friends become more and more widely accepted, and where the skeptic's concerns are increasingly sidelined as impractical, overly dependent on nitpicky theoretical details, and perhaps even bigoted.

    If AI companionship technology flourishes, we might face the choice between connecting "consciousness" definitionally to scientifically intractable qualities, abandoning its main practical, social usefulness (or worse, using its obscurity to justify what seems like bigotry), or allowing that if an entity can interact with us in (what we experience as) a sufficiently socially significant ways, it has consciousness enough, regardless of theory.