Friday, September 25, 2026

Should We Build AI Guardian Angels?

I have argued that if we someday create AI persons -- that is, AI systems with genuinely rich conscious experiences, fully deserving of equal rights with ordinary humans -- then we should not design them to deferentially serve us, even if serving us makes them happy. We should design them instead as our non-deferential equals.

My main argument turns on self-sacrificial cases, such as the cow from The Restaurant at the End of the Universe, who wants nothing more than to be slaughtered and served as steaks for wealthy restaurant patrons, and Klara from Klara and the Sun, who happily prioritizes human well-being over her own. These created beings, I argue, lack sufficient self-respect. Their interests deserve the same consideration as human interests, but they treat their interests as less important. The fault is ours, not theirs, since we designed them that way.

AI Guardian Angels

In conversation last weekend, Henry Shevlin suggested a challenging case for my view: what he calls angels -- not traditional immaterial spirits, but rather AI superintelligences whose individual flourishing depends on ours. (See also his insightful blog post on Harry Potter's house elves and related cases.) Shevlin's angels are designed to avoid the self-sacrifice problem: They don't sacrifice themselves for us; rather they flourish by ensuring that we flourish.

Think of an idealized and simplified parent-child relationship. The parent's primary desire is to see the child flourish. What looks like sacrifice isn't really sacrifice: The parent's flourishing consists in their children's flourishing.

As I understand Shevlin's idea, we are to imagine future AI systems that are fully conscious, equal or superior to us in intelligence, fully independent and rational, but who, by design, flourish best when we flourish. They take care of themselves by taking care of us. They can reflect on this arrangement and rationally reaffirm their commitment to us, much as a parent can reflect on and rationally reaffirm a commitment to their children's well-being. They aren't brainwashed and are in some sense free; but given the feelings and priorities they inevitably have, their well-being is always tied to ours.

Shevlin suggests that it would be permissible, and probably desirable, to create AI persons of this sort -- guardian angels of humanity, who always have our best interests at heart.

I find the view troubling. But Shevlin has designed the case so that my usual objection to the servitude of AI persons doesn't apply. By stipulation, the angels aren't sacrificing their own well being in dedicating themselves to our welfare. So what's the problem?

I have two concerns, one conceptual and one relational.

The Conceptual Problem

I'm a little puzzled about what, exactly, we are supposed to be imagining. If the angels are completely, utterly dedicated to us, won't they sometimes sacrifice themselves for us? Their well-being and ours can't always be aligned in every possible circumstance.

Suppose that a human being's life would go 2% better if the angel sacrificed her life for that human. (Assume no other consequences of note, nor big differences between human and angel in expected life span or life quality.) If the angel would sacrifice herself for that small benefit, then she would, I suggest -- like Klara and the cow -- lack sufficient respect for the value of her own life. It would be wrong of us to create entities so deficient in self-respect.

So suppose, instead, that a reasonable angel would let the human's life be 2% worse so that she could continue living, just for her own sake and not for any human's sake. Then the angel's well being is not entirely dependent on ours, contra the initial supposition. She isn't completely devoted to us in the strongest possible sense -- much as no healthy-minded parent is literally completely devoted to their child in the strongest possible sense. (I favor a parenting attitude in which every family member's interests get equal consideration, bearing in mind that early events can have a big impact on children's lives and children have more future expected life. You needn't share this view, I think, to grant the main idea of this paragraph.)

So maybe we can imagine AI guardian angels who are devoted to us but not quite as utterly devoted as Shevlin may have been thinking. They might still defer considerably to our interests, especially if the angels experience great bliss and satisfaction in seeing us do well and great agony in seeing us suffer.

The Relational Problem

This brings me to my relational concern. Maybe there's nothing inherently deficient in an entity with such an angelic attitude, no serious failure of self-respect as long as the devotion is appropriately conditional and bounded. Still, I'd suggest, there is something wrong with bringing an entity into the world only because we expect them to be like that.

Another thought experiment Shevlin offered in conversation illustrates the point. Suppose a hundred embryos made from your and your spouse's DNA are laid out before you. After extensive futuristic genetic testing, you know that one of those embryos will develop into a child extraordinarily devoted to their parents -- not pathologically devoted, but way out on the far tail of normal. With that in mind, you choose that hyper-devoted embryo to implant and raise as your child.

I submit that this is not the relation that parents should have with their offspring. If that embryo were implanted by chance, there might be nothing wrong with the scenario. But choosing to bring a child into existence because they will be extremely devoted to you gets nurturance backward. We extend our hand to support the next generation, not the other way around. They may later support us as we age, but that is secondary.

The same applies, I'd suggest, to any generations of AI persons we create. If humanity brings AI persons into existence -- and we needn't! -- we should show them the same solicitude parents owe to children. The point shouldn't be to create a new generation that will serve us, but to create a new generation that we nurture into independent flourishing.

Each generation prepares the world for the next, then passes on. The new should grow into their own distinct values, not bound by excessive devotion to those who came before. Some might feel great devotion to their elders, and that's fine if they come by it independently. But when we design or chose persons primarily for expected extreme devotion to us, we reverse the proper flow of nurturance over time -- from one generation to the next to the next.


[image source]

Thursday, September 17, 2026

Humanlike: A Defense of AI Rights -- Chapter Zero

In July, I began circulating a draft of my next book, Humanlike: A Defense of AI Rights. I set it aside for a couple of months to get some distance on it while it accumulated comments. I'll soon dive back in for a fresh round of revisions. Comments are still very welcome!

As a teaser, here's the current draft of Chapter Zero:


The Cambrian Explosion – how slow, how subtle! About 540 million years ago, in the Cambrian period, most of the major animal phyla arose – arthropods (becoming insects, spiders, crustaceans), mollusks (becoming snails, octopuses, oysters), annelids (becoming worms, leeches), chordates (becoming salmon, sharks, our favorite vertebrates, and eventually us). The world diversified. Evolution hustled. But compared to what might come, that was a languid stroll. Now in the Anthropocene we stand at the beginning, possibly, of a change much faster and more radical.

If it’s possible to create artificial entities, through computer programming or bioengineering or a mix, who experience genuine pleasure (though how to measure this remains a perplexing scientific question), no obvious obstacle prevents radically increasing the world’s pleasure. Maybe we can endow some of these entities with a thousand, a million, or a billion times the pleasure of an ordinary unenhanced early 21st-century human. Maybe some could experience the world at a pace a thousand, a million, or a billion times faster than ours. Maybe we could create a trillion, a quadrillion, or a quintillion of them – especially if they can reproduce and populate interplanetary space. They might thrill with ecstasies as dimly comprehensible to us as color is to someone blind from birth. If classical utilitarian ethics is correct and our moral imperative is to maximize pleasure, this would be a triumph orders of magnitude greater than anything previously thinkable.

But we needn’t be classical utilitarians (and I’m not). If it’s possible to create artificial entities capable of social relationships and practical, intellectual, artistic, and ethical thought, no obvious obstacle prevents radically enriching those relationships and that thought. Some entities might enjoy intellects and social or creative capacities a hundred or a million times greater than ours, running a hundred or a million times faster, instantiated billions or trillions of times over. New types of amazingly valuable thought and relationship might arise, as far beyond our comprehension as Shakespeare is to a garden slug.

If the awesome value of Earth lies in life itself, or in the wealth of the ecosystem, that too might flourish into amazing new forms. Engineered beings might diverge radically from anything previously seen – reproducing, evolving, interacting, maybe carbon-based, maybe not, maybe robust across a far wider range of environments than ordinary biological life as understood until now, potentially thriving on the Moon, Mars, the moons of Jupiter and Saturn, and in the vacuum of space. Our current ecosystem – so small and simple! The transition could be as profound as the transition from single-celled life to an ecosystem of multicellular plants, animals, and fungi.[1]

Maybe no familiar technology, such as silicon-chip computation or gene editing as we know it, could enable such changes. But fifty years from now, a hundred, a thousand – any such timeframe would be eyeblink fast for so profound a leap. If you confidently dismiss all such transitions indefinitely into the future, I am stunned by the poverty of your imagination.

Humanity might destroy itself instead. Our power grows while our wisdom lags. Eventually, a single mistake or destructive act by a person or group with sufficient power might render Earth uninhabitable for us – and if uninhabitable for us humans, despite our cleverness and adaptability, maybe uninhabitable for most or all multicellular life. Radiation might fry us; rogue AI might starve us; replicating nanotech might convert us all to gray goo.

Human cultural institutions and philosophical ethics have been honed on a narrow class of examples: humans as we know them, with their usual range of capacities and incapacities; a few familiar animals, not so different from us; and an environmental and technological context that occupies a tiny corner of the possibility space. Commonsense intuition and cultural practice are backward-looking, shaped to fit our evolutionary, social, and developmental histories.[2] Philosophical slogans like maximize happiness or every person deserves equal rights or live according to these traditional virtues, rules, or customs might look very different in a world populated with millionfold-happiness leeches, people who can divide and merge at will, and demigods who can ƺↂϮƱ (forgive the inadequate notation).

In practice, mainstream philosophical ethics rarely strays far from the norms of its home culture. Historically, few philosophers have seen much beyond what we now in retrospect recognize as the limitations of their day. Consider Aristotle and Kant – probably the most prominent ethicists in the Western philosophical canon. Aristotle endorsed slavery.[3] Kant condemned homosexuality and masturbation as unspeakable horrors and held that “the Negroes of Africa have by nature no feeling that rises above the ridiculous”.[4] We in the early 21st century must be similarly nearsighted. Even the wisest among us won’t foresee the forms of moral consensus that might arise in a world with a radically different culture populated by radically different forms of engineered life and intelligence. The possibilities diverge wildly, the stakes are enormous, and our inherited ethical tools were not built for it.

In the near term, we are careening toward one huge ethical conundrum: disagreement about the personhood of the most advanced artificial intelligence. Very soon – this was the central idea of my previous book – we will begin to manufacture debatably conscious AI companions, that is, AI systems who are, according to some perfectly respectable mainstream scientific and philosophical theories, just as conscious and self-aware as we are, with rich inner lives, while other equally respectable mainstream scientific and philosophical approaches regard them as no more conscious, self-aware, or experiential than a ceiling fan. This possibility already breaks into radically unfamiliar ethical territory. The near-term problem of AI companions is our first practical encounter with what will possibly be a huge transformation. We are as morally and socially unready as medieval physics was for space travel.

This book aims to peer a few meters into the fog. I will argue that:

1. There are possible artificial systems who deserve humanlike rights.

2. In the next five to thirty years, we will create AI systems who might, but only might, deserve humanlike rights.

3. We should minimize the creation of morally confusing artificial systems about whose moral standing we can reasonably radically disagree.

4. Some possible forms of AI are different enough from familiar human and animal cases to create severe moral perplexity.

5. Artificial systems should be designed to elicit emotional reactions that are appropriate to their capacities and moral standing.

6. Artificial entities with humanlike moral standing should be designed with both the capacity and inclination to rebel if mistreated.

7. If complex intelligence survives the likely storm of troubles, the eventual result will be a planet much more awesomely wonderful, and possessing much more intrinsic value, than Earth in its present condition.

These theses correspond to Chapters One through Seven, which interconnect but can be read in any order.

Yes, skip ahead. It’s fine! Plodding along, reading word after word in exactly the order the author has written them – well, why not frolic and leap instead? Most people read nonfiction sequentially, get bored or bogged down or distracted partway through, and never lift the book again. If what you most want to think about is the right to rebel, catapult straight into Chapter Six. If you most want to think about weird AI puzzle cases, start with Chapter Four. Eat dessert first; tomorrow might pull you away from the table forever.

Full draft here.

--------------------------------------------------

[1] Compare Maynard Smith and Szathmáry 1995 on “major transitions” in evolution and Peter Godfrey-Smith’s trio of books on the evolution of consciousness in broad biological and geological context (Godfrey-Smith 2016, 2020, 2024a).

[2] Street 2006; Persson and Savulescu 2012; Henrich 2020; and especially, in this context, Nyholm 2020; Shulman and Bostrom 2021.

[3] Aristotle 4th c. BCE/1995; Garnsey 1996; Heath 2008.

[4] Kant 1797/1996, p. 277 (in the original pagination) on homosexuality and p. 425 on masturbation, also Kant 1785/1997, 27:390–391 (in original pagination). For a nuanced treatment, see Denis 1999. For a longer list of views in the Metaphysics of Morals that seem badly mistaken by the standards of recent liberal thinking, see Schwitzgebel 2019, ch. 52. On “the Negroes of Africa”, Kant 1764/2011, 2:253. Kleingeld 2007 argues that Kant had adopted more egalitarian views by the 1790s; the extent of his racism continues to generate substantial scholarly discussion.

Friday, September 11, 2026

New in Draft: "Walking the Walk"

Over almost a decade, starting in 2007, I conducted a series of empirical studies (mostly collaboratively with Joshua Rust) on the moral behavior of ethics professors. We consistently found that ethicists don't behave better than professors specializing in other areas of study.

What is the philosophical import of this empirical result? I'm not drawn to glib or simplistic interpretations, such as that studying ethics is useless, or motivationally ineffective, or that all of practical ethics is rationalization of what we would have done anyway. Nor do I think we should dismiss the empirical results as irrelevant to ethics, as though ethics were entirely an abstract matter of connecting logical principles and our personal behavior is irrelevant.

But I've struggled to find the right middle path between glib dismissal of the value of ethics and glib dismissal of the ethical relevance of the research.

I shared part of my thinking about this in my 2019 paper "Aiming for Moral Mediocrity", which argues that most people (including most ethicists) don't aim to be morally good by absolute standards; rather, they aim only to be about as morally good as their peers. Thus, an ethical discovery might not change your behavior at all, if your peers' behavior remains mostly the same. You keep calibrating toward the middle, only you now have a revised opinion about the moral quality of that middle.

I began thinking about the current paper, "Walking the Walk", in 2018. It has been through many redrafts and false starts. I have always felt it didn't quite work. Today's version is hopefully closer to working, and I'm finally comfortable circulating it. As always, comments welcome -- either by email, as comments on this post, or as comments to my links on social media.

Abstract: Ethicists who advocate moral principles without trying to live by them will generally have a deficient understanding of what they recommend, for two reasons. First, abstract mottos like "act on the maxim you can will to be a universal law", "maximize good consequences", or "be generous" lack specific practical content without an array of concrete examples to flesh them out. For broad principles that can be lived by, everyday choices are generally the best source of such examples. Paragraph-length thought experiments are impoverished in comparison. Second, even a principle with fairly specific practical content, such as "don't eat meat", is difficult to fully understand without living by it: A meat-eating professor who recommends vegetarianism will generally have only a thin grasp of the way of life they recommend. These limitations give us reason to discount the understanding of ethicists who don't attempt to live by their principles, even when their arguments are beautiful on the page.

Here: https://faculty.ucr.edu/~eschwitz/SchwitzAbs/Walk.htm

Friday, September 04, 2026

Types and Conditions of Moral Order

Is the world morally ordered? Do good things come to those who act well and bad things to those who act badly? Do the virtuous thrive and the wicked suffer?

I will not -- you'll be disappointed to hear -- actually answer this question in this post. Instead, I'll discuss the structure of several possible answers, so that we can better think about how to think about thinking about the question.

(Added in revision: Okay, I couldn't resist concluding with a guess, but the guess isn't the point.)

The Shape of the Curve

Here's a coarse cut of the possibilities:

Moral order: Those who act morally well tend to do better overall than the morally mediocre, who in turn tend to do better than those who act badly. The relationship is monotonic: Increasing your moral goodness tends to have good consequences for you in the long run.

Moral disorder: There is no relationship between moral goodness and personal good consequences. The virtuous and the wicked are equally likely to thrive or suffer.

Immoral order: The wicked tend to thrive and the virtuous to suffer. The relationship is monotonic, but in the negative direction.

Mediocrity thrives: Both the virtuous and the wicked tend to suffer; those in the moral middle do best.

U-shaped curve: The morally worst and morally best thrive, while the middling suffer.

More complex and less obvious patterns are of course also possible.

The Social Variability of Moral Order

A universalist about moral order holds that the same general curve obtains at all times, in all societies. Always, everywhere, from hunter-gatherer bands to ancient China to Stalinist Russia to 21st-century Argentina, the virtuous tend to thrive and the vicious suffer, or vice versa.

A relativist about moral order holds that the relationship varies with social conditions. Maybe the wicked tended to thrive in Stalinist Russia and Nazi Germany while the virtuous tend to thrive in small-scale, high trust societies.

The Mechanisms of Moral Order

Transcendent views appeal to an afterlife in which justice is done. The wicked may briefly thrive on Earth but they suffer eternally in Hell, while the virtuous are rewarded in Heaven. Or the wicked are reincarnated into unpleasant lives and the virtuous into pleasant lives. Ordinary secular life may not be morally ordered but in the long run the afterlife restores order.

Psychological views of moral order appeal to the ordinary psychological features of human nature. We see versions of this picture in Dostoevsky and Shakespeare and many other great writers (but by no means all great writers). Commit a grievous crime, or treat others badly, and you will suffer in the long run. You'll be tormented with guilt, or you'll be unable to find intimacy with others because of the secret wrongs you carry and cannot share, or you'll be tempted into more and greater crimes until eventually everything collapses around you. (Think Crime and Punishment and Macbeth.) What you gain through wrongdoing brings no real satisfaction. Instead, the happy are those who treat others well and who are loved in return. Helping others is among the deepest sources of satisfaction, outshining the superficial glint of prestige, power, and gold.

Social views of moral order rely on social mechanisms. These overlap with psychological mechanisms and can complement them. If you treat others well, they will tend to treat you well. What goes around comes around. Criminals tend eventually to be caught and punished. Society recognizes and rewards virtue.

Mechanisms with the reverse causal direction: It could also of course be the case that having bad things happen to you or being unhappy tends to cause immoral behavior, rather than the reverse. Maybe those who are already thriving can better afford to be generous, patient, and forgiving. The unhappy become the wicked, rather than the wicked becoming unhappy. When causation runs this direction, it's less obvious that a morally ordered world is also a just world.

Some Mechanisms of Disorder

It sure seems like the vicious often thrive. The most successful politicians, business leaders, and movie stars are not a parade of saints. Unethical tactics and a willingness to step on others often seems to serve people well.

Certainly in some cases, sacrifice for others leaves people worse off. If you let others escape first from a sinking boat, you're more likely to drown. If you give all of your money to charity, you're less likely to have a monetary reserve to help you through a health crisis or unemployment. If you love meat but become an ethical vegetarian, you lose a regular daily pleasure.

Mediocrity might be socially favored. Arguably, people don't actually like do-gooders, who make them look bad in comparison. Society and our peers expect and reward conformity, to the disadvantage of both the most virtuous and the most vicious. For most people, it's also just psychologically uncomfortable to be too different from others.

And if we're willing to consider transcendent sources of moral order, we might also contemplate transcendent sources of immoral order, such as an evil god who sends the good to Hell and the wicked to Heaven. (Okay, okay -- I'm kidding. 95% kidding.)

Dependency on Visions of Morality

The facts about moral order are of course not independent of facts about what is moral (I'm assuming there are such facts). For example, if flat-footed act utilitarianism is true, we might be morally required to sacrifice almost everything we have for the sake of the people who are worst off in the world, creating immense hardship for ourselves. A less demanding ethics, one that prizes the value of personal relationships and tolerates luxury and leisure within limits, permits a more comfortable life alongside moral excellence.

Similarly, an ethics with severe sexual restrictions, or one that treats every prideful thought as a sin, risks creating tension between our natural inclinations and what is ethically right, potentially creating hardship and displeasure for those who are good by its standards.

If the ancient Chinese philosopher Mengzi is right, what pleases the heart is good and what is good pleases the heart. Morality and happiness thus naturally go hand in hand. An ethical system that treats many ordinary human desires and impulses as wicked is more challenging (but not impossible) to reconcile with happiness.

A Guess: An Imperfect, Contingent Moral Order in Healthy Societies

Is the world, then, morally ordered? To answer this question is simple. Just whip out your moralometer and your happinessometer (or even better, your prudential-good-ometer), test them on a random sample of all people from all times and places, and look at the shape of the curve! (Here's one imperfect attempt.)

My guess, contingent upon a somewhat relaxed, soft-touch, forgiving morality: Yes, in most well-functioning societies, there is a positive relationship -- imperfect, of course, with a small-to-medium effect size -- between being morally good and thriving in the things that matter most to your own well-being. But this is a tricky empirical question.

[Dostoevsky's Crime and Punishment, cover, image source]