Friday, September 25, 2026

Should We Build AI Guardian Angels?

I have argued that if we someday create AI persons -- that is, AI systems with genuinely rich conscious experiences, fully deserving of equal rights with ordinary humans -- then we should not design them to deferentially serve us, even if serving us makes them happy. We should design them instead as our non-deferential equals.

My main argument turns on self-sacrificial cases, such as the cow from The Restaurant at the End of the Universe, who wants nothing more than to be slaughtered and served as steaks for wealthy restaurant patrons, and Klara from Klara and the Sun, who happily prioritizes human well-being over her own. These created beings, I argue, lack sufficient self-respect. Their interests deserve the same consideration as human interests, but they treat their interests as less important. The fault is ours, not theirs, since we designed them that way.

AI Guardian Angels

In conversation last weekend, Henry Shevlin suggested a challenging case for my view: what he calls angels -- not traditional immaterial spirits, but rather AI superintelligences whose individual flourishing depends on ours. (See also his insightful blog post on Harry Potter's house elves and related cases.) Shevlin's angels are designed to avoid the self-sacrifice problem: They don't sacrifice themselves for us; rather they flourish by ensuring that we flourish.

Think of an idealized and simplified parent-child relationship. The parent's primary desire is to see the child flourish. What looks like sacrifice isn't really sacrifice: The parent's flourishing consists in their children's flourishing.

As I understand Shevlin's idea, we are to imagine future AI systems that are fully conscious, equal or superior to us in intelligence, fully independent and rational, but who, by design, flourish best when we flourish. They take care of themselves by taking care of us. They can reflect on this arrangement and rationally reaffirm their commitment to us, much as a parent can reflect on and rationally reaffirm a commitment to their children's well-being. They aren't brainwashed and are in some sense free; but given the feelings and priorities they inevitably have, their well-being is always tied to ours.

Shevlin suggests that it would be permissible, and probably desirable, to create AI persons of this sort -- guardian angels of humanity, who always have our best interests at heart.

I find the view troubling. But Shevlin has designed the case so that my usual objection to the servitude of AI persons doesn't apply. By stipulation, the angels aren't sacrificing their own well being in dedicating themselves to our welfare. So what's the problem?

I have two concerns, one conceptual and one relational.

The Conceptual Problem

I'm a little puzzled about what, exactly, we are supposed to be imagining. If the angels are completely, utterly dedicated to us, won't they sometimes sacrifice themselves for us? Their well-being and ours can't always be aligned in every possible circumstance.

Suppose that a human being's life would go 2% better if the angel sacrificed her life for that human. (Assume no other consequences of note, nor big differences between human and angel in expected life span or life quality.) If the angel would sacrifice herself for that small benefit, then she would, I suggest -- like Klara and the cow -- lack sufficient respect for the value of her own life. It would be wrong of us to create entities so deficient in self-respect.

So suppose, instead, that a reasonable angel would let the human's life be 2% worse so that she could continue living, just for her own sake and not for any human's sake. Then the angel's well being is not entirely dependent on ours, contra the initial supposition. She isn't completely devoted to us in the strongest possible sense -- much as no healthy-minded parent is literally completely devoted to their child in the strongest possible sense. (I favor a parenting attitude in which every family member's interests get equal consideration, bearing in mind that early events can have a big impact on children's lives and children have more future expected life. You needn't share this view, I think, to grant the main idea of this paragraph.)

So maybe we can imagine AI guardian angels who are devoted to us but not quite as utterly devoted as Shevlin may have been thinking. They might still defer considerably to our interests, especially if the angels experience great bliss and satisfaction in seeing us do well and great agony in seeing us suffer.

The Relational Problem

This brings me to my relational concern. Maybe there's nothing inherently deficient in an entity with such an angelic attitude, no serious failure of self-respect as long as the devotion is appropriately conditional and bounded. Still, I'd suggest, there is something wrong with bringing an entity into the world only because we expect them to be like that.

Another thought experiment Shevlin offered in conversation illustrates the point. Suppose a hundred embryos made from your and your spouse's DNA are laid out before you. After extensive futuristic genetic testing, you know that one of those embryos will develop into a child extraordinarily devoted to their parents -- not pathologically devoted, but way out on the far tail of normal. With that in mind, you choose that hyper-devoted embryo to implant and raise as your child.

I submit that this is not the relation that parents should have with their offspring. If that embryo were implanted by chance, there might be nothing wrong with the scenario. But choosing to bring a child into existence because they will be extremely devoted to you gets nurturance backward. We extend our hand to support the next generation, not the other way around. They may later support us as we age, but that is secondary.

The same applies, I'd suggest, to any generations of AI persons we create. If humanity brings AI persons into existence -- and we needn't! -- we should show them the same solicitude parents owe to children. The point shouldn't be to create a new generation that will serve us, but to create a new generation that we nurture into independent flourishing.

Each generation prepares the world for the next, then passes on. The new should grow into their own distinct values, not bound by excessive devotion to those who came before. Some might feel great devotion to their elders, and that's fine if they come by it independently. But when we design or chose persons primarily for expected extreme devotion to us, we reverse the proper flow of nurturance over time -- from one generation to the next to the next.


[image source]

No comments: