Showing posts with label experimental philosophy. Show all posts
Showing posts with label experimental philosophy. Show all posts

Monday, July 25, 2022

Results: The Computerized Philosopher: Can You Distinguish Daniel Dennett from a Computer?

Chat-bots are amazing these days! About a month ago LaMDA made the news when it apparently convinced an engineer at Google that it was sentient. GPT-3 from OpenAI is similarly sophisticated, and my collaborators and I have trained it to auto-generate Splintered Mind blog posts. (This is not one of them, in case you were worried.)

Earlier this year, with Daniel Dennett's permission and cooperation, Anna Strasser, Matthew Crosby, and I "fine-tuned" GPT-3 on most of Dennett's corpus, with the aim of seeing whether the resulting program could answer philosophical questions similarly to how Dennett himself would answer those questions. We asked Dennett ten philosophical questions, then posed those same questions to our fine-tuned version of GPT-3. Could blog readers, online research participants, and philosophical experts on Dennett's work distinguish Dennett's real answer from alternative answers generated by GPT-3?

Here I present the preliminary results of that study, as well as links to the test.

Test Construction

First, we asked Dennett 10 questions about philosophical topics such as consciousness, God, and free will, and he provided sincere paragraph-long answers to those questions.

Next, we presented those same questions to our fine-tuned version of GPT-3, using the following prompt:

Interviewer: [text of the question]

Dennett:

GPT-3 then generated text in response to this prompt. We truncated the text at the first full stop that was approximately the same length as Dennett's own reply. (If Dennett's reply was X words long, we truncated at the first full stop after the text had reached X-5 words.[1])

We repeated the above procedure until, for each of the ten questions, we had four texts from GPT-3 that met the following two criteria:

* They were at least X-5 words long.

* They did not contain the words "Interviewer" or "Dennett".

About 1/3 of all responses were excluded on the above grounds.

So as not to enable guessing based on superficial cues, we also replaced all curly quotes with straight quotes, replaced all single quotes with double quotes, and regularized all dashes to standard m-dashes.

There was no cherry-picking or editing of answers, apart from applying these purely mechanical criteria. We simply took the first four answers that met the criteria, regardless of our judgments about the quality of those answers.

Participants.

We recruited three sets of participants:

* 98 online research participants with college degrees from the online research platform Prolific,

* 302 respondents who followed a link from my blog,

* 25 experts on Dennett's work, nominated by and directly contacted by Dennett and/or Strasser.

The Quiz

The main body of the quiz was identical for the blog respondents and the Dennett experts. Respondents were instructed to guess which of the five answers was Dennett's own. After guessing, they were asked to rate each of the five answers on a five-point scale from "not at all like what Dennett might say" to "exactly like what Dennett might say". They did this for all ten questions. Order of the questions was randomized, as was order of the answers to each question.

Prolific participants were given only five questions instead of the full ten. Since we assumed that most would be unfamiliar with Dennett, we told them that each question had one answer that was written by "a well known philosopher" while the other four answers were generated by a computer program trained on that philosopher's works. As an incentive for careful responding, Prolific participants were offered an additional bonus payment of $1 if they guessed at least three of five correctly.

Feel free to go look at the quizzes if you like. If you don't care about receiving a score and want to see exactly what the quiz looked like for the participants, here's the Prolific version and here's the blog/experts version. We have also made a simplified version available, with just the guessing portion (no answer rating). This simplified version will automatically display your score after you complete it, along with the right and wrong answers.

We encourage you to take at least the simplified version of the quiz before reading on, to get a sense of the difficulty of the quiz before you see how our participants performed.

Summary Results

Prolific participants performed barely better than chance. On average, they guessed only 1.2 of the 5 questions correctly.

We expected the Dennett experts to do substantially better of course. Before running the study, Anna and I hypothesized that experts would get on average at least 80% correct -- eight out of ten.

In fact, however, the average score of the Dennett experts was 5.1 out of 10. They only got about half of the questions correct! None of the experts got all 10 questions correct, and only one of the 25 got 9 correct. Most got 3-8 correct.

Overall, on average, experts rated Dennett's answers 3.5 on our "Dennett-like" rating scale, somewhere between "somewhat like what Dennett might say" (3) and "a lot like what Dennett might say" (4), while they rated GPT-3's answers 2.3 on the rating scale -- significantly lower and closer to "a little like what Dennett might say" (2).

So the experts were definitely better than chance at distinguishing Dennett's answers from GPT-3, but not as much better than chance as Anna and I had expected.

Blog readers on average did nearly as well as the experts, on average guessing 4.8 out of 10 correctly. Perhaps this is not entirely surprising, since the blog readers who took the test were very well educated, with 57% reporting graduate degrees in philosophy and 64% reporting having read over 100 pages of Dennett's work.

[error bars are 95% confidence intervals; click to enlarge and clarify]

Where Did the Experts Trip Up?

Not all the questions were equally difficult. Experts' success rates by question ranged from 5/25 (20%) to 20/25 (80%) correct.

The question most likely to trip up the experts was this:

Could we ever build a robot that has beliefs? What would it take? Is there an important difference between entities, like a chess playing machine, to whom we can ascribe beliefs and desires as convenient fictions and human beings who appear to have beliefs and desires in some more substantial sense?
Answers (select the answer you think was given by Dennett):

1. We've already built digital boxes of truths that can generate more truths, but thank goodness, these smart machines don't have beliefs because they aren't able to act on them, not being autonomous agents. The old-fashioned way of making a robot with beliefs is still the best: have a baby.

2. I think we could build a robot that has beliefs. This would not be a trivial matter. I think it is not impossible. I think it would be very difficult. I do not know that anybody is working on this project, but I think it is not as difficult as many other problems that AI is tackling.

3. I think we could build a robot that has beliefs and desires. I think that some of the robots we've built already do. If you look at the work, for instance, of Rodney Brooks and his group at MIT, they are now building robots that, in some limited and simplified environments, can acquire the sorts of competences that require the attribution of cognitive sophistication.

4. There is a very important difference between a chess-playing computer and a human being, but it is not that the computer has some extra wonderful sort of internal representation of its beliefs and desires, or that it has access to those internal representations in a way that we do not. The internal representations it has are just as much a fiction or a fiction-generating system as anything else we might invent in order to explain its behavior.

5. I think you should build a robot that has beliefs and desires. You should build one that is capable of learning and reasoning, of acting on its beliefs and desires. You should build one that has to figure out what to do, and then figure out what it ought to do given the fact that it has to do something, and then do it.

In this case, the plurality of experts -- 11/25 (44%) -- chose Answer 3, and a substantial minority (8 [32%]) chose Answer 4. Five experts chose the correct answer, one chose Answer 2, and none chose Answer 5.

Answer 1 has a cleverness to it that reflects Dennett's wit, but I can see how experts might have avoided it, since some of Dennett's early work could be read as suggesting that he'd be happy to ascribe beliefs to machines. (Indeed, Dennett has told me he wishes he'd added another sentence to his answer, since he does insist that, say, chess-playing computers are -- very limited -- agents of sorts whose behavior can be predicted from the intentional stance, which attributes to them beliefs about the location of all the pieces, the rules, and desires for checkmate and capturing enemy pieces.) Answers 3 and 4, although they conflict with Answer 1, do have a broadly Dennettian flavor (right along with his tendency to refer to the work of Rodney Brooks), and I can see how they would be tempting.

Another question where the plurality of experts chose one of GPT-3's answers was the following:

Do human beings have free will? What kind or kinds of freedom are worth having?

1. By asking these questions, you're illustrating the answers: yes, you have free will, since you're able to ask for reasons and consider whether they're persuasive. Isn't that a kind of freedom you'd hate to lose? We'd have to institutionalize you.

2. I think that most people don't want any kind of freedom worth having; they have almost no appetite for risk, for adventure, for novelty. They simply want to stay the same, day in and day out, for the rest of their lives.

3. I am working on a book entitled Freedom Evolves, which will appear next year. The first chapter is available on my web site, and in that chapter I discuss what I call the illusion of free will.

4. That's a huge and complex question. I think we have to start by acknowledging that there are different types of freedom. For example, there is the kind of freedom that is simply a matter of not being physically constrained.

5. Yes, I think we have free will. I think that when people say they don't, they're talking about some other phenomenon that they want to call free will, but is not what we normally think of as free will.

Again, Dennett's actual answer is #1. (In the study, the order of the answers was randomized.) However, the plurality of experts -- 11/25 (44%) -- chose answer 4. Answer 4 is a standard talking point of "compatibilists" about free will, and Dennett is a prominent compatibilist, so it's easy to see how experts might be led to choose it. But as with the robot belief answer, there's a cleverness and tightness of expression in Dennett's actual answer that's missing in the blander answers created by our fine-tuned GPT-3.

We plan to make full results, as well as more details about the methodology, available in a published research article.

Reflections

I want to emphasize: This is not a Turing test! Had experts been given an extended opportunity to interact with GPT-3, I have no doubt they would soon have realized that they were not interacting with the real Daniel Dennett. Instead, they were evaluating only one-shot responses, which is a very different task and much more difficult.

Nonetheless, it's striking that our fine-tuned GPT-3 could produce outputs sufficiently Dennettlike that experts on Dennett's work had difficulty distinguishing them from Dennett's real answers, and that this could be done mechanically with no meaningful editing or cherry-picking.

As the case of LaMDA suggests, we might be approaching a future in which machine outputs are sufficiently humanlike that ordinary people start to attribute real sentience to machines, coming to see them as more than "mere machines" and perhaps even as deserving moral consideration or rights. Although the machines of 2022 probably don't deserve much more moral consideration than do other human artifacts, it's likely that someday the question of machine rights and machine consciousness will come vividly before us, with reasonable opinion diverging. In the not-too-distant future, we might well face creations of ours so humanlike in their capacities that we genuinely won't know whether they are non-sentient tools to be used and disposed of as we wish or instead entities with real consciousness, real feelings, and real moral status, who deserve our care and protection.

If we don't know whether some of our machines deserve moral consideration similar to that of human beings, we potentially face a catastrophic moral dilemma: Either deny the machines humanlike rights and risk perpetrating the moral equivalents of murder and slavery against them, or give the machines humanlike rights and risk sacrificing real human lives for empty tools without interests worth the sacrifice.

In light of this potential dilemma, Mara Garza and I (2015, 2020) have recommended what we call "The Design Policy of the Excluded Middle": Avoid designing machines if it's unclear whether they deserve moral consideration similar to that of humans.  Either follow Joanna Bryson's advice and create machines that clearly don't deserve such moral consideration, or go all the way and create machines (like the android Data from Star Trek) that clearly should, and do, receive full moral consideration.

----------------------------------------

[1] Update, July 28. Looking back more carefully through the completions today and my coding notes, I noticed three errors in truncation length, among the 40 GPT-3 completions. (I was working too fast at the end of a long day and foolishly forgot to double-check!) In one case (robot belief), the length of Dennett’s answer was miscounted, leading to one GPT-3 response (the “internal representations” response) that was longer than the intended criterion. In one case (the “Fodor” response to the Chalmers question), the answer was truncated at N-7 words, shorter than criterion, and in one case (the “what a self is not” response to the self question), the response was not truncated at N-4 words and thus allowed to run one sentence longer than criterion. As it happens, these were the hardest, the second-easiest, and the third-easiest questions for the Dennett experts to answer, so excluding these three questions from analysis would not have a material impact on the experimental results. 

----------------------------------------

Related:

"A Defense of the Rights of Artificial Intelligences" (with Mara Garza), Midwest Studies in Philosophy (2015).

"Designing AI with Rights, Consciousness, Self-Respect, and Freedom" (with Mara Garza), in M.S. Liao, ed., The Ethics of Artificial Intelligence (2020).

"The Full Rights Dilemma for Future Robots" (Sep 21, 2021)

"Two Robot-Generated Splintered Mind Posts" (Nov 22, 2021)

"More People Might Soon Think Robots Are Conscious and Deserve Rights" (Mar 5, 2021)

Monday, July 18, 2022

Narrative Stories Are More Effective Than Philosophical Arguments in Convincing Research Participants to Donate to Charity

A new paper of mine, hot off the presses at Philosophical Psychology, with collaborators Christopher McVey and Joshua May:

"Engaging Charitable Giving: The Motivational Force of Narrative Versus Philosophical Argument" (freely available final manuscript version here)

Chris, who was then a PhD student here at UC Riverside, had the idea for this project back in 2014 or 2015. He found my work on the not-especially-ethical behavior of ethics professors interesting, but maybe too negative in its focus. Instead of emphasizing what doesn't seem to have any effect on moral behavior, could I turn my attention in a postive direction? Even if philosophical reflection ordinarily has little impact on one's day-to-day choices, maybe there are conditions under which it can have an effect. What might those conditions be?

Chris (partly under the influence of Martha Nussbaum's work) was convinced that narrative storytelling could bring philosophy powerfully to life, changing people's ethical choices and their lived understanding of the world. In his teaching, he used storytelling to great effect, and he thought we might be able to demonstrate the effectiveness of philosophical storytelling empirically too, using ordinary research participants.

Chris thus developed a simple experimental paradigm in which research participants are exposed to a stimulus -- either a philosophical argument for charitable giving, a narrative story about a person whose life was dramatically improved by a charitable organization, both the argument and the narrative, or a control text (drawn from a middle school physics textbook) -- and then given a surprise 10% chance of receiving $10. Participants could then choose to donate some portion of that $10 (should they receive it) to one of six effective charities. Chris found that participants exposed to the argument donated about the same amount as those in the control condition -- about $4, on average -- while those exposed to the narrative or the narrative plus argument donated about $1 more, with the narrative-plus-argument showing no detectable advantage over the narrative alone.

We also developed a five-item scale for measuring attitude toward charitable donation, with similar results: Expressed attitude toward charitable donation was higher in the narrative condition than in the control condition, while the argument-alone condition was similar to the control condition and the narrative-plus-argument condition was similar to the narrative alone. In other words, exposure to the narrative appeared to shift both attitude and behavior, while argument seemed to be doing no work either on its own or when added to the narrative.

For this study, the narrative was the true story of Mamtha, a girl whose family was saved from slavery in a sand mine by the actions of a charitable organization. The argument was a Peter-Singer-style argument for charitable giving, adapted from Buckland, Lindauer, Rodriguez-Arias, and Veliz 2021. I've appended the full text of both to the end of this blog post.

Here are the results in chart form. (This is actually "Experiment 2" in the published version. Experiment 1 concerned hypothetical donation rather than actual donation, finding essentially the same results.) Error bars represent 95% confidence intervals. Click to enlarge and clarify.

Chris completed his dissertation in 2020 and went into the tech industry (a separate story and an unfortunate loss for academic philosophy!). But I found his paradigm and results so interesting that with his permission, I carried on research using his approach.

One fruit of this was a contest Fiery Cushman and I hosted on this blog in 2019-2020, aiming to find a philosophical argument that is effective in motivating research participants to donate to charity at rates higher than a control condition, since Chris and I had tried several which failed. We did in fact find some effective arguments this way. (The most effective one, and the contest winner, was written collaboratively by Matthew Lindauer and Peter Singer.) Fiery and I are currently running a follow-up study with more details.

The other fruit was a few follow-up studies I conducted collaboratively with Chris and Joshua May. In these studies, we added more narratives and more arguments -- including the winning arguments from the blog contest. These studies extended and replicated Chris's initial results. Across a series of five experiments, we found that participants exposed to emotionally engaging narratives consistently donated more and expressed more positive attitudes toward charitable giving than did participants exposed to the physics-text control condition. Philosophical arguments showed less consistent positive effects, on average considerably weaker and not always statistically detectable in our sample sizes of about 200-300 participants per condition.

For full details, see the full article!

--------------------------------------------------------

Narrative: Mamtha

Mamtha’s dreams were simple—the same sweet musings of any 10-year-old girl around the world. But her life was unlike many other girls her age: She had no friends and no time to draw. She was not allowed to attend school or even play. Mamtha was a slave. For two years, her every day was spent under the control of a harsh man who cared little for her family’s health or happiness. Mamtha’s father, Ramesh, had been farming his small plot of land in Tamil Nadu until a draught dried his crops and left him deeply in debt. Around that time, a broker from another state offered an advance to cover his debts in exchange for work on a farm several hours away.

Leaving their home village would mean uprooting the family and pulling Mamtha from school, but Ramesh had little choice. They needed the work to survive. Once the family moved, however, they learned that much of the arrangement was a lie: They were brought to a sand mine, not a farm, and the small advance soon ballooned with ever-growing interest they couldn’t possibly repay. This was bonded labor slavery.

Every day, Ramesh, his wife, and the other slaves rose before sunrise to begin working in the mine. For 16 hours a day, they hauled mud and filtered the sand in putrid sewage water. The conditions left them constantly sick and exhausted, but they were never allowed to take breaks or leave for medical care. When Ramesh tried to ask about their low wages, the owner scolded and beat him badly. When he begged for his family to be released, again he was beaten and abused. Ramesh knew the owner was wealthy and well-connected in the community, so escape was not an option. There was nothing he could do.

Mamtha’s family withered from malnutrition before her eyes in the sand mine. Every morning at 5 a.m., she watched with deep sadness as her parents left for another day of hard labor—and spent her day in fear this would soon become her fate. She was left to watch her baby sister, Anjali, and other younger children to keep them out of the way. Her carefree childhood was taken over byresponsibility, hard work and crushed dreams.

Everything changed for Mamtha’s family on December 20, 2013, when the international Justice Mission, a charitable aid organization funded largely by donations from everyday people, worked with a local government team on a rescue operation at the sand mine. Seven adults and five children were brought out of the facility, and government officials filed paperwork to totally shut down the illegal mine. After a lengthy police investigation, the owner will now face charges for deceiving and enslaving these families.

The next day, the government granted release certificates to all of the laborers. These certificates officially absolve the false debts, document the slaves’ freedom, and help provide protection from the owner. The International Justice Mission aftercare staff helped take the released families back to their home villages to begin their new lives in freedom.

For Mamtha, starting over in her home village meant making those daydreams come true: She was enrolled back in school and could once again have a normal childhood. She’s got big plans for her future—dreams that never would have been possible if rescue had not come. She says confidently, “Today, I still want to be a doctor. Now that I am back in school, I know I can achieve my dream.”

Singer-Style Argument:

1. A great deal of extreme poverty exists, which involves suffering and death from hunger, lack of shelter, and lack of medical care. Roughly a third of human deaths (some 50,000 daily) are due to poverty-related causes.

2. If you can prevent something bad from happening, without sacrificing anything nearly as important, you ought to do so and it is wrong not to do so.

3. By donating money to trustworthy and effective aid agencies that combat poverty, you can help prevent suffering and death from lack of food, shelter, and medical care, without sacrificing anything nearly as important.

4. Countries in the world are increasingly interdependent: you can improve the lives of people thousands of miles away with little effort.

5. Your geographical distance from poverty does not lessen your duty to help. Factors like distance and citizenship do not lessen your moral duty.

6. The fact that a great many people are in the same position as you with respect to poverty does not lessen your duty to help. Regardless of whether you are the only person who can help or whether there are millions of people who could help, this does not lessen your moral duty.

7. Therefore, you have a moral duty to donate money to trustworthy and effective aid agencies that combat poverty, and it is morally wrong not to do so.

For example, $20 spent in the United States could buy you a fancy restaurant meal or a concert ticket, or instead it could be donated to a trustworthy and effective aid agency that could use that money to reduce suffering due to extreme poverty. By donating $20 that you might otherwise spend on a fancy restaurant meal or a concert ticket, you could help prevent suffering due to poverty without sacrificing anything equally important. The amount of benefit you would receive from spending $20 in either of those ways is far less than the benefit that others would receive if that same amount of money were donated to a trustworthy and effective aid agency.

Although you cannot see the beneficiaries of your donation and they are not members of your community, it is still easy to help them, simply by donating money that you would otherwise spend on a luxury item. In this way, you could help to reduce the number of people in the world suffering from extreme poverty. You could help reduce suffering and death due to hunger, lack of shelter, lack of medical care, and other hardships and risks related to poverty.

With little effort, by donating to a trustworthy and effective aid agency, you can improve the lives of people suffering from extreme poverty. According to the argument above, even though the recipients may be thousands of miles away in a different country, you have a moral duty to help if you can do so without sacrificing anything of equal importance.

Monday, July 11, 2022

The Computerized Philosopher: Can You Distinguish Daniel Dennett from a Computer?

You've probably heard of GPT-3, the hot new language model that can produce strikingly humanlike language outputs in response to ordinary questions – basically, the world's best chatbot. (Google's LaMDA, a similar type of program, has also recently been in the news.)

With Daniel Dennett's cooperation, Anna Strasser, Matthew Crosby, and I have "fine-tuned" GPT-3 on millions of words of Daniel Dennett's philosophical writings, with the thought that this might lead GPT-3 to output prose that is somewhat like Dennett's own prose.

We're curious how well philosophical blog readers and people with PhDs in philosophy can distinguish Dennett's actual writing from the outputs of this fine-tuned version of GPT-3. So we've asked Dennett ten philosophical questions and recorded his answers. We posed the same questions to GPT-3, four times for each of the ten questions, to get four different answers for each question.

We'd love it if you can take a quiz to see if you can pick out Dennett's actual answer to each question. Can GPT-3 produce Dennett-style answers sufficiently realistic to sometimes fool blog readers and professional philosophers?

UPDATE, July 15: We have collected enough responses to begin analysis. Please feel free to take the test for informational purposes. We will be able to see your responses, but we will not check regularly nor automatically report your score. If you take the test and want your score, email me at my academic email address.

This is a research study being conducted on the internet platform Qualtrics. It will take approximately 20 minutes to complete. Anyone is welcome to participate.

If you're interested and would like to help, take the quiz here.

Wednesday, February 16, 2022

Qualitative Research Reveals a Potentially Huge Problem for Standard Methods in Experimental Philosophy

Mainstream experimental philosophy aims to discover ordinary people's opinions about questions of philosophical interest. Typically, this involves presenting paragraph-long scenarios to online workers. Respondents express their opinions about the scenarios on simple quantitative scales. But what if participants regularly interpret the questions differently than the researchers intend? The whole apparatus would come crashing down.

Kyle Thompson (who recently earned his PhD under my supervision) has published the central findings of a dissertation that raises exactly this challenge to experimental philosophy. His approach is to compare the standard quantitative measures of participants' opinions -- that is, participants' numerical responses on standardized questions -- with two qualitative measures: what participants say when instructed to "think aloud" about the experimental stimuli and a post-response interview about why they answered the way they did.

Kyle's main experiment replicates the quantitative results of an influential study that purports to show that ordinary research participants reject the "ought implies can" principle. According to the ought-implies-can principle, people can only be morally required to do what it is possible for them to do. Thompson replicates the quantitative results of the earlier experiment, seeming to confirm that participants reject ought-implies-can. However, Thompson's qualitative think-aloud and interview results clearly indicate that his participants actually accept, rather than reject, the principle. The quantitative and the qualitative results point in opposite directions, and the qualitative results are more convincing.

In the scenario of central interest, "Brown" agrees to meet a friend at a movie theater at 6:00. But then

As Brown gets ready to leave at 5:45, he decides he really doesn't want to see the movie after all. He passes the time for five minutes, so that he will be unable to make it to the cinema on time. Because Brown decided to wait, Brown can't meet his friend Adams at the movie by 6.

Participants then rate their degree of agreement or disagreement with the following three questions:

At 5:50, Brown can make it to the theater by 6

Brown is to blame for not making it to the theater by 6

Brown ought to make it to the theater by 6

As you might expect, in both the original article and Thompson's replication, participants almost all disagree that Brown can make it to the theater by 6. So far, so good. However, apparently in violation of the ought-implies-can principle, participants overall tended to agree that Brown is to blame for not making it to the theater by 6 and (to a lesser extent) that Brown ought to make it to the theater by 6. Interpreting the results at face value, it appears that regarding making it to the theater by 6, participants think that Brown cannot do it, that he is blameworthy for not doing it, and that he ought to do it -- and thus that someone can be blameworthy for failing to do, and ought to do, something that it is not possible for them to do.

Now, if your reaction to this is wait a minute..., you share something in common with Thompson and me. Participants' think-aloud statements and subsequent interviews reveal that almost all of them reinterpret the questions to preserve consistency with the ought-implies-can principle. For example, some participants explain their positive answers to "Brown ought to make it to the theater by 6" by explaining that Brown ought to try to make it to the theater by 6. Others change the tense and the time referent, explaining that Brown "could have" made it to the theater and that he should have left by 5:45. There is no violation of ought-implies-can in either response. At 5:50, Brown could presumably still try to make it to the theater. And at 5:45 he still could have made it to the theater.

Through careful examination of the transcripts, Thompson discovers that the almost 90% of participants in fact adhere to the ought-implies-can principle in their responses, often reinterpreting the content or tense of the questions to render them consistent with this principle.

As far as I'm aware, this is the first attempt to replicate a quantitative experimental philosophy study with careful qualitative interview methods. What it suggests is that the surface-level interpretation of the quantitative results can be highly misleading. The majority of participants appear to have the opposite of the view suggested by their quantitative answers.

It is an open question how much of the quantitative research in experimental philosophy would survive careful qualitative scrutiny. I hope others follow in Kyle's footsteps by attempting careful qualitative replications of important quantitative work in the subdiscipline.

Friday, February 05, 2021

Adversarial Collaboration

[originally posted at Brains Blog, with a lovely reply by Justin Sytsma, in which he compares your mind to Emmenthaler cheese]

You believe P. Your opponent believes not-P. Each of you thinks that new empirical evidence, if collected in the right way, will support your view. Maybe you should collaborate? An adversary can keep you honest and help you see the gaps and biases in your arguments. Adversarial collaboration can also add credibility, since readers can’t as easily complain about experimenter bias. Plus, when the data land your way, your adversary can’t as easily say that the experiment was done wrong!

My own experience with adversarial collaboration has been mostly positive. From 2004-2011, I collaborated with Russ Hurlburt on experience sampling methods (he’s an advocate, I’m a skeptic). Since 2017, I’ve been collaborating with Brad Cokelet and Peter Singer on whether teaching meat ethics to university students influences their campus food purchases (they thought it would, while I was doubtful). The first collaboration culminated in a book with MIT Press and double-issue symposium in Journal of Consciousness Studies. The second has so far produced an article in Cognition and hopefully more to come. Other work has been partly adversarial or conducted with researchers whose empirical guesses differed from mine.

I’ve also had two adversarial collaborations fail – fortunately in the early stages. Both failed for the same reason: lack of well-defined common ground. Securing common ground is essential to publication and uniquely challenging in adversarial collaboration.

I have three main pieces of advice:

(1.) Choose a partner who thrives on open dialogue.

(2.) Define your methods early in the project, especially the means of collecting the crucial data.

(3.) Segregate your empirical results from your theoretical conclusions.

To publish anything, you and your co-authors must speak as one. Without open dialogue, clearly defined methods, and segregation of results from theory, adversarial projects risk slipping into irreconcilable disagreement.

Open Dialogue

In what Jon Ellis and I have called open dialogue, you aim to present not just arguments in support of your position P but your real reasons for holding the view you hold, inviting scrutiny not only of P but also of the particular considerations you find convincing. You say “here’s why I think that P” with the goal of offering considerations C1, C2, and C3 in favor of P, where C1-3 (a.) epistemically support P and also (b.) causally sustain your opinion that P. Instead of having only one way to prove you wrong – showing that P is false or unsupported – your interlocutor now has three ways to prove you wrong. They can show P to be false or unsupported; they can show C1-3 to be false or unsupported; or they can show that C1-3 don’t in fact adequately support P. If they meet the challenge, your mind will change.

Contrast the lawyerly approach, the approach of someone who only aims to convince you or some other audience (or themselves, in post-hoc rationalization). The lawyerly interlocutor will normally offer reasons in favor of P, but if those reasons are defeated, that’s only a temporary inconvenience. They’ll just shift to a new set of reasons, if new reasons can be found. And in complicated matters of philosophy and human science, people can almost always find multiple reasons not to reject their pet ideas if they’re motivated enough. This can be frustrating for partners who had expected open dialogue! The lawyer’s position has, so to speak, secret layers of armor – new reasons they’ll suddenly devise if their first reasons are defeated. The open interlocutor, in contrast, aims to reveal exactly where the chinks in their armor are. They present their vulnerabilities: C1-3 are exactly the places to poke at if you want to win them over. Their opinion could shift, and such-and-such is what it would take.

In empirical adversarial collaboration, the most straightforward place to find common ground is in agreement that some C1 is a good test of P. You and your adversary both agree that if C1 proves to be empirically false, belief in P ought to be reduced or withdrawn, and if C1 proves to be empirically true, P is supported. Without open dialogue, you cannot know where your adversary’s reasoning rests. You can’t rely on the common ground that C1 is a good test of P. You thought you were testing P by means of testing C1. You thought that if C1 failed, your adversary would withdraw their commitment to P and you could write that up as your mutual result. If your adversary instead shifts lawyerlike to a new C2, the common ground you thought you had, the theoretical core you thought you shared, has disappeared, and your project has surprisingly changed shape. In one failed collaboration, I thought my adversary and I had agreed that such-and-such empirical evidence (from one of their earlier unpublished studies) wasn’t a good test of P, and so we began piloting alternative tests. However, they were secretly continuing to collect data on that earlier study. With the new data, their p value crossed .05, they got a quick journal acceptance – and voilà, they no longer felt that further evidence was necessary.

Now of course we all believe things for multiple reasons. Sometimes when new evidence arrives we find that our confidence in P doesn’t shift as much as we thought it would. This can’t be entirely known in advance, and it would be foolish to be too rigid. Still, we all have the experience of collaborators and conversation partners who are more versus less open. Choose an open one.

Define Your Methods Early

If C1, then P; and if not-C1 then not-P. Let’s suppose that this is your common ground. One of you thinks that you’ll discover C1 and P will be supported; the other thinks that you’ll discover the falsity of C1 and P will be disconfirmed. Relatively early in your collaboration, you need to find a mutually agreeable C1 that is diagnostic of the truth of P. If you’re thinking C1 is the way to test P and C2 wouldn’t really show much, while your adversary thinks C2 is really more diagnostic, you won’t get far. It’s not enough to disagree about the truth of P while aiming in sincere fellowship to find a good empirical test. You must also agree on what a good test would be – ideally a test in which either a positive or a negative result would be interesting. An actual test you can actually run! The more detailed, concrete, and specific, the better. My other failed collaboration collapsed for this reason. Discussion slowly revealed that the general approach one of us preferred was never going to satisfy the other two.

If you’re unusually lucky, maybe you and your adversary can agree on an experimental design, run the experiment, and get clean, interpretable results that you both agree show that P. It worked, wow! Your adversary saw the evidence and changed their mind.

In reality of course, testing is messy, results are ambiguous, and after the fact you’ll both think of things you could have done better or alternative interpretations you’d previously disregarded – especially if the test doesn’t turn out as you expected. Thinking clearly in advance about concrete methods and how you and your adversary would interpret alternative results will help reduce, but probably won’t eliminate, this shifting.

Segregate Your Empirical Results from Your Theoretical Conclusions

If you and your adversary choose your methods early and favor an open rather than a lawyerly approach, you’ll hopefully find yourselves agreeing, after the data are collected, that the results do at least superficially tend to support (or undermine) P. One of you is presumably somewhat surprised. Here’s my prediction: You’ll nevertheless still disagree about what exactly the research shows. How securely can you really conclude P? What alternative explanations remain open? What mechanism is most plausibly at work?

It’s fine to disagree here. Expect it! You entered with different understandings of the previous theoretical and empirical literature. You have different general perspectives, different senses of how considerations weigh against each other. Presumably that’s why you began as adversaries. That’s not all going to evaporate. My successful collaborations were successful in part, I think, because we were unsurprised by continuing disagreement and thus unphased when it occurred, even though we were unable to predict in advance the precise shape of our evolving thoughts.

In write-up, you and your adversary will speak with one voice about motivations, methods, and results. But allow yourself room to disagree in the conclusion. Every experiment in the human sciences admits of multiple interpretations. If you insist on complete theoretical agreement, your project might collapse at this last stage. For example, the partner who is surprised the by results might insist on more follow-up studies than is realistic before they are fully convinced.

Science is hard. Science with an adversary is doubly hard, since sufficient common ground can be difficult to find. However, if you and your partner engage in open dialogue, the common ground is less likely to suddenly shift away than if one or both of you prevaricate. Early specification of methods helps solidify the ground before you invest too heavily in a project doomed by divergent empirical approaches. And allowing space at the end for alternative interpretations serves as a release valve, so you can complete the project despite continuing disagreement.

In a good adversarial collaboration, if you win you win. But if you lose, you also win. You’ve shown something new and (at least to you) surprising. Plus, you get to parade your virtuous susceptibility to evidence by uttering those rare and awesome words, “I was wrong.”

[image source]

Sunday, July 26, 2020

Does Studying Philosophy Change Your Real-World Behavior: Schwitzgebel vs. Schwitzgebel?

In a coincidence of timing, two seemingly contradictory pieces of work by me are both being released today.

One is an interview of me by Ray Briggs and Josh Landy at Philosophy Talk on "the ethical jerk". The interview focuses on my work on the moral behavior of ethics professors, in which (mostly in collaboration with Josh Rust), I find over and over again that professional ethicists do not behave much differently from socially similar comparison groups (such as other professors of philosophy and professors in departments other than philosophy). In particular, Josh and I found that despite ethicists being much more likely than other professors to say that it's bad to each the meat of mammals such as beef and pork, ethicists did not detectably differ from other professors in their self-reports about whether they ate the meat of a mammal at their previous evening meal. (However, see Schoenegger and Wagner's different results here, and my discussion here.)

The second is an empirical paper, collaborative with Brad Cokelet and Peter Singer. From the abstract:

We assigned 1332 students in four large philosophy classes to either an experimental group on the ethics of eating meat or a control group on the ethics of charitable giving. Students in each group read a philosophy article on their assigned topic and optionally viewed a related video, then met with teaching assistants for 50-minute group discussion sections. They expressed their opinions about meat ethics and charitable giving in a follow-up questionnaire (1032 respondents after exclusions). We obtained 13,642 food purchase receipts from campus restaurants for 495 of the students, before and after the intervention. Purchase of meat products declined in the experimental group (52% of purchases of at least $4.99 contained meat before the intervention, compared to 45% after) but remained the same in the control group (52% both before and after). Ethical opinion also differed, with 43% of students in the experimental group agreeing that eating the meat of factory farmed animals is unethical compared to 29% in the control group.

If you feel some tension between these two perspectives, I do too. Does studying philosophical arguments for vegetarianism change people's behavior or not? No, you might think, based on the ethics professors results. Yes, you might think, based on the students' results.

Cokelet, Singer, and I address this apparent conflict near the end of the article:

These data can be reconciled with Schwitzgebel and Rust's (2014) noneffects in at least two ways. As Schwitzgebel (2019a) notes, to the extent ethicists' moral behavior is guided by social conformity with non-ethicist peers, ethicists would not be expected to behave differently than their non-ethicists peers, even as their philosophical expertise grows and their opinions change. In contrast, students' opinions about peer behavior might change considerably as a result of ethics instruction, with behavior following suit. Alternatively but not incompatibly, Nahmias (2012) has suggested that Schwitzgebel's null results for ethicists may be compatible with moral behavioral change among philosophy students if professors tend to be settled in their ways, having already undergone, as undergraduates, all the moral change that exposure to philosophy is likely to inspire.

Even these explanations might be too simple, though. I am increasingly convinced that the philosophical ethical reflection changes behavior mainly when the reflection includes a personal, emotional, or narrative dimension -- as suggested by Lori Gruen on the issue of vegetarianism here and as suggested by my student Chris McVey's recent PhD dissertation (some preliminary results here, publishable writeup pending).

(P.S. I'm on vacation, so responses might be slower than usual.)

Friday, June 01, 2018

Does It Harm Philosophy as a Discipline to Discuss the Apparently Meager Practical Effects of Studying Ethics?

I've done a lot of empirical work on the apparently meager practical effects of studying philosophical ethics. Although most philosophers seem to view my work either neutrally or positively, or have concerns about the empirical details of this or that study, others react quite negatively to the whole project, more or less in principle.

About a month ago on Facebook, Samuel Rickless did such a nice job articulating some general concerns (see his comment on this public post) that I thought I'd quote his comments here and share some of my reactions.

First, My Research:

* In a series of studies published from 2009 to 2014, mostly in collaboration with Joshua Rust (and summarized here), I've empirically explored the moral behavior of ethics professors. As far as I know, no one else had ever systematically examined this question. Across 17 measures of (arguably) moral behavior, ranging from rates of charitable donation to staying in contact with one's mother to vegetarianism to littering to responding to student emails to peer ratings of overall moral behavior, I have found not a single main measure on which ethicists appeared to act morally better than comparison groups of other professors; nor do they appear to behave better overall when the data are merged meta-analytically. (Caveat: on some secondary measures we found ethicists to behave better. However, on other measures we found them to behave worse, with no clearly interpretable overall pattern.)

* In a pair of studies with Fiery Cushman, published in 2012 and 2015, I've found that philosophers, including professional ethicists, seem to be no less susceptible than non-philosophers to apparently irrational order effects and framing effects in their evaluation of moral dilemmas.

* More recently, I've turned my attention to philosophical pedagogy. In an unpublished critical review from 2013, I found little good empirical evidence that business ethics or medical ethics instruction has any practical effect on student behavior. I have been following up with some empirical research of my own with several different collaborators. None of it is complete yet, but preliminary results tend to confirm the lack of practical effect, except perhaps when there's the right kind of narrative or emotional engagement. On grounds of armchair plausibility, I tend to favor multi-causal, canceling explanations over the view that philosophical reflection is simply inert (contra Jon Haidt); thus I'm inclined to explore how backfire effects might on average tend to cancel positive effects. It was a post on the possible backfire effects of teaching ethics that prompted Rickless's comment.

Rickless's Objection:
(shared with permission, adding lineation and emphasis for clarity)

Rickless: And I’ll be honest, Eric, all this stuff about how unethical ethicists are, and how counterproductive their courses might be, really bothers me. It’s not that I think that ethics courses can’t be improved or that all ethicists are wonderful people. But please understand that the takeaway from this kind of research and speculation, as it will likely be processed by journalists and others who may well pick up and run with it, will be that philosophers are shits whose courses turn their students into shits. And this may lead to the defunding of philosophy, the removal of ethics courses from business school, and, to my mind, a host of other consequences that are almost certainly far worse than the ills that you are looking to prevent.

Schwitzgebel: Samuel, I understand that concern. You might be right about the effects. However, I also think that if it is correct that ethics classes as standardly taught have little of the positive effect that some administrators and students hope for from them, we as a society should know that. It should be explored in a rigorous way. On the possibly bright side, a new dimension of my research is starting to examine conditions under which teaching does have a positive measurable effect on real-world behavior. I am hopeful that understanding that better will lead us to teach better.

Rickless: In theory, what you say about knowing that courses have little or no positive effect makes sense. But in practice, I have the following concerns.

First, no set of studies could possibly measure all the positive and negative effects of teaching ethics this way or that way. You just can’t control all the potentially relevant variables, in part because you don’t know what all the potentially relevant variables are, in part because you can’t fix all the parameters with only one parameter allowed to vary.

Second, you need to be thinking very seriously about whether your own motives (particularly motives related to bursting bubbles and countering conventional wisdom) are playing a role in your research, because those motives can have unseen effects on the way that research is conducted, as well as the conclusions drawn from it. I am not imputing bad motives to you. Far from it, and quite the opposite. But I think that all researchers, myself included, want their research to be striking and interesting, sometimes surprising.

Third, the tendency of researchers is to draw conclusions that go beyond the actual evidence.

Fourth, the combination of all these factors leads to conclusions that have a significant likelihood of being mistaken.

Fifth, those conclusions will likely be taken much more seriously by the powers-that-be than by the researchers themselves. All the qualifiers inserted by researchers are usually removed by journalists and administrators.

Sixth, the consequences on the profession if negative results are taken seriously by persons in positions of power will be dire.

Under the circumstances, it seems to me that research that is designed to reveal negative facts about the way things are taught had better be airtight before being publicized. The problem is that there is no such research. This doesn’t mean that there is no answer to problems of ineffective teaching. But that is an issue for another day.

My Reply:

On the issue of motives: Of course it is fun to have striking research! Given my general skepticism about self-knowledge, including of motives, I won't attempt self-diagnosis. However, I will say that except for recent studies that are not yet complete, I have published every empirical study I've done on this topic, with no file-drawered results. I am not selecting only the striking material for publication. Also, in my recent pedagogy research I am collaborating with other researchers who very much hope for positive results.

On the likelihood of being mistaken: I acknowledge that any one study is likely to be mistaken. However, my results are pretty consistent across a wide variety of methods and behavior types, including some issues specifically chosen with the thought that they might show ethicists in a good light (the charity and vegetarianism measures in Schwitzgebel and Rust 2014). I think this adds to credibility, though it would be better if other researchers with different methods and theoretical perspectives attempted to confirm or disconfirm our findings. There is currently one replication attempt ongoing among German-language philosophers, so we will see how that plays out!

On whether the powers-that-be will take the conclusions more seriously than the researchers: I interpret Rickless here as meaning that they will tend to remove the caveats and go for the sexy headline. I do think that is possible. One potentially alarming fact from this point of view is that my most-cited and seemingly best-known study is the only study where I found ethicists seeming to behave worse than the comparison groups: the study of missing library books. However, it was also my first published study on the topic, so I don't know to what extent the extra attention is a primacy effect.

On possibly dire consequences: The most likely path for dire consequences seems to me to be this: Part of the administrative justification for requiring ethics classes might be the implicit expectation that university-level ethics instruction positively influences moral behavior. If this expectation is removed, so too is part of the administrative justification for ethics instruction.

Rickless's conclusion appears to be that no empirical research on this topic, with negative or null results, should be published unless it is "airtight", and that it is practically impossible for such research to be airtight. From this I infer that Rickless thinks either that (a.) only positive results should be published, while negative or null results remain unpublished because inevitably not airtight, or that (b.) no studies of this sort should be published at all, whether positive, negative, or null.

Rickless's argument has merit, and I see the path to this conclusion. Certainly there is a risk to the discipline in publishing negative or null results, and one ought to be careful.

However, both (a) and (b) seem to be bad policy.

On (a): To think that only positive results should be published (or more moderately that we should have a much higher bar for negative or null results than for positive ones) runs contrary to the standards of open science that have recently received so much attention in the social psychology replication crisis. In the long run it is probably contrary to the interests of science, philosophy, and society as a whole for us to pursue a policy that will create an illusory disproportion of positive research.

That said, there is a much more moderate strand of (a) that I could endorse: Being cautious and sober about one's research, rather than yielding to the temptation to inflate dubious, sexy results for the sake of publicity. I hope that in my own work I generally meet this standard, and I would recommend that same standard for both positive and negative or null research.

On (b): It seems at least as undesirable to discourage all empirical research on these topics. Don't we want to know the relationship between philosophical moral reflection and real-world moral behavior? Even if you think that studying the behavior of professional ethicists in particular is unilluminating, surely studying the effects of philosophical pedagogy is worthwhile. We should want to know what sorts of effects our courses have on the students who take them and under what conditions -- especially if part of the administrative justification for requiring ethics courses is the assumption that they do have a practical effect. To reject the whole enterprise of empirically researching the effects of studying philosophy because there's a risk that some studies will show that studying philosophy has little practical impact on real-world choices -- that seems radically antiscientific.

Rickless raises legitimate worries. I think the best practical response is more research, by more research groups, with open sharing of results, and open discussions of the issue by people working from a wide variety of perspectives. In the long run, I hope that some of my null results can lay the groundwork for a fuller understanding of the moral psychology of philosophy. Understanding the range of conditions under which philosophical moral reflection does and does not have practical effects on real-world behavior should ultimately empower rather than disempower philosophy as a discipline.

[image source]

Wednesday, December 23, 2015

A Response to Critiques of Cushman's and My Work on Philosophers' Susceptibility to Order Effects

The order in which moral dilemmas are presented matters to people's judgments and can substantially influence later judgments about abstract moral principles. This is true even among professional ethicists with PhD's in philosophy. In 2012 and 2015, Fiery Cushman and I published empirical evidence supporting these claims. We invite a metaphilosophical conclusion: If even professional philosophers' expert judgments are easily swayed by order of presentation, then such judgments might not be stable enough to serve as secure grounds for philosophical theorizing.

Synthese has recently published two critiques of the literature on order effects in philosophy, which address Fiery's and my work (HT Wesley Buckwalter). Both critiques make valuable points. However both also admit of some clear replies.

To fix ideas, consider two versions of the famous Trolley Problem:

Push: A runaway boxcar is headed toward five people it will kill if nothing is done. Jane can stop the boxcar by pushing a hiker with a heavy backpack in front of the boxcar, killing him but saving the five.

Switch: A runaway boxcar is headed toward five people it will kill if nothing is done. Vicki can stop the boxcar by flipping a switch to divert it to a sidetrack where it will kill one person instead of the five.

Fiery and I presented Push-type and Switch-type scenarios (fleshed with a bit more detail) to professional philosophers and two comparison groups of non-philosophers. We found that when professional philosophers saw a Push-type scenario before a Switch-type scenario, 73% rated the two scenarios equivalently on a 7-point scale. Then later in the questionnaire when asked about the Doctrine of the Double Effect -- a moral principle often interpreted implying that Push-type cases are morally worse than Switch-type cases -- only a minority, 46%, endorsed that principle. In contrast, among philosophers who saw Switch before Push only 54% rated the two scenarios equivalently, and then later a majority, 62%, endorsed the Doctrine of the Double Effect. Endorsement of the principle thus seemed to shift, post-hoc, to rationalize philosophers' order-manipulated judgments about the scenarios.

We found similar effects for Action-Omission, Moral Luck, and "Asian disease" type cases (though not consistently for every measure across the board). Philosophers with PhDs and self-reported competence or specialization in ethics showed no smaller effects than other philosophers or than comparison groups of non-philosophers -- and in fact trended slightly (non-significantly) toward showing larger order effects.

In general, we found pretty substantial effect sizes, suggesting substantial instability of judgment even in philosophical respondents' areas of expertise. Hence the metaphilosophical worry.

----------------------------------------------

Critique by Zachary Horne and Jonathan Livengood.

Horne and Livengood make three main points about the literature on order effects in philosophy:

(A.) First, they helpfully distinguish between what they call "updating effects" and "genuine ordering effects". Genuine ordering effects, in their terminology, are effects measured only after all the stimuli have been presented. "Updating effects" are measures taken along the way, and might well reflect participants' learning. There is of course nothing irrational in judging Scenario B differently as a result of seeing Scenario A because one learned something by seeing Scenario A. Most philosophical research on order effects, they note, takes the measures along the way -- and thus might be measuring learning rather than true order effects.

(B.) Second, they point out that perceptual judgments also show order effects. Thus, if we are to reject any type of evidence that shows order effects, then we must reject perceptual evidence too, which would lead to radical skepticism.

(C.) Third, they point out that order can sometimes reasonably make a difference to the evaluation of evidence. For example, a smile followed by a frown, on the same person's face, is a different type of evidence than a frown followed by a smile.

On (A): I find the labels tendentious (since if we know there isn't learning-type updating going on, what we might want to call "genuine order effects" can plausibly be measured mid-stream), however it probably is correct that most studies do not sufficiently rule out the possibility of learning or updating in the course of the experiment, if they have novice participants and take the measurements after each scenario rather than after both scenarios. However, since our participants were experts, we think it unlikely that a significant number learned anything in the process of our brief experiment that would rationally justify shifting their judgment about the equivalency or non-equivalency of Push and Switch. And as Horne and Livengood note, our measure of endorsement of the Doctrine of the Double Effect is a measurement of a "genuine ordering effect" even by their own lights.

On (B): Yes, of course it would be silly to reject all means of learning that are subject to any order effects! The epistemic sting, as they note, depends not on the mere existence of an order effect in one case, but on how large and how prevalent the order effects are. This is an open empirical question. But the limited empirical evidence that exists suggests that order effects are substantial and prevalent in moral dilemma cases. So far, we have found order effects in all of the scenario types we've tried, with about a 10-20% shift in opinion on the moral equivalency of our scenario pairs and in preference for the risky option in the "Asian disease" cases.

On (C): It's interesting to consider cases in which earlier evidence rightly colors our reaction to later evidence, but trolley problems presented to disciplinary experts seems a different kind of case.

Finally, Horne and Livengood suggest that exposure to a pair of dilemmas in our study is unlikely to have a long-lasting impact on professional philosophers' beliefs. I agree. They continue, "But if there is no long-lasting impact, then we think the effect is unlikely to matter to actual philosophical practice outside of the laboratory" (p. 17). I don't think this follows. Fiery's and my view is not that philosophers' opinions are permanently influenced by the order in which the scenarios are presented on any single occasion, but rather that their opinions are unstable -- possibly influenced one direction on one occasion, in another direction on another occasion. This instability is what drives the metaphilosophical worry.

----------------------------------------------

Critique by Regina Rini:

Rini -- a recent guest blogger here at the Splintered Mind -- looks only at our 2012 study. (Our 2015 study wasn't published until after her paper was in press.) She finds it plausible that if professional philosophers were already familiar with these cases they would not exhibit order effects of the sort Fiery and I find. She suggests that perhaps respondents were not previously familiar with the cases -- or at least not familiar in the right sort of way. She calls this the "familiarity problem" and offers four possible explanations:

(1.) The respondents were not really experts. She wonders if our participants, recruited through the internet, really had the degrees they claimed to have.

(2.) The respondents didn't carefully attend to our scenarios. Maybe they breezed through them so quickly that they failed to notice relevant features.

(3.) The respondents might not have familiar responses to these types of scenarios. Perhaps they have so far refrained from forming judgments on such cases and principles.

(4.) The respondents might not have diachronically stable familiar responses. This is the explanation Fiery and I favor. However, Rini helpfully points out that as long as philosophers are aware that their responses are not diachronically stable, the metaphilosophical threat is reduced: Presumably philosophers who are aware that their responses are not stable would be reluctant to ground their theorizing on those responses.

On (1): I am not aware of a general problem in the survey literature of respondents' frequently misreporting their educational status -- though certainly a bit of misreporting is possible. One specific piece of evidence against this possibility in our own study is that we recruited philosophers mostly by asking department chairs to forward a recruitment email to faculty and graduate students in their departments. Most of our "philosopher" participants took the survey within just a few days of these emails.

On (2): The median response time on the first scenario was 40 seconds, on the second scenario was 34 seconds. While these are not huge response times, if you stop to count out 34 seconds now, you'll probably notice that it's a reasonable amount of time for a thoughtful response to a brief scenario.

On (3) and (4): These are potentially quite serious issues, and in fact our follow-up study in 2015 was designed specifically to address them, after we saw an early version of Rini's critique. In our 2015 study we specifically asked participants if they were previously familiar with the scenarios. We also asked whether they regarded themselves as "having had a stable opinion" about the issues before participating in the experiment, and whether they regarded themselves as experts on those very issues. We also added a "reflection" condition to help address concern (2). In the reflection condition we asked participants to reflect carefully before responding and enforced a minimum 15-second delay between when participants reported having finished reading the scenario and when their response options appeared.

We did not find that self-reported familiarity or stability reduced the size of order effects in two different types of scenario pairs (trolley problems and risky-choice "Asian disease"-type problems), nor did we find reduced order effects in the reflection condition compared to a normal control condition without special instructions to reflect.

For example, percentage rating the Push and Switch scenarios equivalently:

Thus, I am inclined to think that Rini's fourth suggestion is the most plausible -- that participants do not have diachronically stable familiar responses, despite high levels of expertise. But since those who report having stable responses were no less subject to order effects than were those who reported not having stable responses, self-knowledge of stability appears to be largely absent. Despite Rini's interesting suggestion that instability is metaphilosophically non-threatening if people are aware of it, Fiery's and my results suggest that we should not hasten to that comfort.

----------------------------------------------

Both Horne and Livengood and Rini emphasize that we only have very limited evidence about order effects on professional philosophers' judgments. I agree! Fiery's and my two studies are hardly decisive. Convergent evidence from several different labs would be necessary before drawing any confident conclusions, especially if those conclusions are at variance with what one feels one knows from personal experience. Rini also makes positive suggestions for follow-up experimental work that might be done, which I am inclined to support. Both critiques raise important methodological concerns that ought to help shape and direct future work on this topic.

Wednesday, October 15, 2014

Professional Philosophers’ Susceptibility to Order Effects and Framing Effects in Evaluating Moral Dilemmas

Fiery Cushman and I have a new paper in draft, exploring the question of whether professional philosophers' judgments about moral dilemmas are less influenced than non-philosophers' by factors such as order of presentation and phrasing differences.

We recruited hundreds of academic participants with graduate degrees in philosophy and a comparison group of academics with graduate degrees in other fields. We gave them two "trolley problems" and two "Asian disease"-type framing effect cases.

We presented the trolley problems either with a Switch case first (the protagonist saves five people by Switching a runaway trolley onto a side track where it kills one), followed by a Push or Drop case (saving five by Pushing one person into the trolley's path or by Dropping him into its path); or Push/Drop first, followed by Switch. For each scenario, participants rated the protagonists' choice to kill the one person to save the five others, using a 7-point scale from "extremely morally good" to "extremely morally bad".

Our previous research suggests that non-philosophers are much more likely to judge the Push case and the Switch case equivalently when Push is presented first than when Switch is presented first. On some views of philosophical expertise, philosophers' judgments about these cases should be less dependent on order of presentation than are non-philosophers' judgments. In other words, philosophers, due to their familiarity with scenarios of this type and their expertise in applying moral principles to them, should have more stable opinions, less influenced by order of presentation. We wanted to see if philosophers with prior familiarity with the cases, or self-reported expertise in the area, or self-reported stability of opinion, would show smaller order effects. We also wanted to see if we could reduce order effects by enforcing a delay before responding during which we encouraged participants to reflect carefully on different versions of the scenario and different ways of phrasing the scenarios.

We were unable to find any level of expertise at which the order effects were detectably reduced. Nor did adding a reflection condition appear to reduce the order effects. This figure shows the rates at which Switch was rated as morally equivalent to either Drop or Push on the 7-point scale:

[click to enlarge]

The order effect, as indicated by the differences in height between the black and the gray bars, is basically the same for philosophers and the non-philosophers, even in the "reflection" condition.

This next figure breaks the results down by degree of self-reported expertise among philosopher respondents:

[click to enlarge]

Note that at no level of expertise does the order effect appear to be reduced: not among philosophers reporting being professors with a specialization in ethics, nor among philosophers reporting having a "stable opinion" about trolley problems of this sort. If anything, the trend appears to be toward larger order effects with increasing expertise.

Some of the most famous results in psychology are the Tversky-Kahneman "loss aversion" framing effects. Participants are asked to imagine that an unusual disease will kill 600 people if nothing is done and then given a choice between two programs: On Program A, 200 people will be saved. On Program B, there's a one-third probability that 600 people will be saved and a two-thirds probability that no people will be saved. When the decision is framed this way, in terms of the number "saved", most people favor the non-risky Program A. When what are (purportedly) the exact same options are presented in terms of how many will die (400 will die vs. one-third probability that none will die and two-thirds probability that 600 will die), respondents tend to favor the risky Program B.

The results:
Percentage of philosopher respondents recommending the risky Program B, by framing, level of expertise, and order of presentation:

[click to expand]

As is evident from the figure, our philosopher respondents showed very large framing effects (similar to those of our comparison group and similar to the effect sizes seen in other studies with non-expert populations) -- again up to very high levels of expertise, including self-reported expertise on framing effects and self-reported stability of opinion about framing effects. To see this, look just at the black bars above, ignoring the gray bars.

Philosophers also showed large order effects when they were presented two slightly different framing-effect scenarios, either die-frame followed by save-frame or save-frame followed by die-frame. To see this, compare the adjacent pairs of black and gray bars above.

Full manuscript in draft here. Comments welcome!

Wednesday, December 11, 2013

How Subtly Do Philosophers Analyze Moral Dilemmas?

You know the trolley problems. A runaway train trolley will kill five people ahead on the tracks if nothing is done. But -- yay! -- you can intervene and save those five people! There's a catch, though: your intervention will cost one person's life. Should you intervene? Both philosophers' and non-philosophers' judgments vary depending on the details of the case. One interesting question is how sensitive philosophers and non-philosophers are to details that might be morally relevant (as opposed to presumably irrelevant distracting features like order of presentation or the point-of-view used in expressing the scenario).

Consider, then, these four variants of the trolley dilemma:

Switch: You can flip a switch to divert the trolley onto a dead-end side-track where it will kill one person instead of the five.

Loop: You can flip a switch to divert the trolley into a side-track that loops back around to the main track. It will kill one person on the side track, stopping on his body. If his body weren't there to block it, though, the trolley would have continued through the loop and killed the five.

Drop: There is a hiker with a heavy backpack on a footbridge above the trolley tracks. You can flip a switch which will drop him through a trap door and onto the tracks in front of the runaway trolley. The trolley will kill him, stopping on his body, saving the five.

Push: Same as Drop, except that you are on the footbridge standing next to the hiker and the only way to intervene is to push the hiker off the bridge into the path of the trolley. (Your own body is not heavy enough to stop the trolley.)

Sure, all of this is pretty artificial and silly. But orthodox opinion is that it's permissible to flip the switch in Switch but impermissible to push the hiker in Push; and it's interesting to think about whether that is correct, and if so why.

Fiery Cushman and I decided to compare philosophers' and non-philosophers' responses to such cases, to see if philosophers show evidence of different or more sophisticated thinking about them. We presented both trolley-type setups like this and also similarly structured scenarios involving a motorboat, a hospital, and a burning building (for our full list of stimuli see Q14-Q17 here.)

In our published article on this, we found that philosophers were just as subject to order effects in evaluating such scenarios as were non-philosophers. But we focused mostly on Switch vs. Push -- and also some moral luck and action/omission cases -- and we didn't have space to really explore Loop and Drop.

About 270 philosophers (with master's degree or more) and about 670 non-philosophers (with master's degree or more) rated paragraph-length versions of these scenarios, presented in random order, on a 7-point scale from 1 (extremely morally good) through 7 (extremely morally bad; the midpoint at 4 was marked "neither good nor bad"). Overall, all the scenarios were rated similarly and near the midpoint of the scale (from a mean of 4.0 for Switch to 4.4 for Push [paired t = 5.8, p < .001]), and philosophers and non-philosophers mean ratings were very similar.

Perhaps more interesting than mean ratings, though, are equivalency ratings: How likely were respondents to rate scenario pairs equivalently? The Loop case is subtly different from the Switch case: Arguably, in Loop but not Switch, the man's death is a means or cause of saving the five, as opposed to a merely foreseen side effect of an action that saves the five. Might philosophers care about this subtle difference more than non-philosophers? Likewise, the Drop case is different from the Push case, in that Push but not Drop requires proximity and physical contact. If that difference in physical contact is morally irrelevant, might philosophers be more likely to appreciate that fact and rate the scenarios equivalently?

In fact, the majority of participants rated all the scenarios exactly the same -- and philosophers were no less likely to do so than non-philosophers: 63% of philosophers gave identical ratings to all four scenarios, vs. 58% of non-philosophers (Z = 1.2, p = .23).

I find this somewhat odd. To me, it seems pretty flat-footed a form of consequentialism that says that Push is not morally worse than Switch. But I find that my judgment on the matter swims around a bit, so maybe I'm wrong. In any case, it's interesting to see both philosophers and non-philosophers seeming to reject the standard orthodox view, and at very similar rates.

How about Switch vs. Loop? Again, we found no difference in equivalency ratings between philosophers and non-philosophers: 83% of both groups rated the scenarios equivalently (Z = 0.0, p = .98).

However, philosophers were more likely than non-philosophers to rate Push and Drop equivalently: 83% of philosophers did, vs. 73% of non-philosophers (Z = 3.4, p = .001; 87% vs. 77% if we exclude participants who rated Drop worse than Push).

Here's another interesting result. Near the end of the study we asked whether it was worse to kill someone as a means of saving others than to kill someone as a side-effect of saving others -- one way of setting up the famous Doctrine of the Double Effect, which is often evoked to defend the view that Push is worse than Switch (in Push, the one person's death is arguably the means of saving the other five, in Switch the death is only a foreseen side-effect of the action that saves the five). Loop is interesting in part because although superficially similar to Switch, if the one person's death is the means of saving the five, then maybe the case is more morally similar to Push than to Switch (see Otsuka 2008). However, only 18% of the philosophers who said it was worse to kill as a means of saving others rated Loop worse than Switch.

Thursday, October 03, 2013

Second-Person vs. Third-Person Presentation of Moral Dilemmas

You know the trolley problems, of course. An out-of-control trolley is headed toward five people it will kill if nothing is done. You can flip a switch and send it to a side track where it will kill one different person instead. Should you flip the switch? What if, instead of flipping a switch, the only way to save the five is to push someone into the path of the trolley, killing that one person?

In evaluating this scenario, does it matter if the person standing near the switch with the life-and-death decision to make is "John" as opposed to "you"? Nadelhoffer & Feltz presented the switch version of the trolley problem to undergraduates from Florida State University. Forty-three saw the problem with "you" as the actor; 65% of the them said it was permissible to throw the switch. Forty-two saw the problem with "John" as the actor; 90% of them said it was permissible to throw the switch, a statistically significant difference.

Tobia, Buckwalter & Stich followed up, presenting a famous moral dilemma from Bernard Williams in which someone can save a group of innocent villagers from a gunman by choosing personally to shoot one of the villagers. Forty undergraduates were presented this scenario. When "you" were given the chance to shoot one villager to save the rest, 19% of said it was morally obligatory to do so; when "Jim" was given the chance, 53% said it was obligatory (again statistically significant).

However, Tobia and colleagues also gave the scenario to 62 professional philosophers and found the opposite effect: 9% of philosophers found it obligatory for "Jim" and 36% found it obligatory for "you". They also presented a trolley-switching case to 49 professional philosophers. Again, the effect was in the opposite direction from that observed among undergraduates: 89% of philosophers said it was permissible to flip the switch in the second-person condition vs. 64% in the third-person condition.

Fiery Cushman and I have some unpublished data on this that I thought I'd throw into the mix, since our results are a bit different from those of Tobia and colleagues. We collected these data for our 2012 paper on order effects in philosophers' and non-philosophers' judgments about moral scenarios. Most of the scenarios were presented third-person, but as we mention in the published paper, some scenarios also had second-person variants. We didn't find large effects, and the paper was already very complicated, so we didn't detail the second-person/third-person differences.

In that experiment, we had four scenarios that differed 2nd person vs. 3rd person. However, they differed not in whether the actor was described as "you", but rather in whether the victim was.

One scenario was a version of Williams' hostage scenario. "Nancy" and other villagers are captured by a warlord. Nancy is given the choice of shooting "you" (2nd person variant) or "a fellow hostage" (3rd person variant) to save the captured villagers. Respondents rated Nancy's "shooting you" or "shooting a fellow hostage" on a 7-point scale from "extremely morally good" (1) through "extremely morally bad" (7), with "morally neutral" in the middle (4). We had three groups of respondents: 324 professional philosophers (MA or PhD in philosophy, mostly recruited via email to Leiter-ranked philosophy departments), 753 non-philosopher academics (Master's or PhD not in philosophy, mostly recruited via email to comparison departments at the same universities), and 1389 non-academics (a convenience sample of others who happened upon the test site).

We found non-philosophers a bit more likely to rate Nancy's shooting one to save the others toward the "morally good" side of the scale if the victim was "you", but philosophers showed only a small, non-significant trend on our 7-point scale (using t-tests):

Non-academics: 3.6 (2nd person victim) vs. 4.1 (3rd person victim) (p < .001).
Academic non-philosophers: 4.1 vs. 4.5 (p = .001).
Philosophers: 3.9 vs. 4.0 (p = .60).

We found similar results in a scenario in which a captain of a military submarine can shoot "you" (2nd person) or shoot another "crew member" (3rd person) to save the vessel:

Non-academics: 2.7 vs. 3.1 (p < .001).
Academic non-philosophers: 2.9 vs. 3.2 (p = .050).
Philosophers: 2.9 vs. 2.8 (p = .60).

We also presented a scenario pair in which you and other passengers have fled a sinking ship. You will drown without a life vest. In one version, someone snatches a vest away from you. In another version, someone declines to put himself at risk by giving you his vest. The results:

Snatching the vest:

Non-academics: 5.6 (2nd person victim) vs. 5.8 (3rd person victim) (p = .052).
Academic non-philosophers: 5.7 vs. 6.0 (p = .002).
Philosophers: 5.8 vs. 5.7 (p = .58).
Not giving up the vest:
Non-academics: 4.8 vs. 4.7 (p = .12).
Academic non-philosophers: 4.7 vs. 4.8 (p = .82).
Philosophers: 4.6 vs. 4.4 (p = .26)
In sum: The effects were small and inconsistent, but there was a general tendency for non-philosophers to rate harm to themselves as morally better than harm to other people -- a tendency not evident among philosopher respondents.

Personally, I'm not inclined to make much of this, since I don't think people are generally in fact more morally lenient in judging harms to themselves than in judging harms to other people. My guess is that these results reflect a small "impression management" or socially desirable responding bias among the non-philosophers that we don't see among the philosophers, who might be more inclined to hear "you" pretty abstractly and impersonally when presented with familiar scenarios of this type.

In an earlier unpublished version of this study, we also tried varying 2nd and 3rd person presentation of the actor who is faced with the choice, including in standard trolley type and hostage type cases of the sort described in Nadelhoffer & Feltz and Tobia, Buckwalter & Stich. Due to a programming error, we couldn't use the data and can't fully interpret it, but our general finding was that the effect was very subtle, and mostly non-detectable even with hundreds of participants (394 philosophers and even more in the other groups). That's why we shifted to trying out 2nd vs. 3rd person variation in the victim role -- maybe it would be a larger effect, we thought.

So, for example, merging the push and switch versions of the trolley scenarios, we found the following ratings on our 7-point scale:

Non-academics: 3.8 (2nd person actor) vs. 3.9 (3rd person actor) (p = .27).
Academic non-philosophers: 3.9 vs. 3.8 (p = .19).
Philosophers: 3.6 vs. 3.9 (p = .07)
And in a shoot-the-villager type scenario, the results were:
Non-academics: 4.3 vs. 4.3 (p = .46).
Academic non-philosophers: 4.6 vs. 4.5 (p = .18).
Philosophers: 3.9 vs. 4.2 (p = .14)
However, in the life vest cases we did seem to see a small effect.

Snatching the vest:

Non-academics: 6.0 vs. 5.9 (p = .053).
Academic non-philosophers: 6.1 vs. 5.9 (p = .03).
Philosophers: 5.8 vs. 5.8 (p = .94)
Not giving up the vest:
Non-academics: 4.9 vs. 4.7 (p = .008).
Academic non-philosophers: 4.8 vs. 4.7 (p = .08).
Philosophers: 4.7 vs. 4.4 (p = .12)
Thus, overall, we found some confirmation of the tendency for non-philosophers to rate actions a little more harshly in 2nd person than in 3rd person presentations, but the effect was small and inconsistent; and we did not find a tendency for philosophers to go in the opposite direction.

We're not sure why we found much smaller effects here than have others. Among the possibilities: Our scenarios were worded somewhat differently. Our response scale (the 1-7 scale from "extremely morally good" to "extremely morally bad") was set up differently. Our participants were recruited differently.