Showing posts with label psychological methods. Show all posts
Showing posts with label psychological methods. Show all posts

Saturday, February 22, 2025

New in Print: The Necessity of Construct and External Validity for Deductive Causal Inference

with Kevin Esterling and David Brady

In deductive causal generalization, internal validity, external validity, and construct validity are *equal* legs of a stool. Internal validity alone is literally meaningless without the other two.

https://www.degruyter.com/document/doi/10.1515/jci-2024-0002/html

Abstract:

The Credibility Revolution advances internally valid research designs intended to identify causal effects from quantitative data. The ensuing emphasis on internal validity, however, has enabled a neglect of construct and external validity. We show that ignoring construct and external validity within identification strategies undermines the Credibility Revolution’s own goal of understanding causality deductively. Without assumptions regarding construct validity, one cannot accurately label the cause or outcome. Without assumptions regarding external validity, one cannot label the conditions enabling the cause to have an effect. If any of the assumptions regarding internal, construct, and external validity are missing, the claim is not deductively supported. The critical role of theoretical and substantive knowledge in deductive causal inference is illuminated by making such assumptions explicit. This article critically reviews approaches to identification in causal inference while developing a framework called causal specification. Causal specification augments existing identification strategies to enable and justify deductive, generalized claims about causes and effects. In the process, we review a variety of developments in the philosophy of science and causality and interdisciplinary social science methodology.

Tuesday, February 04, 2025

A Taxonomy of Validity: Eeek!

There comes a time in everyone's life when their 18-year-old daughter, taking their first psychology class, asks, "Parental-figure-of-mine, what is 'validity'?"

For me that time came last week. Eeek!

Psychologists and social scientists use the term all the time, with a dazzling array of modifiers: internal validity, construct validity, external validity, convergent validity, predictive validity, discriminant validity, face validity, criterion validity.... But ask those same social scientists what validity is exactly, and how all of these notions relate to each other, and most will stumble.

As it happens, I was well positioned to address my daughter's question. I have a new paper, on "validity" in causal inference, forthcoming in the Journal of Causal Inference with social scientists Kevin Esterling and David Brady. This paper has been in progress since (again, eeek!) 2018. In previous posts I've addressed whether validity (in social science usage) is better understood as a property of inferences or as a property of claims (I argue the latter), and the intimate relationship of internal validity, external validity, and construct validity in causal inference.

Today, I'll attempt a brief, theoretically-motivated taxonomy of the better-known types of validity. My aim is more descriptive than argumentative: I'll just outline how I think various "validities" hang together, and maybe some readers will find it to be an attractive and helpful picture.

I start with the assumption that validity is a feature of claims, not of inferences. Philosophers typically describe validity as a property of inferences. Social scientists are all over the map, and even prominent ones are sloppy in their usage. But it best organizes our thinking to address claims primarily and treat inferences as secondary.

I will say that a general causal claim that "A causes B in conditions C" is valid if and only if A does in fact cause B in conditions C. (Compare disquotational theories of truth in philosophy.) Consider for example the causal claim: Enforcement threats on reminder postcards (A) cause increased juror turnout (B) in the 21st-century United States (C).

This statement can be divided into four parts, each of which permits a distinctive type of validity failure:

(i.) A

(ii.) causes

(iii.) B

(iv.) in conditions C.

The four possible failures generate the core taxonomic structure.

Construct validity of the cause: Something might cause B in conditions C, but that something might not be A. A causal generalization has construct validity of the cause if the claim accurately specifies that A in particular (and not, for example, some other related thing) causes B in conditions C. Example of a failure of construct validity of the cause: Increased juror turnout among people who receive postcards might not be due to enforcement threats in particular but simply to being reminded of one's civic duty.

Construct validity of the effect: A might cause something in conditions C, but what it causes might not be B. A causal generalization has construct validity of the effect if the effect of A is accurately specified. A causes specifically B (and not, for example, some other related thing) in conditions C. Example of a failure of construct validity of the effect: Enforcement threats might increase the rates at which jurors who don't show up register a valid excuse without actually increasing turnout rates.

Generalizing: Construct validity is present in a causal generalization when the cause and effect are accurately specified.

External validity: A might cause B, but the conditions might not be correctly specified. A causal generalization has external validity if the claim accurately specifies the range of conditions in which it holds. Example of a failure: Enforcement messages might increase juror turnout not in the U.S. in general but only in low-income neighborhoods. Perfect external validity is probably an unattainable ideal for complex social and psychological processes, since the conditions in which causal generalizations hold will be complex and various.

Note on external validity: Common usage often holds that a claim is externally valid only if it holds across a wide range of contexts or conditions. However, this way of thinking unhelpfully denigrates perfectly accurate causal generalizations as "invalid" if they only hold, and are claimed only to hold, across a narrow range of conditions. Transportability is a better concept for characterizing breadth of applicability. An externally valid causal generalization that is accurately claimed to hold across only a narrow range of contexts is not transportable to those other contexts, but there is no inaccuracy or factual error in the statement "A causes B in conditions C" of the sort required for failure of validity. After all A does cause B in conditions C, just as claimed. So validity in the overarching sense described above is present.

Internal validity: A might be related to B in conditions C, but the relation might not be the directional causal relationship claimed. A causal generalization is internally valid if there is a cause-effect relationship of the type claimed (even if the cause, the effect, and/or the conditions are not accurately specified). Example of a failure: There's a common cause of both A and B, which are not directly causally related. Maybe having a stable address causes potential jurors both to be more likely to be sent the postcards and to be more likely to turn out.

Other types of validity can be understood within the general spirit of this framework.

Convergent validity: Present when two causes claimed to have the same effect in fact have the same effect. In common use, the causes are measures, for example two different measures of extraversion. In this case, A1 (application of the first measure) and A2 (application of the second measure) are claimed to have a common effect B (same normalized extraversion score) in a set of conditions often left unspecified. Convergent validity is present if that claim is true (or to the degree it is true).

Discriminant validity: Present when two causes claimed to have different effects in fact have different effects. A1 is claimed to cause B, and A2 is claimed not to cause B (in a set of conditions that is often left unspecified), and discriminant validity is present when that claim is true (or to the degree it is true). In practice, discriminant validity is often supported by observation of low correlations in appropriately controlled conditions. If A1 and A2 are psychological or social measures (e.g., personality measures of extraversion and openness), then a high correlation between the scores would suggest that there is some common psychological feature both measures are tracking, contrary to the ideal of general discriminant validity.

Predictive validity: Present when A is a common cause of B1 and B2, where B1 is typically the outcome of a measure and B2 is typically an event of practical import conceptually related but not closely physically related to B1. For example, application of a purported measure of recidivism (in this case, application of the measure isn't A but rather an intermediate event A1) among released prisoners has high predictive validity if high scores on the measure (B1) arise from the same cause or set of causes that generate high rates of recidivism (B2).

Note on predictive validity: A simpler characterization of "predictive validity" might be simply that B1 accurately predicts B2, but this isn't the most useful way to conceptualize the issue if the prediction is correct in virtue of B1 causing B2 rather than operating by a common cause. If my wife reliably picks me up from work when I ask, my asking (B1) predicts her picking me up (B2), but my asking does not have "predictive validity" in the intended measurement sense. A better term for this relationship would be casual power.

Face validity: Present when it is intuitively or theoretically plausible that A causes B in conditions C. Notably, face validity needn't require that A in fact causes B in conditions C.

Ecological validity: A type of external validity that emphasizes the importance of generalizing correctly over real-world settings (as opposed to laboratory settings or other artificial settings).

Content validity: A type of construct validity focused on whether the content of a complex measure accurately reflects all aspects of the target measured.

Criterion validity: Present when a measure or intervention satisfies some prespecified criterion of success, regardless of whether the measure or intervention in fact measures what it purports to measure.

Finally, two types of validity where "validity" is a property of the inference rather than in terms of the truth of some part of a causal claim:

Statistical conclusion validity: Present when statistics are appropriately used, regardless of whether A in fact causes B in conditions C.

Logical validity: Present when the conclusion of an argument can't be false if its premises are true.

Tuesday, November 07, 2023

The Prospects and Challenges of Measuring Morality, or: On the Possibility or Impossibility of a "Moralometer"

Could we ever build a "moralometer" -- that is, an instrument that would accurately measure people's overall morality?  If so, what would it take?

Psychologist Jessie Sun and I explore this question in our new paper in draft: "The Prospects and Challenges of Measuring Morality".

Comments and suggestions on the draft warmly welcomed!

Draft available here:

https://osf.io/preprints/psyarxiv/nhvz9

Abstract:

The scientific study of morality requires measurement tools. But can we measure individual differences in something so seemingly subjective, elusive, and difficult to define? This paper will consider the prospects and challenges—both practical and ethical—of measuring how moral a person is. We outline the conceptual requirements for measuring general morality and argue that it would be difficult to operationalize morality in a way that satisfies these requirements. Even if we were able to surmount these conceptual challenges, self-report, informant report, behavioral, and biological measures each have methodological limitations that would substantially undermine their validity or feasibility. These challenges will make it more difficult to develop valid measures of general morality than other psychological traits. But, even if a general measure of morality is not feasible, it does not follow that moral psychological phenomena cannot or should not be measured at all. Instead, there is more promise in developing measures of specific operationalizations of morality (e.g., commonsense morality), specific manifestations of morality (e.g., specific virtues or behaviors), and other aspects of moral functioning that do not necessarily reflect moral goodness (e.g., moral self-perceptions). Still, it is important to be transparent and intellectually humble about what we can and cannot conclude based on various moral assessments—especially given the potential for misuse or misinterpretation of value-laden, contestable, and imperfect measures. Finally, we outline recommendations and future directions for psychological and philosophical inquiry into the development and use of morality measures.

[Below: a "moral-o-meter" given to me for my birthday a few years ago, by my then-13-year-old daughter]

Thursday, January 12, 2023

Further Methodological Troubles for the Moralometer

[This post draws on ideas developed in collaboration with psychologist Jessie Sun.]

If we want to study morality scientifically, we should want to measure it. Imagine trying to study temperature without a thermometer or weight without scales. Of course indirect measures are possible: We can't put a black hole on a scale, but we can measure how it bends the light that passes nearby and thereby infer its mass.

Last month, I raised a challenge for the possibility of developing a "moralometer" (a device that accurately measure's a person's overall morality). The challenge was this: Any moralometer would need to draw on one or more of four methods: self-report, informant report, behavioral measures, or physiological measures. Each one of these methods has serious shortcomings as a basis for general moral measurement of one's overall moral character.

This month, I raise a different (but partly overlapping) set of challenges, concerning how well we can specify the target we're aiming to measure.

Problems with Flexible Measures

Let's call a measure of overall morality flexible if it invites a respondent to apply their own conception of morality, in a flexible way. The respondent might be the target themselves (in self-report measures of morality) or they might be a peer, colleague, acquaintance, or family member of the target (in informant-report measures of morality). The most flexible measures apply "thin" moral concepts in Bernard Williams' sense -- prompts like "Overall, I am a morally good person" [responding on an agree/disagree scale] or "[the target person] behaves ethically".

While flexible measures avoid excessive rigidity and importing researchers' limited and possibly flawed understandings of morality into the rating procedure, the downsides are obvious if we consider how people with noxious worldviews might rate themselves and others. The notorious Nazi Adolf Eichmann, for example, appeared to have thought highly of his own moral character. Alexander "the Great" was admired for millennia, including as a moral exemplar of personal bravery and spreader of civilization, despite his main contribution being conquest through aggressive warfare, including the mass slaughter and enslavement of at least one civilian population.

I see four complications:

Relativism and Particularism. Metaethical moral relativists hold that different moral standards apply to different people or in different cultures. While I would reject extreme relativist views according to which genocide, for example, doesn't warrant universal condemnation, a moderate version of relativism has merit. Cultures might reasonably differ, for example, on the age of sexual consent, and cultures, subcultures, and social groups might reasonably differ in standards of generosity in sharing resources with neighbors and kin. If so, then flexible moralometers, employed by raters who use locally appropriate standards, will have an advantage over inflexible moralometers which might inappropriately import researchers' different standards. However, even flexible moralometers will fail in the face of relativism if they are employed by raters who employ the wrong moral standards.

According to moral particularism, morality isn't about applying consistent rules or following any specifiable code of behavior. Rather, what's morally good or bad, right or wrong, frequently depends on particular features of specific situations which cannot be fully codified in advance. While this isn't the same as relativism, it presents a similar methodological challenge: The farther the researcher or rater stands from the particular situation of the target, the more likely they are to apply inappropriate standards, since they are likely to be ignorant of relevant details. It seems reasonable to accept at least moderate particularism: The moral quality of telling a lie, stealing $20, or stopping to help a stranger, might often depend on fine details difficult to know from outside the situation.

If the most extreme forms of moral relativism or particularism (or moral skepticism) are true, then no moralometer could possibly work, since there won't be stable truths about people's morality, or the truths will be so complicated or situation dependent as to defy any practical attempt at measurement. Moderate relativism and particularism, if correct, provide reason to favor flexible standards as judged by self-ratings or the ratings of highly knowledgeable peers sensitive to relevant local details; but even in such cases all of the relevant adjustments might not be made.

Incommensurability. Goods are incommensurable if there is no fact of the matter about how they should be weighed against each other. Twenty dollar bills and ten dollar bills are commensurable: Two of the latter are worth exactly one of the former. But it's not clear how to weigh, for example, health against money or family versus career. In ethics, if Steven tells a lie in the morning and performs a kindness in the afternoon, how exactly ought these to be weighed against each other? If Tara is stingy but fair, is her overall moral character better, worse, or the same as that of Nicholle, who is generous but plays favorites? Combining different features of morality into a single overall score invites commensurability problems. Plausibly, there's no single determinately best weighting of different factors.

Again, I favor a moderate view. Probably in many cases there is no single best weighting. However, approximate judgments remain possible. Even if health and money can't be precisely weighed against each other, extreme cases permit straightforward decisions. Most of us would gladly accept a scratch on a finger for the sake of a million dollars and would gladly pay $10 to avoid stage IV cancer.  Similarly, Stalin was morally worse than Martin Luther King, even if Stalin had some virtues and King some vices. Severe sexual harassment of an employee is worse than fibbing to your spouse to get out of washing the dishes.

Moderate incommensurability limits the precision of any possible moralometer. Vices and virtues, and rights and wrongs of different types will be amenable only to rough comparison, not precise determination in a single common coin.

Moral error. If we let raters reach independent judgments about what is morally good or bad, right or wrong, they might simply get it wrong. As mentioned above, Eichmann appears to have thought well of himself, and the evidence suggests that he also regarded other Nazi leaders as morally excellent. Raters will disagree about the importance of purity norms (such as norms against sexual promiscuity), the badness of abortion, and the moral importance, or not, of being vegetarian. Bracketing relativism, then at least some of these raters must be factually mistaken about morality, on one side or another, adding substantial error into their ratings.

The error issue is enormously magnified if ordinary people's moral judgments are systematically mistaken. For example, if the philosophically discoverable moral truth is that the potential impact of your choices on future generations morally far outweighs the impact you have on the people around you (see my critiques of "longtermism" here and here), then the person who is an insufferable jerk to everyone around them but donates $5000 to an effective charity might be in fact far morally better than a personally kind and helpful person who donates nothing to charity -- but informants' ratings might very well suggest the reverse. Similar remarks would apply to any moral theory that is sharply at odds with commonsense moral intuition.

Evaluative bias. People are, of course, typically biased in their own favor. Most people (not all!) are reluctant to think of themselves as morally below average, as unkind, unfair, or callous, even if they in fact are. Social desirability bias is the well-known phenomenon that survey respondents will tend to respond to questions in a manner that presents them in a good light. Ratings of friends, family, and peers will also tend to be positively biased: People tend to view their friends and peers positively, and even when not they might be reluctant to "tell on" them to researchers. If the size of evaluative bias were consistent, it could be corrected for, but presumably it can vary considerably from case to case, introducing further noise.

Problems with Inflexible Measures

Given all these problems with flexible measures of morality, it might seem best to build our hypothetical moralometer instead around inflexible measures. Assuming physiological measures are unavailable, the most straightforward way to do this would be to employ researcher-chosen behavioral measures. We could try to measure someone's honesty by seeing whether they will cheat on a puzzle to earn more money in a laboratory setting. We could examine publicly available criminal records. We could see whether they are willing to donate a surprise bonus payment to a charity.

Unfortunately, inflexible measures don't fully escape the troubles that dog flexible measures, and they bring new troubles of their own.

Relativism and particularism. Inflexible measures probably aggravate the problems with relativism and particularism discussed above. With self-report and informant report, there's at least an opportunity for the self or the informant to take into account local standards and particulars of the situation. In contrast, inflexible measures will ordinarily be applied equally to all without adjustment for context. Suppose the measure is something like "gives a surprise bonus of $10 to charity". This might be a morally very different decision for a wealthy participant than for a needy participant. It might be a morally very different decision for a participant who would save that $10 to donate it to a different and maybe better charity than for a participant who would simply pocket the $10. But unless those other factors are being measured, as they normally would not be, they cannot be taken account of.

Incommensurability. Inflexible measures also won't avoid incommensurability problems. Suppose our moralometer includes one measure of honesty, one measure of generosity, and one measure of fairness. The default approach might be for a summary measure simply to average these three, but that might not accurately reflect morality: Maybe a small act of dishonesty in an experimental setting is far less morally important than a small act of unfairness in that same experimental setting. For example, getting an extra $1 from a researcher by lying in a task that transparently appears to demand a lie (and might even be best construed as a game in which telling untruths is just part of the task, in fact pleasing the researcher) might be approximately morally neutral while being unfair to a fellow participant in that same study might substantially hurt the other's feelings.

Sampling and ecological validity. As mentioned in my previous post on moralometers, fixed behavioral measures are also likely to have severe methodological problems concerning sampling and ecological validity. Any realistic behavioral measure is likely to capture only a small and perhaps unrepresentative part of anyone's behavior, and if it's conducted in a laboratory or experimental setting, behavior in that setting might not correlate well with behavior with real stakes in the real world. How much can we really infer about a person's overall moral character from the fact that they give their monetary bonus to charity or lie about a die roll in the lab?

Moral authority. By preferring a fixed measure, the experimenter or the designer of the moralometer takes upon themselves a certain kind of moral authority -- the authority to judge what is right and wrong, moral or immoral, in others' behavior. In some cases, as in the Eichmann case, this authority seems clearly preferable to deferring to the judgment of the target and their friends. But in other cases, it is a source of error -- since of course the experimenter or designer might be wrong about what is in fact morally good or bad.

Being wrong while taking up, at least implicitly, this mantle of moral authority has at least two features that potentially make it worse than the type of error that arises by wrongly deferring to mistaken raters. First, the error is guaranteed to be systematic. The same wrong standards will be applied to every case, rather than scattered in different (and perhaps partly canceling) directions as might be the case with rater error. And second, it risks a lack of respect: Others might reasonably object to being classified as "moral" or "immoral" by an alien set of standards devised by researchers and with which they disagree.

In Sum

The methodological problems with any potential moralometer are extremely daunting. As discussed in December, all moralometers must rely on some combination of self-report, informant report, behavioral measure, or physiological measure, and each of these methods has serious problems. Furthermore, as discussed today, a batch of issues around relativism, particularism, disagreement, incommensurability, error, and moral authority dog both flexible measures of morality (which rely on raters' judgments about what's good and bad) and inflexible measures (which rely on researchers' or designers' judgments).

Coming up... should we even want a moralometer if we could have one?  I discussed the desirability or undesirability of a perfect moralometer in December, but I want to think more carefully about the moral consequences of the more realistic case of an imperfect moralometer.

Thursday, December 22, 2022

The Moral Measurement Problem: Four Flawed Methods

[This post draws on ideas developed in collaboration with psychologist Jessie Sun.]

So you want to build a moralometer -- that is, a device that measures someone's true moral character? Yes, yes. Such a device would be so practically and scientifically useful! (Maybe somewhat dystopian, though? Careful where you point that thing!)

You could try to build a moralometer by one of four methods: self-report, informant report, behavioral measurement, or physiological measurement. Each presents daunting methodological challenges.

Self-report moralometers

To find out how moral a person is, we could simply ask them. For example, Aquino and Reed 2002 ask people how important it is to them to have various moral characteristics, such as being compassionate and fair. More directly, Furr and colleagues 2022 have people rate the extent to which they agree with statements such as "I would say that I am a good person" and "I tend to act morally".

Could this be the basis of a moralometer? That depends on the extent to which people are able and willing to report on their overall morality.

People might be unable to accurately report their overall morality.

Vazire 2010 has argued that self-knowledge of psychological traits tends to be poor when the traits are highly evaluative and not straightforwardly observable (e.g., "intelligent", "creative"), since under those conditions people are (typically) motivated to see themselves favorably and -- due to low observability -- not straightforwardly confronted with the unpleasant news they would prefer to deny.

One's overall moral character is evaluatively loaded if anything is. Nor is it straightforwardly observable. Unlike height or talkativeness, someone motivated not to see themselves as, say, unfair or a jerk can readily find ways to explain away the evidence (e.g., "she deserved it", "I'm in such a hurry").

Furthermore, it sometimes requires a certain amount of moral insight to distinguish morally good from morally bad behavior. Part of being a sexist creep is typically not seeing anything wrong with the kinds of things that sexist creeps typically do. Conversely, people who are highly attuned to how they are treating others might tend to beat themselves up over relatively small violations. We might thus expect a moral Dunning-Kruger effect: People with bad moral character might disproportionately overestimate their moral character, so that people's self-opinions tend to be undiagnostic of the actual underlying trait.

Even to the extent people are able to report their overall morality, people might be unwilling to report it.

It's reasonable to expect that self-reports of moral character would be distorted by socially desirable responding, the tendency for questionnaire respondents to answer in a manner that they believe will reflect well on them. To say that you are extremely immoral seems socially undesirable. We would expect that people (e.g., Sam Bankman-Fried) would tend to want to portray themselves as morally above average. On the flip side, to describe oneself as "extremely moral" (say, 100 on a 0-100 scale from perfect immorality to perfect morality) might come across as immodest. So even people who believe themselves to be tip-top near-saints might not frankly express their high self-opinions when directly asked.

Reputational moralometers

Instead of asking people to report on their own morality, could we ask other people who know them? That is, could we ask their friends, family, neighbors, and co-workers? Presumably, the report would be less distorted by self-serving or ego-protective bias. There's less at stake when judging someone else's morality than when judging your own. Also, we could aggregate across multiple informants, combining several different people's ratings, possibly canceling out some sources of noise and bias.

Unfortunately, reputational moralometers -- while perhaps somewhat better than self-report moralometers -- also present substantial methodological challenges.

The informant advantage of decreased bias could be offset by a corresponding increased in ignorance.

Informants don't observe all of the behavior of the people whose morality they are judging, and they have less access to the thoughts, feelings, and motivations that are relevant to the moral assessment of behavior. Informant reports are thus likely to be based only on a fraction of the evidence that self-report would be based on. Moreover, people tend to hide their immoral behaviors, and presumably some people are better at doing so than others. Also, people play different roles in our lives, and romantic partners, coworkers, friends, and teachers will typically only see us in limited, and perhaps unrepresentative, contexts. A good moralometer would require the correct balancing of a range of informants with complementary patches of ignorance, which is likely to be infeasible.

Informants are also likely to be biased.

Informant reports may be contaminated not by self-serving bias but by "pal-serving bias" (Leising et al 2010). If we rely on people to nominate their own informants, they are likely to nominate people who have a positive perception of them. Furthermore, the informants might be reluctant "tell on" or badly evaluate their friends, especially in contexts (like personnel selection) where the rating could have real consequences for the target. The ideal informant would be someone who knows the target well but isn't positively biased toward you. In reality, however, there's likely a tradeoff between knowledge and bias, so that those who are most likely to be impartial are not the people who know you best.

Positivity bias could in principle be corrected for if every informant was equally biased, but it's likely that some targets will have informants who are more biased than others.

Behavioral moralometers

Given the problems with self-report and informant report, direct behavioral measures might seem promising. Much of my own work on the morality of professional ethicists and the effectiveness of ethics instruction has depended on direct behavioral measures such as courteous and discourteous behavior at philosophy conferences, theft of library books, meat purchases on campus (after attending a class on the ethics of eating meat), charitable giving, and choosing to join the Nazi party in 1930s Germany. Others have measured behavior in dictator games, lying to the experimenter in laboratory settings, criminal behavior, and instances of comforting, helping, and sharing.

Individual behaviors are only a tiny and possibly unrepresentative sample.

Perhaps the biggest problem with behavioral moralometers is that any single, measurable behavior will inevitably be a minuscule fraction of the person's behavior, and might not be at all representative of the person's overall morality. The inference from this person donated $10 in this instance or this person committed petty larceny two years ago to this person's overall moral character is good or bad is a giant leap from a single observation. Given the general variability and inconstancy of most people's behavior, we shouldn't expect a single observation, or even a few related observations, to provide an accurate picture of the person overall.

Although self-report and informant report are likely to be biased, they aggregate many observations of the target into a summary measure, while the typical behavioral study does not.

There is likely a tradeoff between feasibility and validity.

There are some behaviors that are so telling of moral character that a single observation might reveal a lot: If someone commits murder for hire, we can be pretty sure they're no saint. If someone donates a kidney to a stranger, that too might be highly morally diagnostic. But such extreme behaviors will occur at only tiny rates in the general population. Other substantial immoral behaviors, such as underpaying taxes by thousands of dollars or cheating on one's spouse, might occur more commonly, but are likely to be undetectable to researchers (and perhaps unethical to even try to detect).

The most feasible measures are laboratory measures, such as misreporting the roll of a die to an experimenter in order to win a greater payout. But it's unclear what the relationship is between laboratory behaviors for minor stakes and overall moral behavior in the real world.

Individual behaviors can be difficult to interpret.

Another advantage of self-report and to some extent informant report have over direct behavioral measures is that there's an opportunity for contextual information to clarify the moral value or disvalue of behaviors: The morality of donating $10 or the immorality of not returning a library book might depend substantially on one's motives or financial situation, which self-report or informant report can potentially account for but which would be invisible in a simple behavioral measure. (Of course, on the flip side, this flexibility of interpretation is part of what permits bias to creep in.)

[a polygraph from 1937]

Physiological moralometers

A physiological moralometer would attempt to measure someone's morality by measuring something biological like their brain activity under certain conditions or their genetics. Given the current state of technology, no such moralometer is likely to arise soon. The best known candidate might be the polygraph or lie detector test, which is notoriously unreliable and of course doesn't purport to be a general measure of honesty much less of overall moral character.

Any genetic measure would of course omit any environmental influences on morality. Given the likelihood that environmental influences play a major role in people's moral development, no genetic measure could have a high correlation with a person's overall morality.

Brain measures, being potentially closer to measuring the mental states that underlie morality, don't have a similar ceiling accuracy, but currently look less promising than behavioral measures, informant report measures, and probably even self-report measures.

The Inaccuracy of All Methods

It thus seems likely that there is no good method for accurately measuring a person's overall moral character. Self-report, informant report, behavioral measures, and physiological measures all face large methodological difficulties. If a moralometer is something that accurately measures an individual person's morality, like a thermometer accurately (accurately enough) measures a person's body temperature, there's little reason to think we could build one.

It doesn't follow that we can't imprecisely measure someone's moral character. It's reasonable to expect the existence of small correlations between some potential measures and a person's real underlying overall moral character. And maybe such measures could be used to look for trends aggregated across groups.

Now, this whole post has been premised on the idea that it make sense to talk of a person's overall morality as something that could be captured, at least in principle, by a number such as 0 to 100 or -1 to +1. There are a few reasons to doubt this, including moral relativism and moral incommensurability -- but more on that in a future post.

Monday, August 01, 2022

The Nature of Belief From a Philosophical Perspective, With Theoretical and Methodological Implications for Psychology and Cognitive Science

Every so often, I give a brief overview of my perspective on belief to audiences of psychologists. After the 2021 Creditions conference, I was asked to write up my thoughts and publish them in a special issue of Frontiers in Psychology (ed. Rüdiger J. Seitz).

Since it's short enough to fit in a (longish) blog post, I thought I'd post it here. Those who are already familiar with my work on belief won't find much new, but it might be a helpful overview for others. Plus, I direct a few gentle (?) jabs at Eric Mandelbaum, my favorite opponent on this topic.

[output from Dall-E for "belief philosophy psychology in style of Van Gogh"]

Introduction

In recent academic philosophy, representationalism is probably the dominant model of belief. I favor a competing model, dispositionalism. I will briefly describe these views and their contrasting implications, including some theoretical and methodological implications relevant to research psychologists and cognitive scientists.

Representationalism Vs. Dispositionalism, Definitions

According to representationalism, to believe some proposition P (for example, that there's beer in the fridge or that men and women are intellectually equal) is to have a representation with the content P stored in your mind, available to be deployed in relevant reasoning. It's somewhat unclear how literally the “storage” idea is to be taken, but leading representationalists, such as Fodor and Mandelbaum (Fodor, 1987; Mandelbaum, 2014; Quilty-Dunn and Mandelbaum, 2018; Bendaña and Mandelbaum, 2021), appear to take the storage idea rather literally. One might compare to the concept of the “long-term memory store” in theories of memory. The stored representation counts as available to be deployed in relevant reasoning if it can be accessed when relevant. If asked whether men and women differ in intelligence, you'll retrieve the representation that men and women are intellectually equal, engage in some simple theoretical reasoning, and answer “no” (if you want to be honest, etc.). If you feel like drinking a cold beer, you'll retrieve the representation that beer is in the fridge, engage in some simple practical reasoning, and walk toward the kitchen to get the beer.

According to dispositionalism, to believe that P is to be disposed to act and react in ways that are characteristic of believers-that-P. Maybe there's a representation really stored in there; maybe not. If you are disposed to go to the fridge when you want a beer, if you are disposed to say “yes” when asked whether there's beer in the fridge, if you display surprise upon opening the fridge and finding no beer, etc., then you count as believing that there's beer in the fridge, regardless what underlying cognitive architecture enables this. Dispositionalism has its roots in philosophical behaviorism and Ryle (1949). However, I and other recent dispositionalists eschew behaviorism, allowing that some of the relevant dispositions can be “phenomenal” (i.e., pertaining to conscious experience), such as the disposition to feel (and not just exhibit) surprise upon opening the fridge and seeing no beer, and other dispositions can be cognitive (i.e., pertaining to inference or other cognitive transitions), such as the disposition to draw the conclusion that there is beer in the house (Schwitzgebel, 2002, 2021).

Representationalism commits to a particular type of cognitive architecture—the storage of representational contents matching the contents of the believed propositions—and it is to a substantial extent neutral about the extent to which the stored contents are behavior-guiding. Dispositionalism commits to belief as behavior-guiding, while remaining neutral on the underlying architecture. The difference matters to psychological theory and method as I will now explain.

In-Between Believing

On representationalism, it's natural to think of belief as a yes/no matter. P is either stored or it's not. You either believe it or you don't. Representations can't normally be “half-stored.” What would that even mean? If the representation isn't retrieved when relevant, it's a “performance” failure; the underlying “competence” is still there, as long as it could in principle be retrieved in some circumstances. This leads some representationalists, especially Mandelbaum, to unintuitive views about what we believe. For example, if someone tells you “dogs are made of paper,” Mandelbaum holds that you will believe that proposition—even after you reject it as obviously false—because the representation gets stored and starts influencing your cognition. Of course you also simultaneously believe that dogs are not made of paper.

On dispositionalism, believing is more like having a personality trait: You match the dispositional profile to some degree, just like you might match the dispositional profile characteristic of extraversion to some degree. Sometimes, the match might be nearly perfect. I might have all the dispositions characteristic of the belief that there's beer in my fridge. Other times, the match might be far from perfect. Cases of highly imperfect match can be described as in-between cases of belief.

Consider the belief that men and women are intellectually equal. Someone—call him the “implicit sexist”—might be disposed to act and react in some ways that are characteristic of that belief. He might say “men and women are intellectually equal” with a feeling of confidence and sincerity, ready to defend that view passionately in a debate. Other dispositions might tilt the other way. He might feel surprised if a woman makes an intelligent comment at a meeting, and it might take more evidence to convince him that a woman is smart than that a man is smart.

Or consider gradual forgetting. In college, I knew the last name of my roommate's best friend. I could easily recall it. Over time, as memory faded, I would have been able to recognize it, picking it out from nearby alternatives, but recall would have been weaker. As memory continued to fade, I would have recognized it less and less reliably until eventually it was utterly forgotten. During the intermediate phase, I would in some respects act and react like someone would believed his name was (let's say) Guericke, in other respects not. There was no precise moment at which the belief dropped from my mind, instead a long period of gradual, fading in-betweenness.

Dispositionalist views naturally invite us see belief as permitting in-between cases, as personality traits do. Representationalist views have more difficulty accommodating this idea.

Contradictory Belief

Conversely, representationalist views naturally allow for contradictory belief, as discussed in the “dogs are made of paper” example, while dispositionalist views appear to disallow the possibility of having contradictory beliefs. There seems to be no problem in principle in storing both the representation “P” and the representation “not-P.” But one cannot simultaneously have the dispositional structure characteristic of believing that men and women are intellectually equal and the dispositional structure characteristic of believing that women are intellectually inferior. That would be like having the dispositional structure of an extravert and simultaneously the dispositional structure of an introvert—structurally impossible.

Given an implicit sexism case, then, representationalism tends to favor the idea that the sexist believes both that women and men are intellectually equal and that women are intellectually inferior. The two contradictory beliefs are both stored and accessible (perhaps in different cognitive subsystems, retrieved under different conditions). Dispositionalism tends to favor treating such cases as in-between cases of belief. Similarly for other inconsistent or conflicting attitudes: the Sunday theist/weekday atheist; the self-deceived lover who sincerely denies that their partner is cheating but sometimes acts as if they know; the person who would say the road runs north-south if queried in one way but who would say it runs east-west if queried in another way.

Let me briefly defend the dispositionalist stance on this issue. We have no need for contradictory belief. It helps none to say of the implicit sexist that he believes both “men and women are intellectually equal” and “women are intellectually inferior.” To make such a claim comprehensible, we need to present the details: In these respects he acts and reacts like an egalitarian, in these other respects he acts and reacts like a sexist. But now we've just given the dispositional characterization. If necessary—if there are good enough architectural grounds for it—we might still say that he has contradictory representations. But representation is not belief.

Explanatory Depth Vs. Explanatory Superficiality

Quilty-Dunn and Mandelbaum (2018) argue that representationalism has an explanatory depth that coheres well with the aims of cognitive science. If the belief that P is a relation to a stored representational content “P,” we can explain how beliefs cause behavior (retrieving the stored representation does the causal work), we can explain why there's usually such a nice parallel between what we can say and what we can believe (speech and belief involve accessing the same pool of representations), and so forth. The dispositionalist approach, in contrast, is superficial: It points to the dispositional patterns but it does not attempt to explain the causal mechanisms beneath those patterns.

While explanatory depth is a virtue when available, it is not a virtue in this particular case. To think that belief that P always, or typically, involves having an internal representational content “P” is a best empirically unsupported. (Contrast with the empirically well supported claim that the visual system represents motion in regions of the visual field.) At worst, it is a simplistic cartoon sketch of the mind. It's as if someone insisted that having the personality trait of extraversion required having an internal switch flipped to “E,” because otherwise we'd be stuck without an internal causal explanation of extraverted patterns of behavior. Of course there are internal structures that help explain people's extraverted behavior, and of course there are internal structures that help explain people's implicitly sexist behavior and their beer-fetching behavior. But we need not define belief in terms of a simplistic representationalist understanding of those internal structures.

Still, a partial compromise is possible. It might be the case that internal representations of P are present whenever one believes that P. The dispositionalist need not deny this—any more than a personality theorist need not deny that extraversion might involve an heretofore-undiscovered E switch. The dispositionalist just doesn't define belief in terms of such structures, permitting a skeptical neutrality about them.

Intellectualism Vs. Pragmatism

I will now introduce a second philosophical distinction. According to intellectualism about belief, sincere assent or assertion is sufficient or nearly sufficient for belief. According to pragmatism about belief, to really, fully believe you need not just to be ready to say P; you need also to act accordingly.

The intellectualism/pragmatism distinction cross-cuts the representationalism/dispositionalism distinction. However, I submit that the most attractive form of dispositionalism is also pragmatist. To really, fully believe that women are intellectually equal requires more than simply readiness to say they are. It requires not being surprised when a women makes an intelligent remark. It requires treating the women you encounter as if they are just as smart as men in the same circumstances. Alternatively, to really believe that your children's happiness is more important than their academic success it's insufficient to be disposed to say that is the case; you must also to live that way.

The Problem With Questionnaires

I conclude with two methodological implications.

First, if pragmatist dispositionalism is correct, then you might not know what you believe. Do you really believe that men and women are intellectually equal? Do you really believe that your children's happiness is more important than their academic success? You'll say yes and yes. But how do you really live your life? You might be more in-betweenish than you think.

When psychologists want to explore broad, life involving beliefs and values, they often employ questionnaires. Questionnaires are easy! But if pragmatist dispositionalism is correct, questionnaires risk being misleading when asking about beliefs or other attitudes with an important lived component that can diverge from verbal endorsement. Questionnaires get at what you say, not at how you generally act.

A brief example: The Short Schwartz's Values Survey (Lindeman and Verkasalo, 2005) asks participants how important it is to them to achieve “power (social power, authority, wealth)” and various other goods. If intellectualism is the right way to think about values, this is an excellent methodology. However, if pragmatism is better, it's reasonable to doubt how well people know this about themselves.

Developing Beliefs

Developmental psychologists often debate the age children reach various cognitive milestones, such as knowing that objects continue to exist even when they aren't being perceived and knowing that people can have false beliefs. If representationalism is correct, then it's natural to suppose that there is in fact some particular age at which each individual child finally comes to store the relevant representational content. However, if dispositionalism is correct, gradualism is probably more attractive: Such broad beliefs are slowly constructed, involving many relevant dispositions, which might accrete unevenly and unstably over months or years.

In my experience, developmental psychologists often endorse gradualism when explicitly asked. Yet their critiques of each other seem sometimes implicitly to assume the contrary. “Boosters” (who claim that knowledge in some domain tends to come early) reject as too demanding methodologies that appear to reveal later knowledge. “Scoffers” (who claim that knowledge in some domain tends to come late) reject as too easy methodologies that appear to reveal earlier knowledge. Each trusts only the methods that reveal knowledge at the “right” age. But while of course some methodologies might be flawed, the gradualist dispositionalist ought to positively expect that across a variety of equally good methods for discovering whether the child knows P, some should reveal much earlier knowledge than others, though none are flawed—because knowing that P is not a yes-or-no, not an on-or-off thing. There need be no one right age or set of methods. (For more on this issue, see Schwitzgebel, 1999; McGeer and Schwitzgebel 2006.)

------------------------------------

Related:

"Gradual Belief Change in Children", Human Development, 42 (1999), 283-296.

"In-Between Believing", Philosophical Quarterly, 51 (1999), 76-82.

"A Phenomenal, Dispositional Account of Belief", Nous, 36 (2002), 249-275.

"Acting Contrary to Our Professed Beliefs, or the Gulf Between Occurrent Judgment and Dispositional Belief", Pacific Philosophical Quarterly, 91 (2010), 531-553.

"Do You Have Infinitely Many Beliefs about the Number of Planets?", Oct 17, 2012.

"It's Not Just One Thing, To Believe There's a Gas Station on the Corner", Feb 28, 2018.

"Superficialism about Belief", Jul 16, 2020.

"The Pragmatic Metaphysics of Belief" in Cristina Borgoni, Dirk Kindermann, and Andrea Onofri, eds., The Fragmented Mind (Oxford, 2021).

This is just a sample of my work on belief. I've been hacking away on these points since my dissertation 25 years ago!

Wednesday, February 16, 2022

Qualitative Research Reveals a Potentially Huge Problem for Standard Methods in Experimental Philosophy

Mainstream experimental philosophy aims to discover ordinary people's opinions about questions of philosophical interest. Typically, this involves presenting paragraph-long scenarios to online workers. Respondents express their opinions about the scenarios on simple quantitative scales. But what if participants regularly interpret the questions differently than the researchers intend? The whole apparatus would come crashing down.

Kyle Thompson (who recently earned his PhD under my supervision) has published the central findings of a dissertation that raises exactly this challenge to experimental philosophy. His approach is to compare the standard quantitative measures of participants' opinions -- that is, participants' numerical responses on standardized questions -- with two qualitative measures: what participants say when instructed to "think aloud" about the experimental stimuli and a post-response interview about why they answered the way they did.

Kyle's main experiment replicates the quantitative results of an influential study that purports to show that ordinary research participants reject the "ought implies can" principle. According to the ought-implies-can principle, people can only be morally required to do what it is possible for them to do. Thompson replicates the quantitative results of the earlier experiment, seeming to confirm that participants reject ought-implies-can. However, Thompson's qualitative think-aloud and interview results clearly indicate that his participants actually accept, rather than reject, the principle. The quantitative and the qualitative results point in opposite directions, and the qualitative results are more convincing.

In the scenario of central interest, "Brown" agrees to meet a friend at a movie theater at 6:00. But then

As Brown gets ready to leave at 5:45, he decides he really doesn't want to see the movie after all. He passes the time for five minutes, so that he will be unable to make it to the cinema on time. Because Brown decided to wait, Brown can't meet his friend Adams at the movie by 6.

Participants then rate their degree of agreement or disagreement with the following three questions:

At 5:50, Brown can make it to the theater by 6

Brown is to blame for not making it to the theater by 6

Brown ought to make it to the theater by 6

As you might expect, in both the original article and Thompson's replication, participants almost all disagree that Brown can make it to the theater by 6. So far, so good. However, apparently in violation of the ought-implies-can principle, participants overall tended to agree that Brown is to blame for not making it to the theater by 6 and (to a lesser extent) that Brown ought to make it to the theater by 6. Interpreting the results at face value, it appears that regarding making it to the theater by 6, participants think that Brown cannot do it, that he is blameworthy for not doing it, and that he ought to do it -- and thus that someone can be blameworthy for failing to do, and ought to do, something that it is not possible for them to do.

Now, if your reaction to this is wait a minute..., you share something in common with Thompson and me. Participants' think-aloud statements and subsequent interviews reveal that almost all of them reinterpret the questions to preserve consistency with the ought-implies-can principle. For example, some participants explain their positive answers to "Brown ought to make it to the theater by 6" by explaining that Brown ought to try to make it to the theater by 6. Others change the tense and the time referent, explaining that Brown "could have" made it to the theater and that he should have left by 5:45. There is no violation of ought-implies-can in either response. At 5:50, Brown could presumably still try to make it to the theater. And at 5:45 he still could have made it to the theater.

Through careful examination of the transcripts, Thompson discovers that the almost 90% of participants in fact adhere to the ought-implies-can principle in their responses, often reinterpreting the content or tense of the questions to render them consistent with this principle.

As far as I'm aware, this is the first attempt to replicate a quantitative experimental philosophy study with careful qualitative interview methods. What it suggests is that the surface-level interpretation of the quantitative results can be highly misleading. The majority of participants appear to have the opposite of the view suggested by their quantitative answers.

It is an open question how much of the quantitative research in experimental philosophy would survive careful qualitative scrutiny. I hope others follow in Kyle's footsteps by attempting careful qualitative replications of important quantitative work in the subdiscipline.

Friday, April 30, 2021

Are 15 UCLA Anthropology Graduate Students Representative of the Los Angeles Population? More Thoughts on Henrich's The WEIRDest People in the World

Earlier this month, I complained about Joseph Henrich's somewhat loose summaries of scientific research in his recent, influential book The WEIRDest People in the World.  At the time of the post, I had read through Chapter 6.

One of my complaints was that in explaining how research on economic games works, Henrich's paradigmatic fictional example described giving each research participant $20 to $30 per game over the course of ten games, totaling $200-$300 per participant.  However, economic games rarely have stakes that large.  More typically, the stakes are about a tenth of that.  Overstating the amount typically at stake illegitimately prevents the naive reader from forming the skeptical thought that people might behave differently with small amounts of laboratory money than with the larger amounts commonly at stake in real-world situations.

Despite my concerns, I find Henrich's book fascinating, and I am finding much of value in it.  So I kept reading.  Last week, I hit Chapter 9 and, given my complaints about Chapter 6, I was struck by the following paragraph:

These interviews contrasted with those I did in Los Angeles after administering an Ultimatum Game that put $160 on the line.  It was a sum that was calculated to match the Matsigenka stakes. [Matsigenka live in small farming hamlets in the Amazon.]  In this immense urban metropolis, people said they'd feel guilty if they gave less than half.  They conveyed the sense that offering half was the "right" thing to do in this situation.  The one person who made a low offer (25 percent) deliberated for a long time and was clearly worried about rejection.  

Wait, $160 per participant per game?!  

(In the Ultimatum Game, Person A is given a sum of money to split with Person B.  Person A proposes a split -- say, 50/50 or 80/20 -- and then Person B has the choice either to accept the resulting split or reject the offer, in which case neither player gets any money.)

I had to look up the study.  Indeed, Henrich did offer $160 to participants.  But -- understandably given the amounts at stake -- the sample size was very small: only 15 (that is, 15 people in the Person A role, whose offers provided the main data).  And those 15 people were all graduate students in the Anthropology Department at UCLA, paired with 15 other anthropology students.

While it's not exactly wrong for Henrich to summarize the data as he did, his presentation omits details that seem to me quite relevant and which might fuel a skeptical interpretation.  Should we consider 15 UCLA Anthro grad students representative of the Los Angeles population?  Henrich treats their behavior as representative without explicitly flagging for the reader how unusual a group they are.

In the original article, Henrich does make a case for choosing this population.  It's a group of acquaintances, like the Matsigenka population was a group of acquaintances.  Like the Matsigenka participants, the graduate students all personally knew the experimenter, Henrich himself.  That could potentially control for any inclination to be more generous in order to create a favorable impression on a high-status, high-resource acquaintance.  In the original article and in at least one later re-presentation of the work, Henrich explicitly acknowledges some of the potential concerns with taking these students as representative of the larger U.S. urban population.

But of course all of this is hidden beneath Henrich's description in his book of the participants as merely being from "the immense urban metropolis" of Los Angeles.  Given only that description of them, you might reasonably guess that the L.A. participants were strangers recruited off the streets.

Yes, readers can't be told every detail, especially in a book of such sweeping scope as Henrich's.  This creates a situation in which the reader must trust the author.  As an author, part of your job is to warrant that trust.  As a critical reader, part of your job is to assess as best you can whether the author in fact warrants trust.  One tool the reader can use spot checking, especially when the author enters areas where you have some independent sources of knowledge.

If you're inclined to trust Henrich's judgment that these 15 anthropology students were a well-chosen representative sample of Angelenos, then his omission is one you should feel comfortable enough with.  You should think, "I'm in good hands.  He's not distracting me with irrelevant details."  But my own sense is different.  Henrich omits crucial details about his population that I would want to know, and that I think readers in general should want to know so that they can think critically about the presented research.

-----------------------------------------------------------

By the way, Henrich replied on Facebook to my earlier blog post about the book.  If you're curious, check it out.  My sense is that his characterization of my post is inaccurate and that he did not correct that mischaracterization when given an opportunity to do so.  Please feel free to read my earlier post to judge whether I'm being fair in my complaint.

[image source]

Wednesday, April 07, 2021

On Scientific Trust, Loose Summaries, and Henrich's WEIRDest People in the World

Joseph Henrich's ambitious tome, The WEIRDest People in the World, is driving me nuts.  It's good enough and interesting enough that I want to read it.  Henrich's general idea is that people in Western, Educated, Industrial, Rich, Democratic (WEIRD) societies differ psychologically from people in more traditionally structured societies, and that the family policies of the Catholic Church in medieval Europe lie at the historical root of this difference.  It's very cool and I'm almost convinced!

Despite my fascination with his argument, I find that when Henrich touches on topics I know something about, he tends to distort and simplify things.  Maybe this is inevitable in a book of such sweeping scope.  However, it does lead me to mistrust his judgment and wonder how accurate his presentation is on topics where I have no expertise.

Early in reading, I was struck by Henrich's presentation of the famous / notorious "marshmallow test".  Here's his description:

To measure self-control in children, researchers sit them in front of a single marshmallow and explain that if they wait until the experimenter returns to the room, they can have two marshmallows instead of just the one.  The experimenter departs and then secretly watches to see how long it takes for the kid to cave and eat the marshmallow.  Some kids eat the lone marshmallow right away.  A few wait 15 or more minutes until the experimenters gives up and returns with the second marshmallow.  The remainder of the children cave somewhere in between.  A child's self-control is measured by the number of seconds they wait.

Psychological tasks likes these are often powerful predictors of real-life behavior (p. 40).

It's a cute test!  However, I have a graduate student who is currently writing a dissertation chapter on problems with this test.  Maybe the test is a measure of self-control, but it could also be a measure of how much the child trusts the experimenter to actually deliver on the promise, or how much the child desires the social approval of the experimenter, or how comfortable the child is with strange laboratory experiments of this sort, or how hungry they are, how much they want to end the situation so as to reunite with their waiting parent, etc.  Indeed, the a recent conceptual replication of the experiment mostly does not find the types of predictive value that were claimed in early studies, after statistical controls are introduced to account for race, gender, home background, parents' education, vocabulary, and other possible covariates.[1]  

In general, if you've been influenced, as I have, by the "replication crisis" and other recent methodological critiques of social science and medicine, this might be the kind of result that should set off your skeptical tinglers.  The idea that how long a four-year-old waits before eating a marshmallow reveals how much self-control they have, which then "powerfully predicts" real-life behavior outside of the laboratory (e.g., college admission test scores over a decade later, as is sometimes claimed) -- well, it could be true.  I'm not saying it's not.  But I don't think I'd have written it up as Henrich does, without skeptical caveats, as though there's general consensus among psychologists that a child's behavior with a single marshmallow in this peculiar laboratory situation is a valid, powerful measure of self-control with excellent predictive value.  Its prominent placement near the beginning of the book furthermore suggests that Henrich regards this test as part of the general theoretical foundation on which psychological work like his appropriately builds.

In this matter, my knowledgeable judgment and Henrich's differ.  That's fine.  Researchers can differently weigh the considerations.  But if I hadn't had the background knowledge I did, his quick presentation might have led me into a much more optimistic assessment of the value of the marshmallow test than I would have arrived at from a more thorough presentation that acknowledged the caveats.  So there's a sense in which Henrich's presentation is a bad fit for my theoretical inclinations.

Here's another passage that bothered me:

Upon entering the economics laboratory, you are greeted by a friendly student assistant who takes you to a private cubicle.  There, via a computer terminal, you are given $20 and placed into a group with three strangers.  Then, all four of you are given an opportunity to contribute any portion of your endowment -- from nothing at all to $20 -- to a "group project."  After everyone has had an opportunity to contribute, all contributions to the group project are increased by 50 percent and then divided equally among all four group members.  Since players get to keep any money that they don't contribute to the group project, it's obvious that players always make the most money if they give nothing to the project.  But, since any money contributed to the project increases ($20 becomes $30), the group as a whole makes more money when people contribute more of their endowment.  Your group will repeat the interaction for 10 rounds, and you'll receive all of your earnings in cash at the end.  Each round, you'll see the anonymous contributions made by others and your own total income.  If you were a player in this game, how much would you contribute in the first round with this group of strangers?

This is the Public Goods Game (PGG).  It's an experiment designed to capture the basic economic trade-offs faced by individuals when they decide to act in the interest of their broader communities....  societies with more intensive kin-based institutions contribute less on average to the group project in the first round (p. 210-211).

This describes a study in which participants will receive $200-$300 each.  Of course, it's rare to award research participants such large amounts of money.  If you want, say, 200 participants, you'll need a $60,000 budget!  Henrich's endnotes cite two general books, one brief commentary without empirical data, two classic articles in which participants exited the experiment having earned about $30 each on average, and two cross-cultural studies whose payout amounts weren't readily discoverable by me from looking at the materials.  Also in the notes, Henrich says that one study "increased contributions to the group project by 40 percent, not 50 percent.  I'm simplifying" (p. 543).  However, the majority of the cited studies in fact used 40 percent increases, not just the one study to which this caveat was attached.

I'm not seeing why the more accurate 40% is "simpler" than 50%.  This seems to be a gratuitous inaccuracy.  Characterizing the experiment as ten rounds with payoffs of $20-$30 per round is potentially a more serious distortion.  Really, these experiments are run with units that are later exchanged for small amounts of real money.  This is important for at least two reasons: First, these experimental monetary units might be psychologically different from real money, possibly encouraging a more game-like attitude.  And second, when the actual amounts of money at stake are small, the costs of cooperating (and also the benefits) are less, which should amplify concerns about how representative this game-like laboratory behavior is of how the participants would behave in the real world, with more serious stakes.

Suppose that instead of exaggerating the stakes upward by a factor of about 10, Henrich had exaggerated the stakes down by a factor of about 10.  What if, instead of saying that there was $20-$30 at stake per turn, when it's typically more like $2-$3, he had said that $0.20 was at stake per turn?  I suspect this would make an intuitive difference to most ordinary readers of the book.  The leap from "here's how cooperatively research subjects act with $20" to "here's how cooperative people in that culture are with strangers in general" is more attractive than the leap from "here's how cooperatively research subjects act with $0.20" to the same broad conclusion.

In general, I tend to be wary of quick inferences from laboratory behavior to real-world behavior outside the laboratory.  Laboratories are strange social situations and differently familiar to people from different backgrounds.  This is the problem of ecological validity or external validity, and concerns of this sort are why most of my own research on social behavior uses real-world measures.  Other researchers, such as Henrich, might not be as worried about the external validity of laboratory/internet studies.  There's room for legitimate debate.  But in order for us readers to get a sense of whether external validity might be an issue in the studies he cites, at the very least we need an accurate description of what the studies involve.  Henrich's presentation does not provide that, and simplification is a poor motive for this distortion, since $2 is no less simple than $20.

Henrich does not, in my mind, cross over into bald misrepresentation.  He doesn't, for example, say of any particular study that it involves $20 per round.  Rather, the presentation seems to be loose.  He's trying to give the general flavor.  He's writing for a moderately broad audience and aiming to synthesize a huge range of work, unavoidably simplifying and idealizing along the way.  He could respond to my concerns by saying that his best judgment of the conflicting evidence about the marshmallow test is that it's a valid and highly predictive measure of self-control and that his simplified presentation of the material conveys that effectively by avoiding concerns and apparent replication failures that would just (in his judgment) be distracting.  He could say that his best reading of the literature on external validity is that the difference between $2 and $20 doesn't matter and that the quick leap to general conclusions about cooperativeness is justified because we can reasonably expect laboratory studies of this sort to be diagnostic.  He could say that the reader ought to trust that he's done his homework behind the scenes.

We must always trust, to some extent, the scientists we're reading -- that they are reporting their data correctly, that there aren't big problems with the study's execution that they're failing to reveal, and so on.  Part of this involves relying on their inevitably simplified summaries of material with which we are unfamiliar.  We trust the researcher to have digested the material well and fairly, and not to be hiding worries that might legitimately undermine the central claims.  The looser the presentation, the more trust is required.  

This invites the question of whether there are conditions under which more versus less trust is justified.  How much, as a reader, ought you be willing to glide through on trust?

I'd recommend reducing trust under the following three conditions:

(1.) The author has a prior agenda or a big picture theory that might motivate them to interpret and digest results in a biased way.  Most scientists have agendas and theories, of course, and certainly Henrich does.  But there is more and less agenda-driven work, and adversarial collaboration offers the opportunity for bias to be balanced through scientists' opposing agendas.

(2.) The author is not as skeptical as you the reader are about some of the relevant types of research.  If the author is less skeptical than you are, they might be interpreting that research more naively or more at face value than you would if you had read the same research.

(3.) Where the author makes contact with the issues you know best, they seem to be distorting, misinterpreting, or leaping too quickly to broad conclusions.  This might indicate a general bias and sloppiness that might be present but harder for you to see regarding issues about which you know less.

On all three grounds, my trust of Henrich is impaired.

--------------------------------------------------

Update, April 30: See my continuing thoughts about the book here.  See also Henrich's reply to my post here.

--------------------------------------------------

[1] Deep in an endnote, Henrich acknowledges this last concern.  He responds that "it's easy to weaken the relationship between measures of patience and later academic performance by statistically removing all the factors that create variation in patience in the first place" (p. 515).  It's a reasonable, though disputable point.  Regardless, few readers are likely to pick up on something buried in the back half of one among hundreds of endnotes.


Tuesday, March 23, 2021

Empirical Relationships Among Five Types of Well-Being

My new article with Seth Margolis, Daniel Ozer, and Sonja Lyubomirsky is now available as part of a free, open-access anthology on well being with Oxford University Press.

Seth, Dan, Sonja, and I divide philosophical approaches to well being into five broad classes -- hedonic, life satisfaction, desire fulfillment, eudaimonic, and non-eudaimonic objective list. There are many things that a philosopher, psychologist, or ordinary person can mean when they say that someone is "doing well". They're not all the same conceptually, and as we show in the article, they are also empirically distinguishable.

Because there are several types of well-being that are conceptually and empirically different, research findings concerning one type of "well-being" shouldn't automatically be assumed to generalize to other types. For example, what is true about hedonic well-being (having a preponderance of positively valenced over negatively valenced emotions) isn't necessarily true about eudaimonic well being (flourishing in one's distinctively human capacities, such as in friendship and productive activity).

As part of the background for this comparative project, we developed new measures for four of these five types of well-being, including desire fulfillment (how well are you fulfilling the desires you regard as most important), life satisfaction, eudaimonia, and what we call Rich & Sexy Well-Being (wealth, sex, power, and physical beauty; manuscript available on request).  We found positive relationships among all types of well-being (by respondents' self-ratings), but the correlations ran from .50 to .79 (disattenuated), rather than approaching unity.

We also found that the different types of well-being correlated differently with other measures. For example, the "Big Five" personality trait of Openness to Experience has generally not been found to correlate much with measures of well-being. However, we found that it correlated at .45 with our measure of eudaimonic well-being -- a fairly high correlation by social science standards -- and .57 with the "creative imagination" subscale specifically. Openness correlated much less with the other types of well-being, .07 to .21. Thus, a researcher employing a hedonic or life-satisfaction approach to well-being might conclude that the personality trait of Openness to Experience was unrelated to psychological well-being, whereas a researcher who favors a eudaimonic approach might conclude the opposite.

Well-being research is always implicitly philosophical. It always carries contestable assumptions about what well-being consists of. One's choice of well-being measure reflects those implicit assumptions.

Wednesday, February 17, 2021

Three Faces of Validity: Internal, Construct, and External

I have a new draft paper in circulation, "The Necessity of Construct and External Validity for Generalized Causal Claims", co-written with two social scientists, Kevin Esterling and David Brady.  Here's a distillation of the core ideas.


-----------------------------------------


Consider a simple causal claim: "α causes β in γ".  One type of event (say, caffeine after dinner) tends to cause another type of event (disrupted sleep) in a certain range of conditions (among typical North American college students).

Now consider a formal study you could run to test this.  You design an intervention: 20 ounces of Peet's Dark Roast in a white cup, served at 7 p.m.  You design a control condition: 20 ounces of Peet's decaf, served at the same time.  You recruit a population: 400 willing undergrads from Bigfoot Dorm, delighted to have free coffee.  Finally, you design a measure of disrupted sleep: wearable motion sensors that normally go quiet when a person is sleeping soundly.

You do everything right.  Assignment is random and double blind, everyone drinks all and only what's in their cup, etc., and you find a big, statistically significant treatment effect: The motion sensors are 20% more active between 2 and 4 a.m. for the coffee drinkers than the decaf drinkers.  You have what social scientists call internal validity.  The randomness, excellent execution, and large sample size ensure that there are no systematic differences between the treatment and control groups other than the contents of their cups (well...), so you know that your intervention had a causal effect on sleep patterns as measured by the motion sensors.  Yay!

You write it up for the campus newspaper: "Caffeine After Dinner Interferes with Sleep among College Students".

But do you know that?

Of course it's plausible.  And you have excellent internal validity.  But to get to a general claim of that sort, from your observation of 400 undergrads, requires further assumptions that we ought to be careful about.  What we know, based on considerations of internal validity alone, is that this particular intervention (20 oz. of Peet's Dark Roast) caused this particular outcome (more motion from 2 to 4 a.m.) the day and place the experiment was performed (Bigfoot Dorm, February 16, 2021).  In fact, even calling the intervention "20 oz. of Peet's Dark Roast" hides some assumptions -- for of course, the roast was from a particular batch, brewed in a particular way by a particular person, etc.  All you really know based on the methodology, if you're going to be super conservative, is this: Whatever it is that you did that differed between treatment and control had an effect on whatever it was you measured.

Call whatever it was you did in the treatment condition "A" and whatever it was you did differently in the control condition "-A".  Call whatever it was you measured "B".  And call the conditions, including both the environment and everything that was the same or balanced between treatment and control, "C" (that it was among Bigfoot Dorm students, using white cups, brewed an average temperature of 195°F, etc.).

What we know then is that the probability, p, of B (whatever outcome you measured), was greater given A (whatever you did in the treatment condition) than in -A (whatever you did in the control condition), in C (the exact conditions in which the experiment was performed).  In other words:

p(B|A&C) > p(B|-A&C).  [Read this as "The probability of B given A and C is greater than the probability of B given not-A and C."]

But remember, what you claimed was both more specific and more general than that.  You claimed "caffeine after dinner interferes with sleep among college students".  To put it in the Greek-letter format with which we began, you claimed that α (caffeine after dinner) causes β (poor sleep) in γ (among college students, presumably in normal college dining and sleeping contexts in North America, though this was not clearly specified).

In other words, what you think is true is not merely the vague whatever-whatever sentence

p(B|A&C) > p(B|-A&C)

but rather the more ambitious and specific sentence

p(β|α&γ) > p(β|-α&γ).[1]

In order to get from one to the other, you need to do what Esterling, Brady, and I call causal specification.

You need to establish, or at least show plausible, that α is what mattered about A.  You need to establish that it was the caffeine that had the observed effect on B, rather than something else that differed between treatment and control, like tannin levels (which differed slightly between the dark roast and decaf).  The internally valid study tells you that the intervention had causal power, but nothing inside the study could possibly tell you what aspect of the intervention had the causal power.  It may seem likely, based on your prior knowledge, that it would be the caffeine rather than the tannins or any of the potentially infinite number of other things that differ between treatment and control (if you're creative, the list could be endless).

One way to represent this is to say that alongside α (the caffeine) are some presumably inert elements, Î¸ (the tannins, etc.), that also differ between treatment and control.  The intervention A is really a bundle of Î± and Î¸: A = Î±&θ.  Now substituting Î±&θ for A, what the internally valid experiment established was

p(B|(α&θ)&C) > p(B|-(α&θ)&C).

If Î¸ is causally inert, with no influence on the measured outcome B, you can can drop the θ, thus inferring from the sentence above to 

p(B|α&C) > p(B|-α&C).

In this case, you have what Esterling, Brady, and I call construct validity of the cause.  You have correctly specified the element that is doing the causal work.  It's not just A as a whole, but α in particular, the caffeine.  Of course, you can't just assert this.  You ought to establish it somehow.  That's the process of establishing construct validity of the cause.

Analogous reasoning applies to the relationship between B (measured motion-sensor outputs) and β (disrupted sleep).  If you can establish the right kind of relationship between B and β you can move from a claim about B to a conclusion about β, thus moving from 

p(B|α&C) > p(B|-α&C)

to

p(β|α&C) > p(β|-α&C).

If this can be established, you have correctly specified the outcome and have achieved construct validity of the outcome.  You're really measuring disrupted sleep, as you claim to be, rather than something else (like non-disruptive limb movement during sleep).

And finally, if you can establish that the right kind of relationship holds between the actual testing conditions and the conditions to which you generalize (college students in typical North American eating and sleeping environments) -- then you can move from C to γ.  This will be so if your actual population is representative and the situation isn't strange.  More specifically, since what is "representative" and "strange" depends on what causes what, the specification of γ requires knowing what background conditions are required for α to have its effect on β.  If you know that, you can generalize to populations beyond your sample where the relevant conditions γ are present (and refrain from generalizing to cases where the relevant conditions are absent).  You can thus substitute γ for C, generating the causal generalization that you had been hoping for from the beginning:

p(β|α&γ) > p(β|-α&γ).

In this way, internal, construct, and external validity fit together.  Moving from finite, historically particular data to a general causal claim requires all three.  It requires establishing not only internal validity but also establishing construct validity of the cause and outcome and external validity.  Otherwise, you don't have the well-supported generalization you think you have.

Although internal validity is often privileged in social scientists' discussions of causal inference, with internal validity alone, you know only that the particular intervention you made (whatever it was) had the specific effect you measured (whatever that effect amounts to) among the specific population you sampled at the time you ran the study.  You know only that something caused something.  You don't know what causes what.

-----------------------------------------

Here's another way to think about it.  If you claim that "α causes β in γ", there are four ways you could go wrong:

(1.) Something might cause β in γ, but that something might not be α.  (The tannin rather than the caffeine might disrupt sleep.)

(2.) α might cause something in γ, but it might not cause β.  (The caffeine might cause more movement at night without actually disrupting sleep.)

(3.) α might cause β in some set of conditions, but not γ.  (Caffeine might disrupt sleep only in unusual circumstances particular to your school.  Maybe students are excitable because of a recent earthquake and wouldn't normally be bothered.)

(4.) α might have some relationship to β in γ, but it might not be a causal relationship of the sort claimed.  (Maybe, though an error in assignment procedures, only students on the noisy floors got the caffeine.)

Practices that ensure internal validity protect only against errors of Type 4.  To protect against errors of Type 1-3, you need proper causal specification, with both construct and external validity.

-----------------------------------------

Note 1: Throughout the post, I assume that causes monotonically increase the probability of their effects, including the presence of other causes.

-----------------------------------------

Related:



[image modified from source]