#AlwaysLookAtTheItems

“People with higher cognitive ability have weaker moral foundations”

Taken at face value, this is true. The study used a scale called the Moral Foundations Questionnaire 2 (MFQ-2), which says it measures moral foundations, and cognitive ability was negatively correlated with all of its sub-scales, which happen to be labelled “Care”, “Equality”, “Proportionality”, “Loyalty”, “Authority”, and “Purity”.

But looking at just the names psychologists give to the traits they measure is a good way to be misled about reality. Looking at how they define those traits can also be a good way to be misled. If you really want to understand what a psychological scale is measuring, you should look at how it actually measures it: the questions people are being asked.

For example, with the MFQ-2, names like “Equality” and “Loyalty” make you think the sub-scales are measuring broad traits like preferences for equal treatment and rights, or how much one values attachment and obligations to one’s friends, family, and other groups. But if you look at the items, you’ll see that the Equality sub-scale largely measures support for equalizing incomes and resources (e.g., “I believe it would be ideal if everyone in society wound up with roughly the same amount of money”), while the Loyalty sub-scale largely measures patriotism and national identity (e.g., “I think children should be taught to be loyal to their country”).

So, after looking at the items, we realize “people with higher cognitive ability have weaker moral foundations” is rather misleading. It isn’t that people with higher cognitive ability are less morally egalitarian in some broad sense, or that they’re particularly disloyal to their family and friends. What the study actually establishes is quite different: people with higher cognitive ability show less support equalizing income and resources, and have less patriotism and national identity.

The Systemizing Quotient (SQ) gives another example of the same problem. The SQ is supposed to measure a trait called “systemizing”, defined as “the drive to analyze or construct systems”, and men do score higher on the SQ than women do. If you stop there, the natural conclusion is that men have a substantially stronger drive to analyze and construct systems.

But when we actually look at the items, we find that, while many do seem related to thinking analytically, they’re unusually focused on male-typical interests like cars, computers, sports, stocks, technology, etc. When we instead look at factors made up of more gender-neutral content, such as nature, language, or attention to detail, the gender difference decreases dramatically. So at least part of the famous sex difference in “systemizing” is really a sex difference in the particular kinds of systems the questionnaire asks about.

Looking at the items themselves gets us much closer to understanding what a scale measures, but we can go another step. Not every item contributes equally to a scale or factor. We should also look at the item loadings: how strongly each item is related to the latent construct the scale is supposed to be measuring.

For example, taking a look at Tailcalled’s Targeted Personality Test, one might assume that a factor named “Charisma” measures how charismatic the respondent actually is. And, to be fair, there are some items dealing with relatively concrete social behaviors (e.g., “In conversations, I jump straight to the point rather than doing smalltalk and other irrelevant things”). But the highest-loading items are much more directly about how charismatic or socially skilled the respondent perceives themself to be.

This matters if, for example, one wants to look at correlations between Charisma and other variables like self-esteem. If the Charisma score is heavily determined by questions asking people whether they think they’re charismatic or socially competent, then a correlation between Charisma and self-esteem is partly a correlation between two kinds of positive self-evaluation. That’s a rather different result from showing that people who are actually perceived by others as charismatic have higher self-esteem.

There are related issues whenever the method of measurement introduces something important into the construct. If you’re reading a study (or Substack post) that’s looking at the relationship between attractiveness and other variables, but attractiveness is measured not through external raters but through self-reports, then a substantial component of the “attractiveness” variable is going to be how attractive people think they are. So the people who score as most “attractive” won’t necessarily just be the best-looking people; they’ll also tend to be the people with the highest opinions of their own appearance.

There was a recent analysis that found that people who read horoscopes every day report personalities more similar to what we’d expect from stereotypes about their Zodiac sign, while the effect among people who never read horoscopes was approximately zero. But remember how we usually measure personality: through self-reports. One possibility is that there’s a sort of placebo affect where astrology changes people’s actual personalities if people believe in astrology. Another, much more likely possibility is that people who believe in astrology incorporate the stereotypes associated with their Zodiac sign into their self-perception and therefore into how they answer personality questionnaires. If we want to investigate the effect of reading horoscopes on someone’s actual personality, we’d have to look at other-ratings or behavioral measures rather than solely self-ratings.

The broader lesson is this: look at the items. The name of a scale is just the researcher’s summary of what they think the scale is measuring (sometimes not even that!), and researchers are often wrong. The items, on the other hand, are the essence of the scale. So, always look at the items.