Well, Claude just helped me solve this problem satisfactorily to my mind and identified the literature that I would need to find in order to investigate further. Claude gave cautions on potentially hallucinated references since it came from training data. But anyway, here is the output.
"[T]he philosophy of language has actually worked on exactly this problem extensively under the heading of
generic sentences. That literature is directly relevant to what you're trying to do.
Here's a taxonomy that I think captures the distinct kinds:
1. Comparative: "Men are taller than women." Claims group A's distribution is shifted relative to group B's on some measure. Truth conditions are straightforward — compare the distributions. As you noted, can be true even with tiny absolute rates and massive overlap.
2. Statistical majority: "Most Americans own cars." Explicitly quantified. Truth requires >50%. Easy to evaluate.
3. Tendency/disposition: "Men tend to like sports." This sounds comparative but might not be. It could mean "more than women" (hidden comparative) or it could mean "more men like sports than don't" (majority claim) or just "it's a common trait among men" (something weaker). The ambiguity is part of why these are slippery. You should force yourself to ask which one is actually being claimed.
4. Bare generic characterization: "Cretans are liars." "Gen Z is lazy." This is the hard case, and it's the one you're actually wrestling with.
The linguistics and philosophy literature on generics is what you want here. Generics are sentences of the form "Ks are F" — "dogs bark," "mosquitoes carry malaria," "ducks lay eggs." The striking finding is that
there is no single threshold that makes a generic true. The required prevalence shifts depending on the type of property:
For
dangerous or striking properties, people accept generics at very low prevalence. "Mosquitoes carry malaria" is accepted as true even though a tiny fraction of mosquitoes actually carry the parasite. "Pit bulls are dangerous" gets accepted even if most pit bulls never bite anyone. The property being salient, harmful, or attention-grabbing lowers the threshold dramatically.
For
characteristic activities, a moderate threshold is expected. "Birds fly" feels true even though penguins, ostriches, and kiwis don't, because flying is seen as what birds characteristically
do.
For
mundane statistical properties, you basically need a majority. "Americans are right-handed" works. "Americans are left-handed" doesn't, even though millions are.
This maps directly onto your problem. "Cretans are liars" probably functions like the dangerous/striking type — it takes a relatively small number of salient, memorable encounters with Cretan dishonesty to make the generic feel "true" to speakers, because dishonesty is a morally charged, high-salience property. That's not a statistical claim about majority behavior. It's closer to: "the trait is sufficiently prevalent and salient that it's a
characteristic risk of the group."
This is both illuminating and troubling, because it means negative generics about groups can feel (and linguistically function as) "true" at very low actual prevalence, precisely because negative traits are salient. That's a built-in bias in how generics work — not a bias in the perceiver's psychology, but in the semantics of the sentence form itself."
There is also a possibility of representation (my Category 3 from earlier). Here is Claude's more full analysis, though it is more terse than above, which is why I pasted the above first. I would note as well that r = 0.5 in Jussim is remarkably high for social science but in absolute terms leaves plenty of variance still. All who are nerds like me, enjoy!
Taxonomy of Group-Level Generalizations
The Core Question
When can a general statement be made about a group, and what would make that statement justified? The answer depends on what
kind of statement is being made. Different forms of generalization have different truth conditions, different evidence requirements, and different failure modes.
Type 1: Comparative Statements
Form: "Group A is more X than Group B."
Examples: "Men are taller than women." "Blacks are more likely to be incarcerated than whites." "Scandinavian countries have higher social trust than Mediterranean countries."
Truth conditions: Group A's distribution on some measure is shifted relative to Group B's. This is straightforward to evaluate empirically—compare the distributions.
Mathematical framing: If μ_A > μ_B for some measured trait, the comparative statement is true regardless of the absolute values or the degree of overlap between distributions. The
effect size (often expressed as Cohen's d = (μ_A − μ_B) / σ_pooled) captures how large the shift is. For most psychological and behavioral traits, d is in the range 0.2–0.5 (small to moderate), meaning the distributions overlap enormously. Even for a moderate effect size of d = 0.5, roughly 69% of the higher group exceeds the other group's mean—real, but leaving massive individual-level overlap.
Key feature: Can be true even when the absolute rates are tiny (e.g., 5% vs 1%), and even when the vast majority of individuals in both groups are indistinguishable. The statement is about distributional shift, not about what is typical of individuals.
What it does NOT tell you: Whether any given individual from Group A will score higher than a given individual from Group B. For d = 0.5, if you randomly draw one person from each group, the person from the "higher" group scores higher only about 64% of the time.
Type 2: Statistical Majority Claims
Form: "Most Ks are F." / Explicitly quantified claims.
Examples: "Most Americans own cars." "The majority of Norwegians speak English."
Truth conditions: >50% of the group has the trait (or whatever threshold is explicitly stated).
Mathematical framing: P(trait | group member) > 0.5. Straightforward to evaluate by measuring the rate.
Key feature: The simplest and least ambiguous type. Also the rarest in everyday speech—people rarely make explicitly quantified group claims.
Type 3: Tendency/Disposition Claims
Form: "Ks tend to be F." / "Ks are generally F."
Examples: "Men tend to like sports." "Women are generally more empathetic." "Old people tend to be conservative."
Truth conditions: Ambiguous. Could mean any of:
- Hidden comparative: "more than some reference group" (→ Type 1)
- Majority claim: "more of them are F than not" (→ Type 2)
- Something weaker: "it is common among them" (no fixed threshold)
Key feature: The ambiguity is the problem. When evaluating a tendency claim, the first step is forcing precision:
which of the above is actually being asserted? Different evidence applies to each interpretation. Much everyday disagreement about stereotypes comes from people interpreting the same tendency claim differently.
Type 4: Bare Generic Characterizations (Distributional)
Form: "Ks are F." (No quantifier, no comparison group stated.)
Examples: "Cretans are liars." "Gen Z is lazy." "Pit bulls are dangerous." "Birds fly." "Mosquitoes carry malaria."
Truth conditions: This is the hard case. The linguistics and philosophy literature on
generic sentences shows that there is no single prevalence threshold that makes a generic true. The threshold shifts depending on the type of property being attributed:
4a: Dangerous or Striking Properties — Very Low Threshold
Examples: "Mosquitoes carry malaria." "Sharks attack people." "Pit bulls are dangerous."
Accepted as true even at very low prevalence (a tiny fraction of mosquitoes carry malaria; most sharks never attack anyone). The property being harmful, frightening, or attention-grabbing lowers the required threshold dramatically.
Mathematical framing: Something like: P(harm | encounter with K) is low in absolute terms, but the
cost of the harm is high enough that P(harm | K) × Cost(harm) produces a salient expected-damage value. Alternatively, P(trait | K) > P(trait | not-K), and the trait is dangerous enough that even a modest elevation in probability warrants generalization.
Failure mode: This is precisely where negative stereotypes about human groups become most dangerous. Negative, frightening, or morally charged traits get generalized to groups at prevalence levels where positive or neutral traits would not. This is not a bias in the perceiver—it is a feature of how generic sentences
work semantically. The sentence form itself biases toward over-attribution of negative traits.
4b: Characteristic Activities — Moderate Threshold
Examples: "Birds fly." "Dogs bark." "Italians eat pasta."
The trait is seen as characteristic of what the kind
does—part of its nature or identity. Exceptions are tolerated (penguins don't fly, some dogs are mute) because the trait is understood as what the kind is
for or
about.
Mathematical framing: No clean threshold, but roughly: the trait should be modal (more common than any single alternative) and conceptually central to the kind. It need not characterize a majority if it is sufficiently identity-defining.
4c: Mundane Statistical Properties — Majority Threshold
Examples: "Americans are right-handed." (True—~90%.) "Americans are left-handed." (Feels false—~10%, despite millions.)
For properties that are not dangerous, striking, or identity-defining, a bare generic effectively requires a statistical majority to feel true. This is the closest to a simple >50% rule.
Mathematical framing: P(trait | group member) must be high enough that the trait is
typical—roughly >50%, and probably higher for the generic to feel natural.
Type 5: Structural/Representational Characterizations
Form: "Ks are F." (Same surface form as Type 4, but grounded differently.)
Examples: "America is imperialist." "The Catholic Church covered up abuse." "Sparta was militaristic." Possibly: "Cretans are liars" (if driven by Cretan cultural institutions, collective claims like the Zeus tomb story, or piracy as an organized practice).
Truth conditions: The trait is attributed to the group not because of individual prevalence, but because the group's
institutions, representatives, collective decisions, norms, or cultural identity manifest the trait. The group acts
as a group, and the characterization is about what the group does collectively.
Mathematical framing: Individual prevalence is largely irrelevant. P(trait | individual member) could be very low and the characterization still justified. The relevant question is not distributional but structural: Does the group's institutional or representative apparatus produce, incentivize, tolerate, or symbolize the trait?
Key features:
- Bypasses individual prevalence entirely. "America is interventionist" requires zero threshold of individual Americans being interventionist.
- Operates on a different axis from Types 1–4. Types 1–4 all concern the distribution of traits among individuals within the group, just with varying thresholds. Type 5 concerns the group as an agent or entity.
- Can coexist with Type 4 readings. "Cretans are liars" might be simultaneously a Type 4a generic (dangerous/striking trait, low threshold) and a Type 5 structural claim (Cretan institutions or collective cultural acts produced the reputation). Different evidence is relevant for each reading.
Evaluation criteria (distinct from Types 1–4):
- Are the group's institutions, laws, or norms structured in ways that produce or incentivize the trait?
- Do the group's recognized representatives or leaders exhibit the trait in their representative capacity?
- Does the group collectively identify with the trait or the actions that manifest it?
- Is there a mechanism by which the group acts as a unit in ways that exhibit the trait?
Cross-Cutting Considerations
The Diagnosticity Question
Regardless of type, a group-level claim becomes more informative (and arguably more "justified" as a basis for prediction) when the trait is
diagnostic—i.e., when group membership substantially shifts the probability relative to the base rate.
Mathematical framing: What matters is not P(trait | group) alone, but the likelihood ratio P(trait | group) / P(trait | not-group). A trait found in 3% of a group is highly diagnostic if it is found in 0.01% of everyone else, and barely diagnostic if it is found in 2.5% of everyone else.
Conditional Probability and Context
Even when a trait is rare in a group overall, it may be concentrated in subpopulations, regions, or contexts. If you are in such a context, the conditional probability P(trait | group member AND context) may be much higher than the unconditional P(trait | group member). This can make a generalization practically useful in specific situations while being misleading as a blanket statement.
The Gap Between Statistical Intuition and Generic Packaging
The stereotype accuracy literature (Jussim et al.) shows that people's
quantitative estimates of group differences are often roughly calibrated—their comparative intuitions track reality at moderate-to-high correlations (r ≈ .5 for personal stereotypes, higher for consensual ones). But stereotypes as actually held and deployed are typically bare generics (Type 4), not comparative statements (Type 1). The generic form suppresses within-group variation, erases base rates, and lowers thresholds for negative traits. So people can be decent intuitive statisticians
and deploy those intuitions in a sentence form that systematically overclaims.
Individual Prediction
Across nearly all studies, stereotype biases in person perception average r ≈ .10 when individuating information is available, and r ≈ .25 when it is not. This means that once you know something specific about an individual, their group membership adds almost nothing to prediction accuracy. Group-level generalizations, even when accurate at the group level, are weak tools for judging individuals.
Key Sources
Stereotype Accuracy
- Jussim, L. (2012). Social Perception and Social Reality: Why Accuracy Dominates Bias and Self-Fulfilling Prophecy. Oxford University Press. — The primary monograph arguing that stereotype accuracy is large and replicable.
- Jussim, L., Crawford, J.T., Anglin, S.M., et al. (2016). "Stereotype accuracy: One of the largest and most replicable effects in all of social psychology." In T. Nelson (Ed.), Handbook of Prejudice, Stereotyping, and Discrimination (2nd ed.), pp. 31–63. — The chapter making the title claim, reviewing 50+ studies.
- Jussim, L., Crawford, J.T., & Rubinstein, R.S. (2015). "Stereotype (In)Accuracy in Perceptions of Groups and Individuals." Current Directions in Psychological Science, 24(6), 490–497. — Shorter review addressing both group-level accuracy and individual-level application.
Generic Sentences (Philosophy of Language / Cognitive Science)
- Leslie, S.-J. (2007). "Generics and the Structure of the Mind." Philosophical Perspectives, 21, 375–403. — Foundational paper arguing that generics are a cognitively primitive form of generalization with context-dependent truth conditions.
- Leslie, S.-J. (2008). "Generics: Cognition and Acquisition." Philosophical Review, 117(1), 1–47. — Extends the framework to how children acquire generic concepts.
- Leslie, S.-J. (2017). "The Original Sin of Cognition: Fear, Prejudice, and Generalization." Journal of Philosophy, 114(8), 393–421. — Directly connects the generics framework to stereotyping and prejudice, arguing that the cognitive machinery for generics is biased toward over-generalizing dangerous properties. Highly relevant to this taxonomy.
- Prasada, S. & Dillingham, E. (2006). "Principled and statistical connections in common sense conception." Cognition, 99(1), 73–112. — Distinguishes between properties connected to kinds by principle vs. by statistical association.
National Character Stereotypes (A Notable Exception to Accuracy)
- McCrae, R.R., Chan, W., Jussim, L., et al. (2013). "The Inaccuracy of National Character Stereotypes." Journal of Research in Personality, 47(6), 831–842. — Shows that national personality stereotypes specifically do not track measured personality differences well.