Base Rate Neglect
Ignoring general statistics in favor of specific information
What is it?
Base rate neglect (or base rate fallacy) is the tendency to ignore general statistical information (base rates) in favor of specific case information, even when the base rate is more predictive. In Kahneman and Tversky's classic "lawyer-engineer problem," people given a personality description largely ignored whether the person was drawn from a pool of 70% lawyers or 70% engineers—the description dominated their judgment. We are drawn to stories, not statistics. This creates systematic errors. When evaluating a job candidate's impressive interview, we neglect how often strong interviewees turn out to be ordinary performers. When excited about a startup's compelling pitch, we neglect that most startups fail. Medical diagnosis is particularly vulnerable: a positive test result feels definitive, but interpretation requires knowing the base rate of the disease and false positive rate of the test. Base rate neglect leads to overconfidence in individual predictions and underuse of statistical regularities. Correcting it requires consciously asking "what's the base rate?" before being swayed by individual-case information, using Bayesian reasoning that combines base rates with case-specific evidence, and creating decision processes that force consideration of relevant statistics.
Example
Investing in a startup because the founder is charismatic, ignoring that most startups fail. Believing a positive medical test is definitive without considering false positive rates. Hiring after one brilliant interview without asking how often such hires work out.
References
Kahneman, D., & Tversky, A. (1973). On the Psychology of Prediction. Psychological Review, 80(4), 237-251.
Bar-Hillel, M. (1980). The Base-Rate Fallacy in Probability Judgments. Acta Psychologica, 44(3), 211-233.
Tversky, A., & Kahneman, D. (1974). Judgment Under Uncertainty: Heuristics and Biases. Science, 185(4157), 1124-1131.
How to Prevent It
Doxa uses AI and can make mistakes. How it's built
What is the base rate for this outcome in general?
Am I overweighting this specific case over statistics?
Out of 100 cases like this one, how many end the way I expect?
Am I being swayed by vivid details over statistics?
What would a statistician say about this probability?
Always start with base rate data before considering specifics.
Use Bayesian reasoning: combine prior probability with new evidence.
Pick the right reference class: compare with cases truly like yours (same sector, size, stage).
Turn percentages into counts: "of 1,000 people tested, how many are ill and how many are false alarms?"
Treat case-specific information as updating, not replacing base rates.