How faulty statistics arise
Faulty or misleading statistics are not always produced with fraudulent intent. Often the causes are insufficient methodological knowledge, unclear questions, unsuitable data or careless presentation. In other cases, however, numbers are deliberately selected or presented in order to create a desired impression.
Common sources of error
- Unsuitable sample: The group studied is too small or not representative.
- Unclear denominator: Percentages are given without stating the underlying number of cases.
- Misleading charts: Axis scaling, cropping or aspect ratio makes differences appear larger or smaller than they are.
- Missing context: Time period, comparison group or important additional information is omitted.
- Systematic measurement error: The measurement method regularly biases results in the same direction.
- Correlation instead of causation: Two quantities change together without one causing the other.
- Invalid generalisation: Results from a limited study are transferred to other groups, periods or situations.
- Uncertain extrapolation: Short-term trends are projected too far into the future in economic, climate or population forecasts.
- Selective publication: Convenient results are shown while contradictory results are omitted.
Correlation is not causation
A statistical association alone does not establish a causal relationship. Even if beer consumption and the results of an educational study differ between countries, it does not follow that beer consumption causes educational performance. Both quantities may depend on further factors or may simply happen to occur together.
Questions for doubtful statistics
- Who conducted and financed the study?
- How were the data collected?
- How large and representative is the sample?
- What information or comparison values are missing?
- Is causation being inferred improperly from correlation?
- What does the result actually say — and what does it not say?