Statistics & critical thinking

When Statistics Mislead: From Honest Error to Deliberate Manipulation

Statistical distortion does not arise only from bad formulas or poor samples. Expectations, incentives, money, politics, selection and deliberate deception can also shape the statistical picture we see.

Statistical error in a wider sense

When we talk about statistical error, we usually think of mistakes inside an essentially honest scientific process: a poor sample, measurement error, an inappropriate model, confusing correlation with causation, or extrapolating too far. The people involved are trying to describe reality accurately, but they make mistakes.

That is only part of the problem. Numbers can create a false picture even when the arithmetic itself is correct. A researcher's expectations, career incentives, sponsors, political goals, ideology, selective publication or deliberate presentation can determine which numbers become visible in the first place.

The central idea: A statistical statement does not have to be false to be misleading. Sometimes every number quoted is correct — and the deception lies in the choice of numbers.

1. An honest scientist can still be biased

Scientists are human. A researcher who finds a hypothesis plausible may unconsciously make decisions that favour the expected result: Which observation counts as an outlier? Which subgroup is analysed? Which outcome is emphasised? When does data collection stop? Which of several defensible statistical methods is used?

Such choices are not automatically fraud. They become problematic when many analytical possibilities are tried and only the striking result is reported. A related practice is HARKing — “Hypothesizing After the Results are Known”: a hypothesis is developed after seeing the results but later presented as if it had been planned in advance. Norbert Kerr introduced the term in 1998.

This is an important grey zone between honest error and deliberate falsification. A researcher may sincerely believe that the analysis was careful while expectations still influence the selection of analyses and the interpretation of results.

2. When bias becomes research misconduct

A clearer boundary is crossed when data are invented, altered or deliberately omitted so that the research record no longer represents what actually happened. The U.S. Office of Research Integrity explicitly distinguishes fabrication, falsification and plagiarism from honest error.

For that reason, it is better not to use “bad scientists” as one broad category. There is a continuum from unconscious confirmation bias, through questionable research practices, to deliberate deception — and an accusation of intentional fraud requires evidence.

3. The sponsor: ten studies — and only one becomes visible?

Imagine a company wants to know whether its product is harmless or even beneficial. It funds ten studies. Nine find no benefit or contradict the preferred message; one happens to produce a favourable result. If that one study is the result that gets published, advertised or repeatedly cited to politicians and journalists, the public picture can be radically distorted.

Ten honest studies can produce a dishonest message if the public is shown only one of them.

The number “ten” is a thought experiment, not a claim about one specific case. But the underlying mechanisms — publication bias and selective reporting — are well documented. A Cochrane review found that clinical trials with positive findings were substantially more likely to be published than trials with negative or null results. Another Cochrane review found that industry-sponsored drug and device studies more often produced efficacy results and conclusions favourable to the sponsor's product than studies with other sources of funding.

A particularly clear example comes from an analysis of 74 antidepressant trials registered with the U.S. Food and Drug Administration (FDA). In the published literature, 94% of the published trials appeared positive. In the FDA's assessment, only 51% of all registered studies were positive. Twenty-two negative or questionable studies were not published, while another eleven were published in a way the authors judged to convey a positive outcome. Importantly, the investigators also stated that their data could not determine whether the distortion arose from authors, sponsors, journals, or some combination of them.

Tobacco: manufacturing doubt

Tobacco provides an unusually well-documented example because litigation made large collections of internal industry documents public. The World Health Organization describes strategies used to discredit research, fund science and influence policy. A study focusing on Germany grouped the mechanisms visible in internal documents into five categories: suppression, dilution, distraction, concealment and manipulation.

The aim does not always have to be to prove a clearly false counterclaim. Creating the impression that the science is still completely unsettled may be enough. If you can manufacture uncertainty, robust evidence can appear politically weaker than it really is.

Alcohol: conflicts of interest without assuming fraud

There is a similar debate about alcohol-industry involvement in science. A systematic review of perspectives within the alcohol research community discusses concerns including direct funding of university researchers and centres, contract research, support for scientific organisations, and efforts to influence public perceptions of research and alcohol policy. This does not prove that every industry-funded study is wrong. It is a reason to examine funding, study design, complete reporting and scientific independence particularly carefully.

4. A biased system does not need a fraudster

One of the most important points is easy to miss: the scientific literature can become biased even if nobody fabricates data. Researchers may be more inclined to submit surprising findings, journals may prefer positive results, careers and grants reward novelty, and null results may remain unpublished.

The result is a filter. Not every completed study has the same chance of becoming part of the visible literature. Someone who later reads only published papers may therefore see a systematically distorted body of evidence.

One response is the use of Registered Reports. The research question and methods are peer-reviewed before the results are known. In principle, the decision to publish therefore depends less on whether the eventual result is dramatic or “positive”.

5. Political and ideological use of statistics

Political manipulation does not require anyone to invent a number. Definitions, denominators, time windows or comparison countries can be selected so that they support a preferred story. Unemployment can look different under different definitions; public spending looks different as an absolute total and as a percentage of GDP; crime counts tell a different story from rates per 100,000 people; a trend can look very different depending on whether the graph starts in 2019 or 2020.

The quality of official statistics can itself become a political and institutional issue. In 2013 the International Monetary Fund issued a formal declaration of censure against Argentina over inaccurate CPI and GDP data; it lifted the censure in 2016 after methodological improvements. Eurostat published a detailed 2010 report on quality and methodological problems concerning Greek government deficit and debt statistics. These cases illustrate why independent and transparent statistical institutions matter. They do not, by themselves, prove a particular individual's intent; intent must be established separately from the fact that the data were inaccurate or unreliable.

6. Cherry-picking: every number is true, but the story is false

You do not need to falsify a statistic to mislead. You can choose:

  • the convenient start or end year,
  • a favourable subgroup rather than the whole population,
  • relative rather than absolute risk,
  • the mean rather than the median — or vice versa,
  • absolute counts rather than per-capita rates,
  • nominal amounts rather than inflation-adjusted values,
  • one favourable study from a larger contradictory literature.

Every individual statement may be correct. The manipulation then lies not in arithmetic but in the selection of the frame.

7. Manipulation by presentation

After selection comes presentation. A truncated y-axis can make a small difference look enormous. Percentage changes without the starting value can suggest a dramatic effect. A cumulative total can be made to look like a current rate. Three-dimensional bars and areas can distort visual proportions. A very long time axis can hide a short-term collapse; a very short one can make it look catastrophic.

A useful question is: What other, equally correct presentation of the same data would lead to a different immediate impression?

8. The statistic that was never collected

There is an even more basic question: which data are collected at all? What is not measured, classified or published can become almost invisible in public debate. Missing data need not be the result of a conspiracy; costs, poor reporting systems, inconsistent definitions and technical limitations are sufficient. But decisions about what an institution measures determine the boundaries of what can later be discussed as “objective” evidence.

Critical statistical thinking therefore asks not only “Is this number correct?” but also: “Which numbers are missing?”

How can we protect ourselves?

  • Check funding and conflicts of interest. Who benefits from a particular result?
  • Look for preregistration and study registries. They can reveal studies that were conducted but never published.
  • Compare primary outcomes and original analysis plans. Was the question changed after the results were known?
  • Search for null and contradictory findings. One study is rarely the whole evidence base.
  • Compare absolute and relative numbers.
  • Try alternative time windows, denominators and visualisations.
  • Separate methods and data from interpretation. A paper's conclusion may be stronger than its evidence.
  • Prefer independent replication, open data and transparent methods.

From calculation error to information architecture

In the narrow sense, statistical error is a problem of measurement or analysis. In the wider sense, distortion can enter at at least five levels:

  1. Measurement: What is measured, and how?
  2. Analysis: Which models, tests and subgroups are chosen?
  3. Selection: Which studies and results become visible?
  4. Communication: How are the numbers framed and visualised?
  5. Interest and intention: Which personal, commercial, political or ideological goals influence the earlier stages?

The central question therefore changes. It is no longer only “Is the statistic mathematically correct?” but also: “What process produced this particular statistic — and what alternative information am I not seeing?”

References and further reading

  1. U.S. Office of Research Integrity (ORI): Definition of Research Misconduct — Defines research misconduct and explicitly distinguishes fabrication and falsification from honest error.
  2. Lundh et al., Cochrane: Industry sponsorship and research outcome — Systematic review of the association between industry sponsorship and sponsor-favourable results or conclusions.
  3. Hopewell et al., Cochrane: Publication bias in clinical trials — Empirical evidence that positive trial findings are more likely to be published than negative or null findings.
  4. Turner et al. (2008): Selective Publication of Antidepressant Trials and Its Influence on Apparent Efficacy, New England Journal of Medicine — Comparison of 74 FDA-registered antidepressant trials with the published literature.
  5. World Health Organization: Tobacco industry interference with tobacco control — Overview of documented tobacco-industry strategies to influence science and policy.
  6. Grüning et al. (2006): Tobacco Industry Influence on Science and Scientists in Germany — Analysis of internal tobacco-industry documents with a particular focus on Germany.
  7. McCambridge et al. (2018): Alcohol industry involvement in science — Systematic review of concerns about conflicts of interest and industry involvement in alcohol research.
  8. Kerr (1998): HARKing — Hypothesizing After the Results are Known — Foundational paper introducing the concept of HARKing.
  9. International Monetary Fund (2013): Statement on Argentina — Formal censure concerning inaccurate CPI and GDP data.
  10. Eurostat (2010): Report on Greek government deficit and debt statistics — Report on quality and methodological problems in Greek deficit and debt statistics.
  11. Center for Open Science: Registered Reports — Publication model in which the research question and methods are reviewed before results are known.

The examples are intended to illustrate mechanisms. Documented data-quality problems or statistical distortion do not automatically prove deliberate deception by every person involved; intent requires separate evidence.