What Is Heterogeneity (I²) in Meta-Analysis?
Published 2026-08-29 · 3 min read
Heterogeneity is the variability in true effects across the studies in a meta-analysis - the part of the disagreement between study results that chance alone cannot explain. Some heterogeneity is inevitable: included studies differ in populations, doses, co-interventions, outcome definitions, and follow-up. The statistical question is never whether studies differ, but whether they differ enough to change how you pool and interpret them.
The I² statistic is the most widely reported summary of that variability. It describes the percentage of the observed variation in effect estimates that reflects real differences in effects rather than sampling error, ranging from 0% to 100% [2].
How is I² calculated?
I² is derived from Cochran's Q, the standard chi-squared heterogeneity statistic. With k studies, I² = 100% × (Q − df)/Q, where df = k − 1; negative values are set to zero [2]. Conceptually, Q measures total observed variation, df estimates the part expected from chance, and I² reports the remainder as a proportion.
Two properties follow directly. First, I² is a relative measure: it depends on the precision of the included studies, so a set of very large trials can show a high I² even when the absolute differences between their effects are clinically trivial [1]. Second, I² is not the between-study variance itself. That quantity is tau-squared, an absolute measure on the scale of the effect size, estimated by random-effects models [1].
How should you interpret I²?
The Cochrane Handbook offers overlapping bands as a rough guide: 0% to 40% might not be important; 30% to 60% may represent moderate heterogeneity; 50% to 90% may represent substantial heterogeneity; and 75% to 100% represents considerable heterogeneity [3]. The overlaps are deliberate - the thresholds are conventions, and interpretation should weigh the direction and magnitude of effects and the strength of evidence for heterogeneity, not the bare percentage [3].
With few studies, treat I² as unstable. Its uncertainty interval is wide at small study counts, and a pooled analysis of a handful of trials can swing from 0% to a high I² with the addition of one study [2]. Reporting the confidence interval around I², or focusing on tau-squared and prediction intervals, is more honest at low counts [3].
What I² is not
I² is not a measure of clinical importance, not a test of whether pooling is valid, and not a decision rule by itself. A meta-analysis with a low I² can still be misleading if the studies share a common bias, and one with a high I² can still be informative if the effects all point the same way. When heterogeneity is considerable, the productive responses are to check data extraction, explore prespecified subgroups or study-level moderators, and consider whether a single pooled average answers the review question at all [3].
Key takeaways
- Heterogeneity is variability in true effects; I² reports the share of observed variation attributed to it, from 0% to 100% [2].
- I² is relative to study precision; tau-squared is the absolute between-study variance [1].
- Cochrane's bands overlap by design: 0-40% might not matter, while 75-100% is considerable heterogeneity [3].
- I² is imprecise when the number of studies is small, so report its uncertainty or lean on prediction intervals [2].
- A high I² is a prompt to investigate, not an automatic reason to abandon pooling [3].
FAQ
Is 50% heterogeneity high?
A value of 50% sits in the overlap of Cochrane's moderate and substantial bands, so it cannot be labelled high or low on its own [3]. Look at whether effects agree in direction, how wide the prediction interval is, and whether prespecified subgroups explain the spread.
Does I² = 0% mean the studies agree?
Not necessarily. With few or small studies, the analysis has little power to detect real differences, and I² frequently lands at zero even when true effects vary [2]. Absence of detected heterogeneity is not evidence of homogeneity.
Should a high I² make me switch to a random-effects model?
Model choice should follow the review question and the expectation of varying true effects, decided in the protocol - not chased after seeing the statistic [3]. A random-effects model does not explain heterogeneity; it only incorporates it into wider intervals.
Sources
- Higgins JPT, Thompson SG. Quantifying heterogeneity in a meta-analysis. Statistics in Medicine 2002
- Higgins JPT, Thompson SG, Deeks JJ, Altman DG. Measuring inconsistency in meta-analyses. BMJ 2003
- Cochrane Handbook for Systematic Reviews of Interventions, Chapter 10: Analysing data and undertaking meta-analyses
Related articles
How Many Studies Do You Need for a Meta-Analysis?
The honest answer is two - but heterogeneity estimates, prediction intervals, and publication-bias tests each need more. What the methods literature actually says.
4 min read · 2026-08-29
Random-Effects vs Fixed-Effect Meta-Analysis: How to Choose
What the two models actually assume, how their weights and intervals differ, and how to choose in the protocol - not after seeing the heterogeneity statistic.
6 min read · 2026-08-29
How to Do a Meta-Analysis: A Step-by-Step Guide
The seven steps from research question to pooled estimate: PICO and protocol, search, dual screening, extraction, risk of bias, pooling, and GRADE.
7 min read · 2026-08-29