Two problems can undermine a pooled result even when the studies were similar enough to combine. The first is heterogeneity: the trial results disagree with each other by more than chance would explain. Statistical tests such as Cochran's Q and the I-squared statistic quantify this. I-squared is often read as the percentage of variability across trials that is due to real differences rather than sampling error; values above roughly 50 percent are commonly treated as substantial. Heterogeneity is not automatically fatal, but it changes what the summary means: a pooled estimate from highly heterogeneous trials is an average across effects that may differ, and the reasons for the disagreement — different doses, different populations, different outcome definitions — usually matter more than the average itself.
The second problem is publication bias. Studies with positive or statistically significant results are more likely to be published, and more likely to be published quickly, than studies with null results. If a review captures only the published literature, the pooled estimate can be shifted in the favorable direction. A funnel plot is a common visual check: it plots each study's effect against a measure of its precision, and in the absence of bias the points should form a roughly symmetric inverted funnel. Asymmetry — especially missing studies in the region where small null studies would fall — is a signal that publication bias may be present. Funnel-plot inspection is a judgment, not a definitive test, and it is unreliable when fewer than about ten studies are included.