p值是科学研究中的关键统计指标,但它的含义经常被误读。[1]p值衡量的是在不存在真实差异的假设下,由采样误差产生观测结果的概率。[1]例如,一项比较国家间差异的研究中,p值为0.06意味着如果各国之间没有真实差异,采样误差产生该差异大小的概率约为6%。[1]然而,许多研究人员错误地将其解释为研究结果为真实的概率,这是对p值本质的根本性误解。
科学界对p值的使用方式长期存在争议。[1]普遍采用的0.05阈值源自英国统计学家Ronald Fisher在1920年代提出的经验法则,其本身具有任意性。[1]当p值小于0.05时,结果被标记为"统计显著",但这种二分法造成了问题:p=0.049被视为发现,而p=0.051则被认为无效果,两者的实际差异微乎其微。[1]在p<0.05的情况下,即使没有真实效应,研究仍有约1/20的概率产生"显著"结果。[1]
面对这些问题,科学社群已开始采取行动。[1]2019年,超过800名科学家签署了一份声明,呼吁停止使用"统计显著性"这一表述。[1]至少一份学术期刊已完全禁止报告p值。[1]这些改革反映出研究界的共识:p值必须与效应量相结合使用,单独依赖p值作为判断研究价值的标准已不符合现代科学实践的要求。
The p-value remains one of the most frequently misinterpreted concepts in scientific research, despite its central role in determining statistical significance [1]. This metric represents the probability of obtaining observed experimental results due to sampling error alone, assuming no true difference actually exists—a definition often confused with the likelihood that findings are genuinely true [1].
The arbitrary nature of the 0.05 significance threshold has become increasingly problematic [1]. This benchmark originated from an empirical rule proposed by British statistician Ronald Fisher in the 1920s [1], yet its adoption as a hard cutoff has created perverse incentives in research. A p-value of 0.049 qualifies as a publishable discovery, while 0.051 suggests no effect—a binary distinction that lacks principled scientific justification [1]. When interpreted strictly, p<0.05 means only that if no true effect exists, roughly one in twenty studies would still produce "statistically significant" results by chance alone [1].
These concerns have prompted significant pushback from the scientific community. In 2019, over 800 scientists signed a declaration calling for the retirement of statistical significance language entirely [1]. The reform movement has already achieved tangible results, with at least one journal implementing a complete ban on p-value reporting [1]. Such actions reflect growing recognition that p-values must be paired with effect sizes to provide meaningful interpretation of research findings [1].