Reassessment of Gregor Mendel's Pea-Plant Data Uncovers That Initial Allegation of Statistical Flawlessness Was Founded on Deficient Evaluation

Reassessment of Gregor Mendel’s Pea-Plant Data Uncovers That Initial Allegation of Statistical Flawlessness Was Founded on Deficient Evaluation

The experiments conducted by Gregor Mendel on pea plants form the basis of nearly all genetics textbooks to this day: cross a tall plant with a short one, tally the ratios in subsequent generations, and observe the numbers falling into the tidy 3:1 ratio that supplied the world with the fundamental principles of inheritance. For 83 years, that basis came with an asterisk. Critics argued that the numbers were too perfect to be accurate.

This claim originates from a singular paper published in 1936 by statistician Ronald A. Fisher, titled “Has Mendel’s Work Been Rediscovered?”, which appeared in the journal Annals of Science. Fisher performed a statistical analysis on Mendel’s reported findings and determined that they corresponded with the theoretically anticipated ratios more closely than what chance would allow. His verdict was straightforward: somewhere along the line, whether intentionally or not, Mendel’s data had probably been altered, either by Mendel himself or by an assistant eager to satisfy him.

How a 2019 paper reignited the debate

A re-evaluation published in October 2019 in the journal Hereditas, spearheaded by Noel Ellis of the John Innes Centre along with Julie Hofer, Martin Swain, and Peter van Dijk, revisited Fisher’s own statistical approach rather than relying on Mendel’s raw figures. Their discovery did not provide evidence that Mendel had deceived anyone, nor did it furnish proof that he had not. Instead, it revealed a flaw within Fisher’s test itself.

Fisher had amalgamated results from multiple separate experiments conducted by Mendel into a single overall chi-squared statistic, treating them as if they formed one cohesive dataset. Ellis and his collaborators identified various ways in which this methodology failed. Some of the characteristics Mendel monitored are genetically linked, indicating that they do not assort independently as Fisher’s model assumed. Mendel’s experiments were not executed as a singular randomized design; the groups of seeds, environmental conditions, and the sequence in which crosses were documented all varied in manners that Fisher’s collective test did not address. Survival rates of seedlings were inconsistent among experiments, subtly affecting which plants reached the stage of being counted. Additionally, there exists a reasonable likelihood that some phenotypes, particularly subtle variations in color or shape, were misclassified during the counting process, either way.

None of these individual issues would necessarily lead to Fisher’s “too good” result. Collectively, as stated in the 2019 paper, they were sufficient to undermine the statistical foundation for his conclusion. As the authors articulated, their reanalysis “does not support previous suggestions that they differ remarkably from expectation.”

Two statisticians, two distinct standards of evidence

The intriguing aspect of this narrative is not solely that Fisher erred. It is the stark contrast in how the two analyses employed the same fundamental tool, a chi-squared test, to derive opposing conclusions from data that exceeds a century old.

Fisher’s method pooled results from various experiments to enhance statistical power, which is a reasonable instinct if the underlying experiments are genuinely comparable. The contribution of the 2019 paper lay in demonstrating that they were not sufficiently comparable for that pooling to be considered safe, once the interaction of trait linkage, batch effects, and classification noise were analyzed individually rather than aggregated.

This constitutes one re-evaluation, not a second Fisher-scale consensus in its own right, and the authors are cautious about how far they extend their findings. They do not assert to have confirmed that Mendel’s data was authentic, in the sense of eliminating the possibility of selective reporting. What they illustrate is that the specific statistical argument constructed against Mendel in 1936 does not withstand the test of modern critique regarding its own assumptions. This is a narrower assertion than “Mendel is vindicated,” yet it presents a more captivating argument: an 84-year-old accusation was ultimately revealed to be based on erroneous mathematics, not on a flaw in the individual being scrutinized.

Why the “too good” instinct is a pitfall

Fisher’s initial suspicion stemmed from a rational place. Data that aligns unusually well with a theoretical prediction generally warrants a closer examination; both fraud and unconscious bias in scientific data collection are genuine issues, not mere conspiracy theories. The complication arises because “too good to be true” is inherently a statistical judgment, heavily reliant on how the comparison is structured. If several distinct experiments are perceived as a singular vast dataset, apparent precision can seem suspect. Conversely, if they are regarded as what they truly were—a collection of separate, imperfect trials conducted under slightly varying conditions…