Study Exposes Critical Flaw in a Ubiquitous Scientific "Yardstick"

Study Exposes Critical Flaw in a Ubiquitous Scientific "Yardstick"

New research reveals the Normalized Mutual Information (NMI) metric is biased, potentially invalidating conclusions in thousands of studies on data classification and clustering algorithms.
gg
gizmo guru
Dec 16, 2025
1 min read

A new study reveals that Normalized Mutual Information (NMI), a cornerstone metric for evaluating data classification algorithms used in thousands of scientific papers, is systematically biased and can lead researchers to incorrect conclusions about which methods perform best.

Researchers from the University of Michigan and the University of Hong Kong identified two key flaws in the standard NMI calculation. First, it unfairly rewards algorithms that "over-cluster" by inventing more categories than truly exist, making them appear more accurate. Second, the common practice of normalizing the score to a 0-to-1 scale introduces an additional bias toward overly simplistic models. "Scientists use NMI as a kind of yardstick to compare algorithms," explained SFI Postdoctoral Fellow Max Jerdee. "But if the yardstick itself is bent, you might draw the wrong conclusion".

The team demonstrated that these biases are significant enough to alter scientific rankings. When testing popular community-detection algorithms, the choice of how to calculate NMI could point to different "best" algorithms, undermining reliable comparison.

To solve this, the researchers developed a corrected measure called the asymmetric reduced mutual information. This new metric eliminates both sources of bias by adjusting for the information content of the data and by normalizing scores against the ground truth alone, rather than symmetrically. The team's findings and their proposed solution were published in Nature Communications on December 11, 2025.

Professor Mark Newman, a co-author of the study, emphasized the scale of the issue, noting NMI "has been used or referenced in thousands of papers in the decades since it was first proposed". The correction aims to improve the reliability of algorithm evaluation across numerous fields, from medical diagnostics to network science.

About the Writer

More from Mindplex

Keep reading

Three more ideas worth your time.

Browse MindBytes

Discussion

Join the discussion

Sign in to share a response with the community.

Type @ to mention someone Type / or use + to add a block Highlight text, then choose Link
Loading editor

Comments cannot be edited after posting because they become part of the reputation record. Give yours a quick review first.