N23 THE REALITY LAYER
What 90% Accuracy Does Not Tell You
The score is a beginning. The error distribution is the operating manual.
IN THIS NOTE · JANUARY 2026
A high benchmark score can establish progress. It cannot tell a researcher whether the system will fail on the one semantic distinction that controls the next six months of work.
Aggregate scores hide structure
Accuracy compresses easy and hard tasks, stable and unstable runs, benign and consequential mistakes. Two systems can reach the same number while one fails visibly and the other produces confident, internally consistent errors.
The valuable questions concern distribution: which tasks fail, why they fail and whether the system knows when it is outside its competence.
Repetition exposes robustness
A correct result across repeated runs carries different evidence than one correct result among several contradictions. Scientific work depends on continuity, so variance matters even when average performance looks strong.
Evaluation should preserve the trace. Plans, sources, code and intermediate outputs allow users to diagnose whether the answer emerged from a defensible process.
Failure analysis is a product feature
Publishing a failure teaches users how to supervise the system. It also directs engineering toward semantic checks, tool constraints or interface changes that a larger model alone may not solve.
The score attracts attention. The failure analysis earns trust.
Accuracy hides prevalence and consequence
A model can achieve high accuracy by succeeding on common, easy cases while failing on the rare cases that matter. The same headline number can describe very different sensitivity, specificity and calibration depending on class prevalence. In science and health, the cost of a false negative may differ radically from the cost of a false positive, and both costs may change across the workflow.
Evaluation should begin with the decision. What action follows a positive output, what happens after a negative one, and which cases are referred to a human or a second test? The relevant operating point depends on those consequences. A ranking tool for literature triage can tolerate different errors from a system selecting compounds for synthesis or interpreting patient-facing information. Without that context, accuracy is a property of a dataset, not evidence of usefulness.
Test calibration, shift and abstention
A trustworthy probability should correspond to observed frequency under the conditions where it is used. Calibration reveals whether a system's confidence is informative, while subgroup analysis reveals whether aggregate performance hides weak regions. Distribution shift then asks whether those relationships survive new instruments, populations, protocols or time. These tests matter because scientific deployment rarely resembles a frozen benchmark indefinitely.
Abstention is an important output. A system that recognizes unfamiliar inputs and routes them for review can create more value than one forced to answer every case. The evaluation should reward appropriate uncertainty, not only coverage. Monitoring must continue after deployment, with thresholds for investigation and rollback. Ninety percent becomes meaningful only after the denominator, error costs, confidence behavior and recovery process are visible.
- Inspect error classes and consequences, not only averages.
- Test repeated-run stability.
- Publish representative failures with corrective controls.
I would reconsider if one aggregate accuracy score became sufficient to predict reliability across scientific contexts and repeated use.
Primary and institutional sources used as the grounding layer. Interpretation and synthesis are Luca's.
01