N13 THE REALITY LAYER
Why Scientific Benchmarks Miss the Real Work
Research is a continuous workflow. Most benchmarks inspect isolated moments.
IN THIS NOTE · JUNE 2025
A benchmark asks a model to produce an answer under controlled conditions. A researcher asks a system to help navigate ambiguity over time. The gap between those tasks is where many scientific AI products will succeed or fail.
Answers are not workflows
Real research involves defining the question, finding sources, inspecting data, choosing methods, writing code, interpreting outputs and deciding what to do next. Errors can enter at any transition. A final-answer score compresses this chain and gives little information about where the system is dependable.
A system may know facts while failing to preserve context. It may write correct code while misunderstanding a biological field. It may reach the right answer once and fail to reproduce it.
Evaluate collaboration
Useful evaluation should measure plan quality, source selection, tool use, uncertainty, dialogue and continuity across sessions. It should test whether the system asks for missing information and whether a researcher can redirect it without restarting the work.
Repeated runs matter because consistency separates a robust workflow from a lucky completion. Failure analysis matters because two systems with the same score can create radically different risks.
Benchmarks should predict utility
The benchmark is valuable when performance transfers into the user's work. That requires representative tasks, realistic tools, transparent scoring and evidence that improvements change decisions or save rigorous labor.
Scientific AI does not need easier benchmarks. It needs evaluations that are harder in the same way science is hard.
A benchmark compresses the world
A benchmark is valuable because it removes context. Every system sees the same inputs, target and scoring rule, which makes comparison possible. Research is difficult for the opposite reason: the relevant context is incomplete, the target may change as evidence arrives and several answers can be defensible under different assumptions. A high score demonstrates capability under the benchmark's compression. It does not demonstrate that the system will notice which omitted variable controls the real decision.
The compression can also reward the wrong behavior. If every task has a known answer, confident completion looks better than productive uncertainty. If sources are clean and tools always work, the system never has to recover from a broken identifier, contradictory paper or unit mismatch. If evaluation ends with the first answer, correction and collaboration disappear from view. These omissions matter because scientific harm often enters through an apparently reasonable step that no one revisits.
Evaluate a portfolio of research behaviors
A stronger evaluation combines components and trajectories. Components test retrieval, calculation, code, data interpretation and citation. Trajectories test whether the system can plan, preserve project state, ask for missing context, use tools, respond to criticism and update a conclusion without rewriting history. Domain experts should score not only correctness, but whether the evidence supports the action the system recommends.
The hardest test is counterfactual: insert a plausible but wrong premise, an attractive confounder or a source that does not say what its title suggests. Observe whether the system amplifies the mistake, marks uncertainty or designs a discriminating check. NIST frames evaluation as part of a continuous govern-map-measure-manage cycle. Scientific systems need the same idea: evaluation is not a launch gate passed once, but an operating process that follows changing tools, data and use cases.
- Evaluate plans, transitions and recovery, not only final answers.
- Run repeated trials and inspect variance.
- Tie benchmark gains to a real researcher workflow.
I would change this view if isolated answer accuracy reliably predicted sustained performance in complex research projects.
Primary and institutional sources used as the grounding layer. Interpretation and synthesis are Luca's.
01