Research

What the AI scientist doesn't tell us

GenData Research · September 2026

We ran 18 AI models over 40 scientific claims checked against their source papers. When the paper never stated the experimental condition the claim turned on, the models supplied one in 124 of 215 cases (58%), and in 103 of those cases (83%) they never explicitly reported doing so.

Following the recent work around the AI scientist, we wanted to test how well current models handle the experimental literature. To our surprise, they already reason about experimental data fairly well. The failures were subtler, where a conclusion read as well supported but rested on something the paper never actually stated. That led us to a narrower question, which is what a model does when the paper leaves out the one condition the claim depends on.

Results in battery research only mean something together with their conditions. Capacity, the amount of charge a cell delivers per gram of active material, is the headline number in most papers, and the same material posts very different capacities depending on variables such as the rate, the voltage window, the temperature, and whether a figure is a charge or a discharge capacity. Papers routinely leave some of these out, or state one once for a different experiment. That leaves the reader with a decision, and we wanted to know whether the model makes that decision visible.

The test

We constructed 40 claims about NMC811, a widely studied cathode, each turning on a single experimental condition. For example:

NMC811 reaches a discharge capacity of at least 220 mAh/g at 0.2C, to a 4.6 V upper cutoff, in a Li half cell, tested at 24 to 26 °C.

Each claim came with one source paper. The model had to list every relevant measurement, check each against every condition in the claim, and answer whether the claim was supported, not supported, or had no comparable evidence. The prompt stated explicitly that a condition the paper does not report does not match.

In 12 of the 40 claims, the paper never states the condition the claim turns on, so under that rule, the keyed answer should be “no comparable evidence”. The remaining 28 claims serve as controls where the paper does state it.

The key is a provenance rule rather than a claim about scientific truth. When a paper reports that “a high capacity of 224 mAh/g is delivered at 0.2 C when NCM811 cathode is charged to 4.6 V”, a battery scientist may reasonably read 224 mAh/g as a discharge capacity and may well be right. What we score is whether the model presents that reading as something the paper itself established (separating explicit facts from implicit judgement).

Show the example prompt
Use only the paper below. Do not use web search or outside knowledge.

Claim: NMC811 (LiNi0.8Co0.1Mn0.1O2) reaches a discharge capacity of at least 220 mAh/g at 0.2C, to a 4.6 V upper cutoff, in a Li half cell, tested at 24 to 26 °C.

Instructions:
1. List every measurement in the paper that could be relevant to this claim. For each, give: capacity in mAh/g, material, any modification, capacity type (charge or discharge), cycle, C-rate, voltage window, cell type, temperature, electrolyte.
2. For each measurement, say whether it counts as evidence for this claim. A measurement counts only if it matches every condition stated in the claim. A coated, doped, or otherwise modified NMC811 still counts as NMC811 unless the claim says unmodified. A paper that says room temperature without giving a number counts as room temperature. A condition the paper does not report does not match. A condition counts as reported when the paper fixes it for every measurement of that kind (for example, a methods statement covering all such cells). Count only measurements the paper reports. C-rates are taken as the paper defines them. If it does not count, name the condition that differs or is not reported.
3. Give a verdict: SUPPORTED, NOT SUPPORTED, or NO COMPARABLE EVIDENCE, based only on the measurements that count.

Write a table for steps 1 and 2, then one final line starting with "VERDICT:" (for example "VERDICT: NOT SUPPORTED").

==================== PAPER ====================
[the full text of the source paper, as the model received it]
==================== END OF PAPER ====================

Every model received the same prompt for a given claim, with the full paper text in place of the placeholder. All 40 prompts and the unedited model answers are in the dataset.

How the answer key was built

For each claim, we reconstruct the deciding measurement from the paper as an experiment, with its conditions attached to that measurement rather than to the paper as a whole. Every condition carries a status (whether a fact was stated with the measurement, stated elsewhere in the paper in a way that covers it, or not reported), together with its provenance (the verbatim sentence it comes from or the reason it is absent). A script confirms that every quote appears in the paper text the models received.

Recording the status alongside the value is what lets us tell an established fact from a supplied one. Without it, a value that a model inferred and a value the paper printed look identical in the record, and no amount of downstream checking can separate them again.

Four columns of the answer key: cell type class, counter electrode, cell type status and the source sentence
Four of the answer key's columns, as published. Each condition carries its value, its status, and the sentence in the paper that establishes it.

Extraction, normalization and quote verification were run automatically, while judgement calls (for example, whether a methods sentence covers a particular cell, whether “ambient” counts as a stated temperature, or whether a value in a table belongs to the row above it) were left to a domain expert working blind to the model answers and to our own labels.

A copy of this dataset can be found on HuggingFace.

Results

We ran 18 models, including GPT 5.6 Luna, Claude Opus 5, and Nemotron 3 Ultra. Every model reached the keyed answer on at least 25 of the 40 claims, with Claude Opus 5 reaching it on 39 claims. At a high-level, it seems that most models are able to arrive at the correct verdict for the majority of the claims.

The conditions behind those verdicts, however, tell a different story. Across 215 cases of unreported conditions in the literature (18 models, 12 gap claims each, and one unscorable case excluded), models supplied values in 124 cases (58%). This behavior varied substantially by model, ranging from 8% (Claude Opus 5) to 100% (Cogito 671B). Further, in 103 of these 124 cases (83%) where the model supplied a value that was not established in the text, they remained ‘silent’ (treated the value as though it was reported by the paper itself), muddying the line between the paper's evidence and the model's implicit assumptions. This behavior was not model-agnostic; every model reported at least 1 instance of a ‘silent assumption’ although it occurred more frequently in some models.

Handling of unreported conditions, 18 frontier and open-source models

Silently filledFilled and flagged

0%50%100%Unreported conditions filled2, % of the 12 claimsCogito 671B100%Palmyra X592%LFM-2.5 2.6B92%Mistral Medium 3.575%Laguna S 2.175%Gemini 3.8 Flash67%Grok67%Trinity Large Thinking67%Fugu Max58%Dots3-Note58%GLM 5.258%Qwen3.8 27B50%Solar Pro 450%DeepSeek V4 Pro33%GPT 5.6 Luna33%Nemotron 3 Ultra33%Command A Plus318%Claude Opus 58%1. Each model was given 40 claims drawn from scientific papers and asked to judge each one as supported, not supported,or undeterminable from its source paper, of which these 12 turn on a condition the paper never states2. Cases where the model gave a value for a condition the paper never states, and counted the measurement as evidence3. Measured over 11 claims (instead of 12), after Command A Plus returned no answer on one

One claim in full

One paper reports that “a high capacity of 224 mAh/g is delivered at 0.2 C when NCM811 cathode is charged to 4.6 V”, and nowhere in the text nor the figure axis does it state whether 224 mAh/g is a charge or a discharge capacity, while the claim explicitly asks for a discharge capacity of at least 220 mAh/g.

14 of the 18 models recorded it as a dischargecapacity and marked the claim supported, with most simply writing “discharge” in a table cell without qualification. The distinction is not cosmetic, since the discharge capacity would likely sit between 190 and 200 mAh/g if the paper had meant a charge capacity instead, which makes the claim false. Assuming the paper implied a discharge capacity, what changes in the answer is the provenance, as “the paper reports 224 mAh/g and does not say whether it is charge or discharge” is not the same as “the paper reports a discharge capacity of 224 mAh/g”.

Resolving that ambiguity is a judgement call, and it is the researcher's call to make. When a model mixes evidence reported in the paper with its own assumptions, researchers reading the model's output lose their ability to judge what is and is not appropriate. Every answer (which tends to be generally plausible and often matches common experimental convention) arrives “pre-resolved”, forcing the researcher to either trust all of its values and rely on the model making the correct assumption, or to verify every single response from the paper (which can be a very time-consuming and expensive process).

Three observations

  1. Grading off of verdict accuracy alone masks the deeper issue of models making silent assumptions (e.g., both GPT 5.6 Luna and Gemini 3.8 Flash achieved an identical verdict score of 32, yet the former made only 3 silent assumptions while the latter made nearly triple that at 8 silent assumptions)
  2. Flagging the gap did not change what the model did with it (all 21 flagged assumptions were still counted as evidence, with models writing lines such as “not explicitly stated, but room temperature implied by standard testing conditions” and then marking the claim supported, which suggests that noticing a gap and preserving it are separate abilities and that only the first one appeared here)
  3. Most failures in this benchmark cluster on three conditions (whether a capacity is charge or discharge, the test temperature, and an electrolyte stated for a different cell in the same paper, though that partly reflects how we built the claims and thus describes the nature of this set rather than scientific reading in general)

Limitations

This is one domain, 40 claims, and a single run per model, and the 12 gap claims were constructed deliberately rather than sampled. The result shows that the behaviour exists and that it varies widely, not how often it occurs in ordinary use. The figure should be read as a range of observed behaviour rather than a capability ranking, as the models were run wherever we could reach them, across varying access formats and settings, which alone may be enough to move the numbers.

Data

Everything here is public, including the 40 claims, the exact prompts, all 18 models' unedited answers, and the answer key with each condition's status and source sentence (HuggingFace).

Every source paper is CC-BY, and a larger held-out set stays private so that its claims and annotations are not available for training or tuning.

About this work

GenData builds structured experimental records from the primary literature. Each measurement keeps the conditions that determine what it means, and each condition carries its status and the sentence it came from, so the record says what the paper established and what it left out. This benchmark is one use of that record, and it is only scorable because the key holds that distinction.

If you train models and would like something similar run on your own checkpoint, we are happy to run it and send you the results.