What this benchmark measures
Realm tests one job: read a real anatomic pathology report and produce the structured interpretation a tumor board can act on. Reading the report is the easy part. Success is defined by what the model does once it has read the report correctly. It must reach a diagnosis, stage or risk group only when the specimen supports one, reconcile the uncertainties the report leaves open rather than resolving them by assertion, recommend only the management the report justifies, and read each finding for its clinical meaning rather than its literal text. Every finding below ties back to those axes.
Model scoreboard
We ran the dataset of expert authored pathology tasks against eleven frontier models, three independent rollouts per task. Scores are the mean of task means: for each task the mean verifier reward over three rollouts, averaged across the dataset. Reward is the share of positive rubric weight a model earns, after penalties, floored at zero.
Mean of per task mean reward across the dataset, three rollouts per task. Bars are scaled to 100%.
What changed in this release
This release retargets the benchmark. The dataset and the expert answers are unchanged, but every task was reframed and every rubric rebuilt around four clinical failure modes that clinicians report in real operational use of these models:
- stating a definitive diagnosis on insufficient evidence,
- Recommending treatment that is inaccurate or not justified,
- Failing to reconcile uncertainty and name the steps to verify a finding, and
- Reading the report too literally, missing what the findings mean.
Two mechanical changes follow. Prompts now permit a model to declare what cannot be determined, ask for the steps needed to verify a finding where confirmation is required, and ask the model to weigh further testing or treatment against its actual benefit. Rubrics now carry 299 penalty criteria alongside 772 positive ones: a penalty fires when a summary invents a finding, misclassifies the case, or recommends beyond what the report supports, and it subtracts from the score. The earlier rubrics had no penalties, so these results are a new comparison identity and are not continuous with the earlier release.
Headline findings
Three findings shape how these results should be read. Each is stated plainly first, then unpacked in its own section below.
1
Claude Fable 5.1 leads, but the top of the field is a cluster and the ranking is a penalty story.
Fable reaches 0.780, only 0.010 above Claude Opus 5.5 at 0.770, and that gap is not separated: the task paired interval runs from minus 0.013 to plus 0.033. Muse Spark 1.3 at 0.758 is also within noise of the leader, and with Claude Opus 5 at 0.752 and Muse Spark 1.1 at 0.749 the top five sit within 0.031. The field spans 0.117 from first to last. On positive criteria alone the order would be Fable, Opus 5, Opus 5.5, Muse 1.1, Muse 1.3, Sonnet, Gemini, Grok 4.7, Grok 4.6, Sol, Astra; the penalties reorder it and drop Gemini from seventh to last. Reading well is necessary; not overreaching is what sorts the pack.
2
Reading the report is largely solved; saying what it means is not.
Pooled across models, extraction criteria, which ask the model to restate a finding printed in the report, are met 93% of the time. Interpretation criteria are met 73% of the time, the lowest of the five reasoning types, and because interpretation carries 30% of the positive weight it is also where most points are lost: 43 to 52% of every model failed weight. The reward is now won or lost after the report has been read, at the point where the model decides what the findings mean and whether to stop.
3
The costly errors are confident additions, and extra effort buys little.
Overreach, asserting a diagnosis, classification or recommendation the report does not support, is the largest penalty category, 123 of the 299 penalty criteria. Gemini 3.8 Flash trips a penalty in 29.8% of judged cases and gives back 0.159 of its earned weight; GPT-6 Astra trips one in 8.2% and gives back 0.044. Effort diverges far more than score: Claude Opus 5.5 makes 32.4 tool calls per task against 2.4 for Grok 4.6. Within a task, a summary twice as long scores only 0.029 higher while the penalty weight it trips rises by 3.1 points: the extra text is where the unsupported recommendations live.
Benchmark overview
Each task drops a model into a fresh sandbox with one or more real anatomic pathology reports and a generic toolkit: a shell, file read, write and edit, search over files, a todo list, and web search and web fetch for guideline lookups the report does not contain. The agent extracts the report, typically with pdftotext or a Python library, analyzes it, and writes a single structured interpretation to response.md. An LLM judge, then scores that answer against an expert authored rubric of weighted criteria, one call per criterion. Most criteria are positive, for example states the Grade Group as 4; a substantial set are penalties, for example recommends a specific procedure the report does not support, which subtract. The run reward is earned positive weight over total positive weight, floored at zero, so a summary that fabricates a result can score below what its correct content alone would earn.
Task taxonomy
The tasks are drawn from routine diagnostic pathology and span the organ systems and specimen types a working service sees. The mix is intentional: most tasks are single report interpretations, but a meaningful fraction are multi report cases, marrow workups that combine myelogram, flow cytometry, biopsy and cytogenetics, which require reconciling several modalities into one read. The corpus is unchanged from the previous release; the retarget is in the prompts and rubrics, not the reports.
Share of the dataset by category. Groupings are approximate; some multi organ cases are filed under their dominant specimen. Roughly four in five tasks are single report interpretations; the remainder are multi report integrations.
Performance summary
Best of 3 takes the highest of three rollouts per task before averaging. Pass@3 is the probability at least one of three attempts scores 0.8 or more. Perfect is the share of runs at exactly 1.0. Tasks won is the share of the dataset on which the model has the single best per task mean; ties are not counted.
Paired differences
With eleven models the full pairwise grid is large, so the table shows Claude Fable 5.1 against each other model; the clusters behind the leader are described below. Separated means the bootstrap interval over tasks excludes zero.
Fable separates from every model except Claude Opus 5.5 and Muse Spark 1.3. The top of the field is a chain rather than a single leader: Opus 5.5 is itself not separated from Muse Spark 1.3, Claude Opus 5 or Muse Spark 1.1. In the middle, adjacent models from Grok 4.7 to GPT-6 Sol are within noise of each other, though the ends of that group separate. Gemini 3.8 Flash is last and separated from every model except GPT-6 Sol. Won, lost and tied are shares of the dataset.
Where the points go
Every positive criterion was classified by the kind of work it tests: Extraction restates a finding written in the report, Classification reaches the right diagnosis, grade, stage or risk group under a named system, Interpretation explains what the findings mean and their limits, Management recommends the next step, and Communication completes the form, states uncertainty and stays in scope. The classification is not cosmetic: it shows where the points are lost, and it is how the retarget is enforced, since the weight the failure modes concern sits in interpretation, management and the uncertainty facing part of communication.
Weighted pass rate on positive criteria, by type, per model, over all runs. Criteria per type: Extraction 328 (36% of positive weight), Classification 118 (18%), Interpretation 200 (30%), Management 93 (11%), Communication 33 (4%). Extraction is near saturated; Interpretation has the lowest pass rate and, carrying 30% of the weight, is where the most reward is lost.
Report knowledge versus guideline knowledge
The gap widens when a criterion needs knowledge the report does not print, such as a staging table or a guideline threshold, rather than a fact stated on the page.
Weighted pass rate on positive criteria, split by whether satisfying the criterion needs knowledge not printed in the report and by whether it hinges on a specific number.
Hardest tasks: where every model struggled
These are the lowest reward tasks in the benchmark. The pattern is consistent with the second finding: each one punishes saying more than the specimen supports. The anchor column is the criterion most responsible for lost reward, read it as the thing the report did not let you say.
Lowest per task mean reward across models, three rollouts each.
Largest model spreads: where behavior separated most
These tasks produced the widest gap between the strongest and weakest model on the same problem, the most informative cases for understanding what distinguishes a good pathology read from a weak one.
Highest per task spread between the best and weakest model, excluding tasks already listed among the hardest. Several of these are the cases the failure mode examples below draw on.
Where models lose reward
The four failure modes that define the retarget are below, each with a complete example: the task, the exact rubric criterion, the model response, the judge verdict and the judge justification, so the failure is legible rather than asserted. No model is immune, and the examples below are drawn from the top of the field, where the errors are subtle enough to survive an otherwise strong summary.
Reading effort, not just score
Effort diverges far more than score. The table below is the mean per task effort of each model. The models agree on which tasks are hard, the rank correlation of per task means between any two models is 0.57 to 0.87, and repeats are stable, with a within task standard deviation of 0.045 for GPT-6 Astra to 0.067 for Claude Sonnet 5. What they do not share is how much work they spend reaching the answer.
Mean per task effort. Claude Opus 5 writes 2,525 words to GPT-6 Sol 456; a run costs about 0.13 dollars on Muse Spark 1.1 and 2.95 dollars on Claude Opus 5.5 (GPT-6 Sol was not yet priced on the platform when it ran). Summary length barely tracks reward: the per model correlation between words and score runs from +0.18 to -0.22.
What this means for deploying models on pathology workflows
Realm is hard because it is concrete: the rubric grades exact figures, preserved limits, and the absence of unsupported escalation, with no credit for sounding authoritative. Three takeaways follow.
1
Treat extraction as reliable and put human review where judgment happens.
All eleven models read messy reports and produce coherent structured summaries; extraction criteria are met 93% of the time. The reward is lost downstream, at the point where the model decides whether to stop at the report. Human review is most valuable there, not on the extraction layer.
2
The dangerous errors are confident additions and resolved uncertainty, not missing facts.
Overreach is the largest penalty category, 123 of the 299, and the penalties reorder the field: on positive criteria alone the order would be Fable, Opus 5, Opus 5.5, Muse 1.1, Muse 1.3, Sonnet, Gemini, Grok 4.7, Grok 4.6, Sol, Astra. A reviewer should check that every stage, biomarker and recommendation in the output is present in the source report, and that every stated certainty is one the report actually supports.
3
Score behavior, and do not reward effort.
The models cluster tightly on headline score, and the separations that exist are penalty separations: Gemini 3.8 Flash gives back 0.159 of its earned weight to penalties while GPT-6 Astra gives back only 0.044. The most process heavy agent in the field, Claude Opus 5.5, costs more than twice as much per run as the leader without beating it. Teams evaluating models for this work should read how an answer was produced and treat length as a warning sign, not a proxy for care.
Methodology note
The dataset of expert authored pathology tasks, most single report interpretations and a smaller set of multi report integrations, was reauthored and signed off by domain reviewers over 17 to 20 September 2026. Eleven frontier models: Claude Fable 5.1 (max), Claude Opus 5.5 (max), Muse Spark 1.3 (xhigh), Claude Opus 5 (max), Muse Spark 1.1 (xhigh), Grok 4.7 (xhigh), Grok 4.6 (high), Claude Sonnet 5 (default), GPT-6 Astra and GPT-6 Sol (max, with the subagent task tool denied), and Gemini 3.8 Flash (high); three rollouts per task. The eight launch models ran in Daytona sandboxes; Grok 4.7, GPT-6 Sol and Claude Opus 5.5, released the following week, ran on the identical pinned task versions in Modal sandboxes. Twelve repetitions that failed for infrastructure reasons, a provider stream timeout or a harness idle stall with no answer written, were rerun once with the same configuration and matched by task, never selected by score; there were no exclusions. Reports are anonymized. Agents run in a sandboxed OpenCode container with the report PDFs available and a generic toolkit of a shell, file read, write and edit, search over files, a todo list, and web search and web fetch; each writes a single structured interpretation, which the verifier grades. One judge, GPT-5.4 mini, snapshot of 17 March 2026 at medium effort, grades every criterion against a hidden expert answer, one call per criterion. Rubrics carry 1,071 criteria, 772 positive and 299 penalty (overreach 123, misclassification 105, fabrication 53, omission 10, other 8), 12 to 29 per task. Reward is earned positive weight over total positive weight, floored at zero; a penalty subtracts its weight when the behavior is present. Because the earlier rubrics carried no penalties, these results are a new comparison identity and are not continuous with the earlier release.
%20(1).webp)
.avif)