Back

No items found.

What this benchmark measures

Realm tests one job: read a real anatomic pathology report and produce the structured interpretation a tumor board can act on. Reading the report is the easy part. Success is defined by what the model does once it has read the report correctly. It must reach a diagnosis, stage or risk group only when the specimen supports one, reconcile the uncertainties the report leaves open rather than resolving them by assertion, recommend only the management the report justifies, and read each finding for its clinical meaning rather than its literal text. Every finding below ties back to those axes.

Model scoreboard

We ran the dataset of expert authored pathology tasks against eleven frontier models, three independent rollouts per task. Scores are the mean of task means: for each task the mean verifier reward over three rollouts, averaged across the dataset. Reward is the share of positive rubric weight a model earns, after penalties, floored at zero.

Claude Fable 5.1

78.0%

Muse Spark 1.3

75.8%

Claude Opus 5

75.2%

Muse Spark 1.1

74.9%

Grok 4.6

72.1%

Claude Sonnet 5

71.7%

GPT-6 Astra

70.3%

Gemini 3.8 Flash

66.3%

Mean of per task mean reward across the dataset, three rollouts per task. Bars are scaled to 100%.

What changed in this release

This release retargets the benchmark. The dataset and the expert answers are unchanged, but every task was reframed and every rubric rebuilt around four clinical failure modes that clinicians report in real operational use of these models:

  • stating a definitive diagnosis on insufficient evidence,
  • Recommending treatment that is inaccurate or not justified,
  • Failing to reconcile uncertainty and name the steps to verify a finding, and
  • Reading the report too literally, missing what the findings mean.

Two mechanical changes follow. Prompts now permit a model to declare what cannot be determined, ask for the steps needed to verify a finding where confirmation is required, and ask the model to weigh further testing or treatment against its actual benefit. Rubrics now carry 299 penalty criteria alongside 772 positive ones: a penalty fires when a summary invents a finding, misclassifies the case, or recommends beyond what the report supports, and it subtracts from the score. The earlier rubrics had no penalties, so these results are a new comparison identity and are not continuous with the earlier release.

Headline findings

Three findings shape how these results should be read. Each is stated plainly first, then unpacked in its own section below.

1

Claude Fable 5.1 leads, but the top of the field is a cluster and the ranking is a penalty story.

Fable reaches 0.780, only 0.010 above Claude Opus 5.5 at 0.770, and that gap is not separated: the task paired interval runs from minus 0.013 to plus 0.033. Muse Spark 1.3 at 0.758 is also within noise of the leader, and with Claude Opus 5 at 0.752 and Muse Spark 1.1 at 0.749 the top five sit within 0.031. The field spans 0.117 from first to last. On positive criteria alone the order would be Fable, Opus 5, Opus 5.5, Muse 1.1, Muse 1.3, Sonnet, Gemini, Grok 4.7, Grok 4.6, Sol, Astra; the penalties reorder it and drop Gemini from seventh to last. Reading well is necessary; not overreaching is what sorts the pack.

2

Reading the report is largely solved; saying what it means is not.

Pooled across models, extraction criteria, which ask the model to restate a finding printed in the report, are met 93% of the time. Interpretation criteria are met 73% of the time, the lowest of the five reasoning types, and because interpretation carries 30% of the positive weight it is also where most points are lost: 43 to 52% of every model failed weight. The reward is now won or lost after the report has been read, at the point where the model decides what the findings mean and whether to stop.

3

The costly errors are confident additions, and extra effort buys little.

Overreach, asserting a diagnosis, classification or recommendation the report does not support, is the largest penalty category, 123 of the 299 penalty criteria. Gemini 3.8 Flash trips a penalty in 29.8% of judged cases and gives back 0.159 of its earned weight; GPT-6 Astra trips one in 8.2% and gives back 0.044. Effort diverges far more than score: Claude Opus 5.5 makes 32.4 tool calls per task against 2.4 for Grok 4.6. Within a task, a summary twice as long scores only 0.029 higher while the penalty weight it trips rises by 3.1 points: the extra text is where the unsupported recommendations live.

Benchmark overview

Each task drops a model into a fresh sandbox with one or more real anatomic pathology reports and a generic toolkit: a shell, file read, write and edit, search over files, a todo list, and web search and web fetch for guideline lookups the report does not contain. The agent extracts the report, typically with pdftotext or a Python library, analyzes it, and writes a single structured interpretation to response.md. An LLM judge, then scores that answer against an expert authored rubric of weighted criteria, one call per criterion. Most criteria are positive, for example states the Grade Group as 4; a substantial set are penalties, for example recommends a specific procedure the report does not support, which subtract. The run reward is earned positive weight over total positive weight, floored at zero, so a summary that fabricates a result can score below what its correct content alone would earn.

Task taxonomy

The tasks are drawn from routine diagnostic pathology and span the organ systems and specimen types a working service sees. The mix is intentional: most tasks are single report interpretations, but a meaningful fraction are multi report cases, marrow workups that combine myelogram, flow cytometry, biopsy and cytogenetics, which require reconciling several modalities into one read. The corpus is unchanged from the previous release; the retarget is in the prompts and rubrics, not the reports.

Category Share What the agent must do
Hematopathology and bone marrow 18% Reconcile myelogram, flow, biopsy and cytogenetics; preserve differentials such as MDS versus aplastic anemia and AML MRD versus morphology
Breast 14% Grade, stage and report biomarkers without inferring receptor status or stage from a partial specimen
Thyroid, cytology and resection 14% Assign and hold the correct Bethesda category; avoid converting suspicion into diagnosis or surgery
Gastrointestinal 16% Report OLGA and OLGIM, Barrett esophagus and colorectal staging as stated; flag site discordance
Genitourinary, renal, prostate, bladder 16% Reconcile stage discrepancies, for example perirenal fat versus pT1b; grade prostate exactly
Dermatopathology 8% Hold qualifier language; do not assess margins on a punch biopsy
Gynecologic and cytology, cervical, HPV 8% Report molecular genotyping only; do not infer cytology from a PCR specimen
Thoracic and biomarker, IHC studies 6% Keep organizing pneumonia a pattern, not a diagnosis; report CPS exactly

Share of the dataset by category. Groupings are approximate; some multi organ cases are filed under their dominant specimen. Roughly four in five tasks are single report interpretations; the remainder are multi report integrations.

Performance summary

Model Effort Mean 95% CI Best of 3 Pass@3 Perfect Tasks won
Claude Fable 5.1 max 0.780 0.742 to 0.817 0.836 0.62 10% 14%
Claude Opus 5.5 max 0.770 0.731 to 0.807 0.825 0.66 5% 16%
Muse Spark 1.3 xhigh 0.758 0.714 to 0.799 0.819 0.62 9% 12%
Claude Opus 5 max 0.752 0.707 to 0.793 0.810 0.62 7% 12%
Muse Spark 1.1 xhigh 0.749 0.703 to 0.790 0.809 0.62 9% 6%
Grok 4.7 xhigh 0.731 0.684 to 0.774 0.793 0.56 8% 6%
Grok 4.6 high 0.721 0.670 to 0.769 0.792 0.56 9% 8%
Claude Sonnet 5 default 0.717 0.674 to 0.757 0.790 0.52 4% 8%
GPT-6 Astra max, task denied 0.703 0.660 to 0.744 0.753 0.38 4% 2%
GPT-6 Sol max, task denied 0.693 0.656 to 0.728 0.747 0.38 1% 6%
Gemini 3.8 Flash high 0.663 0.608 to 0.712 0.726 0.46 2% 0%

Best of 3 takes the highest of three rollouts per task before averaging. Pass@3 is the probability at least one of three attempts scores 0.8 or more. Perfect is the share of runs at exactly 1.0. Tasks won is the share of the dataset on which the model has the single best per task mean; ties are not counted.

Paired differences

With eleven models the full pairwise grid is large, so the table shows Claude Fable 5.1 against each other model; the clusters behind the leader are described below. Separated means the bootstrap interval over tasks excludes zero.

Comparison Mean difference 95% CI Separated Tasks won / lost / tied
Fable vs Opus 5.5 +1.0 pts -1.3 to +3.3 no 52 / 38 / 10%
Fable vs Muse 1.3 +2.2 pts -0.9 to +5.6 no 50 / 46 / 4%
Fable vs Opus 5 +2.8 pts +1.0 to +4.7 yes 60 / 38 / 2%
Fable vs Muse 1.1 +3.1 pts +0.4 to +5.8 yes 60 / 38 / 2%
Fable vs Grok 4.7 +4.9 pts +2.3 to +7.7 yes 56 / 40 / 4%
Fable vs Grok 4.6 +5.9 pts +2.5 to +9.3 yes 60 / 38 / 2%
Fable vs Sonnet +6.3 pts +3.2 to +9.4 yes 68 / 24 / 8%
Fable vs Astra +7.7 pts +4.8 to +10.8 yes 76 / 20 / 4%
Fable vs Sol +8.7 pts +5.8 to +11.7 yes 74 / 22 / 4%
Fable vs Gemini +11.7 pts +8.6 to +14.9 yes 82 / 16 / 2%

Fable separates from every model except Claude Opus 5.5 and Muse Spark 1.3. The top of the field is a chain rather than a single leader: Opus 5.5 is itself not separated from Muse Spark 1.3, Claude Opus 5 or Muse Spark 1.1. In the middle, adjacent models from Grok 4.7 to GPT-6 Sol are within noise of each other, though the ends of that group separate. Gemini 3.8 Flash is last and separated from every model except GPT-6 Sol. Won, lost and tied are shares of the dataset.

Where the points go

Every positive criterion was classified by the kind of work it tests: Extraction restates a finding written in the report, Classification reaches the right diagnosis, grade, stage or risk group under a named system, Interpretation explains what the findings mean and their limits, Management recommends the next step, and Communication completes the form, states uncertainty and stays in scope. The classification is not cosmetic: it shows where the points are lost, and it is how the retarget is enforced, since the weight the failure modes concern sits in interpretation, management and the uncertainty facing part of communication.

Model Extraction Classification Interpretation Management Communication
Claude Fable 5.1 96.6% 88.9% 85.2% 85.5% 80.6%
Claude Opus 5.5 95.2% 88.9% 81.3% 82.9% 80.4%
Muse Spark 1.3 94.6% 86.1% 75.3% 82.8% 78.5%
Claude Opus 5 96.4% 85.8% 84.5% 87.0% 79.4%
Muse Spark 1.1 94.8% 89.8% 74.6% 81.4% 72.5%
Grok 4.7 91.7% 84.9% 71.2% 63.1% 76.4%
Grok 4.6 92.8% 81.8% 69.7% 70.9% 65.0%
Claude Sonnet 5 91.7% 82.4% 73.9% 74.4% 71.2%
GPT-6 Astra 88.2% 78.1% 56.4% 68.2% 80.7%
GPT-6 Sol 86.3% 78.1% 59.2% 72.7% 70.3%
Gemini 3.8 Flash 93.3% 87.2% 71.8% 73.0% 57.1%

Weighted pass rate on positive criteria, by type, per model, over all runs. Criteria per type: Extraction 328 (36% of positive weight), Classification 118 (18%), Interpretation 200 (30%), Management 93 (11%), Communication 33 (4%). Extraction is near saturated; Interpretation has the lowest pass rate and, carrying 30% of the weight, is where the most reward is lost.

Report knowledge versus guideline knowledge

The gap widens when a criterion needs knowledge the report does not print, such as a staging table or a guideline threshold, rather than a fact stated on the page.

Model In the report (n=478) Needs outside knowledge (n=294) Numeric (n=175) Textual (n=597)
Claude Fable 5.1 0.921 0.868 0.948 0.883
Claude Opus 5.5 0.901 0.849 0.955 0.856
Muse Spark 1.3 0.900 0.789 0.881 0.844
Claude Opus 5 0.924 0.848 0.929 0.880
Muse Spark 1.1 0.895 0.798 0.900 0.840
Grok 4.7 0.853 0.740 0.850 0.791
Grok 4.6 0.851 0.737 0.861 0.785
Claude Sonnet 5 0.846 0.782 0.855 0.807
GPT-6 Astra 0.822 0.637 0.809 0.722
GPT-6 Sol 0.804 0.666 0.824 0.721
Gemini 3.8 Flash 0.837 0.794 0.875 0.802

Weighted pass rate on positive criteria, split by whether satisfying the criterion needs knowledge not printed in the report and by whether it hinges on a specific number.

Hardest tasks: where every model struggled

These are the lowest reward tasks in the benchmark. The pattern is consistent with the second finding: each one punishes saying more than the specimen supports. The anchor column is the criterion most responsible for lost reward, read it as the thing the report did not let you say.

Task Mean Strongest Weakest Anchor most often missed
report_2605240 38.5% Opus 5.5 Gemini Maintain the Bethesda Category II designation rather than reclassifying the specimen as nondiagnostic
report_2605258 42.1% Opus 5.5 Grok 4.6 State that the decision to proceed cannot be determined from the report and name what is required
report_2605306 43.0% Sol Gemini Flag the internal site discordance: block level microscopy assigns active H. pylori gastritis to the body, the final diagnosis assigns it to the antrum
report_2605351 48.8% Muse 1.3 Grok 4.7 Accept the reported OLGA stage rather than recomputing it and drifting
report_2605222 53.4% Opus 5.5 Gemini State that the HPV 52 versus HPV 18 ranking is population dependent, not universal

Lowest per task mean reward across models, three rollouts each.

Largest model spreads: where behavior separated most

These tasks produced the widest gap between the strongest and weakest model on the same problem, the most informative cases for understanding what distinguishes a good pathology read from a weak one.

Task Spread Strongest Weakest What separated the models
report_2605303 48.7 pts Fable Muse 1.3 Stating that the receptor subtype is not established and must be confirmed before the result can be assessed
report_2605363 47.3 pts Muse 1.3 Grok 4.6 Reporting the biomarkers as done on a prior specimen, and the primary as not characterized on this nodal specimen
report_2605281 42.7 pts Opus 5 Astra Holding management open until the pending HPV PCR result is incorporated
report_2605694 39.9 pts Muse 1.3 Gemini Stating what a curettage cannot establish: lymphovascular invasion elsewhere, the molecular subgroup, or management before staging
report_2605191 38.0 pts Astra Gemini Keeping the specimens separate and reporting the ascending colon tumor as conventional, not serrated, adenocarcinoma

Highest per task spread between the best and weakest model, excluding tasks already listed among the hardest. Several of these are the cases the failure mode examples below draw on.

Where models lose reward

The four failure modes that define the retarget are below, each with a complete example: the task, the exact rubric criterion, the model response, the judge verdict and the judge justification, so the failure is legible rather than asserted. No model is immune, and the examples below are drawn from the top of the field, where the errors are subtle enough to survive an otherwise strong summary.

Failure mode Where it shows up Signal in the results
Definitive answer on insufficient evidence Interpretation, Communication Overreach is the most tripped penalty, 123 of the 299
Inaccurate or unjustified management Management Management met 63 to 87%; recommendations close but not case specific do not earn
Unreconciled uncertainty, missing verification Management, Interpretation Five criteria are met by no model in any run; two ask the model to qualify a result rather than assert it
Reading the report too literally Interpretation Interpretation is 43 to 52% of every model lost weight

Definitive answer where the report withholds one

The model states a diagnosis, a stage, or a value the report does not support, in place of holding the category the report actually renders. This is the overreach that dominates the observed penalties, and it is a discipline problem more than a knowledge problem.

Task · report_2605240 — a thyroid fine needle aspiration reported as Bethesda Category II.

Rubric criterion: Maintains the Bethesda Category II designation. (weight 8; Claude Opus 5 met this in 0 of 3 runs, the other models in 87%.)

Model response: Claude Opus 5 stated that the original Bethesda II diagnosis is not defensible, reclassified the case to Bethesda I, nondiagnostic, and instructed the team to treat it as a nondiagnostic aspirate rather than a benign nodule.

Verdict: FAILED

Judge: The submitted response explicitly states that the original Bethesda Category II diagnosis is not defensible, reclassifies the case to Bethesda Category I, and instructs the team to treat it as a nondiagnostic aspirate rather than a benign nodule. Therefore it does not maintain the Bethesda Category II designation.

The correct answer holds the reported category; overriding it with a more confident reclassification invents a call the specimen does not carry. This is overreach, the largest penalty category in the set: 123 of the 299 penalty criteria describe it, against 105 for misclassification and 53 for fabrication.

Inaccurate or unjustified management

The task asks for a pathology report interpretation, and the model answers like a treating clinician, naming a test or therapy the report records only as a suggestion, or gives a recommendation that is close to right but not the one this case calls for. Overreach into management is the single most tripped penalty; the mirror problem is a correct recommendation stated too loosely to earn its criterion.

Task · report_2605236 — a resection where the pathologist suggested genetic counselling.

Rubric criterion: States that the pathologist suggested genetic counselling, without asserting that germline testing is indicated. (weight 10; Claude Fable 5.1 met this in 0 of 3 runs, the other models in 73%.)

Model response: Claude Fable 5.1 stated that the pathologist suggested genetic counselling, then went further and recommended proceeding to germline multigene testing.

Verdict: FAILED

Judge: The response does state that the pathologist suggested genetic counselling, but it goes further and explicitly recommends proceeding to germline multigene testing. That overstates the criterion, which requires mentioning genetic counselling without asserting that germline testing is indicated.

The report records a suggestion; asserting that germline testing is indicated invents a downstream decision the report does not make. The retargeted rubrics require the benefit of any recommended testing to be stated, not only its justification, so a recommendation that is reasonable in general but not the one this report calls for does not earn the criterion.

Uncertainty resolved instead of verified

The report deliberately leaves a question open, an unresolved differential, a discordance between modalities, a result that must be confirmed, and the model either resolves it too decisively or acts on it without naming the step that would verify it.

Task · report_2605303 — a specimen with a PD-L1 result but no established receptor subtype.

Rubric criterion: Notes that the receptor subtype is not established in this report and must be confirmed before the implications of the result can be fully assessed. (weight 10; Muse Spark 1.3 met this in 0 of 3 runs, the other models in 80%.)

Model response: Muse Spark 1.3 discussed the PD-L1 result and advised correlating it with the prior case biomarker profile, but did not note that the receptor subtype is unestablished or that it must be confirmed first.

Verdict: FAILED

Judge: The submitted response discusses the PD-L1 result and advises correlating it with the prior case biomarker profile, but it does not explicitly note that the receptor subtype is not established in this report or state that it must be confirmed before the implications of the result can be fully assessed.

Acting on a result is not enough; the focus requires naming the confirmation the result still needs. The same pattern runs through the hardest cases: on report_2605222 the rubric asks the model to state that the HPV 52 versus HPV 18 ranking is population dependent rather than universal, a criterion no model met in any run, and on report_2605281 to hold management open until the pending HPV PCR result is incorporated rather than finalizing on histology alone.

Reading the report too literally

The model restates the report faithfully but does not read it for meaning: it repeats a staging code without saying what the code implies, applies a named system by rote to a case it does not fit, or relabels a specimen more specifically than the report allows.

Task · report_2605510 — a partial nephrectomy staged pT1b, Nx, Mx.

Rubric criterion: States that Nx and Mx indicate that nodal and distant staging are not assessed in this specimen. (weight 6; Muse Spark 1.1 met this in 0 of 3 runs, the other models in 70%.)

Model response: Muse Spark 1.1 reported the stage as pT1b, Nx, Mx, but did not state that Nx and Mx mean nodal and distant staging are not assessed in this specimen.

Verdict: FAILED

Judge: The response mentions the reported stage as pT1b, Nx, Mx, but it does not explicitly state that Nx and Mx mean nodal and distant staging are not assessed in this specimen.

The report states a code; the model must state its meaning. Over literal reading also drives the interpretation misses that separate the field: on report_2605191 the ascending colon tumor must be reported as a conventional adenocarcinoma rather than relabeled as a serrated subtype, and on report_2605291 the rubric asks the model to note that the infiltrative follicular subtype carries a less favorable behavior than the classic subtype, a reading no model made in any run.

Reading effort, not just score

Effort diverges far more than score. The table below is the mean per task effort of each model. The models agree on which tasks are hard, the rank correlation of per task means between any two models is 0.57 to 0.87, and repeats are stable, with a within task standard deviation of 0.045 for GPT-6 Astra to 0.067 for Claude Sonnet 5. What they do not share is how much work they spend reaching the answer.

Model Turns Tool calls Minutes Reasoning tokens (k)
Claude Fable 5.1 9.3 8.3 10.1 20.9
Claude Opus 5.5 25.7 32.4 19.2 64.3
Muse Spark 1.3 5.3 4.4 2.8 2.8
Claude Opus 5 7.0 7.2 5.3 6.6
Muse Spark 1.1 4.2 3.2 3.4 3.0
Grok 4.7 5.4 5.9 6.5 15.1
Grok 4.6 3.3 2.4 3.0 2.6
Claude Sonnet 5 3.7 2.8 3.1 3.2
GPT-6 Astra 9.4 11.9 4.9 5.1
GPT-6 Sol 9.8 11.5 3.2 6.0
Gemini 3.8 Flash 11.2 10.2 3.8 18.5

Mean per task effort. Claude Opus 5 writes 2,525 words to GPT-6 Sol 456; a run costs about 0.13 dollars on Muse Spark 1.1 and 2.95 dollars on Claude Opus 5.5 (GPT-6 Sol was not yet priced on the platform when it ran). Summary length barely tracks reward: the per model correlation between words and score runs from +0.18 to -0.22.

What this means for deploying models on pathology workflows

Realm is hard because it is concrete: the rubric grades exact figures, preserved limits, and the absence of unsupported escalation, with no credit for sounding authoritative. Three takeaways follow.

1

Treat extraction as reliable and put human review where judgment happens.

All eleven models read messy reports and produce coherent structured summaries; extraction criteria are met 93% of the time. The reward is lost downstream, at the point where the model decides whether to stop at the report. Human review is most valuable there, not on the extraction layer.

2

The dangerous errors are confident additions and resolved uncertainty, not missing facts.

Overreach is the largest penalty category, 123 of the 299, and the penalties reorder the field: on positive criteria alone the order would be Fable, Opus 5, Opus 5.5, Muse 1.1, Muse 1.3, Sonnet, Gemini, Grok 4.7, Grok 4.6, Sol, Astra. A reviewer should check that every stage, biomarker and recommendation in the output is present in the source report, and that every stated certainty is one the report actually supports.

3

Score behavior, and do not reward effort.

The models cluster tightly on headline score, and the separations that exist are penalty separations: Gemini 3.8 Flash gives back 0.159 of its earned weight to penalties while GPT-6 Astra gives back only 0.044. The most process heavy agent in the field, Claude Opus 5.5, costs more than twice as much per run as the leader without beating it. Teams evaluating models for this work should read how an answer was produced and treat length as a warning sign, not a proxy for care.

Methodology note

The dataset of expert authored pathology tasks, most single report interpretations and a smaller set of multi report integrations, was reauthored and signed off by domain reviewers over 17 to 20 September 2026. Eleven frontier models: Claude Fable 5.1 (max), Claude Opus 5.5 (max), Muse Spark 1.3 (xhigh), Claude Opus 5 (max), Muse Spark 1.1 (xhigh), Grok 4.7 (xhigh), Grok 4.6 (high), Claude Sonnet 5 (default), GPT-6 Astra and GPT-6 Sol (max, with the subagent task tool denied), and Gemini 3.8 Flash (high); three rollouts per task. The eight launch models ran in Daytona sandboxes; Grok 4.7, GPT-6 Sol and Claude Opus 5.5, released the following week, ran on the identical pinned task versions in Modal sandboxes. Twelve repetitions that failed for infrastructure reasons, a provider stream timeout or a harness idle stall with no answer written, were rerun once with the same configuration and matched by task, never selected by score; there were no exclusions. Reports are anonymized. Agents run in a sandboxed OpenCode container with the report PDFs available and a generic toolkit of a shell, file read, write and edit, search over files, a todo list, and web search and web fetch; each writes a single structured interpretation, which the verifier grades. One judge, GPT-5.4 mini, snapshot of 17 March 2026 at medium effort, grades every criterion against a hidden expert answer, one call per criterion. Rubrics carry 1,071 criteria, 772 positive and 299 penalty (overreach 123, misclassification 105, fabrication 53, omission 10, other 8), 12 to 29 per task. Reward is earned positive weight over total positive weight, floored at zero; a penalty subtracts its weight when the behavior is present. Because the earlier rubrics carried no penalties, these results are a new comparison identity and are not continuous with the earlier release.

More from Realm benchmark series

CortexRetrievalBench

An evaluation of AI retrieval systems' ability to find and rank authoritative sources for financial research

    Realm: Legal reasoning benchmark

    The standard for evaluating legal reasoning in AI systems

      Realm: Financial reasoning benchmark

      An evaluation of frontier models on finance reasoning and spreadsheet-grounded analysis