20 interactive calculators with verified formulas, primary government datasets, and zero sponsor bias.
Benchmarking ten frontier large language models on 1,477 board-style multiple choice questions in hematology reveals substantial specialist knowledge but

Large Language Models (LLMs) have become increasingly popular among clinicians and patients for medical queries. According to editorial research analyzed by Groundwork, However, their accuracy and safety at the specialist level in hematology remain insufficiently characterised. To address this knowledge gap, we benchmarked ten frontier proprietary and open-weight LLMs across two generations on 1,477 board-style hematology multiple-choice questions (MCQs) derived from five educational datasets spanning nine disease areas and six clinical skill domains.
We utilised five educational datasets to derive 1,477 board-style MCQs, covering nine disease areas and six clinical skill domains. The datasets included text-only and multimodal case vignettes. We then benchmarked ten frontier proprietary and open-weight LLMs across two generations on these MCQs.
The results of our benchmarking study are presented in Table 1. We found that Claude Opus 5 had the highest mean accuracy (92.7% text, 76.9% multimodal), followed closely by Gemini-3.1 Pro (91.4% and 78.7%), Gemini-3.6 Flash (91.0% and 74.8%), and GPT-5.6 Sol (89.9% and 76.7%). Accuracy significantly correlated with model size both for text-only and multimodal MCQs.
| Model | Text Accuracy | Multimodal Accuracy |
|---|---|---|
| Claude Opus 5 | 92.7% | 76.9% |
| Gemini-3.1 Pro | 91.4% | 78.7% |
| Gemini-3.6 Flash | 91.0% | 74.8% |
| GPT-5.6 Sol | 89.9% | 76.7% |
We conducted an error analysis to identify the types of questions that the top-performing models struggled with. Our results suggest that top-performing models exhibited highly concordant failure patterns, suggesting shared limitations on challenging cases.
Our study highlights the substantial specialist hematology knowledge exhibited by frontier LLMs across diverse subspecialist domains and clinical skill sets. However, despite high accuracy on board-style questions in hematology, continuous expert-on-the-loop output monitoring is central. This is because even the most accurate LLMs can make mistakes, and it is essential to verify their responses with human experts.
in summary, our study provides valuable insights into the accuracy and limitations of frontier LLMs in hematology. While these models exhibit substantial specialist knowledge, they are not yet ready to replace human experts in clinical decision-making. Continuous monitoring and verification of LLM outputs are essential to ensure patient safety and optimal clinical outcomes.
Large language models in hematology have shown high accuracy on board-style questions but may struggle with challenging cases. Continuous expert-on-the-loop output monitoring is essential to ensure patient safety and optimal clinical outcomes.
No, large language models are not yet ready to replace human experts in clinical decision-making. They should be used in conjunction with human experts and continuously monitored and verified.
Large language models in hematology can provide rapid and accurate answers to complex medical queries, reducing the workload of human experts and improving patient care.
Large language models in hematology may exhibit biases due to the data they were trained on, which can affect their accuracy and reliability.

Early breast cancer detection saved Katherine LaNasa's life. Learn how to prioritize your health and reduce your risk of developing breast cancer.
Early detection of ovarian cancer remains a clinical challenge because available blood biomarkers lack the sensitivity and specificity required for population
Altered EEG patterns co-vary with specific improvements in mind health at summer camp, suggesting that exposure to device-free natural environments supports
Explore related evidence-based investigations, decision tools, and entity breakdowns:
Contextual evidence and verified documentation referenced in this research guide
Groundwork enforces a strict, independent verification standard. All claims and benchmark figures in this guide are cross-referenced against the primary documentation and regulatory registries listed below:
Maya Okafor (2026). Benchmarking 10 Frontier LLMs on Hematology Exams. Groundwork. Retrieved from https://gworky.com/article/benchmarking-llms-in-hematology
Originally published at https://gworky.com/article/benchmarking-llms-in-hematology — Groundwork Evidence-Based Research.
Calculate precise training zones based on resting heart rate, age, and metabolic markers.
Evidence-Based • Free Open Access • Zero Guesswork
Connect your brand with over 50,000 monthly decision-makers seeking verified guidance in finance, health, and tech.
Find the ideal bedtime for a chosen wake time — or the best wake windows if you go to sleep now — using 90-minute ultradian sleep cycles.
bodyEarly Breast Cancer Detection Saved Katherine LaNasa's Life: A Cautionary Tale
AI Detects a Distributed Blood Metabolomic Systemotype
Altered EEG Patterns Co-Vary with Specific Improvements in Mind Health at Summer Camp
Epigenetic Signatures of Biological Vulnerability and Residual Risk in Heart Failure
Evaluated for performance, privacy protocols, and pricing transparency.
| Solution | Key Benchmark | Pricing | Verdict & Access |
|---|---|---|---|
Check My Body HealthEditor Pick via Check My Body Health | Clinical food sensitivity & biomarker panels | $38 Test Kit | |
AG1 Protocol via Athletic Greens | NSF Certified daily micronutrient complex | $79/mo | |
Clinical Fasting Guide via Metabolic Research Lab | PubMed-backed metabolic autophagy protocol | $37 One-Time |
Health & Tech Writer
Maya Okafor is a Senior Clinical Sciences Analyst focusing on evidence-based dietary interventions, metabolic longevity markers, and pharmaceutical compounding compliance. Her research bridges molecular biology and applied lifestyle medicine, auditing commercial dietary supplements and evaluating peer-reviewed evidence to help readers distinguish scientifically validated regimens from marketing wellness hype.
Health Data Analyst
Sarah Lin heads clinical analysis for the Body & Health Sciences Desk at Groundwork. She directs primary meta-analyses of peer-reviewed randomized controlled trials (RCTs) indexed in PubMed, evaluating metabolic health, cardiovascular biomarkers, and preventative nutrition protocols. Lin ensures Groundwork's health calculators and wellness guides strictly conform to clinical evidence standards and public health guidelines.
This guide underwent secondary data verification to confirm primary source integrity, calculation formulas, and regulatory compliance before publication.