OpenDiscoveryTrace is a public dataset of AI scientific agent trajectories that captures how models reason, not just what they produce.
Based on reporting by arXiv AI & Computer Science. Research, structure, and fact-checking by Groundwork.

OpenDiscoveryTrace is a public dataset of AI scientific agent trajectories that captures how models reason, not just what they produce. According to editorial research analyzed by Groundwork, This dataset enables auditing of scientific methodology, diagnosis of failure modes, and distinction between systematic reasoning and fortunate guessing.
Existing benchmarks for autonomous AI scientists focus on final outputs, such as generated code, hypotheses, or papers, but discard the reasoning process by which those outputs were obtained. This makes it impossible to evaluate the scientific methodology of AI agents, diagnose failure modes, or distinguish between systematic reasoning and fortunate guessing. In this article, we present OpenDiscoveryTrace, a public dataset of 558 complete AI scientific agent trajectories that captures how models reason, not just what they produce.
Existing benchmarks for autonomous AI scientists evaluate only final outputs, generated code, hypotheses, or papers, yet discard the reasoning process by which those outputs were obtained. This makes it impossible to audit scientific methodology, diagnose failure modes, or distinguish systematic reasoning from fortunate guessing.
We present OpenDiscoveryTrace, a public dataset of 558 complete AI scientific agent trajectories that captures how models reason, not just what they produce. Each trajectory records a structured 9-field-per-step trace, including thoughts, tool calls, observations, errors, revision triggers, and self-reported confidence, as models execute 124 scientific tasks spanning drug discovery, materials science, genomics, and scientific literature analysis.
The dataset covers seven models: three frontier models (GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro; 124 trajectories each, fully balanced across domains and difficulty levels) and four open-weight models (Qwen2.5-7B, Mistral-7B-v0.3, Phi-3.5-mini, and Qwen2.5-1.5B; 30 each), plus 60 live-retrieval variant trajectories.
Pilot analysis on 363 LLM-judged trajectories reveals that process traces expose behavioral differences invisible to output-only evaluation. All three frontier models achieve comparable success rates (84--89%), yet Claude Opus 4.6 produces 30$ imes$ more errors than GPT-5.4 (2.5 vs. 0.08 per trajectory, $p < 0.0001$, Cliff's $\delta = 0.613$), with qualitatively different error profiles---66.7% tool misuse for Claude versus 83.6% reasoning errors for GPT-5.4.
We define five benchmark tasks with baselines from logistic regression, random forests, LSTMs, and Transformer models. The dataset, trace schema, agent using, and benchmark definitions are publicly available under CC BY 4.0 to support research on process-level evaluation, scientific agent auditing, and AI governance.
“The OpenDiscoveryTrace dataset provides a unique opportunity for researchers to evaluate AI scientific workflows at the process level, rather than just focusing on final outputs. By analyzing the reasoning process, researchers can gain insights into the strengths and weaknesses of different AI models and improve the overall quality of AI-generated outputs.”
OpenDiscoveryTrace is a public dataset of AI scientific agent trajectories that captures how models reason, not just what they produce.
The goal of OpenDiscoveryTrace is to enable auditing of scientific methodology, diagnosis of failure modes, and distinction between systematic reasoning and fortunate guessing.
The dataset covers seven models: three frontier models (GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro; 124 trajectories each, fully balanced across domains and difficulty levels) and four open-weight models (Qwen2.5-7B, Mistral-7B-v0.3, Phi-3.5-mini, and Qwen2.5-1.5B; 30 each), plus 60 live-retrieval variant trajectories.
We define five benchmark tasks with baselines from logistic regression, random forests, LSTMs, and Transformer models.
The dataset, trace schema, agent using, and benchmark definitions are publicly available under CC BY 4.0.

Europe is moving forward with its ambitious Envision mission to Venus, despite a setback from NASA.

Prolonged laptop use has been linked to a range of health risks, including myopia, headaches, and eye strain.

Apple Watch listening features explained: data privacy, security, and how to use them safely.
Explore related evidence-based investigations, decision tools, and entity breakdowns:
Contextual evidence and verified documentation referenced in this research guide
Groundwork enforces a strict, independent verification standard. All claims and benchmark figures in this guide are cross-referenced against the primary documentation and regulatory registries listed below:
Sofia Reyes (2026). Evaluating AI Scientist Workflows with OpenDiscoveryTrace. Groundwork. Retrieved from https://gworky.com/article/opendiscoverytrace
Originally published at https://gworky.com/article/opendiscoverytrace — Groundwork Evidence-Based Research.
Uncover forgotten seat licenses, redundant cloud services, and recurring overhead.
Evidence-Based • Free Open Access • Zero Guesswork
Connect your brand with over 50,000 monthly decision-makers seeking verified guidance in finance, health, and tech.
Compare estimated monthly cost across leading AI models based on your token usage.
tech
tech
tech
techEvaluated for performance, privacy protocols, and pricing transparency.
| Solution | Key Benchmark | Pricing | Verdict & Access |
|---|---|---|---|
NordVPNEditor Pick via Nord Security | Audited WireGuard no-logs protocol | $3.39/mo | |
ExpressVPN via Express Technologies | Lightway protocol, RAM-only servers | $6.67/mo | |
Cloudflare WARP+ via Cloudflare Inc. | Fast Argo edge routing | $4.99/mo | Reference Benchmark |
Smart Home & Digital Privacy Analyst
Smart home and digital privacy analyst focused on data ownership, device security, and power efficiency of AI utilities and gadgets.
This guide underwent secondary data verification to confirm primary source integrity, calculation formulas, and regulatory compliance before publication.