DiligenceBench · Research

Reports

Short-form analysis from benchmark scores and complete agent trajectories.

01

Model comparison · 2 min read

Why Grok 4.6 edged ahead

Grok 4.6 enters at #2 on the Finance harness.

Grok was much more search-intensive: about 41 tool calls per task versus Opus’s 22. It made 3,293 SEC filing-search calls against 486 for Opus and uniquely satisfied 529 rubric criteria, compared with 470 uniquely satisfied by Opus.

Section scores

SectionGrok 4.6Opus 5
Factual accuracy55.054.2
Analytical reasoning53.251.3
Risk awareness50.849.3

Behavioral differences

Grok 4.6

Broader SEC search, more evidence gathering, more tool failures, and shorter final answers. Source failures appeared on 24 tasks.

Claude Opus 5

More shell-based document processing, longer answers, and fewer failed fetches. Source failures appeared on only four tasks.

Opus wrote longer answers—13,445 versus 10,966 characters—and used more turns, but it more often omitted the rubric’s exact requested formulation. Its cleaner execution did not consistently become the precise assertions rewarded by the judge.

Conclusion. Treat Grok 4.6 and Opus 5 as a statistical tie around 52–53%. Grok’s nominal advantage appears driven by aggressive search and slightly better rubric targeting; Opus was more efficient and sometimes more nuanced, but showed weaker task-specific instruction adherence when the prompt required exact formulations. Sonnet 5’s 46.2% is meaningfully below Grok, with a paired difference of +7.0 points and a confidence interval of roughly +4.5 to +9.4.