Grok was much more search-intensive: about 41 tool calls per task versus Opus’s 22. It made 3,293 SEC filing-search calls against 486 for Opus and uniquely satisfied 529 rubric criteria, compared with 470 uniquely satisfied by Opus.
Section scores
Behavioral differences
Broader SEC search, more evidence gathering, more tool failures, and shorter final answers. Source failures appeared on 24 tasks.
More shell-based document processing, longer answers, and fewer failed fetches. Source failures appeared on only four tasks.
Opus wrote longer answers—13,445 versus 10,966 characters—and used more turns, but it more often omitted the rubric’s exact requested formulation. Its cleaner execution did not consistently become the precise assertions rewarded by the judge.
Conclusion. Treat Grok 4.6 and Opus 5 as a statistical tie around 52–53%. Grok’s nominal advantage appears driven by aggressive search and slightly better rubric targeting; Opus was more efficient and sometimes more nuanced, but showed weaker task-specific instruction adherence when the prompt required exact formulations. Sonnet 5’s 46.2% is meaningfully below Grok, with a paired difference of +7.0 points and a confidence interval of roughly +4.5 to +9.4.