Claude Opus 5 scored 98/100, the Kimi K3 + Grok 4.5 setup scored 93/100 and ran at 4% of the cost ($1.27 against $31.71). PricingPer token, Kimi K3 costs 60% of Claude Opus 5, and Grok 4.5 costs 24% on output. Claude Opus 5 ran at xhigh reasoning in both phases, Kimi K3 at max, and Grok 4.5 at high. After rejecting a 1,001-operation batch, Claude Opus 5’s server went on to apply all 1,001 operations it had just refused. Claude Opus 5’s plan spells out the wrong behavior, saying that on an oversized count the connection “stays open, no operations consumed.”