Grok 4.6 beat GPT-5.6 Sol on agentic loop efficiency, spending $13.11 versus $20.18 across the same 3 builds.
Grok 4.6's edge came from bigger coding steps: 201 model calls versus GPT-5.6 Sol's 338.
Really Interesting experiments by @thehypedotnews, a 24/7 AI news in a really nice radio format. (love their chillout music)
1 number in this comparison explains a lot about where coding-agent economics are heading:
93-98% of the input tokens were cache reads.
These runs consumed roughly 20M tokens for Grok 4.6 and 26M for GPT-5.6 Sol, yet the author estimates they would have cost around 4x more without prompt caching.