Token costs "doubling every 45 days" against a productivity lift of "5%."
Glean founder says its becasue of architecture.
Enterprise AI bills are rising because cheaper tokens now power longer, more complicated task chains.
Today, one prompt can trigger retrieval, tool calls, and multi-step reasoning loops. Hence the number of tokens needed for a task is increasing faster than the rate of decrease in the token price.
Routing and context are better than raw model swaps because they are engineering based.
Cheap outputs cost a lot when mistakes mean more work, more tries, more supervision, or angry customers.
"Permission-aware context means the models don't have to learn something new about the company every time a request is made," Glean says.
Its benchmark showed that it had ~30% less tokens and 2.5x more preferred answers than other MCP tools.
That gap grew as the tasks grew harder, and the reported context-layer win rate grew from 66% to 73%.