Writing · Markets
A map of the thermal stack, the bottlenecks at each layer, and how innovation in one part of the system changes the economics of the rest.
I spent the past few weeks digging into AI cooling, speaking with people across chip packaging, thermal management, and data center infrastructure.
What initially looked like one market turned out to be a stack of very different problems.
The thermal path runs from the silicon itself, through the package, into the cold plate, through the rack liquid loop, and finally out of the building.
Each layer has different technical constraints, buyers, gatekeepers, commercialization cycles, and economics.
More importantly, the layers are not independent.
An improvement in one part of the stack can increase the value of another. It can also make another technology less necessary.
So the useful question is not: which cooling technology is best?
It is: which improvement changes the economics of the rest of the system?
There is another reason the layers need to be separated: the buyer changes as you move through the stack.
Inside the package, accelerator vendors such as NVIDIA and AMD, together with their manufacturing and packaging partners, largely control adoption.
Near the chip, OEMs and system vendors begin to matter.
At the cold plate and rack level, OEMs, rack integrators, hyperscalers and neoclouds have more influence.
At the facility level, the data center owner or operator ultimately controls the architecture.
Two technologies that create similar thermal improvements can therefore have completely different commercialization paths depending on where they sit.
This is the closest you can get to the source of the heat.
Innovation here includes better thermal interface materials, thinner silicon, higher conductivity materials, alternative substrates, direct die cooling and embedded microfluidics.
This layer has potentially enormous technical leverage because every downstream cooling system starts with the heat flux coming out of the package.
But it is also one of the hardest places for a startup to enter.
If your technology sits inside an NVIDIA or AMD package, the accelerator vendor has to support the integration.
That means solving for much more than thermal performance. Stress, bonding, warpage, reliability, yield, manufacturability and cost all matter.
One packaging engineer I spoke with emphasized that conductivity alone is not enough. A material can look excellent thermally and still fail if it cannot be integrated reliably into the package.
Another expert pushed the point further:
“The spreader isn’t the bottleneck. The die to lid interface is.”†
“The real move is dropping the lid and cooling the silicon directly.”†
That distinction matters.
If a meaningful share of thermal resistance sits at the interfaces between the die and the cooling system, then improving heat spreading alone may not be the highest leverage move.
The more radical approach is to remove interfaces and move cooling closer to the silicon itself.
This is where direct-die cooling and embedded microfluidics become interesting. Instead of conducting heat through several layers before reaching the coolant, the architecture brings the coolant physically closer to where the heat is generated.
So the long-term competition at this layer may be less about which material has the highest thermal conductivity and more about: which architecture removes the most meaningful thermal resistance while remaining manufacturable at scale?
Commercially, this layer is unusually gatekept.
You are not selling into a broad market of end users. You are trying to get designed into a platform controlled by a very small number of companies.
The commercialization path can therefore feel relatively binary: you are designed in, or you are not.
Once designed in, however, the position can be highly valuable and sticky.
The next layer sits outside the package but still very close to the accelerator.
This is where a company can potentially improve thermal performance without requiring a full redesign of the GPU package.
That can make commercialization faster.
But it is also one of the more strategically vulnerable parts of the stack because innovation on either side can reduce its incremental value.
There are two broad approaches.
Passive thermal management
Passive technologies improve heat flow continuously.
Examples include high conductivity spreaders, improved thermal interfaces and vapor chambers.
The basic proposition is straightforward: reduce thermal resistance and move heat away from hotspots more efficiently.
The risk is that this only matters if heat spreading is actually the binding constraint.
If package-level interfaces improve substantially, if direct die cooling becomes more common, or if coolant moves deeper into the package, some passive near-chip technologies could become less valuable.
Active thermal management
Active thermal management takes a different approach.
Thermoelectric devices, for example, use electrical power to actively pump heat and can be controlled dynamically.
That potentially allows different parts of a package to operate at different target temperatures.
This matters because modern accelerator packages are heterogeneous. GPU dies, HBM, chiplets and I/O dies do not necessarily have identical thermal limits or generate heat in identical ways.
The proposition therefore becomes less “make everything colder” and more “put cooling where it is needed, when it is needed.”
That could be valuable if it reduces the need to design the entire cooling system around the worst hotspot.
But active cooling consumes power itself.
So the right question is not whether it can lower temperature. It is whether the local intervention produces a meaningful system-level benefit.
This is why I view near-chip thermal management as the contested middle of the stack.
If Layer 1 gets much better, some Layer 2 value may disappear.
If cold plates improve dramatically, Layer 2 can face pressure from the other direction.
A near-chip technology therefore has to prove more than a temperature improvement. It has to show that it changes the economics of the system around it.
This is the center of gravity of AI cooling today.
At some point, heat has to enter a fluid and be carried away. That is what cold plates and the liquid loop ultimately accomplish.
Even if package-level thermal performance improves dramatically, this layer does not disappear. The heat still has to go somewhere.
The main question here is therefore not whether this layer survives. It is: which architecture becomes standard?
Single phase cooling
In single-phase cooling, the coolant remains liquid as it absorbs and transports heat.
It is relatively mature, easier to operate and already widely deployed.
One expert I spoke with put it this way:
“Two phase is the most interesting one long-term, but it’s not where the gains are right now.”†
“Single phase water hasn’t run out of room yet.”†
That is an important point.
There is a tendency in emerging technologies to assume the more exotic architecture must eventually be the better one.
But single phase still has meaningful advantages in operations, serviceability, component maturity, reliability and integration.
Two-phase cooling
In two-phase cooling, the coolant intentionally changes phase.
The liquid absorbs heat and boils. That phase change can absorb a large amount of thermal energy, making the architecture attractive for very high heat flux.
The physics are compelling.
The operational challenges are also real.
Fluid cost, containment, pressure management, reliability, servicing and controls all matter.
So two phase can be technically impressive without necessarily being the near-term deployment winner.
Two-phase cooling should also be distinguished from immersion cooling.
Two-phase refers to the coolant changing from liquid to vapor as part of the heat transfer process.
Immersion describes an architecture in which server hardware is submerged in dielectric fluid.
They can overlap, but they are not the same thing.
Microchannels
Another important area of innovation is geometry.
Instead of moving coolant through relatively large passages, microchannels use many tiny channels to increase surface area and bring fluid closer to the heat source.
Those channels can exist in a conventional cold plate, inside the package substrate, or even directly in silicon.
As they move closer to the die, thermal performance can improve. But integration becomes harder.
That tradeoff appears repeatedly across the stack: the closer cooling gets to the heat source, the more technical leverage it can have, but the harder it becomes to integrate.
Once heat enters the liquid loop, it has to move through the rack.
This is the operational layer: CDUs, pumps, manifolds, flow control, leak detection and the systems that make liquid cooling work reliably across thousands of accelerators.
The innovation question here is less about discovering new thermal physics and more about making the infrastructure scalable, serviceable and reliable.
In that sense, this layer may select more for operators than inventors.
The winners may look less like breakthrough materials companies and more like highly capable infrastructure businesses.
At the other end of the stack, the problem changes completely.
The question is no longer how quickly heat leaves the chip. It is how efficiently the data center gets rid of it.
One expert summarized the opportunity this way:
“Right now, the big lever is at the facility. Getting the water temperature up so you run on dry coolers instead of chillers.”†
“That power goes straight back into GPUs.”†
This is the economic link.
Colder coolant may improve chip temperatures, but producing colder water consumes energy.
If improvements upstream allow the facility to operate with warmer coolant, the data center may be able to reduce compressor-based chilling and redirect more of its power budget toward compute.
The same expert put it more directly:
“Chip level thermal does make the list when you’re power constrained, but only if you use the gain to run warmer coolant rather than a colder chip. Same physics, very different value.”†
The goal is not necessarily to make the GPU as cold as possible.
It is to create enough thermal headroom that the rest of the data center can operate more efficiently.
This is where the cooling market becomes much more interesting.
The layers do not simply coexist. They change each other’s value.
A breakthrough at one layer can cannibalize one category while expanding another.
Suppose direct die cooling or embedded microfluidics dramatically reduces thermal resistance inside the package.
Some passive near-chip technologies may become less valuable because there is less incremental resistance left for them to remove.
But direct liquid cooling may become more important because more heat can now be moved efficiently into the fluid.
And if that package-level improvement allows hotter coolant, facility heat rejection may become cheaper.
One technical breakthrough has therefore reduced the value of one category while increasing the value of others.
The same dynamic applies to active near-chip cooling.
If active thermal control can manage GPU and HBM hotspots precisely, it may reduce the need for some passive spreading and potentially enable warmer coolant.
But it consumes power. Its value depends on whether that local power use creates a larger saving elsewhere.
Two-phase cooling could create another shift.
If it becomes operationally mature and easy to service, it could support much higher heat fluxes and reduce the need for some intermediate thermal enhancements.
At the same time, it could create new value in fluid handling, controls and system integration.
The important point is: value does not simply disappear. It migrates. And the bottleneck moves with it.
That is why cooling should not be analyzed as five independent markets. The economics of each layer depend partly on what happens in the others.
The sharper question is: what system constraint did the improvement actually remove?
I do not think there is a single winning layer. Different layers create different types of value.
Highest technical leverage: chip and package. Improvements close to the silicon can influence the entire downstream thermal system. The upside can be significant if a technology becomes designed in. But access is tightly controlled by chip vendors, and commercialization risk is high.
Clearest near-term deployment: direct liquid cooling. This is already addressing an immediate problem. Demand exists today, buyers understand the value, and increasingly dense AI infrastructure needs some version of it.
Clearest facility-level savings: warmer coolant and efficient heat rejection. If hotter coolant allows a data center to reduce mechanical chilling, the benefit can appear directly in the facility economics.
Highest strategic uncertainty: near-chip thermal management. This layer may create substantial value. But it also faces the greatest substitution risk from innovation on either side. Its durability depends on whether the benefit remains meaningful as package and liquid cooling architectures improve.
Instead of asking: does it cool better?
I would ask five questions.
Inside the package? Above the package? At the cold plate? At the rack? At the facility?
NVIDIA or AMD? OEMs? Rack integrators? Hyperscalers? Neoclouds? Data center operators? The answer changes the commercialization path dramatically.
Not every thermal resistance matters equally. A company can produce a real technical improvement while optimizing a part of the system that was never the limiting factor.
Higher TDP? Warmer coolant? Greater rack density? Lower cooling power? Less throttling? Less mechanical chilling? More compute per megawatt? If the improvement does not change one of those outcomes, its economic value may be limited.
If direct die cooling succeeds, does this technology still matter? If two phase becomes mainstream, does its value increase or disappear? If hotter coolant becomes standard, which facility technologies gain and which lose? This may be the most important test of durability.
The cooling stack appears to be moving in two directions at once.
Closer to the silicon: remove interfaces and move cooling deeper into the package.
At the facility: tolerate warmer coolant and make heat rejection cheaper.
Everything in between remains necessary, but its value depends partly on how those two ends evolve.
Which brings me back to the question I find most useful: which improvement at one layer changes the economics of the rest of the stack?
† Direct quotations are from Chloe’s expert conversations.
Chloe Wang writes about frontier technology, markets, startups, and commercialization. Follow her on LinkedIn, Instagram, and X.
Originally published on Substack.