This is the first part in a series on how AI is crossing out of software and into the physical world. Jan Bosch begins at the bottom of the stack, with the silicon everything else is built on.
There’s a comforting story about artificial intelligence in which the important action happens in the models. Someone has a clever architecture, trains it on enough data and intelligence emerges. In this story, the chips are plumbing: necessary, unglamorous, and interchangeable.
The story is wrong, and the last two years have made it obvious. The binding constraint on AI isn’t ideas and, increasingly, not even data. It’s compute: the physical capacity to train and run large models. And compute isn’t abstract; it’s a specific set of factories, in a specific set of countries, producing a specific set of components that almost nobody can substitute. Before we talk about robots, factories, grids or agents, we have to talk about the substrate they all sit on. Everything downstream inherits its constraints.
For years, the intuitive limit on chips was the leading edge of fabrication – 5nm, 3nm, 2nm, the relentless march that made headlines. That’s no longer where the pressure is greatest.
The frontier AI accelerator isn’t one chip; it’s a package: one or more logic dies built on a leading-edge process, stacks of high-bandwidth memory (HBM) sitting right next to them, all bonded together on an advanced-packaging substrate such as TSMC’s chip-on-wafer-on-substrate, or Cowos. Pull on the supply chain and you find that logic fabrication isn’t the tight link. According to an Epoch AI analysis, the largest AI-chip designers consumed roughly 90 percent of global Cowos capacity and HBM supply in 2025, but only about 12 percent of advanced logic-die production. The scarce resources are the memory and the packaging, not the transistors themselves.
The numbers are stark. TSMC’s Cowos capacity has been sold out through 2025 and into 2026, with the supply-demand gap only expected to narrow from around 20 percent to around 10 percent by the end of 2026. SK Hynix’s CFO said the company had already sold out its entire 2026 HBM supply. TSMC has responded with up to 56 billion dollars in planned 2026 capex and a separate 100-billion-dollar U.S. expansion weighted toward advanced packaging – but new packaging lines take years to qualify, and demand from generative AI keeps outpacing the additions.
This matters more than a normal component shortage, because the bottleneck is structural. HBM is made by a handful of memory vendors. Advanced packaging at the required volume is dominated by one company. The whole system depends on a small number of tools from an even smaller number of suppliers. When your input is genuinely irreplaceable, price doesn’t clear the market. Rather, it’s allocation, and that now becomes a political act.
For a startup, the compute substrate sets a ceiling you didn’t choose and can’t easily raise. The good news is that you almost never touch it directly. The entire point of the cloud is that a four-person team can rent time on hardware it could never build. That abstraction is real, and it’s the reason a small team can train or fine-tune models at all.
The bad news is that renting means you’re last in line during a shortage and your unit economics are set by someone else’s allocation decisions. When capacity is sold out through next year, the marginal GPU hour goes to whoever signed the biggest advance commitment, which is to say: the incumbents. A startup whose product depends on cheap, abundant inference is building on a substrate whose price and availability it doesn’t control.
This creates two viable postures. The first is to treat compute as a commodity you consume as sparingly as possible: Distill models, push inference to the edge and architect so that a doubling of compute cost doesn’t break your business. The second, more contrarian, is to build for the constraint – inference-optimized silicon startups, compilers, memory-efficient architectures and tooling that squeezes more useful work out of each scarce package. The shortage that threatens application startups is the entire market for infrastructure startups. Where you sit on that line should be a deliberate choice, not an accident of your first architecture.
For large incumbents, the calculus inverts. They have the balance sheet to buy their way up the priority list and, increasingly, to stop renting the ceiling and start owning it.
This is the real story behind the hyperscalers designing their own silicon. Custom accelerators aren’t primarily about beating the merchant vendor on raw performance; they’re about escaping a supply chain in which one company captures the margin and controls the allocation. Vertical integration is a hedge against dependence. If your business now runs on compute, being a price-taker on your most important input is an existential risk, and the firms that can integrate are moving to do so.
But integration has a ceiling of its own. You can design your own chip; you still can’t make your own HBM or your own packaging at scale. So, the incumbent’s advantage is real but bounded: it buys priority and margin, not independence. The genuinely uncomfortable position belongs to the mid-tier incumbent: too big to be nimble, too small to command allocation or justify a custom-silicon program. For that company, the compute substrate is a slow squeeze, and “we’ll just buy more GPUs” stops being a strategy somewhere around the point the vendor tells them what they’re allowed to have.
Zoom out far enough and the compute substrate stops being an industry story and becomes a geopolitical one. If frontier AI depends on a supply chain concentrated in a few firms and a few geographies, then controlling that supply chain is a lever of national power, and governments have noticed. US export controls have made advanced accelerators an instrument of foreign policy: The Commerce Department tightened rules in 2026 targeting the most capable processors and closing loopholes that let restricted firms buy through overseas subsidiaries. The policy has whipsawed, such as an H200 allowance late in 2025 and tightened Blackwell restrictions in 2026, but the direction is unmistakable: Chips are treated as strategic materiel, and Nvidia’s share of the Chinese market collapsed toward zero on new shipments as a result, with Jensen Huang conceding much of that market to domestic Chinese rivals.
This is the deepest implication of the substrate. The concentration that makes the AI supply chain efficient also makes it a chokepoint, and chokepoints get contested. Fabs become national-security assets. Packaging capacity becomes leverage. The physical geography of where chips are made starts to shape which countries can build frontier AI at all. And that, in turn, shapes who sets the norms for everything built on top. For a society, the question raised by the compute substrate isn’t “how fast will AI improve?” but “who gets to decide?”
Every topic that follows – robots, factories, grids, autonomous vehicles, scientific labs – describes intelligence moving into some corner of the physical world. Each of those depends, silently, on the substrate we’ve just outlined. When we ask later why edge AI matters, part of the answer is that it relieves pressure on a bottlenecked data center supply chain. When we ask who wins in industrial automation, part of the answer is who can secure the silicon.
The models get the headlines. But the model is only as available, as cheap and as sovereign as the compute beneath it. Start there, and the rest makes more sense.


