AI Compute Landscape
An architecture-led exploration of accelerators, software stacks, rack-scale systems, and the data centers built around them.
Each chapter asks the same questions: what bottleneck the architecture attacks, what unit is deployed, how memory and interconnect work, which software model developers adopt, and what the available evidence proves.
Landscape
Market maps and comparison models for understanding where different compute systems compete.
-
The AI Chips After NVIDIA
Updated:The serious AI chip challengers are not building one replacement for NVIDIA. They are attacking different bottlenecks in memory, latency, power, networking, and software.
NVIDIA
How NVIDIA expanded the accelerator boundary from GPU silicon to software, rack-scale fabrics, power, and cooling.
-
The GPU Stopped Being the Product
NVIDIA's data-center stack evolved from a fast accelerator into a rack-scale computer spanning compute, memory, interconnect, networking, power, and software.
-
The AI Rack Became a Data Center Design Problem
Blackwell moved NVIDIA's top-end scale-up domain from an eight-GPU server into a liquid-cooled rack, making power, cooling, networking, and commissioning part of the computer.
Rack and Data Center Infrastructure
How chips, racks, cooling, electrical equipment, grid power, construction, and operations must arrive together before compute becomes usable.
-
The AI Buildout Is Constrained by Its Slowest Layer
Updated:AI capacity becomes real only when chips, racks, cooling, networks, electrical equipment, grid power, permits, and operators arrive together.
AMD
How AMD combines merchant accelerators, an open software stack, networking, and partner-built systems.
-
AMD Is Rebuilding the GPU Stack in the Open
AMD is assembling a broad merchant alternative to NVIDIA, but its open stack moves more integration responsibility across software, fabrics, system vendors, and operators.
Cloud-Designed Silicon
How Trainium, TPU, Maia, MTIA, and other custom accelerators make the platform operator the system integrator.
-
When the Platform Operator Designs the Chip
Trainium, TPU, Maia, MTIA, and Jalapeño show how custom silicon turns the operator's infrastructure into the computer.
Groq
How compiler-planned execution and distributed SRAM redraw the boundary of low-latency inference.
-
Groq Makes the Compiler Part of the Processor
Groq moves scheduling and data placement from reactive hardware into the compiler, turning a network of SRAM-heavy LPUs into the useful unit of inference capacity.
Cerebras
How wafer-scale compute removes chip boundaries and rebuilds the system around streamed weights and distributed SRAM.
-
Cerebras Moved the Cluster Boundary Onto a Wafer
Cerebras removes many chip boundaries with wafer-scale integration, then rebuilds the system around distributed SRAM, streamed weights, rack-scale I/O, and a compiler.
SambaNova
How dataflow compilation, tiered memory, rack systems, and serving software become one deployable machine.
-
SambaNova Compiles Models Into a Rack
SambaNova's RDU matters because the compiler, three-tier memory system, rack, and serving software turn model placement into one deployable inference machine.
d-Matrix
How digital in-memory compute turns inference into a memory-placement and software-compilation problem.
-
d-Matrix Moves the Math Into Memory
d-Matrix treats inference as a memory-placement problem, combining digital in-memory compute, an SRAM performance tier, LPDDR capacity, and standard data-center fabrics.
Sovereign and Efficient Inference
How alternative platform, card, and NPU systems compete through ownership, efficiency, and deployment fit.
-
Alternative AI Hardware Is a Systems Problem
Huawei, FuriosaAI, and Rebellions show three system boundaries outside the dominant accelerator stack: a complete platform, an efficient inference card, and a scalable NPU system.
Interconnect and Optics
How scale-up fabrics, scale-out networks, collectives, congestion control, copper, and optics determine usable compute.
-
The Interconnect Determines How Much of the Chip You Can Use
At rack and pod scale, accelerator utilization depends on scale-up fabrics, scale-out networks, collective software, congestion control, and the physical path carrying every byte.
Memory, Manufacturing, and Packaging
How foundries, HBM, interposers, substrates, advanced packaging, assembly, and test shape accelerator delivery.
-
The AI Supply Chain Became Part of the Architecture
Accelerator performance and availability now depend on a serial manufacturing system spanning foundry nodes, HBM, chiplets, interposers, substrates, advanced packaging, assembly, and test.
Portable Control Plane
How applications, routing, governance, telemetry, and economic policy can span hardware-specific accelerator pools.
-
There Is No Universal AI Runtime—So Build a Portable Control Plane
AI hardware is not binary-compatible, but applications, governance, routing, observability, and economic policy can still move across separate accelerator pools.