Tag: inference
All the articles with the tag "inference".
-
The AI Chips After NVIDIA
Updated:The serious AI chip challengers are not building one replacement for NVIDIA. They are attacking different bottlenecks in memory, latency, power, networking, and software.
-
Groq Makes the Compiler Part of the Processor
Groq moves scheduling and data placement from reactive hardware into the compiler, turning a network of SRAM-heavy LPUs into the useful unit of inference capacity.
-
d-Matrix Moves the Math Into Memory
d-Matrix treats inference as a memory-placement problem, combining digital in-memory compute, an SRAM performance tier, LPDDR capacity, and standard data-center fabrics.
-
Cerebras Moved the Cluster Boundary Onto a Wafer
Cerebras removes many chip boundaries with wafer-scale integration, then rebuilds the system around distributed SRAM, streamed weights, rack-scale I/O, and a compiler.
-
There Is No Universal AI Runtime—So Build a Portable Control Plane
AI hardware is not binary-compatible, but applications, governance, routing, observability, and economic policy can still move across separate accelerator pools.
-
Alternative AI Hardware Is a Systems Problem
Huawei, FuriosaAI, and Rebellions show three system boundaries outside the dominant accelerator stack: a complete platform, an efficient inference card, and a scalable NPU system.
-
SambaNova Compiles Models Into a Rack
SambaNova's RDU matters because the compiler, three-tier memory system, rack, and serving software turn model placement into one deployable inference machine.