Everyone thinks AI is bottlenecked by compute.
Everyone thinks AI is bottlenecked by compute. The next bottleneck is memory. An NVIDIA H100 delivers around 4 PFLOPS of FP8 compute and 3.35 TB/s of HBM bandwidth. Groq took a completely different approach: 500 MB of on-chip SRAM with 150 TB/s of memory bandw