
0 plays · Sep 18, 2026
In Episode #1028, Anton McGonnell (VP of Product at SambaNova) joins Jon Krohn to explain why the chips running most AI inference today were never designed for the job. Agentic AI has changed the computational profile of inference, with much larger inputs and far heavier caches feeding the token generation that follows, and that shift has exposed where GPU architecture struggles. SambaNova has raised over $2 billion to build an alternative, the reconfigurable dataflow unit, which lays a whole model out spatially across the chip rather than executing it kernel by kernel. In this episode, Anton discusses why the speed that matters is payback, and how speed and concurrency are what turn a fixed hardware cost into a six-month payback. He also walks through the trade-off every inference provider faces between speed per user and throughput per chip, what the RDU architecture changes about scaling and data center deployment, the economics of the new SN50, and why four out of five AI infrastructure leaders say they would pay a premium for faster tokens.
Additional materials: https://www.superdatascience.com/1028
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In this episode you will learn:
1:02:48
19:40
1:06:34
34:27
1:17:590 plays · Sep 1, 2026