Latent Space: The AI Engineer PodcastLatent.SpaceAll episodes
Academia is for Ambition — Alex Zhang, MIT
Last call for regular tickets for AI Engineer NYC! As an exclusive for Latent Space subscribers, the first 30 of you can take a 30% off code if it helps - for new tickets only, no refunds! See you in 2 weeks!
While we tend to cover industry on the pod, every so often we celebrate a clearly emerging superstar PhD. In 2024 we featured Shunyu Yao, who went on to build Operator at OpenAI and is now Chief AI Scientist of Tencent. In 2025 we featured Jack Morris, who went on to cofound Engram at $600m and is now a leading voice on continual learning.
This year we are proud to feature the work of Alex Zhang of MIT.
From GPU kernels and KernelBench to Recursive Language Models, Mismanaged Geniuses, and massive multi-agent swarms, Alex Zhang is exploring how much capability we’re leaving on the table by wrapping increasingly powerful models in primitive systems.
RLMs took over the timeline early this year:
and an RLM based harness was the first to ~solve ARC-AGI-3 before OpenAI’s Astra:
and is even today, influencing new research that has more extreme implications than RLMs:
We go deep on GPU Mode and AI-written kernels, research taste and why academics should take bets industry labs won’t, GEV and alternatives to the standard autoregressive language model, and the idea of harnesses as compositional generalizers. Alex explains RLMs, context offloading, programmatic subagent calling, Prime Agent, persistent subagents, and why the “language model” of the future may actually be an invisible swarm of agents underneath a simple interface. We also discuss OpenAI’s massive agent experiments, Kimi swarms, open-ended research at Sakana AI, speculative programmatic tool calling, capability overhang, Neuralese, and where Alex thinks the next big research opportunities may lie.
We discuss:
* Why AI-generated GPU kernels still leave substantial room for human expertise
* How one expert insight can potentially replace enormous amounts of brute-force token search
* Why PhD students should take research bets that initially look trivial, weird, or pointless
* What SWE-bench, RLMs, ReAct, and Quiet-STaR reveal about research taste
* GEV and why a language model does not have to mean an autoregressive text-to-text decoder
* Why Claude Code, Codex, and Pi are structurally more similar than they look
* How harness design can improve compositional generalization across tasks and domains
* RLMs: context offloading, code execution, recursive subagents, and shared memory
* Prime Agent, continual harnesses, and persistent agent-to-agent communication
* Why the model you query in the future may secretly be an entire swarm or scaffold
* OpenAI’s 10,000-agent experiment, 130B output tokens, and ~$40M-equivalent problem solving
* Why much of an agent swarm may be wasted search — and why convergence is still hard
* Kimi versus OpenAI and different approaches to multi-agent systems
* Open-endedness, Sakana AI, and finding hidden gems in enormous amounts of generated work
* Why current frontier models may already have a large capability overhang
* Speculative programmatic tool calling and overlapping tool execution with generation
* Whether English, code, or an e