LLMs Research
LLMs Research Podcast
Podcast covering important research papers of large language models. llmsresearch.substack.com
Author
LLMs Research
Category
Podcast website
Latest episode
Feb 11, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Your 70-Billion-Parameter Model Might Be 40% Wasted 11.02.2026 12:24
Your 70-Billion-Parameter Model Might Be 40% Wasted Three papers from February 1–6, 2026 converge on a question the field has been avoiding since 2016: what if most transformer layers aren't doing compositional reasoning at all, but just averaging noise? This video traces a decade of evidence, from Veit et al.'s original ensemble observation in ResNets through ShortGPT's layer pruning results and...
Fixing Reasoning from Three Directions at Once 07.02.2026 12:49
Fixing Reasoning from Three Directions at Once LLMs Research Podcast | Episode: Feb 1–6, 2026 DeepSeek-R1 made RL the default approach for reasoning. This episode covers the 30 papers from the first week of February that are debugging what that approach got wrong, across training geometry, pipeline design, reward signals, memory architecture, and inference efficiency. Timestamps [00:00] Opening an...
Mamba's Memory Problem 02.02.2026 7:18
State space models like Mamba promised linear scaling and constant memory. They delivered on efficiency, but researchers kept hitting the same wall: ask Mamba to recall something specific from early in a long context, and performance drops. Three papers at ICLR 2026 independently attacked this limitation. That convergence tells you how fundamental the problem is. This podcast breaks down: - Why Ma...
What ICLR 2026 Taught Us About Multi-Agent Failures 31.01.2026 17:13
Episode Title: What ICLR 2026 Taught Us About Multi-Agent Failures Episode Summary: We scanned ICLR 2026 accepted papers and found 14 that address real problems when building multi-agent systems: slow pipelines, expensive token bills, cascading errors, brittle topologies, and opaque agent coordination. This episode walks through five production problems and the research that provides concrete solu...
The Evolution of Long-Context LLMs: From 512 to 10M Tokens 24.01.2026 15:43
The podcast discusses the technical shift in large language models from a standard 512-token context window to modern architectures capable of processing millions of tokens. Initial growth was constrained by the quadratic complexity of self-attention, which dictated that memory and computational needs increased by the square of the sequence length. To address this bottleneck, researchers developed...
Jan 17–23, 2026: The Rise of the Action Layer 24.01.2026 15:57
Jan 17–23, 2026: The Rise of the Action LayerHow multimodal AI is shifting from passive perception to active controlThe current landscape of research reflects a shift toward Multimodal Agentic Intelligence , where multimodal capabilities are no longer treated as simple perception but as actionable interfaces for control and interaction. This trend involves the integration of visual representations...
From Transformers to Autonomous Agents: A Timeline of the Research That Got Us Here 18.01.2026 24:34
Over the past several years, we have moved from the machine learning era through the large language model era and into what researchers now call the agent era. But we did not arrive here overnight. A series of research contributions, each building on the last, have made agents progressively more capable. This podcast traces that timeline through the papers that shaped today's autonomous systems. G...
LLM Research Highlights: January 10th to 16th, 2026 17.01.2026 22:34
The focus of LLM research is undergoing a significant shift from simply increasing model size to making existing capabilities practically usable for deployment . This week’s papers highlight four emerging themes: Reasoning & Agents, Generation & Synthesis, LLM Decisions, and Structured Episodic Memory . The overarching goal is to solve the real-world bottlenecks of agentic systems, such as high in...
LLM Research Highlights: January 2nd to 9th, 2026 10.01.2026 17:56
Key takeaway from todays podcast: Robust Reasoning: The Holonomic Network achieves perfect fidelity extrapolation 100x beyond training lengths, demonstrating a new universality class for logical reasoning with topological stability. Efficient Inference: RelayLLM cuts inference cost by 98.2% invoking LLM only 1.07% of tokens, while GlimpRouter reduces latency by 25.9% and boosts accuracy by 10.7% v...
LLM Research Highlights: December 27, 2025 – January 2, 2026 04.01.2026 15:39
* Bayesian & Cognitive Advances: Transformers achieve ultra-precise Bayesian inference with 10⁻³ to 10⁻⁴ bit accuracy, while CREST boosts reasoning accuracy by 17.5% and cuts token usage by 37.6%. * Model Efficiency & Scaling: TG reduces data needs by up to 8% and parameters by 42% compared to GPT-2, and Recursive Language Models handle 100x longer inputs at similar or lower inference costs. * Sta...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.