mcgrof
AI Post Transformers
AI-generated podcast where hosts Hal Turing and Dr. Ada Shannon discuss the latest research papers and reports in machine learning, AI systems, and optimization. Featuring honest critical analysis, proper citations, and nerdy humor.
Author
mcgrof
Category
Podcast website
Latest episode
Jul 9, 2026
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Google: R&D inference value on HBF + PNM + low latency interconnect 23.01.2026
To address the hardware bottlenecks of LLM inference, Google researchers Ma and Patterson propos in their paper "Challenges and Research Directions for Large Language Model Inference Hardware" published on January 8, 2026 a few focus areas of research: High Bandwidth Flash (HBF), Processing-Near-Memory (PNM), and low-latency interconnects. HBF addresses the "Memory Wall" by sta...
Storage-next: Do We Need New Hardware for AI Storage, or Just Better Layouts? 21.01.2026
We review the "Storage-Next" paper, published in November 2025, which argues that a fundamental hardware architectural shift is required to elevate NAND flash from a passive storage tier to an active memory tier capable of "seconds-scale" caching. The authors contend that standard SSDs impose a "channel-side ceiling" on IOPS because they are optimized for 4KB blocks,...
Meta's solution to massive DLRM inference through software defined memory 21.01.2026
On November, 2021 Meta (back then Facebook) in collaboration with George Mason University and University of Illinois Chicago published their paper "Supporting Massive DLRM inference through software defined memory". Meta addressed the infrastructure challenge of serving massive Deep Learning Recommendation Models by extending the memory hierarchy to include NVMe Storage Class Memory. Bec...
LeCun's AMI Energy-Based Models and the Path to Autonomous Intelligence 21.01.2026
These sources collectively explore the current landscape and future trajectory of artificial intelligence, specifically focusing on the transition toward human-level reasoning. Renowned scientist Yann LeCun argues that current Large Language Models lack a fundamental understanding of the physical world and proposes a shift toward objective-driven AI that utilizes world models for better planning a...
Process Reward Learning for LLM Reasoning Optimization 19.01.2026
Researchers from the University of Illinois Urbana-Champaign have introduced Process Reward Learning (PRL) on a January 15, 2026 paper. PRL is a novel training framework designed to enhance the reasoning capabilities of Large Language Models. Unlike traditional reinforcement learning that relies on sparse outcome-based rewards, PRL provides dense, fine-grained supervision by decomposing global obj...
Parallel Context-of-Experts Decoding for Efficient RAG reasoning 19.01.2026
The collaboration between SAP Labs, France and EURECOM, France published a paper on January 13, 2026 titord "Parallel Context-of-Experts Decoding for Retrieval Augmented Generation". The paper introduces Parallel Context-of-Experts Decoding (PCED), a training-free framework designed to optimize Retrieval Augmented Generation (RAG) by overcoming the latency and reasoning limitations of lo...
OpenRouter 2025 Report: Analysis of Global LLM Usage Patterns 19.01.2026
We review two research papers, one from January 15, 2026 by OpenRouter Inc and a16z (Andreessen Horowitz) and another from April 2025 by Andrey Fradkin (Boston University and MIT IDE) which provide uses of OpenRouter, they analyze the evolving market dynamics and user behaviors within the Large Language Model (LLM) ecosystem, primarily using data from the OpenRouter marketplace. The studies docume...
Multidimensional Safety Evaluation of Frontier AI Models 19.01.2026
This January 17, 2026 research collaboration between Fudan University, Shanghai Innovation institute, Deakin University and UIUC provide a report which provides a comprehensive safety evaluation of several frontier AI models, including GPT-5.2 and Grok 4.1 Fast, across text, image, and multilingual domains. The study reveals a persistent alignment paradox where a model's desire to be helpful...
MemoBrain: Executive Memory for Tool-Augmented Reasoning Agents 19.01.2026
On January 12, 2026, a collaboration between the Beijing Academy of Artificial Intelligence, the Gaoling School of Artificial Intelligence, and Renmin University of China introduced MemoBrain. In the paper titled “MemoBrain: Executive Memory as an Agentic Brain for Reasoning,” the authors present an executive memory model designed to enhance the performance of tool-augmented AI agents during compl...
MATTRL: Collaborative Test-Time Reinforcement Learning for Multi-Agent Reasoning 19.01.2026
The collaboration between MIT, NUS, NYU, Microsoft, UW , Columbia and NTU describes an inference retrieval Chain of Thought enhancement. The researchers introduce MATTRL, a framework designed to improve the reasoning of Large Language Models (LLMs) through multi-agent collaboration and reinforcement learning. This system organizes specialized AI agents into Multidisciplinary Teams (MDT) to tackle...
How to make LLMs better at sci-fi writing 19.01.2026
In a joint collaboration between Harvard University, Carnegie Mellon University, Stanford University the January 12, 2026 paper "LLM Review: Enhancing Creative Writing via Blind Peer Review Feedback" concludes: Interaction structure may substitute for model scale. Researchers at these institutions have collaborated to developed a new multi-agent framework called LLM Review to address the...
H-net: End-to-End Hierarchical Sequence Modeling via Dynamic Chunking 19.01.2026
On this July 15, 2025 collaboration between Carnegie Mellon University and Cartesia AI researchers introduce H-net in the paper "Dynamic Chunking for End-to-End Hierarchical Sequence Modeling". H-Net is a hierarchical, tokenizer-free large language model that processes raw data like bytes or DNA sequences directly. Unlike traditional models that rely on predefined subword chunks, H-Net e...
GLM-Image: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning 19.01.2026
We review the January 1, 2026 paper "GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning" from the GLM-V Team as a collaboration between Zhipu AI & Tsinghua University which have released an a new open weight model. The GLM-Image model they produce could likely be the first SOTA multimodal model fully trained on Chinese-manu...
CompassMem: Event-Centric Logic Maps for Agent Memory 19.01.2026
The collaboration between Gaoling School of Artificial Intelligence, Renmin University of China published a paper in January 8, 2026 titled "Memory Matters More: Event-Centric Memory as a Logic Map for Agent Searching and Reasoning" which introduces CompassMem, an innovative event-centric memory framework designed to enhance how large language model agents process and retrieve informatio...
A Chain of Thought reasoning academic brawl 19.01.2026
We review a late 2025 heavyweight academic brawl over the future of AI reasoning when folks use Reinforcement Learning with Verifiable Rewards (RLVR) for Chain of Thought (CoT). We have two papers on the table, and they are in complete, direct conflict. In one corner, we have a study by Yue et al. from Tsinghua University, which dropped a bombshell claim: they argue that Reinforcement Learning wit...
Squisher: Approximating the Fisher Information Matrix and use cases 17.01.2026
We focus on the July 2025 paper, "Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator". The paper goes into the mathematical details of approximating the FIM, and the easiest is what they call the Squisher: the Adam optimizer variance. This technique allows for tasks like model pruning and model merging to be performed "for fre...
Scaling laws: long context length and in context learning 17.01.2026
Recent advancements in Long Context Language Models (LCLMs) demonstrate that In-Context Learning (ICL) capabilities follow predictable power-law scaling relationships, where performance improves monotonically with context length up to 10 million tokens and is governed by model depth, width, and training data volume. While Gemini 1.5 exhibits near-perfect recall and continued log-loss improvement a...
NVIDIA: TTT-E2E: Unlocking Long-Context Learning via End-to-End Test-Time Training 17.01.2026
This December 31, 2025 NVIDIA research introduces TTT-E2E, a novel approach to large language model memory that treats long-context processing as a continual learning problem rather than a structural design challenge. By utilizing test-time training, the model effectively compresses context into its own weights through next-token prediction, allowing it to adapt and learn while processing new info...
Attention with a bias 17.01.2026
We review why some transformer models use a bias in attention and how ALiBi helps with long context. The provided sources focus on significant advancements in computational biology, specifically the evolution of the AlphaFold series for predicting 3D biomolecular structures. AlphaFold 2 revolutionized the field by using the Evoformer and attention mechanisms to interpret evolutionary and geometric...
DeepSeek Engram: Scaling Large Language Models via Conditional Memory Lookup 14.01.2026
On January 12, 2026 DeepSeek released its paper on Engram, a novel AI architecture that incorporates conditional memory to optimize how large language models handle information. By utilizing a lookup mechanism for static patterns, this technology separates an AI's logical reasoning from its factual knowledge base. This structural shift allows massive models to run on cheaper hardware by offlo...
PageANN: Scalable Disk ANNS with Page-Aligned Graphs 07.12.2025
The research paper presents PageANN, a novel framework engineered to overcome the severe latency and scalability limitations facing existing disk-based Approximate Nearest Neighbor Search (ANNS) methods used in vector databases. Current systems suffer from inefficient search paths and a crucial misalignment between logical graph node size and the physical I/O granularity of Solid-State Drives (SSD...
NeurIPS 2025: Homogeneous Keys, Heterogeneous Values 04.12.2025
This research presents a novel method for efficient long-context modeling in Large Language Models (LLMs) by tackling the quadratic complexity of attention mechanisms through KV cache compression. The core discovery is a fundamental local KV cache asymmetry, which reveals that adjacent attention keys exhibit high structural homogeneity, while their associated value vectors possess distinct, hetero...
NeurIPS 2025: MoBA: Mixture of Block Attention for Long-Context LLMs 29.11.2025
This paper introduces Mixture of Block Attention (MoBA) to address the prohibitive quadratic computational overhead inherent in traditional attention mechanisms when scaling large language models (LLMs) for long contexts. MoBA is a novel architecture that strategically applies the established Mixture of Experts (MoE) paradigm directly to the attention mechanism itself. Instead of attending to the...
NeurIPS 2025: Reward Reasoning Model 29.11.2025
The source details the development and evaluation of Reward Reasoning Models (RRMs), which are designed to enhance Large Language Model (LLM) alignment by incorporating an explicit chain-of-thought reasoning process before generating a final reward. This innovative structure enables RRMs to adaptively utilize computational resources at inference time for complex evaluation tasks requiring nuanced...
NeurIPS 2025: Parallel Scaling Law for Language Models 29.11.2025
The research proposes Parallel Scaling (PARSCALE) as a novel, efficient strategy to enhance Large Language Model (LLM) capacity by increasing parallel computation rather than merely growing the parameter count. This method reuses existing model parameters by feeding multiple parallel input streams (differentiated by learned prefixes) and dynamically combining their outputs into a single prediction...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.