Mechanical Dirk

Mechanical Dreams

Science EN ↓ 153 episodes

An automatically generated podcast about machine learning and natural language processing. The two fictional hosts talk about papers that I want to learn more about on my way to work. It's not good, but it's useful.

Author

Mechanical Dirk

Category

Science

Latest episode

May 27, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Why Linearly Decaying the Learning Rate to Zero Works Best 16.04.2025
Not All Data Are Unlearned Equally 15.04.2025
A Multi-Power Law for Loss Curve Prediction 14.04.2025
Efficient Training of Ultra-Long Context Large Language Models 11.04.2025
Multi-Token Attention 03.04.2025
From Style to Facts 02.04.2025
Compute Optimal Scaling of Skills 22.03.2025
Predictive Data Selection 15.03.2025
Continual Pre-training of MoEs 12.03.2025
s1 - Simple test-time scaling 06.03.2025
Cognitive Behaviors that Enable Self-Improving Reasoners 05.03.2025
Phi 4 Multimodal Instruct 04.03.2025
Claude 3.7 Sonnet System Card 25.02.2025
Project Sid: Many-agent simulations toward AI civilization 09.02.2025
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training 09.02.2025
Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models 05.02.2025
NExtLong - Toward Effective Long-Context Training without Long Documents 30.01.2025
Over-Tokenized Transformer 30.01.2025
Optimizing Pretraining Data Mixtures with LLM-Estimated Utility 30.01.2025
HashAttention: Semantic Sparsity for Faster Inference 16.01.2025
From Tokens to Words 15.01.2025
DeepSeek V3 07.01.2025
Optimal Linear Decay Learning Rate Schedules and Further Refinements 05.01.2025

In this episode: • The Death of Cosine?: Introduction to the episode and the paper. Professor Norris expresses his skepticism about changing established habits like Cosine Annealing, while Linda teases a shake-up in the status quo. • Theory vs. Reality: A discussion on the massive gap between theoretical learning rates (like 1/t) and what practitioners actually use. Linda explains why the theory h...

Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers 20.12.2024
Efficient and Approximate Per-Example Gradient Norms for Gradient Noise Scale 20.12.2024

Listen to the Mechanical Dreams podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.