Igor Melnyk
Arxiv Papers
Running out of time to catch up with new arXiv papers? We take the most impactful papers and present them as convenient podcasts. If you're a visual learner, we offer these papers in an engaging video format. Our service fills the gap between overly brief paper summaries and time-consuming full paper reads. You gain academic insights in a time-efficient, digestible format. Code behind this work: https://github.com/imelnyk/ArxivPapers
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
[QA] Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning 21.04.2025 7:54
PODS decouples reinforcement learning phases by parallelizing rollouts and selectively updating, using max-variance down-sampling to enhance performance on the GSM8K benchmark compared to standard GRPO. https://arxiv.org/abs//2504.13818 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id169...
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning 21.04.2025 7:09
PODS decouples reinforcement learning phases by parallelizing rollouts and selectively updating, using max-variance down-sampling to enhance performance on the GSM8K benchmark compared to standard GRPO. https://arxiv.org/abs//2504.13818 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id169...
[QA] Let Me Grok for You: Accelerating Grokking via Embedding Transfer from a Weaker Model 21.04.2025 7:38
The paper presents a method to accelerate "grokking" in neural networks by using learned embeddings from a weaker model, enabling direct generalization without delay across various tasks. https://arxiv.org/abs//2504.13292 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify...
Let Me Grok for You: Accelerating Grokking via Embedding Transfer from a Weaker Model 21.04.2025 16:13
The paper presents a method to accelerate "grokking" in neural networks by using learned embeddings from a weaker model, enabling direct generalization without delay across various tasks. https://arxiv.org/abs//2504.13292 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id16924760...
[QA] Reasoning Models Can Be Effective Without Thinking 20.04.2025 7:29
This paper challenges the necessity of lengthy reasoning processes in LLMs, showing that simple prompting (NoThinking) can outperform traditional methods in various reasoning tasks, especially in low-budget scenarios. https://arxiv.org/abs//2504.09858 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arx...
Reasoning Models Can Be Effective Without Thinking 20.04.2025 20:05
This paper challenges the necessity of lengthy reasoning processes in LLMs, showing that simple prompting (NoThinking) can outperform traditional methods in various reasoning tasks, especially in low-budget scenarios. https://arxiv.org/abs//2504.09858 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arx...
[QA] A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce 20.04.2025 8:27
This paper analyzes GRPO in reinforcement learning for language models, revealing that a simple rejection sampling method, RAFT, performs competitively and suggesting improvements for future reward-based training approaches. https://arxiv.org/abs//2504.11343 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podc...
A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce 20.04.2025 14:38
This paper analyzes GRPO in reinforcement learning for language models, revealing that a simple rejection sampling method, RAFT, performs competitively and suggesting improvements for future reward-based training approaches. https://arxiv.org/abs//2504.11343 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podc...
[QA] CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training 19.04.2025 7:14
https://arxiv.org/abs//2504.13161 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training 19.04.2025 20:35
https://arxiv.org/abs//2504.13161 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
[QA] Antidistillation Sampling 18.04.2025 7:21
Antidistillation sampling modifies token probability distributions to weaken reasoning traces for model distillation, enhancing model security while maintaining performance. https://arxiv.org/abs//2504.13146 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podc...
Antidistillation Sampling 18.04.2025 10:44
Antidistillation sampling modifies token probability distributions to weaken reasoning traces for model distillation, enhancing model security while maintaining performance. https://arxiv.org/abs//2504.13146 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podc...
[QA] Position: The Most Expensive Part of an LLM should be its Training Data 18.04.2025 7:16
This paper argues that compensating human labor for training data is the largest cost in developing Large Language Models, significantly exceeding model training expenses, and suggests fairer practices for the future. https://arxiv.org/abs//2504.12427 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arx...
Position: The Most Expensive Part of an LLM should be its Training Data 18.04.2025 20:05
This paper argues that compensating human labor for training data is the largest cost in developing Large Language Models, significantly exceeding model training expenses, and suggests fairer practices for the future. https://arxiv.org/abs//2504.12427 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arx...
[QA] Activated LoRA: Fine-tuned LLMs for Intrinsics 18.04.2025 8:16
Activated LoRA (aLoRA) enhances LoRA by adapting weights only for relevant tokens, allowing instant activation without recomputing the KV cache, improving efficiency in multiturn settings. https://arxiv.org/abs//2504.12397 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotif...
Activated LoRA: Fine-tuned LLMs for Intrinsics 18.04.2025 18:55
Activated LoRA (aLoRA) enhances LoRA by adapting weights only for relevant tokens, allowing instant activation without recomputing the KV cache, improving efficiency in multiturn settings. https://arxiv.org/abs//2504.12397 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotif...
[QA] COLORBENCH: Can VLMs See and Understand the Colorful World? 17.04.2025 7:49
The paper presents COLORBENCH, a benchmark to evaluate vision-language models' color understanding, revealing limitations and emphasizing the need for improved color comprehension in multimodal AI. https://arxiv.org/abs//2504.10514 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692...
COLORBENCH: Can VLMs See and Understand the Colorful World? 17.04.2025 20:40
The paper presents COLORBENCH, a benchmark to evaluate vision-language models' color understanding, revealing limitations and emphasizing the need for improved color comprehension in multimodal AI. https://arxiv.org/abs//2504.10514 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692...
[QA] ReTool: Reinforcement Learning for Strategic Tool Use in LLMs 17.04.2025 8:33
https://arxiv.org/abs//2504.11536 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs 17.04.2025 14:57
https://arxiv.org/abs//2504.11536 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
[QA] Looking beyond the next token 16.04.2025 7:22
The paper presents TRELAWNEY, a method for rearranging training data to improve causal language models' performance in planning and reasoning without altering architecture, enhancing goal generation capabilities. https://arxiv.org/abs//2504.11336 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxi...
Looking beyond the next token 16.04.2025 16:58
The paper presents TRELAWNEY, a method for rearranging training data to improve causal language models' performance in planning and reasoning without altering architecture, enhancing goal generation capabilities. https://arxiv.org/abs//2504.11336 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxi...
[QA] How to Predict Best Pretraining Data with Small Experiments 16.04.2025 8:16
The paper introduces DATADECIDE, a suite for evaluating data selection methods, revealing that small-scale model rankings effectively predict larger model performance, enhancing cost-efficient pretraining decisions. https://arxiv.org/abs//2504.11393 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv...
How to Predict Best Pretraining Data with Small Experiments 16.04.2025 20:22
The paper introduces DATADECIDE, a suite for evaluating data selection methods, revealing that small-scale model rankings effectively predict larger model performance, enhancing cost-efficient pretraining decisions. https://arxiv.org/abs//2504.11393 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv...
[QA] Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability 15.04.2025 7:18
This study evaluates OpenAI's GPT-4o, revealing limitations in semantic synthesis, instruction adherence, and reasoning, challenging assumptions about its multimodal capabilities and calling for improved benchmarks and training strategies. https://arxiv.org/abs//2504.08003 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcast...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.