Igor Melnyk
Arxiv Papers
Running out of time to catch up with new arXiv papers? We take the most impactful papers and present them as convenient podcasts. If you're a visual learner, we offer these papers in an engaging video format. Our service fills the gap between overly brief paper summaries and time-consuming full paper reads. You gain academic insights in a time-efficient, digestible format. Code behind this work: https://github.com/imelnyk/ArxivPapers
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
[QA] Grokking at the Edge of Numerical Stability 09.01.2025 7:50
This paper explores grokking in deep learning, linking delayed generalization to Softmax Collapse and proposing solutions to enable grokking without regularization through new activation functions and training algorithms. https://arxiv.org/abs//2501.04697 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast...
Grokking at the Edge of Numerical Stability 09.01.2025 16:50
This paper explores grokking in deep learning, linking delayed generalization to Softmax Collapse and proposing solutions to enable grokking without regularization through new activation functions and training algorithms. https://arxiv.org/abs//2501.04697 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast...
[QA] ComMer: a Framework for Compressing and Merging User Data for Personalization 08.01.2025 7:21
ComMer is a framework that efficiently personalizes Large Language Models by compressing user documents into compact representations, improving performance in skill learning tasks while facing challenges in knowledge-intensive applications. https://arxiv.org/abs//2501.03276 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.a...
ComMer: a Framework for Compressing and Merging User Data for Personalization 08.01.2025 16:03
ComMer is a framework that efficiently personalizes Large Language Models by compressing user documents into compact representations, improving performance in skill learning tasks while facing challenges in knowledge-intensive applications. https://arxiv.org/abs//2501.03276 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.a...
[QA] Entropy-Guided Attention for Private LLMs 08.01.2025 7:57
This paper addresses privacy concerns in proprietary language models by optimizing transformer architectures for private inference, focusing on the role of nonlinearities and introducing entropy-guided mechanisms for improved performance. https://arxiv.org/abs//2501.03489 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.app...
Entropy-Guided Attention for Private LLMs 08.01.2025 13:20
This paper addresses privacy concerns in proprietary language models by optimizing transformer architectures for private inference, focusing on the role of nonlinearities and introducing entropy-guided mechanisms for improved performance. https://arxiv.org/abs//2501.03489 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.app...
[QA] Easing Optimization Paths: a Circuit Perspective 07.01.2025 7:33
The paper explores using mechanistic interpretability to enhance gradient descent training in AI, aiming to reduce compute costs and mitigate harmful behaviors through efficient learning curricula. https://arxiv.org/abs//2501.02362 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id16924760...
Easing Optimization Paths: a Circuit Perspective 07.01.2025 9:58
The paper explores using mechanistic interpretability to enhance gradient descent training in AI, aiming to reduce compute costs and mitigate harmful behaviors through efficient learning curricula. https://arxiv.org/abs//2501.02362 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id16924760...
[QA] Randomly Sampled Language Reasoning Problems Reveal Limits of LLMs 07.01.2025 7:15
This study evaluates LLMs' language understanding using novel tasks from deterministic finite automata, revealing they struggle compared to basic models when faced with unfamiliar languages. https://arxiv.org/abs//2501.02825 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spot...
Randomly Sampled Language Reasoning Problems Reveal Limits of LLMs 07.01.2025 16:14
This study evaluates LLMs' language understanding using novel tasks from deterministic finite automata, revealing they struggle compared to basic models when faced with unfamiliar languages. https://arxiv.org/abs//2501.02825 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spot...
[QA] Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search 06.01.2025 6:49
This study enhances large language models' reasoning abilities using Monte Carlo Tree Search for process supervision, significantly improving performance on mathematical reasoning tasks and demonstrating transferability across datasets. https://arxiv.org/abs//2501.01478 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple...
Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search 06.01.2025 7:04
This study enhances large language models' reasoning abilities using Monte Carlo Tree Search for process supervision, significantly improving performance on mathematical reasoning tasks and demonstrating transferability across datasets. https://arxiv.org/abs//2501.01478 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple...
[QA] Predicting the Performance of Black-box LLMs through Self-Queries 06.01.2025 7:14
This paper presents a method to predict large language model errors using black-box feature extraction, outperforming traditional approaches and enabling nuanced evaluations of model performance and architecture. https://arxiv.org/abs//2501.01558 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-pa...
Predicting the Performance of Black-box LLMs through Self-Queries 06.01.2025 20:27
This paper presents a method to predict large language model errors using black-box feature extraction, outperforming traditional approaches and enabling nuanced evaluations of model performance and architecture. https://arxiv.org/abs//2501.01558 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-pa...
[QA] On Unifying Video Generation and Camera Pose Estimation 04.01.2025 8:25
The paper introduces JOG3R, a unified model that enhances video generation and camera pose estimation, demonstrating improved accuracy in 3D awareness through fine-tuning and feature repurposing. https://arxiv.org/abs//2501.01409 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016...
On Unifying Video Generation and Camera Pose Estimation 04.01.2025 21:56
The paper introduces JOG3R, a unified model that enhances video generation and camera pose estimation, demonstrating improved accuracy in 3D awareness through fine-tuning and feature repurposing. https://arxiv.org/abs//2501.01409 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016...
[QA] An analytic theory of creativity in convolutional diffusion models 04.01.2025 8:05
The paper presents a predictive theory of creativity in convolutional diffusion models, identifying inductive biases that enable novel image generation beyond training data through local patch combinations. https://arxiv.org/abs//2412.20292 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/i...
An analytic theory of creativity in convolutional diffusion models 04.01.2025 22:28
The paper presents a predictive theory of creativity in convolutional diffusion models, identifying inductive biases that enable novel image generation beyond training data through local patch combinations. https://arxiv.org/abs//2412.20292 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/i...
[QA] Finding Missed Code Size Optimizations in Compilers using LLMs 03.01.2025 7:30
https://arxiv.org/abs//2501.00655 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
Finding Missed Code Size Optimizations in Compilers using LLMs 03.01.2025 18:28
https://arxiv.org/abs//2501.00655 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
[QA] Titans: Learning to Memorize at Test Time 03.01.2025 7:21
The paper introduces Titans, a new architecture combining neural long-term memory and attention, outperforming Transformers in various tasks while efficiently handling larger context windows. https://arxiv.org/abs//2501.00663 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spo...
Titans: Learning to Memorize at Test Time 03.01.2025 30:18
The paper introduces Titans, a new architecture combining neural long-term memory and attention, outperforming Transformers in various tasks while efficiently handling larger context windows. https://arxiv.org/abs//2501.00663 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spo...
[QA] Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism 31.12.2024 7:42
This paper proposes adaptive batch size schedules for large-scale language model training, enhancing efficiency and generalization, while outperforming traditional methods in pretraining models, particularly smaller ones. https://arxiv.org/abs//2412.21124 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast...
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism 31.12.2024 20:11
This paper proposes adaptive batch size schedules for large-scale language model training, enhancing efficiency and generalization, while outperforming traditional methods in pretraining models, particularly smaller ones. https://arxiv.org/abs//2412.21124 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast...
[QA] Functional Risk Minimization 31.12.2024 8:20
The paper introduces Functional Risk Minimization (FRM), a framework improving performance in various learning tasks by comparing functions instead of outputs, enhancing understanding of generalization in over-parameterized models. https://arxiv.org/abs//2412.21149 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.