Igor Melnyk

Arxiv Papers

Science EN ↓ 2489 episodes

Running out of time to catch up with new arXiv papers? We take the most impactful papers and present them as convenient podcasts. If you're a visual learner, we offer these papers in an engaging video format. Our service fills the gap between overly brief paper summaries and time-consuming full paper reads. You gain academic insights in a time-efficient, digestible format. Code behind this work: https://github.com/imelnyk/ArxivPapers

Author

Igor Melnyk

Category

Science

Podcast website

github.com

Latest episode

Sep 1, 2025

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

[QA] Grokking at the Edge of Numerical Stability 09.01.2025

This paper explores grokking in deep learning, linking delayed generalization to Softmax Collapse and proposing solutions to enable grokking without regularization through new activation functions and training algorithms. https://arxiv.org/abs//2501.04697 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast...

Grokking at the Edge of Numerical Stability 09.01.2025

This paper explores grokking in deep learning, linking delayed generalization to Softmax Collapse and proposing solutions to enable grokking without regularization through new activation functions and training algorithms. https://arxiv.org/abs//2501.04697 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast...

[QA] ComMer: a Framework for Compressing and Merging User Data for Personalization 08.01.2025

ComMer is a framework that efficiently personalizes Large Language Models by compressing user documents into compact representations, improving performance in skill learning tasks while facing challenges in knowledge-intensive applications. https://arxiv.org/abs//2501.03276 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.a...

ComMer: a Framework for Compressing and Merging User Data for Personalization 08.01.2025

ComMer is a framework that efficiently personalizes Large Language Models by compressing user documents into compact representations, improving performance in skill learning tasks while facing challenges in knowledge-intensive applications. https://arxiv.org/abs//2501.03276 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.a...

[QA] Entropy-Guided Attention for Private LLMs 08.01.2025

This paper addresses privacy concerns in proprietary language models by optimizing transformer architectures for private inference, focusing on the role of nonlinearities and introducing entropy-guided mechanisms for improved performance. https://arxiv.org/abs//2501.03489 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.app...

Entropy-Guided Attention for Private LLMs 08.01.2025

This paper addresses privacy concerns in proprietary language models by optimizing transformer architectures for private inference, focusing on the role of nonlinearities and introducing entropy-guided mechanisms for improved performance. https://arxiv.org/abs//2501.03489 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.app...

[QA] Easing Optimization Paths: a Circuit Perspective 07.01.2025

The paper explores using mechanistic interpretability to enhance gradient descent training in AI, aiming to reduce compute costs and mitigate harmful behaviors through efficient learning curricula. https://arxiv.org/abs//2501.02362 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id16924760...

Easing Optimization Paths: a Circuit Perspective 07.01.2025

The paper explores using mechanistic interpretability to enhance gradient descent training in AI, aiming to reduce compute costs and mitigate harmful behaviors through efficient learning curricula. https://arxiv.org/abs//2501.02362 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id16924760...

[QA] Randomly Sampled Language Reasoning Problems Reveal Limits of LLMs 07.01.2025

This study evaluates LLMs' language understanding using novel tasks from deterministic finite automata, revealing they struggle compared to basic models when faced with unfamiliar languages. https://arxiv.org/abs//2501.02825 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spot...

Randomly Sampled Language Reasoning Problems Reveal Limits of LLMs 07.01.2025

This study evaluates LLMs' language understanding using novel tasks from deterministic finite automata, revealing they struggle compared to basic models when faced with unfamiliar languages. https://arxiv.org/abs//2501.02825 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spot...

[QA] Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search 06.01.2025

This study enhances large language models' reasoning abilities using Monte Carlo Tree Search for process supervision, significantly improving performance on mathematical reasoning tasks and demonstrating transferability across datasets. https://arxiv.org/abs//2501.01478 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple...

Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search 06.01.2025

This study enhances large language models' reasoning abilities using Monte Carlo Tree Search for process supervision, significantly improving performance on mathematical reasoning tasks and demonstrating transferability across datasets. https://arxiv.org/abs//2501.01478 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple...

[QA] Predicting the Performance of Black-box LLMs through Self-Queries 06.01.2025

This paper presents a method to predict large language model errors using black-box feature extraction, outperforming traditional approaches and enabling nuanced evaluations of model performance and architecture. https://arxiv.org/abs//2501.01558 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-pa...

Predicting the Performance of Black-box LLMs through Self-Queries 06.01.2025

This paper presents a method to predict large language model errors using black-box feature extraction, outperforming traditional approaches and enabling nuanced evaluations of model performance and architecture. https://arxiv.org/abs//2501.01558 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-pa...

[QA] On Unifying Video Generation and Camera Pose Estimation 04.01.2025

The paper introduces JOG3R, a unified model that enhances video generation and camera pose estimation, demonstrating improved accuracy in 3D awareness through fine-tuning and feature repurposing. https://arxiv.org/abs//2501.01409 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016...

On Unifying Video Generation and Camera Pose Estimation 04.01.2025

The paper introduces JOG3R, a unified model that enhances video generation and camera pose estimation, demonstrating improved accuracy in 3D awareness through fine-tuning and feature repurposing. https://arxiv.org/abs//2501.01409 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016...

[QA] An analytic theory of creativity in convolutional diffusion models 04.01.2025

The paper presents a predictive theory of creativity in convolutional diffusion models, identifying inductive biases that enable novel image generation beyond training data through local patch combinations. https://arxiv.org/abs//2412.20292 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/i...

An analytic theory of creativity in convolutional diffusion models 04.01.2025

The paper presents a predictive theory of creativity in convolutional diffusion models, identifying inductive biases that enable novel image generation beyond training data through local patch combinations. https://arxiv.org/abs//2412.20292 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/i...

[QA] Finding Missed Code Size Optimizations in Compilers using LLMs 03.01.2025

https://arxiv.org/abs//2501.00655 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers

Finding Missed Code Size Optimizations in Compilers using LLMs 03.01.2025

https://arxiv.org/abs//2501.00655 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers

[QA] Titans: Learning to Memorize at Test Time 03.01.2025

The paper introduces Titans, a new architecture combining neural long-term memory and attention, outperforming Transformers in various tasks while efficiently handling larger context windows. https://arxiv.org/abs//2501.00663 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spo...

Titans: Learning to Memorize at Test Time 03.01.2025

The paper introduces Titans, a new architecture combining neural long-term memory and attention, outperforming Transformers in various tasks while efficiently handling larger context windows. https://arxiv.org/abs//2501.00663 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spo...

[QA] Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism 31.12.2024

This paper proposes adaptive batch size schedules for large-scale language model training, enhancing efficiency and generalization, while outperforming traditional methods in pretraining models, particularly smaller ones. https://arxiv.org/abs//2412.21124 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast...

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism 31.12.2024

This paper proposes adaptive batch size schedules for large-scale language model training, enhancing efficiency and generalization, while outperforming traditional methods in pretraining models, particularly smaller ones. https://arxiv.org/abs//2412.21124 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast...

[QA] Functional Risk Minimization 31.12.2024

The paper introduces Functional Risk Minimization (FRM), a framework improving performance in various learning tasks by comparing functions instead of outputs, enhancing understanding of generalization in over-parameterized models. https://arxiv.org/abs//2412.21149 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/...

Listen to the Arxiv Papers podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.