Igor Melnyk
Arxiv Papers
Running out of time to catch up with new arXiv papers? We take the most impactful papers and present them as convenient podcasts. If you're a visual learner, we offer these papers in an engaging video format. Our service fills the gap between overly brief paper summaries and time-consuming full paper reads. You gain academic insights in a time-efficient, digestible format. Code behind this work: https://github.com/imelnyk/ArxivPapers
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
[short] Transforming and Combining Rewards for Aligning Large Language Models 02.02.2024 2:14
The paper explores aligning language models to human preferences using reward models and addresses issues of monotone transformations and combining multiple reward models. https://arxiv.org/abs//2402.00742 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcas...
Transforming and Combining Rewards for Aligning Large Language Models 02.02.2024 27:13
The paper explores aligning language models to human preferences using reward models and addresses issues of monotone transformations and combining multiple reward models. https://arxiv.org/abs//2402.00742 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcas...
[short] Efficient Exploration for LLMs 02.02.2024 2:19
https://arxiv.org/abs//2402.00396 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
Efficient Exploration for LLMs 02.02.2024 16:39
https://arxiv.org/abs//2402.00396 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research 02.02.2024 27:13
The paper introduces Dolma, a large English corpus for language model pretraining, and shares insights on data curation practices. It also presents OLMo, a state-of-the-art open language model. https://arxiv.org/abs//2402.00159 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 S...
[short] Infini-gram: Scaling Unbounded -gram Language Models to a Trillion Tokens 01.02.2024 2:06
Language models remain relevant, especially for text analysis and improving neural large language models. Modernizing -gram models involves training at the same scale as neural LLMs and using a new -gram LM with backoff. https://arxiv.org/abs//2401.17377 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/...
Infini-gram: Scaling Unbounded -gram Language Models to a Trillion Tokens 01.02.2024 38:18
Language models remain relevant, especially for text analysis and improving neural large language models. Modernizing -gram models involves training at the same scale as neural LLMs and using a new n-gram LM with backoff. https://arxiv.org/abs//2401.17377 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast...
[short] Arrows of Time for Large Language Models 01.02.2024 2:59
Autoregressive language models exhibit time asymmetry in predicting next versus previous tokens, contrary to information theory expectations. Theoretical framework and implications are discussed. https://arxiv.org/abs//2401.17505 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016...
Arrows of Time for Large Language Models 01.02.2024 30:55
Autoregressive language models exhibit time asymmetry in predicting next versus previous tokens, contrary to information theory expectations. Theoretical framework and implications are discussed. https://arxiv.org/abs//2401.17505 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016...
[short] WEAVER: Foundation Models for Creative Writing 31.01.2024 3:24
https://arxiv.org/abs//2401.17268 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
WEAVER: Foundation Models for Creative Writing 31.01.2024 30:37
https://arxiv.org/abs//2401.17268 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
[short] H2O-Danube-1.8B Technical Report 31.01.2024 3:07
https://arxiv.org/abs//2401.16818 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
H2O-Danube-1.8B Technical Report 31.01.2024 12:22
https://arxiv.org/abs//2401.16818 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
[short] Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling 30.01.2024 2:04
Large language models require abundant compute and data for training, which is infeasible due to costs and data scarcity. The proposed WRAP method uses rephrased web data to improve pre-training efficiency and model performance. https://arxiv.org/abs//2401.16380 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/...
Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling 30.01.2024 28:49
Large language models require abundant compute and data for training, which is infeasible due to costs and data scarcity. The proposed WRAP method uses rephrased web data to improve pre-training efficiency and model performance. https://arxiv.org/abs//2401.16380 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/...
[short] SYNCHFORMER: EFFICIENT SYNCHRONIZATION FROM SPARSE CUES 30.01.2024 2:08
The paper presents a novel audio-visual synchronization model for 'in-the-wild' videos, achieving state-of-the-art performance and exploring interpretability and synchronizability. https://arxiv.org/abs//2401.16423 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotif...
SYNCHFORMER: EFFICIENT SYNCHRONIZATION FROM SPARSE CUES 30.01.2024 14:05
The paper presents a novel audio-visual synchronization model for 'in-the-wild' videos, achieving state-of-the-art performance and exploring interpretability and synchronizability. https://arxiv.org/abs//2401.16423 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotif...
[short] MoE-LLaVA: Mixture of Experts for Large Vision-Language Models 30.01.2024 2:36
Proposed MoE-tuning strategy for Large Vision-Language Models (LVLMs) creates a sparse model with constant computational cost, addressing performance degradation and reducing hallucinations. https://arxiv.org/abs//2401.15947 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spot...
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models 30.01.2024 16:23
Proposed MoE-tuning strategy for Large Vision-Language Models (LVLMs) creates a sparse model with constant computational cost, addressing performance degradation and reducing hallucinations. https://arxiv.org/abs//2401.15947 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spot...
From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities 29.01.2024 4:17
This paper examines the reliability of multi-modal large language models (MLLMs) across text, code, image, and video, aiming to improve transparency and understanding. https://arxiv.org/abs//2401.15071 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters...
[short] From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities 29.01.2024 2:48
This paper examines the reliability of multi-modal large language models (MLLMs) across text, code, image, and video, aiming to improve transparency and understanding. https://arxiv.org/abs//2401.15071 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters...
[short] EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty 29.01.2024 2:15
EAGLE is a lossless acceleration framework for Large Language Models, achieving faster decoding without fine-tuning and maintaining the same text distribution. https://arxiv.org/abs//2401.15077 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify...
EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty 29.01.2024 21:16
EAGLE is a lossless acceleration framework for Large Language Models, achieving faster decoding without fine-tuning and maintaining the same text distribution. https://arxiv.org/abs//2401.15077 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify...
[short] Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities 28.01.2024 2:41
The paper proposes a method, Multimodal Pathway, to improve transformers for a specific modality using irrelevant data from other modalities, resulting in consistent performance improvements. https://arxiv.org/abs//2401.14405 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spo...
Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities 28.01.2024 20:04
The paper proposes a method, Multimodal Pathway, to improve transformers for a specific modality using irrelevant data from other modalities, resulting in consistent performance improvements. https://arxiv.org/abs//2401.14405 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spo...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.