Igor Melnyk

Arxiv Papers

Science EN ↓ 2489 episodes

Running out of time to catch up with new arXiv papers? We take the most impactful papers and present them as convenient podcasts. If you're a visual learner, we offer these papers in an engaging video format. Our service fills the gap between overly brief paper summaries and time-consuming full paper reads. You gain academic insights in a time-efficient, digestible format. Code behind this work: https://github.com/imelnyk/ArxivPapers

Author

Igor Melnyk

Category

Science

Podcast website

github.com

Latest episode

Sep 1, 2025

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Not All Language Model Features Are Linear 24.05.2024

The paper challenges the linear representation hypothesis by exploring multi-dimensional features in language models like GPT-2, identifying circular features for days and months, and demonstrating their role in computational tasks. https://arxiv.org/abs//2405.14860 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com...

[QA] Implicit In-context Learning 24.05.2024

Implicit In-context Learning (I2CL) improves large language models by integrating demonstration examples within the activation space, achieving few-shot performance with zero-shot cost and enhancing transfer learning. https://arxiv.org/abs//2405.14660 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arx...

Implicit In-context Learning 24.05.2024

Implicit In-context Learning (I2CL) improves large language models by integrating demonstration examples within the activation space, achieving few-shot performance with zero-shot cost and enhancing transfer learning. https://arxiv.org/abs//2405.14660 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arx...

[QA] (Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts 23.05.2024

Advancements in machine translation improve quality, but literary translation remains challenging. TRANSAGENTS, a multi-agent framework using large language models, outperforms human translations in literary works. https://arxiv.org/abs//2405.11804 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-...

(Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts 23.05.2024

Advancements in machine translation improve quality, but literary translation remains challenging. TRANSAGENTS, a multi-agent framework using large language models, outperforms human translations in literary works. https://arxiv.org/abs//2405.11804 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-...

[QA] MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning 23.05.2024

The paper introduces MoRA, a method enhancing large language models by employing a square matrix for high-rank updating, outperforming LoRA on memory-intensive tasks. https://arxiv.org/abs//2405.12130 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters....

MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning 23.05.2024

The paper introduces MoRA, a method enhancing large language models by employing a square matrix for high-rank updating, outperforming LoRA on memory-intensive tasks. https://arxiv.org/abs//2405.12130 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters....

[QA] Reducing Transformer Key-Value Cache Size with Cross-Layer Attention 22.05.2024

Key-value caching in large language models is crucial for decoding speed. Multi-Query Attention (MQA) and Cross-Layer Attention (CLA) reduce memory usage while maintaining accuracy, enabling larger models. https://arxiv.org/abs//2405.12981 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id...

Reducing Transformer Key-Value Cache Size with Cross-Layer Attention 22.05.2024

Key-value caching in large language models is crucial for decoding speed. Multi-Query Attention (MQA) and Cross-Layer Attention (CLA) reduce memory usage while maintaining accuracy, enabling larger models. https://arxiv.org/abs//2405.12981 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id...

[QA] Your Transformer is Secretly Linear 22.05.2024

The paper uncovers a linear characteristic in transformer decoders, showing a near-perfect linear relationship between embedding transformations in sequential layers. Regularization reduces linearity and improves performance. https://arxiv.org/abs//2405.12250 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/pod...

Your Transformer is Secretly Linear 22.05.2024

The paper uncovers a linear characteristic in transformer decoders, showing a near-perfect linear relationship between embedding transformations in sequential layers. Regularization reduces linearity and improves performance. https://arxiv.org/abs//2405.12250 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/pod...

[QA] Training Data Attribution via Approximate Unrolled Differentation 21.05.2024

The paper introduces SOURCE, a computationally efficient training data attribution method that combines implicit differentiation and unrolling approaches, outperforming existing techniques in counterfactual prediction. https://arxiv.org/abs//2405.12186 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/ar...

Training Data Attribution via Approximate Unrolled Differentation 21.05.2024

The paper introduces SOURCE, a computationally efficient training data attribution method that combines implicit differentiation and unrolling approaches, outperforming existing techniques in counterfactual prediction. https://arxiv.org/abs//2405.12186 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/ar...

[QA] Information Leakage from Embedding in Large Language Models 21.05.2024

Study investigates privacy risks in large language models. Proposes methods to reconstruct user inputs from embeddings, introduces defense mechanism to safeguard privacy in distributed learning systems. https://arxiv.org/abs//2405.11916 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id169...

Information Leakage from Embedding in Large Language Models 21.05.2024

Study investigates privacy risks in large language models. Proposes methods to reconstruct user inputs from embeddings, introduces defense mechanism to safeguard privacy in distributed learning systems. https://arxiv.org/abs//2405.11916 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id169...

[QA] Layer-Condensed KV Cache for Efficient Inference of Large Language Models 20.05.2024

Proposed method reduces memory consumption in large language models by caching KVs of a small number of layers, improving throughput by up to 26% with competitive performance. https://arxiv.org/abs//2405.10637 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://po...

Layer-Condensed KV Cache for Efficient Inference of Large Language Models 20.05.2024

Proposed method reduces memory consumption in large language models by caching KVs of a small number of layers, improving throughput by up to 26% with competitive performance. https://arxiv.org/abs//2405.10637 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://po...

[QA] Observational Scaling Laws and the Predictability of Language Model Performance 20.05.2024

Proposing an observational approach to understand language model scaling laws using 80 public models, showing predictability of model performance and emergent phenomena without extensive training across scales. https://arxiv.org/abs//2405.10938 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-pape...

Observational Scaling Laws and the Predictability of Language Model Performance 20.05.2024

Proposing an observational approach to understand language model scaling laws using 80 public models, showing predictability of model performance and emergent phenomena without extensive training across scales. https://arxiv.org/abs//2405.10938 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-pape...

[QA] Zero-Shot Tokenizer Transfer 18.05.2024

This paper introduces Zero-Shot Tokenizer Transfer (ZeTT) to enable swapping tokenizers in language models, improving efficiency across languages and coding tasks without performance degradation. https://arxiv.org/abs//2405.07883 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016...

Zero-Shot Tokenizer Transfer 18.05.2024

This paper introduces Zero-Shot Tokenizer Transfer (ZeTT) to enable swapping tokenizers in language models, improving efficiency across languages and coding tasks without performance degradation. https://arxiv.org/abs//2405.07883 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016...

[QA] Many-Shot In-Context Learning in Multimodal Foundation Models 18.05.2024

Multimodal foundation models show improved performance in many-shot in-context learning, with Gemini 1.5 Pro demonstrating higher efficiency and scalability compared to GPT-4o. https://arxiv.org/abs//2405.09798 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://p...

Many-Shot In-Context Learning in Multimodal Foundation Models 18.05.2024

Multimodal foundation models show improved performance in many-shot in-context learning, with Gemini 1.5 Pro demonstrating higher efficiency and scalability compared to GPT-4o. https://arxiv.org/abs//2405.09798 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://p...

[QA] Chameleon: Mixed-Modal Early-Fusion Foundation Models 17.05.2024

https://arxiv.org/abs//2405.09818 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers

Chameleon: Mixed-Modal Early-Fusion Foundation Models 17.05.2024

https://arxiv.org/abs//2405.09818 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers

Listen to the Arxiv Papers podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.