Igor Melnyk

Arxiv Papers

Science EN ↓ 2489 episodes

Running out of time to catch up with new arXiv papers? We take the most impactful papers and present them as convenient podcasts. If you're a visual learner, we offer these papers in an engaging video format. Our service fills the gap between overly brief paper summaries and time-consuming full paper reads. You gain academic insights in a time-efficient, digestible format. Code behind this work: https://github.com/imelnyk/ArxivPapers

Author

Igor Melnyk

Category

Science

Podcast website

github.com

Latest episode

Sep 1, 2025

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

Tweets to Citations: Unveiling the Impact of Social Media Influencers on AI Research Visibility 26.01.2024

This paper investigates the impact of social media influencers on the visibility and citation counts of machine learning research. The study finds that papers endorsed by influencers have significantly higher citation counts, emphasizing the growing influence of social media in scholarly communication. https://arxiv.org/abs//2401.13782 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://...

[short] Rethinking Patch Dependence for Masked Autoencoders 26.01.2024

The paper proposes a novel pretraining framework called Cross-Attention Masked Autoencoders (CrossMAE) that leverages cross-attention between masked and visible tokens, resulting in improved representation learning with reduced decoding compute. CrossMAE outperforms MAE on ImageNet classification and COCO instance segmentation. https://arxiv.org/abs//2401.14391 YouTube: https://www.youtube.com/@Ar...

Rethinking Patch Dependence for Masked Autoencoders 26.01.2024

The paper proposes a novel pretraining framework called Cross-Attention Masked Autoencoders (CrossMAE) that leverages cross-attention between masked and visible tokens, resulting in improved representation learning with reduced decoding compute. CrossMAE outperforms MAE on ImageNet classification and COCO instance segmentation. https://arxiv.org/abs//2401.14391 YouTube: https://www.youtube.com/@Ar...

[short] Deconstructing Denoising Diffusion Models for Self-Supervised Learning 26.01.2024

This study explores the representation learning abilities of Denoising Diffusion Models (DDM) and finds that only a few modern components are critical for learning good representations, leading to a simplified approach resembling a classical Denoising Autoencoder (DAE). https://arxiv.org/abs//2401.14404 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Appl...

Deconstructing Denoising Diffusion Models for Self-Supervised Learning 26.01.2024

This study explores the representation learning abilities of Denoising Diffusion Models (DDM) and finds that only a few modern components are critical for learning good representations, leading to a simplified approach resembling a classical Denoising Autoencoder (DAE). https://arxiv.org/abs//2401.14404 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Appl...

[short] MambaByte: Token-free Selective State Space Model 25.01.2024

MambaByte, a token-free language model trained on byte sequences, demonstrates computational efficiency and competitive performance compared to subword Transformers, making it a viable option for token-free language modeling. https://arxiv.org/abs//2401.13660 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/pod...

MambaByte: Token-free Selective State Space Model 25.01.2024

MambaByte, a token-free language model trained on byte sequences, demonstrates computational efficiency and competitive performance compared to subword Transformers, making it a viable option for token-free language modeling. https://arxiv.org/abs//2401.13660 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/pod...

[short] Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment 24.01.2024

The paper introduces DITTO, a method for improving the role-playing capabilities of large language models (LLMs) by leveraging their extensive knowledge of characters and dialogues. DITTO outperforms open-source baselines and achieves performance comparable to advanced proprietary chatbots. https://arxiv.org/abs//2401.12474 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.c...

Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment 24.01.2024

The paper introduces DITTO, a method for improving the role-playing capabilities of large language models (LLMs) by leveraging their extensive knowledge of characters and dialogues. DITTO outperforms open-source baselines and achieves performance comparable to advanced proprietary chatbots. https://arxiv.org/abs//2401.12474 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.c...

[short] Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding 24.01.2024

Meta-prompting is a technique that enhances the functionality of language models by transforming them into multi-faceted conductors. It guides the models to break down complex tasks into smaller subtasks, which are handled by expert instances of the same model. This approach improves performance across various tasks and simplifies user interaction. The integration of external tools, such as a Pyth...

Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding 24.01.2024

Meta-prompting is a technique that enhances the functionality of language models by transforming them into multi-faceted conductors. It guides the models to break down complex tasks into smaller subtasks, which are handled by expert instances of the same model. This approach improves performance across various tasks and simplifies user interaction. The integration of external tools, such as a Pyth...

[short] WARM: On the Benefits of Weight Averaged Reward Models 23.01.2024

The paper discusses the use of machine learning algorithms to predict the risk of cardiovascular disease in patients with diabetes. https://arxiv.org/abs//2401.12187 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers

WARM: On the Benefits of Weight Averaged Reward Models 23.01.2024

The paper discusses the use of machine learning algorithms to predict the risk of cardiovascular disease in patients with diabetes. https://arxiv.org/abs//2401.12187 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers

[short] Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text 23.01.2024

The paper presents a novel method called Binoculars for detecting machine-generated text from large language models (LLMs) without the need for training data. Binoculars achieves high accuracy in detecting machine-generated text across various sources and situations. https://arxiv.org/abs//2401.12070 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple P...

Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text 23.01.2024

The paper presents a novel method called Binoculars for detecting machine-generated text from large language models (LLMs) without the need for training data. Binoculars achieves high accuracy in detecting machine-generated text across various sources and situations. https://arxiv.org/abs//2401.12070 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple P...

Zero Bubble Pipeline Parallelism 22.01.2024

This paper introduces a scheduling strategy for pipeline parallelism in distributed training that achieves zero pipeline bubbles, resulting in improved performance compared to baseline methods. The authors also develop an algorithm to automatically find optimal schedules and introduce a technique to bypass synchronizations during the optimizer step. Experimental evaluations show significant improv...

[short] MEDUSA: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads 22.01.2024

This paper introduces MEDUSA, a method for improving inference in Large Language Models (LLMs) by adding extra decoding heads to predict multiple tokens in parallel. MEDUSA achieves significant speedup without compromising generation quality. https://arxiv.org/abs//2401.10774 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts...

MEDUSA: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads 22.01.2024

This paper introduces MEDUSA, a method for improving inference in Large Language Models (LLMs) by adding extra decoding heads to predict multiple tokens in parallel. MEDUSA achieves significant speedup without compromising generation quality. https://arxiv.org/abs//2401.10774 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts...

[short] Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning 22.01.2024

Pruning up to 20% of parameters in large language models (LLMs) improves their resistance to "Jailbreaking" prompts, reducing the generation of harmful and illegal content without sacrificing performance. Pruning may also enhance other LLM behaviors and improve safety and reliability. https://arxiv.org/abs//2401.10862 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tikt...

Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning 22.01.2024

Pruning up to 20% of parameters in large language models (LLMs) improves their resistance to "Jailbreaking" prompts, reducing the generation of harmful and illegal content without sacrificing performance. Pruning may also enhance other LLM behaviors and improve safety and reliability. https://arxiv.org/abs//2401.10862 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tikt...

Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability 21.01.2024

The paper introduces a method called Deductive Closure Training (DCT) that uses language models to improve their factuality and coherence. DCT prompts LMs to generate text, reason about its correctness, and fine-tune based on inferred correctness. DCT improves LM fact verification and text generation accuracy. https://arxiv.org/abs//2401.08574 YouTube: https://www.youtube.com/@ArxivPapers TikTok:...

VMamba: Visual State Space Model 21.01.2024

The paper introduces VMamba, a novel architecture that combines the strengths of CNNs and ViTs for visual representation learning. VMamba achieves linear complexity while maintaining global receptive fields and outperforms benchmarks as image resolution increases. https://arxiv.org/abs//2401.10166 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podc...

[short] Training Neural Networks is NP-Hard in Fixed Dimension 20.01.2024

The paper investigates the complexity of training two-layer neural networks with different activation functions. It provides answers to several open questions and establishes the computational hardness of these problems for certain dimensions and activation functions. https://arxiv.org/abs//2303.17045 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple...

Training Neural Networks is NP-Hard in Fixed Dimension 20.01.2024

The paper investigates the complexity of training two-layer neural networks with different activation functions. It provides answers to several open questions and establishes the computational hardness of these problems for certain dimensions and activation functions. https://arxiv.org/abs//2303.17045 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple...

[short] Rethinking FID: Towards a Better Evaluation Metric for Image Generation 20.01.2024

The paper criticizes the Fréchet Inception Distance (FID) as an evaluation metric for generated images and proposes an alternative metric called CMMD, which is based on CLIP embeddings and offers a more reliable assessment of image quality. https://arxiv.org/abs//2401.09603 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.a...

Listen to the Arxiv Papers podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.