Igor Melnyk

Arxiv Papers

Science EN ↓ 2489 episodes

Running out of time to catch up with new arXiv papers? We take the most impactful papers and present them as convenient podcasts. If you're a visual learner, we offer these papers in an engaging video format. Our service fills the gap between overly brief paper summaries and time-consuming full paper reads. You gain academic insights in a time-efficient, digestible format. Code behind this work: https://github.com/imelnyk/ArxivPapers

Author

Igor Melnyk

Category

Science

Podcast website

github.com

Latest episode

Sep 1, 2025

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android almost 10M downloads · 4.8 rating iOS soon

Episodes

Instruction-tuning Aligns LLMs to the Human Brain 04.12.2023

Instruction-tuning improves brain alignment in large language models (LLMs) but does not have the same effect on behavioral alignment. Model size and performance on tasks requiring world knowledge are positively correlated with brain alignment. https://arxiv.org/abs//2312.00575 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcas...

[short] Dataset Distillation in Large Data Era 03.12.2023

This paper introduces a method for distilling large-scale datasets, such as ImageNet-1K/21K, to achieve high accuracy. The proposed method includes curriculum data augmentation and achieves state-of-the-art results. https://arxiv.org/abs//2311.18838 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv...

Dataset Distillation in Large Data Era 03.12.2023

This paper introduces a method for distilling large-scale datasets, such as ImageNet-1K/21K, to achieve high accuracy. The proposed method includes curriculum data augmentation and achieves state-of-the-art results. https://arxiv.org/abs//2311.18838 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv...

[short] One-step Diffusion with Distribution Matching Distillation 01.12.2023

The paper introduces Distribution Matching Distillation (DMD), a method to transform a diffusion model into a one-step image generator with minimal impact on image quality. DMD outperforms other few-step diffusion approaches in terms of speed and performance. https://arxiv.org/abs//2311.18828 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts:...

One-step Diffusion with Distribution Matching Distillation 01.12.2023

The paper introduces Distribution Matching Distillation (DMD), a method to transform a diffusion model into a one-step image generator with minimal impact on image quality. DMD outperforms other few-step diffusion approaches in terms of speed and performance. https://arxiv.org/abs//2311.18828 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts:...

[short] Initializing Models with Larger Ones 01.12.2023

Weight selection is a method for initializing smaller neural network models by selecting a subset of weights from a pretrained larger model, improving performance and reducing training time. It can be used with knowledge distillation and is useful for training small models in resource-constrained settings. https://arxiv.org/abs//2311.18823 YouTube: https://www.youtube.com/@ArxivPapers TikTok: http...

Initializing Models with Larger Ones 01.12.2023

Weight selection is a method for initializing smaller neural network models by selecting a subset of weights from a pretrained larger model, improving performance and reducing training time. It can be used with knowledge distillation and is useful for training small models in resource-constrained settings. https://arxiv.org/abs//2311.18823 YouTube: https://www.youtube.com/@ArxivPapers TikTok: http...

[short] SODA: Bottleneck Diffusion Models for Representation Learning 30.11.2023

SODA is a self-supervised diffusion model that uses an image encoder to generate novel views. It achieves strong representation learning and disentangled latent space, making it promising for image generation and robust representations. https://arxiv.org/abs//2311.17901 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple...

SODA: Bottleneck Diffusion Models for Representation Learning 30.11.2023

SODA is a self-supervised diffusion model that uses an image encoder to generate novel views. It achieves strong representation learning and disentangled latent space, making it promising for image generation and robust representations. https://arxiv.org/abs//2311.17901 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple...

[short] No Representation Rules Them All in Category Discovery 29.11.2023

This paper addresses the problem of Generalized Category Discovery (GCD) and proposes a synthetic dataset called Clevr-4 to evaluate GCD algorithms. The authors propose a new method called GCD, based on mean teachers, which outperforms existing baselines and sets a new state-of-the-art on real data. https://arxiv.org/abs//2311.17055 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www...

No Representation Rules Them All in Category Discovery 29.11.2023

This paper addresses the problem of Generalized Category Discovery (GCD) and proposes a synthetic dataset called Clevr-4 to evaluate GCD algorithms. The authors propose a new method called GCD, based on mean teachers, which outperforms existing baselines and sets a new state-of-the-art on real data. https://arxiv.org/abs//2311.17055 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www...

[short] Manifold Preserving Guided Diffusion 29.11.2023

The paper proposes a training-free conditional image generation framework called MPGD that leverages pretrained diffusion models and off-the-shelf neural networks for various tasks, offering speed-ups and maintaining high sample quality. https://arxiv.org/abs//2311.16424 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.appl...

Manifold Preserving Guided Diffusion 29.11.2023

The paper proposes a training-free conditional image generation framework called MPGD that leverages pretrained diffusion models and off-the-shelf neural networks for various tasks, offering speed-ups and maintaining high sample quality. https://arxiv.org/abs//2311.16424 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.appl...

[short] On the Long Range Abilities of Transformers 29.11.2023

This paper proposes modifications to the transformer architecture that improve its performance on long-range tasks, narrowing the gap with specialized layers. The modifications incorporate an inductive bias towards smoothness and locality, improving results without additional computation or parameters. https://arxiv.org/abs//2311.16620 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://...

On the Long Range Abilities of Transformers 29.11.2023

This paper proposes modifications to the transformer architecture that improve its performance on long-range tasks, narrowing the gap with specialized layers. The modifications incorporate an inductive bias towards smoothness and locality, improving results without additional computation or parameters. https://arxiv.org/abs//2311.16620 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://...

[short] Scalable Extraction of Training Data from (Production) Language Models 29.11.2023

The paper explores the concept of extractable memorization in machine learning models and demonstrates that an adversary can extract large amounts of training data from various types of models. Existing techniques can attack unaligned models, and a new divergence attack is developed for aligned models. The findings suggest that current alignment techniques do not eliminate memorization. https://ar...

Scalable Extraction of Training Data from (Production) Language Models 29.11.2023

The paper explores the concept of extractable memorization in machine learning models and demonstrates that an adversary can extract large amounts of training data from various types of models. Existing techniques can attack unaligned models, and a new divergence attack is developed for aligned models. The findings suggest that current alignment techniques do not eliminate memorization. https://ar...

[short] Scalable AI Safety via Doubly-Efficient Debate 27.11.2023

The paper proposes a new debate method for AI safety, addressing challenges in the original framework. The new protocols allow the honest strategy to succeed using a polynomial number of steps and verify alignment of stochastic AI systems. https://arxiv.org/abs//2311.14125 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.ap...

Scalable AI Safety via Doubly-Efficient Debate 27.11.2023

The paper proposes a new debate method for AI safety, addressing challenges in the original framework. The new protocols allow the honest strategy to succeed using a polynomial number of steps and verify alignment of stochastic AI systems. https://arxiv.org/abs//2311.14125 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.ap...

[short] Let's Verify Step by Step 24.11.2023

Process supervision is shown to significantly outperform outcome supervision in training language models to solve complex reasoning problems, as demonstrated on the MATH dataset. Active learning is also shown to improve the efficacy of process supervision. https://arxiv.org/abs//2305.20050 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: ht...

Let's Verify Step by Step 24.11.2023

Process supervision is shown to significantly outperform outcome supervision in training language models to solve complex reasoning problems, as demonstrated on the MATH dataset. Active learning is also shown to improve the efficacy of process supervision. https://arxiv.org/abs//2305.20050 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: ht...

PaSS: Parallel Speculative Sampling 24.11.2023

The paper discusses the challenges of generating tokens in large language models and proposes a method called parallel decoding that allows for drafting multiple tokens simultaneously using a single model, resulting in improved performance. https://arxiv.org/abs//2311.13581 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.a...

[short] In-Context Learning Functions with Varying Number of Minima 22.11.2023

This paper explores the interplay between In-Context Learning (ICL) and the properties of functions it approximates. The study finds that increasing the number of minima in the functions degrades ICL performance, but ICL still outperforms a 2-layer Neural Network model and learns faster. https://arxiv.org/abs//2311.12538 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/...

In-Context Learning Functions with Varying Number of Minima 22.11.2023

This paper explores the interplay between In-Context Learning (ICL) and the properties of functions it approximates. The study finds that increasing the number of minima in the functions degrades ICL performance, but ICL still outperforms a 2-layer Neural Network model and learns faster. https://arxiv.org/abs//2311.12538 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/...

[short] Orca 2: Teaching Small Language Models How to Reason 21.11.2023

Orca 2, a small language model, is trained using various reasoning techniques and is able to outperform larger models on complex tasks. The model's performance is similar to models 5-10x larger, demonstrating the potential of smaller models in advanced reasoning. https://arxiv.org/abs//2311.11045 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple P...

Listen to the Arxiv Papers podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.