Igor Melnyk
Arxiv Papers
Running out of time to catch up with new arXiv papers? We take the most impactful papers and present them as convenient podcasts. If you're a visual learner, we offer these papers in an engaging video format. Our service fills the gap between overly brief paper summaries and time-consuming full paper reads. You gain academic insights in a time-efficient, digestible format. Code behind this work: https://github.com/imelnyk/ArxivPapers
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
[QA] Sapiens: Foundation for Human Vision Models 24.08.2024 7:49
Sapiens is a versatile model family for human-centric vision tasks, achieving state-of-the-art performance through self-supervised pretraining and scalable design, excelling in pose estimation, segmentation, depth, and normal prediction. https://arxiv.org/abs//2408.12569 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.appl...
Sapiens: Foundation for Human Vision Models 24.08.2024 22:52
Sapiens is a versatile model family for human-centric vision tasks, achieving state-of-the-art performance through self-supervised pretraining and scalable design, excelling in pose estimation, segmentation, depth, and normal prediction. https://arxiv.org/abs//2408.12569 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.appl...
[QA] Show-o: One Single Transformer to Unify Multimodal Understanding and Generation 24.08.2024 7:25
Show-o is a unified transformer model that integrates multimodal understanding and generation, outperforming existing models in various vision-language tasks while supporting diverse input-output modalities. https://arxiv.org/abs//2408.12528 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/...
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation 24.08.2024 28:14
Show-o is a unified transformer model that integrates multimodal understanding and generation, outperforming existing models in various vision-language tasks while supporting diverse input-output modalities. https://arxiv.org/abs//2408.12528 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/...
[QA] Jamba-1.5: Hybrid Transformer-Mamba Models at Scale 23.08.2024 7:22
Jamba-1.5 introduces instruction-tuned large language models with high throughput, low memory usage, and extensive context length, outperforming competitors while being publicly available under an open model license. https://arxiv.org/abs//2408.12570 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxi...
Jamba-1.5: Hybrid Transformer-Mamba Models at Scale 23.08.2024 16:53
Jamba-1.5 introduces instruction-tuned large language models with high throughput, low memory usage, and extensive context length, outperforming competitors while being publicly available under an open model license. https://arxiv.org/abs//2408.12570 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxi...
[QA] Hermes 3 Technical Report 23.08.2024 7:45
Hermes 3 is a neutrally-aligned instruct-tuned model with strong reasoning and creativity, achieving state-of-the-art performance on benchmarks, with weights available on Hugging Face. https://arxiv.org/abs//2408.11857 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: h...
Hermes 3 Technical Report 23.08.2024 11:21
Hermes 3 is a neutrally-aligned instruct-tuned model with strong reasoning and creativity, achieving state-of-the-art performance on benchmarks, with weights available on Hugging Face. https://arxiv.org/abs//2408.11857 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: h...
[QA] LLM Pruning and Distillation in Practice: The Minitron Approach 22.08.2024 7:23
https://arxiv.org/abs//2408.11796 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
LLM Pruning and Distillation in Practice: The Minitron Approach 22.08.2024 12:36
https://arxiv.org/abs//2408.11796 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
[QA] Approaching Deep Learning through the Spectral Dynamics of Weights 22.08.2024 7:26
This paper explores spectral dynamics of weights in deep learning, revealing optimization biases, enhancing weight decay effects, and distinguishing between memorizing and generalizing networks across various tasks. https://arxiv.org/abs//2408.11804 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv...
Approaching Deep Learning through the Spectral Dynamics of Weights 22.08.2024 26:45
This paper explores spectral dynamics of weights in deep learning, revealing optimization biases, enhancing weight decay effects, and distinguishing between memorizing and generalizing networks across various tasks. https://arxiv.org/abs//2408.11804 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv...
[QA] Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations 21.08.2024 8:04
The paper challenges the Linear Representation Hypothesis, showing that gated recurrent neural networks encode token sequences using magnitude rather than direction, suggesting broader interpretability in neural network research. https://arxiv.org/abs//2408.10920 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us...
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations 21.08.2024 21:03
The paper challenges the Linear Representation Hypothesis, showing that gated recurrent neural networks encode token sequences using magnitude rather than direction, suggesting broader interpretability in neural network research. https://arxiv.org/abs//2408.10920 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us...
[QA] Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model 21.08.2024 7:53
Transfusion is a multi-modal training method combining language modeling and diffusion, achieving superior performance in generating images and text with models up to 7B parameters. https://arxiv.org/abs//2408.11039 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: http...
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model 21.08.2024 24:23
Transfusion is a multi-modal training method combining language modeling and diffusion, achieving superior performance in generating images and text with models up to 7B parameters. https://arxiv.org/abs//2408.11039 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: http...
[QA] Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models 20.08.2024 8:37
The paper presents MOHAWK, a method for distilling Transformers into state space models, achieving strong performance with significantly less training data and computational resources. https://arxiv.org/abs//2408.10189 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: h...
Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models 20.08.2024 31:52
The paper presents MOHAWK, a method for distilling Transformers into state space models, achieving strong performance with significantly less training data and computational resources. https://arxiv.org/abs//2408.10189 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: h...
[QA] JPEG-LM: LLMs as Image Generators with Canonical Codec Representations 19.08.2024 7:47
This paper proposes using canonical codecs for image and video generation in autoregressive models, demonstrating improved efficiency and effectiveness over traditional pixel-based and vector quantization methods. https://arxiv.org/abs//2408.08459 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-p...
JPEG-LM: LLMs as Image Generators with Canonical Codec Representations 19.08.2024 20:04
This paper proposes using canonical codecs for image and video generation in autoregressive models, demonstrating improved efficiency and effectiveness over traditional pixel-based and vector quantization methods. https://arxiv.org/abs//2408.08459 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-p...
[QA] TextCAVs: Debugging vision models using text 19.08.2024 7:19
TextCAVs is a novel method for generating concept activation vectors using text descriptions, reducing the need for labeled image data in deep learning model interpretability, particularly in medical applications. https://arxiv.org/abs//2408.08652 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-p...
TextCAVs: Debugging vision models using text 19.08.2024 9:33
TextCAVs is a novel method for generating concept activation vectors using text descriptions, reducing the need for labeled image data in deep learning model interpretability, particularly in medical applications. https://arxiv.org/abs//2408.08652 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-p...
[QA] Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers 18.08.2024 9:08
The paper presents rStar, a self-play mutual reasoning method that enhances small language models' reasoning abilities without fine-tuning, achieving significant accuracy improvements across various reasoning tasks. https://arxiv.org/abs//2408.06195 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/a...
Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers 18.08.2024 16:46
The paper presents rStar, a self-play mutual reasoning method that enhances small language models' reasoning abilities without fine-tuning, achieving significant accuracy improvements across various reasoning tasks. https://arxiv.org/abs//2408.06195 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/a...
[QA] Towards flexible perception with visual memory 17.08.2024 7:18
https://arxiv.org/abs//2408.08172 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.