Igor Melnyk
Arxiv Papers
Running out of time to catch up with new arXiv papers? We take the most impactful papers and present them as convenient podcasts. If you're a visual learner, we offer these papers in an engaging video format. Our service fills the gap between overly brief paper summaries and time-consuming full paper reads. You gain academic insights in a time-efficient, digestible format. Code behind this work: https://github.com/imelnyk/ArxivPapers
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
[short] The Impact of Depth and Width on Transformer Language Model Generalization 01.11.2023 3:09
Deeper transformer language models tend to generalize better for compositional tasks, even when the total number of parameters is kept constant. The benefits of depth for generalization cannot be solely attributed to better performance on language modeling or in-distribution data. https://arxiv.org/abs//2310.19956 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_...
The Impact of Depth and Width on Transformer Language Model Generalization 01.11.2023 14:58
Deeper transformer language models tend to generalize better for compositional tasks, even when the total number of parameters is kept constant. The benefits of depth for generalization cannot be solely attributed to better performance on language modeling or in-distribution data. https://arxiv.org/abs//2310.19956 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_...
[short] Battle of the Backbones: A Large-Scale Comparison of Pretrained Models across Computer Vision Tasks 01.11.2023 2:27
The paper presents Battle of the Backbones (BoB), a benchmarking framework that evaluates the performance of various pretrained models on computer vision tasks. The results show that supervised pretrained convolutional neural networks still perform best, but self-supervised learning backbones are competitive and should be explored further. https://arxiv.org/abs//2310.19909 YouTube: https://www.you...
Battle of the Backbones: A Large-Scale Comparison of Pretrained Models across Computer Vision Tasks 01.11.2023 25:55
The paper presents Battle of the Backbones (BoB), a benchmarking framework that evaluates the performance of various pretrained models on computer vision tasks. The results show that supervised pretrained convolutional neural networks still perform best, but self-supervised learning backbones are competitive and should be explored further. https://arxiv.org/abs//2310.19909 YouTube: https://www.you...
[short] A Survey on Knowledge Editing of Neural Networks 31.10.2023 2:10
This paper surveys the emerging field of knowledge editing in deep neural networks, which aims to enable efficient and effective changes to pre-trained models without affecting previous tasks. It reviews relevant approaches and outlines potential future directions. https://arxiv.org/abs//2310.19704 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Pod...
A Survey on Knowledge Editing of Neural Networks 31.10.2023 37:10
This paper surveys the emerging field of knowledge editing in deep neural networks, which aims to enable efficient and effective changes to pre-trained models without affecting previous tasks. It reviews relevant approaches and outlines potential future directions. https://arxiv.org/abs//2310.19704 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Pod...
[short] MM-VID : Advancing Video Understanding with GPT-4V(ision) 31.10.2023 2:27
MM-VID is an integrated system that combines GPT-4V with specialized tools to facilitate advanced video understanding, addressing challenges in long-form videos and complex tasks. It uses video-to-script generation to transcribe multimodal elements and enables advanced capabilities in video comprehension. Experimental results demonstrate its effectiveness across different video genres and lengths....
MM-VID : Advancing Video Understanding with GPT-4V(ision) 31.10.2023 16:40
MM-VID is an integrated system that combines GPT-4V with specialized tools to facilitate advanced video understanding, addressing challenges in long-form videos and complex tasks. It uses video-to-script generation to transcribe multimodal elements and enables advanced capabilities in video comprehension. Experimental results demonstrate its effectiveness across different video genres and lengths....
CodeFusion: A Pre-trained Diffusion Model for Code Generation 30.10.2023 11:42
The paper introduces \system, a pre-trained diffusion code generation model that addresses the limitation of auto-regressive models by iteratively denoising a complete program. It performs on par with state-of-the-art models in accuracy and outperforms them in diversity versus quality. https://arxiv.org/abs//2310.17680 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@a...
[short] Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-Image Generation 30.10.2023 2:01
The paper introduces Davidsonian Scene Graph (DSG), an evaluation framework for text-to-image models. DSG addresses reliability challenges in question generation and visual question answering, and includes an open-sourced evaluation benchmark called DSG-1k. https://arxiv.org/abs//2310.18235 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: h...
Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-Image Generation 30.10.2023 19:42
The paper introduces Davidsonian Scene Graph (DSG), an evaluation framework for text-to-image models. DSG addresses reliability challenges in question generation and visual question answering, and includes an open-sourced evaluation benchmark called DSG-1k. https://arxiv.org/abs//2310.18235 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: h...
[small] Zephyr: Direct Distillation of LM Alignment 29.10.2023 2:06
The paper presents ZEPHYR-7B, a language model that achieves state-of-the-art performance on chat benchmarks by using distilled direct preference optimization (dDPO) and AI Feedback (AIF) data. https://arxiv.org/abs//2310.16944 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 S...
Zephyr: Direct Distillation of LM Alignment 29.10.2023 12:09
The paper presents ZEPHYR-7B, a language model that achieves state-of-the-art performance on chat benchmarks by using distilled direct preference optimization (dDPO) and AI Feedback (AIF) data. https://arxiv.org/abs//2310.16944 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 S...
[short] JudgeLM : Fine-tuned Large Language Models are Scalable Judges 28.10.2023 2:25
The paper proposes a method called JudgeLM to evaluate large language models (LLMs) in open-ended scenarios. They fine-tune LLMs as scalable judges and introduce techniques to address biases. JudgeLM achieves state-of-the-art performance and high agreement with human judges. https://arxiv.org/abs//2310.17631 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers...
JudgeLM : Fine-tuned Large Language Models are Scalable Judges 28.10.2023 17:05
The paper proposes a method called JudgeLM to evaluate large language models (LLMs) in open-ended scenarios. They fine-tune LLMs as scalable judges and introduce techniques to address biases. JudgeLM achieves state-of-the-art performance and high agreement with human judges. https://arxiv.org/abs//2310.17631 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers...
How do Language Models Bind Entities in Context? 27.10.2023 18:23
The paper analyzes language models and identifies a mechanism for binding entities to their attributes. It shows that language models represent binding information through internal activations and that binding vectors form a continuous subspace. This provides insights into how language models represent symbolic knowledge in-context. https://arxiv.org/abs//2310.17191 YouTube: https://www.youtube.co...
[short] Controlled Decoding from Language Models 27.10.2023 1:57
Controlled decoding (CD) is a novel off-policy reinforcement learning method that uses a value function called a prefix scorer to steer autoregressive generation towards high reward outcomes. CD is effective in controlling language models and can handle multiple rewards without additional complexity. It can also be applied in a blockwise fashion at inference-time, making it a promising approach fo...
Controlled Decoding from Language Models 27.10.2023 14:52
Controlled decoding (CD) is a novel off-policy reinforcement learning method that uses a value function called a prefix scorer to steer autoregressive generation towards high reward outcomes. CD is effective in controlling language models and can handle multiple rewards without additional complexity. It can also be applied in a blockwise fashion at inference-time, making it a promising approach fo...
[short] Fantastic Gains and Where to Find Them: On the Existence and Prospect of General Knowledge Transfer between Any Pretrained Model 27.10.2023 2:30
This paper explores the transfer of "complementary" knowledge between pretrained deep learning models without performance degradation. The authors propose a data partitioning technique for successful transfer and assess the scalability and impact of model properties on knowledge transfer. https://arxiv.org/abs//2310.17653 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www....
Fantastic Gains and Where to Find Them: On the Existence and Prospect of General Knowledge Transfer between Any Pretrained Model 27.10.2023 20:56
This paper explores the transfer of "complementary" knowledge between pretrained deep learning models without performance degradation. The authors propose a data partitioning technique for successful transfer and assess the scalability and impact of model properties on knowledge transfer. https://arxiv.org/abs//2310.17653 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www....
[short] A Picture is Worth a Thousand Words: Principled Recaptioning Improves Image Generation 26.10.2023 2:25
https://arxiv.org/abs//2310.16656 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
A Picture is Worth a Thousand Words: Principled Recaptioning Improves Image Generation 26.10.2023 18:32
https://arxiv.org/abs//2310.16656 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers
[short] Detecting Pretraining Data from Large Language Models 26.10.2023 1:45
This paper introduces a method, MIN-K% PROB, to detect if a large language model was trained on a given text without knowing the pretraining data. It achieves better results than previous methods and is effective in various real-world scenarios. https://arxiv.org/abs//2310.16789 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podca...
Detecting Pretraining Data from Large Language Models 26.10.2023 28:08
This paper introduces a method, MIN-K% PROB, to detect if a large language model was trained on a given text without knowing the pretraining data. It achieves better results than previous methods and is effective in various real-world scenarios. https://arxiv.org/abs//2310.16789 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podca...
[short] QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter Models 26.10.2023 1:56
The paper introduces QMoE, a compression and execution framework that allows trillion-parameter language models to be run efficiently on affordable hardware with minimal accuracy loss. https://arxiv.org/abs//2310.16795 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: h...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.