Igor Melnyk
Arxiv Papers
Running out of time to catch up with new arXiv papers? We take the most impactful papers and present them as convenient podcasts. If you're a visual learner, we offer these papers in an engaging video format. Our service fills the gap between overly brief paper summaries and time-consuming full paper reads. You gain academic insights in a time-efficient, digestible format. Code behind this work: https://github.com/imelnyk/ArxivPapers
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ 03.10.2023 21:01
The paper proposes using an abstract graphics language as an intermediate representation for generating vector graphics from text. They introduce a large-scale dataset and a new model that outperforms existing models in terms of similarity to human-created figures. The framework and datasets are publicly available. https://arxiv.org/abs//2310.00367 YouTube: https://www.youtube.com/@ArxivPapers Tik...
Representation Engineering: A Top-Down Approach to AI Transparency 03.10.2023 10:21
This paper introduces representation engineering (RepE), an approach that enhances transparency in AI systems using insights from cognitive neuroscience. RepE focuses on population-level representations and offers effective solutions for understanding and controlling large language models, addressing safety concerns. https://arxiv.org/abs//2310.01405 YouTube: https://www.youtube.com/@ArxivPapers T...
The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) 02.10.2023 18:17
The paper analyzes GPT-4V, a large multimodal model, and explores its capabilities, inputs, working modes, and prompts. It demonstrates GPT-4V's ability to process multimodal inputs and discusses potential applications and future research directions. https://arxiv.org/abs//2309.17421 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: http...
[short] Efficient Streaming Language Models with Attention Sinks 02.10.2023 3:15
This paper introduces StreamingLLM, an efficient framework that allows large language models to generalize to infinite sequence length in streaming applications without fine-tuning. It addresses challenges related to memory consumption and text length, and achieves stable and efficient language modeling. https://arxiv.org/abs//2309.17453 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https:...
Efficient Streaming Language Models with Attention Sinks 02.10.2023 23:32
This paper introduces StreamingLLM, an efficient framework that allows large language models to generalize to infinite sequence length in streaming applications without fine-tuning. It addresses challenges related to memory consumption and text length, and achieves stable and efficient language modeling. https://arxiv.org/abs//2309.17453 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https:...
[short] Tree Cross Attention 02.10.2023 2:48
This paper introduces Tree Cross Attention (TCA), a module that retrieves information from a logarithmic number of tokens for efficient inference. The proposed architecture, ReTreever, outperforms Perceiver IO while using the same number of tokens. https://arxiv.org/abs//2309.17388 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://po...
Tree Cross Attention 02.10.2023 18:27
This paper introduces Tree Cross Attention (TCA), a module that retrieves information from a logarithmic number of tokens for efficient inference. The proposed architecture, ReTreever, outperforms Perceiver IO while using the same number of tokens. https://arxiv.org/abs//2309.17388 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://po...
[short] Demystifying CLIP Data 30.09.2023 3:29
The paper introduces MetaCLIP, a data curation approach that outperforms CLIP's data on standard benchmarks. MetaCLIP achieves higher accuracy in zero-shot ImageNet classification and scales well with larger datasets. https://arxiv.org/abs//2309.16671 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast...
Demystifying CLIP Data 30.09.2023 22:49
The paper introduces MetaCLIP, a data curation approach that outperforms CLIP's data on standard benchmarks. MetaCLIP achieves higher accuracy in zero-shot ImageNet classification and scales well with larger datasets. https://arxiv.org/abs//2309.16671 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast...
[short] Effective Long-Context Scaling of Foundation Models 30.09.2023 2:44
The paper presents a series of long-context language models (LLMs) that achieve effective context windows of up to 32,768 tokens. The models are built through continual pretraining and achieve consistent improvements on various language modeling tasks and research benchmarks. The paper also analyzes the components and design choices of the models. https://arxiv.org/abs//2309.16039 YouTube: https:/...
Effective Long-Context Scaling of Foundation Models 30.09.2023 21:06
The paper presents a series of long-context language models (LLMs) that achieve effective context windows of up to 32,768 tokens. The models are built through continual pretraining and achieve consistent improvements on various language modeling tasks and research benchmarks. The paper also analyzes the components and design choices of the models. https://arxiv.org/abs//2309.16039 YouTube: https:/...
[short] GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond 29.09.2023 2:55
This paper introduces GPT-Fathom, an open-source evaluation suite for large language models (LLMs). It evaluates 10+ LLMs on 20+ benchmarks, providing insights into the evolution from GPT-3 to GPT-4 and improving transparency. https://arxiv.org/abs//2309.16583 YouTube: https://www.youtube.com/@ArxivPapers PODCASTS: Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spo...
GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond 29.09.2023 13:25
This paper introduces GPT-Fathom, an open-source evaluation suite for large language models (LLMs). It evaluates 10+ LLMs on 20+ benchmarks, providing insights into the evolution from GPT-3 to GPT-4 and improving transparency. https://arxiv.org/abs//2309.16583 YouTube: https://www.youtube.com/@ArxivPapers PODCASTS: Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spo...
[short] MotionLM: Multi-Agent Motion Forecasting as Language Modeling 29.09.2023 3:18
The paper introduces MotionLM, a model for multi-agent motion prediction in autonomous vehicles. It uses language modeling to represent trajectories and achieves state-of-the-art performance on the Waymo Open Motion Dataset. https://arxiv.org/abs//2309.16534 YouTube: https://www.youtube.com/@ArxivPapers PODCASTS: Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spoti...
MotionLM: Multi-Agent Motion Forecasting as Language Modeling 29.09.2023 18:31
The paper introduces MotionLM, a model for multi-agent motion prediction in autonomous vehicles. It uses language modeling to represent trajectories and achieves state-of-the-art performance on the Waymo Open Motion Dataset. https://arxiv.org/abs//2309.16534 YouTube: https://www.youtube.com/@ArxivPapers PODCASTS: Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spoti...
[short] Jointly Training Large Autoregressive Multimodal Models 28.09.2023 3:01
The paper introduces the Joint Autoregressive Mixture (JAM) framework, which combines text and image generation models to create high-quality multimodal outputs. It also presents a data-efficient instruction-tuning strategy for mixed-modal generation tasks. https://arxiv.org/abs//2309.15564 YouTube: https://www.youtube.com/@ArxivPapers PODCASTS: Apple Podcasts: https://podcasts.apple.com/us/podcas...
Jointly Training Large Autoregressive Multimodal Models 28.09.2023 25:12
The paper introduces the Joint Autoregressive Mixture (JAM) framework, which combines text and image generation models to create high-quality multimodal outputs. It also presents a data-efficient instruction-tuning strategy for mixed-modal generation tasks. https://arxiv.org/abs//2309.15564 YouTube: https://www.youtube.com/@ArxivPapers PODCASTS: Apple Podcasts: https://podcasts.apple.com/us/podcas...
Deep Model Fusion: A Survey 28.09.2023 17:50
This paper presents a comprehensive survey on deep model fusion, a technique that merges the parameters or predictions of multiple deep learning models. It categorizes existing methods and discusses challenges and future research directions. https://arxiv.org/abs//2309.15698 YouTube: https://www.youtube.com/@ArxivPapers PODCASTS: Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/i...
DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models 27.09.2023 13:53
DeepSpeed-Ulysses is a methodology for efficient and scalable training of large language models with long sequences. It uses sequence parallelism and efficient communication to achieve faster training with longer sequences compared to existing methods. https://arxiv.org/abs//2309.14509 YouTube: https://www.youtube.com/@ArxivPapers PODCASTS: Apple Podcasts: https://podcasts.apple.com/us/podcast/arx...
Large Language Model Alignment: A Survey 27.09.2023 23:38
This survey explores alignment techniques for large language models (LLMs) to ensure their behavior aligns with human values. It categorizes methods, discusses interpretability and vulnerabilities, presents benchmarks, and outlines future research directions. https://arxiv.org/abs//2309.15025 YouTube: https://www.youtube.com/@ArxivPapers PODCASTS: Apple Podcasts: https://podcasts.apple.com/us/podc...
[short] Aligning Large Multimodal Models with Factually Augmented RLHF 27.09.2023 3:15
The paper introduces a new alignment algorithm called Factually Augmented RLHF to address multimodal misalignment in large multimodal models. The proposed approach achieves significant improvement in performance and is available as open-source code and data. https://arxiv.org/abs//2309.14525 YouTube: https://www.youtube.com/@ArxivPapers PODCASTS: Apple Podcasts: https://podcasts.apple.com/us/podca...
Aligning Large Multimodal Models with Factually Augmented RLHF 27.09.2023 23:56
The paper introduces a new alignment algorithm called Factually Augmented RLHF to address multimodal misalignment in large multimodal models. The proposed approach achieves significant improvement in performance and is available as open-source code and data. https://arxiv.org/abs//2309.14525 YouTube: https://www.youtube.com/@ArxivPapers PODCASTS: Apple Podcasts: https://podcasts.apple.com/us/podca...
[short] SCREWS : A Modular Framework for Reasoning with Revisions 26.09.2023 3:33
Large language models (LLMs) can improve their accuracy on tasks by refining and revising their output based on feedback. The SCREWS framework enables exploration in this space by providing modules for sampling, conditional resampling, and selection, leading to improved reasoning strategies for various tasks. https://arxiv.org/abs//2309.13075 YouTube: https://www.youtube.com/@ArxivPapers PODCASTS:...
SCREWS : A Modular Framework for Reasoning with Revisions 26.09.2023 25:13
Large language models (LLMs) can improve their accuracy on tasks by refining and revising their output based on feedback. The SCREWS framework enables exploration in this space by providing modules for sampling, conditional resampling, and selection, leading to improved reasoning strategies for various tasks. https://arxiv.org/abs//2309.13075 YouTube: https://www.youtube.com/@ArxivPapers PODCASTS:...
[small] Small-scale proxies for large-scale Transformer training instabilities 26.09.2023 3:03
The paper investigates training instabilities in large Transformer-based models and explores ways to reproduce and study these instabilities at smaller scales. It examines sources of instability, explores the impact of learning rate and other interventions, and studies cases where instabilities can be predicted. https://arxiv.org/abs//2309.14322 YouTube: https://www.youtube.com/@ArxivPapers PODCAS...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.