Igor Melnyk
Arxiv Papers
Running out of time to catch up with new arXiv papers? We take the most impactful papers and present them as convenient podcasts. If you're a visual learner, we offer these papers in an engaging video format. Our service fills the gap between overly brief paper summaries and time-consuming full paper reads. You gain academic insights in a time-efficient, digestible format. Code behind this work: https://github.com/imelnyk/ArxivPapers
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Orca 2: Teaching Small Language Models How to Reason 21.11.2023 35:52
Orca 2, a small language model, is trained using various reasoning techniques and is able to outperform larger models on complex tasks. The model's performance is similar to models 5-10x larger, demonstrating the potential of smaller models in advanced reasoning. https://arxiv.org/abs//2311.11045 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple P...
[short] Exponentially Faster Language Modeling 21.11.2023 2:51
FastBERT is a BERT variant that uses only 0.3% of its neurons during inference, achieving similar performance to other BERT models. It selectively engages 12 out of 4095 neurons per layer using fast feedforward networks, resulting in significant speedup. https://arxiv.org/abs//2311.10770 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: http...
Exponentially Faster Language Modeling 21.11.2023 15:08
FastBERT is a BERT variant that uses only 0.3% of its neurons during inference, achieving similar performance to other BERT models. It selectively engages 12 out of 4095 neurons per layer using fast feedforward networks, resulting in significant speedup. https://arxiv.org/abs//2311.10770 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: http...
[short] Video-LLaVA: Learning United Visual Representation by Alignment Before Projection 20.11.2023 5:13
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection This paper proposes a unified vision-language model that combines images and videos into a single feature space, resulting in improved performance on various visual-language tasks compared to models that treat images and videos separately. https://arxiv.org/abs//2311.10122 YouTube: https://www.youtube.com/@ArxivPaper...
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection 20.11.2023 21:11
This paper proposes a unified vision-language model that combines images and videos into a single feature space, resulting in improved performance on various visual-language tasks compared to models that treat images and videos separately. https://arxiv.org/abs//2311.10122 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.ap...
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers 20.11.2023 3:46
The paper analyzes the effectiveness of using shallow feed-forward networks to mimic the attention mechanism in the Transformer model. Results show that these "attentionless Transformers" can rival the performance of the original architecture, highlighting the potential to streamline complex architectures for sequence-to-sequence tasks. https://arxiv.org/abs//2311.10642 YouTube: https://...
[short] The Chosen One: Consistent Characters in Text-to-Image Diffusion Models 19.11.2023 2:10
This paper presents a fully automated solution for generating consistent characters from text prompts. The proposed method outperforms baseline methods in terms of prompt alignment and identity consistency, as demonstrated through quantitative analysis and a user study. Several practical applications are showcased. https://arxiv.org/abs//2311.10093 YouTube: https://www.youtube.com/@ArxivPapers Tik...
The Chosen One: Consistent Characters in Text-to-Image Diffusion Models 19.11.2023 15:23
This paper presents a fully automated solution for generating consistent characters from text prompts. The proposed method outperforms baseline methods in terms of prompt alignment and identity consistency, as demonstrated through quantitative analysis and a user study. Several practical applications are showcased. https://arxiv.org/abs//2311.10093 YouTube: https://www.youtube.com/@ArxivPapers Tik...
[short] Striped Attention: Faster Ring Attention for Causal Transformers 17.11.2023 2:36
The paper proposes Striped Attention, an extension to the Ring Attention algorithm, to address workload imbalance in causal transformer models. It achieves significant throughput improvements on GPUs and TPUs at long sequence lengths. https://arxiv.org/abs//2311.09431 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.c...
Striped Attention: Faster Ring Attention for Causal Transformers 17.11.2023 20:37
The paper proposes Striped Attention, an extension to the Ring Attention algorithm, to address workload imbalance in causal transformer models. It achieves significant throughput improvements on GPUs and TPUs at long sequence lengths. https://arxiv.org/abs//2311.09431 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.c...
[short] SiRA: Sparse Mixture of Low Rank Adaptation 16.11.2023 1:44
SiRA is a sparse mixture of low rank adaption approach that leverages sparse computation to improve the performance of large language models on downstream tasks. It outperforms other approaches in single task and multitask settings. https://arxiv.org/abs//2311.09179 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com...
SiRA: Sparse Mixture of Low Rank Adaptation 16.11.2023 10:03
SiRA is a sparse mixture of low rank adaption approach that leverages sparse computation to improve the performance of large language models on downstream tasks. It outperforms other approaches in single task and multitask settings. https://arxiv.org/abs//2311.09179 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com...
[short] Fast Chain-of-Thought: A Glance of Future from Parallel Decoding Leads to Answers Faster 15.11.2023 2:07
FastCoT is a model-agnostic framework that uses parallel decoding and auto-regressive decoding simultaneously to improve inference time without sacrificing performance in language models. https://arxiv.org/abs//2311.08263 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify...
Fast Chain-of-Thought: A Glance of Future from Parallel Decoding Leads to Answers Faster 15.11.2023 15:51
FastCoT is a model-agnostic framework that uses parallel decoding and auto-regressive decoding simultaneously to improve inference time without sacrificing performance in language models. https://arxiv.org/abs//2311.08263 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify...
[short] Frontier Language Models are not Robust to Adversarial Arithmetic, or “What do I need to say so you agree 2+2=5?” 15.11.2023 2:32
The paper introduces the problem of adversarial arithmetic, where language models are tested on arithmetic questions with adversarial prompts. The authors propose an algorithm for finding successful attacks and show that models can be partially hardened against these attacks, but not fully robust. https://arxiv.org/abs//2311.07587 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.t...
Frontier Language Models are not Robust to Adversarial Arithmetic, or “What do I need to say so you agree 2+2=5?” 15.11.2023 18:51
The paper introduces the problem of adversarial arithmetic, where language models are tested on arithmetic questions with adversarial prompts. The authors propose an algorithm for finding successful attacks and show that models can be partially hardened against these attacks, but not fully robust. https://arxiv.org/abs//2311.07587 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.t...
[short] The ART of LLM Refinement: Ask, Refine, and Trust 15.11.2023 2:14
The paper explores the ability of large language models (LLMs) to judge the quality of their own generations. It introduces a reasoning with refinement strategy called ART, which improves performance on multistep reasoning tasks by asking necessary questions and ranking refinements. The approach achieves a 5-point gain over self-refinement baselines while using a smaller model as the decision make...
The ART of LLM Refinement: Ask, Refine, and Trust 15.11.2023 17:59
The paper explores the ability of large language models (LLMs) to judge the quality of their own generations. It introduces a reasoning with refinement strategy called ART, which improves performance on multistep reasoning tasks by asking necessary questions and ranking refinements. The approach achieves a 5-point gain over self-refinement baselines while using a smaller model as the decision make...
[short] LUMOS: Learning Agents with Unified Data, Modular Design, and Open-Source LLMs 13.11.2023 2:40
LUMOS is a framework for training language agents using a unified data format and modular architecture. It achieves comparable or superior performance to state-of-the-art agents and exhibits advantages in complex question answering, web tasks, math problems, and generalization to unseen tasks. https://arxiv.org/abs//2311.05657 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tikto...
LUMOS: Learning Agents with Unified Data, Modular Design, and Open-Source LLMs 13.11.2023 18:37
LUMOS is a framework for training language agents using a unified data format and modular architecture. It achieves comparable or superior performance to state-of-the-art agents and exhibits advantages in complex question answering, web tasks, math problems, and generalization to unseen tasks. https://arxiv.org/abs//2311.05657 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tikto...
[short] The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models 13.11.2023 2:38
This study investigates the anisotropy dynamics and intrinsic dimension of embeddings in transformer architectures, revealing distinct patterns in encoders and decoders. Initial training expands dimensionality, while later training refines into more compact representations. https://arxiv.org/abs//2311.05928 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers...
The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models 13.11.2023 8:22
This study investigates the anisotropy dynamics and intrinsic dimension of embeddings in transformer architectures, revealing distinct patterns in encoders and decoders. Initial training expands dimensionality, while later training refines into more compact representations. https://arxiv.org/abs//2311.05928 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers...
[short] Prompt Cache: Modular Attention Reuse for Low-Latency Inference 10.11.2023 2:26
Prompt Cache is a method for accelerating inference in large language models by reusing attention states of frequently occurring text segments. It significantly reduces latency without compromising output accuracy or requiring modifications to the model parameters. https://arxiv.org/abs//2311.04934 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Pod...
Prompt Cache: Modular Attention Reuse for Low-Latency Inference 10.11.2023 29:46
Prompt Cache is a method for accelerating inference in large language models by reusing attention states of frequently occurring text segments. It significantly reduces latency without compromising output accuracy or requiring modifications to the model parameters. https://arxiv.org/abs//2311.04934 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Pod...
Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO 10.11.2023 9:38
The paper explores the use of speculative sampling to reduce latency in text generation, comparing it to autoregressive sampling. The authors also discuss the use of model-based optimizations and provide a Jupyter notebook and sample executions. https://arxiv.org/abs//2311.04951 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podca...
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.