Igor Melnyk

Arxiv Papers

Science EN ↓ 2489 episodes

Running out of time to catch up with new arXiv papers? We take the most impactful papers and present them as convenient podcasts. If you're a visual learner, we offer these papers in an engaging video format. Our service fills the gap between overly brief paper summaries and time-consuming full paper reads. You gain academic insights in a time-efficient, digestible format. Code behind this work: https://github.com/imelnyk/ArxivPapers

Author

Igor Melnyk

Category

Science

Podcast website

github.com

Latest episode

Sep 1, 2025

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android almost 10M downloads · 4.8 rating iOS soon

Episodes

Self-Evaluation Improves Selective Generation in Large Language Models 18.12.2023

This paper explores the use of token-level self-evaluation to improve the accuracy and quality of generated content by large language models. Experimental results show that self-evaluation based scores are effective in selective generation. https://arxiv.org/abs//2312.09300 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.a...

[short] Weight Subcloning: Direct Initialization of Transformers Using Larger Pretrained Ones 18.12.2023

The paper introduces a technique called weight subcloning to transfer knowledge from a pretrained model to smaller variants, improving training speed by initializing weights from larger models. https://arxiv.org/abs//2312.09299 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 S...

Weight Subcloning: Direct Initialization of Transformers Using Larger Pretrained Ones 18.12.2023

The paper introduces a technique called weight subcloning to transfer knowledge from a pretrained model to smaller variants, improving training speed by initializing weights from larger models. https://arxiv.org/abs//2312.09299 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 S...

[short] TinyGSM: achieving on GSM8k with small language models 16.12.2023

The paper explores the use of small language models for solving grade school math problems. By using a high-quality dataset and a verifier model, they achieve 81.5% accuracy, outperforming larger models. https://arxiv.org/abs//2312.09241 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id16...

TinyGSM: achieving on GSM8k with small language models 16.12.2023

The paper explores the use of small language models for solving grade school math problems. By using a high-quality dataset and a verifier model, they achieve 81.5% accuracy, outperforming larger models. https://arxiv.org/abs//2312.09241 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id16...

[short] Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking 15.12.2023

Reward models are important for aligning language models with human preferences, but they can be exploited by the model. Training an ensemble of reward models can mitigate this issue, but it does not completely eliminate reward hacking. https://arxiv.org/abs//2312.09244 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple...

Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking 15.12.2023

Reward models are important for aligning language models with human preferences, but they can be exploited by the model. Training an ensemble of reward models can mitigate this issue, but it does not completely eliminate reward hacking. https://arxiv.org/abs//2312.09244 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple...

[short] Vision-Language Models as a Source of Rewards 15.12.2023

The paper explores using off-the-shelf vision-language models as sources of rewards for reinforcement learning agents. It demonstrates how rewards for visual achievement of language goals can be derived from CLIP models, leading to more capable RL agents. https://arxiv.org/abs//2312.09187 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: htt...

Vision-Language Models as a Source of Rewards 15.12.2023

The paper explores using off-the-shelf vision-language models as sources of rewards for reinforcement learning agents. It demonstrates how rewards for visual achievement of language goals can be derived from CLIP models, leading to more capable RL agents. https://arxiv.org/abs//2312.09187 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: htt...

[short] SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention 14.12.2023

SwitchHead is a novel method that reduces compute and memory requirements in Transformers, achieving speedup while maintaining language modeling performance. It uses Mixture-of-Experts layers and requires fewer attention matrices than standard Transformers. https://arxiv.org/abs//2312.07987 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: h...

SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention 14.12.2023

SwitchHead is a novel method that reduces compute and memory requirements in Transformers, achieving speedup while maintaining language modeling performance. It uses Mixture-of-Experts layers and requires fewer attention matrices than standard Transformers. https://arxiv.org/abs//2312.07987 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: h...

[short] Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations 13.12.2023

The paper discusses the use of machine learning algorithms to predict the risk of cardiovascular disease in patients with diabetes. https://arxiv.org/abs//2312.06674 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations 13.12.2023

The paper discusses the use of machine learning algorithms to predict the risk of cardiovascular disease in patients with diabetes. https://arxiv.org/abs//2312.06674 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers

[short] Steering Llama 2 via Contrastive Activation Addition 13.12.2023

Contrastive Activation Addition (CAA) is a method for steering language models by modifying activations during forward passes. It outperforms traditional methods and provides insights into how concepts are represented in Large Language Models (LLMs). https://arxiv.org/abs//2312.06681 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://...

Steering Llama 2 via Contrastive Activation Addition 13.12.2023

Contrastive Activation Addition (CAA) is a method for steering language models by modifying activations during forward passes. It outperforms traditional methods and provides insights into how concepts are represented in Large Language Models (LLMs). https://arxiv.org/abs//2312.06681 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://...

[short] LLM360: Towards Fully Transparent Open-Source LLMs 12.12.2023

LLM360 is an initiative to fully open-source Large Language Models (LLMs) by providing all training code, data, model checkpoints, and intermediate results. The goal is to support open and collaborative AI research and transparency in the LLM training process. https://arxiv.org/abs//2312.06550 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts...

LLM360: Towards Fully Transparent Open-Source LLMs 12.12.2023

LLM360 is an initiative to fully open-source Large Language Models (LLMs) by providing all training code, data, model checkpoints, and intermediate results. The goal is to support open and collaborative AI research and transparency in the LLM training process. https://arxiv.org/abs//2312.06550 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts...

[short] Unlocking Anticipatory Text Generation: A Constrained Approach for Faithful Decoding with Large Language Models 12.12.2023

This paper proposes a method for minimizing undesirable behaviors and ensuring faithfulness to instructions in text generation using large language models. The approach is effective across various text generation tasks. https://arxiv.org/abs//2312.06149 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/a...

Unlocking Anticipatory Text Generation: A Constrained Approach for Faithful Decoding with Large Language Models 12.12.2023

This paper proposes a method for minimizing undesirable behaviors and ensuring faithfulness to instructions in text generation using large language models. The approach is effective across various text generation tasks. https://arxiv.org/abs//2312.06149 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/a...

[short] Context Tuning for Retrieval Augmented Generation 12.12.2023

The paper proposes Context Tuning for Retrieval Augmented Generation (RAG), which improves tool retrieval and plan generation by using a smart context retrieval system. Empirical results show significant improvements in recall, accuracy, and reduction in hallucination compared to existing methods. https://arxiv.org/abs//2312.05708 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.t...

Context Tuning for Retrieval Augmented Generation 12.12.2023

The paper proposes Context Tuning for Retrieval Augmented Generation (RAG), which improves tool retrieval and plan generation by using a smart context retrieval system. Empirical results show significant improvements in recall, accuracy, and reduction in hallucination compared to existing methods. https://arxiv.org/abs//2312.05708 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.t...

[short] Using Captum to Explain Generative Language Models 12.12.2023

The paper introduces new features in the Captum library for model explainability in PyTorch, specifically designed for analyzing generative language models. It provides an overview of the functionalities and example applications for understanding learned associations in these models. https://arxiv.org/abs//2312.05491 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arx...

Using Captum to Explain Generative Language Models 12.12.2023

The paper introduces new features in the Captum library for model explainability in PyTorch, specifically designed for analyzing generative language models. It provides an overview of the functionalities and example applications for understanding learned associations in these models. https://arxiv.org/abs//2312.05491 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arx...

[short] HALO: An Ontology for Representing Hallucinations in Generative Models 11.12.2023

The paper introduces HALO, a formal ontology for describing and representing hallucinations in large language models like ChatGPT, addressing the lack of a formal model for this issue. https://arxiv.org/abs//2312.05209 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: h...

HALO: An Ontology for Representing Hallucinations in Generative Models 11.12.2023

The paper introduces HALO, a formal ontology for describing and representing hallucinations in large language models like ChatGPT, addressing the lack of a formal model for this issue. https://arxiv.org/abs//2312.05209 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: h...

Listen to the Arxiv Papers podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.