Igor Melnyk

Arxiv Papers

Science EN ↓ 2489 episodes

Running out of time to catch up with new arXiv papers? We take the most impactful papers and present them as convenient podcasts. If you're a visual learner, we offer these papers in an engaging video format. Our service fills the gap between overly brief paper summaries and time-consuming full paper reads. You gain academic insights in a time-efficient, digestible format. Code behind this work: https://github.com/imelnyk/ArxivPapers

Author

Igor Melnyk

Category

Science

Podcast website

github.com

Latest episode

Sep 1, 2025

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android almost 10M downloads · 4.8 rating iOS soon

Episodes

Can AI Be as Creative as Humans? 06.01.2024

This paper introduces the concept of Relative Creativity to evaluate the creative potential of AI models. It proposes Statistical Creativity as a quantifiable measure and provides a training guideline for fostering creativity in AI. https://arxiv.org/abs//2401.01623 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com...

[short] Improving Text Embeddings with Large Language Models 06.01.2024

The paper introduces a simple method for obtaining high-quality text embeddings using synthetic data and minimal training steps. The method outperforms existing approaches on text embedding benchmarks without using labeled data, and achieves state-of-the-art results when combined with labeled data. https://arxiv.org/abs//2401.00368 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www....

Improving Text Embeddings with Large Language Models 06.01.2024

The paper introduces a simple method for obtaining high-quality text embeddings using synthetic data and minimal training steps. The method outperforms existing approaches on text embedding benchmarks without using labeled data, and achieves state-of-the-art results when combined with labeled data. https://arxiv.org/abs//2401.00368 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www....

[short] LLAMA PRO: Progressive LLaMA with Block Expansion 06.01.2024

The paper proposes a new method for improving Large Language Models (LLMs) without forgetting previous knowledge. The method involves expanding Transformer blocks and tuning them using new corpus. The resulting model, LLAMA PRO-8.3B, performs well in general tasks, programming, and mathematics, surpassing existing models in the LLaMA family. The findings contribute to the development of advanced l...

LLAMA PRO: Progressive LLaMA with Block Expansion 06.01.2024

 The paper proposes a new method for improving Large Language Models (LLMs) without forgetting previous knowledge. The method involves expanding Transformer blocks and tuning them using new corpus. The resulting model, LLAMA PRO-8.3B, performs well in general tasks, programming, and mathematics, surpassing existing models in the LLaMA family. The findings contribute to the development of advanced...

[short] Instruct-Imagen: Image Generation with Multi-modal Instruction 05.01.2024

The paper introduces Instruct-Imagen, a model that uses natural language to generate images based on various instructions. It outperforms previous models and shows promise for generalizing to new tasks. https://arxiv.org/abs//2401.01952 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id169...

Instruct-Imagen: Image Generation with Multi-modal Instruction 05.01.2024

The paper introduces Instruct-Imagen, a model that uses natural language to generate images based on various instructions. It outperforms previous models and shows promise for generalizing to new tasks. https://arxiv.org/abs//2401.01952 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id169...

[short] LLM Augmented LLMs: Expanding Capabilities through Composition 05.01.2024

This paper introduces CALM, a method for efficiently and practically combining existing language models with more specific models to enhance their capabilities. CALM allows for the augmentation of language models on new tasks while preserving their existing capabilities. Experimental results show significant improvements in translation, arithmetic reasoning, code generation, and explanation tasks....

LLM Augmented LLMs: Expanding Capabilities through Composition 05.01.2024

This paper introduces CALM, a method for efficiently and practically combining existing language models with more specific models to enhance their capabilities. CALM allows for the augmentation of language models on new tasks while preserving their existing capabilities. Experimental results show significant improvements in translation, arithmetic reasoning, code generation, and explanation tasks....

[short] TinyLlama: An Open-Source Small Language Model 05.01.2024

TinyLlama is a compact language model pretrained on 1 trillion tokens. It achieves better computational efficiency and outperforms existing models of similar size in downstream tasks. https://arxiv.org/abs//2401.02385 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: ht...

TinyLlama: An Open-Source Small Language Model 05.01.2024

TinyLlama is a compact language model pretrained on 1 trillion tokens. It achieves better computational efficiency and outperforms existing models of similar size in downstream tasks. https://arxiv.org/abs//2401.02385 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: ht...

[short] GPT-4V(ision) is a Generalist Web Agent, if Grounded 04.01.2024

The paper explores the potential of large multimodal models (LMMs) as generalist web agents that can complete tasks on websites. They propose a web agent called SEEACT and evaluate its performance on the MIND2WEB benchmark. The results show that LMMs like GPT-4V have the potential to complete tasks on live websites, but grounding remains a challenge. https://arxiv.org/abs//2401.01614 YouTube: http...

GPT-4V(ision) is a Generalist Web Agent, if Grounded 04.01.2024

The paper explores the potential of large multimodal models (LMMs) as generalist web agents that can complete tasks on websites. They propose a web agent called SEEACT and evaluate its performance on the MIND2WEB benchmark. The results show that LMMs like GPT-4V have the potential to complete tasks on live websites, but grounding remains a challenge. https://arxiv.org/abs//2401.01614 YouTube: http...

[short] aMUSEd: An open MUSE reproduction 04.01.2024

aMUSEd is a lightweight masked image model (MIM) for text-to-image generation that is faster and more interpretable than latent diffusion. It can be fine-tuned to learn additional styles and is effective for large-scale text-to-image generation. https://arxiv.org/abs//2401.01808 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podca...

aMUSEd: An open MUSE reproduction 04.01.2024

aMUSEd is a lightweight masked image model (MIM) for text-to-image generation that is faster and more interpretable than latent diffusion. It can be fine-tuned to learn additional styles and is effective for large-scale text-to-image generation. https://arxiv.org/abs//2401.01808 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podca...

[short] Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models 03.01.2024

The paper introduces a new fine-tuning method called SPIN, which uses self-play to improve the performance of Large Language Models (LLMs) without the need for additional human-annotated data. The method is shown to outperform models trained with direct preference optimization (DPO) and extra GPT-4 preference data. https://arxiv.org/abs//2401.01335 YouTube: https://www.youtube.com/@ArxivPapers Tik...

Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models 03.01.2024

The paper introduces a new fine-tuning method called SPIN, which uses self-play to improve the performance of Large Language Models (LLMs) without the need for additional human-annotated data. The method is shown to outperform models trained with direct preference optimization (DPO) and extra GPT-4 preference data. https://arxiv.org/abs//2401.01335 YouTube: https://www.youtube.com/@ArxivPapers Tik...

[short] Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws 02.01.2024

The paper modifies the DeepMind Chinchilla scaling laws for large language models (LLMs) to include the cost of inference. The analysis suggests that LLM researchers should train smaller and longer models than the Chinchilla-optimal for large inference demand. https://arxiv.org/abs//2401.00448 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts...

Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws 02.01.2024

The paper modifies the DeepMind Chinchilla scaling laws for large language models (LLMs) to include the cost of inference. The analysis suggests that LLM researchers should train smaller and longer models than the Chinchilla-optimal for large inference demand. https://arxiv.org/abs//2401.00448 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts...

[short] MosaicBERT: A Bidirectional Encoder Optimized for Fast Pretraining 02.01.2024

The paper introduces MosaicBERT, an optimized BERT-style encoder architecture and training recipe that enables fast and cost-effective pretraining of custom models. The approach achieves high accuracy and pretraining speed, making it a valuable tool for researchers and engineers. https://arxiv.org/abs//2312.17482 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_p...

MosaicBERT: A Bidirectional Encoder Optimized for Fast Pretraining 02.01.2024

The paper introduces MosaicBERT, an optimized BERT-style encoder architecture and training recipe that enables fast and cost-effective pretraining of custom models. The approach achieves high accuracy and pretraining speed, making it a valuable tool for researchers and engineers. https://arxiv.org/abs//2312.17482 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_p...

[short] Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models 01.01.2024

The paper evaluates the performance of Google's Gemini, a Multimodal Large Language Model (MLLM), in complex reasoning tasks that require the integration of commonsense knowledge across modalities. The study finds that Gemini demonstrates competitive commonsense reasoning capabilities compared to other models. https://arxiv.org/abs//2312.17661 YouTube: https://www.youtube.com/@ArxivPapers TikT...

Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models 01.01.2024

The paper evaluates the performance of Google's Gemini, a Multimodal Large Language Model (MLLM), in complex reasoning tasks that require the integration of commonsense knowledge across modalities. The study finds that Gemini demonstrates competitive commonsense reasoning capabilities compared to other models. https://arxiv.org/abs//2312.17661 YouTube: https://www.youtube.com/@ArxivPapers TikT...

[short] Fast Inference of Mixture-of-Experts Language Models with Offloading 30.12.2023

The paper explores strategies for running large sparse Mixture-of-Experts (MoE) language models on consumer hardware with limited accelerator memory, proposing a novel offloading strategy that allows for efficient execution on desktop hardware and free-tier Google Colab instances. https://arxiv.org/abs//2312.17238 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_...

Fast Inference of Mixture-of-Experts Language Models with Offloading 30.12.2023

The paper explores strategies for running large sparse Mixture-of-Experts (MoE) language models on consumer hardware with limited accelerator memory, proposing a novel offloading strategy that allows for efficient execution on desktop hardware and free-tier Google Colab instances. https://arxiv.org/abs//2312.17238 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_...

Listen to the Arxiv Papers podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.