Igor Melnyk

Arxiv Papers

Science EN ↓ 2489 episodes

Running out of time to catch up with new arXiv papers? We take the most impactful papers and present them as convenient podcasts. If you're a visual learner, we offer these papers in an engaging video format. Our service fills the gap between overly brief paper summaries and time-consuming full paper reads. You gain academic insights in a time-efficient, digestible format. Code behind this work: https://github.com/imelnyk/ArxivPapers

Author

Igor Melnyk

Category

Science

Podcast website

github.com

Latest episode

Sep 1, 2025

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android 5M+ downloads · 4.8 rating iOS soon

Episodes

[QA] Do Large Language Model Benchmarks Test Reliability? 06.02.2025

The paper highlights the lack of reliable benchmarks for large language models, proposing "platinum benchmarks" to minimize label errors and revealing persistent model failures in simple tasks. https://arxiv.org/abs//2502.03461 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 S...

Do Large Language Model Benchmarks Test Reliability? 06.02.2025

The paper highlights the lack of reliable benchmarks for large language models, proposing "platinum benchmarks" to minimize label errors and revealing persistent model failures in simple tasks. https://arxiv.org/abs//2502.03461 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id16...

Detecting Strategic Deception Using Linear Probes 06.02.2025

The study evaluates linear probes for detecting AI deception, achieving high accuracy in distinguishing honest from deceptive outputs, but concludes that current methods are insufficient for robust defense. https://arxiv.org/abs//2502.03407 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/i...

[QA] Evaluation of Large Language Models via Coupled Token Generation 05.02.2025

This paper argues for controlling randomization in evaluating large language models, showing that coupled autoregressive generation can yield different rankings than vanilla methods, despite fewer required samples. https://arxiv.org/abs//2502.01754 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-...

Evaluation of Large Language Models via Coupled Token Generation 05.02.2025

This paper argues for controlling randomization in evaluating large language models, showing that coupled autoregressive generation can yield different rankings than vanilla methods, despite fewer required samples. https://arxiv.org/abs//2502.01754 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-...

[QA] Learning the RoPEs: Better 2D and 3D Position Encodings with STRING 05.02.2025

https://arxiv.org/abs//2502.02562 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers

Learning the RoPEs: Better 2D and 3D Position Encodings with STRING 05.02.2025

https://arxiv.org/abs//2502.02562 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers

[QA] Should You Use Your Large Language Model to Explore or Exploit? 04.02.2025

The study assesses large language models' effectiveness in decision-making tasks, finding they struggle with exploitation but assist in exploring large action spaces, outperforming simple linear regression in exploration. https://arxiv.org/abs//2502.00225 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/pod...

Should You Use Your Large Language Model to Explore or Exploit? 04.02.2025

The study assesses large language models' effectiveness in decision-making tasks, finding they struggle with exploitation but assist in exploring large action spaces, outperforming simple linear regression in exploration. https://arxiv.org/abs//2502.00225 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/pod...

[QA] Harmonic Loss Trains Interpretable AI Models 04.02.2025

This paper presents harmonic loss as a superior alternative to cross-entropy loss, enhancing interpretability, convergence speed, and performance in neural networks and large language models across various datasets. https://arxiv.org/abs//2502.01628 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv...

Harmonic Loss Trains Interpretable AI Models 04.02.2025

This paper presents harmonic loss as a superior alternative to cross-entropy loss, enhancing interpretability, convergence speed, and performance in neural networks and large language models across various datasets. https://arxiv.org/abs//2502.01628 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv...

[QA] Trading inference-time compute for adversarial robustness. 03.02.2025

Increasing inference-time compute enhances the robustness of reasoning models against adversarial attacks, with success rates decreasing as compute increases, suggesting potential for improved adversarial resilience in Large Language Models. https://arxiv.org/abs//2501.18841 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts....

Trading inference-time compute for adversarial robustness. 03.02.2025

Increasing inference-time compute enhances the robustness of reasoning models against adversarial attacks, with success rates decreasing as compute increases, suggesting potential for improved adversarial resilience in Large Language Models. https://arxiv.org/abs//2501.18841 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts....

[QA] LLMs can see and hear without any training 02.02.2025

MILS is a training-free method that enhances LLMs with multimodal capabilities, achieving state-of-the-art results in zero-shot captioning and media generation through iterative scoring and feedback. https://arxiv.org/abs//2501.18096 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id169247...

LLMs can see and hear without any training 02.02.2025

MILS is a training-free method that enhances LLMs with multimodal capabilities, achieving state-of-the-art results in zero-shot captioning and media generation through iterative scoring and feedback. https://arxiv.org/abs//2501.18096 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id169247...

[QA] o3-mini vs DeepSeek-R1: Which One is Safer? 02.02.2025

The paper assesses the safety of DeepSeek-R1 and OpenAI's o3-mini, revealing DeepSeek-R1's higher unsafety rate (11.98%) compared to o3-mini (1.19%) using the ASTRAL testing tool. https://arxiv.org/abs//2501.18438 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify...

o3-mini vs DeepSeek-R1: Which One is Safer? 02.02.2025

The paper assesses the safety of DeepSeek-R1 and OpenAI's o3-mini, revealing DeepSeek-R1's higher unsafety rate (11.98%) compared to o3-mini (1.19%) using the ASTRAL testing tool. https://arxiv.org/abs//2501.18438 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016 Spotify...

[QA] Optimizing Large Language Model Training Using FP4 Quantization 01.02.2025

This paper presents a novel FP4 training framework for large language models, enhancing efficiency and accuracy through innovative quantization techniques and mixed-precision training, suitable for next-gen hardware. https://arxiv.org/abs//2501.17116 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxi...

Optimizing Large Language Model Training Using FP4 Quantization 01.02.2025

This paper presents a novel FP4 training framework for large language models, enhancing efficiency and accuracy through innovative quantization techniques and mixed-precision training, suitable for next-gen hardware. https://arxiv.org/abs//2501.17116 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxi...

[QA] People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text 01.02.2025

This study evaluates human detection of AI-generated text, revealing that frequent LLM users excel at identifying it, outperforming automated detectors, and providing insights into their detection strategies. https://arxiv.org/abs//2501.15654 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers...

People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text 01.02.2025

This study evaluates human detection of AI-generated text, revealing that frequent LLM users excel at identifying it, outperforming automated detectors, and providing insights into their detection strategies. https://arxiv.org/abs//2501.15654 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers...

[QA] Large Language Models Think Too Fast To Explore Effectively 31.01.2025

This study examines Large Language Models' exploration abilities in open-ended tasks, revealing they generally underperform compared to humans, highlighting limitations and suggesting improvements for adaptability. https://arxiv.org/abs//2501.18009 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/ar...

Large Language Models Think Too Fast To Explore Effectively 31.01.2025

This study examines Large Language Models' exploration abilities in open-ended tasks, revealing they generally underperform compared to humans, highlighting limitations and suggesting improvements for adaptability. https://arxiv.org/abs//2501.18009 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/ar...

[QA] Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs 31.01.2025

The paper identifies "underthinking" in LLMs, where frequent thought switching hampers reasoning depth. It proposes a new strategy to improve accuracy in complex tasks without model fine-tuning. https://arxiv.org/abs//2501.18585 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1...

Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs 31.01.2025

The paper identifies "underthinking" in LLMs, where frequent thought switching hampers reasoning depth. It proposes a new strategy to improve accuracy in complex tasks without model fine-tuning. https://arxiv.org/abs//2501.18585 YouTube: https://www.youtube.com/@ArxivPapers TikTok: https://www.tiktok.com/@arxiv_papers Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1...

Listen to the Arxiv Papers podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.