Dr. Satya Mallick
Artificial Intelligence : Papers & Concepts
This podcast is for AI engineers and researchers. We utilize AI to explain papers and concepts in AI.
Autor
Dr. Satya Mallick
Kategorie
Podcast-Website
Neueste Folge
23. Apr 2026
Wo hören?
Podcasts in der App Replaio Radio Bald verfügbarPodcasts kommen bald in die App. Installiere sie jetzt und erlebe als Erster einen ganz neuen Blick auf Podcasts
Folgen
DeepSeek mHC Video 05.01.2026 12:01
Why do some large AI models suddenly collapse during training—and how can geometry prevent it? In this episode of Artificial Intelligence: Papers and Concepts , we break down DeepSeek AI's Manifold-Constrained Hyperconnections (mHC) , a new architectural approach that fixes training instability in large language models. We explore why traditional hyperconnections caused catastrophic signal explosi...
Chinchilla Scaling Law Video 18.12.2025 12:50
In this episode of Artificial Intelligence: Papers and Concepts , curated by Dr. Satya Mallick , we break down DeepMind's 2022 paper "Training Compute-Optimal Large Language Models" —the work that challenged the "bigger is always better" era of LLM scaling. You'll learn why many famous models were under-trained , what it means to be compute-optimal , and why the best performance comes from scalin...
Gradient-Based Planning Video 13.12.2025 12:56
How should an AI or robot decide what to do next? In this episode, we explore a new approach to planning that rethinks how world models are trained. The episode is based on the paper "Closing the Train-Test Gap in World Models for Gradient-Based Planning" Many AI systems can predict the future accurately, yet struggle when asked to plan actions efficiently. We explain why this train–test mismatch...
SAM3D: The Next Leap in 3D Understanding 10.12.2025 13:56
Forget flat photos—SAM3D is rewriting how machines understand the world. In this episode, we break down the groundbreaking new model that takes the core ideas of Meta's Segment Anything Model and expands them into the third dimension , enabling instant 3D segmentation from just a single image. We start with the limitations of traditional 2D vision systems and explain why 3D understanding has alway...
DINOv3 : A new Self-Supervised Learning (SSL) Vision Language Model (VLM) 29.10.2025 13:37
In this episode, we explore DINOv3, a new self-supervised learning (SSL) vision foundation model from Meta AI Research, emphasizing its ability to scale effortlessly to massive datasets and large architectures without relying on manual data annotation. The core innovations are scaling model and dataset size , introducing Gram anchoring to prevent the degradation of dense feature maps during lon...
dots.ocr SOTA Document Parsing in a Compact VLM 28.10.2025 12:49
dots.ocr is a powerful, multilingual document parsing model from rednote-hilab that achieves state-of-the-art performance by unifying layout detection and content recognition within a single, efficient vision-language model (VLM). Built upon a compact 1.7B paramete r Large Language Model (LLM), it offers a streamlined alternative to complex, multi-model pipelines, enabling faster inference spee...
DeepSeek-OCR : A Revolutionary Idea 23.10.2025 14:33
In this episode, we dive deep into DeepSeek-OCR , a cutting-edge open-source Optical Character Recognition (OCR) / Text Recognition model that's redefining accuracy and efficiency in document understanding. DeepSeek-OCR flips long-context processing on its head by rendering text as images and then decoding it back—shrinking context length by 7–20× while preserving high fidelity. We break down how...
nanochat by Karpathy - How to build your own ChatGPT for $100 21.10.2025 12:18
"The best ChatGPT that $100 can buy. " That's Andrej Karpathy's positioning for nanochat —a compact, end ‑ to ‑ end stack that goes from tokenizer training to a ChatGPT ‑ style web UI in a few thousand lines of Python (plus a tiny Rust tokenizer). It's meant to be read , hacked , and run so students, researchers, and tech enthusiats can understand the entire pipeline needed to train a baby versio...
SmolVLM: Small Yet Mighty Vision Language Model 01.10.2025 14:26
In this episode of Artificial Intelligence: Papers and Concepts, we explore SmolVLM, a family of compact yet powerful vision language models (VLMs) designed for efficiency. Unlike large VLMs that require significant computational resources, SmolVLM is engineered to run on everyday devices like smartphones and laptops. We dive into the research paper SmolVLM: Redefining Small and Efficient Multimod...
Common Pitfalls in Computer Vision & AI Projects (and How to Avoid Them) 01.10.2025 17:48
In this episode, we dig deep into the unglamorous side of AI and computer vision projects — the mistakes, misfires, and blind spots that too often derail even the most promising teams. Based on BigVision.ai's playbook "Common Pitfalls in Computer Vision & AI Projects" , we walk through a field-tested catalog of pitfalls drawn from real failures and successes. We cover: Why ambiguous problem statem...
Ähnliche Podcasts
Replaio ist kein Herausgeber von Podcasts; die Namen der Sendungen, Cover und Audioinhalte gehören ihren Autoren und werden über öffentliche RSS-Feeds verbreitet