fr

The AI Illuminators

AI Illuminated

A new way to keep up with AI research. Delivered to your ears. Illuminated by AI.Part of the GenAI4Good initiative.

N'hésitez pas à visiter le site du podcast et à soutenir son créateur : podcasters.spotify.com

Auteur

The AI Illuminators

Catégorie

Education

Site du podcast

podcasters.spotify.com

Dernier épisode

7 déc. 2024

Où écouter ?

Les podcasts dans l'appli Replaio Radio Bientôt disponible

Les podcasts arrivent très bientôt dans l'appli. Installe-la dès maintenant et découvre en avant-première une toute nouvelle façon de vivre les podcasts

Télécharger sur Google Play Installe-la gratuitement Android près de 10 M de téléchargements · note de 4,8 iOS bientôt

Épisodes

MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos 07.12.2024

0:00 Introduction  0:20 Limitations of traditional SfM and SLAM techniques. 0:57 Shortcomings of existing neural network methods. 1:07 MegaSaM's approach: balance of accuracy, speed, and robustness. 1:31 Differentiable bundle adjustment (BA) layer. 2:03 Integration of monocular depth priors and motion probability maps. 2:37 Uncertainty-aware global BA scheme. 3:14 Two-stage training scheme. 3:45 C...

SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models 12.11.2024

[00:00] SVD-Quant: 4-bit diffusion model quantization [00:27] Challenge: Outlier sensitivity in 4-bit quantization [00:59] Solution: Smoothing + SVD approach [01:37] Technical: SVD's role in low-rank approximation [02:08] Nunchuku: New inference engine with kernel fusion [02:35] Comparison: INT4 vs FP4 quantization methods [03:00] Results: 3.5x memory reduction on Flux-1.0 [03:44] Feature: Seamles...

D3RoMa: Disparity Diffusion-based Depth Sensing for Material-Agnostic Robotic Manipulation 11.11.2024

[00:00] Intro [00:18] Current limitations in depth-sensing technology [00:56] D3RoMa's diffusion model approach to depth estimation [01:47] Integration of geometric constraints in the model [02:27] HiSS: New dataset for transparent/specular objects [03:18] Benchmark results showing major accuracy improvements [04:02] Current limitations and future development areas [05:34] Technical details of HiS...

Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers 09.11.2024

[00:00] Intro [00:21] Key problem: Poor generalization in robotic learning [00:51] HPT: New transformer architecture for robotics [00:59] Core components of HPT architecture [01:44] Scale analysis: Data and model size impacts [02:16] Training data: Real robots, simulations, human videos [02:54] Results: 20% improvement on new tasks [04:04] Real-world testing limitations [05:18] Future additions: T...

HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots 04.11.2024

[00:00] Introduction to Hover: Neural Whole Body Controller for Humanoids [00:15] Problem: Current controllers lack versatility across tasks [00:50] Human motion imitation as a unified control approach [01:23] Policy distillation: Learning from an oracle policy [02:01] Command space: Kinematic, joint angle, and root tracking modes [02:34] Motion retargeting: From human data to robot movements [03:...

Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning 03.11.2024

[00:00] Intro [00:24] Tackles RL challenges using a visual backbone, efficient RL, and human feedback. [01:20] Pretrained backbone boosts stability and exploration efficiency. [02:06] RLPD combines offline data and human corrections effectively. [02:57] Human-guided interventions reduce errors, enabling gradual autonomy. [03:42] System choices aid spatial generalization and safe exploration. [04:4...

Local Policies Enable Zero-shot Long Horizon Manipulation 02.11.2024

[00:00] Paper intro: Zero-shot robotic manipulation via local policies [00:26] Key challenges: Limited generalization and sim-to-real transfer [01:03] Local policies: Task decomposition through localized focus regions [01:38] Foundation models: VLMs for task understanding [02:07] Training approach: Simulation-based RL + visuomotor policy distillation [02:46] Implementation: Depth maps and impedanc...

MENTOR: Mixture-of-Experts Network with Task-Oriented Perturbation for Visual Reinforcement Learning 30.10.2024

[00:00] Introduction to Mentor system for visual RL [00:29] Problem: Sample inefficiency in robotic learning [00:59] Innovation: Mixture of Experts (MoE) architecture [01:55] Results: MoE achieves 100% success in multi-task testing [02:33] Feature: Task-oriented perturbation for exploration [03:55] Real-world testing: 83% success in robotic tasks [04:33] Study: MoE and perturbation each boost perf...

SkillMimicGen: Automated Demonstration Generation for Efficient Skill Learning and Deployment 29.10.2024

[00:00] SkillGen: AI system for robotic learning and automation. [00:19] Core Innovation: Automated dataset generation from minimal human input. [01:12] Skill Segmentation: Smart system for breaking down and adapting complex tasks. [01:59] Hybrid Skill Policy: Framework for controlling robot actions and task completion. [02:50] Performance Results: 75.4% success rate, generating 24,000+ demonstrat...

LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias 28.10.2024

[00:00] Intro to LVSM: Novel transformer for view synthesis [00:14] Problems with existing 3D synthesis methods [00:59] LVSM architecture: encoder-decoder vs decoder-only [01:41] Performance trade-offs between architectures [02:13] Using Pluecker rays for implicit 3D geometry [02:49] Zero-shot capabilities with varying input views [03:23] Training stability and technical solutions [03:59] Training...

Dynamic 3D Gaussian Tracking for Graph-Based Neural Dynamics Modeling 27.10.2024

[00:00] Introduction to 3D Gaussian tracking for robotic manipulation [00:26] Limitations of current video prediction methods [01:11] Advantages of 3D Gaussian representation [02:04] Graph Neural Networks for modeling object dynamics [02:54] Control particle implementation and computation reduction [03:42] Physics-based optimization for prediction stability [04:25] Integration with real-world robo...

SPIRE: Synergistic Planning, Imitation, and Reinforcement Learning for Long-Horizon Manipulation 25.10.2024

[00:00] Introduction [00:20] Core limitations in robot manipulation: challenges with RL and IL [01:08] SPIRE's hybrid approach: combining task planning with learning methods [01:44] TAMP-gated learning: selective application of learned policies [02:20] Training innovations: warm-starting RL and KL-divergence implementation [02:59] Results: 35-50% performance gain, 6x more data efficient [04:04] Mu...

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation 24.10.2024

[00:00] VILA-U: A unified visual AI model [00:29] Problem: Inefficiency of separate visual modules [01:11] Vision tower: Novel quantization approach [02:09] Training strategy: CLIP-based staged learning [03:03] RVQ technique: Enhanced visual representation [03:47] Multi-modal training: Text-image-video fusion [04:35] Performance: Results and current limitations [05:23] Impact: Contrastive loss eff...

CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object Dynamics 23.10.2024

[00:00] Intro [00:31] Challenge: Limited multi-humanoid training data [00:55] CooHOI's two-phase learning framework [01:49] Object dynamics as implicit agent communication [02:25] Bounding box strategy for long objects [03:07] Results: Superior performance vs baselines [03:42] Ablation study findings: Key system components [04:18] Limitation: Basic hand manipulation only [04:57] Impact: New approa...

SynFlowNet: Design of Diverse and Novel Molecules with Synthesis Constraints 22.10.2024

[00:00] Introduction to SynFlowNet [00:29] Problem: AI-generated molecules often can't be synthesized [01:17] Solution: SynFlowNet - uses real chemical reactions [02:03] GFlowNets: Enables diverse molecule generation [02:47] Scalability: Morgan fingerprints handle 200K+ compounds [03:14] Challenge: Solving backward trajectory issues [04:14] Results: Better synthesis rates and molecular diversity [...

L3DG: Latent 3D Gaussian Diffusion 21.10.2024

[00:00] Intro to L3DG for 3D modeling [00:32] Solving room-sized 3D scene complexity [01:36] VQ-VAE compresses 3D Gaussian representation [02:41] Generative sparse transpose convolution [03:20] Latent diffusion for scene generation [04:30] Visual improvements over baselines [05:14] Scalability challenges for room-sized scenes [06:13] Spherical harmonics for view dependence [06:58] RGB and perceptu...

The Ingredients for Robotic Diffusion Transformers 20.10.2024

[00:00] Intro [00:33] Combining transformers & diffusion models [01:12] Key design: Scalable attention blocks (AdaLN) [02:30] Efficient observation tokenization [03:45] DiT Block policy architecture overview [04:20] BiPlay dataset introduction [04:53] Performance improvements over baselines [05:30] Key findings from ablations [06:08] Generalization to different robot types [06:43] Simulation v...

Estimating Body and Hand Motion in an Ego-sensed World 19.10.2024

[00:00] Introduction to EgoAllo system [00:38] Challenges in egocentric motion estimation [01:20] Importance of spatial/temporal invariance [02:11] Comparison of conditioning parameterizations [02:57] Integration of hand observations [03:50] Global alignment phase [04:28] Guidance losses in sampling [05:03] Handling longer sequences [05:35] Evaluation results [06:30] System limitations and future...

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation 19.10.2024

[00:00] Intro [00:28] Limitation of existing unified models [00:57] Janus's decoupled visual encoding solution [01:18] Advantages of decoupling [02:03] Janus architecture [02:50] Three-stage training [03:41] Ablation studies [04:23] Extensions for Janus [05:10] Performance gains [05:47] Current limitations [06:31] Impact of simplicity and extensibility [07:10] Qualitative results [08:18] Potential...

One Step Diffusion via Shortcut Models 19.10.2024

[00:00] Introduction [00:23] Computational cost of traditional diffusion models [00:59] Reducing iterations in image generation [01:06] Shortcut models [01:39] Training process and self-consistency property [02:22] Advantages over other methods [03:05] Results on image generation benchmarks [03:45] Application to robotic control [04:15] Limitations and future work [04:54] Best practices Authors: K...

Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning 18.10.2024

[00:00] Intro: Research on multi-task learning in LLMs [00:38] Balancing safety and performance in multilingual settings [01:17] Model merging techniques explored [02:16] Model merging outperforms data mixing [02:49] Merging monolingual models improves multilingual capabilities [03:28] Key ablation studies [04:14] Safety and performance evaluation metrics [04:52] Effectiveness variations across la...

CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos 18.10.2024

[00:00] Intro: CoTracker 3 paper [00:17] Challenge: Point tracker training [00:51] Existing methods' limitations [01:05] CoTracker 3 innovations [01:41] Novel semi-supervised training [02:24] Online vs offline versions [03:02] Performance improvements [03:47] Ablation study insights [04:29] Limitations and future work Authors: Nikita Karaev, Iurii Makarov, Jianyuan Wang, Natalia Neverova, Andrea V...

nGPT: Normalized Transformer with Representation Learning on the Hypersphere 18.10.2024

[00:00] Introduction [00:30] Consistent unit norm normalization in NGPT [01:08] Mathematical mechanism behind faster convergence [01:52] Elimination of weight decay in NGPT [02:21] Role of learnable eigen learning rates in optimization [03:04] Discussion on training speedup vs. per-step computation time [03:46] Condition number differences between GPT and NGPT [04:18] Ablation studies on scaling f...

RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation 18.10.2024

[00:00] Intro [00:20] Challenges in bimanual manipulation models [01:00] Robotics Diffusion Transformer (RDT) approach [01:34] RDT architecture design [02:22] Data scarcity and unified action space [03:07] Multi-task bimanual dataset [03:49] RDT's experimental results [04:30] Benefits of large-scale pre-training [05:12] Diffusion models in robotics [06:02] RDT's architectural modifications [06:57]...

Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models 18.10.2024

[00:00] Continuous-time consistency models (CTMs) [00:20] Limitations of existing CTMs [00:51] TrigFlow: New CTM formulation [01:42] CTM training instability [02:21] Training objective modifications [02:55] Scaling CTMs to 1.5 billion parameters [03:37] Comparison with state-of-the-art models [04:14] Consistency training vs. distillation [04:52] CTMs vs. variational score distillation [05:20] Key...

Écoute le podcast AI Illuminated sur Replaio

La radio et les podcasts dans une seule appli - gratuite, sans inscription. Installe-la dès aujourd'hui et ne rate pas le lancement

Télécharger sur Google Play

Replaio n'est pas éditeur de podcasts ; les noms des émissions, les visuels et l'audio appartiennent à leurs auteurs et sont diffusés via des flux RSS publics