The AI Illuminators
AI Illuminated
A new way to keep up with AI research. Delivered to your ears. Illuminated by AI.Part of the GenAI4Good initiative.
No dejes de visitar la web del podcast y apoyar a su creador: podcasters.spotify.com
Autor
The AI Illuminators
Categoría
Web del podcast
Último episodio
7 de dic. de 2024
¿Dónde escuchar?
Podcasts en la app Replaio Radio Muy prontoLos podcasts llegarán muy pronto a la app. Instálala ahora y sé el primero en descubrir una forma totalmente nueva de vivir los podcasts
Episodios
MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos 07.12.2024 7:17
0:00 Introduction 0:20 Limitations of traditional SfM and SLAM techniques. 0:57 Shortcomings of existing neural network methods. 1:07 MegaSaM's approach: balance of accuracy, speed, and robustness. 1:31 Differentiable bundle adjustment (BA) layer. 2:03 Integration of monocular depth priors and motion probability maps. 2:37 Uncertainty-aware global BA scheme. 3:14 Two-stage training scheme. 3:45 C...
SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models 12.11.2024 7:20
[00:00] SVD-Quant: 4-bit diffusion model quantization [00:27] Challenge: Outlier sensitivity in 4-bit quantization [00:59] Solution: Smoothing + SVD approach [01:37] Technical: SVD's role in low-rank approximation [02:08] Nunchuku: New inference engine with kernel fusion [02:35] Comparison: INT4 vs FP4 quantization methods [03:00] Results: 3.5x memory reduction on Flux-1.0 [03:44] Feature: Seamles...
D3RoMa: Disparity Diffusion-based Depth Sensing for Material-Agnostic Robotic Manipulation 11.11.2024 15:07
[00:00] Intro [00:18] Current limitations in depth-sensing technology [00:56] D3RoMa's diffusion model approach to depth estimation [01:47] Integration of geometric constraints in the model [02:27] HiSS: New dataset for transparent/specular objects [03:18] Benchmark results showing major accuracy improvements [04:02] Current limitations and future development areas [05:34] Technical details of HiS...
Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers 09.11.2024 11:02
[00:00] Intro [00:21] Key problem: Poor generalization in robotic learning [00:51] HPT: New transformer architecture for robotics [00:59] Core components of HPT architecture [01:44] Scale analysis: Data and model size impacts [02:16] Training data: Real robots, simulations, human videos [02:54] Results: 20% improvement on new tasks [04:04] Real-world testing limitations [05:18] Future additions: T...
HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots 04.11.2024 7:21
[00:00] Introduction to Hover: Neural Whole Body Controller for Humanoids [00:15] Problem: Current controllers lack versatility across tasks [00:50] Human motion imitation as a unified control approach [01:23] Policy distillation: Learning from an oracle policy [02:01] Command space: Kinematic, joint angle, and root tracking modes [02:34] Motion retargeting: From human data to robot movements [03:...
Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning 03.11.2024 9:29
[00:00] Intro [00:24] Tackles RL challenges using a visual backbone, efficient RL, and human feedback. [01:20] Pretrained backbone boosts stability and exploration efficiency. [02:06] RLPD combines offline data and human corrections effectively. [02:57] Human-guided interventions reduce errors, enabling gradual autonomy. [03:42] System choices aid spatial generalization and safe exploration. [04:4...
Local Policies Enable Zero-shot Long Horizon Manipulation 02.11.2024 8:12
[00:00] Paper intro: Zero-shot robotic manipulation via local policies [00:26] Key challenges: Limited generalization and sim-to-real transfer [01:03] Local policies: Task decomposition through localized focus regions [01:38] Foundation models: VLMs for task understanding [02:07] Training approach: Simulation-based RL + visuomotor policy distillation [02:46] Implementation: Depth maps and impedanc...
MENTOR: Mixture-of-Experts Network with Task-Oriented Perturbation for Visual Reinforcement Learning 30.10.2024 8:52
[00:00] Introduction to Mentor system for visual RL [00:29] Problem: Sample inefficiency in robotic learning [00:59] Innovation: Mixture of Experts (MoE) architecture [01:55] Results: MoE achieves 100% success in multi-task testing [02:33] Feature: Task-oriented perturbation for exploration [03:55] Real-world testing: 83% success in robotic tasks [04:33] Study: MoE and perturbation each boost perf...
SkillMimicGen: Automated Demonstration Generation for Efficient Skill Learning and Deployment 29.10.2024 9:11
[00:00] SkillGen: AI system for robotic learning and automation. [00:19] Core Innovation: Automated dataset generation from minimal human input. [01:12] Skill Segmentation: Smart system for breaking down and adapting complex tasks. [01:59] Hybrid Skill Policy: Framework for controlling robot actions and task completion. [02:50] Performance Results: 75.4% success rate, generating 24,000+ demonstrat...
LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias 28.10.2024 7:13
[00:00] Intro to LVSM: Novel transformer for view synthesis [00:14] Problems with existing 3D synthesis methods [00:59] LVSM architecture: encoder-decoder vs decoder-only [01:41] Performance trade-offs between architectures [02:13] Using Pluecker rays for implicit 3D geometry [02:49] Zero-shot capabilities with varying input views [03:23] Training stability and technical solutions [03:59] Training...
Dynamic 3D Gaussian Tracking for Graph-Based Neural Dynamics Modeling 27.10.2024 15:31
[00:00] Introduction to 3D Gaussian tracking for robotic manipulation [00:26] Limitations of current video prediction methods [01:11] Advantages of 3D Gaussian representation [02:04] Graph Neural Networks for modeling object dynamics [02:54] Control particle implementation and computation reduction [03:42] Physics-based optimization for prediction stability [04:25] Integration with real-world robo...
SPIRE: Synergistic Planning, Imitation, and Reinforcement Learning for Long-Horizon Manipulation 25.10.2024 7:46
[00:00] Introduction [00:20] Core limitations in robot manipulation: challenges with RL and IL [01:08] SPIRE's hybrid approach: combining task planning with learning methods [01:44] TAMP-gated learning: selective application of learned policies [02:20] Training innovations: warm-starting RL and KL-divergence implementation [02:59] Results: 35-50% performance gain, 6x more data efficient [04:04] Mu...
VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation 24.10.2024 10:26
[00:00] VILA-U: A unified visual AI model [00:29] Problem: Inefficiency of separate visual modules [01:11] Vision tower: Novel quantization approach [02:09] Training strategy: CLIP-based staged learning [03:03] RVQ technique: Enhanced visual representation [03:47] Multi-modal training: Text-image-video fusion [04:35] Performance: Results and current limitations [05:23] Impact: Contrastive loss eff...
CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object Dynamics 23.10.2024 5:28
[00:00] Intro [00:31] Challenge: Limited multi-humanoid training data [00:55] CooHOI's two-phase learning framework [01:49] Object dynamics as implicit agent communication [02:25] Bounding box strategy for long objects [03:07] Results: Superior performance vs baselines [03:42] Ablation study findings: Key system components [04:18] Limitation: Basic hand manipulation only [04:57] Impact: New approa...
SynFlowNet: Design of Diverse and Novel Molecules with Synthesis Constraints 22.10.2024 7:02
[00:00] Introduction to SynFlowNet [00:29] Problem: AI-generated molecules often can't be synthesized [01:17] Solution: SynFlowNet - uses real chemical reactions [02:03] GFlowNets: Enables diverse molecule generation [02:47] Scalability: Morgan fingerprints handle 200K+ compounds [03:14] Challenge: Solving backward trajectory issues [04:14] Results: Better synthesis rates and molecular diversity [...
L3DG: Latent 3D Gaussian Diffusion 21.10.2024 18:18
[00:00] Intro to L3DG for 3D modeling [00:32] Solving room-sized 3D scene complexity [01:36] VQ-VAE compresses 3D Gaussian representation [02:41] Generative sparse transpose convolution [03:20] Latent diffusion for scene generation [04:30] Visual improvements over baselines [05:14] Scalability challenges for room-sized scenes [06:13] Spherical harmonics for view dependence [06:58] RGB and perceptu...
The Ingredients for Robotic Diffusion Transformers 20.10.2024 7:53
[00:00] Intro [00:33] Combining transformers & diffusion models [01:12] Key design: Scalable attention blocks (AdaLN) [02:30] Efficient observation tokenization [03:45] DiT Block policy architecture overview [04:20] BiPlay dataset introduction [04:53] Performance improvements over baselines [05:30] Key findings from ablations [06:08] Generalization to different robot types [06:43] Simulation v...
Estimating Body and Hand Motion in an Ego-sensed World 19.10.2024 15:56
[00:00] Introduction to EgoAllo system [00:38] Challenges in egocentric motion estimation [01:20] Importance of spatial/temporal invariance [02:11] Comparison of conditioning parameterizations [02:57] Integration of hand observations [03:50] Global alignment phase [04:28] Guidance losses in sampling [05:03] Handling longer sequences [05:35] Evaluation results [06:30] System limitations and future...
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation 19.10.2024 9:21
[00:00] Intro [00:28] Limitation of existing unified models [00:57] Janus's decoupled visual encoding solution [01:18] Advantages of decoupling [02:03] Janus architecture [02:50] Three-stage training [03:41] Ablation studies [04:23] Extensions for Janus [05:10] Performance gains [05:47] Current limitations [06:31] Impact of simplicity and extensibility [07:10] Qualitative results [08:18] Potential...
One Step Diffusion via Shortcut Models 19.10.2024 5:26
[00:00] Introduction [00:23] Computational cost of traditional diffusion models [00:59] Reducing iterations in image generation [01:06] Shortcut models [01:39] Training process and self-consistency property [02:22] Advantages over other methods [03:05] Results on image generation benchmarks [03:45] Application to robotic control [04:15] Limitations and future work [04:54] Best practices Authors: K...
Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning 18.10.2024 7:58
[00:00] Intro: Research on multi-task learning in LLMs [00:38] Balancing safety and performance in multilingual settings [01:17] Model merging techniques explored [02:16] Model merging outperforms data mixing [02:49] Merging monolingual models improves multilingual capabilities [03:28] Key ablation studies [04:14] Safety and performance evaluation metrics [04:52] Effectiveness variations across la...
CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos 18.10.2024 5:07
[00:00] Intro: CoTracker 3 paper [00:17] Challenge: Point tracker training [00:51] Existing methods' limitations [01:05] CoTracker 3 innovations [01:41] Novel semi-supervised training [02:24] Online vs offline versions [03:02] Performance improvements [03:47] Ablation study insights [04:29] Limitations and future work Authors: Nikita Karaev, Iurii Makarov, Jianyuan Wang, Natalia Neverova, Andrea V...
nGPT: Normalized Transformer with Representation Learning on the Hypersphere 18.10.2024 6:44
[00:00] Introduction [00:30] Consistent unit norm normalization in NGPT [01:08] Mathematical mechanism behind faster convergence [01:52] Elimination of weight decay in NGPT [02:21] Role of learnable eigen learning rates in optimization [03:04] Discussion on training speedup vs. per-step computation time [03:46] Condition number differences between GPT and NGPT [04:18] Ablation studies on scaling f...
RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation 18.10.2024 9:57
[00:00] Intro [00:20] Challenges in bimanual manipulation models [01:00] Robotics Diffusion Transformer (RDT) approach [01:34] RDT architecture design [02:22] Data scarcity and unified action space [03:07] Multi-task bimanual dataset [03:49] RDT's experimental results [04:30] Benefits of large-scale pre-training [05:12] Diffusion models in robotics [06:02] RDT's architectural modifications [06:57]...
Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models 18.10.2024 9:02
[00:00] Continuous-time consistency models (CTMs) [00:20] Limitations of existing CTMs [00:51] TrigFlow: New CTM formulation [01:42] CTM training instability [02:21] Training objective modifications [02:55] Scaling CTMs to 1.5 billion parameters [03:37] Comparison with state-of-the-art models [04:14] Consistency training vs. distillation [04:52] CTMs vs. variational score distillation [05:20] Key...
Podcasts similares
Replaio no es editor de podcasts; los nombres de los programas, las portadas y el audio pertenecen a sus autores y se distribuyen a través de canales RSS públicos