weedge
AI Podcast
Latest podcasts about AI Technology and Papers.
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
AI技术前沿:Phi-4 大型语言模型的突破 09.01.2025 8:08
深入探讨微软最新发布的Phi-4大型语言模型,了解其在数据质量、合成数据、训练方法和后训练优化方面的创新。 我们将分析其在各种基准测试上的表现,以及它如何超越以往的Phi系列模型,甚至在某些领域超越更大的模型。
StyleTTS 2: Towards Human-Level Text-to-Speech 09.01.2025 6:16
A podcast discussion about the StyleTTS 2 model for text-to-speech synthesis, focusing on its innovative use of style diffusion and adversarial training with large speech language models to achieve human-level performance.
WavChat:语音对话模型调查 09.01.2025 13:40
本播客深入探讨了语音对话模型的最新进展,包括其功能、表示形式、训练范式以及流媒体和交互能力。
宇宙世界基础模型平台:物理人工智能的未来 07.01.2025 6:26
深入探讨NVIDIA Cosmos世界基础模型平台,该平台旨在促进物理人工智能的发展,通过数字孪生和世界模型,加速人工智能在现实世界中的应用。
AI Radio FM - Technology Channel, Your Personal Generative AI Podcast 07.01.2025 4:43
A podcast discussing the Story-Adapter framework for long story visualization.
AI科技前沿:故事扩散模型深度解析 07.01.2025 9:18
本期播客深入探讨故事扩散模型,一种用于生成连贯图像和视频的新方法。我们将详细分析其核心技术,包括一致性自注意力机制和语义运动预测器,并讨论其在视觉故事生成方面的应用和潜力。
智能格林:基于潜在扩散模型的开放式视觉故事讲述 07.01.2025 7:16
本期播客讨论了一篇关于使用潜在扩散模型进行开放式视觉故事讲述的论文。我们深入探讨了该模型的技术细节、数据集构建以及实验结果,展示了其在生成连贯图像序列方面的卓越能力。
AI Radio FM - Technology Channel 07.01.2025 7:54
A podcast discussing the IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
PowerInfer-2: Fast Large Language Model Inference on a Smartphone 07.01.2025 12:03
A podcast discussion about PowerInfer-2, a framework for running large language models on smartphones, focusing on its neuron cluster design, adaptive computation strategies, and I/O optimizations.
LatentSync: Audio Conditioned Latent Diffusion Models for Lip Sync 06.01.2025 9:19
A deep dive into LatentSync, an innovative lip-sync framework using audio-conditioned latent diffusion models, its methodology, experiments, and the resolution of SyncNet convergence issues.
Infinity: Scaling Bitwise Autoregressive Modeling for High-Resolution Image Synthesis 06.01.2025 6:12
A podcast discussing the groundbreaking research paper on Infinity, a novel autoregressive model for high-resolution image synthesis.
AI驱动的交互式头部生成 06.01.2025 5:30
本期播客深入探讨了INFP,一个用于双人对话的音频驱动的头部生成框架。我们将探讨其创新方法、数据集以及实验结果,展示其在自然人机交互中的潜力。
CosyVoice 2: 使用大型语言模型实现可扩展的流式语音合成 06.01.2025 8:01
一个关于 CosyVoice 2 的播客,这是一个改进的流式语音合成模型,它利用大型语言模型,实现了接近人类水平的自然度,最小的响应延迟,以及在流模式下几乎无损的合成质量。
Flow Matching for Generative Modeling 06.01.2025 6:03
A podcast discussing the new paradigm for generative modeling using Continuous Normalizing Flows (CNFs) called Flow Matching (FM). FM offers a simulation-free approach for training CNFs by regressing vector fields of fixed conditional probability paths, which enables training CNFs at unprecedented scale and allows for the use of different probability paths.
Swin Transformer: A New Vision Transformer 05.01.2025 7:42
A podcast discussing the Swin Transformer, a hierarchical vision transformer using shifted windows for computer vision tasks.
ConvNeXt: A Modern ConvNet for the 2020s 05.01.2025 6:28
A podcast discussing the architecture and performance of ConvNeXt, a modern ConvNet model that challenges the dominance of Vision Transformers.
AI Vision Podcast: Masked Autoencoders for Scalable Vision Learning 05.01.2025 5:25
A deep dive into Masked Autoencoders (MAE) and their impact on computer vision, discussing their architecture, training efficiency, and performance on ImageNet and downstream tasks.
AI Radio FM - Technology Channel, Your Personal Generative AI Podcast 04.01.2025 5:24
A podcast discussing the auxiliary-loss-free load balancing strategy for mixture-of-experts models.
混合专家模型(MoE)技术综述 04.01.2025 5:21
本播客深入探讨了混合专家模型(MoE)的最新进展、算法设计、系统实现以及实际应用。从稀疏和密集MoE的背景知识开始,我们提出了一个创新的MoE分类法,并探讨了选通函数、专家网络、训练方案和系统设计方面的复杂性,从而全面了解MoE。
零气泡流水线并行 04.01.2025 6:34
本期播客深入探讨了零气泡流水线并行技术,这是一种旨在提高大规模分布式训练效率的创新方法。我们分析了传统流水线并行方法中的气泡问题,并介绍了如何通过精细化调度和优化器同步绕过技术来实现零气泡。此外,我们还讨论了自动调度算法、内存优化策略以及实验结果,旨在为听众提供一个全面而深入的技术解析。
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding 04.01.2025 6:09
A podcast discussion about GShard, a module for scaling neural networks using conditional computation and automatic sharding, focusing on its application to multilingual machine translation.
AI Radio FM - Technology Channel: GShard and Giant Models 04.01.2025 8:38
A deep dive into GShard, a module for scaling giant neural networks, focusing on its application to multilingual machine translation and its impact on training efficiency and model quality.
混合张量专家数据并行方法优化混合专家训练 04.01.2025 5:19
深入探讨 DeepSpeed-TED,一种新颖的三维混合并行框架,用于训练具有大型基础模型的混合专家模型。我们讨论了内存优化、通信优化以及与现有方法的性能比较。
统一序列并行方法:为长上下文生成式AI赋能 04.01.2025 7:16
本播客深入探讨了统一序列并行(Unified Sequence Parallelism,简称USP)方法,这是一种用于训练具有极长上下文的生成式AI模型的先进技术。我们分析了现有的序列并行方法,如DeepSpeed-Ulysses和Ring-Attention,并提出了一个统一的框架,该框架结合了两者的优点,同时克服了它们的局限性。通过详细讨论,我们将深入了解USP如何与数据并行、张量并行、ZeRO和流水线并行等现有并行技术相结合,从而为4D混合并行系统提供最佳实践。...
LoongTrain: 高效长序列大语言模型训练 04.01.2025 7:32
本期播客深入探讨LoongTrain,一个为长序列大语言模型设计的高效训练框架。我们将讨论其核心的2D注意力机制,以及它如何结合头并行和上下文并行来克服扩展性限制并保持效率。此外,还将分析Double-Ring-Attention机制,以及设备放置策略对训练速度的影响。
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.