weedge
AI Podcast
Latest podcasts about AI Technology and Papers.
Where to listen?
Podcasts in the app Replaio Radio Coming soonPodcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts
Episodes
Ring Attention with Blockwise Transformers for Near-Infinite Context 04.01.2025 6:21
A podcast discussing a novel approach to scale transformer models to handle near-infinite context lengths.
FlashAttention-3: Revolutionizing Attention Mechanisms on GPUs 04.01.2025 4:35
A podcast discussing the FlashAttention-3 algorithm, its improvements over previous versions, and its impact on large language models.
AI FlashAttention-2 Podcast 04.01.2025 6:17
A fast-paced discussion on FlashAttention-2, a faster attention mechanism for Transformers, exploring its algorithms, parallelism, and performance benefits.
FlashAttention: 高效且内存优化的精确注意力机制 04.01.2025 7:05
探讨 FlashAttention 算法,一种在 GPU 上实现快速、内存高效精确注意力机制的新方法。深入分析其 IO 复杂度,并与现有的注意力机制进行性能比较。
DeepSpeed Ulysses: 极端长序列Transformer模型训练的系统优化 04.01.2025 7:17
本播客深入探讨了DeepSpeed Ulysses,一种用于训练具有极长序列长度的Transformer模型的创新方法,它通过优化序列并行性和通信效率,显著提升了训练速度和可扩展性。我们将讨论其核心设计、通信分析、内存效率以及与现有方法的比较。
DistFlashAttn: 分布式长文本大语言模型训练的内存高效注意力机制 04.01.2025 6:53
本播客深入探讨 DistFlashAttn,一种专为长文本大语言模型训练设计的分布式内存高效注意力机制,详细解析其核心技术和性能优势。
大型Transformer模型中减少激活重计算 04.01.2025 6:05
本播客讨论了一种加速大型Transformer模型训练的新方法,通过减少激活重计算来实现。我们将深入探讨序列并行和选择性激活重计算技术。
序列并行:从系统角度进行长序列训练 04.01.2025 5:23
探讨一种名为“序列并行”的内存高效并行方法,该方法旨在突破输入序列长度的限制,并能在GPU上高效训练更长的序列。该方法与现有的并行技术兼容,并能实现4D并行。核心思想是将输入序列分割成多个块,并分配给不同的GPU进行处理。为了计算注意力输出,引入了环形自注意力机制。
AI驱动的大规模语言模型训练:Megatron-LM在GPU集群上的高效实践 04.01.2025 7:11
本期播客深入探讨了如何使用Megatron-LM在GPU集群上高效训练大规模语言模型,重点关注张量并行、流水线并行和数据并行的组合应用,以及创新的交错流水线调度方法。
AI Radio FM - Technology Channel: PagedAttention for Large Language Model Serving 04.01.2025 7:31
A podcast discussing PagedAttention, a novel memory management technique for serving large language models, and its implementation in vLLM.
ORCA: 分布式Transformer生成模型服务系统 04.01.2025 5:40
本期播客深入探讨了ORCA,一个为Transformer模型设计的分布式服务系统。我们将详细介绍其创新的迭代级调度和选择性批处理技术,以及它们如何显著提升模型服务的性能。
LLM推理优化:连续批处理实现23倍吞吐量提升 04.01.2025 6:04
本期播客深入探讨了大型语言模型(LLM)推理中的连续批处理技术,揭示了其如何显著提高吞吐量并降低延迟。我们将讨论传统批处理的局限性,并详细介绍连续批处理的原理及其在实际应用中的优势,尤其是在使用vLLM时的卓越性能表现。
Mooncake:一种以KVCache为中心的LLM服务解耦架构 04.01.2025 5:33
本播客深入探讨Mooncake的创新架构,这是一种专为高效服务大型语言模型而设计的解耦系统。
AI Radio FM - Technology Channel, Your Personal Generative AI Podcast 02.01.2025 7:32
A podcast discussing the InternLM-XComposer2 model, its architecture, and capabilities in free-form text-image composition and comprehension.
AI Radio FM - Technology Channel, Your Personal Generative AI Podcast 02.01.2025 6:28
A fast-paced, enthusiastic podcast discussing the latest advancements in AI, focusing on the InternLM-XComposer2-4KHD model.
AI Radio FM - Technology Channel 02.01.2025 6:05
A podcast discussing InternLM-XComposer-2.5, a versatile large vision language model.
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System 02.01.2025 6:28
A podcast discussing the InternLM-XComposer2.5-OmniLive system, a novel multimodal AI for long-term video and audio interaction.
DeepSeekMoE: 超越专家混合模型的终极专业化 28.12.2024 6:40
本期播客深入探讨了DeepSeekMoE这一创新的混合专家模型架构,旨在实现专家知识的终极专业化。我们将讨论其核心策略、实验验证以及与现有模型的对比,揭示其在大型语言模型领域的优势。
DeepSeek-V3: A Deep Dive into a Powerful Mixture-of-Experts Model 27.12.2024 5:36
A podcast discussion analyzing the DeepSeek-V3 technical report, covering its architecture, training, and performance.
E2 TTS: 令人惊讶的简单零样本文本到语音技术 27.12.2024 7:30
本期节目深入探讨了E2 TTS,一种完全非自回归的零样本文本到语音系统,它在自然度、说话人相似度和可懂度方面都达到了最先进的水平。我们将详细讨论其训练过程、推理方法以及如何通过其扩展来提升用户体验。
BigVGAN: 通用神经声码器大规模训练 27.12.2024 5:03
本播客讨论了BigVGAN,一种通用的神经声码器,它通过大规模训练实现高保真音频合成,并在各种分布外场景中表现出色。
F5-TTS: 突破性文本到语音技术 27.12.2024 5:42
深入探讨 F5-TTS,一种基于流匹配的非自回归文本到语音系统,该系统在零样本语音合成方面表现出色。
深入浅出:注意力机制的演变与应用 25.12.2024 7:04
本期播客将深入探讨注意力机制在深度学习领域的演变与应用,从Seq2Seq模型的局限性到Transformer的创新,再到Self-Attention GAN的强大功能,我们将一步步揭开注意力机制的神秘面纱,带您领略其在自然语言处理、计算机视觉等领域的卓越表现。
Speech and Language Processing 24.12.2024 10:06
A podcast discussing the content from Daniel Jurafsky and James H. Martin's "Speech and Language Processing" textbook, Third Edition draft, specifically focusing on fundamental algorithms for NLP, NLP applications, and annotating linguistic structure.
从慢速双向到快速因果视频生成器 23.12.2024 6:02
本播客讨论了一种新的视频生成方法,该方法通过将预训练的双向扩散模型转化为因果模型,并结合分布匹配蒸馏技术,实现了快速、高质量的视频生成。该方法支持流式视频生成、视频到视频的转换、图像到视频的生成以及动态提示。
Similar podcasts
Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.