Pawan K Jha
Architecting Intelligence
AI systems, LLM engineering, ML infrastructure, and production AI architecture.
N'hésitez pas à visiter le site du podcast et à soutenir son créateur : podcasters.spotify.com
Auteur
Pawan K Jha
Catégorie
Site du podcast
Dernier épisode
28 juin 2026
Où écouter ?
Les podcasts dans l'appli Replaio Radio Bientôt disponibleLes podcasts arrivent très bientôt dans l'appli. Installe-la dès maintenant et découvre en avant-première une toute nouvelle façon de vivre les podcasts
Épisodes
Self-Attention, QKV, and the Foundation of KV Cache | Architecting LLM Inference — Part 2A 28.06.2026 16:29
Before you can understand the KV cache, you need to understand how LLMs build context. That's what this episode is about. In Part 2A of Architecting LLM Inference, we build the foundation for everything that follows — starting with what context means inside a language model, and ending with exactly why the KV cache exists. What we cover: What context and contextual representation mean in practice...
Flashcard #001: LLM Inference Lifecycle 04.06.2026 0:54
Part of the Architecting Intelligence Flashcards series — short visual notes on AI, LLM systems, and ML infrastructure.
Podcasts similaires
Replaio n'est pas éditeur de podcasts ; les noms des émissions, les visuels et l'audio appartiennent à leurs auteurs et sont diffusés via des flux RSS publics