Pawan K Jha

Architecting Intelligence

AI systems, LLM engineering, ML infrastructure, and production AI architecture.

Não deixe de visitar o site do podcast e apoiar quem o produz: podcasters.spotify.com

Autor

Pawan K Jha

Categoria

Technology

Site do podcast

podcasters.spotify.com

Último episódio

28 de jun de 2026

Onde ouvir?

Podcasts no app Replaio Radio Em breve

Os podcasts estão chegando ao app. Instale agora e seja o primeiro a descobrir um jeito totalmente novo de curtir podcasts

Baixe no Google Play Instale grátis Android 5 mi+ downloads · nota 4,8 iOS em breve

Episódios

Self-Attention, QKV, and the Foundation of KV Cache | Architecting LLM Inference — Part 2A 28.06.2026

Before you can understand the KV cache, you need to understand how LLMs build context. That's what this episode is about. In Part 2A of Architecting LLM Inference, we build the foundation for everything that follows — starting with what context means inside a language model, and ending with exactly why the KV cache exists. What we cover: What context and contextual representation mean in practice...

Flashcard #001: LLM Inference Lifecycle 04.06.2026

Part of the Architecting Intelligence Flashcards series — short visual notes on AI, LLM systems, and ML infrastructure.

Ouça o podcast Architecting Intelligence no Replaio

Rádio e podcasts em um só app - grátis e sem cadastro. Instale hoje e não perca o lançamento

Baixe no Google Play

O Replaio não é o publicador dos podcasts; os nomes dos programas, as capas e o áudio pertencem aos seus autores e são distribuídos por feeds RSS públicos