Pawan K Jha

Architecting Intelligence

Technology EN ↓ 2 episodes

AI systems, LLM engineering, ML infrastructure, and production AI architecture.

Be sure to visit the podcast's website and support the creator: podcasters.spotify.com

Author

Pawan K Jha

Category

Technology

Podcast website

podcasters.spotify.com

Latest episode

Jun 28, 2026

Where to listen?

Podcasts in the app Replaio Radio Coming soon

Podcasts are coming to the app soon. Install now and be the first to see a whole new take on podcasts

Get it on Google Play Install for free Android almost 10M downloads · 4.8 rating iOS soon

Episodes

Self-Attention, QKV, and the Foundation of KV Cache | Architecting LLM Inference — Part 2A 28.06.2026

Before you can understand the KV cache, you need to understand how LLMs build context. That's what this episode is about. In Part 2A of Architecting LLM Inference, we build the foundation for everything that follows — starting with what context means inside a language model, and ending with exactly why the KV cache exists. What we cover: What context and contextual representation mean in practice...

Flashcard #001: LLM Inference Lifecycle 04.06.2026

Part of the Architecting Intelligence Flashcards series — short visual notes on AI, LLM systems, and ML infrastructure.

Listen to the Architecting Intelligence podcast in Replaio

Radio and podcasts in one app - free, with no sign-up. Install today and do not miss the launch

Get it on Google Play

Replaio is not a podcast publisher; show names, artwork and audio belong to their authors and are distributed through public RSS feeds.