Snorkel AI

Benchtalks

Science EN ↓ Επεισόδια: 3

Benchtalks is Snorkel AI's podcast series at the intersection of AI evaluation, data quality, and real-world impact. Hosted by the Snorkel team, each episode brings together researchers, practitioners, and leaders to dig into the questions that matter most as AI benchmarks grow more sophisticated, dynamic, and reflective of the complexity found in real-world deployments. We explore the full stack of what it takes to build AI that actually works — from the design of rigorous, open benchmarks that close the gap between what we measure and what we encounter in production, to the expert-in-the-loo...

Μην παραλείψεις να επισκεφτείς τη σελίδα του podcast και να στηρίξεις τον δημιουργό: www.buzzsprout.com

Δημιουργός

Snorkel AI

Κατηγορία

Science

Ιστοσελίδα του podcast

www.buzzsprout.com

Τελευταίο επεισόδιο

25 Ιουν 2026

Πού να ακούσεις;

Podcast στην εφαρμογή Replaio Radio Έρχεται σύντομα

Τα podcast έρχονται σύντομα στην εφαρμογή. Εγκατάστησέ την τώρα και δες πρώτος μια εντελώς νέα προσέγγιση στα podcast

Κατέβασέ το από το Google Play Δωρεάν εγκατάσταση Android σχεδόν 10 εκατ. λήψεις · βαθμολογία 4,8 iOS σύντομα

Επεισόδια

Benchtalks #3: Parth Asawa (Continual Learning Bench) - We Taught AI Everything Except How to Learn 25.06.2026

For our third Benchtalks, the series dedicated to the researchers building the measurement toolkits that frontier labs hill-climb on, Snorkel AI co-founder Vincent Sunn Chen sat down with Parth Asawa , a UC Berkeley PhD student advised by Matei Zaharia and Joey Gonzalez , and the creator of Continual Learning Bench , the first standardized benchmark for measuring whether AI systems actually learn...

Benchtalks #2: John Yang (SWE-bench, ProgramBench) - The future of coding benchmarks 03.06.2026

For our second Benchtalks , the series dedicated to the researchers building the measurement toolkits that frontier labs hill-climb on, Snorkel AI co-founder Vincent Sunn Chen sat down with John Yang , a Stanford PhD student and creator of the SWE-bench franchise, SWE-smith, CodeClash, and most recently ProgramBench . This interview covers:  Why every frontier model scored 0% at launch — until GPT...

Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) - Building the benchmark factory 31.03.2026

In this inaugural episode of Benchtalks, Snorkel AI co-founder Vincent Chen sits down with Alex Shaw, MTS at Laude Institute and co-creator of Terminal-Bench, to unpack what the rapid hill-climbing on TB2 reveals about the state of AI agent evaluation — and where the field needs to go. This interview covers:  Why TB2 went from 20–30% during development to 75–80% at the frontier today The bet on th...

Άκου το podcast Benchtalks στο Replaio

Ραδιόφωνο και podcast σε μία εφαρμογή - δωρεάν, χωρίς εγγραφή. Εγκατάστησέ την σήμερα και μη χάσεις την πρεμιέρα

Κατέβασέ το από το Google Play

Το Replaio δεν είναι εκδότης podcast - τα ονόματα των εκπομπών, τα εξώφυλλα και ο ήχος ανήκουν στους δημιουργούς τους και διανέμονται μέσω δημόσιων ροών RSS