40 subscribers
Go offline with the Player FM app!
Podcasts Worth a Listen
SPONSORED


The Challenge of AI Model Evaluations with Ankur Goyal
Manage episode 487909891 series 2455731
Evaluations are critical for assessing the quality, performance, and effectiveness of software during development. Common evaluation methods include code reviews and automated testing, and can help identify bugs, ensure compliance with requirements, and measure software reliability.
However, evaluating LLMs presents unique challenges due to their complexity, versatility, and potential for unpredictable behavior.
Ankur Goyal is the CEO and Founder of Braintrust Data, which provides an end-to-end platform for AI application development, and has a focus on making LLM development robust and iterative. Ankur previously founded Impira which was acquired by Figma, and he later ran the AI team at Figma. Ankur joins the show to talk about Braintrust and the unique challenges of developing evaluations in a non-deterministic context.
Sean's been an academic, startup founder, and Googler. He has published works covering a wide range of topics from AI to quantum computing. Currently, Sean is an AI Entrepreneur in Residence at Confluent where he works on AI strategy and thought leadership. You can connect with Sean on LinkedIn.
Please click here to see the transcript of this episode.
Sponsorship inquiries: sponsor@softwareengineeringdaily.com
2108 episodes
Manage episode 487909891 series 2455731
Evaluations are critical for assessing the quality, performance, and effectiveness of software during development. Common evaluation methods include code reviews and automated testing, and can help identify bugs, ensure compliance with requirements, and measure software reliability.
However, evaluating LLMs presents unique challenges due to their complexity, versatility, and potential for unpredictable behavior.
Ankur Goyal is the CEO and Founder of Braintrust Data, which provides an end-to-end platform for AI application development, and has a focus on making LLM development robust and iterative. Ankur previously founded Impira which was acquired by Figma, and he later ran the AI team at Figma. Ankur joins the show to talk about Braintrust and the unique challenges of developing evaluations in a non-deterministic context.
Sean's been an academic, startup founder, and Googler. He has published works covering a wide range of topics from AI to quantum computing. Currently, Sean is an AI Entrepreneur in Residence at Confluent where he works on AI strategy and thought leadership. You can connect with Sean on LinkedIn.
Please click here to see the transcript of this episode.
Sponsorship inquiries: sponsor@softwareengineeringdaily.com
2108 episodes
All episodes
×

1 SED News: Data Land Grabs, Copyright Fights, and the Great AI Talent War 47:14


1 ByteDance’s Container Networking Stack with Chen Tang 47:57

1 WayForward Games with Tomm Hulett and Voldi Way 46:02

1 CodeRabbit and RAG for Code Review with Harjot Gill 48:42

1 Emulating Retro Games on Modern Consoles with Robin Lavallée and Bill Litshauer 1:01:34

1 SED News: Corporate Spies, Postgres, and the Weird Life of Devs Right Now 44:38

1 TanStack and the Future of Frontend with Tanner Linsley 55:13

1 The Challenge of AI Model Evaluations with Ankur Goyal 45:22

1 Modern Distributed Applications with Stephan Ewen 41:20


1 Chip Design in the AI Era with Thomas Andersen 50:33

1 OpenTofu with Cory O’Daniel and Malcolm Matalka 48:58

1 Mojo and Building a CUDA Replacement with Chris Lattner 56:14
Welcome to Player FM!
Player FM is scanning the web for high-quality podcasts for you to enjoy right now. It's the best podcast app and works on Android, iPhone, and the web. Signup to sync subscriptions across devices.