We're sunsetting PodQuest on 2025-07-28. Thank you for your support!
Export Podcast Subscriptions
cover of episode Product Metrics are LLM Evals // Raza Habib CEO of Humanloop // #320

Product Metrics are LLM Evals // Raza Habib CEO of Humanloop // #320

2025/6/3
logo of podcast MLOps.community

MLOps.community

AI Deep Dive AI Chapters Transcript
People
R
Raza Habib
Topics
Raza Habib: 我认为产品指标和LLM评估之间没有真正的区别,或者说,最好的评估与产品指标是相同的。评估的目的是衡量AI系统的质量,而生成式AI的出现使得衡量性能变得更加主观。在某些情况下,最好的评估是最终用户的体验,即用户是否达到了他们想要的结果。评估是为了在开发过程中提供替代方案,因为在产品投入生产之前,不一定能获得最终用户体验。我们使用LLM创建的许多评估,如LLM作为评判者或自动化评估,实际上也在作为生产指标运行。如果在与AI互动后完成了某个步骤,那就是一次成功,并且可以很容易地判断是否采取了最佳行动。

Deep Dive

Shownotes Transcript

Raza Habib, the CEO of LLM Eval platform Humanloop), talks to us about how to make your AI products more accurate and reliable by shortening the feedback loop of your evals. Quickly iterating on prompts and testing what works, along with some of his favorite Dario from Anthropic AI) Quotes.

// Bio

Raza is the CEO and Co-founder at Humanloop. He has a PhD in Machine Learning from UCL, was the founding engineer of Monolith AI, and has built speech systems at Google. For the last 4 years, he has led Humanloop and supported leading technology companies such as Duolingo, Vanta, and Gusto to build products with large language models. Raza was featured in the Forbes 30 Under 30 technology list in 2022, and Sifted recently named him one of the most influential Gen AI founders in Europe.

// Related Links

Websites: https://humanloop.com


Catch all episodes, blogs, newsletters, and more: https://go.mlops.community/TYExplore

MLOps Swag/Merch: [https://shop.mlops.community/]

Connect with Demetrios on LinkedIn: /dpbrinkm

Connect with Raza on LinkedIn: /humanloop-raza



Timestamps:



[00:00] Cracking Open System Failures and How We Fix Them

[05:44] LLMs in the Wild — First Steps and Growing Pains

[08:28] Building the Backbone of Tracing and Observability

[13:02] Tuning the Dials for Peak Model Performance

[13:51] From Growing Pains to Glowing Gains in AI Systems

[17:26] Where Prompts Meet Psychology and Code

[22:40] Why Data Experts Deserve a Seat at the Table

[24:59] Humanloop and the Art of Configuration Taming

[28:23] What Actually Matters in Customer-Facing AI

[33:43] Starting Fresh with Private Models That Deliver

[34:58] How LLM Agents Are Changing the Way We Talk

[39:23] The Secret Lives of Prompts Inside Frameworks

[42:58] Streaming Showdowns — Creativity vs. Convenience

[46:26] Meet Our Auto-Tuning AI Prototype

[49:25] Building the Blueprint for Smarter AI

[51:24] Feedback Isn’t Optional — It’s Everything