AI Evals Done Right: From Vibes to Confident Decisions
The provided transcript contains no technical content, data, or descriptions of methodologies regarding AI evaluations. It consists entirely of repeated expressions of gratitude and does not address a specific problem, approach, or set of key takeaways. Consequently, there is no subject matter available to summarize.
This description was generated by Open-Source AI using the transcript of the session and the original submission contents.
This session took place in track Generative AI & Synthetic Data and was classified suitable for intermediate domain by the speaker.
Submission
The proposal as submitted by the speaker before the conference.
Testing traditional software is "simple"... same input, same output. LLMs? Not so much. Same prompt, different result every time. So how do you actually know if your AI product is good?
Spoiler: Most teams don't. They ship on vibes and hope for the best.
This talk takes you through our real journey at Blue Yonder, where we built an LLM-powered analytics system and needed a way to actually measure its quality. You'll see how we went from "feels okay-ish" to concrete numbers that let us make real decisions - with actual examples from production along the way.
The methodology is called Error Analysis: collect traces, annotate them from the user's perspective, group similar issues into failure modes, and turn those into automated evals. Along the way, we'll share practical best practices like why binary Pass/Fail beats rating scales, and why 100% pass rate means your evals are broken.
The payoff? When a new model drops, we run our pipeline and know within hours - not weeks - whether it's better or worse for our specific use case. Real percentages. Real trade-offs. Real decisions.
Expect a meme-powered walkthrough and a clear path to implement this yourself starting with just 20 traces.
Outline:
- Introduction: The challenge of testing stochastic systems, why we needed a better approach
- Collecting and Annotating Traces: Every trace is a user experiencing your product, Open Coding from the user perspective, real examples of failure modes we discovered
- Building the Failure Taxonomy: Grouping observations into categories, Axial Coding, turning scattered comments into actionable failure modes
- Writing Evals That Work: LLM-as-judge setup, binary scores vs rating scales, validating against human judgment
- From Vibes to Decisions: Prioritizing what to fix, measuring improvement, 24-hour model benchmarking
- Wrap-up: Your action plan, start with 20 traces
Transcript (auto)
Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.
Speaker 1 [00:00]
Thank you. you you you you you you you you you you you you you Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you very much. Thank you. Thank you. Thank you. Thank you.