Hamel Husain
13 stories
Hamel Husain
Hamel HusainVoicesAI Product Engineering Notes Notes from 13 sessions on evals, context, and systems. 9.5 hours of talks distilled into about 20 minutes of reading.

Hamel HusainVoices“It’s Hard to Eval” Is a Product Smell For the past 3 years, AI evals have been my professional focus. 1 The most common objection I hear to evals is “our product is hard to eval”. This objection is a product smell. Artifacts that are hard for you to verify a

Hamel HusainVoicesThe Revenge of the Data Scientist Is the heyday of the data scientist over? The Harvard Business Review once called it “The Sexiest Job of the 21st Century.” 1 In tech, data scientist roles were often among the best paid. 2 The job also demanded an unusu

Hamel HusainVoicesEvals Skills for Coding Agents Today, Shreya Shankar and I are publishing evals skills , a set of skills for AI product evals 1 . Eval tools often get in the way. They nudge you toward generic off-the-shelf metrics and fully automated evals before you

Hamel HusainVoicesWhy I Stopped Using nbdev Programmers love to proclaim they’ve found the best tool. Paul Graham called Lisp his “ secret weapon .” DHH described Ruby as “ a magical glove that just fit my brain perfectly .” Pieter Levels ships million-dollar prod

Hamel HusainVoicesSelecting The Right AI Evals Tool Over the past year, I’ve focused heavily on AI Evals , both in my consulting work and teaching. A question I get constantly is, “What’s the best tool for evals?”. I’ve always resisted answering directly for two reasons.

Hamel HusainVoicesAI Evals: Everything You Need to Know This document curates the most common questions Shreya and I received while teaching 700+ engineers & PMs AI Evals. Warning: These are sharp opinions about what works in most cases. They are not universal truths. Use you

Hamel HusainVoicesA Field Guide to Rapidly Improving AI Products Most AI teams focus on the wrong things. Here’s a common scene from my consulting work: AI TEAM Here’s our agent architecture – we’ve got RAG here, a router there, and we’re using this new framework for… ME [Holding up m

Hamel HusainVoicesBuilding an Audience Through Technical Writing: Strategies and Mistakes People often find me through my writing on AI and tech. This creates an interesting pattern. Nearly every week, vendors reach out asking me to write about their products. While I appreciate their interest and love learni

Hamel HusainVoicesUsing LLM-as-a-Judge For Evaluation: A Complete Guide Earlier this year, I wrote Your AI product needs evals . Many of you asked, “How do I get started with LLM-as-a-judge?” This guide shares what I’ve learned after helping over 30 companies set up their evaluation systems.

Hamel HusainVoicesAn Open Course on LLMs, Led by Practitioners Today, we are releasing Mastering LLMs , a set of workshops and talks from practitioners on topics like evals, retrieval-augmented-generation (RAG), fine-tuning and more. This course is unique because it is: Taught by 25

Hamel HusainVoicesDebugging AI With Adversarial Validation For years, I’ve relied on a straightforward method to identify sudden changes in model inputs or training data, known as “drift.” This method, Adversarial Validation 1 , is both simple and effective. The best part? It re

Hamel HusainVoicesYour AI Product Needs Evals Motivation I started working with language models five years ago when I led the team that created CodeSearchNet , a precursor to GitHub CoPilot. Since then, I’ve seen many successful and unsuccessful approaches to buildi
Nothing matches this filter yet.