Making “this doesn’t sound right” testable
I found the failure modes, wrote the assertions and a response-quality rubric, and calibrated an LLM judge against human review.
I found the failure modes, wrote the assertions and a response-quality rubric, and calibrated an LLM judge against human review.
A live tool that puts the evals method to work: choose the standards, apply them, watch an LLM judge grade the lift.
A controlled vocabulary and in-product messaging, written for people and encoded as an agent-readable skill.
UI and labels that frame AI insights and recommendations, with assertions to measure the output.
A more relevant, engaging push-notification system, grounded in user research and shipped as updates.
Reading user intent, naming a tradeoff, and recovering from a bad guess.
Iterating a system prompt and updating a Python script to lift output quality.
My case studies are password-protected. If you’re a recruiter or hiring manager, reach out and I’ll walk you through them.