One vocabulary, three surfaces
Controlled vocabulary and terminology, written for people and encoded as an agent-readable skill.
I do content and model design. I work upstream, where decisions get made before anyone writes a word. I design the pattern the AI follows: the materials, the needle, the stitch count, what the finished thing should look like.
Your first scarf never comes out right. That’s okay. Next time you adjust the tension. You create a gauge swatch and check it as you work. I measure AI outputs to figure out what needs adjusting. And I’m learning to knit, so yarn is on my mind.
Principal Content Designer · Content systems and model design · 11 years in UX · Aspiring knitter
Content models and taxonomies: families of words and phrases that hang together, so AI writing feels coherent and on brand.
Repeatable workflows that show AI how to handle key scenarios, and help design and product partners make tough calls.
Review outputs, measure performance, direct the LLM judge, and write the assertions and rubrics that set the standard.
Built evals for AI-generated content, wrote UX copy for AI-powered search, and shaped a more effective push-notification framework.
Rider education and trust-building for an autonomous taxi service.
Led content strategy and writing for clients including Amazon, Philz Coffee, and Rocket Mortgage.
Onboarding and growth experiments for Dropbox Paper and integrations like Dropbox for Slack. Mentored and managed writers.
Stop AI slop
I build the skills and rubrics that decide which of these a model ships.
The prompt · same for all three
“This is my first scarf and I’m about 12 rows in. There’s a hole a couple rows down and I can’t tell if I did something wrong. Should I pull it all out and start over?”
“Knitting is such a rewarding journey, and every project teaches you something new! Your scarf is coming along beautifully — small imperfections are what make handmade things special.”
Mock dataFails all three. It has never seen your knitting.
“Dropped stitches are common for beginners and they’re usually fixable. Most knitters can pick them back up without unraveling the whole project.”
Mock dataReads better. Still won’t say what it doesn’t know.
“That looks like a dropped stitch two rows down, and you don’t need to start over. I can’t see your tension from here — pick it up with a crochet hook and check whether the row still lies flat.”
Mock dataPasses. The shortest version that still answers the question.
Made up content. A real method. This is a tiny version of the tests I run on content skills: write the assertions, pick cases that stress them, measure what changes.
Models write clean, confident sentences all day long. Fluency isn’t the issue. What’s missing is usually a standard AI can apply. At scale a confidently wrong sentence costs brand trust the same way a factual error does.
The fix? Decide what good looks and feels like, define specific criteria to get there, study outputs, track failure modes. Repeat.
Controlled vocabulary and terminology, written for people and encoded as an agent-readable skill.
Assertions, a response-quality rubric, and calibrating an LLM judge against human review.
A provenance model separating what a user said from what the system inferred.
AI output that’s almost right, and nobody can put their finger on what’s wrong. That’s my favorite place to start.