The Archive · 2013 — Now
Product archive.
6 posts in this category
6 essays in the archive · 13 languages · 54 categories
I Evaluated 40 AI Podcasts Across 7 Languages. Six Pipeline Fixes Later, Here's What Improved.
I built an eval framework to grade my AI podcast pipeline across 40 real episodes in 7 languages. The review step was making scripts worse (3.34 → 2.91). Citation URLs all looked fake (score: 1.28). Chinese was the weakest locale (2.88). Six fixes, one day, and a new model later — here are the before/after numbers and every change I made.
Read the analysis
How I Write System Prompts That Actually Get Enforced (After My AI Hosts Said 'Trời ơi' 13 Times)
I spent days crafting a system prompt for my AI podcast hosts. They read every word and ignored almost all of it. The fix wasn't better words — it was understanding that LLMs process rules differently depending on where they live in your pipeline.

The One Question That Killed an Entire Curriculum Architecture in a Week
The product I shipped in April was right about the engine. The product I settled on in June was right about the spine. In between: two complete curriculum rewrites, a framework that killed an architecture in a week, and the question I wish I had been asking since March.
The Full Index
Earlier entries
Sort: Newest
Subscribe
Get the next useful note when there is something to say
I write about AI, work, expat life, and building products. Pick the topics you care about, and I'll send the next one when it's worth your time.
