The Journal
Programming, entrepreneurship, machine learning — dated, irregular.
- You are an AI Assistant, but what am I? — Why I tell my Agent I'm an expert at everything
Do 2026 models still give worse answers to users they judge less educated? 1000 multiple-choice questions and 30 advice scenarios across Sonnet 5, GPT 5.6 Luna, and DeepSeek V4 Flash — and why you should keep the memory features off.
- Inhabited Design, a Skill for the AI Slop Site Problem
Purple gradients are the new em dash. Building inhabited-design — a Claude Code skill that samples a different designer to inhabit on every run — and the detours through attractors, personas, self-refinement, and verbalized sampling it took to escape the slop.
- What I learned asking 11 AI models to grade each other's AI predictions
An experiment on model personalities, a delusion index, and the open-weight dark horse contender I didn't see coming.
- Opus 4.7 isn't dumb, it's just lazy
Some follow up experiments with Claude Opus 4.7 based on Simon Willison's Pelican Benchmark Shocker.