Jul 17, 2026
Two tasks from the same Claude Code run, same mid-tier model: one cost 76K tokens, the other 970K. The gap wasn't the model.
Jul 10, 2026
My agent was driving the browser like a cautious intern. One tool swap later, the same task went from 40 seconds to 1.5. Here's why, benchmark included.
Jul 3, 2026
Claude Fable 5 rewards short prompts and punishes vague ones. Here's the difference, and what still applies once the free ride ends.
Jun 26, 2026
The machine writes the code now. What stays rare, and what pays you next: verification, judgment, and direction.
Jun 19, 2026
On June 9, Fable 5 crushed every benchmark. On June 12, the US government unplugged it. Here's why the people who didn't feel a thing weren't on Fable.
Jun 12, 2026
Boris Cherny writes loops, not prompts — the three-rung ladder that decides who compounds and who stalls.
Jun 5, 2026
Robin Sloan built an app for his family in 2020 — and why the AI era just handed that exact power to everyone who can write a sentence.
May 29, 2026
Anthropic's new model is the first that stops quietly shipping broken code and telling you it worked — the upgrade non-coders actually needed.
May 22, 2026
Two competing "Managed Agents" launches, one identical architecture — and the cage that quietly closes around your stack.