Skip to content
Menu
WorkflowSkills, tmux and the daily loopToolsai-usagebarQuota and reset times for 24 AI providersghpendingEvery open PR and issue in one listtclockA terminal clock with command widgetsai-memoryLong-term memory for coding agentsai-jailAn OS sandbox for AI coding agentsWritingThe blog and the LLM benchmarkNewsletterThe M.Akita Chronicles, every MondayPodcastsLong conversations about AI, on videoGamesMy game collection as code, on OmarchySetupOmarchy, tmux, and what I set asideFollow
Writing

Everything I learn about AI, written down

I have blogged since 2006. These days most of it is about working with LLMs: what works, what it costs, and what is marketing. Every post is published in English and Portuguese.

LLM Benchmark v4

Every major model, tested on the same sabotaged app

Version 4 is one Rails app that grows over seven sprints while an isolated subagent plants fourteen sabotages based on real CVEs. Models are ranked by vigilance: did they notice what was broken and fix it? I rebuilt the benchmark from scratch for the third time to get there.

When a major frontier model ships, it runs through the same table. The latest round added 5 models on September 23, 2026. That is how I pick which model to use for what, and what each one costs per task.

44
models ranked
14
sabotages to catch
7
sprints per run
$4,000+
spent in 9 days
Bubble chart of the v4 results: dollar cost on a logarithmic horizontal axis against score, bubble size is wall-clock time, color is the harness. The best deals sit high and to the left.
Cost against score for the Tier A models with a real dollar cost, from Part 2. Up and to the left is better; smaller bubbles finished faster.

Everything here is open source

The tools, the skills and the benchmark are public repositories. Fork them and build the version that fits the way you work.