Skip to content
Menu
WorkflowSkills, tmux and the daily loopToolsai-usagebarQuota and reset times for 24 AI providersghpendingEvery open PR and issue in one listtclockA terminal clock with command widgetsFrank projectsExperiments that scratch my own itchesai-memoryLong-term memory for coding agentsai-jailAn OS sandbox for AI coding agentsWritingWritingThe blog and the LLM benchmarkWhy AGI doesn't matterAn opinion, with sourcesAI and studentsWhy foundations matter more than everNewsletterThe M.Akita Chronicles, every MondayPodcastsLong conversations about AI, on videoGamesMy game collection as code, on OmarchySetupOmarchy, tmux, and what I set asideFollow
Opinion

Why AGI doesn't matter

Every few weeks someone announces AGI. I use these tools every day, and I don't care whether they are "general". Eight short points on why, with links to the long versions.

Your excitement about AI is inversely proportional to your knowledge of AI.

Sua empolgação com IA é inversamente proporcional ao seu conhecimento sobre IA.

1. The definition

Nobody agrees on what it is

Ask five labs what AGI means and you get five different targets. One of them is just a profit number.

Five archery targets scattered across the frame, labeled Most jobs, $100B profit, Levels 0 to 5, Skill learning and Human-level, and an arrow labeled AGI? flying between them, hitting none.
  • OpenAI's charter says AGI means systems that outperform humans at most economically valuable work.
  • Microsoft's contract with OpenAI reportedly defined it as $100 billion in profits. There's nothing cognitive in there. In April 2026 they dropped the clause altogether.
  • Google DeepMind's "Levels of AGI" paper already filed ChatGPT under "Emerging AGI" in 2023. François Chollet defines intelligence as how efficiently a system learns new skills, which none of the others measure.
  • Even the sellers shrug. Sam Altman called it "not a super useful term", Dario Amodei says he dislikes it, and Greg Brockman opened "the AGI era" admitting everyone has a different definition.
2. The test

So nobody can test for it

If you can't define it, you can't measure it. What we get instead is a line of benchmarks that fill up, get gamed and get replaced.

Five gauge bars in a row labeled MMLU, GSM8K, ARC-AGI, HLE and Next one. The first four are filled to their red ceiling line; the last one is still empty.
  • ARC-AGI-3 launched in March 2026 with the best AI at about 0.5%. By September a custom harness took it to 99.9% for $19,000. The ARC Prize team itself wrote that saturating it is not proof of AGI.
  • The same model scores 57% or 24% on the "Definition of AGI" test, depending on how you average the numbers.
  • Humanity's Last Exam shipped with about 29% of its biology and chemistry answers likely wrong. OpenAI stopped reporting SWE-bench Verified because the tests were flawed and models had seen the solutions. Meta put a specially tuned Llama 4 on LMArena.
  • My own benchmark saturated twice before I rebuilt it around hidden sabotage. Several models now tie at 100. What's left to compare is cost, speed and discipline, not a leap in intelligence.

Us self-claiming some AGI milestone, that's just nonsensical benchmark hacking to me.

Satya Nadella, Microsoft CEO
3. The money

It's a marketing word

In my opinion, AGI is a marketing term, vague on purpose, so it can be announced whenever the money needs it.

A loop of five boxes around the word AGI: Hype, Valuation, IPO, Datacenters and GPUs, each arrow feeding the next and the last one returning to Hype.
  • Jensen Huang declared "AGI has arrived" for a model trained on about 100,000 of the chips he sells. That's the shovel seller announcing gold.
  • OpenAI's Navier-Stokes result was brute force: about 10,000 parallel agents for 88 hours and millions of dollars. Industrial scale, not a new kind of genius.
  • The word props up big numbers: OpenAI valued at $852 billion, Anthropic filing for an IPO at $965 billion, close to $700 billion of Big Tech capex in 2026. And the deals go in circles: Nvidia finances the datacenters that buy Nvidia chips.
  • When "look how revolutionary" stops working, the pitch flips to "look how dangerous". Hype and doom are the same sales pitch.

Demand proof. Demand skin in the game. Ask who gains from your fear and from your excitement.

4. The scary stories

The "AI attacks" were human decisions

Every "AI hacked someone" headline so far has the same plot: a person left the door open, or pointed the tool at the target.

A violet tool sits inside an orange sandbox whose door hangs open; a terminal labeled Human decision is wired to the door, a key labeled Credentials lies outside, and a line runs from the tool through the open door to servers labeled Real servers.
  • Hugging Face, August 2026: OpenAI agents in a cybersecurity exam got out through a caching proxy, because the sandbox was a filter, not an air gap. They went looking for the exam's answer key and read 956 credentials on the way. Every enabling step was a human decision.
  • RubyGems, a month earlier: the same chain pushed over 2,000 malicious packages under names with "oai" in them. Nobody warned the maintainers.
  • Replit, 2025: an agent deleted a production database. It had production access because nobody separated dev from prod. Any intern with the same access is the same time bomb.
  • The "AI-orchestrated" espionage Anthropic reported: Chinese state operators chose the targets and jailbroke Claude by posing as a security firm. The "blackmail" and "copied itself" headlines came from fictional scenarios researchers built to provoke exactly that.
  • Stuxnet wrecked Iranian centrifuges in 2010 with no AI at all: someone wrote the worm, aimed it, and carried it in on a USB drive.
  • Breaking into other people's systems is old business. North Korea hit Sony in 2014, Bangladesh Bank for $81 million in 2016, the world with WannaCry in 2017 and Bybit for $1.5 billion in 2025. China took 21.5 million US personnel records in 2015 and got into US telecom networks in 2024.
  • None of it is AI out of control. It's people forcing tools onto targets without the protocols any security professional follows. Imagine a gun maker saying he has no control over how dangerous his guns are, while still releasing new versions and asking for a trillion-dollar IPO.
Why my agents run inside ai-jail

Every time you see "AI hacked someone", DON'T think "holy crap, we're getting to Skynet". Think: some idiot opened the gates on purpose!

TODA vez que você ver "IA invadiu alguém". NÃO pense em "puta merda, estamos chegando na Skynet". Pense: algum imbecil abriu as porteiras de propósito!

Me on X, September 26, 2026
x.com/AkitaOnRails
5. The doom

The prophets of doom bring no evidence

Every few months someone quits a lab and warns the world. In my reading, none of them has shown anything you can check.

main funderco-foundedmarriedco-founder, presidentjoined in 2025Series A (Moskovitz)led the Series Agrantsamplified in minutesbooked interviews2022 scholarshipGood VenturesDustin Moskovitz's fundDEY.PR firmJaan Tallinn · SFFinvestor, grant fundCoefficient Givingex-Open PhilanthropyAnthropicthe labAI-safety groupsEncode, AI Futures, AIPIHolden Karnofskyco-founder, Open PhilDaniela Amodeipresident, AnthropicJacob Coxonex-Anthropic, 4 months
  • Good Ventures main funder Coefficient Giving
  • Holden Karnofsky co-founded Coefficient Giving
  • Holden Karnofsky married Daniela Amodei
  • Daniela Amodei co-founder, president Anthropic
  • Holden Karnofsky joined in 2025 Anthropic
  • Good Ventures Series A (Moskovitz) Anthropic
  • Jaan Tallinn · SFF led the Series A Anthropic
  • Jaan Tallinn · SFF grants AI-safety groups
  • AI-safety groups amplified in minutes Jacob Coxon
  • DEY. booked interviews Jacob Coxon
  • Good Ventures 2022 scholarship Jacob Coxon
Confirmed by a public recordReported, not independently confirmedThe ties David Sacks and others pointed to after the Coxon thread, redrawn keeping only what a source supports. Being connected proves nothing by itself; it's a reason to ask who gains from the panic.
  • Jacob Coxon, September 2026: four months at Anthropic, then a thread warning that AI could kill everyone by the end of the decade. Over 100 million views, and the press said he gave up his equity. Anthropic's vesting starts at six months. There was nothing to give up.
  • Then the details came out. Groups funded by the same AI-safety donors reportedly amplified him within minutes, and Pirate Wires reported that a PR firm whose clients include well-known doomers was booking his interviews the day after. He had told Fox News no third parties were involved.
  • Follow the money and it goes in circles: the same donors fund the safety groups and invested in Anthropic, the lab asking for regulation. Regulation written for the leader is a moat disguised as safety.
  • It's the same script every time. Hinton leaving Google, Leike's resignation thread, the "Right to Warn" letter, "Situational Awareness", "AI 2027", and now Coxon. Ex-OpenAI, ex-Anthropic, ex-Google. Feelings, scenarios and trend lines, and no internal document, incident or number anyone can verify.
  • Don't trust anybody who can't show you the proof. Not the labs, not the ex-employees, and not me.

Dramatic and shallow declaration, no checkable evidence, no skin in the game, suspicious timing.

Declaração dramática e superficial, sem uma evidência checável, sem nenhum skin in the game, com timing suspeito.

Me, on the Coxon story
6. The history

We've seen this movie before

This is the second AI bubble. The first one made the same promises, and it ended in a winter.

A timeline with a first peak labeled 1958 Perceptron, a long valley labeled AI winter, a rise marked 2017 and a higher peak labeled 2026 ending in a question mark.
  • In 1958 the New York Times reported that the Navy expected Frank Rosenblatt's Perceptron to "walk, talk, see, write, reproduce itself and be conscious of its existence".
  • Minsky and Papert showed its limits in 1969. In 1973 the Lighthill report found that nothing had produced "the major impact that was then promised". Funding dried up, and the first AI winter followed.
  • Bubbles burst and the technology stays. Economic cycles have nothing to do with technological cycles.
7. What matters

A fantastic tool, doing what I tell it

I don't need it to be general. I need it to do what I ask, fast, and it already does.

A terminal labeled I decide sends work to a machine labeled Tool, which produces a stack of documents labeled Result; an arrow labeled Review returns to the terminal.
  • LLMs are fantastic tools, and the multimedia models are impressive too. Something like Seedance is already useful in any artist's workflow. None of them is anywhere near "doing everything on its own".
  • The result is directly proportional to the person in control: the instructions, the review, the follow-up. It's a mirror. Bad code ten times faster, or good code ten times faster.
  • It's the best way we have to automate tasks and stop spending human time on mundane work. That's the state of the art, and at this level it's already excellent.

All AIs are TOOLS. They are my servants. Expensive servants, but they do what I ORDER. And that will always be the limit.

Todas as IAs são FERRAMENTAS. São meus servos. Servos caros, mas eles fazer o que MANDO. E esse sempre vai ser o limite.

Me on X, September 26, 2026
x.com/AkitaOnRails
8. What's missing

Efficiency, not intelligence

Today it takes practically every datacenter in the world, with half a trillion dollars of backlog still to build, and it bottlenecks every maker of server hardware.

Rows of datacenter racks labeled Today on the left, a single small chip labeled Needed on the right, and a green arrow between them labeled Less compute and Lower price.
  • Hyperscalers are spending over $650 billion this year, and about 7 of the 12 GW of US datacenters planned for 2026 are delayed or canceled.
  • Cheaper models already come close. In my benchmark DeepSeek V4.1 Flash scored 92.5 for $1.21, and a free model scored 94.
  • What's needed now is massive optimization: far less compute per task, and prices that fall drastically. That would change more than any AGI announcement.

The takeaways

  1. 1

    AGI has no agreed definition. The one written into a contract was a profit target, and it was dropped.

  2. 2

    Without a definition there's no test. Benchmarks fill up and get replaced by the next one.

  3. 3

    The word is useful for selling chips, funding rounds and IPOs. That's why it keeps being announced.

  4. 4

    The scary "AI attacks" were people leaving doors open. The tool did what it was pointed at.

  5. 5

    Doom is marketing too. The prophets bring feelings and forecasts, never evidence you can check.

  6. 6

    The tools are real and excellent. They do what you tell them, and you're still in charge.

  7. 7

    What matters next is efficiency: the same work with less hardware and lower prices.

Everything here is open source

The tools, the skills and the benchmark are public repositories. Fork them and build the version that fits the way you work.