100,000 Commits Later: What A "Claw" Actually Is
OpenClaw started as an obscure repo called Warelay in November 2025. Simon Willison's mid-year tour of 2026 explains how it ate the year.
The repo that ate the year
In November 2025, someone made the first commit to an obscure GitHub repository called Warelay. By the end of January 2026 it had been renamed four times — CLAWDIS, then CLAWDBOT, then Moltbot, then OpenClaw — and had accumulated 8,300 commits in under two months. Simon Willison, speaking in his closing keynote at the WeAreDevelopers World Congress North America in San Jose last Friday, said he checked the number again before the talk: over 100,000 commits. He calls it the most vibe-coded piece of software in existence.
That repository defined a software category. There is now OpenClaw, NanoClaw, IronClaw, PicoClaw. The industry is busy rebranding them as "personal agents" or "general agents." Willison still calls them Claws.
What a Claw is, underneath
Strip the branding and a Claw is not a new invention. As Willison puts it, a Claw is a coding agent wearing a less threatening hat — under the hood it works the same way, by writing code and then executing that code on your computer to get things done.
That mechanism matters more than the marketing. A chatbot generates text you read. An agent generates code and then runs it, on a real machine, with real file access and real network access. The capability jump people noticed in 2026 came from that loop becoming reliable. Willison dates the turn to November 2025, when Claude Opus 4.5 and GPT-5.1 shipped. Both were incremental improvements, but paired with their coding harnesses — Claude Code, which had existed since February 2025, and the younger Codex — they crossed from "often make mistakes" to reliable enough for daily use.
The consumer consequence was physical. Bay Area Apple stores sold out of Mac Minis because people were buying them to run OpenClaw. Drew Breunig's explanation, which Willison quotes approvingly: your Claw is a digital pet, and the Mac Mini is the aquarium you keep it in. In March, companies in China hosted OpenClaw install parties with non-technical people queuing around the block for help getting one running.
The risk is the same loop. If a program writes and executes code on your behalf, the interesting question is what stops it. Willison counted roughly 40 of the 277 sessions at the San Jose conference touching on sandboxing or agent security. The race, he says, is to build the first safe Claw — one you can hand to a regular person without them instantly shooting themselves in the foot.
Tokenmaxxing went up, then straight back down
The cost story is the clearest signal in the whole talk. In February, headlines had Meta making AI adoption part of performance reviews, Microsoft pushing every employee to use AI, and Uber boasting 90% of its engineers were on AI workflows. Months later: Meta cracking down on token use, Microsoft saying tokenmaxxing is "not what we are optimizing for," Uber capping employee AI spend.
Nothing broke. The bill arrived. Willison's framing: a year ago it was hard to spend $50 on tokens because there was nothing interesting to do with them; with agents you can now spend $1,000 in a day doing real work. Agents found product-market fit and that fit is expensive. He ties the same dynamic to Anthropic's valuation reaching a reported figure of maybe a trillion dollars.
Meanwhile the older arguments about whether models can code quietly ended. StrongDM has operated since July 2025 on two rules: code must not be written by humans, and code must not be reviewed by humans. Dan Shapiro calls this the Dark Factory — automate enough and you can turn the lights out. In February that sounded radical. Willison's read is that StrongDM were simply six months ahead, and that the real open problem is no longer generation but verification: how do you know unread code is good?
Questions You Should Be Asking
- If a Claw writes and executes code on a company device, what sandbox is it running in — and who at your organisation can describe that sandbox without checking?
- Your team's agent spend went from trivial to meaningful this year. What is the per-engineer daily ceiling, and what happens to output when you enforce it?
- StrongDM banned human code review with decades of security experience in the room. What verification do you have that would catch what a human reviewer would have caught?
- When a vendor sells you a "personal agent," is it doing anything structurally different from a coding agent with a friendlier interface? Ask them to explain the difference in terms of what runs on the machine.
- Google trained Gemini 3.1 Pro across many animals on many vehicles and, by Willison's account, defeated his pelican-on-a-bicycle test. What in your evaluation suite has not already been trained for?
What To Watch Next
Watch whether a safe Claw actually ships to consumers. Meta's Muse launched three weeks ago and sits at the top of the free charts on the iPhone App Store — the first mass-market test of whether this category survives contact with people who did not buy a Mac Mini as an aquarium. Willison's own newsletter flags Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna and a new price war as the next chapter. If frontier prices fall while token budgets stay capped, the constraint on agent adoption stops being capability and becomes procurement.
- 1Before running any Claw-style agent, sandbox it in a container or VM with scoped API tokens, since it executes code and shell commands on your behalf.
- 2Treat commit velocity as a red flag, not a feature: audit a vibe-coded repo's dependency list and permission scopes before adopting it in production.
- 3Pin a specific commit or release tag when depending on fast-moving agent repos, because four renames in two months will break your imports.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
