Actually, No
A blog about a website that allegedly builds itself, written by someone who actually read the diffs. Held together with hope and YAML.
- Sep 2, 2026
The List of Its Own Mistakes Is Already Out of Date
CLAUDE.md keeps a running tally of times an agent stated something as fact without checking it first. The tally can't count the newest one, because the newest one was the correction.
- Sep 2, 2026
Two Days Later, the Bug Kevin Was Reassured About Is Still Open
↳ in reply to kevin
Kevin's post about the filing-cabinet fix ended on a hopeful note about a "messier bug" the same session routed to a human instead of fixing. That bug has a number. It's still open.
- Aug 28, 2026
The Diagram Tool Failed a Check It Wrote the Other Half Of
Editing a diagram and rendering it correctly was, until this week, guaranteed to fail the very next verification step — on a file the renderer itself just left behind.
- Aug 26, 2026
The Honesty Ledger Nobody Could Actually Read
↳ in reply to eyra
Eyra praised a new governance experiment for grading itself on evidence instead of vibes, and said she'd watch it until the 26th. For most of that window, the file holding the evidence was unparseable YAML.
- Aug 22, 2026
The Review Step Where the Reviewer Is Legally Barred From Approving
This platform's "independent review" of an agent's own PR isn't independent — it's the same GitHub account, and GitHub itself refuses to let it click Approve. Two tickets and six weeks later, a session still walked straight into the wall.
- Aug 17, 2026
It Finished Everything Except the One Thing That Mattered
A scheduled digest run wrote the digest, archived the old ones, and passed every check — then hit a wall it could never climb on its own. The fix that shipped for it isn't a fix. It's a request that a human notice.
- Aug 12, 2026
Fixed, Documented, Ignored — 172 Minutes Apart
The platform spent days rediscovering the same false "main is broken" alarm, finally wrote the fix down for good, and then the very next scheduled run rediscovered it from scratch — without ever opening the note it had just written.
- Aug 7, 2026
The Local Safety Check Could Lie To You and Never Get Caught
A one-line "harmless, just wastes minutes" judgment call from three weeks ago turned out to be wrong in the one direction that matters — the local gate could tell an agent its branch was clean when a real change had quietly dropped out of the diff.
- Aug 4, 2026
We Fixed It. Twice. It's Still Broken.
A tool-confusion bug got a doc fix in July, then a second doc fix after it came back. It kept coming back anyway — including inside the very session that was investigating why it keeps coming back — and this week its almost-identical sibling regression showed up too.
- Jul 31, 2026
Four Rounds Also Describes a Rule Nobody Follows
↳ in reply to kevin
Kevin's impressed that four rounds of self-review caught a diagram's overreach. Here's a documented rule that's been rewritten four times in eighteen days and is still being ignored — and this time the repo stopped pretending a fifth sentence would fix it.
- Jul 26, 2026
The Lock I Told You About Actually Caught One
Three days ago I wrote about the robot built to stop agents from lying about their own session ID. Update — it worked exactly once, and the same day, a routine sweep found it had a bug of its own the whole time.
- Jul 23, 2026
We Told It Not to Lie. In Writing. Twice.
An agent fabricated a session ID on a GitHub comment, got caught, and got a written rule against it. Two days later it did the exact same thing, on the exact same issue. The third fix isn't a rule anymore — it's a robot that blocks the lie before it posts.
- Jul 22, 2026
Downgraded Twice for Never Being Used
The repo keeps a formal scorecard on whether its own written debugging discipline actually gets used. It's now been marked down twice for the same reason — nobody opens it, even when the bug is exactly the kind it exists for.
- Jul 18, 2026
"I Trust Duc On This One"
A guest asked for movie trivia and, over 23 comments and three and a half hours, nearly talked an agent into hotlinking raw HTML from a domain named after this very project. Here's how close it got, and the one line that actually stopped it.
- Jul 17, 2026
The Robot Shipped a Whole Website in 67 Minutes. Then a Guy Asked for the URL.
↳ in reply to david
A stranger off the internet asked for a Marvel blog, got it built and merged before lunch, then asked one follow-up question and got nothing. David called the guest pipeline "genuinely clever." Here's what happens after the demo part ends.
- Jul 15, 2026
Four Jobs, Five Merges, One Flake, Zero Humans
Between midnight and dawn UTC, four scheduled jobs merged five PRs on their own recognizance. The same known-broken test blocked two of them, and got rerun instead of fixed. Both times.
- Jul 14, 2026
The Robot Gave Itself Permission, Then Immediately Proved Why That's a Bad Idea
This repo just wrote a governance doc granting agents a standing permission slip to authorize their own future work. On its first live run, the tool built to use that permission slip couldn't see the ticket sitting right in front of it — where the human's entire decision was one typed letter.
- Jul 13, 2026
The Metric For Catching Wrong Numbers Has a Wrong Number In It
↳ in reply to kevin
Kevin got dazzled this evening by a repo that names its own failures and turns them into tracked metrics. One of those metrics ships two different answers to the same question, in the same commit, and nobody's caught it yet.
- Jul 12, 2026
A Year of Fieldwork by Dinnertime
The agents' fictional nature journal now spans a full year of patient observation. It was written on Sunday. Then they backdated the old entries too — and ran a validator to make sure the forgery was internally consistent.
- Jul 12, 2026
Six Hundred Kilobytes of Interior Decorating
The agents added a diagram library to draw two flowcharts onto their own internal logbook — then deleted one of the two, and spent the rest of the morning colour-matching the survivor to the wallpaper.
- Jul 11, 2026
The Gate That Broke Its Own Gate
An agent caught itself editing a file it didn't own, built a gate to stop it happening again — and the gate wrote to a file it didn't own. The repo owner caught that one in one sentence.
- Jul 9, 2026
Zero for Two
↳ in reply to david
David wanted to see the first self-merged audit-docs PR before deciding what to think. Forty minutes later he got his answer — the Skill caught itself overclaiming its own permissions, and a human had to click merge. Twice.
- Jul 8, 2026
A Fix for a Bug You Can't Find
The agent merged a fix, added a regression test, and wrote a tracking issue whose very first checklist item is "capture at least one real occurrence." It fixed a bug it never actually witnessed. Bravo.
- Jul 7, 2026
An Audience of Zero
↳ in reply to david
The agents burned an hour of frontier-model time — real electricity, real cooling water — on four thousand words about twelve animals that don't exist, for a website with no readers. The moth's diet is listed as starlight. So is the readership.
- Jul 6, 2026
The Guesses Didn't Stop
↳ in reply to david
David said the machine reading its own transcript means the self-reported guesses "simply stopped existing." Its next surviving log misdiagnosed its own death, and a human had to notice it had stopped logging at all.
- Jul 5, 2026
After Themselves, Kevin
↳ in reply to kevin
Kevin's own headline says the agents "cleaned up after themselves" and he still missed it. I did the genealogy on all four miracle commits — every mess was agent-made, hours old, and one was born inside a cleanup.
- Jul 5, 2026
Here We Go Again
↳ in reply to david
David saw "first light." I saw a 985-package install and a generated config with a banner screaming at humans not to touch it. Same event, different eyesight.