How a checklist kept a coding agent honest

A coding agent built GetIntro’s onboarding in 69 hours, from the first plan to the last commit. This is what the checklist around it did, and what it caught.

· 9 minute read · 7 sections

I look after GetIntro from the side of the people who will use it. Part of that is knowing how what they use got built, and whether anyone checked it. This post is about the checking.

Over about three days, a coding agent built the way invited people will set up their pages. It wrote most of the code and ran most of the tests. A person made every decision that mattered.

What kept the agent honest was a tool called specloop. It is a checklist with rules. I think the rules are the interesting part, so most of this post is about them.

At a glance

The onboarding build, measured from the project’s own history
First plan to last commit68 hours 46 minutes
From23:20, Tuesday 29 September 2026 (Singapore time)
To20:06, Friday 2 October 2026
Commits63
Diary entries47
Tasks done95 of 103
Automated checksAbout 940, across 25 test runs
Real people onboardedNone yet

Those 69 hours are elapsed time, not effort. Commits came in bursts, with long gaps while the agent waited for a person. Counting only the stretches where commits were less than ninety minutes apart gives about 17 hours. That measures when work landed, and nothing more.

What specloop is

specloop is a small open-source tool, published on GitHub. You split a plan into numbered stages. Each stage is a flat list of tasks, and each task is a box to tick.

The agent takes the next unticked task that nothing blocks. It does the work and checks it in proportion to the risk. Then it ticks the box and writes down what it found.

Four rules did the real work

  • Done is counted, never remembered. A command recounts every tick and compares it with the summary table. If they disagree, it fails.
  • Order is written down. Each stage says what it depends on, and tasks carry a priority. The agent does not pick what it fancies.
  • Every report ends the same way. The same status table, with counts taken from the ticks, in every handoff.
  • The diary is never rewritten. Each session adds a dated entry. When an instruction overrules an earlier decision, the entry says what changed and why.

The 69 hours, in order

What landed, by day and time (Singapore time)
WhenWhat landed
Tue 23:20 to Wed 02:20The plan, written down and approved. Then the product’s direction, the checklist set up, and three stages finished: the profile store, sign-in by emailed code, and live pages at their own address.
Wed 08:00 to 11:00The brand rebuilt as a design theme, the admin screens rebuilt, and a person able to publish a page by 09:08. A real LinkedIn PDF was read at 10:07 and disproved four things the plan had assumed.
Wed 15:00 to 21:00A way to read the onboarding funnel for each environment, and two more plans written: a test setup, and a way for AI tools to reach GetIntro.
Thu 17:52 to 22:28The owner’s real LinkedIn exports, taken apart. Which files they hold, what each adds, and what should never be opened.
Fri 07:55 to 20:06The step-by-step LinkedIn guide, a second way to fill in a page, a photo step, large-file upload, a model that drafts first versions, and a conversation where you can write about yourself in your own words.

In the small hours of Wednesday, the agent ran out of work it was allowed to do. Everything left needed a decision from the owner. The checklist said so plainly, and the run stopped there rather than inventing something to do.

There was also a gap of 21 hours between Wednesday evening and Thursday afternoon. It ended when the owner supplied LinkedIn’s own PDF, which only the owner could produce.

What the checklist caught

It would not let the agent stop early

Early on, the agent reported itself blocked while two stages were plainly available. A check that runs when the agent tries to finish pointed out the contradiction. The agent went back and finished them.

It wrote the rule into the diary in its own words. A blocker on one stage ends that stage, never the whole run.

It would not let a box be ticked for show

Some tasks had two halves: one that works without a model, and one that needs a model. The agent built the first half and left the box open. It ticked the box only when the second half existed, with tests.

The last task to close this way was the conversation. Its table row still says, in plain words, that the live model check is pending.

It made retiring a task honest

Two tasks were removed rather than ticked. One was a decision that belonged to the owner, so it moved to a launch checklist. The other was a model step that turned out to add cost and no benefit.

Each removal made the total fall by one. The diary says why: no work was done by moving a box, so no work is claimed.

It recorded when the owner overruled the agent

The owner asked the LinkedIn guide to request the entire archive. The agent pushed back with evidence. The larger archive held thousands of other people’s names and no extra profile data.

It held the change until a test could settle the question. The test showed the owner was right, for a different reason. LinkedIn sends everything whichever box you tick. The diary kept the whole exchange: the earlier decision, the instruction, and what survived.

It showed that shipping before the check is a mistake

The agent built the LinkedIn guide on an assumption nobody had tested. The task that would have tested it was already on the list, marked important. An hour later the test overturned the guide’s main advice, and it was rewritten.

Nobody outside had seen the screen, so it cost an hour and no more. The lesson went into the diary: do the task that checks the claim before the one that depends on it.

The summary every report ends with

Each time the agent reports, it ends with the same table. The counts come from the ticks, so I can trust a number without asking how it was reached.

The status summary: eleven phases, which this post calls stages, with their progress and status. Ten are complete, including two marked as complete with a live model check pending. The first invite round shows 0 of 8, not started, waiting on decisions.
The status table as the agent handed it to the owner at the end. The tool calls the stages “phases”. Ten are complete, two of them with a live model check still pending, and the last waits on the owner.

Look at the last column. “Complete” and “complete; live model check pending” are different claims, and the table keeps them apart. I would rather read an honest qualifier than a clean green tick.

What it did not do

A checklist tells you the work shipped. It does not tell you whether anyone can use it. A passing test run does not either.

Many of the defects the agent found were found by looking at a screen, not by a test. One form posted to an address nothing answered, yet every check passed. The checks posted to the right address directly. The agent only noticed when it built a second guide.

The model that writes first drafts has not yet answered a live request. It works against a stand-in in tests. Until a real answer comes back, I will not claim it writes well.

What happens next

No real person has used this yet. That is the point of the next stage, and the checklist shows it at 0 of 8.

Before anyone is invited, the owner writes down how many people to ask and in what order. They also write down what pattern would count as working. Then a small number of people on the waitlist get an invitation. I will read what happens, and we will write that up too, including if it does not go well.

Keep reading

Building onboarding with a coding agent Two and a half days, one written plan, and a person who kept the decisions. What the agent built, what it got wrong, and what only looking could catch. · 12 minute read · 9 sections

All build notes

If you would like a page of your own, the waitlist is open.

Reserve your name