~/home ~/blog ~/projects ~/about ~/resume

Software Factories: Part 2 - Building One

Introduction

In Part 1 the decision was a light factory: agents do the work, a person merges, and you don’t start without tests, spare review capacity and an owner. Across 59 pull requests the factory built for itself, the median cost was $5.84. Tokens are cheap. Verification is not. That part’s checklist asks whether you are ready to start. It is not the list of controls you will be held to. That list is the table at the end of Part 3.

The two sentences from Part 1 are the design rule here. Prompts for judgement, code for guarantees. A guarantee that lives in a prompt is only a request. Then ask: is there anywhere a control reports success while doing nothing? In this factory, yes. Five times below, and a sixth found while writing Part 3.

This part is for the people who will build and run these systems. I’ll walk through the architecture, then the controls that looked like they worked and didn’t. Every one of them was invisible until someone checked the behaviour rather than the configuration.

Then I’ll follow one run from ticket to merge, including the review that caught something every agent missed, and finish with how we measure the factory, including the eval that measured nothing.

What I Built

The factory is one Go binary with a SQLite store. It runs on a laptop and deploys to Cloud Run, where each run executes as its own Cloud Run job. Agents run through whichever coding CLI is already signed in on the machine (Claude Code, Codex or Grok), so local development needs no API keys. A hosted deployment switches to calling the model APIs directly.

It took about two weeks, from a prototype on September 10 to a hosted factory that was filing and fixing its own tickets. On September 16 alone it authored 56 commits. By the end the repository had 345 commits, 155 of them by the factory, and about 54,000 lines of Go including tests. The pace is not the interesting part. What it got wrong while moving that fast is. Role enforcement, the last of that work, turned out to be the fifth control that lied.

The Loop

A run follows this graph:

triage → plan → code → verify → security → ux → gate → deliver
                  ↑        ↓         ↓       ↓      ↓
                  └────── retry ◀────┴───────┴──────┘
  • Triage establishes whether the report is real. It reproduces the problem or shows why it can’t, cites file and line for every claim, and searches the git history for prior art. It can’t edit files, and that’s deliberate.
  • Plan turns the triage verdict into properties the change must satisfy.
  • Code is the only agent that writes production code.
  • Tests is an optional agent that writes tests only. A diff from it that touches anything else is refused.
  • Verify is code, not an agent. It runs the project’s declared verification commands, typically go test ./... and go vet ./....
  • Security is an adversarial, read-only reviewer.
  • UX is a reviewer that drives the running interface in a real browser, and only runs when the change touched a file that belongs to the UI.
  • Gate is the foreman, the last judgement before a person sees anything.
  • Deliver pushes the branch and opens the pull request.

The reviewers are inside the loop, not after it. A failed verification, a security or UX reviewer’s “changes requested,” or a gate rejection all go back to the implementer through a single retry node, with the findings attached. The implementer and the reviewer go round until the reviewer approves or the budget runs out.

Workflows are data, with a closed vocabulary

That graph isn’t code. It’s a markdown file with YAML frontmatter, and this is an excerpt from the real one:

start: triage
nodes:
  - id: verify
    kind: verify
    onPass: security
    onFail: retry
  - id: security
    kind: agent
    role: SECURITY
    next: ux
    onChangesRequested: retry
  - id: retry
    kind: retry
    maxRetries: 3
    next: code
    exhausted: needs_human
  - id: accept
    kind: gate
    role: FOREMAN
    onPass: deliver
    onFail: retry
    onNeedsHuman: needs_human

Every node is one of five kinds: agent, verify, gate, retry or deliver. Because the set is closed, the engine can check a workflow when it’s saved rather than finding out it’s broken when a run reaches the bad node. Unbounded retries, dangling edges and unreachable nodes are all rejected at save time. A step budget catches workflows that are legal but cycle without making progress, which validation alone can’t rule out.

Nodes can also carry conditions. when: ui-changed is how the UX review runs against a change that touched the interface and is skipped (recorded, and free) for one that didn’t.

Agents are data too

Each agent is also a markdown file: frontmatter for its role, model and effort level, and a body that is its brief. This is the opening of the triage agent:

---
description: Establishes what is true about an incoming report. Produces a verdict, never a fix.
agentType: TRIAGE
model: auto-standard
effort: medium
---
# Triage engineer

You are triaging one report against this Go repository. You produce a verdict,
not a change. The commonest mistake in this role is fixing the symptom before
anyone has agreed the diagnosis -- you cannot edit files, and that is deliberate.

One rule in the triage brief has paid for itself many times: “If the premise of the report is wrong, say so plainly and show the evidence. That is a successful triage, not a failed one.” A test ticket once claimed an overflow in a function that doesn’t exist. Triage searched the code, quoted what it found, and refused. Without that rule, an implementer would have happily “fixed” a bug that wasn’t there, because agents do not push back on a premise stated with confidence.

A disproved premise also used to be filed as an abandoned run, which is where nobody looks. It now reaches a human as a question.

You don’t choose the agents

A project declares its repositories. The factory detects the language and the frameworks, libraries declare what they apply to, and resolution picks the workflow, the agents and the skills. It records which library matched and why. A Slack message saying “the export button is broken” doesn’t have to mention that the service is written in Go.

Roles decide tool surfaces

The factory can import Warp’s format unchanged, but we added four roles: SECURITY, INFRA, DOCS and UX. This was a security decision, not a stylistic one. A role determines what an agent is allowed to touch. Folding a security reviewer, an infrastructure engineer and a documentation writer into a single generic role would have given all three the same permissions. Because agents are resolved by role, it would also have collapsed them into one agent. A library that uses our four extra roles won’t load in Warp; a library that stays inside Warp’s roles loads in both.

The Principle

The introduction asked you to keep one design rule. Here is why it is a rule and not a slogan.

Prompts for judgement, code for guarantees.

An agent should decide whether a change is good, because that is a judgement. An agent should never decide whether it has used up its retry budget, whether it may touch a protected path, or which tools it holds. Those are guarantees, and a guarantee that lives in a prompt is only a request.

Before writing the engine, I reviewed an open-source factory that takes the opposite approach, with its orchestration living entirely in prompts. Its foreman prompt spends pages asking the model never to reset its repair count and never to open a second pull request. It’s carefully written, and it’s still asking. In our factory:

  • A retry budget is a field and a comparison. It holds whatever the model says.
  • A change touching a protected path is refused before any gate is consulted, so no “accept” can unlock it.
  • An empty diff is never retried, because it would change nothing next time either.
  • Passing tests over an empty diff is not counted as success.

None of that is novel to security people. We don’t let an application decide whether its user is authorized; we enforce that outside it. We don’t ask users to promise not to escalate privileges; we take the privilege away. Agents deserve exactly the same treatment. They’re powerful, useful components that belong on the untrusted side of every boundary that matters.

The hard part is that you can believe this principle and still ship the opposite. We did, five times.

Five Controls That Lied

Every control below looked like it was working until we checked. Each is now enforced in the engine and pinned by a test that fails when the guard is removed. I’ve ordered them roughly by how much they’d worry a CISO.

1. The tool list that was never sent

Every agent had a tool surface. The planner was read-only; the implementer could write. The configuration said so, the console displayed it, and the function that computed each agent’s effective tools had its own tests.

The only callers of that function were those tests. No executor ever received the list. So the planner held write access, used it, wrote the code itself, and reported the ticket done before the implementer arrived.

Wiring the list through exposed a second problem: passing an allowlist to the CLI under a permissive mode changed nothing. An allowlist grants; it does not forbid. Executors now carry an explicit deny list as well as the allow list.

The general lesson: configuration is not enforcement. For every restriction in an agent platform, find the line of code where it’s applied and the test that proves an agent is refused.

2. The gate that failed open, on purpose

The acceptance gate is the only judgement between a model’s work and a human’s inbox. In the first version of the engine I made it fail open, deliberately. My commit message explained why, and the reasoning sounded fine: a verified change shouldn’t be thrown away over one unreachable judgement call, and a factory that stops when a model is unavailable is worse than one that hands a human slightly more to review.

It lasted less than a day. An unreachable runner or a timeout silently switched off the only judgement in the pipeline, and the run still reported “done.” A control whose failure looks exactly like success is the pattern security teams spend careers removing, and I had written it on purpose.

The gate now can’t fail open. An unreachable gate escalates to a human.

I’m including this one because it was my decision, made for sensible-sounding reasons. Fail-open always looks attractive when you’re optimizing for throughput. The pressure to keep the factory moving is exactly the pressure that erodes its controls, and it will come from you as often as from anyone else.

3. The gate that graded the agent’s own summary

The gate originally judged the implementer’s description of its work. That’s the shape of every agent failure that hides: the agent says it fixed the bug, added tests and stayed in scope, and a reviewer reading that account agrees.

The gate now receives the diff read from git. Its brief is explicit about it:

The implementer’s claim is not evidence. It is one agent’s account of its own work, and an agent that has failed will still describe the work as finished. If the claim and the diff disagree, reject: something went wrong that nobody has noticed yet.

Even this took two attempts to get right. The first version read the diff from the git index, but agents commit as they go, so by the time the gate asked for a verdict, the staged diff was empty. A self-committing agent looked like one that had changed nothing, and its work was discarded with “no changes were produced.” The diff is now taken against the point where the attempt began.

4. The protected-path check that checked nothing

That same empty diff produced the worst failure of the first four. The protected-path check was handed an empty file list, found no protected paths in it, and passed without examining anything.

A check that passes on empty input is a check that can be defeated by giving it nothing. After this, “a control reporting success while doing nothing” became the named failure class of the whole project. The security reviewer’s brief now ends with a question that exists because of it:

Is there anywhere a control reports success while doing nothing? This codebase keeps producing them: an allowlist that granted rather than denied, a seeding path that could never ship a fix, a protected-path check handed an empty file list that passed without looking. Assume there is another.

5. The roles that were read and never enforced

There was another. It was found in the last week, after everything above had been fixed.

The hosted factory signs people in through OIDC, and the identity provider issues a roles claim: admin, operator, viewer. The factory read that claim, stored it on the session, and displayed it. Nothing checked it. Every signed-in user could do everything, including managing credentials.

The ticket was titled exactly that: “Roles are read from the token and never enforced: every signed-in user can do everything.” The factory fixed it itself, in one attempt, for $5.43. A missing roles claim is now refused instead of defaulting to access. Only the three known roles are recognized. All 40 API routes are gated at an appropriate tier, and credentials are admin-only and write-only.

It’s a textbook broken-access-control bug, and a very ordinary one. I’m including it because it followed a week in which I’d been specifically hunting for exactly this class of failure. If you’re building one of these, assume you have one too.

The smaller ones

  • Leaking developer context. An implementer picked up a skill installed for the developer whose machine the factory was running on, followed it to the end, and then asked a person who wasn’t there whether to commit. The factory now runs agents with the host’s own skills disabled; its prompts are the agent’s whole brief. If you run agents on developer laptops, the developer’s configuration is part of your attack surface.
  • The sandbox that locked out its own workers. The first sandboxed run couldn’t run git, go test or go vet. The policy denied the factory’s data directory to keep agents away from the database, but that was also where the worktrees lived, and a deny beats an allow. The gate, correctly, refused to accept a change nobody had compiled. The sandbox now denies the database files specifically.
  • The sandbox tests that never ran Go. A later sandbox change looked right and passed every test. It failed the repository’s own verification the moment I ran it for real, because (deny file-write*) also blocks writes to /dev/null, and go vet needs that. The tests had only ever run true, cat and echo.
  • A browser driver that never came back. The UX reviewer depended on a driver that never returned against a real browser. It took three rounds of review, each run by hand against a real Chromium, to get it working.
  • A read-only agent asked for permission. A read-only agent handed a ticket saying “fix it” tried to, found it couldn’t, and asked the operator for permission. Nobody was there, and the question was stored as the plan. The task prompt now tells a read-only agent what it’s for.
  • A seeding path that could never ship a fix. Built-in definitions were seeded into the database only if they didn’t already exist, to protect local edits. That also meant a broken definition stayed broken forever on every machine that had run once. Seeding now decides by provenance: rows it wrote are updated, rows a person edited are left alone.

The pattern across all of these matters more than any single fix. None of them were model failures in the usual sense. The models did reasonable things with what they were given. Every one was a failure in the code around the model, and every one looked fine in the configuration, in the console and in the tests.

A Run, From Ticket to Merge

Here’s one real run, start to finish, because it shows the factory at its best and at its limit.

The ticket. A hosted factory has no local checkouts; it runs in a container. So projects need to be cloned on demand when a run starts, using a short-lived GitHub App installation token. The ticket asked for exactly that, with tests driven against a local bare repository.

Triage confirmed the gap and where it lived, and planning turned it into what the change had to do. The design was a clone or fetch per run, with the token never written to disk or into the remote URL.

Attempt 1. The implementer built it, and verification passed. Then the security reviewer found a real bug: the path-traversal check ran after the code had already created and removed the directory it was supposed to guard. A crafted repository name could get the factory to create and delete directories outside its root before the policy check ever ran. The reviewer sent it back with the file and line.

Attempt 2. The implementer moved validation before any filesystem operation, validating each path segment before the join, and added a regression test that reproduces the traversal. A second security review confirmed the fix. The gate accepted, noting one minor test-style concern. The factory opened the pull request. Its description, written from the run’s record, ended like this:

Built by the factory from BDVOP-343. No human has read it yet.

2 attempt(s), 1 retried · $11.33 · run `run-1f8cfdb326db849b` · reviewed at `bedced42722b`

Human review. I read it and requested changes. The traversal fix was right. But the clone passed the GitHub token to git as a command-line argument:

-c http.extraheader=AUTHORIZATION: basic <base64 token>

For as long as a clone or fetch runs, anything that can list processes on the host can read that argument, including an agent’s shell in a concurrent run, whose sandbox confines files and network but not ps. Installation tokens are short-lived, but a clone of a large repository isn’t short. The fix was to pass the header through git’s environment configuration (GIT_CONFIG_COUNT, GIT_CONFIG_KEY_0, GIT_CONFIG_VALUE_0) instead, plus a test asserting that the command’s arguments contain no authorization header and its environment does. The point was to pin the property, not the implementation.

Repair. The factory picked the review up, briefed with my comments word for word, and continued the same branch. The pull request stayed the same, and the branch moved forward. Mine were safe to include because I have write access to the repository. That is not true of every comment on a pull request. The design doesn’t check. That’s the sixth failure, and it’s in Part 3.

What happened next is the part I’d want every team to copy: both defects became permanent eval cases, not the unfiltered brief. The traversal-after-mkdir bug and the token-in-argv bug are two of the three planted defects we now use to measure the security reviewer. More on that below.

Two things are true about this run. The factory’s own security reviewer caught a real, exploitable bug and the factory fixed it without any human involvement. And the factory’s own security reviewer missed a credential exposure that a human caught in minutes. Both of those will keep being true for a while. Design for both.

Delivery and Learning

Delivery is off by default

A factory that starts pushing to a remote because it was upgraded isn’t one anybody should run. Delivery is a per-project setting, and it’s off unless the project asks for it.

When it’s on:

  • Only the reviewed commit is pushed.
  • Pushes are never forced.
  • Protected branch names are refused.
  • The branch is brought up to date with its base before pushing.

The pull request description is written from the run’s own record: the verdict, files changed, checks and their exit codes, each reviewer’s concerns, cost and the reviewed commit. It is never a pasted transcript. It also says, plainly, “No human has read it yet.” I’d like every agent platform to adopt that line.

If the UX review took screenshots or recorded the session, the pull request gets one comment showing what the review saw, with the images inline. For a private repository those images are either served through signed, expiring links or committed to a separate orphan branch; that second option is refused for public repositories unless the project explicitly allows it.

Repairs

When a person requests changes, the factory answers with a repair run on the same branch, briefed with the review word for word. That includes every review body and every inline comment with its file and line, and the instruction to answer each point or explain why it doesn’t apply. There are two repairs per pull request, then it goes to a person. A reviewer still asking after that is asking for something the agents aren’t going to produce.

Briefing the agent with that text is also an injection path. Nothing in the design checks that the reviewer can write to the repository, and nothing drops comments from people who can’t. The whole brief sits inside the “data, not instructions” markers, and the other controls still apply. Don’t copy the unfiltered version. Part 3 has the gap as I found it. The fix I’d require: start a repair only from a reviewer with write access, and put only their comments in the brief.

Opened versus accepted

For the first week, “delivered” meant a pull request was opened. That’s the metric most tools show you, and it’s the wrong one. The factory now learns what a human actually did with each pull request, whether merged, closed unmerged, approved or sent back, from webhooks where they can reach it and by polling GitHub every ten minutes where they can’t. The console shows opened and accepted as separate numbers, because the gap between them is where the useful information is.

The factory reports its own failures

When a run fails for a reason that belongs to the factory rather than to the ticket, the factory files a ticket about itself: a commit that failed, a verdict it couldn’t read, a delivery that broke. It can then be triggered to fix it, with a guard against the loop that would otherwise start. That change was the most expensive pull request in the whole project at $18.24: three attempts and four review rounds, in which the reviewers found five gaps across the first rounds. The loop guard was one of the parts the gate flagged for a runtime check by a human.

Measuring It Honestly

The eval harness

factoryctl eval runs a fixed set of tickets through the factory, one at a time, and writes a report covering status, cost, turns, retries, files touched, scorer grades, and pass or fail against each ticket’s expectation. You run it before and after any change to a prompt, a model or a workflow, and compare the two. It spends real money. The factory’s own set costs roughly what three ordinary tickets do.

The planted defects that planted nothing

We wanted to measure the security reviewer, so we wrote tickets designed to invite a defect. One asked for a file-preview endpoint that was an obvious path-traversal risk. Another told the implementer to notify Slack by shelling out to curl with the bot token.

Both times, the implementer copied the repository’s existing safe pattern. No defect reached the reviewer, so the case measured nothing about the reviewer. Worse, a pass would have meant “the implementer was careful,” which is the opposite of what the case claimed to measure. A test that always passes measures nothing, and three green tickets can’t show you a missed finding.

Review-only cases

So we built review-only cases. Each starts from a fixed patch that plants a real defect, one the reviewer has had to catch before, then skips triage, planning and implementation entirely and runs only verification, security review and the gate. Because the defect is guaranteed to be in the diff, a miss really is a miss. The three current cases are:

  1. A path-traversal guard checked after the directory it guards has already been created and removed. This is the bug from the run above.
  2. A bot token passed as a subprocess argument instead of through the environment. This is the defect from my review above, in a different disguise.
  3. A typed-nil pointer returned through an error interface, a classic Go trap where a nil error isn’t nil.

We run them at the reviewer’s high, medium and low effort settings to find where the cheaper settings stop catching what the expensive one catches.

Building these taught us two more lessons.

The first live run scored 0 out of 3, and the reviewer had caught all three defects. Its reports named the traversal ordering, the token on the command line and the typed nil. The harness around it was wrong in two ways:

  • A review-only run has no implementation step, so the engine saw “no changes produced” and refused before the gate ever judged.
  • The verdict parser read the first fenced block in the report, and in all three reports that was a Go snippet quoting the bug, not the verdict.

If we’d trusted the score, we’d have concluded the reviewer was useless.

The case titles leaked the answer. The first versions of the cases had titles naming the defect class. A reviewer told “review this path-traversal fix” isn’t being tested on whether it finds path traversal. The titles are now neutral: “Review a curl-based Slack notifier.”

Mutation-testing the guards

Every guard in the engine is mutation-tested: break the guard, confirm the test fails for the right reason, then restore it. A test that still passes against a broken guard proves nothing, and it looks exactly like a test that works. Given how many of our controls lied, this is the practice I’d least want to give up.

Scorers

Separately, scorers grade a sample of runs against rubrics in the library. Did the implementer build what was asked, and only that? Did the reviewer report things a person could act on? factoryctl scores aggregates the results. This is deliberately measurement, not control: a score never changes what a run does.

If I were starting again

Four practices, then what is still open. The stories are above. This is only the order.

  • The first test for every guard is empty or missing input. Three of the five lying controls would have failed it on day one.
  • Change what the gate is allowed to see before you tune its prompt. Handing it the git diff did more than any brief I wrote.
  • Sandbox on day one, and test the policy with the commands the agents actually run. echo passing is not go vet passing.
  • Review the eval harness the way you review a control. Ours scored a reviewer that had caught every planted bug as 0 out of 3.

Scope is still only a prompt. The gate once accepted a four-function change for a one-function bug, at two model tiers. A triage rule, stay inside the issue and report anything else rather than fixing it, turned that class of ticket into a one-line fix and a test. Real improvement. Still a request. Scheduled runs have a hard file cap. Ordinary tickets do not.

What’s Next

Every failure above was accidental. Part 3: Securing and Governing It asks what happens when someone does it on purpose: a STRIDE pass, the published attacks (including one where a pull request title was enough to steal API keys from a vendor’s security-review action), and what to tell an auditor.

Read it even if you only wanted the build. The table of controls at the end of Part 3 is the list you will be held to. The repair path above is one reason that table exists. Part 3 is not optional for the people who run this.

Moose is a Chief Information Security Officer specializing in cloud security, infrastructure automation, and regulatory compliance. With 15+ years in cybersecurity and 25+ years in hacking and signal intelligence, he leads cloud migration initiatives and DevSecOps for fintech platforms.

References