// software factory IN DAILY USE

The Software Factory

One small piece at a time. More checks for riskier work. Low-risk work ships unattended.

This is how Fulltrace builds itself. I keep the roadmap and set the order. An AI coding session does the building, one small planned piece of work at a time, and I call each piece a slice. Low and medium risk slices run unattended, day and night. Riskier slices and all design work run with me in the session, and for risky work the rules call for a second AI to be given the design and told to break it before anything is built. Every slice starts with a written plan and ends with a written closeout. The build log is what the factory produced, and the roadmap is where it is headed.

550+ slices shipped since 30 june 300+ shipped unattended since 19 august 06 risk classes 19 standing rules

Unattended and Attended

Low and medium risk slices get built and shipped with nobody watching. I put rows in the queue in the order I want them and set the runner tier: which Claude model, how hard it thinks, and a cheaper setting for low-risk rows. A session with nobody in it then works the queue one slice at a time, day and night. Each slice runs the whole protocol on its own, from kickoff through its named checks to closeout and a commit, and the next one starts straight after. Over 300 slices have shipped this way since 19 August, and one overnight run took on 17 slices and shipped all 17.

It only takes work it is allowed to take

A row qualifies when it is proposed, its prerequisites have shipped and its risk class is R3 or lower. Rows for design or review work stay with me. Each slice gets four hours a try, unsaved changes in the working tree stop the run before it starts, and a slice that changes the runner's own files stops the chain until I have looked at it.

A fork is set aside and the queue keeps moving

A slice that meets a question only I can answer is parked in an Awaiting ruling lane, and the chain moves on to the next row. A slice that fails because the model was not capable enough may retry one step up a short published list of models, never more than one step, and never past a model I named by hand.

Shipped means the register says so

The runner does not take a session's word for it. A slice counts as shipped only once its row is in the shipped register. Each slice commits its work on this machine as it goes.

The phone hears about it

The phone gets a heads-up before the chain runs out of work, a note when a closeout step ticks, and the reason when the chain stops. A fork that needs me arrives as an urgent push. Every slice records which model and effort it ran on, so I can compare what a run cost against what it was given to work with.

R4 and R5 slices change how the system writes, approves or is set up, so they never join the queue. They need a written design before any code, and they run in a session I start and watch from Build mode in Studio. The limit is fixed in the runner's settings and guarded by tests, so raising it to R4 turns those tests red.

A Light on the Dark Factory

The factory runs dark, so Fulltrace gives me ways to switch the lights on. The board, a local page opened from Studio, shows the chain at work: the slice on the line, every round with its time and cost, and each commit as it lands. The stats page shows how much of the machine's working time the chain did. And when a run ends, Fulltrace points me at exactly what needs me: the checks to click through and the decisions waiting.

view demo →

Opens full screen. Escape closes it.

Demo available on desktop or tablet only.

Fulltrace board
The Fulltrace board: what is running now and six live numbers in the header, the chain's rounds below, the write trace and the activity feed.
0:00

Seven Steps for Every Slice

Every slice goes through the same seven steps, from a one-line docs fix to a new way of writing files. Only the design step changes with the risk.

Seven steps in a cycle. 01 select, 02 kickoff, 03 design gate, 04 implement, 05 prove, 06 validate, 07 closeout. The closeout leads back to select, because it recommends the next slice. closeout recommends the next slice 01 Select 02 Kickoff 03 Design gate 04 Implement 05 Prove 06 Validate 07 Closeout
01 Select I pick a slice, or the chain takes the next one in my queue. Its row is marked in progress.
02 Kickoff A written plan made before any work starts: what the slice will do, the rules it must follow, and the checks that will prove it.
03 Design gate Sized to the risk. Low-risk work goes straight to code. Risky work needs a written design first.
04 Implement Small commits, each one easy to undo.
05 Prove Each new check is proven to work: the code is broken on purpose, and the check must fail. A slice that adds no check skips this.
06 Validate Exactly the checks the kickoff named. They cannot be changed at the end.
07 Closeout A short written record of what happened, with a recommendation for what to do next.
01Select: I pick a slice, or the chain takes the next one in my queue. Its row is marked in progress.
02Kickoff: a written plan made before any work starts. It says what the slice will do, the rules it must follow, and the checks that will prove it.
03Design gate: sized to the risk. Low-risk work goes straight to code. Risky work needs a written design first.
04Implement: small commits, each one easy to undo.
05Prove: each new check is proven to work. The code is broken on purpose, and the check must fail. A slice that adds no check skips this step.
06Validate: exactly the checks the kickoff named. They cannot be changed at the end.
07Closeout: a short written record of what happened, with a recommendation for what to do next.

↺Step 07 leads back to 01, because each closeout recommends the next slice.

I decide

Which slices run and in what order, approvals, and what happens to each review finding. Since August the AI settles routine questions itself, under a standing instruction of mine, and records each ruling as mine. A question that needs me stops the slice, and in the chain that slice is set aside so the queue keeps moving.

The AI does the work

It writes the kickoff, builds the slice, runs the checks, writes the records and suggests the next step, with one clear recommendation. In the chain it also closes each slice and commits it.

A third role sits outside the pair: a separate AI whose only job is to find the holes in a design. See the adversarial review below.

Slices and Risk Classes

There is no sprint board and no ticket system. The unit of work is the slice: small enough to commit on its own, big enough to ship something real, with an ID that is never reused. The kickoff and closeout share that ID, so every slice sits between two written records. Each slice has a risk class, and the checks grow with the risk.

R0: docs only

Design notes, plans and register changes. No code.

R1: read-only code

Code that calls no model, writes nothing and changes no screen.

R2: screens that only show

Screens and views that display information and change nothing. Any new endpoint only reads.

R3: model spend

Work that costs real money. The kickoff sets a spending limit before anything runs.

R4: write or authority

Anything that can change data or how the system behaves: writes, approvals, settings. Design first, and only with me in the session.

R5: platform boundary

Opening the system to other machines or other people. A decision of its own, with everything R4 needs on top.

R0 to R2 go straight from kickoff to work. R3 must name its spending limit first. R0 to R3 can run unattended. R4 and R5 need a written design before any code, the rulebook recommends an adversarial review, and they only run with me in the session. Design and building are always separate slices: a shipped design adds new building rows at its closeout.

One File Steers the Work

The roadmap is a single markdown file with a hard rule: nothing longer than a paragraph goes in it. It changes only by editing rows and flipping statuses, and the long write-ups live in separate phase documents. That keeps it readable as a steering board.

01Now: short, dated notes on what just shipped.
02Next: suggested slices, grouped by role. The order is a suggestion. I decide.
03Decision gates: standing decisions that bind named rows. A kickoff that touches a bound row quotes its gate word for word.
04Backlog register: live rows only, each with an ID, role, risk class, prerequisites and status.
05Shipped register: one line per shipped slice, linked to its closeout.
06Standing rules: the nineteen rules, named by ID at kickoff. The full list is in the next section.
07Slice protocol: the steps, the risk classes, and the kickoff and closeout templates.
08Phase index: every track, its status, and the document that owns it.
The live and shipped registers together are the only record of which slice IDs are taken, and an exported copy never counts. Any new session, on any machine, picks up the exact state of play from the documents alone, because nothing important lives only in chat history.

Nineteen Standing Rules

These rules always apply: what can never widen, what never grants power, and what stays read-only. Each kickoff names the rules it touches by ID, and each closeout checks them. Since August each rule also has a record in code of how it is enforced: a guard that refuses, a test that fails, a checklist step, or nothing yet. Studio lists them under Config > Governance. In one sitting I ruled eighteen of them permanent and one temporary.

the_nineteen_invariantsexpand the full list
IDInvariant
G1No free models for reliability-critical audit, challenge, or apply runs.
G2No apply widening: structural, architectural, cross-surface, medium-confidence and ambiguous findings stay manual.
G3No portfolio-wide apply: every write is single-target, verifier-gated, build-gated, and ledgered.
G4No apply-write through the queue: approvals bind to an exact apply fingerprint, granted out of band.
G5The challenger stays advisory unless a separate policy decision says otherwise: a clear verdict never masks a failure.
G6The list of failure domains is not widened. It is the one list locked by the smoke tests.
G7After any server-side deploy, restart the server with the safe-restart wrapper before live validation.
G8No new platform code beyond accepted slices: new gateway endpoints are GET-only, and the few calls that change anything are named in the rule.
G9Config never grants authority: JSON cannot define new executable behaviour.
G10Secrets live in env or SecretStorage only: never in config, queues, ledgers, or reports.
G11Runtime artefacts are not source: no broad deletes over tracked report folders.
G12Every slice can be committed on its own, compiles, and passes the smoke tests where the server is touched.
G13Proposal-first: no model output is applied without deterministic validation and human review.
G14Screens that observe stay read-only, apart from confirmed actions the rule names one by one. It has been widened five times, each time by amending the rule before any code.
G15A standing fence over hosted, multi-user, write-surface, and queued-apply expansions: opening it is its own decision.
G16Every grant record carries originating-surface and grantor-auth provenance from first write.
G17Decision-quality telemetry is descriptive only and never gates.
G18A design gate that refuses a real operator need must name the lawful alternative in the same document.
G19Bytes leave the machine only through registered egress seams. A message to a person needs a single-use grant minted by hand; social posting is never allowed.
Alongside the rules sit thirteen decision gates: standing design decisions, each owned by a document, that bind named rows. A bound row cannot start without quoting its gate word for word. Two examples: memory can inform a decision but never make it, and an approval from my phone creates the same grant, by the same path, as one from my desk.

Written Records at Fixed Points

AI work rarely crashes. It trails off. So every slice produces written records at fixed points and in a fixed shape, which turns "done" into a claim I can check.

Kickoff record

Written before work starts: what the slice will do, what is in and out of scope, the rules it must follow (quoted by ID), how success will be checked, the named checks, and a spending limit when models are involved. The closeout is judged against exactly this.

Closeout record

Short by rule, about 200 words: the outcome (shipped, partial or abandoned), the commits, what changed, and each check with its result. Two fields can never be left out. Manual check says what to click or run to see the result myself, and User docs says what changed in the user guide. A slice with nothing to show writes N/A and says why, and a slice that changed Studio's code cannot use N/A for Manual check.

After the closeout

Nine fixed steps: file the record in its phase document, update the phase index, the user guide, memory and the changelog, then move the row to the shipped register, tidy the queue, commit, restart the server if it changed, and check that the click pass can be built. The row moves near the end, because updating the docs is part of the slice. New rows are added the moment they exist, including every deferral. Since August a command drafts the four follow-up entries from the closeout record, and marks every sentence that needs me.

Folding in review findings

The reviewer's verdict is pasted into the reviewed document word for word, and every finding gets a written decision: fold it in (with the edit), reject it (with the reason), or defer it (with the row that now owns it). The AI suggests each decision. I make it.

The Adversarial Review

For risky work, the rules recommend a review of the design before any code. The session that wrote the design starts a second AI with read-only access and one fixed instruction: break the design. It sees none of the first session's reasoning, so it cannot share its blind spots. Codex is the first choice. When Codex is out of quota, which the rules expect, the review steps down a short list: a manual upload to Microsoft Copilot, another vendor's command-line tool or a one-off script over OpenRouter, and last, a separate session from the same vendor. The switch is always recorded.

Keeping it independent
01A fresh start: the reviewer gets the document and its named sources, and none of the author's reasoning.
02Break it: the prompt asks for the strongest way to break the design. A review with no findings must list what it tried.
03A fixed prompt: the prompt is a recorded template, so the AI that wrote the design cannot soften the review by choosing the words.
04On the record: the reviewer's model and tool are written into the reviewed document.
What comes back

Every review returns a verdict (GO | CONDITIONAL-GO | NO-GO), findings ranked most severe first, answers to the document's open questions, and a list of what it tried and failed to break. The verdict does not decide on its own. A NO-GO whose findings I reject can still ship, and the record says so. What the reviews have returned is on the codex page.

The first live run was on 10 July 2026, on the design for chat and model routing. Two Codex rounds both returned NO-GO, with 13 findings between them. Twelve were folded into the design and I rejected the other one. A Microsoft Copilot pass then accepted the folded document.

Three Backlogs

Not everything needs a risk class. Work waits in the backlog that matches how formal it needs to be.

Backlog register

The real work queue, inside the roadmap. Every row has an ID, a role, a risk class and its gates, and every slice starts from a row. The roadmap page shows it in full.

UI fix register

Small fixes to how Studio looks and behaves, with no ceremony: an entry goes in whenever something is spotted. Screenshots attach by file name, so an entry never needs editing to gain an image. Fixes launch from Build mode one at a time or as a batch.

Docs backlog

Changes that never reached the docs, found by comparing the git log with the changelog, then worked as ordinary doc updates.

Build Mode in Studio

The whole workflow shows inside Studio, my VSCode extension. Since 1 September the panel has two modes: Operate, for using Fulltrace, and Build, for building it. Build has eight tabs (Slices, Register, Build log, UI/UX, Inbox, Gates, Reports and Ideas), and they read the live documents straight off disk.

From the Slices tab I kick off a slice in the coding assistant I pick, which marks its row in progress, and I start or stop the unattended runner and set its tier. Build also saves UI issues and new ideas. It grants no new power, and the session still does the work. Most ceremony ticks are set by a file or a commit, so they show that a step really happened, and no tick can start or stop anything. The panel is toured on the studio page.

Why This Shape

01
More checks for more risk

Docs-only work moves at the speed of conversation. Work that spends money must set a limit first, and work that writes or grants power needs a written design before any code. The risk classes make that a rule, so it never rests on a judgement made in a hurry.

02
Every slice states its outcome

Every slice ends with a clear outcome and a recommended next step. The Manual check field means I can always test a claim myself, and an N/A has to come with a reason.

03
The documents hold the state

The documents hold the whole state of play. The AI's memory points at them and never replaces them. A new session on a new machine can pick up from the repo alone.

04
A second AI checks the design

An author reviewing its own design shares its own blind spots. A fresh reviewer with a fixed prompt does not. Disagreements are recorded where I can see them, and I make the call.

Taken together, this is a software factory in the plain sense: standard steps, checks that can stop the line, and every slice leaving with its record attached. Two limits apply. The rules are only partly machine-readable: the nineteen rules have a register in code, the steps and checklist are a file per project, and a workflow I write down must be accepted by its hash before it can run, but the kickoff and closeout records are still prose, and the rest is on the roadmap. And the line runs one slice at a time. Once I have queued the work it runs dark, and I come back to rule on anything it set aside, read the closeouts and click through the checks.
See what it produced

The build log is the record of this workflow in action, in order. The roadmap is its live steering board.

Read the build log →