Codex is OpenAI's coding assistant. In Fulltrace its main job is reviewing designs: when risky work needs a second AI to try to break the design, Codex is asked first. It connects to the same local server as my other assistants, but only to the read-only tools, so it can read my notes and files and cannot change them. For pay-per-use work, three role profiles send planning, coding and longer jobs to low-cost models on OpenRouter.
Before risky work is built, Codex is handed the design and told to find the holes. The reviewer is a different model from the author, and a different vendor where possible, so it does not share the author's blind spots. When Codex hit its quota in August, the review fell back to a model from the same vendor as the author, and the record says so.
The verdicts change what gets built. The design for writing to my local model fleet drew a NO-GO with three blocking findings before any code existed, and all three were folded in: a ledger write was reordered to close a crash window, tag checks moved out of the webview, and a guard now stops a double submission.
For the local-model design, even the plan for writing the document was reviewed first: two rounds and sixteen findings, all folded in. The document itself then drew two more rounds and thirteen findings, and the second round raised nothing already seen. A round that raises nothing new is my signal that a design is ready.
A finding is a case to answer. One review in July returned NO-GO with nine findings. I rejected three of the blocking ones after weighing them, still made the small changes they prompted, and took the rest. Another review had twelve of its thirteen findings folded in. The reviewer has no veto.
The rulebook records this review as recommended rather than mandatory. On 21 September an R4 slice shipped unattended without one, because an unattended chain cannot run it. The same day, R4 and R5 work came out of the unattended queue, so risky slices now run with me in the session, where the review happens. The rules live on the factory page. This page records what the reviews actually returned.
There is no separate server for Codex. It connects to the same local server as Claude Code and Cursor, over HTTP, with the same login token. It does not get the second token that allows changes, so the server refuses any tool that writes, commits, pushes or installs. Codex can read my notes, list and read files, and search the docs.
Registered in ~/.codex/config.toml under [mcp_servers.agentOS-portfolio].
Stateless StreamableHTTP, the same endpoint and transport as Claude Code. The server must be running under pm2 before Codex can connect.
The server has fourteen tools. The six read-only ones are open to Codex. The eight that change things need a second token, which the Claude Code and Cursor setups carry and Codex does not. The file tools only work inside a fixed list of allowed folders.
| Tool | What it does | Codex |
|---|---|---|
| list_portfolio_files | Lists the files in AgentOS/context-portfolio/, with their metadata. | yes |
| read_portfolio_file | Reads a portfolio file by name or portfolio:// address, straight from disk, so it is always current. | yes |
| agentOS_fs_list | Lists a folder inside the allowed roots. Every file tool refuses paths outside them. | yes |
| agentOS_fs_read | Reads a file inside the allowed roots. | yes |
| agentOS_fs_search | Searches for text across files inside the allowed roots. | yes |
| agentOS_docs_search | Searches a local index of the project docs by meaning, at $0 a query. Every answer says how much of the index it searched, and it checks the index is current on every read. | yes |
| agentOS_fs_write | Writes a file inside the allowed roots. Used where plain code makes a write that a model proposed. | no |
| agentOS_fs_replace | Replaces an exact snippet in a file. If the snippet no longer matches, the edit fails and the file is left alone. | no |
| memory_write | Writes an assistant memory note into the repo's memory folder, and checks the path again at the moment of writing. | no |
| agentOS_git_commit | Stages and commits everything in a given folder, with the message supplied. | no |
| agentOS_git_push | Pushes a given folder, and refuses any remote that is not on github.com. | no |
| agentOS_ext_compile | Runs npm run compile for the VSCode extension. | no |
| agentOS_ext_package | Runs npm run package to build a .vsix file. | no |
| agentOS_ext_install | Finds the newest .vsix and installs it with code.cmd. | no |
The repo keeps a managed block of Codex config. The sync script puts it between two markers in ~/.codex/config.toml instead of replacing the whole file, so settings for this machine outside the markers are left alone.
The server must be running under pm2 before Codex can use these tools. If it is not, every tool call fails with a connection error. Run pm2 status before starting a Codex session.
The Codex command-line tool is free, open-source software. The managed block adds OpenRouter as a pay-per-token provider, so one key reaches DeepSeek, Qwen and Kimi. Three role profiles match each kind of job to a model, and I pick one with codex --profile <name>. The key lives in the OPENROUTER_API_KEY environment variable, never in the repo or its config. Codex can also run on my OpenAI subscription, and that is how the design reviews run by default.
The three roles are a planner for thinking a job through, a workhorse for everyday coding, and a third model for long multi-step jobs. Each profile is its own file, which the sync script copies into place. A fourth, free profile was retired on 2 September, after OpenRouter withdrew its model and the profile started failing.
| Profile | Model | Role |
|---|---|---|
| planner | deepseek/deepseek-v4-flash | Planning and reasoning. 1M context at a fraction of frontier prices. |
| tools | qwen/qwen3-coder | The workhorse for coding with tools. 1M context. |
| agentic | moonshotai/kimi-k2.5 | Long multi-step jobs. 262K context. |
Studio reads the log every Codex session writes, so OpenRouter usage shows up there too, with token counts and cost estimates per model beside Claude's.
Codex keeps its memory as plain files. The shared memory that Claude Code and Cursor use (the Claude page maps it file by file) is copied from clients/shared/memory/ to ~/.codex/memories/ on every sync.
Studio's Memory tab reads the same folder, so Codex notes appear beside Claude's, each with a Codex badge, and I can see every assistant's memory without leaving VSCode.
A PowerShell script keeps Codex in step with the repo. It pushes the config, profiles, memory and commands out, pulls memory back to inspect, and reports when the two have drifted apart.
| Command | What it does |
|---|---|
| push-runtime | Deploys the managed Codex block (between its markers), the three profile files, AGENTS.md, the shared memories from clients/shared/memory/ and the shared commands to ~/.codex/. Each command's placeholders are filled in for Codex as it deploys. |
| pull-runtime | Copies the runtime memories back to clients/shared/memory/ and the managed block back to config-fragment.toml, for inspection. Commands are not pulled back. |
| clean-runtime | Removes the Fulltrace-managed Codex state from ~/.codex/ and redeploys the repo copy. Used after a restructure or to clear drift. |
| check-sync | Reports drift between the repo copy and ~/.codex/, and exits non-zero when something is out of step, so it can run before a session or in CI. |
This setup is built for one trusted workstation.
The server listens on 127.0.0.1 only, so other machines on the network cannot reach it.
The tool endpoint and, since 5 September, the telemetry endpoints all need the token in ~/.agentos/token. The sync script keeps the AGENTOS_TOKEN environment variable in step with it.
A pull copies ~/.codex/memories back into the repo, so I review those files before committing, in case they hold client data, secrets or details of this machine.