I tested the connection with a small FastAPI project. Codex chose a Python version, installed the dependencies and built a health endpoint in a Linux virtual machine, or VM. The independent checks passed. The original project stayed untouched.
I then added persistent VMs to the harness. Each feature or fix gets its own environment, with files and installed tools that survive a pause.
In the previous post, I described how I wanted Codex and Muse to work together. The version I’ve built so far uses one Codex worker per change.
Start with one change
I call a feature or fix a code mod. It has its own conversation and plan. I can switch to another mod without mixing the messages or losing what I was typing.
When I describe the change, Codex reads the project and writes a plan. I get a numbered list showing what each task should achieve and which tasks need to happen first. I can review it before starting the worker.
The harness gives Codex one task at a time. Rust tracks the dependencies and saves progress as the worker finishes each task.
Give the agent its own computer
Apple Container runs a separate Linux VM for each mod.
The VM acts like a separate computer on my Mac. The harness copies the source files into it, then Codex installs the tools it needs, edits the code and runs tests there. My original files stay where they are.
container exec --interactive \
--workdir /workspace "$vm" \
env -i HOME=/home/sprowt \
CODEX_HOME=/opt/codex-home \
PATH=/usr/local/bin:/usr/bin:/bin \
/usr/local/bin/codex exec-server \
--listen stdioWhen I reopen the harness, it restores the saved plan and progress. It waits for me to resume. Completed tasks stay done, and I can retry unfinished work.
Check the work, then let me review it
When Codex finishes a task, it returns commands for checking the result. The harness runs those commands again, inside the same VM. If a check fails or is missing, the task stays unfinished.
One Codex command interface would have run the checks on the Mac. I connected verification directly to the VM’s tool runner so it checks the environment where the worker made the changes.
After all tasks finish, the harness reruns their checks together. I can then inspect the diff, which shows the files added, changed or removed, and confirm applying it. If I’ve edited an affected file in the original project meanwhile, the harness blocks the apply.
The to-do app in the screenshot hit this boundary. Chromium downloaded after I updated the allowlist, but couldn’t start because the VM lacked Linux libraries, including GLib and NSS. The worker reported three browser checks as blocked.
What works and what’s still missing
The terminal labels my messages and the worker’s replies separately. It shows when a worker is active, along with its model and reasoning level. I can expand the plan to see file scopes and check commands.
Laya, the small model running locally, recommends a model and reasoning level for the planner. Its initial results were uncertain or wrong. The harness falls back to Sol with high reasoning when it can’t use the recommendation. I haven’t measured a benefit from routing yet.
The probe validated the connection and checked isolation from the host. I haven’t completed a security audit of the harness.
Muse, multiple workers on one mod and shared memory are still planned. Different mods can already run at the same time. Execution uses Linux, so native iOS and macOS builds need a future macOS VM backend.
A few engineering details
Rust validates the plan’s dependencies and file scopes. SQLite saves conversations, queues, drafts, plans and check results.
The harness excludes common credential files when it copies source. Downloads go through a domain allowlist set on the host. I added Playwright’s browser download domains and redirects. Allowed services can still receive data.
The harness applies files one at a time. It rolls back earlier replacements if it handles an error, but a crash can leave a partial apply. It keeps the starting copies.
The persistent backend has a separate live VM test for network policy, source export and reconnecting. Ordinary unit tests skip it because it starts a real VM.
