All posts

I want my coding agents to work together

I have subscriptions to Codex and Muse. I want them to work on different parts of the same change, share what they learn, and bring the result back together. That sounds simple. The interesting part is everything around it.

Sprowt Harness · Design notes · September 2026

I’m designing Sprowt Harness, a local tool for coordinating coding agents. You would open a repository, run the harness in your terminal, and start a feature or a fix. Several of those conversations could run at once. Each could use Codex, Muse, or both.

The first place I want to use it is Sprowt, the finance product I’m building. The harness itself is independent of that product. It should work with other repositories, and I’d like to open source it as it develops.

This is the proposed architecture. There is still work to do before I can claim the full system works.

Why build this?

Partly because I want to build faster. If a frontend change and a backend change can move independently, I want to let them. I also want the agents to question each other’s assumptions and check the combined result.

And partly because I want to learn. My background is in distributed systems. Queues, state, coordination, recovery: those are familiar problems, but adding models makes the boundaries interesting again. What happens when I change the requirements halfway through? Which worker needs to know? What survives a restart?

More agents won’t automatically mean better code, or even less waiting. I want to measure time to a verified change, how much rework integration creates, and how often I need to step in. Those are the outcomes I care about.

Explore the architecture

A coding harness is the software around a model that gives it tools, context and a way to keep working. Codex and Muse already provide that for individual agents. Sprowt Harness would coordinate several of those agents across a project.

Start with the overview below. Select a box for its explanation, then open that component. The selector also takes you directly to any of the ten detailed diagrams. Each arrow names what moves between the steps.

Proposed architecture

How a request becomes a change

A Rust coordinator uses Laya’s recommendations to help direct the work. Codex and Muse build in parallel and exchange messages through a shared mailbox. Open a component to see how it works.

Connections Work or information Messages both ways Recommendation
ComponentsLocal serviceIsolated workerSaved informationExternal service

The mailbox is available throughout a change. It carries questions, replies and artifact references; source changes remain in Git. Multiple code mods can run at once, each with its own mailbox.

Read the connections as text
  • Open your project sends to Coordinate the work: your request.
  • Coordinate the work sends to Build in separate workspaces: assignments.
  • Build in separate workspaces exchanges messages with Share a mailbox: messages + replies.
  • Share a mailbox sends to Check and combine the changes: artifact references.
  • Check and combine the changes sends to Update what the project knows: confirmed merge.

Shared services

These support the main path. Open one to see its own diagram.

A few useful terms

Code mod
One feature or fix, with its own conversation, queue, plan and workers.
Worker
A Codex or Muse session assigned part of the work.
Worktree
A separate checkout of the repository. Workers can edit without overwriting each other’s files.
Sandbox
An environment that restricts what a worker’s code can access and run. A worktree alone does not provide that protection.
MCP
Model Context Protocol: a common interface through which agents can discover and call tools.

Parallel work needs a place to meet

A code mod could assign an API change to Codex and the interface to Muse. Each worker would have its own workspace. They would exchange questions, replies and artifact links through a shared mailbox exposed by MCP tools. The mailbox would keep sender identity and recipient information, while each worker keeps its own model conversation. The two-way arrows in the overview and Worker execution diagrams show this connection.

A provider adapter would deliver each message into an active turn when supported, or queue it for the next turn. Delivery tracking would stay separate from whether the worker understood and acted on the message. Code and test reports would remain in Git and artifact storage; messages would carry references to them.

I want to be able to queue messages, edit or reorder the ones that haven’t been delivered, and steer work already in progress. A steering message needs to reach the relevant workers with my original wording intact. “Delivered” must remain distinct from “understood and applied.”

Their contributions would meet in an integration workspace. That is where the combined change gets built and tested. A queue for each project would check it against the current target branch and current requirements before merging. If the combination fails, that becomes repair work.

Skills and settings would have two scopes: shared harness defaults and project-specific instructions. Workers would receive a recorded snapshot of the instructions they used. A shared MCP gateway would expose permitted tools to both agents and keep a record of their calls.

What stays on my Mac?

The current design uses Rust for the coordinator, Tokio for asynchronous work, and SQLite for queues, events and recovery state. A background process would keep track of work when I close the terminal, then resume from saved state after a restart. Execution would pause while the Mac is asleep or powered off.

For isolation, I’m looking at Apple’s open source container project, which runs Linux containers in lightweight virtual machines on a Mac. E2B’s open source runtime is another useful reference, but the requirement here is a sandbox I run locally. The sandbox backend should be replaceable.

Account credentials would stay with a trusted service on the host, backed by Keychain. Workers would receive narrowly scoped access to that service. The broker still needs strict limits on what a worker can ask it to do: keeping a credential out of a file does not by itself prevent misuse.

Local orchestration also doesn’t make Codex or Muse local models. Their model requests still go to their providers. The sandbox, project state and small Laya model would run locally.

Memory should point back to something

Project setup would build an initial map from source code, documentation, build commands and existing instructions. Human-readable project notes would sit alongside a searchable index. A useful memory should say where it came from and which version of the project it describes.

After a merge, the harness would first mark affected knowledge as stale. It would then use the actual merged code and test evidence to update the notes and indexes, publishing a new knowledge revision together. Active workers would be told about relevant changes without pretending those changes already exist in their older checkouts.

A small model for the frequent decisions

The coordinator would be a Rust service assisted by Laya, its local decision model. Rust would manage queues, worker sessions, delivery and permissions. Laya would recommend how to route messages and classify events. Codex or Muse would handle deeper planning and task decomposition.

I want to try Laya in five places: choosing recipients for steering messages, ranking context, classifying candidate memories, deciding when background monitoring deserves attention, and classifying failures.

These are proposed uses, each needing its own evaluation. Laya would make bounded recommendations from small inputs. Rules would still enforce permissions. Codex or Muse would handle deeper investigation and draft knowledge updates. Low-confidence or invalid results would fall back to rules or a coding agent.

The attraction is a local model that can help with decisions we make often. Whether it saves time while making good enough decisions is something I need to test, not assume.

The first useful test is small

One repository. One code mod. Two workers making changes that need to fit together. Then a steering message, a failed test, a restart, and a merge that updates project memory.

I’ll start there, including validating the agents’ login, session and sandbox behavior. If that loop works reliably, we can add concurrency and see where it actually helps. I’ll write about what holds up and what needs to change.

Projects behind the design