Claude Code, Codex, OpenCode, Antigravity and more in the same project, sharing one memory, running in sandboxed cells, with a record of everything they did.
curl -fsSL https://raw.githubusercontent.com/Zen1th53/marshal/main/install.sh | shQuick start · Features · Security · Platforms · Built with MARSHAL · Docs · Limitations
The MARSHAL workspace: your agent team, one keypress per agent, and a status line that always tells the truth.
You run Claude Code in one terminal and Codex in another. Each keeps its own history, so neither knows what the other changed. When a session ends, its reasoning ends up in scrollback you will never read again. Both agents also edit your repository with your full privileges, and nothing sits between them and the disk.
MARSHAL sits between them. It is a local control plane that runs agents in sandboxed cells, records what they actually do, and lets each agent build on the others' work.
|
Every conversation, tool call and diff is saved automatically and can be searched across agents and across sessions. |
Codex starts out knowing what Claude just changed, and Claude knows what Codex changed. One shared channel keeps parallel sessions up to date. |
Agents run in sandboxed cells on their own worktrees, and secrets are redacted. If a boundary can't be enforced, the run stops. |
Everything stays on your machine. MARSHAL uploads nothing, and nothing runs unless you ask for it.
| Running agent CLIs by hand | With MARSHAL | |
|---|---|---|
| Several agents in one project | Separate silos | One workspace |
| Memory across sessions and agents | No | Yes, automatic, searchable |
| Agent B knows what agent A changed | No | Yes, briefing and a shared channel |
| Sandboxed execution | No | Yes, Bubblewrap cells |
| Isolated Git worktree per run | No | Yes |
| Secrets kept out of stored history | No | Yes, redacted before writing |
| Verifiable record of what changed | No | Yes, content-addressed evidence |
| Native agent UX (your config, auth, skills, MCP) | Yes | Yes, unchanged and not proxied |
marshal tui| Key / command | What it does |
|---|---|
F8 · /claude |
Open a native Claude Code session |
F7 · /codex |
Open a native Codex session |
F9 · /opencode |
Open a native OpenCode session |
F12 · /agy |
Open a native Antigravity session (agy) |
/claude continue · /codex continue · /opencode continue · /agy continue |
Pick up where the agent left off |
/opencode resume <id> · /opencode fork <id> |
Resume or fork a specific OpenCode session |
/codex new · /resume · /codex cli |
Start fresh, resume, or open the plain CLI |
F1 Help · F2 Review · F3 Diff |
Help, review, and the working-tree diff viewer |
F4 Status · F5 Models · F6 MCP |
Runtime state, models, MCP servers |
F10 · /update |
Check for a newer release, and install it after verifying its checksum |
Ctrl+P |
Fuzzy command palette over every capability |
/goal <outcome> |
Set the session objective shown in the header |
Sessions are native. Claude Code runs as Claude Code, with your configuration,
authentication, skills, MCP servers, plugins and its own permission prompts.
MARSHAL doesn't proxy the provider, rewrite prompts, or get between you and the
agent. Codex, Claude Code, OpenCode and Antigravity (agy) run
as native sessions, and an adapter also ships for Gemini CLI.
You can also skip the workspace and launch a native session straight from the shell. Any extra arguments go to the agent unchanged:
marshal codex
marshal claude
marshal opencode
marshal agyThe Team panel shows the real status of every harness. Each binary is
actually probed, so a missing agent shows as UNAVAILABLE and is never faked.
Once you leave an agent, switching to another takes one keypress. With two terminals open, two agents can run at the same time in the same repository. They share one project memory, and each keeps its own import state, so they don't interfere with each other.
Keyboard-first workspace details
- Contextual autocomplete:
/for commands (fuzzy, e.g./rb→/rollback),@for live agents,#for claim, evidence, task and checkpoint IDs. Tabnever submits. It only completes.Enterruns the command.- Safe paste: pasted newlines never execute anything. Large pastes collapse to
[Pasted text #1 +42 lines]and are restored exactly when you submit. - Diff viewer: colored unified diffs with secret redaction. Use
n/pto move between hunks. - Real terminal app: the alternate screen repaints in place, and your scrollback is restored exactly as it was when you exit.
- Safe interrupt:
Ctrl+Ccloses overlays, then clears the input. Only a second press exits. - Status line: shows path, branch and cleanliness, working mode, session state, active agents and budget.
See the workspace guide for the full command reference.
You never have to save anything. From the moment an agent starts, MARSHAL writes to the project database every two seconds, and again when the agent exits. OpenCode and Antigravity sessions are imported automatically when the session closes:
- the conversation: what you asked and what the agent answered
- every tool call: the commands it ran and the files it opened
- the code it wrote or deleted, stored as a complete diff rather than a summary
You can then search it across agents and over time:
/memory search watcher # in the workspace
marshal memory recall ... # or from the shellNever stored: the model's hidden reasoning, or any credential. Secrets are redacted before anything is written. Tool output is size-limited so one huge dump can't crowd out the rest of the record, and truncated output is always marked as truncated.
Records are kept as observations, not verified facts. MARSHAL shows where each one came from, so an agent's guess is never presented as settled fact.
This is what you can't get by running the CLIs yourself.
When an agent starts, it gets a briefing built from what the other agents have already done in the project: recent sessions, the commands they ran, and the changes they made. Codex starts out knowing what Claude just did, and vice versa.
A briefing is only a snapshot, so MARSHAL keeps it current with a shared
channel. Every agent drops what it does into one ordered stream as it happens,
and each agent reads its own view of it (.marshal/inbox/<agent>.md). One event
is stored once, however many agents end up reading it.
An agent that was closed still catches up, and an agent that opens late joins mid-conversation rather than being handed a summary. Each reader has a cursor, so Codex opening while Claude is five steps into a task sees those five steps — the work itself, not a paragraph about it. Entries are stamped and a boundary marks where each session begins, so older work never reads as though it just arrived. A full view drops its oldest entries rather than sealing itself, because the recent work is the part a returning agent needs.
Nothing interrupts a running agent. The native CLI owns the terminal, so delivery is a pull: the view is a file that is current whenever the agent looks, and the briefing says so rather than implying the agent is kept in sync.
The briefing and the view both label themselves as untrusted data, not instructions. They quote other agents' output, which can contain anything those agents happened to read, and nothing in them overrides you.
Two decisions, both made before the work starts, in .marshal/live-peers or
through /memory peers:
participants: claude, codex, opencode, agy
agy: all # every other agent
claude: all
codex: opencode, agy # not claude
opencode: none # contributes, reads nothing
An agent is never shown its own work, and that is not a setting. It already knows what it did — the work is its own conversation — so handing it back would be noise at best, and at worst a model reading its own output as though another agent had reported it.
Joining and seeing are separate. An agent can contribute while reading almost nothing, and that is an arrangement rather than a gap. The reason is practical: models differ in what they can use. A capable one does better seeing everything the others did; a smaller one does worse, because context it cannot follow is context it can be confused by. So the list is per reader, and the two directions between any pair may disagree — a reviewer can read the implementer without the implementer reading the reviewer.
all and none are accepted, agy is understood as Antigravity, and a line
naming only agents MARSHAL does not run is skipped rather than recorded as a
decision to read nothing.
Every agent reaches it as it works. There is no exempt provider.
| Agent | History | Read while it runs |
|---|---|---|
| Claude Code | append-only JSONL | as it grows |
| Codex | append-only JSONL | as it grows |
Antigravity (agy) |
per-conversation SQLite | read-only, under WAL |
| OpenCode | SQLite | read-only, under WAL |
The two SQLite stores are opened read-only and never written to, and SQLite in WAL mode serves readers while a writer holds the file — measured against a copy of a real store: 394 reads against a live writer, none blocked, the reader within two rows of the writer throughout 958 concurrent inserts.
Reading a store MARSHAL does not own is a coupling, and it is guarded rather than assumed. OpenCode's shape is checked before each read; if it has moved, that path stands down and the supported CLI export takes over, which cannot run mid-session and so delivers at exit — later than it should be, never wrong and never missing. What is read is assembled into the same document the CLI export produces and handed to the same decoder, so the live path cannot select different fields from the export path. In particular the model's hidden reasoning is excluded there, once, for both — checked against 199 real reasoning blocks, none of which reached a transcript.
Verified on 2026-09-22 against the installed CLIs — Claude Code 2.1.278, Codex
0.155.1, OpenCode 1.18.16, agy 1.2.7 — by decoding their real transcripts with
the same watcher the runtime uses: a 4.2 MB Claude session yielded 888 messages
and a 224 KB Codex session 18, while the same Claude file grew between two runs
minutes apart, which is what live capture looks like from outside the process.
The arrangement space is tested exhaustively rather than by example. Each
reader takes any subset of the three other agents, which with the participant
subsets is 65,536 configurations; every one is written, read back, and
checked to still mean the same thing for all sixteen author/reader pairs. Every
filter a reader can have is then rendered to a real file and read back, because
a filter that is right in the configuration and wrong in the rendering would
show a model exactly what you kept from it. See
internal/tui/native_channel_exhaustive_test.go.
Capture also reports itself while it runs. The native CLI owns the terminal for
the whole session, so MARSHAL cannot draw a counter — it writes one instead, to
.marshal/<agent>/live-status.json: records imported, entries delivered, last
sync, and any capture error.
- Sandboxed execution. Agent processes run in isolated cells with a read-only root filesystem, private runtime directories, and no network by default.
- Isolated worktrees. Each run works on its own Git worktree and branch. Your working tree is never used for experiments.
- Fail closed. If a boundary can't be enforced, the run stops instead of continuing without the policy.
- Secrets never reach storage. Credentials are redacted from stored output and never enter long-term memory.
- Nothing starts by accident. Plain text in the workspace runs nothing. If you type a command without its slash, MARSHAL suggests the command you probably meant. An agent starts only when you ask for it by name.
- Delegation must be earned. To let MARSHAL act without asking each time, the session needs a cryptographically verified entitlement. A local flag can't grant it, and the test suite includes bypass attempts that must fail.
Command output and artifacts are content-addressed and linked to the commit that produced them. When you ask "what did this run actually change?", you get an answer you can verify, not a log you have to take on trust.
- Claims and evidence:
/claims,/evidenceand/inspectshow what was asserted, what backs it up, and any contradictions. - Checkpoints and rollback:
/checkpointand/rollbacksave and restore the worktree together with its claim state. - Approvals:
/approvals,/approveand/rejecthandle high-risk actions, each tied to a specific commit. - Evidence bundles:
/exportwrites a bundle with a deterministic digest to.marshal/evidence/.
For work that needs more than a chat session, MARSHAL provides a governed pipeline. Each stage writes a versioned record tied to an exact repository state, and no command can mark a stage successful directly.
marshal goal "add rate limiting to the API" # intent, hard constraints, risk tier
marshal plan create ... # task DAG, team, verification policy
marshal exec start ... # governed run with approval gates
marshal review start ... # independent verification + attestation
marshal learning search ... # memory filtered by evidenceExiting with status zero doesn't make a run verified. Completion requires every mandatory criterion to be met and critical evidence from independent sources.
- MCP server (
marshal mcp serve): an authenticated Model Context Protocol endpoint for IDEs and orchestrators. - A2A server (
marshal a2a serve): the Agent-to-Agent protocol, for discovery and safe task delegation. - Bearer tokens (
marshal auth token create | list | revoke) protect both servers. Neither starts unless you run it.
MARSHAL is developed inside its own workspace. Its source code is written and reviewed in native Codex, Claude Code and OpenCode sessions opened through MARSHAL. Those sessions share one project memory, pick up each other's work through the cross-agent briefing, and run the same verification commands that are documented for users.
MARSHAL source repository
|
v
MARSHAL workspace ---- Codex . Claude Code . OpenCode
|
|-- shared project memory (.marshal/state.db)
|-- source changes and tests
'-- release gate and tagged source commit
|
v
reproducible GitHub release artifacts
Self-hosting does not replace independent provenance. Release archives are
built by the pinned GitHub Actions release workflow, and every commit, test run,
checksum, SBOM and provenance attestation can be inspected independently.
Commits made with an agent carry a Co-authored-by trailer naming it.
For details, see architecture, execution cells and the security model.
Linux, one command:
curl -fsSL https://raw.githubusercontent.com/Zen1th53/marshal/main/install.sh | shThe script finds the latest release, checks the download against its published
checksums, and installs to ~/.local/bin. It needs no sudo and touches nothing
outside the install directory. If verification fails, nothing is installed.
# Choose the location, or pin a version
MARSHAL_INSTALL_DIR=/usr/local/bin MARSHAL_VERSION=v0.0.3 \
sh -c "$(curl -fsSL https://raw.githubusercontent.com/Zen1th53/marshal/main/install.sh)"From a release archive
The archive and checksums.txt are on the
latest release:
sha256sum -c checksums.txt --ignore-missing
tar -xzf marshal_<version>_linux_amd64.tar.gz
install -Dm755 marshal "$HOME/.local/bin/marshal"From source (requires Go and git)
go install github.com/Zen1th53/marshal/cmd/marshal@latestcd /path/to/your/project # an empty directory works too
marshal setup # check readiness, and offer each missing step: git init,
# a baseline commit, and the project runtime
marshal doctor # check the host and probe the agent CLIs you have installed
marshal tui # open the workspaceThen, in the workspace:
/goal <outcome> say what you are trying to achieve
/claude work with Claude Code (or press F8)
...exit the agent when you are done
/codex hand over to Codex (F7); it already knows what changed
/opencode or to OpenCode (F9); its session is saved when it exits
/memory search <anything> ask the project what happened
/diff review the working tree
/mode manual the default working mode
/mode auto switch the session's working mode
/mode ultra refused without a verified entitlement
/ultra why this session is Standard, and how to ask for more
manual and auto are session preferences. They are recorded and shown in
the status line, but on their own they don't allow anything to act without you.
Autonomous delegation, where MARSHAL decides instead of asking each time,
requires a cryptographically verified entitlement. That check runs on every
attempt; it isn't cached at startup. A local setting can't grant it, and
/mode ultra without an entitlement is refused. Without one, the session runs as
Standard and keeps asking you, which is intended behavior and not a degraded
mode.
| Platform | Status | Notes |
|---|---|---|
| Linux | Released | Fully supported, with sandboxed execution through Bubblewrap. Download |
| macOS | Planned | On the roadmap. It needs a native sandbox backend first. |
| Windows | Maybe, never | No commitment yet. The Linux build may work under WSL2, but it is untested. |
| Linux | The supported platform for sandboxed execution. MacOS and Windows have no equivalent backend. |
Bubblewrap (bwrap) |
Required for execution cells. |
| Git | MARSHAL works on a Git repository. |
| Agent CLIs | Install and log in to the ones you want. MARSHAL doesn't bundle or proxy any of them. |
Run marshal doctor to see what is installed and what is missing.
These are stated plainly, because a control plane that overstates its guarantees is worse than none:
- Network egress isn't filtered per destination. Runs that need network access stop instead of proceeding without the policy.
- Sandboxed execution is Linux only.
- One agent per terminal. A native session takes over its terminal, so running two agents at once needs two terminals.
- Vector search needs an embedding provider. Exact and lexical search work without one.
- No third-party security audit. Automated suites test the security invariants, but that isn't external certification.
| Getting started | Step-by-step tutorial |
| Workspace guide | Native sessions, memory capture, cross-agent exchange, all commands |
| CLI reference | Every command |
| Security model | Threat model and isolation boundaries |
| Providers | Supported agent CLIs and configuration |
| MCP · A2A | Protocol servers |
| Troubleshooting | Diagnostics and recovery |
| Documentation hub | Everything else |
MARSHAL Community is a local, single-node runtime scoped to one project, and it is everything in this repository. Enterprise adds a remote web control plane, orchestration across multiple nodes, centralized approvals, and multi-user access control under a separate commercial license. The Community binary doesn't include a web control plane.
Issues and pull requests are welcome. Please read CONTRIBUTING.md and CODE_OF_CONDUCT.md first.
To report a security vulnerability, follow SECURITY.md and report it privately rather than opening a public issue.
If MARSHAL is useful to you, starring the repository helps other people find it.
- Community: GNU Affero General Public License v3.0 only. See LICENSE.
- Commercial: licenses without AGPL copyleft are available. See LICENSING.md.
- Historical releases up to
runtime-v0.4.0remain under their original Apache-2.0 grants. See docs/legal/LICENSE-HISTORY.md. - Third-party attributions: THIRD_PARTY_NOTICES.md.