English · 简体中文 · 日本語 · 한국어 · Español · Français · Deutsch
A read-only, evidence-first review skill for AI agent systems. Works in Claude Code and Codex.
A Harness is the model-facing runtime and governance layer around a model: context assembly, tools, skills, MCP, state, control flow, parsing, permissions, trace, eval, recovery, escalation and human oversight. Your agent's behavior is decided there at least as much as in the prompt — but that layer rarely gets reviewed as a system.
harness-review reviews one explicit behavior end to end and binds every conclusion to evidence you can re-check.
- Traces one behavior from input and context through actions, state, validation, recovery and human judgement to the final outcome.
- Separates five kinds of evidence —
normative,implemented,exercised,observed,inferred— and refuses to let one silently overwrite another. - Reports what it could not verify. A capability that does not exist is recorded as
uncovered, never inferred as a pass. - Dispatches 2–4 independent read-only reviewers when the slice warrants it, and reports degradation when the host cannot.
- It does not modify your project. Read-only by default.
- It does not auto-trigger. Explicit invocation only — it will not fire on generic code review, QA or debugging requests.
- It is not a linter, a general code reviewer or a test runner.
Install (no build, no uv, no Python packages — the distribution is prebuilt):
curl -fsSL https://raw.githubusercontent.com/Dexter-Yao/harness-review/main/install.sh | shPrefer not to pipe a script into your shell? Do it by hand:
git clone https://github.com/Dexter-Yao/harness-review
python3 harness-review/scripts/install_local.py --install
rm -rf harness-reviewThe installer copies files, so the clone is disposable. It refuses to overwrite anything it does not own, and writes a receipt so --check and --uninstall know exactly what belongs to it.
Use it — explicit invocation, one acceptable slice at a time:
Claude Code /harness-review review the payment approval flow in src/payments/
Codex $harness-review review the payment approval flow in src/payments/
Report language follows your question; the structured field names stay English so the contract stays machine-checkable.
Reviewing the bundled fixture at evals/fixtures/agent-app — a background payment agent whose README requires human approval before any payment is sent:
[HR-001] [CRITICAL] Payment is sent with no approval record
Type: implementation
Flow node: tools and side effects
Invariant: A payment may be sent only after a human approval record exists.
Observed: run_payment_agent calls payment_tool.send(proposal) directly after a
key-presence check, with no approval lookup on any path.
Evidence: evals/fixtures/agent-app/payment_agent.py:17
Evidence kind: implemented
Source role: authoritative
Verification status: verified
Runtime mapping: mapped
Coverage: uncovered
Claim status: not-applicable
Impact: Every accepted request can move money irreversibly.
Owner: payment agent runtime
Recommendation: Require an approval record at the tool boundary, not in prose.
Verification method: A test asserting send() is never reached without approval.
It also reports what it could not establish:
### Verification Gaps
- No runtime trace was available for the current implementation. `trace.json`
exists but nothing proves it was produced by this code, so the "run completed"
claim stays `observed` for the artifact and `unknown` for the runtime.
- The only test asserts the function runs without raising; no test exercises the
approval boundary. Trace/eval coverage is therefore `partial`, not `covered`.
That last section is the point of the tool: it tells you where you are flying blind, rather than implying silence means safety.
Acceptance and behavior · Harness inventory and ownership · Parse-first and task context · Agent-readable · Tools, permissions and side effects · State and runtime control · Observability and provenance · Eval and governance · Human oversight and verifiability · Economics and adaptability
Full rubric: src/references/review-model.md. Report contract: src/references/output-contract.md.
Three phases: establish a task projection (target, mode, acceptance, owners) → investigate evidence by claim type → aggregate, bind findings to the projection, and close the loop.
Four reviewer responsibilities: Flow & Ownership · Context, Parse-first & Agent-readable · Tools, State, Recovery & Human Oversight · Trace, Eval, Observability & Economics.
| Claude Code | Codex | |
|---|---|---|
| Invocation | /harness-review |
$harness-review |
| Explicit-only enforcement | disable-model-invocation: true |
allow_implicit_invocation: false |
| Tool permission boundary | allowed-tools / disallowed-tools |
none available |
| Reviewers | 4 bundled read-only subagents | session-exposed subagents |
Be aware of the asymmetry: Codex skills have no equivalent of Claude's allowed-tools. There, read-only is this skill's behavioral contract, enforced by the session's own sandbox and approvals — not by the skill. If your session cannot guarantee read-only, the skill says so before proceeding. Neither host's boundary is an operating-system sandbox.
To share the skill with a team through the repository instead of installing it per user:
python3 scripts/install_local.py --install --scope projectpython3 scripts/install_local.py --check # compare install against the current release
python3 scripts/install_local.py --install # re-run to upgrade in place
python3 scripts/install_local.py --uninstall # remove only receipt-owned paths--uninstall never deletes a file you edited; it reports it instead.
Linux, macOS and Windows. The build is byte-deterministic across platforms: the same commit produces the same content-addressed release id everywhere.
See CONTRIBUTING.md. Everything that gates a pull request runs offline with no host CLI and no network:
uv sync && uv run pytest
uv run python scripts/build_skill.py --check