An agent harness for the post-framework era. From Nicia — every general needs a victory title.
Six entities. No orchestration graphs. Evals first. Execution state stored as a graph using TypeGraph. Sandboxed workspace and execution graph share one SQLite database — file provenance is graph-native. Agent behavior evaluated as structural assertions on the graph.
→ Read the thesis before looking at the code.
→ Read the entity design decisions before reading the implementation.
→ Understand the storage model for how nodes and edges record execution.
→ Read the eval methodology before running benchmarks — especially the graph-based behavioral eval section.
# Install
pnpm install
# Configure environment
cp .env.example .env
# Then edit .env with your keys:
# ANTHROPIC_API_KEY=sk-ant-... (required)
# BRAVE_API_KEY=BSA... (optional — web-search uses mock if absent)
# Build all packages
pnpm build
# Run a local agent
pnpm cli run \
--definition fixtures/definitions/research-assistant.json \
--input "What will be the impact of AGI on GDP?"Caution: evals use a lot of tokens
# Run the benchmark eval suite
pnpm eval
# Run 5 times for statistical significance
pnpm eval --runs 5
# Behavioral evals — graph assertions, no LLM judge
pnpm eval --category dispatch --no-judge
pnpm eval --category hitl --no-judge
# Generate reports
pnpm eval:report # single-run summary
pnpm eval:multi-run --last 5 # cross-run variance and significance├── packages/
│ ├── core/ # Zod schemas, TypeGraph graph schema, Repository
│ ├── sdk/ # Anthropic SDK wrapper, subagent loop
│ ├── harness/ # Run loop, task dispatch, skill execution
│ └── workspace/ # Virtual bash shell (just-bash) + agentfs filesystem
├── apps/
│ └── cli/ # Local development runner
├── tools/
│ ├── web-search/ # Brave Search tool implementation
│ └── web-fetch/ # URL fetch tool implementation
├── evals/
│ ├── tasks/ # Eval task definitions (YAML)
│ ├── results/ # Eval run results and reports
│ ├── calibration/ # Judge calibration data
│ └── llm-judge/ # LLM judge prompts and rubric
├── fixtures/
│ ├── definitions/ # AgentDefinition fixtures
│ └── skills/ # Skill fixture data (seeded into graph)
├── archive/
│ └── cloudflare-worker/ # Archived Worker/Durable Object adapter
└── docs/ # Design documentation
An agent needs one file: a definition (JSON). Skills are optional — an agent with just a system prompt and tools is a valid starting point.
Create fixtures/definitions/my-agent.json:
{
"id": "a1b2c3d4-...",
"version": 1,
"name": "My Agent",
"description": "What this agent does.",
"systemPrompt": "You are a research assistant. Use web-search to find...",
"skills": [],
"limits": {
"maxTasksPerRun": 20,
"maxOperationsPerTask": 3,
"maxTokensPerRun": 100000
},
"createdAt": "2025-01-01T00:00:00.000Z"
}The systemPrompt is what the model sees. Every agent gets web-search,
web-fetch, and bash (sandboxed workspace) as built-in tools. limits
control the run's resource envelope. Optional workspace config can seed
initial files and declare outputPaths globs for auto-capture.
pnpm cli run \
--definition fixtures/definitions/my-agent.json \
--input "Your task here"The CLI prints an execution tree on completion showing tasks, operations,
and artifacts. State is persisted to nicator.db (SQLite, created in the
working directory).
Skills give an agent reusable, scoped reasoning capabilities — each skill
runs its own LLM loop with the parent's tools. Create
fixtures/skills/my-skill.json:
{
"name": "my-skill",
"version": "1.0.0",
"description": "One sentence the coordinator sees when deciding to invoke this.",
"maxIterations": 10,
"prompt": "# Skill Prompt\n\nYou are a specialist. Use web-search and web-fetch to..."
}description is what the model reads to decide when to use the skill —
make it specific. maxIterations caps the skill's tool-call loop (set to
1 for pure-analysis skills that don't need tools). prompt is the skill's
full system prompt. See docs/entities.md § Skill for
design details.
Then reference it in your definition's skills array:
"skills": [
{ "name": "my-skill", "version": "1.0.0", "policy": { "type": "always" } }
]References must match a fixture by name + version exactly. Each skill
can have a policy: always (default), never,
require_hitl_approval (human gate per invocation), or
max_calls_per_run (budget cap). If a skill has require_hitl_approval
policy, the CLI pauses for stdin approval.
The model can spawn subagents on the fly — no declarative topology
required. The systemPrompt tells it how to coordinate:
agent(name, prompt, task_input)— ad-hoc dispatch: model constructs the promptskill(skill_name, task_input)— activates a pre-registered skill with its fixture prompt
See fixtures/definitions/debate-assistant.json for a working example
that spawns advocate and judge subagents.
Each run gets a sandboxed shell (just-bash — pure TypeScript, 79+ builtins) backed by a virtual filesystem (agentfs — SQLite, same database as TypeGraph). Agents get CLI capabilities without host access. Workspace files promoted to artifacts become content-addressed, versioned graph nodes with full provenance. See docs/entities.md § Artifact and docs/graph-model.md for the storage details.
This is not a production framework. It is a reference implementation for the
local CLI harness, graph-native provenance model, workspace, skills, HITL
abstraction, and eval methodology. The previous Cloudflare Workers deployment
path has been archived under archive/cloudflare-worker/ so the mainline stays
focused and reproducible. See docs/why.md.
pnpm build # Build all packages (turbo)
pnpm typecheck # TypeScript strict check across all workspaces
pnpm test # Run tests (vitest)
pnpm lint # Lint (eslint)
pnpm fix # Auto-fix lint + format issues
pnpm check # Lint + typecheck + format check (CI gate)
pnpm dev # Watch mode for all packagesMIT