A team workflow for building prototypes with AI coding agents — spec-driven, phase-based, repeatable.
A working starter kit for our team to adopt spec-driven development with AI coding agents. It's Claude Code-first (the agent team lives in .claude/agents/), but everything it produces is plain Markdown — AGENTS.md, constitution.md, spec.md, plan.md, PROGRESS.md — so the artifacts port to Cursor, Copilot, and the rest. The automation — auto-discovered agents and /-skills — is Claude Code-specific; in other tools you run the same workflow by hand (see the quickstart). Fork it per project and you'll ship prototypes faster without sacrificing code quality or team handoff.
The mental model: you decide what and why; the agents do the typing. Nothing gets built until a spec and plan are written and approved. Work happens in small phases that each fit one AI session, and
PROGRESS.mdcarries the memory between them.
- Spec before code, always. No agent writes a line until the human and the AI agree on what's being built.
- Phases that fit one context window. Long sessions drift; short, well-scoped phases stay sharp.
- A small team of specialised agents, not one generalist that does everything.
The agents do the typing. The humans do the deciding. That balance is the whole point.
┌─────────────────────────────────────────────────────────────────┐
│ Loop 1: SHAPE (one-time per feature) │
│ │
│ 1. Constitution → 2. Specify → 3. Clarify → 4. Plan │
│ (rules) (what/why) (edge cases) (how/phases) │
│ ↑ ↓ │
│ └────────── human review at each gate ───────┘ │
└─────────────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────────────┐
│ Loop 2: BUILD (repeats per phase) │
│ │
│ 5. Tasks → 6. Implement → 7. Verify → PROGRESS.md updated │
│ │
│ ↑ human approves phase before next starts ↓ │
└─────────────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────────────┐
│ Loop 3: SHIP (once all phases done) │
│ │
│ Demo it → harden only if you're keeping it │
└─────────────────────────────────────────────────────────────────┘
ai-prototyping-workflow/
├── README.md ← you are here
├── constitution.md ← fork this per project; defines the rules
├── AGENTS.md / CLAUDE.md ← agent contract (AGENTS.md is source; CLAUDE.md points to it)
├── init.sh ← one-command fork into a new repo
├── .claude/agents/ ← 6 agents, one job each
│ ├── analyst.md
│ ├── architect.md
│ ├── planner.md
│ ├── doc-fetcher.md
│ ├── implementer.md
│ └── reviewer.md
├── .claude/skills/ ← conditional steps, invoked by hand
│ ├── visual-check/ ← brand alignment for UI work
│ └── harden/ ← refactor passes, if you keep the prototype
├── specs/_template/ ← copy this into specs/<your-feature>/
│ ├── spec.md
│ ├── plan.md
│ └── phases/phase-template.md
├── specs/example-link-checker/ ← a filled-in worked example (lean path)
├── templates/
│ ├── PROGRESS.md ← living state + final outcome decision
│ └── ARCHITECTURE.md ← generated at the end (if the prototype continues)
└── docs/
└── quickstart.md ← walkthrough of your first prototype
Six agents, each with one clear job that produces a concrete, reviewable output — a file on disk or a structured report. Two conditional steps — visual checks and hardening — are skills (/visual-check, /harden), invoked by hand when you need them.
| Agent | Job | When to invoke | Produces |
|---|---|---|---|
| analyst | Validate the user story; ask clarifying questions | Specify & Clarify | spec.md, clarifications.md |
| architect | Decide tech stack, types, state shape, folder structure | Plan | plan.md architecture sections |
| planner | Break the plan into phases that each fit one session | Plan & Tasks | Phase breakdown + phase-N-tasks.md |
| doc-fetcher | Fetch and summarise external library APIs | Before any phase with a non-trivial dependency | specs/<feature>/refs/<lib>.md |
| implementer | Build one phase end-to-end with tests | Implement | Code, tests, PROGRESS.md update |
| reviewer | Challenge the plan; spec-drift check per phase (review-only) | Gate 3, then Verify | Review notes |
Why six and not twenty: fewer, sharper agents beat more, fuzzier ones — start simple and add an agent only when a real gap appears.
The two skills:
| Skill | Job | When to run |
|---|---|---|
/visual-check |
Screenshot + brand alignment for UI work | After any UI-touching phase (frontend only) |
/harden |
Three refactor passes + ARCHITECTURE.md |
After the demo, only if you're keeping the prototype |
Both are manual-only so they never fire on their own or bypass a human gate.
These are the four moments where the workflow stops and waits for a human. Skip any of them and the workflow degrades into vibe coding with extra steps.
- Constitution approval — the rules of the road
- Spec sign-off — what we're building
- Plan review — how we're building it (expect 2–4 revisions; this is real engineering time)
- Per-phase approval — proof this phase works before the next one starts
- Fork this repo into your project.
- Edit
constitution.mdto match your stack (Python? TypeScript? Kedro?). - Copy
specs/_template/tospecs/<your-feature>/and fill inspec.md. - Open Claude Code in the repo — the agents in
.claude/agents/are auto-discovered. (On Cursor/Copilot, see Running it outside Claude Code.) - Run the loop:
analystvalidates the spec →architectplans →plannerphases it →implementerbuilds each phase (thereviewerchecks each) → demo it → run/hardenonly if you're keeping it. - Update
PROGRESS.mdafter every phase. This is the single most important habit.
New to the workflow? Read docs/quickstart.md for the exact commands and the recovery playbook for when something goes wrong. A complete worked example lives in specs/example-link-checker/ — read it to see what each artifact should look like.
Tiny, throwaway prototype? Use the lean path — same four gates, far less ceremony (skip Clarify, one phase, short PROGRESS.md). See docs/quickstart.md.