Skip to content

About

Production-grade agent skills for Claude Code, Cursor and Codex: evidence research, product ops, QA, release gates and decision workflows.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

1 watching

Forks

Latest commit

 

History

111 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CometWeb Agent Skills

Tested Agent Skills for research, product decisions, QA and release gates, installable in Cursor, Claude Code, Codex and other hosts that read SKILL.md packages.

A request passes through a specialist skill to a structured result; tools and approvals stay with the host.

Validate MIT Skills

This repository contains 32 reusable skill packages. Each one is a SKILL.md entry point plus the references, scripts and eval cases it needs, and every package is covered by the test suite. Skills supply instructions and output contracts. They do not supply a model, connector accounts or permission to act: browser access, connectors and external side effects stay with your host and your authorization.

Quickstart

Clone once, then run the installer for your host. Installers symlink each package from skills/ into the host's skills directory, so keep the clone where it is; git pull updates every installed skill in place.

git clone https://github.com/CometWeb-io/agent-skills.git
cd agent-skills
./scripts/install-all.sh      # all six hosts, previewed before anything is written
./scripts/install-claude.sh   # or one host: install-cursor.sh, install-codex.sh,
                              # install-qwen.sh, install-qoder.sh, install-lingma.sh

A successful run ends with a line such as OK: 32 Claude Code skills installed in /home/you/.claude/skills. Installers are safe to rerun, stop before changing anything when a target path is already taken, and accept --dry-run and --uninstall. INSTALL.md lists each host's target directory and override variable, conflict backups, upgrades, and the plugin marketplaces for Claude Code, ChatGPT and Codex.

Claude Code can install the same skills as a plugin instead, without a clone:

claude plugin marketplace add CometWeb-io/agent-skills
claude plugin install cometweb-agent-skills@cometweb-agent-skills

Then ask your assistant:

Use product-operator to compare this repo with the roadmap.
Give me the three most useful next actions and how to verify each one.

Pick a starting point

You need to… Start with You get
Check a claim evidence-researcher Sources, contradictions and an Evidence Pack
Decide what to do next product-operator A bounded now / next / stop list
Inspect an app web-app-auditor Findings tied to observed behavior
Review a release release-readiness A verdict for a specific candidate
Combine specialists skill-orchestrator Ordered steps and structured handoffs

For isolated specialist runs, use skill-orchestrator-multiagent.

Skill catalog

32 skills. Versions come from each package's VERSION; the compatibility matrix lists host support.

Foundation and orchestration

Skill Use it for Version
ai-council Evidence-governed decisions, risk gates, forecasts, and GO / TEST / DEFER verdicts. (runs only when named) 5.2.4
cometweb-context Fresh, provenance-aware context snapshots before work that depends on current project state. 1.5.2
evidence-researcher Claim decomposition, source verification, falsifiers, contradictions, and Evidence Packs. 1.0.6
portfolio-operator Cross-project focus, capacity conflicts, and pause / delegate decisions. 1.2.4
skill-orchestrator Multi-skill workflows with ordered steps and CW-AIP handoffs. 1.1.5
skill-orchestrator-multiagent Isolated subagent execution for multi-skill workflows. 1.1.6
benchmark-curator Benchmark corpora, holdouts, contamination controls, and revision hashes. 1.7.4
feedback-integrator Recurring failure patterns, improvement proposals, and regression tests. 1.8.0
quality-loop-operator Briefing, review, repair, acceptance, measurement, and quality lifecycle control. 1.7.4
rubric-designer Observable evaluation criteria, evidence floors, blocker rules, and rubric locks. 1.7.4
skill-auditor Skill routing, portability, package hygiene, and supply-chain audits. 1.7.4
skill-evaluator Fair skill experiments, behavioral lift, resource cost, and host comparisons. 1.7.3

Product, research, and partnerships

Skill Use it for Version
ai-humanize Natural English and Polish rewrites that preserve meaning and voice. 2.6.1
competitive-intelligence Competitor watchlists, change detection, and recurring delta digests. 1.1.2
design-partner-finder Finding, qualifying, and managing design partners and early adopters. 1.2.3
product-operator Weekly product control loops, roadmap drift, and now / next / later / stop actions. 2.4.2
product-teardown Evidence-backed product, UX, architecture, and implementation pattern analysis. 1.2.1
repo-to-roadmap Whole-project baselines, gap inventories, dependencies, and target-state roadmaps. 1.1.1
brief-architect Explicit artifact contracts, evidence policies, risks, and acceptance criteria. 1.7.4
content-writer Evidence-aware reader-facing articles, guides, reports, and documentation. 1.7.4
content-reviewer Constructive editorial QA with evidence-backed, actionable findings. 1.7.4
content-roaster Adversarial content review, proof debt, objections, and repair verification. 6.1.3
science-roaster Reviewer #2-style critique of methods, inference, validity, and reproducibility. 6.1.3
repo-roaster Adversarial repository review with invariants, reachability, and repair contracts. 6.1.3
repair-operator Minimal dependency-aware repairs and fresh verification of closed findings. 1.7.4
artifact-acceptance Final evidence-backed acceptance gates for knowledge artifacts. 1.7.4

Publication, operations, QA, and release

Skill Use it for Version
customer-ops Support triage, incidents, account risk, commitments, and engineering handoffs. 2.2.1
ebook-publisher Research-backed ebooks, white papers, workbooks, and publication QA. 1.0.3
longform-publisher Canonical long-form manuscripts and release-ready derived documents. (frozen) 1.1.4
release-readiness Candidate-bound production gates and GO / GO_WITH_CONTROLS / NO_GO / DEFER verdicts. 1.3.2
seo-geo-aeo-maxxing Multi-pillar SEO / GEO / AEO visibility audits. 1.3.4
web-app-auditor Evidence-driven click-through QA for websites and web applications. 1.4.2

What is inside a skill

skills/<id>/
  SKILL.md        front door: frontmatter description (what routes it) + workflow
  VERSION         semantic version, mirrored in registry/skills.json
  references/     detail the host reads only when the skill opens it
  scripts/        deterministic kernels, validators, run_evals.py
  evals/, tests/  cases that run under pytest
  agents/         generated host adapters (do not hand-edit)

registry/skills.json is the source of truth for descriptions, versions, ownership boundaries and routing signals. The Cursor routing rule, the compatibility matrix and the catalog above are generated from it. A passing validator checks a contract; it does not prove that an AI answer is correct.

How skills hand off: CW-AIP

Skills stay standalone, but when one feeds another the handoff is a typed JSON envelope rather than prose. That is the CometWeb Agent Interchange Protocol.

question ──▶ evidence-researcher ──▶ EvidenceEnvelope ──▶ ai-council ──▶ DecisionEnvelope
web-app-auditor ──▶ FindingEnvelope[] ──▶ release-readiness ──▶ release verdict
repo-to-roadmap ──▶ roadmap baseline ──▶ product-operator ──▶ handoff to a specialist
  • Small core, typed payload. Every envelope carries identity, producer, as_of time, dependencies and a payload with its hash; each payload kind has its own JSON Schema.
  • Producer truth. Only the emitting skill sets producer and timestamps; a consumer that disagrees emits a new envelope instead of editing one.
  • Status is local. Verified for research is not verified for a decision, which is not verified for a release. Consumers may apply stricter gates.

protocol/cw-aip-v2/ is current; protocol/cw-interchange-v1.md stays valid for skills that have not migrated. skill-orchestrator runs multi-step flows and passes the envelopes between steps. Check an envelope with:

uv run python tooling/validate_envelope.py --final fixtures/cwaip-v2/evidence-final.json

Write and evaluate a new skill

You need uv and Python 3.12+.

uv sync --group dev
uv run python tooling/new_skill.py my-skill \
  --description "80-1024 characters: what the skill does, when to use it, and when not to" \
  --summary "One line for the README catalog."
uv run python tooling/check_all.py --fast

The scaffold is registered and passes every fast gate straight away, but its content is placeholder. CONTRIBUTING.md walks through replacing it, making the evals and routing cases real, and bumping the plugin version. Deterministic evals prove a package keeps its contract; whether a skill makes a model behave better is a separate experiment, described in docs/CONTRACT-TRACE-AND-SKILL-EVALS.md and run with the skill-evaluator skill.

Contributing

Read CONTRIBUTING.md for the check flow and the quality bar, and SECURITY.md for reporting a vulnerability. The documentation index lists everything else, from how routing works to evaluating and releasing.

MIT license · Changelog

About

Production-grade agent skills for Claude Code, Cursor and Codex: evidence research, product ops, QA, release gates and decision workflows.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages