A set of skills for Kotlin, Jetpack Compose, Android development, and grounded writing.
The repository is also a portable Agent Plugins
v1.0.0 package. Conforming clients discover the root plugin.json
and the immediate skill directories under skills/.
With the skills CLI:
npx skills add chrisbanes/skills
Or install as a Claude Code plugin:
/plugin marketplace add chrisbanes/skills
/plugin install chrisbanes-skills@chrisbanes-skills
Or install as a Codex plugin:
codex plugin marketplace add chrisbanes/skills --ref main
codex plugin add chrisbanes-skills@chrisbanes-skills
Or install as an OpenCode V2 plugin (2.0.22 or newer):
{
"plugins": ["chrisbanes-skills@git+https://github.com/chrisbanes/skills.git"]
}OpenCode V1 1.18.29 or newer uses plugin instead of plugins.
See .opencode/INSTALL.md for details.
Most skills in this repository are self-contained. The workflows below compose
with skills from other repositories; installing chrisbanes/skills does not
install them, and the workflows never install them implicitly.
| Consumer | Requirement | Provider | Source and install |
|---|---|---|---|
deliver-spec GitHub readiness |
Conditional | triage |
Matt Pocock's skills: npx skills add mattpocock/skills --skill triage |
implement-with-subagents behavior tasks |
Conditional | tdd at approved test seams |
Matt Pocock's skills: npx skills add mattpocock/skills --skill tdd |
run-github-project explicit/repository-required test-first work |
Conditional | tdd |
Matt Pocock's skills: npx skills add mattpocock/skills --skill tdd |
run-github-project Todo triage |
Conditional | triage |
Matt Pocock's skills: npx skills add mattpocock/skills --skill triage |
run-github-project Wayfinder lane |
Conditional | wayfinder, research |
Matt Pocock's skills: use --skill wayfinder or --skill research; see the workflow provider matrix |
Delivery and parallel implementation use one fresh independent read-only reviewer
against the approved source and repository standards; no external review skill
or issue-tracker setup is required. Reviews include the production paths and
boundary cases examined and material uncertainty. The authoritative
behavioral review and repair procedure covers
related boundary repairs, repeated external findings and browser diagnostics.
npm run references:sync supplies identical local copies to its consumers;
edit the authoritative source, not the copies. The same mechanism supplies
to-plan's boundary-validation rule
to implementation consumers.
Review mode in implement-with-subagents, and review or setup mode in
run-github-project, do not require these external skills. See the
run-github-project provider matrix
for its lane-specific fallback and blocking behavior.
- Working on Compose state or effects? Start with
compose-state-and-effects. - Investigating recomposition, stability, or jank? Start with
compose-performance. - Comparing Android benchmark configurations or a measured Android default? Start with
android-benchmark-comparison. - Reviewing Flow or coroutine architecture? Start with
kotlin-concurrency-and-flow.
using-chrisbanes-skills— route Kotlin and Jetpack Compose work to focused skills, adding a second only for an independent decision in the same change.
android-benchmark-comparison— compare physical Android benchmark configurations with verified coverage, controlled conditions, trace-backed diagnosis of unstable rankings, and bounded conclusions.
compose-state-and-effects— decide state ownership and effect lifecycle for local UI state, screen state holders, Flow collection, callbacks, cleanup, navigation, snackbar, analytics, and focus requests.
compose-performance— diagnose stability, deferred reads, composition contracts, and cross-phase back-writing from concrete runtime evidence.
compose-component-design— design caller-placeable Compose APIs whose variable visual regions are caller-provided slots.compose-animations— choose Compose animation APIs for visibility, value targets, coordinated transitions, and content swaps; align with official quick guide and decision tree.compose-focus-navigation— design and test keyboard, TV, D-pad, and focus-first Compose navigation behavior.
compose-ui-testing-patterns— choose between plain UI tests, semantics assertions, key/focus tests, interaction state tests with MutableInteractionSource, screenshot tests, and integration tests.
kotlin-concurrency-and-flow— review coroutine, rawThread, andExecutorownership, cancellation, Flow state/event modeling, sharing, replay, and one-shot delivery.kotlin-control-flow— write and review Kotlin branching with subjectwhen, guard conditions, sealed exhaustiveness, smart casts, nullable branching, and early returns.kotlin-api-design— choose function owners, semantic domain types, and Kotlin Multiplatform platform boundaries.
grounded-writing— draft or review public developer documentation and other user-owned text, including internal report reviews, while preserving evidence, format, and material-edit restraint.
deliver-spec— coordinate one approved spec or verified Project-controller handoff through solo or justified delegated implementation, reusable scope approval and evidence, independent integrated-candidate review and focused repair review, and PR shepherding; validate applicable live-qualification harness paths offline (isolated production stores for persistent paths, fake runtimes for asynchronous paths), then reconcile failures under bounded approved allowances.release-kotlin-library— assess readiness, prepare, and verify Kotlin library releases; check thegradle-maven-publish-pluginprerequisite, reconcile changelogs and Metalava API snapshots, and follow repository checks and publication gates.gradle-run— run every agent-initiated Gradle command through a compact-output wrapper; the implementation owner diagnoses and fixes failures, with optional read-only investigation helpers.implement-with-subagents— coordinate one or more worker-owned implementation tasks through dependency-aware dispatch, task acceptance, integration, and same-owner repair; dispatch independent ready work concurrently within actual capacity, reuse sequential worker slots, and support read-only orchestration review.to-plan— create risk-scaled implementation plans with observable slices and consequential boundary validation before dependent work, plus authenticated integration amendments and explicit consumer handoffs; decide separate PRs by coherent deliverable boundaries.run-github-project— set up, review, or operate a repository's GitHub Project workflow with current-column authorization, human-only Backlog promotion, and deterministic configured agent selection; deliver ordinary implementation tickets throughdeliver-spec, grant run-level merging withdrain --auto-merge, preflight offline live-qualification harness evidence, preserve asynchronous or persistent failures in durable checkpoints, prioritize active delivery through safe planner checkpoints, and record ticket-local pauses while independent work continues.shepherd— autonomously poll open PRs and MRs, or follow a host PR monitor's wake events (such as Claude desktop's Auto-fix) instead of polling, triage review comments, and repair CI failures with affected checks and required integrated-candidate validation.
Workflows that delegate agents share the subagent selection and handoff
reference. Each workflow retains its own
delegation trigger, authority, and acceptance rules. Each consuming skill also
contains a copy of the reference for standalone installation.
deliver-spec bundles its own solo-or-delegated implementation procedure;
implement-with-subagents retains its explicitly delegated workflow.
Edit references/subagent-selection.md, then run
npm run references:sync and include the updated bundled copies in the change.
Run npm run build to check that all copies are regular files matching their
sources; it fails on drift without rewriting files.
This is a breaking taxonomy change. Replace the removed entrypoints as follows:
| Removed skills | Replacement |
|---|---|
compose-state-authoring, compose-state-hoisting, compose-side-effects |
compose-state-and-effects |
compose-recomposition-performance, compose-stability-diagnostics, compose-state-deferred-reads |
compose-performance |
compose-modifier-and-layout-style, compose-slot-api-pattern |
compose-component-design |
kotlin-coroutines-structured-concurrency, kotlin-flow-state-event-modeling |
kotlin-concurrency-and-flow |
kotlin-functions, kotlin-types-value-class, kotlin-multiplatform-expect-actual |
kotlin-api-design |
Skills live at skills/<skill-name>/SKILL.md, flat (no language nesting). The name: in the SKILL.md frontmatter must match the directory name.
Frontmatter is validated against skills.schema.json, which
tracks the core Agent Skills specification
and permits disable-model-invocation for Claude Code compatibility.
name and description are required; portable optional fields are license,
compatibility, metadata, and allowed-tools. Explicit-only workflow skills
also mirror that policy in Codex's agents/openai.yaml.
Release versions use CalVer: YYYY.M.D, YYYY.M.D.N, or YYYY.M.D.NN, without
zero-padded month or day values. For example, use 2026.6.17 for the first
release of the day and 2026.6.17.1 or 2026.6.17.01 for another release.
Single-digit daily release numbers are normalized to the padded form, so both
inputs produce 2026.6.17.01.
Keep root plugin.json, .claude-plugin/plugin.json,
.codex-plugin/plugin.json, and new Git release tags on the same version.
Existing zero-padded tags from before this policy map to the non-padded manifest
version, so 2026.06.16 maps to 2026.6.16. Only bump versions when publishing
an installable release.
To publish a release, run the Release workflow from GitHub Actions. Leave the
version input empty to choose today's next unused UTC version, or provide a specific
CalVer value, optionally with a one- or two-digit daily release number. Use the
dry-run option to validate without creating a commit, tag, or GitHub release.
Automatic numbering checks existing tags and releases: the first release uses
YYYY.M.D, then subsequent releases increment the highest daily suffix through
.99. Gaps are not reused. If all daily numbers are exhausted, the workflow
fails and requires an explicit version or a release on another day.
The same workflow also runs nightly (02:17 UTC). It releases only if skills/,
the plugin manifests, .opencode/, .agents/, plugin.json, or package.json
changed since the latest tag; otherwise it exits without a release.
Before pushing, lint skills (frontmatter schema + markdown):
npm install
npm run lint
This also runs on CI for all PRs.
For a taxonomy change, also run the durable cluster behavior evaluation. It checks routing, required references, safeguards, exceptions, and finish gates at the public agent-facing seam.
Before publishing a release, manually run the advisory evaluations for the changed skills. Select each affected suite; when shared evaluation machinery changes, include every suite affected by that change. Preview each suite with the intended filters and current subject and judge cost assumptions:
python3 evals/run.py plan \
--suite <suite> \
--skill <changed-skill> \
--model <model> \
--reasoning <effort> \
--judge-model <judge-model> \
--judge-reasoning <effort> \
--subject-cost-per-call-usd <amount> \
--judge-cost-per-call-usd <amount> \
--jsonRepeat --skill for each changed skill. Without a --case filter, omitting
--skill selects all non-calibration cases in the suite. Calibration cases are
excluded unless you pass their explicit IDs with --case. Inspect case_ids in
the JSON output and confirm the selected cases match the intended scope before
reviewing call counts and estimated cost. The plan preview includes only the
counts and cost for first attempts. Execution may automatically retry each
subject and judge call once. Before --execute, get explicit cost approval for
up to twice the previewed subject calls, judge calls, and estimated cost, to
cover one retry per call. Execute the matching run command manually. Report
invalid or inconclusive results. Before any rerun, preview again, inspect its
case_ids, review the calls and cost, and get new approval for the same retry
headroom. Results remain advisory, not release gates. See
evals/README.md for suite selection and command options.
The advisory evaluator tests concrete scenarios modelled on real-world coding
work, with expected outcomes and no-change controls. It compares no-skill,
forced-skill, and automatic-routing runs. Baseline and automatic use
the cases eligible for automatic activation; restraint checks that a skill
does not make an unnecessary change. Scorecards also compare subject-side tokens,
tool calls, completed turns, elapsed time, and total attempted work per
successful outcome. The
table reports the latest available result for each skill and correctness metric.
These scores were produced using
gpt-6-luna
with high reasoning, judged by
gpt-6-sol with
high reasoning. Results are model- and reasoning-specific; other configurations
may perform differently. The human audit queue remains open. These are not merge
or release gates. See
evals/README.md for evaluation setup and reproducibility.
Skill-revision compatibility checks compare old and revised instructions within each model, separately from these benchmark scores. See the workflow compatibility record for Astra and 5.6 coverage and its current evidence limits.
| Skill | Baseline | Automatic | Restraint |
|---|---|---|---|
compose-animations |
75.0% | 100.0% | 100.0% |
compose-component-design |
86.7% | 100.0% | 100.0% |
compose-focus-navigation |
33.3% | 100.0% | 100.0% |
compose-performance |
83.3% | 100.0% | 100.0% |
compose-state-and-effects |
83.3% | 100.0% | 100.0% |
compose-ui-testing-patterns |
55.6% | 100.0% | 100.0% |
gradle-run |
41.7% | 100.0% | 100.0% |
kotlin-api-design |
58.3% | 100.0% | 100.0% |
kotlin-concurrency-and-flow |
44.4% | 100.0% | 100.0% |
kotlin-control-flow |
33.3% | 100.0% | 100.0% |
android-benchmark-comparison |
33.3% | 100.0% | 100.0% |
deliver-spec |
— | — | — |
grounded-writing |
0.0% | 100.0% | 100.0% |
implement-with-subagents |
— | — | 100.0% |
release-kotlin-library |
0.0% | 100.0% | 100.0% |
run-github-project |
— | — | 100.0% |
shepherd |
— | — | 100.0% |
to-plan |
— | — | 100.0% |
The android-benchmark-comparison, compose-state-and-effects,
compose-ui-testing-patterns, gradle-run, grounded-writing, kotlin-api-design,
kotlin-concurrency-and-flow, kotlin-control-flow, and
release-kotlin-library automatic cells, and the to-plan restraint cell,
use later focused evidence. Baseline and efficiency values use the complete
suite. The improvement result record,
targeted probe record,
and inline repair record
give provenance and remaining failures.
Values are per-run medians, baseline → automatic, followed by the automatic percentage change. These subject-only measurements use the latest complete, same-run evidence available for each suite and include failed runs and negative controls. Baseline-to-automatic efficiency comparisons use only cases eligible for automatic activation. Multi-skill scenarios contribute to every targeted skill row. A turn is one completed Codex turn; time remains environment-sensitive. Run provenance and local scorecard paths are in the GPT-6 improvement result record.
| Skill | Tokens / run | Tool calls / run | Turns / run | Time / run |
|---|---|---|---|---|
compose-animations |
58.6k → 85.4k (+46%) | 4 → 5 (+25%) | 1 → 1 (+0%) | 28.2s → 34.8s (+23%) |
compose-component-design |
48.7k → 73.4k (+51%) | 4 → 5 (+25%) | 1 → 1 (+0%) | 24.4s → 29.5s (+21%) |
compose-focus-navigation |
59.1k → 84.9k (+44%) | 5 → 5 (+0%) | 1 → 1 (+0%) | 28.9s → 36.6s (+27%) |
compose-performance |
49.1k → 85.0k (+73%) | 4 → 5 (+25%) | 1 → 1 (+0%) | 26.4s → 31.8s (+21%) |
compose-state-and-effects |
60.1k → 89.0k (+48%) | 4 → 5 (+25%) | 1 → 1 (+0%) | 28.2s → 37.8s (+34%) |
compose-ui-testing-patterns |
60.2k → 84.2k (+40%) | 5.5 → 5 (-9%) | 1 → 1 (+0%) | 27.1s → 26.8s (-1%) |
gradle-run |
59.8k → 104.3k (+75%) | 4 → 7 (+75%) | 1 → 1 (+0%) | 27.6s → 45.3s (+64%) |
kotlin-api-design |
60.1k → 83.9k (+40%) | 4 → 5 (+25%) | 1 → 1 (+0%) | 32.8s → 38.8s (+18%) |
kotlin-concurrency-and-flow |
60.0k → 74.0k (+23%) | 4 → 5 (+25%) | 1 → 1 (+0%) | 26.5s → 28.0s (+6%) |
kotlin-control-flow |
60.9k → 83.9k (+38%) | 5 → 5 (+0%) | 1 → 1 (+0%) | 26.2s → 38.5s (+47%) |
android-benchmark-comparison |
46.9k → 53.8k (+15%) | 3 → 3 (+0%) | 1 → 1 (+0%) | 23.2s → 35.0s (+51%) |
grounded-writing |
35.6k → 56.9k (+60%) | 2 → 4 (+100%) | 1 → 1 (+0%) | 23.2s → 22.1s (-5%) |
release-kotlin-library |
70.2k → 88.3k (+26%) | 4 → 5 (+25%) | 1 → 1 (+0%) | 42.1s → 43.9s (+4%) |
