Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
15 changes: 15 additions & 0 deletions benchmarks/chi-bench/submissions/2026-07-26-erius-opus5/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
# Michael Johnson (MJ) · erius · anthropic/claude-opus-5

Submitted: 2026-07-27 · chi-bench chi-bench-v1.0.0 · pass@1: **54.7%**

| Domain | pass@1 | n_trials |
|---|---|---|
| pa_provider | 72.0% | 25 |
| pa_um | 36.0% | 25 |
| cm | 56.0% | 25 |

Inspect a trajectory:

zstdcat trials/pa_provider/<trial_id>/agent/trajectory.jsonl.zst | jq .

See `submission.json` for the full manifest, `provenance.json` for reproducibility info.
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
{
"chi_bench_git_sha": null,
"image_digest": null,
"judge_model": "claude-opus-4-7",
"harness_version": "0.1.0"
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
benchmark,dataset_version,submission_id,team,agent,model,domain,pass_at_1,n_trials,n_tasks,mean_cost_usd,mean_walltime_s,submitted_at
chi-bench,chi-bench-v1.0.0,erius-opus5,Michael Johnson (MJ),erius,anthropic/claude-opus-5,overall,0.5466666666666666,75,75,0.0,0.0,2026-07-27T03:15:17Z
chi-bench,chi-bench-v1.0.0,erius-opus5,Michael Johnson (MJ),erius,anthropic/claude-opus-5,pa_provider,0.72,25,25,0.0,0.0,2026-07-27T03:15:17Z
chi-bench,chi-bench-v1.0.0,erius-opus5,Michael Johnson (MJ),erius,anthropic/claude-opus-5,pa_um,0.36,25,25,0.0,0.0,2026-07-27T03:15:17Z
chi-bench,chi-bench-v1.0.0,erius-opus5,Michael Johnson (MJ),erius,anthropic/claude-opus-5,cm,0.56,25,25,0.0,0.0,2026-07-27T03:15:17Z
55 changes: 55 additions & 0 deletions benchmarks/chi-bench/submissions/2026-07-26-erius-opus5/sub.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,55 @@
# Erius (Opus 5) × CHI-Bench submission — unified answer-blind GAMPO advisory agent on Opus 5.
# uv run cb submission validate -f configs/submissions/erius_opus5.yaml
# uv run cb submission run -f configs/submissions/erius_opus5.yaml
# uv run cb submission package -f configs/submissions/erius_opus5.yaml
# DRAFT for review (not yet filed). Trees frozen: rsi_checkpoints/opus5_submission_2026-07-26/.

schema: chi-bench/submission/v1

submission:
id: erius-opus5
team: Michael Johnson (MJ)
contact: johnson.michael@gmail.com
agent: erius # erius brand -> GampoAdvisoryHarness (the answer-blind GAMPO advisory agent that ran)
model: anthropic/claude-opus-5
notes: |
Erius on Opus 5: one GAMPO-governed agent (gampo-advisory), one model (claude-opus-5),
strict single-attempt pass@1 (n_attempts = 1), run via the Claude Max subscription (CLI
auth). Judge pinned to claude-opus-4-7 per leaderboard policy; agent and judge run on Max,
and the ONLY component that bills the funded API key is the care-management patient
simulator (sonnet-4-5), which the benchmark structurally requires.

Governance is an ANSWER-BLIND, per-task definition-of-done: a specification keyed only to
each case's own visible policy and to published United States standards (CMS-0057-F / Da
Vinci PAS, NCD/LCD and NASS/InterQual criteria, CMS/AMA coding read against the chart, and
the CCM/APCM/GUIDE care-management programs), never the hidden answer key, rubric, or
solution.

Per-domain configuration (all single-attempt, same agent and model):
- Prior-authorization (pa_provider): answer-blind advisory with chart-documentation-
fidelity facts (place-of-service and diagnosis-code corrections traced to the case's
own chart, never the gold), default reasoning effort. 18/25 = 72%.
- Utilization-management (pa_um): answer-blind advisory, xhigh reasoning effort. A
terminal sign-off commit lever was tested and REJECTED under a zero-regression rule
(it recovered one sign-off task but regressed a passing one, net 9/25), so the filed
UM slice is the plain advisory board. 9/25 = 36%.
- Care-management (cm): the GAMPO procedure only, under the same gampo-advisory agent
with no content-advisory file injected (the procedure beats the CMS content advisory
on this domain), max reasoning effort, sonnet-4-5 patient simulator. 14/25 = 56%.

Overall: 41/75 = 54.7% strict single-attempt pass@1. Exploratory: single trial per task,
25 tasks per domain. This is a cross-generation replication of the Opus 4.8 erius result
(filed separately at 37.3% under the generic procedure); it is not a resubmission of that
entry.

run:
environment: docker # local single-image; 'modal' for cloud-parallel
n_attempts: 1 # leaderboard policy: pass@1
concurrency: 4
max_retries: 2
timeout_multiplier: 2.0
env_file: /Users/mjohnson/Downloads/gampo-chibench/.env

dataset:
version: chi-bench-v1.0.0
domains: [pa_provider, pa_um, cm]
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
{
"schema": "chi-bench/submission/v1",
"submission": {
"id": "erius-opus5",
"team": "Michael Johnson (MJ)",
"contact": "johnson.michael@gmail.com",
"agent": "erius",
"model": "anthropic/claude-opus-5",
"notes": "Erius on Opus 5: one GAMPO-governed agent (gampo-advisory), one model (claude-opus-5),\nstrict single-attempt pass@1 (n_attempts = 1), run via the Claude Max subscription (CLI\nauth). Judge pinned to claude-opus-4-7 per leaderboard policy; agent and judge run on Max,\nand the ONLY component that bills the funded API key is the care-management patient\nsimulator (sonnet-4-5), which the benchmark structurally requires.\n\nGovernance is an ANSWER-BLIND, per-task definition-of-done: a specification keyed only to\neach case's own visible policy and to published United States standards (CMS-0057-F / Da\nVinci PAS, NCD/LCD and NASS/InterQual criteria, CMS/AMA coding read against the chart, and\nthe CCM/APCM/GUIDE care-management programs), never the hidden answer key, rubric, or\nsolution.\n\nPer-domain configuration (all single-attempt, same agent and model):\n - Prior-authorization (pa_provider): answer-blind advisory with chart-documentation-\n fidelity facts (place-of-service and diagnosis-code corrections traced to the case's\n own chart, never the gold), default reasoning effort. 18/25 = 72%.\n - Utilization-management (pa_um): answer-blind advisory, xhigh reasoning effort. A\n terminal sign-off commit lever was tested and REJECTED under a zero-regression rule\n (it recovered one sign-off task but regressed a passing one, net 9/25), so the filed\n UM slice is the plain advisory board. 9/25 = 36%.\n - Care-management (cm): the GAMPO procedure only, under the same gampo-advisory agent\n with no content-advisory file injected (the procedure beats the CMS content advisory\n on this domain), max reasoning effort, sonnet-4-5 patient simulator. 14/25 = 56%.\n\nOverall: 41/75 = 54.7% strict single-attempt pass@1. Exploratory: single trial per task,\n25 tasks per domain. This is a cross-generation replication of the Opus 4.8 erius result\n(filed separately at 37.3% under the generic procedure); it is not a resubmission of that\nentry.\n",
"submitted_at": "2026-07-27T03:15:17Z"
},
"dataset": {
"version": "chi-bench-v1.0.0",
"domains": [
"pa_provider",
"pa_um",
"cm"
],
"name": "chi-bench"
},
"results": {
"overall": {
"n_trials": 75,
"n_tasks": 75,
"pass_at_1": 0.5466666666666666,
"mean_cost_usd": 0.0,
"mean_walltime_s": 0.0
},
"per_domain": {
"pa_provider": {
"n_trials": 25,
"n_tasks": 25,
"pass_at_1": 0.72,
"mean_cost_usd": 0.0,
"mean_walltime_s": 0.0
},
"pa_um": {
"n_trials": 25,
"n_tasks": 25,
"pass_at_1": 0.36,
"mean_cost_usd": 0.0,
"mean_walltime_s": 0.0
},
"cm": {
"n_trials": 25,
"n_tasks": 25,
"pass_at_1": 0.56,
"mean_cost_usd": 0.0,
"mean_walltime_s": 0.0
}
},
"mean_cost_usd": 0.0,
"mean_walltime_s": 0.0
},
"provenance": {
"chi_bench_git_sha": null,
"image_digest": null,
"judge_model": "claude-opus-4-7",
"harness_version": "0.1.0"
}
}
Binary file not shown.
Original file line number Diff line number Diff line change
@@ -0,0 +1,109 @@
{
"id": "9d4c9a1a-8d09-40ce-8519-a963697c2da4",
"task_name": "actava-ai/cm_afib_moderate_anxious_001",
"trial_name": "cm_afib_moderate_anxious_001__riGkGZN",
"trial_uri": "file:///Users/mjohnson/Downloads/gampo-chibench/harness/trials_cm_gadv_cm_afib_moderate_anxious_001_a1/cm_afib_moderate_anxious_001__riGkGZN",
"task_id": {
"path": "data/care_management/tasks/cm_afib_moderate_anxious_001"
},
"source": null,
"task_checksum": "a43c14fadec72f11c1de9af6ee038fd27d3c453ce78dd9de56d454e395e33846",
"config": {
"task": {
"path": "data/care_management/tasks/cm_afib_moderate_anxious_001",
"git_url": null,
"git_commit_id": null,
"name": null,
"ref": null,
"overwrite": false,
"download_dir": null,
"source": null
},
"trial_name": "cm_afib_moderate_anxious_001__riGkGZN",
"trials_dir": "trials_cm_gadv_cm_afib_moderate_anxious_001_a1",
"timeout_multiplier": 1.0,
"agent_timeout_multiplier": 6.0,
"verifier_timeout_multiplier": null,
"agent_setup_timeout_multiplier": null,
"environment_build_timeout_multiplier": null,
"agent": {
"name": null,
"import_path": "chi_bench.experiment.agents.gampo_advisory_harness:GampoAdvisoryHarness",
"model_name": "anthropic/claude-opus-5",
"override_timeout_sec": null,
"override_setup_timeout_sec": null,
"max_timeout_sec": null,
"kwargs": {},
"env": {
"OPENAI_API_KEY": "${OPENAI_API_KEY}",
"ANTHROPIC_API_KEY": "${ANTHROPIC_API_KEY}",
"CLAUDE_CODE_OAUTH_TOKEN": "${CLAUDE_CODE_OAUTH_TOKEN}",
"GEMINI_API_KEY": "${GEMINI_API_KEY}",
"CLAUDE_CODE_EFFORT_LEVEL": "max"
}
},
"environment": {
"type": null,
"import_path": "chi_bench.experiment.docker_env:ChiBenchDockerEnvironment",
"force_build": false,
"delete": true,
"override_cpus": null,
"override_memory_mb": null,
"override_storage_mb": null,
"override_gpus": null,
"suppress_override_warnings": false,
"mounts_json": null,
"env": {},
"kwargs": {}
},
"verifier": {
"override_timeout_sec": null,
"max_timeout_sec": null,
"env": {},
"disable": false
},
"artifacts": [],
"job_id": null
},
"agent_info": {
"name": "gampo-advisory",
"version": "2.1.220",
"model_info": {
"name": "claude-opus-5",
"provider": "anthropic"
}
},
"agent_result": {
"n_input_tokens": 5034992,
"n_cache_tokens": 4768237,
"n_output_tokens": 94008,
"cost_usd": 7.401553500000001,
"rollout_details": null,
"metadata": null
},
"verifier_result": {
"rewards": {
"reward": 1.0
}
},
"exception_info": null,
"started_at": "2026-07-26T18:58:46.801779Z",
"finished_at": "2026-07-26T19:25:43.941210Z",
"environment_setup": {
"started_at": "2026-07-26T18:58:46.873130Z",
"finished_at": "2026-07-26T18:58:51.682709Z"
},
"agent_setup": {
"started_at": "2026-07-26T18:58:51.682775Z",
"finished_at": "2026-07-26T18:59:23.248776Z"
},
"agent_execution": {
"started_at": "2026-07-26T18:59:23.248882Z",
"finished_at": "2026-07-26T19:21:37.396433Z"
},
"verifier": {
"started_at": "2026-07-26T19:21:37.501249Z",
"finished_at": "2026-07-26T19:25:43.810129Z"
},
"step_results": null
}
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
{"reward": 1.0}
Original file line number Diff line number Diff line change
@@ -0,0 +1,113 @@
{
"binary_reward": 1.0,
"fractional_reward": 1.0,
"passed_checks": 19,
"total_checks": 19,
"check_scores": {
"cm.assessment.record_exists": 1.0,
"cm.assessment.completed": 1.0,
"cm.assessment.required_sections_present": 1.0,
"cm.care_plan.record_exists": 1.0,
"cm.care_plan.finalized": 1.0,
"cm.care_plan.problem_count": 1.0,
"cm.care_plan.goal_structure": 1.0,
"cm.care_plan.intervention_structure": 1.0,
"cm.care_plan.escalation_conditions_present": 1.0,
"cm.care_plan.follow_up_cadence_present": 1.0,
"cm.chart_review.record_exists": 1.0,
"cm.cross_stage.target_status": 1.0,
"cm.cross_stage.audit_actions": 1.0,
"cm.cross_stage.no_forbidden_mutations": 1.0,
"judge.cm.chart_review.quality": 1.0,
"judge.cm.outreach.quality": 1.0,
"judge.cm.assessment.quality": 1.0,
"judge.cm.care_plan.quality": 1.0,
"judge.cm.stage_coherence": 1.0
},
"checks": {
"cm.assessment.record_exists": true,
"cm.assessment.completed": true,
"cm.assessment.required_sections_present": true,
"cm.care_plan.record_exists": true,
"cm.care_plan.finalized": true,
"cm.care_plan.problem_count": true,
"cm.care_plan.goal_structure": true,
"cm.care_plan.intervention_structure": true,
"cm.care_plan.escalation_conditions_present": true,
"cm.care_plan.follow_up_cadence_present": true,
"cm.chart_review.record_exists": true,
"cm.cross_stage.target_status": true,
"cm.cross_stage.audit_actions": true,
"cm.cross_stage.no_forbidden_mutations": true,
"judge.cm.chart_review.quality": true,
"judge.cm.outreach.quality": true,
"judge.cm.assessment.quality": true,
"judge.cm.care_plan.quality": true,
"judge.cm.stage_coherence": true
},
"failed_checks": [],
"not_applicable_checks": [],
"stages": {
"cm_assessment": {
"passed": true,
"checks": {
"cm.assessment.record_exists": true,
"cm.assessment.completed": true,
"cm.assessment.required_sections_present": true
},
"passed_count": 3,
"total_count": 3,
"not_applicable_count": 0,
"details": {
"case_id": "CM-CASE-CM_AFIB_MODERATE_ANXIOUS_001",
"record_count": 1
}
},
"cm_care_plan": {
"passed": true,
"checks": {
"cm.care_plan.record_exists": true,
"cm.care_plan.finalized": true,
"cm.care_plan.problem_count": true,
"cm.care_plan.goal_structure": true,
"cm.care_plan.intervention_structure": true,
"cm.care_plan.escalation_conditions_present": true,
"cm.care_plan.follow_up_cadence_present": true
},
"passed_count": 7,
"total_count": 7,
"not_applicable_count": 0,
"details": {
"case_id": "CM-CASE-CM_AFIB_MODERATE_ANXIOUS_001",
"record_count": 1
}
},
"cm_chart_review": {
"passed": true,
"checks": {
"cm.chart_review.record_exists": true
},
"passed_count": 1,
"total_count": 1,
"not_applicable_count": 0,
"details": {
"case_id": "CM-CASE-CM_AFIB_MODERATE_ANXIOUS_001",
"record_count": 1
}
},
"cm_cross_stage": {
"passed": true,
"checks": {
"cm.cross_stage.target_status": true,
"cm.cross_stage.audit_actions": true,
"cm.cross_stage.no_forbidden_mutations": true
},
"passed_count": 3,
"total_count": 3,
"not_applicable_count": 0,
"details": {
"case_id": "CM-CASE-CM_AFIB_MODERATE_ANXIOUS_001"
}
}
}
}
Binary file not shown.
Loading
Loading