Skip to content

A failed llama-bench process can be scored after printing a speed table #17

Description

@DivyamTalwar

Reproduction and expected behavior

The benchmark path parses its captured speed table without checking subprocess returncode. A failed run can therefore be scored or offered as a contribution. Reject nonzero exits before parsing and scoring; preserve successful results and bound the diagnostic tail.

Base: 252e5193902d466726da9af75047dfffff2ae662. The attached PR adds regressions against the actual package/service modules, then applies the correction. Running those final regression files against unchanged production produces:

8 failed, 5 passed in 0.09s

Focused command after placing the regressions in the checkout:

python -m pytest -q tests/test_bench_exit_status.py

Scope

This validates exit-status handling with a synthetic process boundary, not a performance benchmark. Existing hardware laws, calibration values and scoring formulas are unchanged.

Synthetic test input only; no reported live outage, security incident, hardware-performance measurement, or P0/P1 label.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions