Skip to content
35 changes: 35 additions & 0 deletions docs/components/geak.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,41 @@ GEAK's `e2e_workflow` recursively drives `kernel_workflow` to author and tune th
individual hot kernels worth fixing. See
[Hyperloom optimization loop](../conceptual/optimization-loop.md).

## GPU pinning in the handoff

GEAK launches full servers out-of-process (baseline, profile, config-tuning
validation) and writes a visible-devices mask for each one, so the handoff has
to say which cards the run owns. Two fields carry that, in two different
coordinate systems:

| Field | Coordinate system | Value |
|-------|-------------------|-------|
| `gpu_ids` | HIP-level device list — HIP indexes into the ROCr-visible set | logical positions inside an *inherited* `ROCR_VISIBLE_DEVICES` mask, capped at `tp` and at the mask width (`ROCR=6` → `"0"`); a `HIP`/`CUDA` mask uncapped (`HIP=4,5` → `"4,5"`); `0..tp-1` when the run is unpinned. Ids are re-serialized from the parsed mask, so whitespace and repeats are normalized |
| `gpu_pin` | absolute device ids | `{"var", "value", "ids", "count", "source"}` for the winning mask — omitted entirely when no mask is set anywhere, which means "whole machine visible", not "pinned to card 0". `value` is the mask verbatim (so a UUID mask can be re-exported); `ids` is empty for a non-numeric mask, hence `count` |

The mask is resolved variable-major — `ROCR_VISIBLE_DEVICES` before
`HIP_VISIBLE_DEVICES` before `CUDA_VISIBLE_DEVICES`, the same order as
`orchestrator/bus/gpu_pool.py` and `orchestrator/policy/gate.py` — and within
each variable the process environment before the materialized baseline
recipe's `benchmark.envs`.

The process env comes first because the recipe's ROCR key is not evidence of a
pin: `materialize_config_with_envs` autofills `ROCR_VISIBLE_DEVICES=0..tp-1`
into every materialized recipe when the mask is absent or narrower than `TP`.
A recipe ROCR value byte-identical to that default is therefore ignored, so a
`HIP`-pinned or genuinely unpinned run is not silently re-pinned to cards
`0..tp-1`. A recipe mask that differs from the default *is* honoured — but its
ids are forwarded absolute, not logical, because the GEAK child inherits the
process environment and never sees that mask.

`tp` in the handoff is read from the same resolved recipe as `gpu_ids`, so the
two cannot disagree when the materializer clamps `TP` to the visible GPU count.

A consumer that re-exports `HIP_VISIBLE_DEVICES` should use `gpu_ids`; one that
writes `ROCR_VISIBLE_DEVICES` itself must use `gpu_pin["value"]`, because
writing `gpu_ids` into `ROCR_VISIBLE_DEVICES` resets the child to physical card
0 regardless of the run's pin.

## GEAK documentation

For detailed documentation on GEAK, see [GEAK on ROCm Docs](https://rocm.docs.amd.com/projects/geak/en/latest/).
44 changes: 44 additions & 0 deletions src/hyperloom/inference_optimizer/breakdown/collectors/geak.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@

from __future__ import annotations

from collections.abc import Mapping
from datetime import datetime, timezone
from pathlib import Path
from typing import Any
Expand Down Expand Up @@ -614,6 +615,34 @@ def _geak_accepted_kernels_from_integrate_results(
return accepted


def _handoff_gpu_fields(handoff: Mapping[str, Any] | None) -> dict[str, Any]:
"""Pull the device set + absolute GPU pin out of a handoff, if it has them.

A GEAK baseline that reads ``no_gain``/``incomplete`` because its servers
landed on a foreign tenant's card is otherwise indistinguishable from a
real result, so the breakdown records what the handoff told GEAK about the
run's cards (issue #1312).

Keys are inserted only when present, mirroring the writer's ``if gpu_pin:``
guard. An explicit ``null`` would conflate three different things: a
genuinely unpinned run, a pin that resolved empty, and a pre-v3 handoff on
disk (still produced by any session resumed from before this change).

Args:
handoff: The parsed ``geak/handoff.json``, or ``None``.

Returns:
A mapping with ``gpu_ids`` / ``gpu_pin`` for whichever keys the handoff
carries; ``{}`` when it carries neither.
"""
out: dict[str, Any] = {}
for key in ("gpu_ids", "gpu_pin"):
val = (handoff or {}).get(key)
if val is not None:
out[key] = val
return out


def _geak_reconstruct_from_disk(
session_dir: Path,
warnings: list[str],
Expand Down Expand Up @@ -671,6 +700,7 @@ def _load_json(p: Path) -> dict[str, Any]:
"workload": handoff.get("workload"),
"accepted_flags": handoff.get("accepted_flags"),
"raw_baseline_tput": _to_float(handoff.get("raw_baseline_tput")),
**_handoff_gpu_fields(handoff),
}

# 2) a flushed-but-unpromoted result.json (absent or non-ok status).
Expand Down Expand Up @@ -1003,9 +1033,23 @@ def _rel_if_under(p: Any) -> Any:
else:
accepted_kernels = []

# The cards GEAK was told to use. Read from the handoff on disk rather than
# from ``geak_result``, which never carried them — and recorded on THIS
# path, not just the crash-recovery one, because the outcome that needs
# disambiguating (`no_gain`) is a completed run.
_gpu_fields = _handoff_gpu_fields(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

on_error fires on a missing geak/handoff.json, adding a spurious warning to healthy runs.

FileNotFoundError is an OSError, and common/jsonio.read_json invokes on_error(exc) for OSError before returning the default. Any resumed pre-v3 session, or a run whose GEAK artifacts live under an external exp_root, reaches this unconditional read on the SUCCESS collect path and gets geak: handoff read failed: [Errno 2] No such file or directory appended to the breakdown warnings.

It is also unconditional disk I/O on a path that previously touched no files, for two optional fields — worth gating on the file existing, or passing an on_error that ignores FileNotFoundError.

read_json(
session_dir / "geak" / "handoff.json",
default={},
require_dict=True,
on_error=lambda exc: warnings.append(f"geak: handoff read failed: {exc}"),
)
)

section: dict[str, Any] = {
"engaged": True,
"status": status,
**_gpu_fields,
# Failure provenance (None on success).
"error_class": result.get("error_class"),
"error": result.get("error"),
Expand Down
255 changes: 255 additions & 0 deletions src/hyperloom/inference_optimizer/tests/test_geak_handoff_gpu_pin.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,255 @@
# SPDX-FileCopyrightText: 2026 Advanced Micro Devices, Inc.
# SPDX-License-Identifier: MIT
"""GPU-pin forwarding in the GEAK handoff (issue #1312).

GEAK launches full servers out-of-process and writes a visible-devices mask for
each one. When the handoff carries no pin it falls back to ``0..tp-1``, so every
server lands on physical GPU 0 no matter where the run was pinned — on a shared
host that collides with a foreign tenant and the resulting OOM reads like a real
regression.

These tests guard both halves of the contract:

* ``gpu_ids`` stays in the coordinate system the consumer applies it in (HIP
indexes into the ROCr-visible set), so existing pins keep working;
* ``gpu_pin`` carries the ABSOLUTE mask plus the variable it came from, so a
consumer that writes ``ROCR_VISIBLE_DEVICES`` re-applies the pin instead of
resetting the child to card 0.
"""

from __future__ import annotations

import pytest

from hyperloom.orchestrator.loop.coordinator_helpers import (
_is_autofilled_rocr,
_parse_device_list,
_resolve_gpu_pin,
_resolve_handoff_gpu_ids,
)


def _autofilled(tp: int) -> dict[str, object]:
"""The ``benchmark.envs`` every materialized recipe carries.

``materialize_config_with_envs`` writes ``ROCR_VISIBLE_DEVICES=0..tp-1``
unconditionally when the mask is absent or narrower than TP, so this shape
— not an empty mapping — is what the resolver sees in production.
"""
return {"TP": tp, "ROCR_VISIBLE_DEVICES": ",".join(str(i) for i in range(tp))}


_MASK_VARS = ("ROCR_VISIBLE_DEVICES", "HIP_VISIBLE_DEVICES", "CUDA_VISIBLE_DEVICES")


@pytest.fixture(autouse=True)
def _clear_masks(monkeypatch: pytest.MonkeyPatch) -> None:
"""Run every case against a known-unpinned environment."""
for var in _MASK_VARS:
monkeypatch.delenv(var, raising=False)


# --------------------------------------------------------------------------- #
# _parse_device_list
# --------------------------------------------------------------------------- #


def test_parse_device_list_forms() -> None:
assert _parse_device_list("4,5,6,7") == [4, 5, 6, 7]
assert _parse_device_list(" 6 ") == [6]
assert _parse_device_list("0;1") == [0, 1]
assert _parse_device_list("3,3,2") == [3, 2]


def test_parse_device_list_tolerates_junk_and_empty() -> None:
assert _parse_device_list("") == []
assert _parse_device_list(None) == []
assert _parse_device_list("a,,-1,2") == [2]


# --------------------------------------------------------------------------- #
# _resolve_gpu_pin
# --------------------------------------------------------------------------- #


def test_pin_unset_everywhere_is_empty() -> None:
"""No mask anywhere means "whole machine visible", NOT "pinned to 0"."""
assert _resolve_gpu_pin(recipe_envs={}, environ={}) == {}


def test_pin_from_process_rocr() -> None:
"""The case issue #1312 hit: ROCm's canonical mask, previously ignored."""
out = _resolve_gpu_pin(recipe_envs={}, environ={"ROCR_VISIBLE_DEVICES": "7"})
assert out == {
"var": "ROCR_VISIBLE_DEVICES",
"value": "7",
"ids": [7],
"count": 1,
"source": "process_env",
}


def test_pin_prefers_rocr_over_hip_and_cuda() -> None:
env = {
"CUDA_VISIBLE_DEVICES": "0",
"HIP_VISIBLE_DEVICES": "1",
"ROCR_VISIBLE_DEVICES": "4,5",
}
out = _resolve_gpu_pin(recipe_envs={}, environ=env)
assert out["var"] == "ROCR_VISIBLE_DEVICES"
assert out["ids"] == [4, 5]


def test_pin_prefers_process_env_over_recipe() -> None:
"""The process mask is the one the GEAK child actually inherits."""
out = _resolve_gpu_pin(
recipe_envs={"TP": 1, "ROCR_VISIBLE_DEVICES": "6"},
environ={"ROCR_VISIBLE_DEVICES": "3"},
)
assert out["source"] == "process_env"
assert out["ids"] == [3]


def test_pin_uses_recipe_when_the_process_is_unmasked() -> None:
"""A hand-authored recipe mask is still a pin when nothing else says otherwise."""
out = _resolve_gpu_pin(recipe_envs={"TP": 2, "ROCR_VISIBLE_DEVICES": "6,7"}, environ={})
assert out["source"] == "baseline_recipe"
assert out["ids"] == [6, 7]


def test_pin_skips_blank_values() -> None:
out = _resolve_gpu_pin(
recipe_envs={"ROCR_VISIBLE_DEVICES": " "},
environ={"HIP_VISIBLE_DEVICES": "2,3"},
)
assert out["var"] == "HIP_VISIBLE_DEVICES"
assert out["ids"] == [2, 3]


# --------------------------------------------------------------------------- #
# The materializer's autofilled ROCR mask (PR #1321 review)
# --------------------------------------------------------------------------- #


def test_autofilled_recipe_rocr_does_not_override_a_hip_pin() -> None:
"""Regression: recipe-first made every HIP-pinned run report cards 0..tp-1.

``materialize_config_with_envs`` synthesizes ``ROCR_VISIBLE_DEVICES=0,1``
into the recipe for a ``TP=2`` run that has no ROCR anywhere. Honouring
that as a pin overrode the real ``HIP_VISIBLE_DEVICES=4,5`` and told a
ROCR-writing consumer to hard-pin physical cards 0 and 1 — recreating the
card-0 collision this whole change exists to remove.
"""
out = _resolve_gpu_pin(
recipe_envs=_autofilled(2),
environ={"HIP_VISIBLE_DEVICES": "4,5"},
)
assert out["var"] == "HIP_VISIBLE_DEVICES"
assert out["ids"] == [4, 5]
assert _resolve_handoff_gpu_ids(gpu_pin=out, tp=2) == "4,5" # pre-PR value, preserved


def test_autofilled_recipe_rocr_leaves_an_unpinned_run_unpinned() -> None:
"""The documented ``{}`` contract has to be reachable in production."""
assert _resolve_gpu_pin(recipe_envs=_autofilled(4), environ={}) == {}
assert _resolve_handoff_gpu_ids(gpu_pin={}, tp=4) == "0,1,2,3"


def test_a_real_recipe_rocr_pin_survives_the_autofill_check() -> None:
assert not _is_autofilled_rocr(value="4,5", recipe_envs={"TP": 2})
assert _is_autofilled_rocr(value="0,1", recipe_envs={"TP": 2})
# No resolved TP => cannot claim it was synthesized; keep the mask.
assert not _is_autofilled_rocr(value="0,1", recipe_envs={})


def test_variable_precedence_is_global_not_per_source() -> None:
"""A leftover recipe CUDA key must not outrank a real process ROCR pin."""
out = _resolve_gpu_pin(
recipe_envs={"TP": 2, "CUDA_VISIBLE_DEVICES": "0"},
environ={"ROCR_VISIBLE_DEVICES": "6,7"},
)
assert out["var"] == "ROCR_VISIBLE_DEVICES"
assert out["ids"] == [6, 7]


def test_pin_reads_process_env_by_default(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setenv("ROCR_VISIBLE_DEVICES", "5")
assert _resolve_gpu_pin()["ids"] == [5]


# --------------------------------------------------------------------------- #
# _resolve_handoff_gpu_ids
# --------------------------------------------------------------------------- #


def test_gpu_ids_unpinned_is_range_tp() -> None:
"""Unchanged legacy behaviour for an unpinned run."""
assert _resolve_handoff_gpu_ids(gpu_pin={}, tp=4) == "0,1,2,3"
assert _resolve_handoff_gpu_ids(gpu_pin=None, tp=1) == "0"
assert _resolve_handoff_gpu_ids(gpu_pin={}, tp=0) == "0"


def test_gpu_ids_rocr_pin_is_logical() -> None:
"""HIP indexes into the ROCr-visible set, so ROCR=6 is HIP index 0."""
pin = _resolve_gpu_pin(recipe_envs={}, environ={"ROCR_VISIBLE_DEVICES": "6"})
assert _resolve_handoff_gpu_ids(gpu_pin=pin, tp=1) == "0"

pin4 = _resolve_gpu_pin(recipe_envs={}, environ={"ROCR_VISIBLE_DEVICES": "4,5,6,7"})
assert _resolve_handoff_gpu_ids(gpu_pin=pin4, tp=4) == "0,1,2,3"
# Capped at tp, as the unpinned path always was.
assert _resolve_handoff_gpu_ids(gpu_pin=pin4, tp=2) == "0,1"
# ...and at the mask when tp overshoots it: you cannot serve on cards you
# cannot see.
assert _resolve_handoff_gpu_ids(gpu_pin=pin4, tp=8) == "0,1,2,3"


def test_gpu_ids_hip_pin_is_verbatim() -> None:
"""No ROCr mask => ROCr shows every card, so HIP ids are absolute."""
pin = _resolve_gpu_pin(recipe_envs={}, environ={"HIP_VISIBLE_DEVICES": "4,5"})
assert _resolve_handoff_gpu_ids(gpu_pin=pin, tp=2) == "4,5"


def test_gpu_ids_cuda_pin_is_verbatim() -> None:
pin = _resolve_gpu_pin(recipe_envs={}, environ={"CUDA_VISIBLE_DEVICES": "3"})
assert _resolve_handoff_gpu_ids(gpu_pin=pin, tp=1) == "3"


def test_gpu_ids_never_empty_for_a_blank_mask() -> None:
"""A present-but-empty mask must not produce an empty device list."""
pin = {"var": "ROCR_VISIBLE_DEVICES", "value": "", "ids": [], "count": 0, "source": "process_env"}
assert _resolve_handoff_gpu_ids(gpu_pin=pin, tp=2) == "0,1"


def test_gpu_ids_counts_a_uuid_mask_instead_of_falling_back_to_card_0() -> None:
"""ROCm accepts UUID masks; they parse to zero numeric ids but N devices.

Counting ``ids`` here would see an empty list, read the run as unpinned and
emit ``0..tp-1`` — landing every GEAK server on card 0, the exact default
this change exists to eliminate.
"""
pin = _resolve_gpu_pin(
recipe_envs={},
environ={"ROCR_VISIBLE_DEVICES": "GPU-a1b2c3,GPU-d4e5f6"},
)
assert pin["ids"] == []
assert pin["count"] == 2
assert pin["value"] == "GPU-a1b2c3,GPU-d4e5f6" # re-exportable as-is
assert _resolve_handoff_gpu_ids(gpu_pin=pin, tp=2) == "0,1"


def test_a_yaml_sequence_mask_is_not_stringified_into_junk() -> None:
"""``ROCR_VISIBLE_DEVICES: [4, 5]`` in a recipe is a list, not a string."""
pin = _resolve_gpu_pin(recipe_envs={"TP": 2, "ROCR_VISIBLE_DEVICES": [4, 5]}, environ={})
assert pin["value"] == "4,5"
assert pin["ids"] == [4, 5]


def test_gpu_ids_are_absolute_for_a_recipe_only_rocr_pin() -> None:
"""A mask the child does not inherit cannot be indexed logically.

The phase launches GEAK with ``dict(os.environ)``, so a mask that exists
only in the recipe never reaches the child; ROCr shows it every card and
the absolute ids are the correct HIP indices.
"""
pin = _resolve_gpu_pin(recipe_envs={"TP": 2, "ROCR_VISIBLE_DEVICES": "6,7"}, environ={})
assert _resolve_handoff_gpu_ids(gpu_pin=pin, tp=2) == "6,7"
Loading
Loading