Skip to content
Open
Show file tree
Hide file tree
Changes from 60 commits
Commits
Show all changes
64 commits
Select commit Hold shift + click to select a range
034e83e
Support v1 raw multimodal offload
eligotts Jun 24, 2026
b7903f6
Update raw image renderer pin
eligotts Jun 25, 2026
6c0157a
Simplify v1 raw multimodal serving: always materialize refs
S1ro1 Jun 27, 2026
5ec8148
serving: use canonical split_raw_mm_ref + bump renderers pin
S1ro1 Jun 27, 2026
bcc93c2
Bump renderers pin (drop orphaned image_cache_max)
S1ro1 Jun 27, 2026
2c5af36
Merge remote-tracking branch 'origin/main' into feat/v1-raw-mm-offload
eligotts Jun 28, 2026
75b9e7c
Bump raw image dependency pins
eligotts Jun 28, 2026
99b590a
feat: support inline multimodal image storage
eligotts Jun 28, 2026
3518763
fix: preserve launcher image asset env
eligotts Jun 28, 2026
f37ceac
Simplify raw multimodal offload path
eligotts Jun 29, 2026
a66f1b1
Merge remote-tracking branch 'origin/main' into feat/v1-raw-mm-offload
eligotts Jun 29, 2026
80dc59f
Simplify raw multimodal validation
eligotts Jun 29, 2026
bf2bc8e
Clarify raw image asset env wiring
eligotts Jun 29, 2026
805486f
Use URI-complete raw multimodal refs
eligotts Jun 29, 2026
2f000f3
Merge branch 'main' into feat/v1-raw-mm-offload
eligotts Jun 29, 2026
dce2759
Drop raw multimodal descriptor version
eligotts Jun 29, 2026
0b0f1fb
Bump verifiers raw multimodal client prep
eligotts Jun 29, 2026
d8218c3
Merge remote-tracking branch 'origin/main' into feat/v1-raw-mm-offload
eligotts Jun 29, 2026
d18ed10
Bump renderers for processed multimodal output
eligotts Jun 30, 2026
0da60a8
Bump renderers lock cleanup
eligotts Jun 30, 2026
c7deb3b
Bump verifiers for main sync
eligotts Jul 1, 2026
89fcbc1
Merge remote-tracking branch 'origin/main' into feat/v1-raw-mm-offload
eligotts Jul 1, 2026
36b5042
fix: truncate raw multimodal refs safely
eligotts Jul 3, 2026
53d9ed8
Tighten raw multimodal validation and trainer mm plumbing
eligotts Jul 4, 2026
a13a214
Bump renderers and verifiers for raw multimodal hardening
eligotts Jul 4, 2026
7ea71a9
Merge remote-tracking branch 'origin/main' into feat/v1-raw-mm-offload
eligotts Jul 5, 2026
1f290dc
Document multimodal offload monitoring in the monitor-run skill
eligotts Jul 5, 2026
a3f5251
Bump renderers and verifiers for main reconciliation
eligotts Jul 23, 2026
2119bdb
Merge remote-tracking branch 'origin/main' into feat/v1-raw-mm-offload
eligotts Jul 23, 2026
a2fdf6a
Default SFT renderers to processed multimodal output
eligotts Jul 23, 2026
7676155
Cache materialized raw image refs in the serving layer
eligotts Jul 23, 2026
d00ddac
Bump verifiers: drop the orphaned prepare_messages hook
eligotts Jul 23, 2026
30cdce0
Merge remote-tracking branch 'origin/main' into feat/v1-raw-mm-offload
eligotts Jul 23, 2026
e0c22ee
Audit fixes: SFT mm position ids, cache cancellation, CI test adaptat…
eligotts Jul 23, 2026
6a1a54b
Pack raw-ref multimodal samples like main packs eager ones
eligotts Jul 24, 2026
23e5980
Merge remote-tracking branch 'origin/main' into feat/v1-raw-mm-offload
eligotts Jul 24, 2026
3bfaf11
Bump renderers for the unwrapped ref payload; harden the cache key
eligotts Jul 24, 2026
fd3f379
Inherit the parent env explicitly at the orchestrator launch
eligotts Jul 24, 2026
2c3f6d6
Merge remote-tracking branch 'origin/main' into feat/v1-raw-mm-offload
eligotts Jul 31, 2026
3b2dffd
Relock against the merged research-environments pin; bump verifiers
eligotts Jul 31, 2026
95afd1f
Adapt orchestrator timing metrics to the verifiers agent-span rename
eligotts Jul 31, 2026
37ff96a
Bump renderers: raw mode accepts file:// URLs only
eligotts Aug 1, 2026
743a022
refactor: derive processor fingerprints from renderer layout specs
eligotts Aug 3, 2026
ccb0cd2
Merge remote-tracking branch 'origin/main' into feat/v1-raw-mm-offload
eligotts Aug 4, 2026
4b4f9e3
refactor: trust the train-time checkpoint contract at materialize
eligotts Aug 4, 2026
52adc59
refactor: prune RawMMItem to the fields materialize reads
eligotts Aug 4, 2026
f753800
chore: bump renderers for shared mm_image helpers
eligotts Aug 4, 2026
21e0430
chore: bump renderers after dropping image cache
eligotts Aug 4, 2026
8e00291
chore: bump renderers for local-only processed image sources
eligotts Aug 4, 2026
97c08a2
Merge remote-tracking branch 'origin/main' into feat/v1-raw-mm-offload
eligotts Aug 5, 2026
0e54032
chore: bump renderers for the Pillow-only vision extra
eligotts Aug 6, 2026
98480ba
chore: bump verifiers to latest raw-image-offload head
eligotts Aug 7, 2026
8bcd811
refactor: env servers resolve their own image-asset dir
eligotts Aug 7, 2026
3888738
chore: relock for the verifiers pin; exclude nemo-gym-weather-v1
eligotts Aug 7, 2026
b92dbd5
chore: bump verifiers for the nemo-gym relock
eligotts Aug 7, 2026
ab09483
refactor(trainer): always fail-fast on missing MM images
eligotts Aug 7, 2026
bbb9dce
fix(inference): retain strong refs to mm materialize tasks
eligotts Aug 7, 2026
49fa709
docs: drop advanced.md VLM limitation churn from this PR
eligotts Aug 7, 2026
0ea481a
docs: correct VF_RENDERER_IMAGE_OFFLOAD_DIR wiring notes
eligotts Aug 7, 2026
454b438
docs: drop configuration.md multimodal offload note
eligotts Aug 7, 2026
17b1735
docs: shorten SFT multimodal_output validator messaging
eligotts Aug 7, 2026
4cde3d8
refactor: lean PROTECTED_ENV_VARS change for image offload dir
eligotts Aug 7, 2026
f92bb05
docs: drop monitor-run multimodal offload skill section
eligotts Aug 7, 2026
b57c4ac
refactor(mm): share image load/verify; trim serving_tokens helpers
eligotts Aug 7, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion deps/verifiers
Submodule verifiers updated 89 files
+4 −0 .github/workflows/test.yml
+0 −1 docs/mint.json
+3 −104 docs/v1/agent.md
+1 −1 docs/v1/architecture.md
+11 −82 docs/v1/env.md
+0 −96 docs/v1/evaluation.md
+2 −1 docs/v1/gepa.md
+1 −14 docs/v1/harnesses.md
+2 −2 docs/v1/overview.md
+0 −2 docs/v1/tasksets.md
+20 −0 environments/nemo_gym_weather_v1/README.md
+4 −0 environments/nemo_gym_weather_v1/nemo_gym_weather_v1/__init__.py
+5 −0 environments/nemo_gym_weather_v1/nemo_gym_weather_v1/example.jsonl
+22 −0 environments/nemo_gym_weather_v1/nemo_gym_weather_v1/taskset.py
+13 −0 environments/nemo_gym_weather_v1/pyproject.toml
+32 −1 pyproject.toml
+1 −1 skills/evaluate-environments/SKILL.md
+10 −5 tests/v1/conftest.py
+8 −0 tests/v1/fixtures/echo_agentic_v1.py
+4 −5 tests/v1/fixtures/echo_tool_v1.py
+113 −60 tests/v1/test_e2e.py
+3 −5 tests/v1/test_envs.py
+24 −4 tests/v1/test_judges.py
+515 −71 uv.lock
+2 −3 verifiers/envs/experimental/composable/harnesses/rlm.py
+2 −1 verifiers/envs/experimental/utils/git_checkout_cache.py
+2 −2 verifiers/types.py
+66 −0 verifiers/utils/multimodal.py
+12 −0 verifiers/utils/path_utils.py
+4 −1 verifiers/v1/__init__.py
+252 −6 verifiers/v1/acp/__init__.py
+0 −213 verifiers/v1/acp/_runner.py
+449 −0 verifiers/v1/acp/runner.py
+3 −2 verifiers/v1/cli/eval/main.py
+1 −38 verifiers/v1/cli/eval/resume.py
+7 −7 verifiers/v1/cli/eval/runner.py
+6 −2 verifiers/v1/cli/output.py
+42 −0 verifiers/v1/cli/resume.py
+253 −27 verifiers/v1/cli/validate.py
+9 −0 verifiers/v1/clients/client.py
+17 −17 verifiers/v1/clients/train.py
+1 −1 verifiers/v1/configs/cli/env.py
+14 −0 verifiers/v1/configs/cli/validate.py
+3 −1 verifiers/v1/configs/harness.py
+2 −2 verifiers/v1/configs/judge.py
+5 −1 verifiers/v1/dialects/chat.py
+2 −2 verifiers/v1/envs/agentic_judge/__init__.py
+65 −120 verifiers/v1/envs/agentic_judge/env.py
+6 −0 verifiers/v1/envs/shared_agentic_judge/__init__.py
+2 −1 verifiers/v1/envs/user_sim/env.py
+87 −31 verifiers/v1/graph.py
+106 −15 verifiers/v1/harness.py
+2 −1 verifiers/v1/harnesses/bash/harness.py
+3 −2 verifiers/v1/harnesses/bash/program.py
+2 −1 verifiers/v1/harnesses/browser_use/harness.py
+3 −2 verifiers/v1/harnesses/browser_use/program.py
+121 −61 verifiers/v1/harnesses/claude_code/harness.py
+152 −238 verifiers/v1/harnesses/codex/harness.py
+54 −0 verifiers/v1/harnesses/node.py
+2 −1 verifiers/v1/harnesses/null/harness.py
+3 −2 verifiers/v1/harnesses/null/program.py
+3 −5 verifiers/v1/harnesses/openclaw/harness.py
+6 −31 verifiers/v1/harnesses/pi/harness.py
+79 −42 verifiers/v1/harnesses/rlm/harness.py
+17 −4 verifiers/v1/interception/server.py
+2 −2 verifiers/v1/interception/tunnel/prime.py
+76 −62 verifiers/v1/judges/rubric.py
+40 −14 verifiers/v1/legacy.py
+22 −9 verifiers/v1/mcp/launch.py
+3 −2 verifiers/v1/mcp/server.py
+1 −1 verifiers/v1/mcp/toolset.py
+33 −11 verifiers/v1/rollout.py
+2 −0 verifiers/v1/runtimes/__init__.py
+42 −7 verifiers/v1/runtimes/base.py
+170 −2 verifiers/v1/runtimes/docker/__init__.py
+25 −16 verifiers/v1/runtimes/limiters.py
+100 −1 verifiers/v1/runtimes/modal.py
+70 −21 verifiers/v1/runtimes/prime.py
+67 −13 verifiers/v1/runtimes/subprocess.py
+9 −1 verifiers/v1/session.py
+3 −0 verifiers/v1/tasksets/__init__.py
+33 −3 verifiers/v1/tasksets/harbor/taskset.py
+17 −0 verifiers/v1/tasksets/nemo_gym/__init__.py
+95 −0 verifiers/v1/tasksets/nemo_gym/response.py
+42 −0 verifiers/v1/tasksets/nemo_gym/server.py
+168 −0 verifiers/v1/tasksets/nemo_gym/taskset.py
+103 −0 verifiers/v1/tasksets/nemo_gym/toolset.py
+31 −12 verifiers/v1/trace.py
+8 −2 verifiers/v1/types.py
5 changes: 4 additions & 1 deletion packages/prime-rl-configs/src/prime_rl/configs/env_server.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
import verifiers.v1 as vf
from pydantic import SerializeAsAny, model_validator

from prime_rl.configs.shared import LogConfig
from prime_rl.configs.shared import LogConfig, MultimodalConfig
from prime_rl.utils.config import BaseConfig


Expand All @@ -26,6 +26,9 @@ class EnvServerConfig(BaseConfig):
output_dir: Path = Path("outputs")
"""Directory to write outputs to — logs and any generated artifacts are written as subdirectories."""

multimodal: MultimodalConfig = MultimodalConfig()
"""Raw multimodal image offload settings; with ``output_dir``, resolves the run image-asset dir this server offloads into."""

@model_validator(mode="before")
@classmethod
def _resolve_env(cls, data):
Expand Down
5 changes: 4 additions & 1 deletion packages/prime-rl-configs/src/prime_rl/configs/inference.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@
from pydantic import Field, model_validator
from pydantic_config import BaseConfig

from prime_rl.configs.shared import BaseModelConfig, EnvVars, LogConfig, SlurmConfig
from prime_rl.configs.shared import BaseModelConfig, EnvVars, LogConfig, MultimodalConfig, SlurmConfig
from prime_rl.utils.config import find_package_resource, rgetattr, rsetattr
from prime_rl.utils.parsers import resolve_reasoning_parser, resolve_tool_call_parser

Expand Down Expand Up @@ -423,6 +423,9 @@ class InferenceConfig(BaseConfig):
dry_run: bool = False
"""Only validate and dump resolved configs, then exit early."""

multimodal: MultimodalConfig = MultimodalConfig()
"""Raw multimodal image offload settings shared with trainer and orchestrator."""

@model_validator(mode="after")
def validate_multi_node_requires_slurm(self):
if self.deployment.type in ("multi_node", "disaggregated") and self.slurm is None:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,7 @@
FileMonitorConfig,
HeartbeatConfig,
LogConfig,
MultimodalConfig,
PrimeMonitorConfig,
TransportConfig,
WandbWithExtrasConfig,
Expand Down Expand Up @@ -571,6 +572,9 @@ class OrchestratorConfig(BaseConfig):
heartbeat: HeartbeatConfig | None = None
"""BetterStack heartbeat configuration for monitoring training progress."""

multimodal: MultimodalConfig = MultimodalConfig()
"""Raw multimodal image offload settings shared with trainer and inference."""

@model_validator(mode="after")
def auto_setup_tokenizer(self):
if self.tokenizer.name is None:
Expand Down
3 changes: 3 additions & 0 deletions packages/prime-rl-configs/src/prime_rl/configs/rl.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@
from prime_rl.configs.shared import (
EnvVars,
FileMonitorConfig,
MultimodalConfig,
SlurmConfig,
TransportConfig,
VLMConfig,
Expand Down Expand Up @@ -261,6 +262,8 @@ class RLConfig(BaseConfig):

weight_broadcast: SharedWeightBroadcastConfig | None = None

multimodal: MultimodalConfig = MultimodalConfig()
"""Shared raw multimodal image offload settings. Propagated to trainer, orchestrator, and inference."""
rollout_transport: TransportConfig | None = None

bench: bool = False
Expand Down
17 changes: 17 additions & 0 deletions packages/prime-rl-configs/src/prime_rl/configs/sft.py
Original file line number Diff line number Diff line change
Expand Up @@ -248,6 +248,23 @@ def normalize_deployment(cls, data):

### Validate configs (e.g. raise for unsupported (combinations of) configs)

@model_validator(mode="after")
def renderer_emits_processed_multimodal(self):
"""SFT consumes processed pixel tensors straight from the renderer — it has no
raw-ref materializer, so the renderers-library default of ``multimodal_output='raw'``
(built for the RL offload path) would silently ship JSON descriptors into training.
Default SFT renderers to ``'processed'`` and reject an explicit ``'raw'``."""
if "multimodal_output" in self.renderer.model_fields_set:
if self.renderer.multimodal_output == "raw":
raise ValueError(
"multimodal_output='raw' is unsupported for SFT: the SFT data path materializes "
"images from processed renderer output, not raw refs. Remove the override "
"(SFT defaults to 'processed')."
)
else:
self.renderer = self.renderer.model_copy(update={"multimodal_output": "processed"})
return self

@model_validator(mode="after")
def deepep_disables_grad_clipping(self):
if self.model.ep_comm_backend == "deepep" and self.optim.max_norm is not None:
Expand Down
25 changes: 20 additions & 5 deletions packages/prime-rl-configs/src/prime_rl/configs/shared.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,19 +6,29 @@

from prime_rl.utils.config import BaseConfig

# Launcher-managed env vars that a component's `env_vars` must not set: GPU partitioning
# and the single shared W&B run. The launcher always sets these last, so allowing them in
# `env_vars` would be a silent no-op (or, on multi-node, a footgun) — reject them instead.
# Env vars that a component's `env_vars` must not set:
# - CUDA_VISIBLE_DEVICES / WANDB_SHARED_*: the launcher always sets these last.
# - VF_RENDERER_IMAGE_OFFLOAD_DIR: each env-server resolves this from
# ``[multimodal].offload_dir`` (operator process env wins via setdefault).
# Allowing them in ``env_vars`` would be a silent no-op or a footgun — reject instead.
PROTECTED_ENV_VARS = frozenset(
{"CUDA_VISIBLE_DEVICES", "WANDB_SHARED_MODE", "WANDB_SHARED_RUN_ID", "WANDB_SHARED_LABEL"}
{
"CUDA_VISIBLE_DEVICES",
"VF_RENDERER_IMAGE_OFFLOAD_DIR",
"WANDB_SHARED_MODE",
"WANDB_SHARED_RUN_ID",
"WANDB_SHARED_LABEL",
}
)


def reject_protected_env_vars(env_vars: dict[str, str]) -> dict[str, str]:
clobbered = sorted(PROTECTED_ENV_VARS & env_vars.keys())
if clobbered:
raise ValueError(
f"env_vars cannot set launcher-managed vars {clobbered} — set by the launcher, not overridable"
f"env_vars cannot set protected vars {clobbered}; "
"use [multimodal].offload_dir for VF_RENDERER_IMAGE_OFFLOAD_DIR, "
"and leave CUDA_VISIBLE_DEVICES / WANDB_SHARED_* to the launcher"
)
return env_vars

Expand Down Expand Up @@ -86,6 +96,11 @@ def resolve_project_dir(self):
ServerType = Literal["vllm", "openai"]


class MultimodalConfig(BaseConfig):
offload_dir: Path | None = None
"""Directory for offloaded image assets. Supports environment expansion such as ``/data/outputs/run_${RUN_ID}/assets/images``. When unset, prime-rl resolves a run-scoped default."""


class VLMConfig(BaseConfig):
vision_encoder_attr: str
"""Dotted attribute path to the vision encoder module (e.g. ``model.visual``)."""
Expand Down
4 changes: 4 additions & 0 deletions packages/prime-rl-configs/src/prime_rl/configs/trainer.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@
FileMonitorConfig,
HeartbeatConfig,
MetricsServerConfig,
MultimodalConfig,
TrainerLogConfig,
TransportConfig,
WandbConfig,
Expand Down Expand Up @@ -628,6 +629,9 @@ class TrainerConfig(BaseConfig):
max_concurrent_runs: int = Field(1, ge=1)
"""Maximum number of concurrent runs to allow. If 1, only one run may run at a time."""

multimodal: MultimodalConfig = MultimodalConfig()
"""Raw multimodal image offload settings shared with orchestrator and inference."""

enable_token_export: bool = False
"""Opt-in per-token JSONL export for rollout debugging. When enabled, writes token ids and aligned trainer metrics after each forward pass."""

Expand Down
1 change: 1 addition & 0 deletions packages/prime-rl-configs/src/prime_rl/utils/validation.py
Original file line number Diff line number Diff line change
Expand Up @@ -132,6 +132,7 @@ def propagate(shared_path: str, *targets: str) -> None:
# Top-level scalars.
propagate("max_steps", "trainer.max_steps", "orchestrator.max_steps")
propagate("seq_len", "trainer.model.seq_len", "orchestrator.seq_len")
propagate("multimodal", "trainer.multimodal", "orchestrator.multimodal", "inference.multimodal")

# [slurm] → inference: a multi-node RL run drives its inference deployment under
# the same SLURM allocation, so the nested inference inherits [slurm]. This is
Expand Down
3 changes: 3 additions & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -133,9 +133,12 @@ members = [
]
# tau2-synth-v1 and tau3-bench-v1 pin `tau2` forks/revs that conflict with the
# canonical tau2-bench-v1; a workspace shares one lock, so these can't coexist.
# nemo-gym-weather-v1 needs verifiers[nemo-gym], whose nemo-gym==0.4.0 pins
# openai<=2.7.2 against verifiers' own openai>=2.9.0.
exclude = [
"deps/research-environments/environments/tool_use/tau2_synth_v1",
"deps/research-environments/environments/tool_use/tau3_bench_v1",
"deps/verifiers/environments/nemo_gym_weather_v1",
]

# ModelExpress 0.3.0 publishes protobuf<6 metadata, but its generated proto is
Expand Down
31 changes: 31 additions & 0 deletions skills/training/monitor-run/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -173,6 +173,37 @@ A few warnings are normal. Escalate when errors are persistent, growing, or hit
- **Trainer**: NCCL/CUDA errors, OOM, NaN loss or gradients.
- **Inference**: NCCL/CUDA errors, OOM, request timeouts.

### Multimodal image offload checks

v1 multimodal RL offloads images exactly once, at verifiers ingress: every image
content part is rewritten to a `file://` asset under the run image directory
before rendering, and renderers/inference/trainer all work from those refs. A
`data:` URL reaching a renderer ("requires offloaded file:// image assets")
means ingress was bypassed.

- The image directory comes from `[multimodal].offload_dir` in the resolved
config; unset, it defaults to a run-scoped path (`{output_dir}/assets/images`
or the hosted `RUN_ID` path). Each env-server resolves that into
`VF_RENDERER_IMAGE_OFFLOAD_DIR` for verifiers/renderers (protected — set via
`[multimodal].offload_dir`, not `env_vars`).
- While multimodal rollouts are in flight, image files should accumulate under
that directory. Zero files means the offload path isn't being exercised or
image preparation failed before request submission.
- Inference rejects bad refs with `invalid_mm_image_ref` 400s (hash mismatch,
fingerprint mismatch, unreadable asset) — grep the inference log.
- Inference caches materialized refs and logs
`mm materialize cache: hits=X misses=Y hit_rate=Z% bytes=A/B evictions=C`
every 1000 lookups. Hit rate should climb after turn 1 of multi-turn
multimodal rollouts; a stuck-at-zero hit rate with repeat images means the
cache is disabled or thrashing (sized by `PRIME_RL_MM_MATERIALIZE_CACHE_GB`,
default 2.0, `0` disables).
- The orchestrator raises on placeholder/token drift ("does not cover
image-typed tokens") before a sample ships — treat any occurrence as a bug,
not noise.
- Trainer metrics: `mm/images_materialized` and `time/mm_materialize`. Missing
image files fail the trainer hard ("raw image materialization failed") —
check whether something cleaned the offload directory mid-run.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

eh do we need this in skills ?

### Process tree

All processes use `setproctitle` so they're visible in `ps`/`htop`/`pstree`:
Expand Down
7 changes: 7 additions & 0 deletions src/prime_rl/entrypoints/env_server.py
Original file line number Diff line number Diff line change
@@ -1,3 +1,4 @@
import os
from functools import partial

from verifiers.v1 import pool_serve_kwargs
Expand All @@ -7,11 +8,17 @@
from prime_rl.orchestrator.utils import setup_env_server_logging
from prime_rl.utils.config import cli
from prime_rl.utils.process import set_proc_title
from prime_rl.utils.run_assets import IMAGE_OFFLOAD_DIR_ENV, resolve_image_offload_dir
from prime_rl.utils.utils import clean_exit


@clean_exit
def run_server(config: EnvServerConfig):
# Renderers offload images to the dir named by this env var; resolve it from
# this server's own config, letting an operator-set env win (multi-node).
os.environ.setdefault(
IMAGE_OFFLOAD_DIR_ENV, str(resolve_image_offload_dir(config.output_dir, config.multimodal, os.environ))
)
# ``serve.pool`` (static or elastic) sizes the server; a v0/legacy env runs through
# the bridge, a v1 env is a native env block — both speak the same serve protocol,
# so the orchestrator is agnostic. serve_env applies the logging setup in this process
Expand Down
15 changes: 12 additions & 3 deletions src/prime_rl/entrypoints/rl.py
Original file line number Diff line number Diff line change
Expand Up @@ -105,6 +105,8 @@ def write_subconfigs(config: RLConfig, output_dir: Path) -> None:
"serve": {**source_dict.get("serve", {}), "address": address},
"legacy": source_dict.get("legacy", {}),
"log": {"level": config.orchestrator.log.vf_level, "json_logging": config.orchestrator.log.json_logging},
"output_dir": str(config.orchestrator.output_dir),
"multimodal": to_toml_dict(config.multimodal),
}
with open(env_dir / f"{source.resolved_name}.toml", "wb") as f:
tomli_w.dump(env_server_dict, f)
Expand Down Expand Up @@ -157,6 +159,7 @@ def rl_local(config: RLConfig):
"WANDB_SHARED_MODE": "1",
"WANDB_SHARED_RUN_ID": os.environ.get("WANDB_SHARED_RUN_ID", uuid.uuid4().hex),
}
inherited_env = dict(os.environ)

# Validate client port matches inference server port
if config.inference is not None and not config.orchestrator.model.client.is_elastic:
Expand Down Expand Up @@ -200,7 +203,7 @@ def sigterm_handler(signum, frame):
inference_process = Popen(
inference_cmd,
env={
**os.environ,
**inherited_env,
**DEFAULT_COMMON_ENV_VARS,
**DEFAULT_INFERENCE_ENV_VARS,
**config.env_vars,
Expand Down Expand Up @@ -290,7 +293,7 @@ def sigterm_handler(signum, frame):
stdout=log_file,
stderr=log_file,
env={
**os.environ,
**inherited_env,
**DEFAULT_COMMON_ENV_VARS,
"LOGURU_FORCE_COLORS": "1",
"WANDB_PROGRAM": "uv run rl",
Expand Down Expand Up @@ -339,7 +342,7 @@ def sigterm_handler(signum, frame):
trainer_process = Popen(
trainer_cmd,
env={
**os.environ,
**inherited_env,
**DEFAULT_COMMON_ENV_VARS,
**DEFAULT_TRAINER_ENV_VARS,
"LOGURU_FORCE_COLORS": "1",
Expand Down Expand Up @@ -437,6 +440,9 @@ def write_slurm_script(config: RLConfig, config_dir: Path, script_path: Path) ->
kv_offload_disk_path=str(offload.disk.path) if (is_mooncake and offload.disk is not None) else "",
kv_offload_device_name=offload.device_name if is_mooncake else "",
)
image_offload_dir = (
os.path.expanduser(str(config.multimodal.offload_dir)) if config.multimodal.offload_dir is not None else ""
)
Comment thread
cursor[bot] marked this conversation as resolved.

# Per-component env vars: launcher defaults (shared + multi-node-specific) with the
# user's config merged on top. Runtime wiring stays in the template.
Expand All @@ -463,6 +469,7 @@ def write_slurm_script(config: RLConfig, config_dir: Path, script_path: Path) ->
**config.slurm.template_vars,
config_path=config_dir / RL_TOML,
output_dir=config.output_dir,
image_offload_dir=image_offload_dir,
gpus_per_node=config.deployment.gpus_per_node,
)
elif config.inference is not None and config.inference.deployment.type == "disaggregated":
Expand All @@ -474,6 +481,7 @@ def write_slurm_script(config: RLConfig, config_dir: Path, script_path: Path) ->
config_dir=config_dir,
output_dir=config.output_dir,
orchestrator_output_dir=config.orchestrator.output_dir,
image_offload_dir=image_offload_dir,
num_train_nodes=config.deployment.num_train_nodes,
num_infer_nodes=infer_deploy.num_nodes * config.deployment.num_infer_replicas,
nodes_per_infer_replica=infer_deploy.num_nodes,
Expand Down Expand Up @@ -515,6 +523,7 @@ def write_slurm_script(config: RLConfig, config_dir: Path, script_path: Path) ->
config_dir=config_dir, # TODO: should prob have each subconfig path separately
output_dir=config.output_dir,
orchestrator_output_dir=config.orchestrator.output_dir,
image_offload_dir=image_offload_dir,
num_train_nodes=config.deployment.num_train_nodes,
num_infer_nodes=config.deployment.total_infer_nodes,
nodes_per_infer_replica=config.deployment.infer_nodes_per_replica,
Expand Down
Loading