Repository navigation
Feature: LLM model router as a first-class runtime - #511
Open
leprachuan wants to merge 6 commits into
Open
leprachuan wants to merge 6 commits into
leprachuan wants to merge 6 commits into
Conversation
Adds a "router" runtime that picks a target runtime+model per message using a configurable "brain" LLM call (itself dispatched through the existing runtime infrastructure), with a configurable routing prompt, an explicit allowlist, per-runtime cooldown on infra failures, safe fallback to the agent's primary pair, and stickiness that resumes the previously-routed runtime's own sub-session so prompt caching still applies. Explicit /runtime or /model selection simply switches the session away from "router" (and back), so nothing is silently overridden. - llm_router.py: config load/save/validate, decision parsing, cooldown tracking, and LLMRouter.route() (pure, agent_manager-free, fully unit-testable). - config_schemas.py: RouterConfigSchema for the PUT /api/v1/router-config API validator. - agent_manager.py: "router" runtime registration, run_router() + supporting helpers, dispatch branch, and GET/PUT /api/v1/router-config, POST /api/v1/router/test, GET /api/v1/router/status. - tests/test_issue506_llm_router.py: 50 tests covering valid/invalid decisions, timeouts, allowlist enforcement, cooldown/fallback retry, stickiness, config validation, and the #84 "auto"-runtime guard. - webui/: reference RouterSettingsPanel.tsx + API client, following the existing AgentSettingsPanel.tsx convention (live WebUI wiring into webui/dist/app.js is a follow-up). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Found via live testing on dev (192.168.1.100): /runtime set router was rejected as "Unknown runtime" because _slash_runtime() validates against its own hardcoded list, separate from get_available_runtimes() / check_runtime_available(). Also adds 'router' to GET /api/v1/models' known_runtimes set and the CLI --runtime argparse choices, both of which had the same gap. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Owner
Author
Live validation on dev (192.168.1.100)Deployed
Pre-existing issues found along the way (unrelated to this PR, not fixed here)
Happy to file separate issues for these if wanted. |
This was referenced Aug 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Implements #506 β a configurable LLM-based model router. Foster's refinement of the issue's design: the router is a first-class runtime (
router), selectable anywhere any other runtime is (sessions, background tasks, delegation, agents.json primaries). Its decision-making "brain" is itself invoked via a full runtime dispatch β pick any existing runtime+model as the brain (e.g.wee+ a local Ollama model, orcopilot+ haiku).What it does
routerruntime: registered inget_all_runtimes()/get_available_runtimes()/check_runtime_available(), available only when routing is enabled (config orWEE_ROUTER_ENABLED) and its configured brain runtime is itself available.SessionManager.run_router(): builds an eligible allowlist (configured pairs minus cooled-down/disabled/unavailable runtimes, always excludingrouteritself), short-circuits when 0 or 1 pairs are eligible, otherwise calls the brain LLM with a configurable prompt template and validates its strict-JSON reply against the allowlist.session_data["router_sessions"]) from the session's top-levelsession_id, so repeated routing to the same runtime resumes its own conversation and reuses its prompt cache β without corrupting the top-levelsession_idsome runtimes (copilot/claude/codex) write to as a side effect.primary_runtime/primary_modelas a last resort. Target-dispatch failures matching the same infra-failure patterns used for background-task fallback (429/quota/5xx/timeout) trigger a cooldown for that runtime and exactly one retry with the fallback pair.route()never raises.router./runtime <x>or/model <x>switches away from routing entirely (pinned until/runtime router); nothing is silently overridden.GET/PUT /api/v1/router-config,POST /api/v1/router/test(dry-run a routing decision without touching any session),GET /api/v1/router/status(cooldowns, brain reachability).config_schemas.RouterConfigSchemafor API-facing validation (recursion guards, required prompt placeholders, positive timeouts).RouterSettingsPanel.tsx+ API client + CSS), matching the existingAgentSettingsPanel.tsx"reference/future-build" convention β live wiring intowebui/dist/app.jsis a follow-up, same status as that existing component today.Regression Test Required
A test reproducing this feature's acceptance criteria is included:
tests/test_issue506_llm_router.py(50 tests) covers valid/invalid/timeout decisions, allowlist enforcement, cooldown + single fallback retry, stickiness, config validation, and a guard against reintroducing theautoruntime name removed in #84.Test plan
pytest tests/test_issue506_llm_router.pyβ 50/50 passingkeyring)routerdisappears cleanly when disabledCo-Authored-By: Claude Sonnet 5 noreply@anthropic.com