Skip to content

Feature: LLM model router as a first-class runtime - #511

Open
leprachuan wants to merge 6 commits into
devfrom
issue-506-llm-router
Open

leprachuan wants to merge 6 commits into
devfrom
issue-506-llm-router

Conversation

@leprachuan

Copy link
Copy Markdown
Owner

Summary

Implements #506 β€” a configurable LLM-based model router. Foster's refinement of the issue's design: the router is a first-class runtime (router), selectable anywhere any other runtime is (sessions, background tasks, delegation, agents.json primaries). Its decision-making "brain" is itself invoked via a full runtime dispatch β€” pick any existing runtime+model as the brain (e.g. wee + a local Ollama model, or copilot + haiku).

What it does

  • New router runtime: registered in get_all_runtimes() / get_available_runtimes() / check_runtime_available(), available only when routing is enabled (config or WEE_ROUTER_ENABLED) and its configured brain runtime is itself available.
  • SessionManager.run_router(): builds an eligible allowlist (configured pairs minus cooled-down/disabled/unavailable runtimes, always excluding router itself), short-circuits when 0 or 1 pairs are eligible, otherwise calls the brain LLM with a configurable prompt template and validates its strict-JSON reply against the allowlist.
  • Stickiness: each router-selected target's own sub-session id is tracked separately (session_data["router_sessions"]) from the session's top-level session_id, so repeated routing to the same runtime resumes its own conversation and reuses its prompt cache β€” without corrupting the top-level session_id some runtimes (copilot/claude/codex) write to as a side effect.
  • Failure handling: brain timeout/error/invalid output β†’ safe fallback pair β†’ agent's primary_runtime/primary_model as a last resort. Target-dispatch failures matching the same infra-failure patterns used for background-task fallback (429/quota/5xx/timeout) trigger a cooldown for that runtime and exactly one retry with the fallback pair. route() never raises.
  • Precedence: routing only happens while a session's runtime is literally router. /runtime <x> or /model <x> switches away from routing entirely (pinned until /runtime router); nothing is silently overridden.
  • New endpoints: GET/PUT /api/v1/router-config, POST /api/v1/router/test (dry-run a routing decision without touching any session), GET /api/v1/router/status (cooldowns, brain reachability).
  • config_schemas.RouterConfigSchema for API-facing validation (recursion guards, required prompt placeholders, positive timeouts).
  • Reference WebUI panel (RouterSettingsPanel.tsx + API client + CSS), matching the existing AgentSettingsPanel.tsx "reference/future-build" convention β€” live wiring into webui/dist/app.js is a follow-up, same status as that existing component today.

Regression Test Required

A test reproducing this feature's acceptance criteria is included: tests/test_issue506_llm_router.py (50 tests) covers valid/invalid/timeout decisions, allowlist enforcement, cooldown + single fallback retry, stickiness, config validation, and a guard against reintroducing the auto runtime name removed in #84.

Test plan

  • pytest tests/test_issue506_llm_router.py β€” 50/50 passing
  • Full suite run before/after in a clean venv to confirm no regressions (baseline: 587 failed/2353 passed; this branch: 578 failed/2362 passed β€” the swing is exactly these new tests; remaining failures are pre-existing local-venv gaps e.g. missing keyring)
  • Deploy to dev (192.168.1.100) and validate live: enable with a local Ollama brain, confirm routing to different pairs for different prompt types, cooldown/fallback under a simulated failure, stickiness across repeated messages, and that router disappears cleanly when disabled

Co-Authored-By: Claude Sonnet 5 noreply@anthropic.com

leprachuan and others added 2 commits August 26, 2026 10:03
Adds a "router" runtime that picks a target runtime+model per message
using a configurable "brain" LLM call (itself dispatched through the
existing runtime infrastructure), with a configurable routing prompt,
an explicit allowlist, per-runtime cooldown on infra failures, safe
fallback to the agent's primary pair, and stickiness that resumes the
previously-routed runtime's own sub-session so prompt caching still
applies. Explicit /runtime or /model selection simply switches the
session away from "router" (and back), so nothing is silently
overridden.

- llm_router.py: config load/save/validate, decision parsing,
  cooldown tracking, and LLMRouter.route() (pure, agent_manager-free,
  fully unit-testable).
- config_schemas.py: RouterConfigSchema for the PUT /api/v1/router-config
  API validator.
- agent_manager.py: "router" runtime registration, run_router() +
  supporting helpers, dispatch branch, and
  GET/PUT /api/v1/router-config, POST /api/v1/router/test,
  GET /api/v1/router/status.
- tests/test_issue506_llm_router.py: 50 tests covering valid/invalid
  decisions, timeouts, allowlist enforcement, cooldown/fallback retry,
  stickiness, config validation, and the #84 "auto"-runtime guard.
- webui/: reference RouterSettingsPanel.tsx + API client, following
  the existing AgentSettingsPanel.tsx convention (live WebUI wiring
  into webui/dist/app.js is a follow-up).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Found via live testing on dev (192.168.1.100): /runtime set router was
rejected as "Unknown runtime" because _slash_runtime() validates
against its own hardcoded list, separate from get_available_runtimes()
/ check_runtime_available(). Also adds 'router' to GET /api/v1/models'
known_runtimes set and the CLI --runtime argparse choices, both of
which had the same gap.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@leprachuan

Copy link
Copy Markdown
Owner Author

Live validation on dev (192.168.1.100)

Deployed issue-506-llm-router to the dev API (port 8001) and validated end-to-end:

  • Disabled by default, no regression: router absent from GET /api/v1/runtimes until enabled via PUT /api/v1/router-config.
  • Real routing discrimination: with brain=copilot/auto and allowlist {claude-sdk/haiku: coding} + {copilot/auto: trivial}, a quicksort request routed to claude-sdk/haiku with a sensible LLM-authored reason; a "hey whats up" greeting routed to copilot/auto.
  • Full session dispatch + stickiness: created a session pinned to runtime=router, sent two turns. Turn 1 log: [Router] decision=copilot/auto source=router latency_ms=7550 reason_len=165 β€” no prompt content logged. Turn 2 correctly resumed the same underlying Copilot session ([Session] Resuming Copilot session: 431cfd78-...) instead of starting fresh, and the reply correctly referenced turn 1's code β€” confirms prompt-cache reuse via the router-tracked sub-session actually works.
  • No single point of failure: pointed the brain at a runtime with an unrelated pre-existing bug (see below) to force a brain failure. The real request still completed via the fallback pair (decision=copilot/auto source=fallback) rather than erroring out.
  • Precedence: /runtime set router enables routing; a routed message dispatched correctly; /runtime set copilot switched the session back to a pinned concrete runtime, confirmed via session status. Nothing silently overridden either direction.
  • Bug found and fixed during live testing (now in 9503633): /runtime set router was initially rejected as "Unknown runtime" β€” _slash_runtime() validates against its own hardcoded runtime list, separate from get_available_runtimes()/check_runtime_available(). Also added router to GET /api/v1/models's known-runtimes set and the CLI --runtime argparse choices, which had the same gap.

Pre-existing issues found along the way (unrelated to this PR, not fixed here)

  • The wee runtime currently fails all dispatches on this dev box with Session error: External tool "bash" conflicts with a built-in tool of the same name (Copilot SDK). Separately, Ollama models without a large num_ctx baked in silently return empty content against Wee's ~14k-char system prompt (matches an existing internal note on this). Worth its own issue if wee is meant to be usable as a router brain/target today.
  • telegram-bot-listener.service / webex-connector.service on this host are in a long-running crash loop (Failed to determine user credentials: No such process, restart counter >282,000) β€” looks like a stale systemd User= directive, unrelated to this change.
  • /opt/n8n-copilot-shim-dev on this host is a symlink to /opt/n8n-copilot-shim, not a separate checkout β€” so "dev" and the locally-named "prod" path are the same directory here (real prod runs elsewhere, on lepbuntu, and was not touched). Also, /opt/n8n-copilot-shim had no config/ directory and the service's n8n user lacked write permission to create one, which blocked the first PUT /api/v1/router-config until I created it with correct ownership.

Happy to file separate issues for these if wanted.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant