Description
llm: routes are wired correctly. The selector parses, hook_pair_for_entity maps it to cmf.llm_input / cmf.llm_output, the visitor installs a handler on each, and visitor_e2e pins that an llm: route fires on the LLM hooks and not on the tool hooks. What is missing is not the routing, it is what the route can see, and that splits into three cases that want three different answers.
The proposal is to leave the hook payload alone and put the missing request context on extensions. MessagePayload { message: Message } is shared by every CMF entity, so changing it to hold a conversation ripples into tools, prompts, and resources for a reason that only concerns LLMs. Extensions are already the "everything else about this request" channel, they are capability-gated, which is what a transcript needs, and adding one is a pattern this codebase has executed twelve times. Nothing here is a breaking change for an existing plugin.
Three tiers, and only one of them is a missing carrier
Reachable from APL today. llm.model_id, llm.provider, llm.capabilities. agent.input, agent.session_id, agent.conversation_id, agent.turn, agent.agent_id, agent.parent_agent_id, agent.conversation.summary, agent.conversation.topics. On the post side completion.stop_reason, completion.model, completion.tokens.input / output / total, completion.latency_ms. Plus the concatenated text of the message itself as args or result, which is what makes whole-message redaction work. This is a real surface and it is not written down anywhere, which is how the first draft of this issue managed to claim an llm: route sees nothing.
Carried, but not reachable from policy. AgentExtension.conversation.history is Vec<serde_json::Value>, and the CMF bridge deliberately does not flatten it: the comment in crates/ppe-apl-cmf/src/agent.rs says it is too unstructured, and that a policy wanting history should call a plugin. So a plugin holding read_agent can read the turns and no APL predicate can. The bridge's position is right and the type is what is wrong. The field doc calls it "recent conversation history (lightweight summaries)", so today the host, the plugin, and the policy author agree on its contents by convention rather than by type.
No carrier at all. The system prompt. The tool definitions offered to the model. max_tokens, temperature, top_p, tool_choice, stop sequences, the streaming flag. LLMExtension is model_id, provider, and capabilities, and that is the whole of it. So "this model may not be offered this tool" and "no request over N output tokens" cannot be written at any layer, plugin included.
Proposal
1. Carry the request
Add the missing fields, either on LLMExtension or on a new request-side extension next to it, and bridge the scalars into the bag:
llm.max_tokens, llm.temperature, llm.top_p, llm.stream, llm.tool_choice
llm.offered_tools StringSet
llm.system_prompt String, or a digest of it if carrying the text is too much
llm.offered_tools as a StringSet is the highest value item in this issue. The payload walk already promotes scalar arrays to sets and the language already has contains, so
- "llm.offered_tools contains 'send_email' and not subject.roles contains 'finance': deny"
is writable with no change to the language, no change to the evaluator, and no change to the payload. Tool-definition governance is the LLM policy with no workaround at any layer today, and this turns it into a config line.
2. Give history a type, and keep it off the bag
Change ConversationContext.history from Vec<serde_json::Value> to a typed sequence. Vec<Message> is the natural choice: CMF already defines Message, hosts already build them, and a scanner plugin walks its ContentParts with the same code it uses on the current turn. Extensions are Option<Arc<..>>, so carrying it costs a refcount bump per hook rather than a clone.
Keep the bridge's decision not to flatten it. Once the type is real, "too unstructured, call a plugin" stops being a limitation and becomes a documented contract, and the deliverable is a reference scanner plugin under reference/plugins/ alongside pii-scanner, not new bag keys. Cheap summaries over the history can still land on the bag if they earn their place: a turn count, the set of roles present.
3. Say what already works
docs/ gains the attribute surface an llm: route can read, per phase, with a worked config. The tier-one list above is the content.
Open questions
- Can the host see any of this? For a gateway sitting on a provider HTTP API, the system prompt, the offered tool definitions, and the sampling parameters are all in the request body, so yes. For an in-process SDK hook it depends where the hook sits. I have not checked the praxis filter. This decides whether part 1 is worth building, so answer it first.
- Does the system prompt belong in full, or as a digest? Full text makes
llm.system_prompt contains '...' writable and puts a possibly large string on every request's bag. A digest supports pinning ("the system prompt is the one we shipped") and nothing else. Pinning is the policy I would expect to want more often.
- Typing
history as Vec<Message> is a breaking change to ConversationContext. It is pre-1.0 and the field appears to have no in-tree producer, so the blast radius is hosts rather than this workspace. Worth confirming against praxis before committing.
- What does
entity_name mean for an llm: route? Today it matches the model id, so llm: gpt-4 and llm: "gpt-*" are the natural spellings, and the second does not currently work. See issue-glob-selector-never-authorizes.md, which blocks this.
restrict emits a backend candidate constraint, and constraining which model backend serves a call is the obvious LLM use of it. Whether the host honours it on the LLM path is untested here.
- Elicitation on an LLM call. The pending-elicitation protocol answers
-32120 and expects the caller to resend echoing X-Policy-Elicitation-Id. On a tool call the caller is an MCP client that can do that; on an LLM call it is the agent runtime. Is human approval on an LLM call in scope?
Acceptance criteria
- A carrier for the system prompt, the offered tool definitions, and the sampling parameters, with the scalars bridged, or a recorded decision that they stay out of scope and why.
llm.offered_tools supports a contains predicate, with an e2e test that denies on it.
ConversationContext.history has a defined element type, and a stated answer on whether policy reaches it directly or through a plugin.
- A reference plugin that scans typed history, if the answer is "through a plugin".
- The attribute surface an
llm: route can read is documented per phase, with a worked config.
- e2e tests assert on a deny. An allow cannot distinguish a route that evaluated from one that never fired.
Description
llm:routes are wired correctly. The selector parses,hook_pair_for_entitymaps it tocmf.llm_input/cmf.llm_output, the visitor installs a handler on each, andvisitor_e2epins that anllm:route fires on the LLM hooks and not on the tool hooks. What is missing is not the routing, it is what the route can see, and that splits into three cases that want three different answers.The proposal is to leave the hook payload alone and put the missing request context on extensions.
MessagePayload { message: Message }is shared by every CMF entity, so changing it to hold a conversation ripples into tools, prompts, and resources for a reason that only concerns LLMs. Extensions are already the "everything else about this request" channel, they are capability-gated, which is what a transcript needs, and adding one is a pattern this codebase has executed twelve times. Nothing here is a breaking change for an existing plugin.Three tiers, and only one of them is a missing carrier
Reachable from APL today.
llm.model_id,llm.provider,llm.capabilities.agent.input,agent.session_id,agent.conversation_id,agent.turn,agent.agent_id,agent.parent_agent_id,agent.conversation.summary,agent.conversation.topics. On the post sidecompletion.stop_reason,completion.model,completion.tokens.input/output/total,completion.latency_ms. Plus the concatenated text of the message itself asargsorresult, which is what makes whole-message redaction work. This is a real surface and it is not written down anywhere, which is how the first draft of this issue managed to claim anllm:route sees nothing.Carried, but not reachable from policy.
AgentExtension.conversation.historyisVec<serde_json::Value>, and the CMF bridge deliberately does not flatten it: the comment incrates/ppe-apl-cmf/src/agent.rssays it is too unstructured, and that a policy wanting history should call a plugin. So a plugin holdingread_agentcan read the turns and no APL predicate can. The bridge's position is right and the type is what is wrong. The field doc calls it "recent conversation history (lightweight summaries)", so today the host, the plugin, and the policy author agree on its contents by convention rather than by type.No carrier at all. The system prompt. The tool definitions offered to the model.
max_tokens,temperature,top_p,tool_choice, stop sequences, the streaming flag.LLMExtensionismodel_id,provider, andcapabilities, and that is the whole of it. So "this model may not be offered this tool" and "no request over N output tokens" cannot be written at any layer, plugin included.Proposal
1. Carry the request
Add the missing fields, either on
LLMExtensionor on a new request-side extension next to it, and bridge the scalars into the bag:llm.offered_toolsas aStringSetis the highest value item in this issue. The payload walk already promotes scalar arrays to sets and the language already hascontains, so- "llm.offered_tools contains 'send_email' and not subject.roles contains 'finance': deny"is writable with no change to the language, no change to the evaluator, and no change to the payload. Tool-definition governance is the LLM policy with no workaround at any layer today, and this turns it into a config line.
2. Give history a type, and keep it off the bag
Change
ConversationContext.historyfromVec<serde_json::Value>to a typed sequence.Vec<Message>is the natural choice: CMF already definesMessage, hosts already build them, and a scanner plugin walks itsContentParts with the same code it uses on the current turn. Extensions areOption<Arc<..>>, so carrying it costs a refcount bump per hook rather than a clone.Keep the bridge's decision not to flatten it. Once the type is real, "too unstructured, call a plugin" stops being a limitation and becomes a documented contract, and the deliverable is a reference scanner plugin under
reference/plugins/alongsidepii-scanner, not new bag keys. Cheap summaries over the history can still land on the bag if they earn their place: a turn count, the set of roles present.3. Say what already works
docs/gains the attribute surface anllm:route can read, per phase, with a worked config. The tier-one list above is the content.Open questions
llm.system_prompt contains '...'writable and puts a possibly large string on every request's bag. A digest supports pinning ("the system prompt is the one we shipped") and nothing else. Pinning is the policy I would expect to want more often.historyasVec<Message>is a breaking change toConversationContext. It is pre-1.0 and the field appears to have no in-tree producer, so the blast radius is hosts rather than this workspace. Worth confirming against praxis before committing.entity_namemean for anllm:route? Today it matches the model id, sollm: gpt-4andllm: "gpt-*"are the natural spellings, and the second does not currently work. Seeissue-glob-selector-never-authorizes.md, which blocks this.restrictemits a backend candidate constraint, and constraining which model backend serves a call is the obvious LLM use of it. Whether the host honours it on the LLM path is untested here.-32120and expects the caller to resend echoingX-Policy-Elicitation-Id. On a tool call the caller is an MCP client that can do that; on an LLM call it is the agent runtime. Is human approval on an LLM call in scope?Acceptance criteria
llm.offered_toolssupports acontainspredicate, with an e2e test that denies on it.ConversationContext.historyhas a defined element type, and a stated answer on whether policy reaches it directly or through a plugin.llm:route can read is documented per phase, with a worked config.