Problem
The Agent Worker has two materially different completion paths:
- POST /v1/complete builds the active skill tools and delegate_to_agent, then runs the multi-round agent loop through workers/agent/src/agents/runner.ts.
- POST /v1/stream injects the prompt and memory but forwards the request directly to the LLM gateway without tools and without the agent runner.
The deployed CLI uses /v1/stream, so Agent Worker mode can stream ordinary text but cannot use Google, GitHub, MantisHub, or Todoist tools and cannot delegate to specialist agents. Telegram uses /v1/complete and therefore has a different capability set. This makes the same VeeClaw deployment behave differently depending on the channel.
Desired behavior
Streaming through the Agent Worker should preserve the same agent capabilities as non-streaming completion while retaining incremental text delivery for normal responses. Tool calls and delegation should be handled server-side and should not leak as raw protocol artifacts to end users.
Design considerations
Choose and document one implementation strategy:
- Implement a streaming agent loop that forwards content deltas, buffers tool-call deltas until arguments are complete, executes connector and delegation calls, appends assistant/tool messages, and continues for the configured maximum number of rounds.
- Use a hybrid path: stream directly when no tool call is needed, but fall back to an internal non-streaming agent loop when the model requests tools, then stream the final response. This is simpler but may delay the first token for tool-driven requests.
- Explicitly route tool-capable requests to /v1/complete and make the limitation visible to clients. This is the minimum fallback, but it does not provide true parity.
The implementation should account for partial SSE chunks, parallel tool execution, delegate calls, connector failures, maximum rounds, cancellation, and the final response capture used by background memory processing.
Acceptance criteria
- A no-tool request to /v1/stream still delivers incremental text and a clean completion signal.
- A request that needs Google, GitHub, MantisHub, or Todoist invokes the same routing and execution behavior as /v1/complete.
- A request that needs specialist delegation works from the deployed CLI path.
- Tool calls are not executed twice when SSE frames are split or retried.
- Connector errors and maximum-round exhaustion produce a useful user-facing response.
- The final assistant response is captured once for memory updates.
- Regression tests cover plain streaming, tool calls, delegation, multi-round execution, partial SSE frames, and connector failures.
- The CLI and Telegram paths document their supported model/tool behavior consistently.
Problem
The Agent Worker has two materially different completion paths:
The deployed CLI uses /v1/stream, so Agent Worker mode can stream ordinary text but cannot use Google, GitHub, MantisHub, or Todoist tools and cannot delegate to specialist agents. Telegram uses /v1/complete and therefore has a different capability set. This makes the same VeeClaw deployment behave differently depending on the channel.
Desired behavior
Streaming through the Agent Worker should preserve the same agent capabilities as non-streaming completion while retaining incremental text delivery for normal responses. Tool calls and delegation should be handled server-side and should not leak as raw protocol artifacts to end users.
Design considerations
Choose and document one implementation strategy:
The implementation should account for partial SSE chunks, parallel tool execution, delegate calls, connector failures, maximum rounds, cancellation, and the final response capture used by background memory processing.
Acceptance criteria