Skip to content

Restore tool and delegation parity for streaming responses #8

Description

@vboctor

Problem

The Agent Worker has two materially different completion paths:

  • POST /v1/complete builds the active skill tools and delegate_to_agent, then runs the multi-round agent loop through workers/agent/src/agents/runner.ts.
  • POST /v1/stream injects the prompt and memory but forwards the request directly to the LLM gateway without tools and without the agent runner.

The deployed CLI uses /v1/stream, so Agent Worker mode can stream ordinary text but cannot use Google, GitHub, MantisHub, or Todoist tools and cannot delegate to specialist agents. Telegram uses /v1/complete and therefore has a different capability set. This makes the same VeeClaw deployment behave differently depending on the channel.

Desired behavior

Streaming through the Agent Worker should preserve the same agent capabilities as non-streaming completion while retaining incremental text delivery for normal responses. Tool calls and delegation should be handled server-side and should not leak as raw protocol artifacts to end users.

Design considerations

Choose and document one implementation strategy:

  1. Implement a streaming agent loop that forwards content deltas, buffers tool-call deltas until arguments are complete, executes connector and delegation calls, appends assistant/tool messages, and continues for the configured maximum number of rounds.
  2. Use a hybrid path: stream directly when no tool call is needed, but fall back to an internal non-streaming agent loop when the model requests tools, then stream the final response. This is simpler but may delay the first token for tool-driven requests.
  3. Explicitly route tool-capable requests to /v1/complete and make the limitation visible to clients. This is the minimum fallback, but it does not provide true parity.

The implementation should account for partial SSE chunks, parallel tool execution, delegate calls, connector failures, maximum rounds, cancellation, and the final response capture used by background memory processing.

Acceptance criteria

  • A no-tool request to /v1/stream still delivers incremental text and a clean completion signal.
  • A request that needs Google, GitHub, MantisHub, or Todoist invokes the same routing and execution behavior as /v1/complete.
  • A request that needs specialist delegation works from the deployed CLI path.
  • Tool calls are not executed twice when SSE frames are split or retried.
  • Connector errors and maximum-round exhaustion produce a useful user-facing response.
  • The final assistant response is captured once for memory updates.
  • Regression tests cover plain streaming, tool calls, delegation, multi-round execution, partial SSE frames, and connector failures.
  • The CLI and Telegram paths document their supported model/tool behavior consistently.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions