Repository navigation
Fall back when a topic loses every leader, count the drop, and refresh again (release-0.1) - #95
Conversation
…h again (release-0.1)
stephclay
left a comment
There was a problem hiding this comment.
Two inline findings on the new leader-drop path. The knownTopics change is correct. A topic goes into the set only when Metadata listed partitions for it and the client kept none. Topics with a topic-level error or no partitions stay out, and this matches the README.
— via Claude Code
…Leader below 0 as a drop
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The exported AgentPool.Refresh signature introduces a source-incompatible API change, and the affected configuration documentation is stale.
Review effort: Balanced
Findings: 1
Open (2)
What changed in this PR
Adds resilient routing when metadata drops all leaders for a known topic, with observability and follow-up refreshes.
Changes:
- Falls back to live agents while preserving unknown/no-leader behavior.
- Counts, logs, and refreshes after excluded leaders.
- Adds routing, metadata, metric, and integration coverage.
| File | Description |
|---|---|
README.md |
Documents fallback behavior. |
pkg/wgo/partition_assignment.go |
Tracks known topics and leaderless partitions. |
pkg/wgo/partition_assignment_test.go |
Tests fallback selection. |
pkg/wgo/metrics.go |
Adds the dropped-leader counter. |
pkg/wgo/demoter_test.go |
Updates constructor usage. |
pkg/wgo/client.go |
Records drops and schedules refreshes. |
pkg/wgo/client_test.go |
Tests routing, metrics, and nudging. |
pkg/wgo/agentpool.go |
Detects and reports excluded leaders. |
pkg/wgo/agentpool_test.go |
Tests metadata classification. |
docs/internal/metrics.md |
Documents the metric category. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
…interval's leader-drop use


Summary
Follow-up to the live-agent fallback in #88 (
main) and #93 (release-0.1). This PR is againstrelease-0.1. Port tomainafter it merges.#93 falls back only when some other partition of the topic still has a leader. If every leader for a topic is excluded in one refresh,
Candidatesstill returns nil andProduceSyncfails the whole batch, including healthy topics. This change treats that topic as known and hashes it onto a live agent. A topic Metadata has never returned, and a topic-level error, stay unknown so they still trigger an on-demand refresh.A refresh that excludes a leader increments
warpstream_agentpool_leader_dropped_total(once per excluded leader, including the constructor) and logs the count plus the first excluded leader (first_topic,first_partition,first_node_id). The background refresh nudges another fetch. After a periodic tick that follow-up starts at once; later repeats wait outOnDemandMetadataRefreshInterval(default 1s). The constructor counts and logs, and does not nudge. Produce does not wait.This also covers the exact broker/topic mismatch RequestCachedMetadata can return, described upstream in twmb/franz-go#1481: a leader id present in Topics but missing from that same response's Brokers now hashes onto a live agent instead of failing the batch.
Also addresses the #88 review comment on the benchmark sink: the comment now describes the benchmark's own allocation count, not how
AgentPool.Refreshuses the result.Left for follow-up PRs, still part of the leader-map fallback:
ProduceSyncfail the whole batch. The next PR routes the partitions that have a candidate and fails only the ones that do not.