Skip to content

Fall back when a topic loses every leader, count the drop, and refresh again - #96

Merged
koloss2001 merged 1 commit into
mainfrom
koloss2001/leader-drop-followup
Oct 2, 2026
Merged

koloss2001 merged 1 commit into
mainfrom
koloss2001/leader-drop-followup

Conversation

@koloss2001

Copy link
Copy Markdown
Contributor

Summary

Port of #95 onto main.

#88 falls back only when some other partition of the topic still has a leader. If every named leader for a topic is excluded in one refresh, Candidates still returns nil and ProduceSync fails the whole batch, including healthy topics. This change treats that topic as known and hashes it onto a live agent. A topic Metadata has never returned, and a topic-level error, stay unknown so they still trigger an on-demand refresh. A partition Metadata reported with no leader (Leader below 0) does not hash and is not counted as a drop. It fails, and the existing nil-candidate path refreshes.

A refresh that excludes a leader increments warpstream_agentpool_leader_dropped_total (once per excluded leader, including the constructor) and logs the count plus the first excluded leader (first_topic, first_partition, first_node_id). The background refresh nudges another fetch. After a periodic tick that follow-up starts at once; later repeats wait out OnDemandMetadataRefreshInterval (default 1s). The constructor counts and logs, and does not nudge. Produce does not wait. AgentPool.Refresh stays ([]int32, error). The drop details are on the unexported refresh.

This also covers the broker/topic mismatch RequestCachedMetadata can return, described upstream in twmb/franz-go#1481: a leader id present in Topics but missing from that same response's Brokers now hashes onto a live agent instead of failing the batch.

Also addresses the #88 review comment on the benchmark sink: the comment now describes the benchmark's own allocation count, not how AgentPool.Refresh uses the result.

Left for follow-up PRs, still part of the leader-map fallback:

  • Partition isolation. A topic Metadata has never returned, a topic-level error, an empty agent pool, or a Leader below 0 still makes ProduceSync fail the whole batch. The next PR routes the partitions that have a candidate and fails only the ones that do not.
  • Stale leader still in the map. The named leader is present but the agent is gone. A later PR refreshes on unknown broker, dial failure, and a short list of topology errors. It does not retry the write already in flight. Timeouts and connection resets stay out.

@koloss2001
koloss2001 requested review from a team as code owners October 2, 2026 18:01
@koloss2001
koloss2001 merged commit a0eabf2 into main Oct 2, 2026
16 checks passed
@koloss2001
koloss2001 deleted the koloss2001/leader-drop-followup branch October 2, 2026 18:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants