Skip to content

kgo: return cached metadata brokers with their topics - #1481

Open
koloss2001 wants to merge 1 commit into
twmb:masterfrom
koloss2001:kgo-cached-metadata-brokers
Open

koloss2001 wants to merge 1 commit into
twmb:masterfrom
koloss2001:kgo-cached-metadata-brokers

Conversation

@koloss2001

Copy link
Copy Markdown

What

RequestCachedMetadata returns Brokers and Topics from the same metadata response.

When the call returns topics, Brokers, ControllerID, and ClusterID come from the broker list saved with those topics. A cache hit whose topics come from two responses fetches that set once more and uses that response. A request with no topics still returns the live connection table. cl.brokers is unchanged, and there is no new API.

Why

We're building warpstream-go, a produce-only Kafka client on top of franz-go. Its refresh calls RequestCachedMetadata and uses that single return value as the snapshot for the call: dialable brokers come from Brokers, and each partition leader comes from Topics in the same struct. A non-negative Leader that is absent from Brokers is not a broker that response can name, so the partition cannot be routed from it. kadm.Metadata returns the same pair. kgo's own producer does not read this struct.

A Metadata response is one cluster view. These two answers are each consistent:

Response A, before node 3 is gone:

  • Brokers: 1, 2, 3
  • topic orders, partition 0, leader 3

Response B, after:

  • Brokers: 1, 2
  • topic orders, partition 0, leader 1

RequestCachedMetadata could return a third shape, which neither response contained:

  • Brokers: 1, 2
  • topic orders, partition 0, leader 3

Two ways that happened. The call fetched A and stored A's topics, then another fetch (the metadata loop, Ping, or fetchBrokerMetadata) installed B's broker list into cl.brokers before the helper copied it. Or the call did not fetch: topics were still inside limit, and a brokers-only Metadata had already replaced cl.brokers and left the topic cache alone.

ControllerID is the value that response sent, including -1. updateMetadataBrokers still ignores -1 so the client can dial the last controller it knew. That id is not copied into this snapshot, because it may not be in this response's broker list.

Happy to adjust the approach.

RequestCachedMetadata copies topics from the metadata cache and brokers
from the live connection table, so the two halves can come from different
Metadata responses. A caller that routes from that one struct can see a
partition leader that is not in Brokers.
@twmb

twmb commented Oct 7, 2026

Copy link
Copy Markdown
Owner

Simpler alternative - wdyt? #1482

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants