You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Add hedge-trigger, final-outcome, attempted-payload and excluded-leader metrics - #87
The existing metrics cannot show how these paths behave. This PR adds metrics to quantify the changes, monitor the new behaviour and troubleshoot it, fine-tune wgo, and measure how well it mitigates WarpStream issues.
It does not change routing, hedging, timing or retry behaviour.
Metric
Labels
What it tells you
warpstream_produce_hedge_triggers_total
trigger: latency, primary_failure, demoted_probe
What starts extra hedging.
warpstream_produce_hedge_trigger_wins_total
trigger
How often each trigger's hedge wins. Wins / triggers is the win rate.
Why a Hedger call failed. Counted once per failed call, not per Produce call or wire request. Successes and cancellations are not counted; successes are in warpstream_produce_requests_attempts.
warpstream_agentpool_agents_changed_total
direction: added, removed
Agents added or removed by a metadata refresh.
warpstream_agentpool_excluded_leaders (gauge)
none
Partitions whose leader is currently excluded. Goes back to 0 when it clears.
Records and compressed bytes sent, including failed and canceled attempts.
Things to know
warpstream_hedge_wins_total now counts only wins that were actually returned. Before, a fallback that finished but lost to the primary could still count. The existing hedge win-ratio panel will need to be updated.
excluded_leaders replaces warpstream_agentpool_leader_dropped_total. The old counter is not in v0.1.1 or any deployed version. Its rate depended on how often we refresh, so it fell as the backoff slowed refreshes, which looked like recovery.
Failure-reason order: terminal error first (from the primary or a retry), then work deadline, then candidates exhausted. Cancellation is not counted. A deadline that has been reached counts as expired even if its timer has not fired yet.
This PR adds 3 new metrics (warpstream_produce_hedge_triggers_total, warpstream_produce_final_outcome_total, warpstream_agentpool_agents_changed_total) but does not update docs/internal/metrics.md. pkg/AGENTS.md requires this doc to stay current with metric changes.
on the doc update:
This was resolved in #68 : metrics.md documents parity/category design, not individual custom metrics.
I clarified AGENTS.md accordingly.
pkg/wgo/hedger.go:144-148 (unchanged by this PR, already on main) returns early on a routing mismatch, before any of the attempt/final-outcome metrics closures run. If this guard ever triggers, the failure is invisible in warpstream_produce_final_outcome_total, breaking the one-increment-per-failed-produce invariant this PR adds. Low likelihood since it requires an internal routing bug, but worth a follow-up since it's adjacent to the metrics work here.
Conflicts in client.go and client_test.go: take main's per-partition
rejection handling from #100 and keep one no_agent_assigned final
outcome per rejected group until the final-outcome reasons are reworked.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Pre-existing gap: routing-mismatch guard skips final-outcome metrics
pkg/wgo/hedger.go:144-148 (unchanged by this PR, already on main) returns early on a routing mismatch, before any of the attempt/final-outcome metrics closures run. If this guard ever triggers, the failure is invisible in warpstream_produce_final_outcome_total, breaking the one-increment-per-failed-produce invariant this PR adds. Low likelihood since it requires an internal routing bug, but worth a follow-up since it's adjacent to the metrics work here.
Fixed: the guard now counts one internal_error and returns before any dispatch. It is kept out of the attempts histogram because no attempt ran. Test asserts the outcome, no dispatch, and no attempt observation.
koloss2001
changed the title
Add hedge-trigger, produce-failure, and agent-pool churn metrics
Add hedge-trigger, final-outcome, attempted-payload and excluded-leader metrics
Oct 8, 2026
Preserve primary terminal error in merged hedged stop reason
pkg/wgo/hedger.go:266
hedged.stop only describes the fallback accumulator. If the primary returns a non-retriable/unknown error and the fallback merely exhausts its candidates, this records candidates_exhausted even though the documented precedence says terminal_error wins. The no-proactive and racing branches have the same gap; carry the primary failure classification into the merged stop before observing the final outcome.
🧠 Review effort: Balanced
Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.
The reason will be displayed to describe this comment to others. Learn more.
diffAgentMembership assumes sorted, unique input. refresh sorts newAgents but doesn't dedupe it. If a Metadata response ever lists a NodeID twice, [5,5] followed by [5] counts one removed with no real membership change, which inflates agents_changed_total. diffRemovedAgents is unaffected because it goes through agentSet. Probably rare in practice, but a slices.Compact after the sort would make the documented precondition true. (Or derive the counts from the same set-based diff refresh already does.)
The reason will be displayed to describe this comment to others. Learn more.
Fixed. refresh now deduplicates NodeIDs after sorting. It was worse than described: the pool itself held the duplicate and membership_changed was also counted. Added a test that duplicates a broker in the Metadata response.
The reason will be displayed to describe this comment to others. Learn more.
I'm reverting this. A correct Metadata response contains one entry per NodeID, and a duplicate is treated as malformed, so the sorted list already satisfies diffAgentMembership. Compacting also changes fallback routing for a malformed response (hash % len(agents) and the secondary walk), which this PR otherwise doesn't touch. I've removed the slices.Compact call and its test.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Since v0.1.1 wgo has gained several bug fixes and behaviour changes, including:
ProduceSynccall (Reject only the unroutable partitions in a ProduceSync call #100), and out-of-range partitions (Reject an out-of-range partition instead of guessing a fallback agent #107)The existing metrics cannot show how these paths behave. This PR adds metrics to quantify the changes, monitor the new behaviour and troubleshoot it, fine-tune wgo, and measure how well it mitigates WarpStream issues.
It does not change routing, hedging, timing or retry behaviour.
warpstream_produce_hedge_triggers_totaltrigger:latency,primary_failure,demoted_probewarpstream_produce_hedge_trigger_wins_totaltriggerwarpstream_produce_requests_failed_totalreason:candidates_exhausted,terminal_error,write_timeout,internal_errorProducecall or wire request. Successes and cancellations are not counted; successes are inwarpstream_produce_requests_attempts.warpstream_agentpool_agents_changed_totaldirection:added,removedwarpstream_agentpool_excluded_leaders(gauge)warpstream_produce_attempt_records_total,warpstream_produce_attempt_bytes_totalattempt:primary,hedgeThings to know
warpstream_hedge_wins_totalnow counts only wins that were actually returned. Before, a fallback that finished but lost to the primary could still count. The existing hedge win-ratio panel will need to be updated.excluded_leadersreplaceswarpstream_agentpool_leader_dropped_total. The old counter is not in v0.1.1 or any deployed version. Its rate depended on how often we refresh, so it fell as the backoff slowed refreshes, which looked like recovery.