kfake: make faults consistent across requests; let When call the cluster - #1479
Merged
Merged
Conversation
twmb
force-pushed
the
kfake-fault-rules
branch
5 times, most recently
from
September 27, 2026 18:53
d6d729e to
00c4a5c
Compare
What a fault selector matched depended on the request. TopicID was filled at a few sites only, nested keys dropped their group or txnID, and faults fired before the coordinator check on some requests and after it on others, before ACLs on some, and on non-leaders for partitions. Faults now follow one rule: the key names every identifier known at the site, topic name and ID filled in centrally, and a fault is checked after ACLs and routing (NOT_COORDINATOR, NOT_CONTROLLER, NOT_LEADER_OR_FOLLOWER) and before existence checks. The ACL and the fault stay one deny call; a misrouted key skips faults that answer. An Observe fault still counts a misrouted request, and a fault that names Nodes still fires there, to simulate a leader or coordinator that has not learned it was replaced. A TopLevel fault ignored its selectors and fired on versions whose response has no top-level ErrorCode, which the client saw as an empty success. It now matches Group and TxnID against the request, rejects other selectors, and fires only when the code reaches the wire. OffsetFetch v0-1, DescribeLogDirs v0-2, and ElectLeaders v0 put a request-level fault on each entity instead of an unserialized field. When ran on the cluster goroutine, so a When that called TopicInfo deadlocked. While a When runs, admin now serves Cluster methods inline. Handler fixes found along the way: AddPartitionsToTxn v4-5 was advertised but never read the batched Transactions field; it now adds and verifies batches as Kafka does. TxnOffsetCommit v6 dropped unknown-ID topics from an error response. ConsumerGroupHeartbeat created a group before rejecting the heartbeat. Group configs had no ACL or fault check. Timed-out AlterUserSCRAMCredentials and share offset alter/delete faults were counted but never answered.
twmb
force-pushed
the
kfake-fault-rules
branch
from
September 27, 2026 18:54
00c4a5c to
f77970d
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Faults behaved differently depending on the request. This makes them follow one set of rules, documents those rules, and fixes handler bugs an audit of all 62 handlers turned up.
Rules
Group, one in a transaction byTxnID. Unset selectors still match anything, soPartitionsalone still matches that partition in every request. The topic name and ID are filled in centrally, soTopicandTopicIDboth work whether a request carries names or IDs.Observefault counts the request on any broker, and a fault that setsNodesfires on those brokers even when they are the wrong one, to simulate a leader or coordinator that has not learned it was replaced.TopLevelmatchesGroup/TxnIDagainst the request and rejects other selectors. It fires only on versions whose response serializes a top-levelErrorCode. OffsetFetch v0-1, DescribeLogDirs v0-2, and ElectLeaders v0 put a request-level fault on each entity instead.Whencan callClustermethods. While aWhenruns,adminruns its function inline under a mutex. Methods that wait for the cluster to make progress (WaitGroupInfo,FaultHandle.Wait) andClosestill can't be called fromWhen.Handler fixes
Transactionsfield. It now handles batches and VerifyOnly as Kafka does, with CLUSTER_ACTION required for v4+. At every version, an unauthorized, unknown, or faulted partition fails the transaction, and the rest get OPERATION_NOT_ATTEMPTED.Test changes
TestFaultObserveAndWhen: itsWhennow callsTopicInfo.TestFaultTopLevel: a v6 Fetch must not be faulted.TestTxnDescribeTransactions: runs a v4 VerifyOnly check.TestFaultNode: a fault withoutNodesleaves a non-leader to answer NOT_LEADER; one naming the old leader fires there.Behavior changes for existing kfake users
Changelog
TopLevelfaults honor Group/TxnID and fire only when the code reaches the wire,Fault.Whencan call Cluster methods, and AddPartitionsToTxn supports v4-5 batches and VerifyOnly.