Preserve input validation under Python optimization - #202
Open
Mr-Neutr0n wants to merge 1 commit into
Open
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
The public Python wrappers use
assertfor required input validation. Python removes those statements underpython -O, so invalid inputs can reach the CUDA extension.A reproduced example is sparse decode with
causal=Trueandis_fp8_kvcache=False: current main calls the backend in optimized mode instead of rejecting the unsupported combination. Reuse-consistency checks, the legacynum_splitsguard, and unsupported prefill options are removed in the same way.Invalid first calls can also set
have_initializedbefore the later mode check fails, leaving scheduler state behind for a request that never ran.Fix
assertstatements with explicit argument exceptions that survive optimizationTypeErrorfor the scheduler object type,ValueErrorfor unsupported arguments and reuse mismatches, andRuntimeErrorfor an impossible internal scheduler stateRegression coverage
The new CPU-only tests isolate the wrapper from the CUDA extension and verify under both regular and optimized Python that:
Nonelegacynum_splitsplaceholder is rejectedValidation
uv run --with torch --with numpy python -m unittest tests/test_flash_mla_input_validation.pyuv run --with torch --with numpy python -O -m unittest tests/test_flash_mla_input_validation.pyuvx ruff check tests/test_flash_mla_input_validation.pypython3 -m compileall -q flash_mla tests/test_flash_mla_input_validation.pyassertstatements remain inflash_mla_interface.pygit diff --checkThis is independent of #201: that PR makes scheduler initialization transactional when the backend itself raises; this PR ensures invalid arguments cannot bypass validation or commit state before a backend call.
Prepared with OpenAI Codex assistance; I reproduced the optimized-mode failure and reviewed the change.