Repository navigation
Phase 7a follow-up: re-prove a hand-off report's lineage before reading it (0.261.252) - #1647
Merged
Paul Lizer (paullizer) merged 6 commits intoOct 6, 2026
Conversation
Capture using --write-golden in a fresh detached worktree with no production changes. Copy generated JSON verbatim. Changes reflect upstream merge capabilities, export profiles, composition mapping and planner/workflow guidance; Phase 4/5 route snapshots are unchanged. Existing V2 goldens are untouched. Hand-off off-golden suite: 10 passed under both normal and optimized pytest, exit 0. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> (cherry picked from commit 5a625599c623d17633427ac363966438ba793582)
The hand-off branch of read_workflow_result loaded the report manifest by its exact node selectors and checked its identity, text output and completion, but never walked its lineage. A report whose consumed-input receipts were malformed, or named a missing or corrupt parent, still read as available, although the general path refuses the same lineage. _read_handoff_result now calls the shared node lineage authorizer (authorize_workflow_node_result_read) on the saved workflow, seeded with the loaded manifest and the same memoized loader, before the result is described or excerpted. Sources aren't re-resolved. read_workflow_result passes through the user id it has already proved. The reader tests add a strong valid control whose report consumed a real collect records result, built through the runner's result contract, and cover a malformed receipt (the repro), a missing parent, a parent whose output disagrees with the receipt, a parent whose bytes don't match its hash, and edited or re-enabled definitions with lineage. A descriptor-only read walks the lineage but reads no section. Refs microsoft#1549, microsoft#1543. Follow-up to microsoft#1640. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Add the Workflow Hand-off Result Lineage Fix doc, covering the reader's missing lineage walk, the change and its failure codes, the recaptured off-golden, validation and limitations. In the hand-off feature doc, the reader now re-proves the report's lineage with the shared node lineage authorizer, an edited or re-enabled workflow still fails closed with workflow_result_invalid, and the test, file and performance notes match. Add a 0.261.252 release-notes section with two bug fixes. Existing sections are unchanged, and the docs inventory regenerates unchanged. Part of microsoft#1549. Part of microsoft#1543. Follow-up to microsoft#1640. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Paul Lizer (paullizer)
force-pushed
the
paullizer-7a-follow-up-handoff-lineage
branch
from
October 6, 2026 04:24
27d5b59 to
c07b911
Compare
test_the_phase_4_blueprint_and_payloads_are_unchanged failed at V2 15feec6 on the schema digest alone. Its pin (ab12fb1d) predated microsoft#1641, which added the merge task schema and was merged into Phase 7a at b8e4730 without a refresh. The pin now holds 520139cb, the value V2 computes at a5a5b1c, after microsoft#1641 and before hand-off merged, so the test still proves hand-off left the Phase 4 schema unchanged. The payload, validation and dry-run pins are unchanged. The fix doc adds the builder test to Files modified, records the pin under Validation, and gives the 200-document read time as about 1.7 seconds, to match the c07b911 timings. Refs microsoft#1549, microsoft#1543. Follow-up to microsoft#1640. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Paul Lizer (paullizer)
force-pushed
the
paullizer-7a-follow-up-handoff-lineage
branch
from
October 6, 2026 06:21
c07b911 to
8328678
Compare
Bring in origin/paullizer-react-v2-ui at 6bd12cf, the merge of microsoft#1639 (0.261.251), so microsoft#1647 lands second as agreed. Only two files conflicted: - config.py keeps VERSION = "0.261.252", one above V2. - release_notes.md keeps the v0.261.252 section directly above V2's v0.261.251 section, with every section byte-exact. Every other file matches the side that changed it. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Paul Lizer (paullizer)
marked this pull request as ready for review
October 6, 2026 12:01
Paul Lizer (paullizer)
merged commit Oct 6, 2026
16f42e4
into
microsoft:paullizer-react-v2-ui
11 checks passed
Paul Lizer (paullizer)
added a commit
that referenced
this pull request
Oct 6, 2026
…w-up - Status date 2026-10-06, with paullizer-react-v2-ui at 0.261.252. - Phase 6: 6b-2 done in #1639 (v0.261.251), with how it shipped. - Phase 7: 7a done in #1640 (v0.261.250) and its lineage follow-up in #1647 (v0.261.252); 7b, the V2 hand-off card, is next. Admins should leave the hand-off setting off until 7b ships. - Section 9: the run deep-link and one-time workflow decisions are settled. - Release-notes integrity follow-up updated to 187 missing releases. Docs only; no version bump. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
This was referenced Oct 6, 2026
Paul Lizer (paullizer)
added a commit
that referenced
this pull request
Oct 8, 2026
…lows-capability Update roadmap status: 6b-2 and 7a done, with the #1647 lineage follow-up
This was referenced Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
_read_handoff_resultnow runs the shared node lineage authorizer (authorize_workflow_node_result_read) on the report before it builds a descriptor or loads any text. If the report's consumed-input receipts don't chain to real parent results of this run, the read is refused. That applies to the descriptor-only read, the excerpt read and the stored-context re-check. Before this fix, a report manifest withconsumed_inputs: [{"malformed_receipt": true}], saved under its own valid hash, read asavailable=Trueand returned its text.-x), which recapturesorchestration_workflow_handoff_off_golden.jsonon unmodified V2 a5a5b1c (.248). The merged fixture was captured on f1aeef1 (.233), before Merge CSV, Excel, PDF, Word and PowerPoint files in V2 chat and workflows #1641 changed planner content, and it failed 4/10 at 15feec6. The golden blob isf93f8e4c3dcc263cc5a71186f2cecc8c60574098, byte-identical to the coordinator's independent recapture. Neither file was edited or regenerated here.test_workflow_handoff_builder.pypinned the blueprint schema digest from before Merge CSV, Excel, PDF, Word and PowerPoint files in V2 chat and workflows #1641 added themergetask schema, so it failed at 15feec6. Theschemapin now holds the value V2 computes at a5a5b1c, after Merge CSV, Excel, PDF, Word and PowerPoint files in V2 chat and workflows #1641 and before hand-off, so the test still proves hand-off left the Phase 4 schema unchanged. The other four pins are unchanged.enable_chat_orchestration_workflow_handoffis off, which is the default. With it on, a valid hand-off reads exactly as before, with the same descriptor and text, but each read now also loads the report's lineage (see E2E cost). Only a report whose lineage can't be re-proved is refused.Changes since review
The coordinator reviewed
c07b91151and asked for one more commit, to refresh the stale Phase 4 schema digest pin in this PR. Then V2 moved, so commit 6 merges the new V2 tip. Nothing else changed.024004035("Refresh the Phase 4 schema digest pin after Merge CSV, Excel, PDF, Word and PowerPoint files in V2 chat and workflows #1641"):functional_tests/test_workflow_handoff_builder.py:PHASE4_DIGESTS["schema"]goes fromab12fb1d…to520139cbce120ca017106bf200832712913e395c6cd3e6fa9b530eae01cbd66e, and the comment above it says why. No other pin changed. See decision (n) for the file's Version header.docs/explanation/fixes/WORKFLOW_HANDOFF_RESULT_LINEAGE_FIX.md: a Files modified row for the builder test, a "Phase 4 digest pin" validation note, and "about 1.9 seconds" is now "about 1.7 seconds", matching thec07b91151timings. The release notes and the feature doc are unchanged._digest(WORKFLOW_BLUEPRINT_SCHEMA)is520139cb…, the same as at 15feec6 and this PR's heads. Its$defsarefile_sync_schedule,handle,merge,merge_optionsandrunner.83286787e, and itsgit patch-id --stable(8e25d4bb…) and message matchc07b91151's.git diff --stat c07b91151 83286787eis the two files above, +17/−2, andapplication/is identical. V2 was still 15feec6 then. The push used--force-with-leasepinned toc07b91151.83286787e. The builder passes 32/32 in all four modes, and the harness set passes 285 under pytest and pytest -O. Docs coverage is 7/7, site quality 6/6, and guardrails exit 0 with every check passing. Every other result stands fromc07b91151, because no other file changed.6d45bb8a7("Merge V2 6bd12cf (Phase 6b-2: show chat-started workflow runs and their posted results in V2 #1639) into the hand-off lineage follow-up"). V2 moved to 6bd12cf when Phase 6b-2: show chat-started workflow runs and their posted results in V2 #1639 (6b-2, .251) merged.git -c rerere.enabled=false merge; no recorded resolution was applied. Two files conflicted.config.pykeepsVERSION = "0.261.252". The release notes put this PR's .252 section first, then V2's .251 section, then the rest; removing either section reproduces the other side's file byte for byte.83286787e. The 49 only Phase 6b-2: show chat-started workflow runs and their posted results in V2 #1639 changed match V2:application/v2_ui25, docs 7,ui_tests9, and V2 JavaScript and XSS-guardrail tests 8. Nothing underapplication/single_appchanged.3df18e47363e5763d7e36de30391c25d6c2c10a1, the same as the coordinator's independent trial merge. The push was a fast-forward from83286787e. See Reruns at the merge.Root cause
read_workflow_resultsends a structured hand-off run down_read_handoff_result. That function reads one result: the report that the run'sworkflow_outputsreceipt names. It does three things:load_node_result, which recomputes the node identity from the saved workflow. That's why an edited definition already failed closed before any load.It never called the lineage authorizer. The general (non-structured) path calls
_guarded("authorize", lambda: authorize_workflow_run_read(...)), and the feature doc (L567–568) and the reader test's docstring promised the hand-off did the same. So nothing checked that the report's consumed-input receipts chain to real parent results of this run.Fix
The fix goes after
_require_completed_resultand before the result is built, so it covers the descriptor-only read and the excerpt read:workflow, notworkflow_runtime_store(...).run_definition(). That matchesload_node_resultin the same function and keeps the edited-workflow fail-closed behavior.user_idis passed in fromread_workflow_result, which has already provedworkflow["user_id"] == user_id._ManifestMemoloader, and passesinclude_sources=Falsebecause the reader discards the access summary. The authorizer never re-resolves sources.authorize_workflow_result_contextre-reads through the same function, so stored chat contexts get the walk too.functions_workflow_node_results.pyandfunctions_workflow_results.pyaren't touched, and the authorizer itself is unchanged.Failure codes
workflow_result_invalidauthorizeAnalysisResultUnavailable(analysis_lineage_invalid)workflow_result_not_found(404)authorizeCosmosResourceNotFoundErroroutput_ref(receipt side and parent side)workflow_result_invalidauthorizeAnalysisResultUnavailableworkflow_result_invalidauthorizeWorkflowResultIntegrityErrorworkflow_result_invalidThe general path gives the same code and status for a missing parent (
workflow_result_not_found, 404), and a test asserts that.Commits
00ba92039: cherry-pick of 5a625599c, the off-golden recapture. Message kept, plus-x.3a2bd7a0c: the fix and its tests.ad06270e5: docs, including the fix doc, the feature doc and the release notes.024004035: the Phase 4 schema digest pin refresh and its fix-doc note, added after review.83286787e:VERSION = "0.261.252", the last of this PR's own commits.6d45bb8a7: the merge of V2 6bd12cf (Phase 6b-2: show chat-started workflow runs and their posted results in V2 #1639), with rerere off.The branch started from V2 15feec6. V2 hadn't moved when this opened, nor when the CodeQL fix or the pin refresh was pushed; then #1639 merged, and commit 6 merges it. The branch is on paullizer/simplechat because a push to microsoft/simplechat returned 403; the kickoff gives that fallback.
This PR first opened at
27d5b5939. CodeQL flagged one warning in this PR's own test fixture, so I folded a fixture fix into commit 2 and force-pushed with a lease (see decision (m) and CodeQL).git range-diffshowed commit 2 changed, and commit 3 and the VERSION commit patch-identical (=). Commit 1 wasn't rewritten.git diff --stat 27d5b5939 c07b91151is one file:functional_tests/test_workflow_handoff_result_reader.py, +40/−26.Linked issue
Part of #1549
Part of #1543
Follow-up to #1640
Release Notes & Latest Features
Is this visible to end users?
Is this admin-facing (Admin Settings, governance, deployment, config)?
Should this become a Latest Feature card?
Screenshot needed for the card?
Version bump
application/single_app/config.pyVERSIONthird segment bumped, or not needed because this is docs-only. It goes from0.261.250to0.261.252in commit 5. .251 is 6b-2 (Phase 6b-2: show chat-started workflow runs and their posted results in V2 #1639), which merged first; the V2 merge keeps .252, so no renumber was needed.deployers/version.txtbumped, or not needed becausedeployers/was not changed.deployers/isn't changed.Testing / validation
Environment. Every job ran from the worktree root, one process at a time, with a 1,800 s timeout. The full set ran at the review head
c07b91151. Cells marked (at83286787e) reran at the new head after the pin refresh. Between the two heads onlytest_workflow_handoff_builder.pyand the fix doc differ, so every other row stands.<py>isC:\Users\paullizer\.copilot\session-state\aa1bf1e0-b10d-4312-8dbb-143e9e9f9966\files\venv\Scripts\python.exe(Flask 3.1.3). The environment isPYTHONIOENCODING=utf-8,PYTHONUTF8=1andPYTHONPATH=application/single_app;functional_tests. An exit of0xC000026Bcounts as interrupted, not as a pass.Compared with the first head. The same 106 jobs first ran at
27d5b5939, before the CodeQL fixture fix. Every row's exit code and pass/fail counts are identical in both runs.These are the command forms;
<file>is the file in each row, and every exact command is listed at the end:<py> -u -m pytest <file> -p no:langsmith_plugin -p no:cacheprovider -q -rfE<py> -O -u -m pytest <file> -p no:langsmith_plugin -p no:cacheprovider -q -rfE<py> -u <file><py> -O -u <file>In the full command list at the end,
<flags>stands for exactly-p no:langsmith_plugin -p no:cacheprovider -q -rfE.Totals. At
c07b91151: 106 jobs; 102 exited 0; 4 didn't, all on the stale builder pin, which was pre-existing at 15feec6 (see Baseline classification). At83286787e: 8 reruns of what changed since review (the builder in four modes, the harness set in both pytest modes, and the two docs checks); all 8 exited 0. At the V2 merge6d45bb8a7: 10 reruns, all exited 0; see Reruns at the merge.Reruns at the merge
6d45bb8a7merges V2 6bd12cf (#1639); see Changes since review. It changes nothing underapplication/single_app, and none of the reader, off-golden, builder or E2E test files or the Python test support. So I reran this PR's own test files, the docs checks and the guardrails, one at a time from the worktree root with the environment above:<py> -u -m pytest functional_tests/test_workflow_handoff_result_reader.py <flags>: exit 0, 81 passed<py> -O -u -m pytest functional_tests/test_workflow_handoff_result_reader.py <flags>: exit 0, 81 passed<py> -u -m pytest functional_tests/test_orchestration_workflow_handoff_off_golden.py <flags>: exit 0, 10 passed<py> -O -u -m pytest functional_tests/test_orchestration_workflow_handoff_off_golden.py <flags>: exit 0, 10 passed<py> -u -m pytest functional_tests/test_workflow_handoff_builder.py <flags>: exit 0, 32 passed<py> -O -u -m pytest functional_tests/test_workflow_handoff_builder.py <flags>: exit 0, 32 passed<py> -u .\scripts\build_docs_inventory.py: exit 0; the inventory is unchanged apart from line endings, restored withgit checkout<py> -u functional_tests/test_docs_app_surface_coverage.py: exit 0, 7/7 checks passed<py> -u functional_tests/test_docs_site_quality.py: exit 0, 6/6 checks passed6bd12cfcdfefe83df58c70bff64e67d7e130363d,HEADandr1: exit 0. It reviewed the same 9 files as at83286787e, and every check passed. The malicious-PR review has 268 findings and 0 blockers, with the same per-file counts.Two docs checks outside this PR's list fail at the merge. Each fails the same way at V2 6bd12cf, rerun on a temporary detached worktree that was removed afterwards, so both are pre-existing:
6d45bb8a7<py> -u functional_tests/test_docs_release_notes_integrity.py<py> -u functional_tests/test_docs_link_integrity.pyfeatures.ymllinks. None is in this PR's docs. The relative-link total goes from 1,349 to 1,354 with this PR's 5 new links, which all resolve.Repro, before and after
The repro is a session script, not a committed test. It builds
HandoffFixture()from the worktree under test, with the network blocked, and then:consumed_inputs = [{"malformed_receipt": True}];fixture.store.save(...), so it gets a fresh, valid content hash;receipt["result_ref"]at the new reference;authorize_workflow_node_result_readon the same reference.Each run used the environment above, from the root of the worktree under test:
<py> -u repro_r1.py <worktree>and<py> -O -u repro_r1.py <worktree>.include_excerpts=Trueavailable=True, report text returned, 2 loadsavailable=True, 1 loadanalysis_lineage_invalid, mapped toworkflow_result_invalidavailable=True, report text returned, 2 loadsavailable=True, 1 loadanalysis_lineage_invalid, mapped toworkflow_result_invalidworkflow_result_invalid, 1 loadworkflow_result_invalid, 1 loadanalysis_lineage_invalid, mapped toworkflow_result_invalidworkflow_result_invalid, 1 loadworkflow_result_invalid, 1 loadanalysis_lineage_invalid, mapped toworkflow_result_invalidThe "after" rows ran at the review head, c07b911;
application/single_appand the reader test file are identical at83286787eand6d45bb8a7. They first ran at 27d5b59 with identical output, and were rerun because the repro importsHandoffFixture, which the CodeQL fix restructured. After the fix, the single load is the report manifest; the text section is never read. The reader logs{'code': 'workflow_result_invalid', 'stage': 'authorize', 'error_type': 'AnalysisResultUnavailable'}, with no identifiers.Raw output
The
[LOG] [WorkflowResults] Workflow result unavailablelines from the two "after" runs are omitted above; the paragraph before this block quotes their payload.Four modes
test_workflow_handoff_result_reader.pytest_orchestration_workflow_handoff_off_golden.pytest_workflow_handoff_builder.py83286787e)83286787e)83286787e)83286787e)Goldens
At 15feec6 the coordinator saw 9, 5, 5 and 6 passes for the first four in both pytest modes. The counts here match.
test_orchestration_workflow_setting_off_golden.pytest_orchestration_workflow_runs_off_golden.pytest_orchestration_workflow_results_off_golden.pytest_workflow_chat_delivery_off_golden.pytest_orchestration_reference_authorizer_golden.pytest_workflow_run_time_context.pyHand-off set
test_orchestration_workflow_handoff_adapter.pytest_orchestration_workflow_handoff_capability.pytest_orchestration_workflow_handoff_imports.pytest_orchestration_workflow_handoff_planner.pytest_orchestration_workflow_handoff_routes.pytest_workflow_handoff_builder.py83286787e)83286787e)test_workflow_handoff_end_to_end.pytest_workflow_handoff_lifecycle.pytest_workflow_handoff_origin.pyroute_tests/test_route_blueprint_policy_inventory.pyroute_tests/test_route_unauthenticated_policy_contract.pyroute_tests/test_route_policy_test_coverage.pyResult readers and delivery
test_workflow_result_reader.pytest_workflow_result_orchestration_lineage.pytest_workflow_result_review_paths.pytest_workflow_result_routes.pytest_workflow_result_chat_routes.pytest_workflow_result_followup.pytest_workflow_result_privacy_import_cycle.pytest_workflow_result_contract.pytest_workflow_result_store.pytest_workflow_result_masking.pyroute_tests/test_workflow_result_context_policy.pytest_workflow_chat_delivery_concurrency.pytest_workflow_chat_delivery_contract.pytest_workflow_chat_delivery_control_pins.pytest_workflow_chat_delivery_imports.pytest_workflow_chat_delivery_loop.pytest_workflow_chat_delivery_notice_and_unread.pytest_workflow_chat_delivery_placement_and_masking.pytest_workflow_chat_delivery_projection_and_seed.pytest_workflow_chat_delivery_refusals.pytest_workflow_chat_delivery_save_guard.pytest_workflow_chat_delivery_status_route.pytest_workflow_chat_delivery_worker.pytest_orchestration_workflow_results_aliases.pytest_orchestration_workflow_results_answer.pytest_orchestration_workflow_results_capability.pytest_orchestration_workflow_results_imports.pytest_orchestration_workflow_results_reads.pytest_orchestration_workflow_results_time_zone.pytest_chat_workflow_results_admin.pyHarness set
test_m365_run_as_self_authored.py,test_workflow_assist_dry_run_parity.py,test_workflow_draft_save_parity.py,test_workflow_draft_service.py,test_workflow_draft_v2_round_trip.py,test_workflow_handoff_builder.py,test_workflow_handoff_origin.py,test_workflow_origin_provenance.py83286787e)83286787e)Baseline classification
At the review head
c07b91151, four jobs didn't exit 0, all on the same test. I reran each one on a temporary detached worktree at 15feec6, with the same commands, environment and runner. The worktree was created withgit worktree add --detach <session files>\base_15feec650 15feec6509f0846238362838103be6e76d19f8c1(nogit stash) and removed afterwards. The first run at27d5b5939gave the same counts asc07b91151in all four rows. Commit 4 removes the cause; the last result column is from the new head.c07b9115183286787etest_workflow_handoff_builder.pytest_workflow_handoff_builder.pyAll 8 failing runs failed the same test,
test_workflow_handoff_builder.py::test_the_phase_4_blueprint_and_payloads_are_unchanged.schema, which is_digest(drafts.WORKFLOW_BLUEPRINT_SCHEMA). It was pinned toab12fb1d29c7eaa028a534efa4f610795a31e25638643f30a8ec399b493d9b26and computes to520139cbce120ca017106bf200832712913e395c6cd3e6fa9b530eae01cbd66eat a5a5b1c, at 15feec6 and at every head of this PR. Thepayload_email,payload_review,validateanddry_reviewdigests all match their pins.mergetask property to the blueprint schema and didn't refresh the pin.No job was interrupted or timed out, so no job needed a rerun on its own.
Baseline commands, run from the 15feec6 worktree
<py> -u -m pytest functional_tests/test_workflow_handoff_builder.py -p no:langsmith_plugin -p no:cacheprovider -q -rfE: exit 1, 1 failed, 31 passed<py> -O -u -m pytest functional_tests/test_workflow_handoff_builder.py -p no:langsmith_plugin -p no:cacheprovider -q -rfE: exit 1, 1 failed, 31 passed<py> -u -m pytest functional_tests/test_m365_run_as_self_authored.py functional_tests/test_workflow_assist_dry_run_parity.py functional_tests/test_workflow_draft_save_parity.py functional_tests/test_workflow_draft_service.py functional_tests/test_workflow_draft_v2_round_trip.py functional_tests/test_workflow_handoff_builder.py functional_tests/test_workflow_handoff_origin.py functional_tests/test_workflow_origin_provenance.py -p no:langsmith_plugin -p no:cacheprovider -q -rfE: exit 1, 1 failed, 284 passed-O: exit 1, 1 failed, 284 passedMutations and E2E probe
See the Kill table and E2E cost sections below.
Docs
The inventory steps ran at the review head
c07b91151, from the worktree root, with the environment above; they first ran at27d5b5939with the same results. The two docs checks reran at83286787e, after the fix-doc update, with the same results. The inventory step and both checks reran at the merge6d45bb8a7, with the same results again.<py> -u .\scripts\build_docs_inventory.pygit diff --quietafter regeneratinggit statusflagsdocs/_data/app_surface.ymlonly because the generator writes LF and the working copy is CRLF;git diff --ignore-cr-at-eol --statis empty. I restored the file withgit checkout.<py> -u functional_tests/test_docs_app_surface_coverage.py(__main__)<py> -u functional_tests/test_docs_site_quality.py(__main__)The inventory is unchanged because this PR adds no setting, admin tab, action, chat control or page.
Guardrails
Command:
<py> C:\Users\paullizer\.copilot\session-state\151a0b09-0034-492f-bac4-4a1ca884309b\files\run_v2_guardrails.py <worktree> 15feec6509f0846238362838103be6e76d19f8c1 HEAD r1. I ran it at the review headc07b91151and again at the new head83286787e, and it exited 0 both times. It first ran at27d5b5939, which gave the same results asc07b91151, including the finding totals by severity and by file. At the merge6d45bb8a7it ran against base 6bd12cf; see Reruns at the merge.At
83286787ethe script reviewed 9 changed files: 2 app.py, 2 on the XSS surface, 2 route.py, and 0 V2ts/tsx. Atc07b91151it reviewed 8, because the builder test wasn't changed yet.single_app.pyfiles compile83286787e, and 267 atc07b91151. The report's final verdict is "Needs investigation"ts/tsxfiles changedBreakdown of the malicious-PR-review findings at
83286787e:c07b91151it was 264 Important and 3 Moderate.functional_tests/test_support/orchestration_workflow_handoff_off_golden.json, the cherry-picked recapture. Its blob isf93f8e4c3dcc263cc5a71186f2cecc8c60574098, byte-identical to the coordinator's independent recapture.functional_tests/test_workflow_handoff_result_reader.py.c07b91151it was 10. The new one is the "external connection" marker on the#1641link at L178, in the Phase 4 digest pin note.functions_workflow_result_reader.py,config.pyortest_workflow_handoff_builder.py.HandoffFixture();monkeypatch.setattr(reader, "authorize_workflow_node_result_read", ...)calls. One shows the malformed manifest would read as a normal report without the walk, so the refusal comes from the walk. The other shows the strong control's result is identical without the walk and only the parent's load disappears, and it pins the walk's arguments;instructionsin the edited-definition test.Every validation command
All 106 jobs at `c07b91151` and 8 at `83286787e`, in run order, with exit code and pass/fail counts (timings and warning counts omitted)
At the review head
c07b91151:<py> -u -m pytest functional_tests/test_workflow_handoff_result_reader.py <flags>: exit 0, 81 passed<py> -O -u -m pytest functional_tests/test_workflow_handoff_result_reader.py <flags>: exit 0, 81 passed<py> -u functional_tests/test_workflow_handoff_result_reader.py: exit 0, 81 passed<py> -O -u functional_tests/test_workflow_handoff_result_reader.py: exit 0, 81 passed<py> -u -m pytest functional_tests/test_orchestration_workflow_handoff_off_golden.py <flags>: exit 0, 10 passed<py> -O -u -m pytest functional_tests/test_orchestration_workflow_handoff_off_golden.py <flags>: exit 0, 10 passed<py> -u functional_tests/test_orchestration_workflow_handoff_off_golden.py: exit 0, 10 passed<py> -O -u functional_tests/test_orchestration_workflow_handoff_off_golden.py: exit 0, 10 passed<py> -u -m pytest functional_tests/test_orchestration_workflow_setting_off_golden.py <flags>: exit 0, 9 passed<py> -O -u -m pytest functional_tests/test_orchestration_workflow_setting_off_golden.py <flags>: exit 0, 9 passed<py> -u -m pytest functional_tests/test_orchestration_workflow_runs_off_golden.py <flags>: exit 0, 5 passed<py> -O -u -m pytest functional_tests/test_orchestration_workflow_runs_off_golden.py <flags>: exit 0, 5 passed<py> -u -m pytest functional_tests/test_orchestration_workflow_results_off_golden.py <flags>: exit 0, 5 passed<py> -O -u -m pytest functional_tests/test_orchestration_workflow_results_off_golden.py <flags>: exit 0, 5 passed<py> -u -m pytest functional_tests/test_workflow_chat_delivery_off_golden.py <flags>: exit 0, 6 passed<py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_off_golden.py <flags>: exit 0, 6 passed<py> -u -m pytest functional_tests/test_orchestration_reference_authorizer_golden.py <flags>: exit 0, 2 passed, 74 subtests passed<py> -O -u -m pytest functional_tests/test_orchestration_reference_authorizer_golden.py <flags>: exit 0, 2 passed, 74 subtests passed<py> -u -m pytest functional_tests/test_workflow_run_time_context.py <flags>: exit 0, 13 passed<py> -O -u -m pytest functional_tests/test_workflow_run_time_context.py <flags>: exit 0, 13 passed<py> -u -m pytest functional_tests/test_orchestration_workflow_handoff_adapter.py <flags>: exit 0, 23 passed<py> -O -u -m pytest functional_tests/test_orchestration_workflow_handoff_adapter.py <flags>: exit 0, 23 passed<py> -u -m pytest functional_tests/test_orchestration_workflow_handoff_capability.py <flags>: exit 0, 90 passed<py> -O -u -m pytest functional_tests/test_orchestration_workflow_handoff_capability.py <flags>: exit 0, 90 passed<py> -u -m pytest functional_tests/test_orchestration_workflow_handoff_imports.py <flags>: exit 0, 31 passed<py> -O -u -m pytest functional_tests/test_orchestration_workflow_handoff_imports.py <flags>: exit 0, 31 passed<py> -u -m pytest functional_tests/test_orchestration_workflow_handoff_planner.py <flags>: exit 0, 52 passed<py> -O -u -m pytest functional_tests/test_orchestration_workflow_handoff_planner.py <flags>: exit 0, 52 passed<py> -u -m pytest functional_tests/test_orchestration_workflow_handoff_routes.py <flags>: exit 0, 61 passed<py> -O -u -m pytest functional_tests/test_orchestration_workflow_handoff_routes.py <flags>: exit 0, 61 passed<py> -u -m pytest functional_tests/test_workflow_handoff_builder.py <flags>: exit 1, 1 failed, 31 passedFAILED functional_tests/test_workflow_handoff_builder.py::test_the_phase_4_blueprint_and_payloads_are_unchanged<py> -O -u -m pytest functional_tests/test_workflow_handoff_builder.py <flags>: exit 1, 1 failed, 31 passedFAILED functional_tests/test_workflow_handoff_builder.py::test_the_phase_4_blueprint_and_payloads_are_unchanged<py> -u -m pytest functional_tests/test_workflow_handoff_end_to_end.py <flags>: exit 0, 3 passed<py> -O -u -m pytest functional_tests/test_workflow_handoff_end_to_end.py <flags>: exit 0, 3 passed<py> -u -m pytest functional_tests/test_workflow_handoff_lifecycle.py <flags>: exit 0, 26 passed<py> -O -u -m pytest functional_tests/test_workflow_handoff_lifecycle.py <flags>: exit 0, 26 passed<py> -u -m pytest functional_tests/test_workflow_handoff_origin.py <flags>: exit 0, 34 passed<py> -O -u -m pytest functional_tests/test_workflow_handoff_origin.py <flags>: exit 0, 34 passed<py> -u -m pytest functional_tests/route_tests/test_route_blueprint_policy_inventory.py <flags>: exit 0, 12 passed<py> -O -u -m pytest functional_tests/route_tests/test_route_blueprint_policy_inventory.py <flags>: exit 0, 12 passed<py> -u -m pytest functional_tests/route_tests/test_route_unauthenticated_policy_contract.py <flags>: exit 0, 7 passed<py> -O -u -m pytest functional_tests/route_tests/test_route_unauthenticated_policy_contract.py <flags>: exit 0, 7 passed<py> -u -m pytest functional_tests/route_tests/test_route_policy_test_coverage.py <flags>: exit 0, 2 passed<py> -O -u -m pytest functional_tests/route_tests/test_route_policy_test_coverage.py <flags>: exit 0, 2 passed<py> -u -m pytest functional_tests/test_workflow_result_reader.py <flags>: exit 0, 99 passed<py> -O -u -m pytest functional_tests/test_workflow_result_reader.py <flags>: exit 0, 99 passed<py> -u -m pytest functional_tests/test_workflow_result_orchestration_lineage.py <flags>: exit 0, 11 passed<py> -O -u -m pytest functional_tests/test_workflow_result_orchestration_lineage.py <flags>: exit 0, 11 passed<py> -u -m pytest functional_tests/test_workflow_result_review_paths.py <flags>: exit 0, 2 passed<py> -O -u -m pytest functional_tests/test_workflow_result_review_paths.py <flags>: exit 0, 2 passed<py> -u -m pytest functional_tests/test_workflow_result_routes.py <flags>: exit 0, 26 passed<py> -O -u -m pytest functional_tests/test_workflow_result_routes.py <flags>: exit 0, 26 passed<py> -u -m pytest functional_tests/test_workflow_result_chat_routes.py <flags>: exit 0, 12 passed<py> -O -u -m pytest functional_tests/test_workflow_result_chat_routes.py <flags>: exit 0, 12 passed<py> -u -m pytest functional_tests/test_workflow_result_followup.py <flags>: exit 0, 91 passed<py> -O -u -m pytest functional_tests/test_workflow_result_followup.py <flags>: exit 0, 91 passed<py> -u -m pytest functional_tests/test_workflow_result_privacy_import_cycle.py <flags>: exit 0, 11 passed<py> -O -u -m pytest functional_tests/test_workflow_result_privacy_import_cycle.py <flags>: exit 0, 11 passed<py> -u -m pytest functional_tests/test_workflow_result_contract.py <flags>: exit 0, 10 passed<py> -O -u -m pytest functional_tests/test_workflow_result_contract.py <flags>: exit 0, 10 passed<py> -u -m pytest functional_tests/test_workflow_result_store.py <flags>: exit 0, 38 passed, 93 subtests passed<py> -O -u -m pytest functional_tests/test_workflow_result_store.py <flags>: exit 0, 38 passed, 93 subtests passed<py> -u -m pytest functional_tests/test_workflow_result_masking.py <flags>: exit 0, 49 passed<py> -O -u -m pytest functional_tests/test_workflow_result_masking.py <flags>: exit 0, 49 passed<py> -u -m pytest functional_tests/route_tests/test_workflow_result_context_policy.py <flags>: exit 0, 38 passed<py> -O -u -m pytest functional_tests/route_tests/test_workflow_result_context_policy.py <flags>: exit 0, 38 passed<py> -u -m pytest functional_tests/test_workflow_chat_delivery_concurrency.py <flags>: exit 0, 5 passed<py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_concurrency.py <flags>: exit 0, 5 passed<py> -u -m pytest functional_tests/test_workflow_chat_delivery_contract.py <flags>: exit 0, 72 passed<py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_contract.py <flags>: exit 0, 72 passed<py> -u -m pytest functional_tests/test_workflow_chat_delivery_control_pins.py <flags>: exit 0, 35 passed<py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_control_pins.py <flags>: exit 0, 35 passed<py> -u -m pytest functional_tests/test_workflow_chat_delivery_imports.py <flags>: exit 0, 31 passed<py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_imports.py <flags>: exit 0, 31 passed<py> -u -m pytest functional_tests/test_workflow_chat_delivery_loop.py <flags>: exit 0, 9 passed<py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_loop.py <flags>: exit 0, 9 passed<py> -u -m pytest functional_tests/test_workflow_chat_delivery_notice_and_unread.py <flags>: exit 0, 25 passed<py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_notice_and_unread.py <flags>: exit 0, 25 passed<py> -u -m pytest functional_tests/test_workflow_chat_delivery_placement_and_masking.py <flags>: exit 0, 8 passed<py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_placement_and_masking.py <flags>: exit 0, 8 passed<py> -u -m pytest functional_tests/test_workflow_chat_delivery_projection_and_seed.py <flags>: exit 0, 16 passed<py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_projection_and_seed.py <flags>: exit 0, 16 passed<py> -u -m pytest functional_tests/test_workflow_chat_delivery_refusals.py <flags>: exit 0, 2 passed<py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_refusals.py <flags>: exit 0, 2 passed<py> -u -m pytest functional_tests/test_workflow_chat_delivery_save_guard.py <flags>: exit 0, 14 passed<py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_save_guard.py <flags>: exit 0, 14 passed<py> -u -m pytest functional_tests/test_workflow_chat_delivery_status_route.py <flags>: exit 0, 48 passed<py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_status_route.py <flags>: exit 0, 48 passed<py> -u -m pytest functional_tests/test_workflow_chat_delivery_worker.py <flags>: exit 0, 112 passed<py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_worker.py <flags>: exit 0, 112 passed<py> -u -m pytest functional_tests/test_orchestration_workflow_results_aliases.py <flags>: exit 0, 10 passed<py> -O -u -m pytest functional_tests/test_orchestration_workflow_results_aliases.py <flags>: exit 0, 10 passed<py> -u -m pytest functional_tests/test_orchestration_workflow_results_answer.py <flags>: exit 0, 18 passed<py> -O -u -m pytest functional_tests/test_orchestration_workflow_results_answer.py <flags>: exit 0, 18 passed<py> -u -m pytest functional_tests/test_orchestration_workflow_results_capability.py <flags>: exit 0, 70 passed<py> -O -u -m pytest functional_tests/test_orchestration_workflow_results_capability.py <flags>: exit 0, 70 passed<py> -u -m pytest functional_tests/test_orchestration_workflow_results_imports.py <flags>: exit 0, 26 passed<py> -O -u -m pytest functional_tests/test_orchestration_workflow_results_imports.py <flags>: exit 0, 26 passed<py> -u -m pytest functional_tests/test_orchestration_workflow_results_reads.py <flags>: exit 0, 53 passed<py> -O -u -m pytest functional_tests/test_orchestration_workflow_results_reads.py <flags>: exit 0, 53 passed<py> -u -m pytest functional_tests/test_orchestration_workflow_results_time_zone.py <flags>: exit 0, 8 passed<py> -O -u -m pytest functional_tests/test_orchestration_workflow_results_time_zone.py <flags>: exit 0, 8 passed<py> -u -m pytest functional_tests/test_chat_workflow_results_admin.py <flags>: exit 0, 15 passed<py> -O -u -m pytest functional_tests/test_chat_workflow_results_admin.py <flags>: exit 0, 15 passed<py> -u -m pytest functional_tests/test_m365_run_as_self_authored.py functional_tests/test_workflow_assist_dry_run_parity.py functional_tests/test_workflow_draft_save_parity.py functional_tests/test_workflow_draft_service.py functional_tests/test_workflow_draft_v2_round_trip.py functional_tests/test_workflow_handoff_builder.py functional_tests/test_workflow_handoff_origin.py functional_tests/test_workflow_origin_provenance.py <flags>: exit 1, 1 failed, 284 passed-
FAILED functional_tests/test_workflow_handoff_builder.py::test_the_phase_4_blueprint_and_payloads_are_unchanged<py> -O -u -m pytest functional_tests/test_m365_run_as_self_authored.py functional_tests/test_workflow_assist_dry_run_parity.py functional_tests/test_workflow_draft_save_parity.py functional_tests/test_workflow_draft_service.py functional_tests/test_workflow_draft_v2_round_trip.py functional_tests/test_workflow_handoff_builder.py functional_tests/test_workflow_handoff_origin.py functional_tests/test_workflow_origin_provenance.py <flags>: exit 1, 1 failed, 284 passed-
FAILED functional_tests/test_workflow_handoff_builder.py::test_the_phase_4_blueprint_and_payloads_are_unchangedAt the new head
83286787e:<py> -u -m pytest functional_tests/test_workflow_handoff_builder.py <flags>: exit 0, 32 passed<py> -O -u -m pytest functional_tests/test_workflow_handoff_builder.py <flags>: exit 0, 32 passed<py> -u functional_tests/test_workflow_handoff_builder.py: exit 0, 32 passed<py> -O -u functional_tests/test_workflow_handoff_builder.py: exit 0, 32 passed<py> -u -m pytest functional_tests/test_m365_run_as_self_authored.py functional_tests/test_workflow_assist_dry_run_parity.py functional_tests/test_workflow_draft_save_parity.py functional_tests/test_workflow_draft_service.py functional_tests/test_workflow_draft_v2_round_trip.py functional_tests/test_workflow_handoff_builder.py functional_tests/test_workflow_handoff_origin.py functional_tests/test_workflow_origin_provenance.py <flags>: exit 0, 285 passed<py> -O -u -m pytest functional_tests/test_m365_run_as_self_authored.py functional_tests/test_workflow_assist_dry_run_parity.py functional_tests/test_workflow_draft_save_parity.py functional_tests/test_workflow_draft_service.py functional_tests/test_workflow_draft_v2_round_trip.py functional_tests/test_workflow_handoff_builder.py functional_tests/test_workflow_handoff_origin.py functional_tests/test_workflow_origin_provenance.py <flags>: exit 0, 285 passed<py> -u functional_tests/test_docs_app_surface_coverage.py: exit 0, 7/7 checks passed<py> -u functional_tests/test_docs_site_quality.py: exit 0, 6/6 checks passedDecisions made autonomously
None of these deviate from the kickoff or the review. Each one is a choice they left open.
(a) How the strong control's parent is built. It uses the runner's contract, not hand-written receipts.
collectis an engine node with no task, soLineageFixturebuilds its result the way the runner builds a control node's result:workflow_node_identityandworkflow_execution_id, then the runner's private_build_task_result(..., "workflow-result-v2");persist_workflow_task_result;open_workflow_record_input(..., output_name="records")receipt, plus theinput_namethe runner adds;workflow_execution_scope,build_workflow_task_resultandpersist_workflow_task_result.There's no frozen loop in this fixture. The E2E covers the real
collect,eachand frozen-loop lineage.(b) How the corrupt-bytes case is checked. The test fixture's in-memory store doesn't verify digests on load, but the real store does. So this case wraps the loader with the real
_verify_payloadand proves the walk turnsWorkflowResultIntegrityErrorintoworkflow_result_invalid.(c) A test was red between commits.
test_version_is_at_least_the_lineage_release(assert_app_version_at_least("0.261.252")) failed from Commit 2 until the VERSION commit, because the kickoff puts that bump last. It passes at HEAD.(d) How M4 is killed.
_guarded("manifest")and_guarded("authorize")map the same exceptions to the same codes. Only the logged closed stage tells them apart, so the new refusal tests assert that the logged stage isauthorize.(e) An opt-in lineage parameter on the old fixture.
HandoffFixture(lineage=None): whenlineageis unset, every existing test keeps its exact fixture, withconsumed_inputs: None.LineageFixturepassesCollectedParent(...).receipts. That callable saves the realcollectparent and returns its receipt, which becomes the report'sconsumed_inputsjust before the report is persisted. An earlier revision used an overridablelineage()method called fromHandoffFixture.__init__; CodeQL flagged that, see (m).(f) Private imports in the tests. The tests import the private
_build_task_result(runner) and_verify_payload(store) so they use the real contract and real verification.(g) The general-path missing-ancestor comparison. It removes both the parent's task row and its stored result. That way only the lineage walk, not the row loop, can reach the missing parent.
(h) The new edited-definition test. It mirrors all 4 variants of the existing test: definition edit, enable toggle, alert edit and run-as change.
(i) What the strong control pins. It pins the walk's run, identity, reference,
reader_user_idand manifest arguments, but notinclude_sources, because M3 is equivalent.(j) The E2E probe wasn't committed. It was a temporary test module,
functional_tests/test_zz_tmp_probe_handoff_lineage_e2e.py. It re-ran the E2E test with an autouse fixture that counted node-result loads and page reads and wrapped the lineage walk. Awalkmode kept the real walk, and anowalkmode replaced it with a no-op as an M1 stand-in. I deleted it, with its bytecode, after each use.(k) What "script mode" means. It's
python -u <file>. Both four-mode files'__main__runspytest.main([__file__, "-q"]), which is the files' own runner and is unchanged. That means script mode runs without-p no:langsmith_plugin.(l) Where the walk's cost is documented. It's in the feature doc's Performance section, since the Known limitations section doesn't cover lineage.
(m) Fixing this PR's own CodeQL alert by rewriting commit 2. CodeQL's first analysis of
27d5b5939reported one newpy/init-calls-subclasswarning:HandoffFixture.__init__calledself.lineage(...), whichLineageFixtureoverrides.The fix. I removed the pattern instead of dismissing the alert or suppressing the rule. The hook became the
lineage=parameter from (e), and the parent's construction moved to a smallCollectedParentclass. The two tamper callbacks now receive(receipt, parent, save)instead of reaching into a half-built fixture.Why commit 2. I folded the fix in with
git commit --fixupandgit -c rerere.enabled=false rebase --autosquash, so the VERSION commit stays last as the kickoff requires. Then I pushed with--force-with-leasepinned to27d5b5939. The old head is still in the local reflog.Equivalence proof. A session script loaded the old and new fixture modules side by side and built six variants:
For each it compared stored results, manifests, references, receipts, run rows, loads and selectors, plus descriptor-only and excerpt reads. All six were identical under
pythonandpython -O.Rerun. Every validation job, the four modes, the mutations, the repro, the guardrails, the docs checks and the E2E probe reran at
c07b91151(see Testing / validation, Kill table and E2E cost).(n) The builder test's Version header. Commit 4 changes only the
schemapin and its comment, as the review asked. The file'sVersion: 0.261.250header is unchanged, because the review didn't ask for it and no test in the file changed.(o) Landing. The user asked to submit and merge this PR, so once CI on
6d45bb8a7passed I marked it ready and merged it. I used a merge commit, as Phase 6b-2: show chat-started workflow runs and their posted results in V2 #1639 and Phase 7a: hand off large chat requests to a one-time workflow (server, 0.261.238) #1640 did, pinned with--match-head-commit. I didn't use--adminor--delete-branch. V2's .251 sits right below this PR's .252, so no renumber was needed.Kill table
Each mutation was applied to
functions_workflow_result_reader.pyat the review headc07b91151. The reader file then ran under pytest (<py> -u -m pytest functional_tests/test_workflow_handoff_result_reader.py -p no:langsmith_plugin -p no:cacheprovider -q -rf) with nothing deselected, and the file was restored. Afterwards the restored file was byte-identical andgit diff --quiet -- application/single_appexited 0.The same four mutations first ran at
27d5b5939, before the CodeQL fixture fix. Both runs gave the same counts and the same killing test IDs. The reader, its test file and the rest ofapplication/single_appare identical at83286787eand at the merge6d45bb8a7.[receipt]and[parent], and bytes mismatch (test 4); descriptor-only lineage read (test 6)include_excerptsis trueinclude_sources=Trueinclude_sourcesonly decides whether the walk keeps each source snapshot and returns them inaccess()["sources"]. Every source is still snapshotted and counted (source_seen), and the reader discardsaccess(), so the reader's result, loads and failures can't differ.Falsejust avoids keeping the list._guarded("manifest", ...)instead of"authorize"authorize.These are the killing test IDs, all in
functional_tests/test_workflow_handoff_result_reader.py:test_a_descriptor_alone_walks_the_lineage_but_reads_no_section,test_a_malformed_consumed_input_receipt_is_refused_before_the_report_is_read,test_a_missing_parent_is_refused_with_the_general_paths_code,test_a_parent_whose_bytes_do_not_match_its_hash_is_refused,test_a_parent_whose_output_disagrees_with_the_receipt_is_refused[parent],[receipt], andtest_a_report_with_real_lineage_reads_after_walking_its_parenttest_a_report_with_real_lineage_reads_after_walking_its_parenttest_a_malformed_consumed_input_receipt_is_refused_before_the_report_is_read,test_a_missing_parent_is_refused_with_the_general_paths_code,test_a_parent_whose_bytes_do_not_match_its_hash_is_refused, andtest_a_parent_whose_output_disagrees_with_the_receipt_is_refused[parent]and[receipt]E2E cost
This was measured in
functional_tests/test_workflow_handoff_end_to_end.py,test_an_accepted_handoff_runs_once_and_posts_its_summary_to_chat, at both sizes, using the file's in-memory canonical store. The figures below are from the review headc07b91151.application/single_appand the E2E file are identical at83286787eand at the merge6d45bb8a7, so they still apply.functional_tests/test_zz_tmp_probe_handoff_lineage_e2e.py. It re-collected that E2E test and added an autouse fixture that counted every node-result load and page read during the hand-off read. The fixture also wrappedauthorize_workflow_node_result_readto count calls and outcomes. The module was deleted after each use, andgit status --porcelainwas empty afterwards. It was never committed.reauthorizestub is for input reauthorization, not this walk. The walk below ran for real: one call per read, and it succeeded every time.-O-O-OFirst measurement. The probe first ran at the original fix commit,
39655e906, before the CodeQL fixture fix rewrote it. Apart fromVERSIONinconfig.py, the application code and the E2E file are identical atc07b91151. That run gave the same load and page counts in every row, and 200-document read times of 1.8534 s (normal) and 1.8789 s (-O). Only wall time varies between runs.The 200-document loads break down as follows. They come to about 7 point reads per document, the same cost model as the general path's per-row walk.
collectmanifesteachmanifestThe report manifest loads twice because, by the end of a 200-document walk, it has been evicted from both the authorizer's LRU (32) and the reader's
_ManifestMemo(64); see Follow-ups. The worker's proof cache applies only inside a worker execution, so it doesn't help this read.Observable. The E2E file still passes without the walk, because it doesn't assert loads. The probe's load counts are the observable that M1 changes: 27 against 2 for 3 documents, and 1,409 against 2 for 200. The unmodified E2E file passes at HEAD in both pytest modes; see Hand-off set.
Probe commands, run from the worktree root at c07b911
Each command ran with the environment above, plus
PROBE_MODEandPROBE_OUT=<session files>\probe_e2e_final.jsonl. They ran one at a time in this order:PROBE_MODE=walk:<py> -u -m pytest functional_tests/test_zz_tmp_probe_handoff_lineage_e2e.py -p no:langsmith_plugin -p no:cacheprovider -q -rfE: exit 0PROBE_MODE=nowalk: the same command: exit 0PROBE_MODE=walk:<py> -O -u -m pytest functional_tests/test_zz_tmp_probe_handoff_lineage_e2e.py -p no:langsmith_plugin -p no:cacheprovider -q -rfE: exit 0PROBE_MODE=nowalk: the same-Ocommand: exit 0CodeQL
I read CodeQL with
gh api "repos/microsoft/simplechat/commits/<FULL SHA>/check-runs?check_name=CodeQL"and thengh api repos/microsoft/simplechat/check-runs/<id>/annotations, once for each head.27d5b5939af137e9e8acbe2e95867b544bf90758(first head)c07b91151bab479e767e74d10398ea59803abb2b(review head)83286787ecd6707ae2e60b668feaa426ab0e2500(pin head)6d45bb8a7419c2abfc7fdd4f4122cfe002cae3e5(final head, V2 merge)Alerts at the final head: none.
The alert on the first head:
__init__method calls overridden method" (py/init-calls-subclass), atfunctional_tests/test_workflow_handoff_result_reader.pyL166.c07b91151reported no new alerts.Other notes:
83286787eand again at6d45bb8a7, the CodeQL check first showedneutral("configuration not found") while theAnalyzejobs were still running. The same check run changed to success when they finished.Analyze (python)succeeded (112250769245 at6d45bb8a7, 112137073045 at83286787e, 112107020803 atc07b91151). Its only annotation is a runner-image notice that theubuntu-latestlabel is migrating, at path.github, levelnotice. It isn't a code alert.83286787eand6d45bb8a7: the runner notice, plus 10 keyword-marker warnings on changed lines of the fix doc (docs/explanation/fixes/WORKFLOW_HANDOFF_RESULT_LINEAGE_FIX.md). They're heuristic text markers, not findings, and the doc has no code.repos/microsoft/simplechat/code-scanning/alerts?pr=1647) returns 403 for this token. It needs a scope I may not request, so the check runs' output and annotations above are the record.Every other check on
6d45bb8a7succeeded, as it did on83286787eandc07b91151:Follow-ups
_ManifestMemo(64). This is harmless but wasteful; pinning the root would fix it.test_docs_release_notes_integrity.pyandtest_docs_link_integrity.pyfail at V2 6bd12cf and at the merge, the same way (see Reruns at the merge). Neither is on this PR's validation list, and this PR doesn't fix them.test_workflow_handoff_builder.py::test_the_phase_4_blueprint_and_payloads_are_unchangedfailed at 15feec6 and atc07b91151, in pytest and pytest -O, alone and in the harness process.schemapin was stale:ab12fb1d…(from 5255406) was expected, and520139cb…was computed. b8e4730 ("Merge V2 a5a5b1c (Merge CSV, Excel, PDF, Word and PowerPoint files in V2 chat and workflows #1641) into Phase 7a") added themergetask schema and didn't refresh the pin.83286787ethe builder passes 32/32 in all four modes, and the harness set passes 285 in both pytest modes.Documentation
### **(v0.261.252)**section with two#### Bug Fixesentries: the lineage re-proof and the recaptured off-golden. It's a pure 15-line insertion, so every existing section is byte-exact. After the merge it's still +15/−0 against V2 6bd12cf.CHAT_ORCHESTRATION_WORKFLOW_HANDOFF.md, "Result reading and delivery" now describes the identity check and the lineage re-proof, which apply to descriptor-only and excerpt reads. The Files table row, the test-table row and Performance are updated too.docs/explanation/fixes/WORKFLOW_HANDOFF_RESULT_LINEAGE_FIX.mdcovers the issue, root cause, files, tests, before and after, and the stale off-golden and its verified recapture. Since review, it also covers the Phase 4 digest pin.Security checklist
@swagger_route(security=get_auth_security()). N/A: no routes are added or changed.sanitize_settings_for_user(). N/A: no settings are sent.