Skip to content

Phase 7a follow-up: re-prove a hand-off report's lineage before reading it (0.261.252) - #1647

Merged
Paul Lizer (paullizer) merged 6 commits into
microsoft:paullizer-react-v2-uifrom
paullizer:paullizer-7a-follow-up-handoff-lineage
Oct 6, 2026
Merged

Paul Lizer (paullizer) merged 6 commits into
microsoft:paullizer-react-v2-uifrom
paullizer:paullizer-7a-follow-up-handoff-lineage

Conversation

@paullizer

@paullizer Paul Lizer (paullizer) commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • The hand-off reader re-proves the report's lineage before reading it. _read_handoff_result now runs the shared node lineage authorizer (authorize_workflow_node_result_read) on the report before it builds a descriptor or loads any text. If the report's consumed-input receipts don't chain to real parent results of this run, the read is refused. That applies to the descriptor-only read, the excerpt read and the stored-context re-check. Before this fix, a report manifest with consumed_inputs: [{"malformed_receipt": true}], saved under its own valid hash, read as available=True and returned its text.
  • The hand-off off-golden passes again. This cherry-picks 7a's unpushed commit 5a625599c (-x), which recaptures orchestration_workflow_handoff_off_golden.json on unmodified V2 a5a5b1c (.248). The merged fixture was captured on f1aeef1 (.233), before Merge CSV, Excel, PDF, Word and PowerPoint files in V2 chat and workflows #1641 changed planner content, and it failed 4/10 at 15feec6. The golden blob is f93f8e4c3dcc263cc5a71186f2cecc8c60574098, byte-identical to the coordinator's independent recapture. Neither file was edited or regenerated here.
  • The Phase 4 digest test passes again (added after review). test_workflow_handoff_builder.py pinned the blueprint schema digest from before Merge CSV, Excel, PDF, Word and PowerPoint files in V2 chat and workflows #1641 added the merge task schema, so it failed at 15feec6. The schema pin now holds the value V2 computes at a5a5b1c, after Merge CSV, Excel, PDF, Word and PowerPoint files in V2 chat and workflows #1641 and before hand-off, so the test still proves hand-off left the Phase 4 schema unchanged. The other four pins are unchanged.
  • Impact: nothing changes while enable_chat_orchestration_workflow_handoff is off, which is the default. With it on, a valid hand-off reads exactly as before, with the same descriptor and text, but each read now also loads the report's lineage (see E2E cost). Only a report whose lineage can't be re-proved is refused.

Changes since review

The coordinator reviewed c07b91151 and asked for one more commit, to refresh the stale Phase 4 schema digest pin in this PR. Then V2 moved, so commit 6 merges the new V2 tip. Nothing else changed.

  • New commit 4, 024004035 ("Refresh the Phase 4 schema digest pin after Merge CSV, Excel, PDF, Word and PowerPoint files in V2 chat and workflows #1641"):
    • functional_tests/test_workflow_handoff_builder.py: PHASE4_DIGESTS["schema"] goes from ab12fb1d… to 520139cbce120ca017106bf200832712913e395c6cd3e6fa9b530eae01cbd66e, and the comment above it says why. No other pin changed. See decision (n) for the file's Version header.
    • docs/explanation/fixes/WORKFLOW_HANDOFF_RESULT_LINEAGE_FIX.md: a Files modified row for the builder test, a "Phase 4 digest pin" validation note, and "about 1.9 seconds" is now "about 1.7 seconds", matching the c07b91151 timings. The release notes and the feature doc are unchanged.
  • The digest value. On a temporary detached worktree at a5a5b1c, _digest(WORKFLOW_BLUEPRINT_SCHEMA) is 520139cb…, the same as at 15feec6 and this PR's heads. Its $defs are file_sync_schedule, handle, merge, merge_options and runner.
  • History. Commits 1–3 keep their SHAs. The VERSION commit was re-applied last as 83286787e, and its git patch-id --stable (8e25d4bb…) and message match c07b91151's. git diff --stat c07b91151 83286787e is the two files above, +17/−2, and application/ is identical. V2 was still 15feec6 then. The push used --force-with-lease pinned to c07b91151.
  • Reruns at 83286787e. The builder passes 32/32 in all four modes, and the harness set passes 285 under pytest and pytest -O. Docs coverage is 7/7, site quality 6/6, and guardrails exit 0 with every check passing. Every other result stands from c07b91151, because no other file changed.
  • New commit 6, 6d45bb8a7 ("Merge V2 6bd12cf (Phase 6b-2: show chat-started workflow runs and their posted results in V2 #1639) into the hand-off lineage follow-up"). V2 moved to 6bd12cf when Phase 6b-2: show chat-started workflow runs and their posted results in V2 #1639 (6b-2, .251) merged.
    • It ran with git -c rerere.enabled=false merge; no recorded resolution was applied. Two files conflicted. config.py keeps VERSION = "0.261.252". The release notes put this PR's .252 section first, then V2's .251 section, then the rest; removing either section reproduces the other side's file byte for byte.
    • Every other file comes unchanged from one side. The 7 files only this PR changed match 83286787e. The 49 only Phase 6b-2: show chat-started workflow runs and their posted results in V2 #1639 changed match V2: application/v2_ui 25, docs 7, ui_tests 9, and V2 JavaScript and XSS-guardrail tests 8. Nothing under application/single_app changed.
    • The merged tree is 3df18e47363e5763d7e36de30391c25d6c2c10a1, the same as the coordinator's independent trial merge. The push was a fast-forward from 83286787e. See Reruns at the merge.

Root cause

read_workflow_result sends a structured hand-off run down _read_handoff_result. That function reads one result: the report that the run's workflow_outputs receipt names. It does three things:

  • It loads the report manifest through load_node_result, which recomputes the node identity from the saved workflow. That's why an edited definition already failed closed before any load.
  • It checks the text output and identity.
  • It requires a completed result.

It never called the lineage authorizer. The general (non-structured) path calls _guarded("authorize", lambda: authorize_workflow_run_read(...)), and the feature doc (L567–568) and the reader test's docstring promised the hand-off did the same. So nothing checked that the report's consumed-input receipts chain to real parent results of this run.

Fix

The fix goes after _require_completed_result and before the result is built, so it covers the descriptor-only read and the excerpt read:

_guarded("authorize", lambda: authorize_workflow_node_result_read(
    workflow, run_id, producer, receipt["result_ref"], reader_user_id=user_id,
    manifest=manifest, load_result=loader, include_sources=False,
))
  • Saved workflow. It walks the saved workflow, not workflow_runtime_store(...).run_definition(). That matches load_node_result in the same function and keeps the edited-workflow fail-closed behavior.
  • User ID. user_id is passed in from read_workflow_result, which has already proved workflow["user_id"] == user_id.
  • Loader. It reuses the reader's _ManifestMemo loader, and passes include_sources=False because the reader discards the access summary. The authorizer never re-resolves sources.
  • Re-check path. authorize_workflow_result_context re-reads through the same function, so stored chat contexts get the walk too.
  • Unchanged. functions_workflow_node_results.py and functions_workflow_results.py aren't touched, and the authorizer itself is unchanged.

Failure codes

Failure Code Logged stage Logged error type
Malformed consumed-input receipt (the repro) workflow_result_invalid authorize AnalysisResultUnavailable (analysis_lineage_invalid)
Missing parent workflow_result_not_found (404) authorize CosmosResourceNotFoundError
Parent output disagrees with the receipt's output_ref (receipt side and parent side) workflow_result_invalid authorize AnalysisResultUnavailable
Parent bytes don't match their hash (same size) workflow_result_invalid authorize WorkflowResultIntegrityError
Edited or re-enabled definition (4 variants, both fixtures) workflow_result_invalid before any load n/a

The general path gives the same code and status for a missing parent (workflow_result_not_found, 404), and a test asserts that.

Commits

  1. 00ba92039: cherry-pick of 5a625599c, the off-golden recapture. Message kept, plus -x.
  2. 3a2bd7a0c: the fix and its tests.
  3. ad06270e5: docs, including the fix doc, the feature doc and the release notes.
  4. 024004035: the Phase 4 schema digest pin refresh and its fix-doc note, added after review.
  5. 83286787e: VERSION = "0.261.252", the last of this PR's own commits.
  6. 6d45bb8a7: the merge of V2 6bd12cf (Phase 6b-2: show chat-started workflow runs and their posted results in V2 #1639), with rerere off.

The branch started from V2 15feec6. V2 hadn't moved when this opened, nor when the CodeQL fix or the pin refresh was pushed; then #1639 merged, and commit 6 merges it. The branch is on paullizer/simplechat because a push to microsoft/simplechat returned 403; the kickoff gives that fallback.

This PR first opened at 27d5b5939. CodeQL flagged one warning in this PR's own test fixture, so I folded a fixture fix into commit 2 and force-pushed with a lease (see decision (m) and CodeQL).

  • git range-diff showed commit 2 changed, and commit 3 and the VERSION commit patch-identical (=). Commit 1 wasn't rewritten.
  • git diff --stat 27d5b5939 c07b91151 is one file: functional_tests/test_workflow_handoff_result_reader.py, +40/−26.

Linked issue

Part of #1549
Part of #1543
Follow-up to #1640

Release Notes & Latest Features

  • New Feature
  • Bug Fix
  • UI Enhancement
  • Breaking Change
  • Internal only

Is this visible to end users?

  • Yes
  • No

Is this admin-facing (Admin Settings, governance, deployment, config)?

  • Yes
  • No

Should this become a Latest Feature card?

  • Yes
  • No
  • Already added

Screenshot needed for the card?

  • Yes
  • No
  • Attached

Version bump

  • application/single_app/config.py VERSION third segment bumped, or not needed because this is docs-only. It goes from 0.261.250 to 0.261.252 in commit 5. .251 is 6b-2 (Phase 6b-2: show chat-started workflow runs and their posted results in V2 #1639), which merged first; the V2 merge keeps .252, so no renumber was needed.
  • deployers/version.txt bumped, or not needed because deployers/ was not changed. deployers/ isn't changed.

Testing / validation

Environment. Every job ran from the worktree root, one process at a time, with a 1,800 s timeout. The full set ran at the review head c07b91151. Cells marked (at 83286787e) reran at the new head after the pin refresh. Between the two heads only test_workflow_handoff_builder.py and the fix doc differ, so every other row stands. <py> is C:\Users\paullizer\.copilot\session-state\aa1bf1e0-b10d-4312-8dbb-143e9e9f9966\files\venv\Scripts\python.exe (Flask 3.1.3). The environment is PYTHONIOENCODING=utf-8, PYTHONUTF8=1 and PYTHONPATH=application/single_app;functional_tests. An exit of 0xC000026B counts as interrupted, not as a pass.

Compared with the first head. The same 106 jobs first ran at 27d5b5939, before the CodeQL fixture fix. Every row's exit code and pass/fail counts are identical in both runs.

These are the command forms; <file> is the file in each row, and every exact command is listed at the end:

  • pytest: <py> -u -m pytest <file> -p no:langsmith_plugin -p no:cacheprovider -q -rfE
  • pytest -O: <py> -O -u -m pytest <file> -p no:langsmith_plugin -p no:cacheprovider -q -rfE
  • script: <py> -u <file>
  • script -O: <py> -O -u <file>

In the full command list at the end, <flags> stands for exactly -p no:langsmith_plugin -p no:cacheprovider -q -rfE.

Totals. At c07b91151: 106 jobs; 102 exited 0; 4 didn't, all on the stale builder pin, which was pre-existing at 15feec6 (see Baseline classification). At 83286787e: 8 reruns of what changed since review (the builder in four modes, the harness set in both pytest modes, and the two docs checks); all 8 exited 0. At the V2 merge 6d45bb8a7: 10 reruns, all exited 0; see Reruns at the merge.

Reruns at the merge

6d45bb8a7 merges V2 6bd12cf (#1639); see Changes since review. It changes nothing under application/single_app, and none of the reader, off-golden, builder or E2E test files or the Python test support. So I reran this PR's own test files, the docs checks and the guardrails, one at a time from the worktree root with the environment above:

  1. <py> -u -m pytest functional_tests/test_workflow_handoff_result_reader.py <flags>: exit 0, 81 passed
  2. <py> -O -u -m pytest functional_tests/test_workflow_handoff_result_reader.py <flags>: exit 0, 81 passed
  3. <py> -u -m pytest functional_tests/test_orchestration_workflow_handoff_off_golden.py <flags>: exit 0, 10 passed
  4. <py> -O -u -m pytest functional_tests/test_orchestration_workflow_handoff_off_golden.py <flags>: exit 0, 10 passed
  5. <py> -u -m pytest functional_tests/test_workflow_handoff_builder.py <flags>: exit 0, 32 passed
  6. <py> -O -u -m pytest functional_tests/test_workflow_handoff_builder.py <flags>: exit 0, 32 passed
  7. <py> -u .\scripts\build_docs_inventory.py: exit 0; the inventory is unchanged apart from line endings, restored with git checkout
  8. <py> -u functional_tests/test_docs_app_surface_coverage.py: exit 0, 7/7 checks passed
  9. <py> -u functional_tests/test_docs_site_quality.py: exit 0, 6/6 checks passed
  10. The guardrails script above, with base 6bd12cfcdfefe83df58c70bff64e67d7e130363d, HEAD and r1: exit 0. It reviewed the same 9 files as at 83286787e, and every check passed. The malicious-PR review has 268 findings and 0 blockers, with the same per-file counts.

Two docs checks outside this PR's list fail at the merge. Each fails the same way at V2 6bd12cf, rerun on a temporary detached worktree that was removed afterwards, so both are pre-existing:

Command V2 6bd12cf 6d45bb8a7
<py> -u functional_tests/test_docs_release_notes_integrity.py exit 1, 0/1: 186 releases aren't on any generated page exit 1, 0/1: 187, adding v0.261.252
<py> -u functional_tests/test_docs_link_integrity.py exit 1, 2/5 exit 1, 2/5
  • Release-notes pages. 185 releases were already missing at 15feec6, and no V2 PR regenerates the pages; .251 and .250 are missing too.
  • Links. The same 23 broken links at both: 11 relative markdown links, 10 site-page links and 2 features.yml links. None is in this PR's docs. The relative-link total goes from 1,349 to 1,354 with this PR's 5 new links, which all resolve.

Repro, before and after

The repro is a session script, not a committed test. It builds HandoffFixture() from the worktree under test, with the network blocked, and then:

  1. deep-copies the report manifest and sets consumed_inputs = [{"malformed_receipt": True}];
  2. saves it with fixture.store.save(...), so it gets a fresh, valid content hash;
  3. points receipt["result_ref"] at the new reference;
  4. reads it with excerpts and descriptor-only;
  5. calls the real authorize_workflow_node_result_read on the same reference.

Each run used the environment above, from the root of the worktree under test: <py> -u repro_r1.py <worktree> and <py> -O -u repro_r1.py <worktree>.

Commit Mode include_excerpts=True Descriptor only Authorizer on the same reference
15feec6 (before) python available=True, report text returned, 2 loads available=True, 1 load analysis_lineage_invalid, mapped to workflow_result_invalid
15feec6 (before) python -O available=True, report text returned, 2 loads available=True, 1 load analysis_lineage_invalid, mapped to workflow_result_invalid
c07b911 (after) python refused with workflow_result_invalid, 1 load refused with workflow_result_invalid, 1 load analysis_lineage_invalid, mapped to workflow_result_invalid
c07b911 (after) python -O refused with workflow_result_invalid, 1 load refused with workflow_result_invalid, 1 load analysis_lineage_invalid, mapped to workflow_result_invalid

The "after" rows ran at the review head, c07b911; application/single_app and the reader test file are identical at 83286787e and 6d45bb8a7. They first ran at 27d5b59 with identical output, and were rerun because the repro imports HandoffFixture, which the CodeQL fix restructured. After the fix, the single load is the report manifest; the text section is never read. The reader logs {'code': 'workflow_result_invalid', 'stage': 'authorize', 'error_type': 'AnalysisResultUnavailable'}, with no identifiers.

Raw output
=== before 15feec650 (15feec650) python
include_excerpts=True: available=True excerpts=['The review found two contracts that need an owner.'] loads=2
include_excerpts=False: available=True excerpts=[] loads=1
authorizer: AnalysisResultUnavailable(analysis_lineage_invalid) -> workflow_result_invalid
exit=0
=== before 15feec650 (15feec650) python -O
include_excerpts=True: available=True excerpts=['The review found two contracts that need an owner.'] loads=2
include_excerpts=False: available=True excerpts=[] loads=1
authorizer: AnalysisResultUnavailable(analysis_lineage_invalid) -> workflow_result_invalid
exit=0
=== after HEAD (c07b91151) python
include_excerpts=True: closed workflow_result_invalid loads=1
include_excerpts=False: closed workflow_result_invalid loads=1
authorizer: AnalysisResultUnavailable(analysis_lineage_invalid) -> workflow_result_invalid
exit=0
=== after HEAD (c07b91151) python -O
include_excerpts=True: closed workflow_result_invalid loads=1
include_excerpts=False: closed workflow_result_invalid loads=1
authorizer: AnalysisResultUnavailable(analysis_lineage_invalid) -> workflow_result_invalid
exit=0

The [LOG] [WorkflowResults] Workflow result unavailable lines from the two "after" runs are omitted above; the paragraph before this block quotes their payload.

Four modes

File pytest pytest -O script script -O
test_workflow_handoff_result_reader.py 81 passed 81 passed 81 passed 81 passed
test_orchestration_workflow_handoff_off_golden.py 10 passed 10 passed 10 passed 10 passed
test_workflow_handoff_builder.py 32 passed (at 83286787e) 32 passed (at 83286787e) 32 passed (at 83286787e) 32 passed (at 83286787e)

Goldens

At 15feec6 the coordinator saw 9, 5, 5 and 6 passes for the first four in both pytest modes. The counts here match.

File pytest pytest -O
test_orchestration_workflow_setting_off_golden.py 9 passed 9 passed
test_orchestration_workflow_runs_off_golden.py 5 passed 5 passed
test_orchestration_workflow_results_off_golden.py 5 passed 5 passed
test_workflow_chat_delivery_off_golden.py 6 passed 6 passed
test_orchestration_reference_authorizer_golden.py 2 passed, 74 subtests passed 2 passed, 74 subtests passed
test_workflow_run_time_context.py 13 passed 13 passed

Hand-off set

File pytest pytest -O
test_orchestration_workflow_handoff_adapter.py 23 passed 23 passed
test_orchestration_workflow_handoff_capability.py 90 passed 90 passed
test_orchestration_workflow_handoff_imports.py 31 passed 31 passed
test_orchestration_workflow_handoff_planner.py 52 passed 52 passed
test_orchestration_workflow_handoff_routes.py 61 passed 61 passed
test_workflow_handoff_builder.py 32 passed (at 83286787e) 32 passed (at 83286787e)
test_workflow_handoff_end_to_end.py 3 passed 3 passed
test_workflow_handoff_lifecycle.py 26 passed 26 passed
test_workflow_handoff_origin.py 34 passed 34 passed
route_tests/test_route_blueprint_policy_inventory.py 12 passed 12 passed
route_tests/test_route_unauthenticated_policy_contract.py 7 passed 7 passed
route_tests/test_route_policy_test_coverage.py 2 passed 2 passed

Result readers and delivery

File pytest pytest -O
test_workflow_result_reader.py 99 passed 99 passed
test_workflow_result_orchestration_lineage.py 11 passed 11 passed
test_workflow_result_review_paths.py 2 passed 2 passed
test_workflow_result_routes.py 26 passed 26 passed
test_workflow_result_chat_routes.py 12 passed 12 passed
test_workflow_result_followup.py 91 passed 91 passed
test_workflow_result_privacy_import_cycle.py 11 passed 11 passed
test_workflow_result_contract.py 10 passed 10 passed
test_workflow_result_store.py 38 passed, 93 subtests passed 38 passed, 93 subtests passed
test_workflow_result_masking.py 49 passed 49 passed
route_tests/test_workflow_result_context_policy.py 38 passed 38 passed
test_workflow_chat_delivery_concurrency.py 5 passed 5 passed
test_workflow_chat_delivery_contract.py 72 passed 72 passed
test_workflow_chat_delivery_control_pins.py 35 passed 35 passed
test_workflow_chat_delivery_imports.py 31 passed 31 passed
test_workflow_chat_delivery_loop.py 9 passed 9 passed
test_workflow_chat_delivery_notice_and_unread.py 25 passed 25 passed
test_workflow_chat_delivery_placement_and_masking.py 8 passed 8 passed
test_workflow_chat_delivery_projection_and_seed.py 16 passed 16 passed
test_workflow_chat_delivery_refusals.py 2 passed 2 passed
test_workflow_chat_delivery_save_guard.py 14 passed 14 passed
test_workflow_chat_delivery_status_route.py 48 passed 48 passed
test_workflow_chat_delivery_worker.py 112 passed 112 passed
test_orchestration_workflow_results_aliases.py 10 passed 10 passed
test_orchestration_workflow_results_answer.py 18 passed 18 passed
test_orchestration_workflow_results_capability.py 70 passed 70 passed
test_orchestration_workflow_results_imports.py 26 passed 26 passed
test_orchestration_workflow_results_reads.py 53 passed 53 passed
test_orchestration_workflow_results_time_zone.py 8 passed 8 passed
test_chat_workflow_results_admin.py 15 passed 15 passed

Harness set

Files pytest pytest -O
One process: test_m365_run_as_self_authored.py, test_workflow_assist_dry_run_parity.py, test_workflow_draft_save_parity.py, test_workflow_draft_service.py, test_workflow_draft_v2_round_trip.py, test_workflow_handoff_builder.py, test_workflow_handoff_origin.py, test_workflow_origin_provenance.py 285 passed (at 83286787e) 285 passed (at 83286787e)

Baseline classification

At the review head c07b91151, four jobs didn't exit 0, all on the same test. I reran each one on a temporary detached worktree at 15feec6, with the same commands, environment and runner. The worktree was created with git worktree add --detach <session files>\base_15feec650 15feec6509f0846238362838103be6e76d19f8c1 (no git stash) and removed afterwards. The first run at 27d5b5939 gave the same counts as c07b91151 in all four rows. Commit 4 removes the cause; the last result column is from the new head.

File Mode c07b91151 15feec6 83286787e Classification
test_workflow_handoff_builder.py pytest 1 failed, 31 passed 1 failed, 31 passed 32 passed Pre-existing; fixed by commit 4
test_workflow_handoff_builder.py pytest -O 1 failed, 31 passed 1 failed, 31 passed 32 passed Pre-existing; fixed by commit 4
Harness set, one process pytest 1 failed, 284 passed 1 failed, 284 passed 285 passed Pre-existing (same test); fixed by commit 4
Harness set, one process pytest -O 1 failed, 284 passed 1 failed, 284 passed 285 passed Pre-existing (same test); fixed by commit 4

All 8 failing runs failed the same test, test_workflow_handoff_builder.py::test_the_phase_4_blueprint_and_payloads_are_unchanged.

  • Which digest: only schema, which is _digest(drafts.WORKFLOW_BLUEPRINT_SCHEMA). It was pinned to ab12fb1d29c7eaa028a534efa4f610795a31e25638643f30a8ec399b493d9b26 and computes to 520139cbce120ca017106bf200832712913e395c6cd3e6fa9b530eae01cbd66e at a5a5b1c, at 15feec6 and at every head of this PR. The payload_email, payload_review, validate and dry_review digests all match their pins.
  • Why it was stale: the pin dates from 5255406. Later, b8e4730 ("Merge V2 a5a5b1c (Merge CSV, Excel, PDF, Word and PowerPoint files in V2 chat and workflows #1641) into Phase 7a") added the merge task property to the blueprint schema and didn't refresh the pin.
  • Resolution: after review, the coordinator decided to refresh the pin in this PR. Commit 4 sets it to the a5a5b1c value; see Changes since review.

No job was interrupted or timed out, so no job needed a rerun on its own.

Baseline commands, run from the 15feec6 worktree
  1. <py> -u -m pytest functional_tests/test_workflow_handoff_builder.py -p no:langsmith_plugin -p no:cacheprovider -q -rfE: exit 1, 1 failed, 31 passed
  2. <py> -O -u -m pytest functional_tests/test_workflow_handoff_builder.py -p no:langsmith_plugin -p no:cacheprovider -q -rfE: exit 1, 1 failed, 31 passed
  3. <py> -u -m pytest functional_tests/test_m365_run_as_self_authored.py functional_tests/test_workflow_assist_dry_run_parity.py functional_tests/test_workflow_draft_save_parity.py functional_tests/test_workflow_draft_service.py functional_tests/test_workflow_draft_v2_round_trip.py functional_tests/test_workflow_handoff_builder.py functional_tests/test_workflow_handoff_origin.py functional_tests/test_workflow_origin_provenance.py -p no:langsmith_plugin -p no:cacheprovider -q -rfE: exit 1, 1 failed, 284 passed
  4. The same as 3 with -O: exit 1, 1 failed, 284 passed

Mutations and E2E probe

See the Kill table and E2E cost sections below.

Docs

The inventory steps ran at the review head c07b91151, from the worktree root, with the environment above; they first ran at 27d5b5939 with the same results. The two docs checks reran at 83286787e, after the fix-doc update, with the same results. The inventory step and both checks reran at the merge 6d45bb8a7, with the same results again.

Command Result
<py> -u .\scripts\build_docs_inventory.py exit 0. Counts: capabilities 126, admin_groups 15, admin_tabs 49, admin_sections 110, actions 32, chat_controls 47, sub_elements 3, app_pages 29, feature_surfaces 25.
git diff --quiet after regenerating exit 0, so the inventory is unchanged. git status flags docs/_data/app_surface.yml only because the generator writes LF and the working copy is CRLF; git diff --ignore-cr-at-eol --stat is empty. I restored the file with git checkout.
<py> -u functional_tests/test_docs_app_surface_coverage.py (__main__) exit 0, 7/7 checks passed, at both heads
<py> -u functional_tests/test_docs_site_quality.py (__main__) exit 0, 6/6 checks passed, at both heads

The inventory is unchanged because this PR adds no setting, admin tab, action, chat control or page.

Guardrails

Command: <py> C:\Users\paullizer\.copilot\session-state\151a0b09-0034-492f-bac4-4a1ca884309b\files\run_v2_guardrails.py <worktree> 15feec6509f0846238362838103be6e76d19f8c1 HEAD r1. I ran it at the review head c07b91151 and again at the new head 83286787e, and it exited 0 both times. It first ran at 27d5b5939, which gave the same results as c07b91151, including the finding totals by severity and by file. At the merge 6d45bb8a7 it ran against base 6bd12cf; see Reruns at the merge.

At 83286787e the script reviewed 9 changed files: 2 app .py, 2 on the XSS surface, 2 route .py, and 0 V2 ts/tsx. At c07b91151 it reviewed 8, because the builder test wasn't changed yet.

Check Result (same at both heads unless noted)
broken-access-control PASS: passed for 2 files
xss-sinks PASS: passed for 2 files
swagger-routes PASS: the changed Python files define no Flask routes, so the check was skipped
python-syntax PASS: 513 top-level single_app .py files compile
malicious-pr-review PASS: 0 blockers; 268 heuristic findings at 83286787e, and 267 at c07b91151. The report's final verdict is "Needs investigation"
v2-ts-sinks (added lines) PASS: no ts/tsx files changed

Breakdown of the malicious-PR-review findings at 83286787e:

  • Severity: 265 Important, 3 Moderate. At c07b91151 it was 264 Important and 3 Moderate.
  • By file:
    • 240 in functional_tests/test_support/orchestration_workflow_handoff_off_golden.json, the cherry-picked recapture. Its blob is f93f8e4c3dcc263cc5a71186f2cecc8c60574098, byte-identical to the coordinator's independent recapture.
    • 16 in functional_tests/test_workflow_handoff_result_reader.py.
    • 11 in the new fix doc. At c07b91151 it was 10. The new one is the "external connection" marker on the #1641 link at L178, in the Phase 4 digest pin note.
    • 1 in the release notes.
    • None in functions_workflow_result_reader.py, config.py or test_workflow_handoff_builder.py.
  • What the test-file findings are: keyword markers on test code. Examples:
    • HandoffFixture();
    • the two monkeypatch.setattr(reader, "authorize_workflow_node_result_read", ...) calls. One shows the malformed manifest would read as a normal report without the walk, so the refusal comes from the walk. The other shows the strong control's result is identical without the walk and only the parent's load disappears, and it pins the walk's arguments;
    • a comment;
    • the expected failure-code dict;
    • an edited task's instructions in the edited-definition test.
  • External URLs: only microsoft/simplechat issue and PR links in the docs (Let chat orchestration propose, create, run and hand off saved workflows #1543, Hand off big one-time jobs from chat to a durable workflow #1549, Phase 7a: hand off large chat requests to a one-time workflow (server, 0.261.238) #1640, Merge CSV, Excel, PDF, Word and PowerPoint files in V2 chat and workflows #1641).

Every validation command

All 106 jobs at `c07b91151` and 8 at `83286787e`, in run order, with exit code and pass/fail counts (timings and warning counts omitted)

At the review head c07b91151:

  1. <py> -u -m pytest functional_tests/test_workflow_handoff_result_reader.py <flags>: exit 0, 81 passed
  2. <py> -O -u -m pytest functional_tests/test_workflow_handoff_result_reader.py <flags>: exit 0, 81 passed
  3. <py> -u functional_tests/test_workflow_handoff_result_reader.py: exit 0, 81 passed
  4. <py> -O -u functional_tests/test_workflow_handoff_result_reader.py: exit 0, 81 passed
  5. <py> -u -m pytest functional_tests/test_orchestration_workflow_handoff_off_golden.py <flags>: exit 0, 10 passed
  6. <py> -O -u -m pytest functional_tests/test_orchestration_workflow_handoff_off_golden.py <flags>: exit 0, 10 passed
  7. <py> -u functional_tests/test_orchestration_workflow_handoff_off_golden.py: exit 0, 10 passed
  8. <py> -O -u functional_tests/test_orchestration_workflow_handoff_off_golden.py: exit 0, 10 passed
  9. <py> -u -m pytest functional_tests/test_orchestration_workflow_setting_off_golden.py <flags>: exit 0, 9 passed
  10. <py> -O -u -m pytest functional_tests/test_orchestration_workflow_setting_off_golden.py <flags>: exit 0, 9 passed
  11. <py> -u -m pytest functional_tests/test_orchestration_workflow_runs_off_golden.py <flags>: exit 0, 5 passed
  12. <py> -O -u -m pytest functional_tests/test_orchestration_workflow_runs_off_golden.py <flags>: exit 0, 5 passed
  13. <py> -u -m pytest functional_tests/test_orchestration_workflow_results_off_golden.py <flags>: exit 0, 5 passed
  14. <py> -O -u -m pytest functional_tests/test_orchestration_workflow_results_off_golden.py <flags>: exit 0, 5 passed
  15. <py> -u -m pytest functional_tests/test_workflow_chat_delivery_off_golden.py <flags>: exit 0, 6 passed
  16. <py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_off_golden.py <flags>: exit 0, 6 passed
  17. <py> -u -m pytest functional_tests/test_orchestration_reference_authorizer_golden.py <flags>: exit 0, 2 passed, 74 subtests passed
  18. <py> -O -u -m pytest functional_tests/test_orchestration_reference_authorizer_golden.py <flags>: exit 0, 2 passed, 74 subtests passed
  19. <py> -u -m pytest functional_tests/test_workflow_run_time_context.py <flags>: exit 0, 13 passed
  20. <py> -O -u -m pytest functional_tests/test_workflow_run_time_context.py <flags>: exit 0, 13 passed
  21. <py> -u -m pytest functional_tests/test_orchestration_workflow_handoff_adapter.py <flags>: exit 0, 23 passed
  22. <py> -O -u -m pytest functional_tests/test_orchestration_workflow_handoff_adapter.py <flags>: exit 0, 23 passed
  23. <py> -u -m pytest functional_tests/test_orchestration_workflow_handoff_capability.py <flags>: exit 0, 90 passed
  24. <py> -O -u -m pytest functional_tests/test_orchestration_workflow_handoff_capability.py <flags>: exit 0, 90 passed
  25. <py> -u -m pytest functional_tests/test_orchestration_workflow_handoff_imports.py <flags>: exit 0, 31 passed
  26. <py> -O -u -m pytest functional_tests/test_orchestration_workflow_handoff_imports.py <flags>: exit 0, 31 passed
  27. <py> -u -m pytest functional_tests/test_orchestration_workflow_handoff_planner.py <flags>: exit 0, 52 passed
  28. <py> -O -u -m pytest functional_tests/test_orchestration_workflow_handoff_planner.py <flags>: exit 0, 52 passed
  29. <py> -u -m pytest functional_tests/test_orchestration_workflow_handoff_routes.py <flags>: exit 0, 61 passed
  30. <py> -O -u -m pytest functional_tests/test_orchestration_workflow_handoff_routes.py <flags>: exit 0, 61 passed
  31. <py> -u -m pytest functional_tests/test_workflow_handoff_builder.py <flags>: exit 1, 1 failed, 31 passed
    • FAILED functional_tests/test_workflow_handoff_builder.py::test_the_phase_4_blueprint_and_payloads_are_unchanged
  32. <py> -O -u -m pytest functional_tests/test_workflow_handoff_builder.py <flags>: exit 1, 1 failed, 31 passed
    • FAILED functional_tests/test_workflow_handoff_builder.py::test_the_phase_4_blueprint_and_payloads_are_unchanged
  33. <py> -u -m pytest functional_tests/test_workflow_handoff_end_to_end.py <flags>: exit 0, 3 passed
  34. <py> -O -u -m pytest functional_tests/test_workflow_handoff_end_to_end.py <flags>: exit 0, 3 passed
  35. <py> -u -m pytest functional_tests/test_workflow_handoff_lifecycle.py <flags>: exit 0, 26 passed
  36. <py> -O -u -m pytest functional_tests/test_workflow_handoff_lifecycle.py <flags>: exit 0, 26 passed
  37. <py> -u -m pytest functional_tests/test_workflow_handoff_origin.py <flags>: exit 0, 34 passed
  38. <py> -O -u -m pytest functional_tests/test_workflow_handoff_origin.py <flags>: exit 0, 34 passed
  39. <py> -u -m pytest functional_tests/route_tests/test_route_blueprint_policy_inventory.py <flags>: exit 0, 12 passed
  40. <py> -O -u -m pytest functional_tests/route_tests/test_route_blueprint_policy_inventory.py <flags>: exit 0, 12 passed
  41. <py> -u -m pytest functional_tests/route_tests/test_route_unauthenticated_policy_contract.py <flags>: exit 0, 7 passed
  42. <py> -O -u -m pytest functional_tests/route_tests/test_route_unauthenticated_policy_contract.py <flags>: exit 0, 7 passed
  43. <py> -u -m pytest functional_tests/route_tests/test_route_policy_test_coverage.py <flags>: exit 0, 2 passed
  44. <py> -O -u -m pytest functional_tests/route_tests/test_route_policy_test_coverage.py <flags>: exit 0, 2 passed
  45. <py> -u -m pytest functional_tests/test_workflow_result_reader.py <flags>: exit 0, 99 passed
  46. <py> -O -u -m pytest functional_tests/test_workflow_result_reader.py <flags>: exit 0, 99 passed
  47. <py> -u -m pytest functional_tests/test_workflow_result_orchestration_lineage.py <flags>: exit 0, 11 passed
  48. <py> -O -u -m pytest functional_tests/test_workflow_result_orchestration_lineage.py <flags>: exit 0, 11 passed
  49. <py> -u -m pytest functional_tests/test_workflow_result_review_paths.py <flags>: exit 0, 2 passed
  50. <py> -O -u -m pytest functional_tests/test_workflow_result_review_paths.py <flags>: exit 0, 2 passed
  51. <py> -u -m pytest functional_tests/test_workflow_result_routes.py <flags>: exit 0, 26 passed
  52. <py> -O -u -m pytest functional_tests/test_workflow_result_routes.py <flags>: exit 0, 26 passed
  53. <py> -u -m pytest functional_tests/test_workflow_result_chat_routes.py <flags>: exit 0, 12 passed
  54. <py> -O -u -m pytest functional_tests/test_workflow_result_chat_routes.py <flags>: exit 0, 12 passed
  55. <py> -u -m pytest functional_tests/test_workflow_result_followup.py <flags>: exit 0, 91 passed
  56. <py> -O -u -m pytest functional_tests/test_workflow_result_followup.py <flags>: exit 0, 91 passed
  57. <py> -u -m pytest functional_tests/test_workflow_result_privacy_import_cycle.py <flags>: exit 0, 11 passed
  58. <py> -O -u -m pytest functional_tests/test_workflow_result_privacy_import_cycle.py <flags>: exit 0, 11 passed
  59. <py> -u -m pytest functional_tests/test_workflow_result_contract.py <flags>: exit 0, 10 passed
  60. <py> -O -u -m pytest functional_tests/test_workflow_result_contract.py <flags>: exit 0, 10 passed
  61. <py> -u -m pytest functional_tests/test_workflow_result_store.py <flags>: exit 0, 38 passed, 93 subtests passed
  62. <py> -O -u -m pytest functional_tests/test_workflow_result_store.py <flags>: exit 0, 38 passed, 93 subtests passed
  63. <py> -u -m pytest functional_tests/test_workflow_result_masking.py <flags>: exit 0, 49 passed
  64. <py> -O -u -m pytest functional_tests/test_workflow_result_masking.py <flags>: exit 0, 49 passed
  65. <py> -u -m pytest functional_tests/route_tests/test_workflow_result_context_policy.py <flags>: exit 0, 38 passed
  66. <py> -O -u -m pytest functional_tests/route_tests/test_workflow_result_context_policy.py <flags>: exit 0, 38 passed
  67. <py> -u -m pytest functional_tests/test_workflow_chat_delivery_concurrency.py <flags>: exit 0, 5 passed
  68. <py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_concurrency.py <flags>: exit 0, 5 passed
  69. <py> -u -m pytest functional_tests/test_workflow_chat_delivery_contract.py <flags>: exit 0, 72 passed
  70. <py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_contract.py <flags>: exit 0, 72 passed
  71. <py> -u -m pytest functional_tests/test_workflow_chat_delivery_control_pins.py <flags>: exit 0, 35 passed
  72. <py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_control_pins.py <flags>: exit 0, 35 passed
  73. <py> -u -m pytest functional_tests/test_workflow_chat_delivery_imports.py <flags>: exit 0, 31 passed
  74. <py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_imports.py <flags>: exit 0, 31 passed
  75. <py> -u -m pytest functional_tests/test_workflow_chat_delivery_loop.py <flags>: exit 0, 9 passed
  76. <py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_loop.py <flags>: exit 0, 9 passed
  77. <py> -u -m pytest functional_tests/test_workflow_chat_delivery_notice_and_unread.py <flags>: exit 0, 25 passed
  78. <py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_notice_and_unread.py <flags>: exit 0, 25 passed
  79. <py> -u -m pytest functional_tests/test_workflow_chat_delivery_placement_and_masking.py <flags>: exit 0, 8 passed
  80. <py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_placement_and_masking.py <flags>: exit 0, 8 passed
  81. <py> -u -m pytest functional_tests/test_workflow_chat_delivery_projection_and_seed.py <flags>: exit 0, 16 passed
  82. <py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_projection_and_seed.py <flags>: exit 0, 16 passed
  83. <py> -u -m pytest functional_tests/test_workflow_chat_delivery_refusals.py <flags>: exit 0, 2 passed
  84. <py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_refusals.py <flags>: exit 0, 2 passed
  85. <py> -u -m pytest functional_tests/test_workflow_chat_delivery_save_guard.py <flags>: exit 0, 14 passed
  86. <py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_save_guard.py <flags>: exit 0, 14 passed
  87. <py> -u -m pytest functional_tests/test_workflow_chat_delivery_status_route.py <flags>: exit 0, 48 passed
  88. <py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_status_route.py <flags>: exit 0, 48 passed
  89. <py> -u -m pytest functional_tests/test_workflow_chat_delivery_worker.py <flags>: exit 0, 112 passed
  90. <py> -O -u -m pytest functional_tests/test_workflow_chat_delivery_worker.py <flags>: exit 0, 112 passed
  91. <py> -u -m pytest functional_tests/test_orchestration_workflow_results_aliases.py <flags>: exit 0, 10 passed
  92. <py> -O -u -m pytest functional_tests/test_orchestration_workflow_results_aliases.py <flags>: exit 0, 10 passed
  93. <py> -u -m pytest functional_tests/test_orchestration_workflow_results_answer.py <flags>: exit 0, 18 passed
  94. <py> -O -u -m pytest functional_tests/test_orchestration_workflow_results_answer.py <flags>: exit 0, 18 passed
  95. <py> -u -m pytest functional_tests/test_orchestration_workflow_results_capability.py <flags>: exit 0, 70 passed
  96. <py> -O -u -m pytest functional_tests/test_orchestration_workflow_results_capability.py <flags>: exit 0, 70 passed
  97. <py> -u -m pytest functional_tests/test_orchestration_workflow_results_imports.py <flags>: exit 0, 26 passed
  98. <py> -O -u -m pytest functional_tests/test_orchestration_workflow_results_imports.py <flags>: exit 0, 26 passed
  99. <py> -u -m pytest functional_tests/test_orchestration_workflow_results_reads.py <flags>: exit 0, 53 passed
  100. <py> -O -u -m pytest functional_tests/test_orchestration_workflow_results_reads.py <flags>: exit 0, 53 passed
  101. <py> -u -m pytest functional_tests/test_orchestration_workflow_results_time_zone.py <flags>: exit 0, 8 passed
  102. <py> -O -u -m pytest functional_tests/test_orchestration_workflow_results_time_zone.py <flags>: exit 0, 8 passed
  103. <py> -u -m pytest functional_tests/test_chat_workflow_results_admin.py <flags>: exit 0, 15 passed
  104. <py> -O -u -m pytest functional_tests/test_chat_workflow_results_admin.py <flags>: exit 0, 15 passed
  105. <py> -u -m pytest functional_tests/test_m365_run_as_self_authored.py functional_tests/test_workflow_assist_dry_run_parity.py functional_tests/test_workflow_draft_save_parity.py functional_tests/test_workflow_draft_service.py functional_tests/test_workflow_draft_v2_round_trip.py functional_tests/test_workflow_handoff_builder.py functional_tests/test_workflow_handoff_origin.py functional_tests/test_workflow_origin_provenance.py <flags>: exit 1, 1 failed, 284 passed
    - FAILED functional_tests/test_workflow_handoff_builder.py::test_the_phase_4_blueprint_and_payloads_are_unchanged
  106. <py> -O -u -m pytest functional_tests/test_m365_run_as_self_authored.py functional_tests/test_workflow_assist_dry_run_parity.py functional_tests/test_workflow_draft_save_parity.py functional_tests/test_workflow_draft_service.py functional_tests/test_workflow_draft_v2_round_trip.py functional_tests/test_workflow_handoff_builder.py functional_tests/test_workflow_handoff_origin.py functional_tests/test_workflow_origin_provenance.py <flags>: exit 1, 1 failed, 284 passed
    - FAILED functional_tests/test_workflow_handoff_builder.py::test_the_phase_4_blueprint_and_payloads_are_unchanged

At the new head 83286787e:

  1. <py> -u -m pytest functional_tests/test_workflow_handoff_builder.py <flags>: exit 0, 32 passed
  2. <py> -O -u -m pytest functional_tests/test_workflow_handoff_builder.py <flags>: exit 0, 32 passed
  3. <py> -u functional_tests/test_workflow_handoff_builder.py: exit 0, 32 passed
  4. <py> -O -u functional_tests/test_workflow_handoff_builder.py: exit 0, 32 passed
  5. <py> -u -m pytest functional_tests/test_m365_run_as_self_authored.py functional_tests/test_workflow_assist_dry_run_parity.py functional_tests/test_workflow_draft_save_parity.py functional_tests/test_workflow_draft_service.py functional_tests/test_workflow_draft_v2_round_trip.py functional_tests/test_workflow_handoff_builder.py functional_tests/test_workflow_handoff_origin.py functional_tests/test_workflow_origin_provenance.py <flags>: exit 0, 285 passed
  6. <py> -O -u -m pytest functional_tests/test_m365_run_as_self_authored.py functional_tests/test_workflow_assist_dry_run_parity.py functional_tests/test_workflow_draft_save_parity.py functional_tests/test_workflow_draft_service.py functional_tests/test_workflow_draft_v2_round_trip.py functional_tests/test_workflow_handoff_builder.py functional_tests/test_workflow_handoff_origin.py functional_tests/test_workflow_origin_provenance.py <flags>: exit 0, 285 passed
  7. <py> -u functional_tests/test_docs_app_surface_coverage.py: exit 0, 7/7 checks passed
  8. <py> -u functional_tests/test_docs_site_quality.py: exit 0, 6/6 checks passed

Decisions made autonomously

None of these deviate from the kickoff or the review. Each one is a choice they left open.

  • (a) How the strong control's parent is built. It uses the runner's contract, not hand-written receipts. collect is an engine node with no task, so LineageFixture builds its result the way the runner builds a control node's result:

    • workflow_node_identity and workflow_execution_id, then the runner's private _build_task_result(..., "workflow-result-v2");
    • saved through the real persist_workflow_task_result;
    • the receipt is the real open_workflow_record_input(..., output_name="records") receipt, plus the input_name the runner adds;
    • the report itself is still built with workflow_execution_scope, build_workflow_task_result and persist_workflow_task_result.

    There's no frozen loop in this fixture. The E2E covers the real collect, each and frozen-loop lineage.

  • (b) How the corrupt-bytes case is checked. The test fixture's in-memory store doesn't verify digests on load, but the real store does. So this case wraps the loader with the real _verify_payload and proves the walk turns WorkflowResultIntegrityError into workflow_result_invalid.

  • (c) A test was red between commits. test_version_is_at_least_the_lineage_release (assert_app_version_at_least("0.261.252")) failed from Commit 2 until the VERSION commit, because the kickoff puts that bump last. It passes at HEAD.

  • (d) How M4 is killed. _guarded("manifest") and _guarded("authorize") map the same exceptions to the same codes. Only the logged closed stage tells them apart, so the new refusal tests assert that the logged stage is authorize.

  • (e) An opt-in lineage parameter on the old fixture. HandoffFixture(lineage=None): when lineage is unset, every existing test keeps its exact fixture, with consumed_inputs: None. LineageFixture passes CollectedParent(...).receipts. That callable saves the real collect parent and returns its receipt, which becomes the report's consumed_inputs just before the report is persisted. An earlier revision used an overridable lineage() method called from HandoffFixture.__init__; CodeQL flagged that, see (m).

  • (f) Private imports in the tests. The tests import the private _build_task_result (runner) and _verify_payload (store) so they use the real contract and real verification.

  • (g) The general-path missing-ancestor comparison. It removes both the parent's task row and its stored result. That way only the lineage walk, not the row loop, can reach the missing parent.

  • (h) The new edited-definition test. It mirrors all 4 variants of the existing test: definition edit, enable toggle, alert edit and run-as change.

  • (i) What the strong control pins. It pins the walk's run, identity, reference, reader_user_id and manifest arguments, but not include_sources, because M3 is equivalent.

  • (j) The E2E probe wasn't committed. It was a temporary test module, functional_tests/test_zz_tmp_probe_handoff_lineage_e2e.py. It re-ran the E2E test with an autouse fixture that counted node-result loads and page reads and wrapped the lineage walk. A walk mode kept the real walk, and a nowalk mode replaced it with a no-op as an M1 stand-in. I deleted it, with its bytecode, after each use.

  • (k) What "script mode" means. It's python -u <file>. Both four-mode files' __main__ runs pytest.main([__file__, "-q"]), which is the files' own runner and is unchanged. That means script mode runs without -p no:langsmith_plugin.

  • (l) Where the walk's cost is documented. It's in the feature doc's Performance section, since the Known limitations section doesn't cover lineage.

  • (m) Fixing this PR's own CodeQL alert by rewriting commit 2. CodeQL's first analysis of 27d5b5939 reported one new py/init-calls-subclass warning: HandoffFixture.__init__ called self.lineage(...), which LineageFixture overrides.

    • The fix. I removed the pattern instead of dismissing the alert or suppressing the rule. The hook became the lineage= parameter from (e), and the parent's construction moved to a small CollectedParent class. The two tamper callbacks now receive (receipt, parent, save) instead of reaching into a half-built fixture.

    • Why commit 2. I folded the fix in with git commit --fixup and git -c rerere.enabled=false rebase --autosquash, so the VERSION commit stays last as the kickoff requires. Then I pushed with --force-with-lease pinned to 27d5b5939. The old head is still in the local reflog.

    • Equivalence proof. A session script loaded the old and new fixture modules side by side and built six variants:

      • plain hand-off;
      • plain hand-off, partial run;
      • lineage;
      • lineage, partial run;
      • lineage with the receipt-side tamper;
      • lineage with the parent-side tamper.

      For each it compared stored results, manifests, references, receipts, run rows, loads and selectors, plus descriptor-only and excerpt reads. All six were identical under python and python -O.

    • Rerun. Every validation job, the four modes, the mutations, the repro, the guardrails, the docs checks and the E2E probe reran at c07b91151 (see Testing / validation, Kill table and E2E cost).

  • (n) The builder test's Version header. Commit 4 changes only the schema pin and its comment, as the review asked. The file's Version: 0.261.250 header is unchanged, because the review didn't ask for it and no test in the file changed.

  • (o) Landing. The user asked to submit and merge this PR, so once CI on 6d45bb8a7 passed I marked it ready and merged it. I used a merge commit, as Phase 6b-2: show chat-started workflow runs and their posted results in V2 #1639 and Phase 7a: hand off large chat requests to a one-time workflow (server, 0.261.238) #1640 did, pinned with --match-head-commit. I didn't use --admin or --delete-branch. V2's .251 sits right below this PR's .252, so no renumber was needed.

Kill table

Each mutation was applied to functions_workflow_result_reader.py at the review head c07b91151. The reader file then ran under pytest (<py> -u -m pytest functional_tests/test_workflow_handoff_result_reader.py -p no:langsmith_plugin -p no:cacheprovider -q -rf) with nothing deselected, and the file was restored. Afterwards the restored file was byte-identical and git diff --quiet -- application/single_app exited 0.

The same four mutations first ran at 27d5b5939, before the CodeQL fixture fix. Both runs gave the same counts and the same killing test IDs. The reader, its test file and the rest of application/single_app are identical at 83286787e and at the merge 6d45bb8a7.

Mutation Change Result Tests that killed it
none baseline 81 passed n/a
M1 delete the new authorizer call Killed: 7 failed, 74 passed malformed receipt (test 1); strong control's exact loads (test 2); missing parent (test 3); output disagrees [receipt] and [parent], and bytes mismatch (test 4); descriptor-only lineage read (test 6)
M2 walk only when include_excerpts is true Killed: 6 failed, 75 passed M1's set minus the strong control: the descriptor-only lineage read, plus the descriptor-only halves of the malformed, missing-parent, output-disagrees ×2 and bytes tests
M3 include_sources=True Survived: 81 passed none. Equivalent: include_sources only decides whether the walk keeps each source snapshot and returns them in access()["sources"]. Every source is still snapshotted and counted (source_seen), and the reader discards access(), so the reader's result, loads and failures can't differ. False just avoids keeping the list.
M4 _guarded("manifest", ...) instead of "authorize" Killed: 5 failed, 76 passed malformed, missing parent, bytes, and output disagrees ×2. Both stages map the same exceptions to the same codes, so these are killed by the logged closed stage, which these tests assert is authorize.

These are the killing test IDs, all in functional_tests/test_workflow_handoff_result_reader.py:

  • M1: test_a_descriptor_alone_walks_the_lineage_but_reads_no_section, test_a_malformed_consumed_input_receipt_is_refused_before_the_report_is_read, test_a_missing_parent_is_refused_with_the_general_paths_code, test_a_parent_whose_bytes_do_not_match_its_hash_is_refused, test_a_parent_whose_output_disagrees_with_the_receipt_is_refused[parent], [receipt], and test_a_report_with_real_lineage_reads_after_walking_its_parent
  • M2: the same set as M1, without test_a_report_with_real_lineage_reads_after_walking_its_parent
  • M4: test_a_malformed_consumed_input_receipt_is_refused_before_the_report_is_read, test_a_missing_parent_is_refused_with_the_general_paths_code, test_a_parent_whose_bytes_do_not_match_its_hash_is_refused, and test_a_parent_whose_output_disagrees_with_the_receipt_is_refused[parent] and [receipt]

E2E cost

This was measured in functional_tests/test_workflow_handoff_end_to_end.py, test_an_accepted_handoff_runs_once_and_posts_its_summary_to_chat, at both sizes, using the file's in-memory canonical store. The figures below are from the review head c07b91151. application/single_app and the E2E file are identical at 83286787e and at the merge 6d45bb8a7, so they still apply.

  • The probe was a temporary test module, functional_tests/test_zz_tmp_probe_handoff_lineage_e2e.py. It re-collected that E2E test and added an autouse fixture that counted every node-result load and page read during the hand-off read. The fixture also wrapped authorize_workflow_node_result_read to count calls and outcomes. The module was deleted after each use, and git status --porcelain was empty afterwards. It was never committed.
  • The "no walk" rows replace the authorizer with a no-op, which stands in for M1.
  • The file's reauthorize stub is for input reauthorization, not this walk. The walk below ran for real: one call per read, and it succeeded every time.
Mode Documents Node-result loads Page reads Read time Walk time Walk calls (ok)
walk, normal 3 27 0 0.0106 s 0.0098 s 1 (1)
walk, normal 200 1,409 0 1.6724 s 1.6709 s 1 (1)
walk, -O 3 27 0 0.0104 s 0.0096 s 1 (1)
walk, -O 200 1,409 0 1.6645 s 1.6632 s 1 (1)
no walk, normal and -O 3 and 200 2 0 0.0007 to 0.0009 s n/a n/a

First measurement. The probe first ran at the original fix commit, 39655e906, before the CodeQL fixture fix rewrote it. Apart from VERSION in config.py, the application code and the E2E file are identical at c07b91151. That run gave the same load and page counts in every row, and 200-document read times of 1.8534 s (normal) and 1.8789 s (-O). Only wall time varies between runs.

The 200-document loads break down as follows. They come to about 7 point reads per document, the same cost model as the general path's per-row walk.

What's loaded Loads
report manifest 2
collect manifest 1
each manifest 1
per-document body-node manifests 200
sections 1,205

The report manifest loads twice because, by the end of a 200-document walk, it has been evicted from both the authorizer's LRU (32) and the reader's _ManifestMemo (64); see Follow-ups. The worker's proof cache applies only inside a worker execution, so it doesn't help this read.

Observable. The E2E file still passes without the walk, because it doesn't assert loads. The probe's load counts are the observable that M1 changes: 27 against 2 for 3 documents, and 1,409 against 2 for 200. The unmodified E2E file passes at HEAD in both pytest modes; see Hand-off set.

Probe commands, run from the worktree root at c07b911

Each command ran with the environment above, plus PROBE_MODE and PROBE_OUT=<session files>\probe_e2e_final.jsonl. They ran one at a time in this order:

  1. PROBE_MODE=walk: <py> -u -m pytest functional_tests/test_zz_tmp_probe_handoff_lineage_e2e.py -p no:langsmith_plugin -p no:cacheprovider -q -rfE: exit 0
  2. PROBE_MODE=nowalk: the same command: exit 0
  3. PROBE_MODE=walk: <py> -O -u -m pytest functional_tests/test_zz_tmp_probe_handoff_lineage_e2e.py -p no:langsmith_plugin -p no:cacheprovider -q -rfE: exit 0
  4. PROBE_MODE=nowalk: the same -O command: exit 0

CodeQL

I read CodeQL with gh api "repos/microsoft/simplechat/commits/<FULL SHA>/check-runs?check_name=CodeQL" and then gh api repos/microsoft/simplechat/check-runs/<id>/annotations, once for each head.

Head CodeQL check run Conclusion Title Annotations
27d5b5939af137e9e8acbe2e95867b544bf90758 (first head) 112102042747 success 1 new alert 1
c07b91151bab479e767e74d10398ea59803abb2b (review head) 112107188549 success No new alerts in code changed by this pull request 0
83286787ecd6707ae2e60b668feaa426ab0e2500 (pin head) 112137438992 success No new alerts in code changed by this pull request 0
6d45bb8a7419c2abfc7fdd4f4122cfe002cae3e5 (final head, V2 merge) 112251009769 success No new alerts in code changed by this pull request 0

Alerts at the final head: none.

The alert on the first head:

  • Annotation: a warning, "__init__ method calls overridden method" (py/init-calls-subclass), at functional_tests/test_workflow_handoff_result_reader.py L166.
  • Message: "This call to HandoffFixture.lineage in an initialization method is overridden by LineageFixture.lineage."
  • Scope: it was in this PR's own test fixture. No application code was involved.
  • Fix: I fixed it by removing the pattern, not by dismissing the alert or suppressing the rule. See decision (m). After the fix, the re-analysis of c07b91151 reported no new alerts.

Other notes:

  • Timing. At 83286787e and again at 6d45bb8a7, the CodeQL check first showed neutral ("configuration not found") while the Analyze jobs were still running. The same check run changed to success when they finished.
  • Runner notices. At every head, Analyze (python) succeeded (112250769245 at 6d45bb8a7, 112137073045 at 83286787e, 112107020803 at c07b91151). Its only annotation is a runner-image notice that the ubuntu-latest label is migrating, at path .github, level notice. It isn't a code alert.
  • malicious-pr-security-review succeeded with 11 annotations, identical at 83286787e and 6d45bb8a7: the runner notice, plus 10 keyword-marker warnings on changed lines of the fix doc (docs/explanation/fixes/WORKFLOW_HANDOFF_RESULT_LINEAGE_FIX.md). They're heuristic text markers, not findings, and the doc has no code.
  • The code-scanning alerts API (repos/microsoft/simplechat/code-scanning/alerts?pr=1647) returns 403 for this token. It needs a scope I may not request, so the check runs' output and annotations above are the record.
  • Nothing was dismissed.

Every other check on 6d45bb8a7 succeeded, as it did on 83286787e and c07b91151:

  • license/cla
  • Analyze (actions)
  • Analyze (javascript-typescript)
  • Analyze (python)
  • xss-sink-check
  • swagger-route-check
  • malicious-pr-security-review
  • syntax-check
  • enforce-branch-flow
  • broken-access-control-check

Follow-ups

  • Per-read walk cost. Each hand-off read now makes about 7 point reads per document: 1,409 loads for a 200-document report. That's the same cost model as the general path's per-row walk, but it's paid on every read, including stored-context re-checks. A per-run or per-result proof cache keyed by the report's reference would make repeat reads cheap. It's out of scope here.
  • Double root-manifest load. At 200 items the report manifest is loaded twice, because by the end of the walk it has been evicted from both the authorizer's LRU (32) and the reader's _ManifestMemo (64). This is harmless but wasteful; pinning the root would fix it.
  • No check against declared inputs. The authorizer doesn't check consumed inputs against the compiled flow's declared inputs. That's by design, and it's unchanged here.
  • The E2E doesn't assert load counts. It passes with or without the walk. The reader unit tests now pin the walk, and an E2E load assertion could pin it end to end too.
  • Off-golden regeneration policy now that the feature is in V2. The coordinator is tracking this, and the test's regeneration instructions are unchanged.
  • Pre-existing V2 docs failures. test_docs_release_notes_integrity.py and test_docs_link_integrity.py fail at V2 6bd12cf and at the merge, the same way (see Reruns at the merge). Neither is on this PR's validation list, and this PR doesn't fix them.
  • Resolved: stale builder digest. test_workflow_handoff_builder.py::test_the_phase_4_blueprint_and_payloads_are_unchanged failed at 15feec6 and at c07b91151, in pytest and pytest -O, alone and in the harness process.

Documentation

  • Release notes updated, or not needed. There's a new ### **(v0.261.252)** section with two #### Bug Fixes entries: the lineage re-proof and the recaptured off-golden. It's a pure 15-line insertion, so every existing section is byte-exact. After the merge it's still +15/−0 against V2 6bd12cf.
  • Feature documentation updated, or not needed. In CHAT_ORCHESTRATION_WORKFLOW_HANDOFF.md, "Result reading and delivery" now describes the identity check and the lineage re-proof, which apply to descriptor-only and excerpt reads. The Files table row, the test-table row and Performance are updated too.
  • Fix documentation updated, or not needed. New docs/explanation/fixes/WORKFLOW_HANDOFF_RESULT_LINEAGE_FIX.md covers the issue, root cause, files, tests, before and after, and the stale off-golden and its verified recapture. Since review, it also covers the Phase 4 digest pin.

Security checklist

  • New Flask routes include @swagger_route(security=get_auth_security()). N/A: no routes are added or changed.
  • Settings sent to non-admin frontends use sanitize_settings_for_user(). N/A: no settings are sent.
  • Browser JavaScript is served from local SimpleChat static assets only; no CDN-hosted JS. N/A: no JavaScript or templates are changed.
  • No secrets, keys, connection strings, or local-only artifacts are included.

Capture using --write-golden in a fresh detached worktree with no production changes. Copy generated JSON verbatim. Changes reflect upstream merge capabilities, export profiles, composition mapping and planner/workflow guidance; Phase 4/5 route snapshots are unchanged. Existing V2 goldens are untouched. Hand-off off-golden suite: 10 passed under both normal and optimized pytest, exit 0.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
(cherry picked from commit 5a625599c623d17633427ac363966438ba793582)
Comment thread functional_tests/test_workflow_handoff_result_reader.py Fixed
Paul Lizer (paullizer) and others added 2 commits October 6, 2026 00:21
The hand-off branch of read_workflow_result loaded the report manifest by its exact node
selectors and checked its identity, text output and completion, but never walked its lineage.
A report whose consumed-input receipts were malformed, or named a missing or corrupt parent,
still read as available, although the general path refuses the same lineage.

_read_handoff_result now calls the shared node lineage authorizer
(authorize_workflow_node_result_read) on the saved workflow, seeded with the loaded manifest and
the same memoized loader, before the result is described or excerpted. Sources aren't
re-resolved. read_workflow_result passes through the user id it has already proved.

The reader tests add a strong valid control whose report consumed a real collect records result,
built through the runner's result contract, and cover a malformed receipt (the repro), a missing
parent, a parent whose output disagrees with the receipt, a parent whose bytes don't match its
hash, and edited or re-enabled definitions with lineage. A descriptor-only read walks the
lineage but reads no section.

Refs microsoft#1549, microsoft#1543. Follow-up to microsoft#1640.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Add the Workflow Hand-off Result Lineage Fix doc, covering the reader's
missing lineage walk, the change and its failure codes, the recaptured
off-golden, validation and limitations.

In the hand-off feature doc, the reader now re-proves the report's lineage
with the shared node lineage authorizer, an edited or re-enabled workflow
still fails closed with workflow_result_invalid, and the test, file and
performance notes match.

Add a 0.261.252 release-notes section with two bug fixes. Existing
sections are unchanged, and the docs inventory regenerates unchanged.

Part of microsoft#1549. Part of microsoft#1543. Follow-up to microsoft#1640.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@paullizer
Paul Lizer (paullizer) force-pushed the paullizer-7a-follow-up-handoff-lineage branch from 27d5b59 to c07b911 Compare October 6, 2026 04:24
Paul Lizer (paullizer) and others added 2 commits October 6, 2026 02:19
test_the_phase_4_blueprint_and_payloads_are_unchanged failed at V2
15feec6 on the schema digest alone. Its pin (ab12fb1d) predated microsoft#1641,
which added the merge task schema and was merged into Phase 7a at
b8e4730 without a refresh.

The pin now holds 520139cb, the value V2 computes at a5a5b1c, after microsoft#1641
and before hand-off merged, so the test still proves hand-off left the
Phase 4 schema unchanged. The payload, validation and dry-run pins are
unchanged.

The fix doc adds the builder test to Files modified, records the pin under
Validation, and gives the 200-document read time as about 1.7 seconds, to
match the c07b911 timings.

Refs microsoft#1549, microsoft#1543. Follow-up to microsoft#1640.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@paullizer
Paul Lizer (paullizer) force-pushed the paullizer-7a-follow-up-handoff-lineage branch from c07b911 to 8328678 Compare October 6, 2026 06:21
Bring in origin/paullizer-react-v2-ui at 6bd12cf, the merge of microsoft#1639
(0.261.251), so microsoft#1647 lands second as agreed.

Only two files conflicted:
- config.py keeps VERSION = "0.261.252", one above V2.
- release_notes.md keeps the v0.261.252 section directly above V2's
  v0.261.251 section, with every section byte-exact.

Every other file matches the side that changed it.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@paullizer
Paul Lizer (paullizer) marked this pull request as ready for review October 6, 2026 12:01
@paullizer
Paul Lizer (paullizer) merged commit 16f42e4 into microsoft:paullizer-react-v2-ui Oct 6, 2026
11 checks passed
Paul Lizer (paullizer) added a commit that referenced this pull request Oct 6, 2026
…w-up

- Status date 2026-10-06, with paullizer-react-v2-ui at 0.261.252.
- Phase 6: 6b-2 done in #1639 (v0.261.251), with how it shipped.
- Phase 7: 7a done in #1640 (v0.261.250) and its lineage follow-up in #1647
  (v0.261.252); 7b, the V2 hand-off card, is next. Admins should leave the
  hand-off setting off until 7b ships.
- Section 9: the run deep-link and one-time workflow decisions are settled.
- Release-notes integrity follow-up updated to 187 missing releases.

Docs only; no version bump.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Paul Lizer (paullizer) added a commit that referenced this pull request Oct 8, 2026
…lows-capability

Update roadmap status: 6b-2 and 7a done, with the #1647 lineage follow-up
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants