Repository navigation
Phase 6b-2: show chat-started workflow runs and their posted results in V2 - #1639
Merged
Paul Lizer (paullizer) merged 34 commits intoOct 6, 2026
Conversation
The status parser, the tab's one tracker engine and its store and hook, the run-level cancel and durable-resume clients, the delivered-message helpers, the scope-aware run deep link, and the bell's workflow_chat_delivery case. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…run summary - WorkflowRunCard replaces the Started workflows links when chat workflow runs are on: phase-driven states, Cancel with confirmation, Retry through the durable runtime resume, Review and approve via the run inspector, Reconnect Microsoft 365, Open run, and Check now with a polite Checked time. - WorkflowDeliveryFooter on messages a workflow run posted: Follow up (puts the run's result in this chat's composer), Retry workflow run only when a fresh status row allows it, and Open run. - Plain chat Retry is hidden on delivered messages (the server refuses it). - WorkflowRunningTag next to the generating-images tag, built from tracker state only. - WorkflowProposalRunSummary on a created recurring workflow: next run and last run status, Open latest results and Follow up. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Pull the delivered-message helpers out of the shell hook into pure functions (workflowDelivery.ts) so the landing plan, the busy check, the reply shape and the footer's Retry gate can be tested directly. Give the tracker a per-chat first-read baseline, so a chat read never moves the global one, and make v2WorkflowRunPath refuse scope types it doesn't know. Add node tests for the status parser, the tracker engine, the run-link routing, the run action clients and the delivered-message helpers, with shared fixtures in the status route's exact shape. Server facts the client mirrors are pinned against the modules that write them. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…r-phase-6b-2-v2-workflow-run-card-and-trac
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
- Run-card harness (ui_tests/fixtures/workflow_run_tracking) mounting the real stores, tracker, card, footer and running tag against a mocked API. - test_v2_workflow_run_card.py: phases, actions, fail-closed states, delivery settle in the open and other chats, baseline, footers, tag, XSS. - test_v2_workflow_run_tracker_spa.py: one tracker per tab in the built SPA, flags-off zero requests, Check now, reload never re-announces. - Version headers moved to 0.261.233. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
- Bell: a workflow_chat_delivery notice reads as "Workflow results" and opens its run in V2, with or without a link; a notice without its workspace gets no link, and a Microsoft 365 notice keeps its classic page. - Proposal card: a created card shows its next run in the reader's time zone, the newest run's status, Open latest results and Follow up, never guesses an unknown status, honours the workflow and results flags, and says run details are unavailable when a read fails. Refs microsoft#1546 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The new run card, delivery footer, running tag, recurring run lines, tracker, store and helpers pass check_xss_sinks.py in full, as do the changed files that had no older findings. No changed V2 file writes raw HTML, every run-tracking link is workflowRunHref over a run's ids or the fixed Microsoft 365 path, and status rows carry ids, never a URL. Synthetic snippets show the checker still flags a link or HTML taken from a status row. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The fake status route answers only when a test says so, so a change that sends a request a test doesn't expect left the test waiting with no end. Each test now has a 10-second limit, which turns that into a failure. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Adds the V2 section to the 6b-1 feature doc: when V2 tracks runs, the run card's states and actions, Check now, the tracker's cadence and baseline, how a posted result lands, the running tag, the posted-message footer, the recurring-workflow card, run links and V2 limitations. Splits File structure and Testing into Server (6b-1) and V2 (6b-2). Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
- chat-controls: next run, last run, open latest results and follow up on the recurring-workflow card; live status, check now, cancel run, retry, review and approve, reconnect Microsoft 365, results posted below and the running tag on started workflows; a new "Results posted to the chat" section. - trigger-a-workflow: live status card, cancel and retry from chat, posted results with follow up, needs-you waits, two troubleshooting rows. - Phase 5 and 6a feature docs point to the 6b-2 V2 experience. Refs microsoft#1546 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
New Features: live run status under a plan's answer, one tracker per browser tab, Follow up / Retry / Open run on posted results, next and last run on the recurring-workflow card, and workflow notices that open the run in V2. UI enhancements: the running tag in the chat list and the Workflow results bell label and alert-card Open run. Refs microsoft#1546 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Brings in microsoft#1626 (orchestration file render permission fix, released as 0.261.233) and microsoft#1633 (roadmap updates). The only conflict was the top of release_notes.md, where both sides added a v0.261.233 section; both are kept intact under their own headings. 6b-2 is renumbered to 0.261.234 in the next commit. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
microsoft#1626 released 0.261.233 first, so 6b-2 moves to the next patch version: config.py, its release notes heading, and every 6b-2 version reference in its feature docs, test headers and version assertions. microsoft#1626's own 0.261.233 references are unchanged. Refs microsoft#1546 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Cancel and Retry on the run card already wait while another action on the same run is under way, but Review and approve did not. A read that landed mid-Retry and found the run waiting for approval rendered a live link to the gate while the resume was still in flight. The link now carries aria-disabled and swallows click and Enter while any action on that run is pending, exactly like Cancel and Retry, and comes back once the action is answered and the chat's runs have been re-read. The new UI test holds the resume, lands a waiting-for-approval read with Check now, and checks the link is held for both a click and Enter, then opens the run once the resume is answered. Mutations M29a (busy always false), M29b (no preventDefault) and M29c (no aria-disabled) are all killed. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
… card Adds three ad-hoc media slots to docs/reference/chat-controls.md, following the page's existing media.html pattern: the Started workflows run card with the chat-list running tag, a result posted to the chat with Follow up and Open run, and the created proposal card's Next run and Last run. No images are committed; the slots render placeholders until someone captures them. Docs-only, so the version is unchanged. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Both run tracking XSS guardrail files ran their tests from a plain loop, so python -O removed every assert and the optimized script run passed even with a planted HTML sink. Their __main__ now runs pytest.main, which rewrites the asserts into explicit checks, the same runner as the V2 run card and tracker UI suites. A planted regression now fails in all four modes. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The card suite's posted-message helper now opens More actions on the posted message and on the question in the same chat. It proves the menu opened, then checks that only the question offers Edit. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CodeQL flagged two py/overwritten-inherited-attribute warnings on microsoft#1639: RunHarness reassigned `conversations` after NotificationApi set it, and RecurringApi reassigned `proposal` after ProposalApi set it. Following the query's advice, NotificationApi and the bell Harness now take an optional `titles` map and ProposalApi an optional starting record, which the subclasses pass instead of overwriting. The constructed state is unchanged. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…during the delivery re-read - WorkflowRunCard appends the tracker's runs for the answer that the Phase 5 list doesn't name, oldest request first, once the list has loaded or failed. A run or step the list already names is left out, so a step the list says can't open stays closed. A plan run that's gone (404) still shows nothing. - reloadMessages takes an opt-in onlyIfUnchanged guard, used only by the delivery re-read: it is dropped if a reply started or the messages changed while it was out, and its results wait for the next quiet moment. Every other caller behaves as before. - Tracker test for the seen-undelivered shortcut, a request-count check for the 10 s dedupe window, and card tests for both fixes. The card fixture removes the bundle's CSS sidecar on teardown. - CHAT_WORKFLOW_RESULT_DELIVERY.md describes both fixes and the new tests. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
V2 is still at 0.261.233. microsoft#1635, microsoft#1636, microsoft#1637 and microsoft#1638 claim .234 to .236, so this branch takes .237. Every 0.261.234 reference moves to 0.261.237: config.py, the release notes header, the feature docs, the chat-controls reference and the test headers. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
14 of 24 tasks
Brings in microsoft#1635 (clearer saved-result error messages and repaired workflow test harnesses, released as 0.261.234). The only conflicts were VERSION in config.py, kept at 0.261.237 because that is still above V2's 0.261.234, and the top of release_notes.md, where both sides added a section; both are kept intact under their own headings, v0.261.237 above v0.261.234. Refs microsoft#1546 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
microsoft#1635 split _workflow_runtime_response's PermissionError branch: a saved record that fails its own check (AnalysisResultUnavailable) now answers its own 403 text before the access 403. The action-client suite pinned the old single-return shape, so it failed after the V2 merge. The pin now reads the outer branch and requires every return in it to be a 403, which V2 shows as its fixed no-access sentence whatever the server says. Refs microsoft#1546 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Brings in microsoft#1636 (V2 workflow alerts: require acknowledgment, repeating sound, size options and team delivery, released as 0.261.235). Three conflicts: - config.py: VERSION kept at 0.261.237, which is still above V2's 0.261.235. - release_notes.md: both sides added a section at the top. Both are kept intact under their own headings, v0.261.237 above v0.261.235, so against V2 the file only gains 6b-2's section. - ui_tests/test_v2_workflow_alert_notices.py: the last line of test_open_run_goes_to_the_run_in_its_workspace keeps 6b-2's read_calls == ["o1", "o2", "o4"], and every must-acknowledge test microsoft#1636 appended after it is kept unchanged. Refs microsoft#1546 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
microsoft#1636 changed BootstrapPayload.features to BootstrapFeatures, which allows undefined values, so App.tsx and MessageList.tsx no longer typechecked against workflowRunTrackerShouldRun's Record<string, boolean> parameter after the merge. The gate already compares each flag with === true, so only the parameter type widens; the tracker test now also covers an explicit undefined flag. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Brings in V2 25470e7 (0.261.236, the V2 sidebar conversation scroll fix). Conflicts resolved: - config.py: keep 0.261.237. - ConversationRail.tsx: keep the WorkflowRunningTag import and mount with V2's widened FocusEvent/ReactNode/RefObject type import. - V2_WORKFLOW_ALERT_NOTICES.md: keep both since-version lines in order; V2's 60-test row with Open run wording; Open run provenance row. - test_v2_workflow_alert_notices.py: Version 0.261.237 with every history line from both sides. - release_notes.md: 0.261.237 section above V2's 0.261.236 section. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Brings in V2 4f3eacb (0.261.237 orchestration settings document type fix and 0.261.238 Microsoft 365 actions in orchestration plans). Conflicts resolved: - config.py: keep this branch's version; it is renumbered above V2's 0.261.238 in the next commit. - release_notes.md: this branch's section on top, then V2's 0.261.238 and 0.261.237 sections unchanged. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
V2 is 0.261.238 after microsoft#1644, which also used 0.261.237. Moves this branch's version, release-notes heading and every 0.261.237 line it adds to 0.261.239. V2's own 0.261.237 section and files are unchanged. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Paul Lizer (paullizer)
added a commit
that referenced
this pull request
Oct 5, 2026
V2 reached 0.261.238 when #1644 merged, so this branch moves one above it. Sets config.py VERSION to 0.261.239 and moves this branch's own 0.261.238 lines to 0.261.239: the hand-off modules' and tests' Version, Implemented in, MINIMUM_VERSION and assert_app_version_at_least literals, the two route-test coverage notes, the run adapter test's hand-off refusal note, the two hand-off rows in docs/admin/orchestration.md, the feature doc and this branch's release-notes header. #1644's own 0.261.238 lines (its schema docstring, its Microsoft 365 section and troubleshooting row in orchestration.md, its fix docs and its release-notes section) are unchanged. The open V2 PRs claim 0.261.234 (#1639), 0.261.235 (#1637), 0.261.236 (#1638) and 0.261.246 (#1641), none of them 0.261.239. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve both proposal-test histories and every V2 release-note section. The version bump follows separately. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Keep all base version history; update only 6b-2's version, tests and documentation after the V2 microsoft#1641 merge. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Take V2's config.py (0.261.250) and keep 6b-2's release-note section above V2's 0.261.250 section with every V2 byte preserved. The renumber to 0.261.251 follows separately. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Keep all base version history; update only 6b-2's version, tests and documentation after the V2 microsoft#1640 merge, which took 0.261.250: config.py moves from 0.261.250 and 51 lines in 21 other files move from 0.261.249, all to 0.261.251. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CodeQL flagged py/overwritten-inherited-attribute at ui_tests/test_v2_workflow_run_card.py L252: RunHarness.__init__ assigned self.streams = 100 after NotificationApi.__init__ had already set it to 0. NotificationApi and the bell Harness now take a keyword-only streams argument (default 0), and RunHarness passes streams=100 to super().__init__. Every harness the two suites build has the same attributes, values and attribute order as before; the bell suite's own harness still starts at 0. Test-only change; VERSION stays 0.261.251. Refs microsoft#1546, microsoft#1543 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
14 of 24 tasks
Paul Lizer (paullizer)
marked this pull request as ready for review
October 6, 2026 11:31
Paul Lizer (paullizer)
merged commit Oct 6, 2026
6bd12cf
into
microsoft:paullizer-react-v2-ui
11 checks passed
Paul Lizer (paullizer)
added a commit
to paullizer/simplechat
that referenced
this pull request
Oct 6, 2026
Bring in origin/paullizer-react-v2-ui at 6bd12cf, the merge of microsoft#1639 (0.261.251), so microsoft#1647 lands second as agreed. Only two files conflicted: - config.py keeps VERSION = "0.261.252", one above V2. - release_notes.md keeps the v0.261.252 section directly above V2's v0.261.251 section, with every section byte-exact. Every other file matches the side that changed it. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Paul Lizer (paullizer)
added a commit
that referenced
this pull request
Oct 6, 2026
…w-up - Status date 2026-10-06, with paullizer-react-v2-ui at 0.261.252. - Phase 6: 6b-2 done in #1639 (v0.261.251), with how it shipped. - Phase 7: 7a done in #1640 (v0.261.250) and its lineage follow-up in #1647 (v0.261.252); 7b, the V2 hand-off card, is next. Admins should leave the hand-off setting off until 7b ships. - Section 9: the run deep-link and one-time workflow decisions are settled. - Release-notes integrity follow-up updated to 187 missing releases. Docs only; no version bump. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
This was referenced Oct 6, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
config.py's version.Linked issue
Refs #1546
Refs #1543
paullizer-react-v2-ui. This PR builds only on its status route and itsmetadata.workflow_deliverycontract, and doesn't change its server code.Release Notes & Latest Features
Is this visible to end users?
Is this admin-facing (Admin Settings, governance, deployment, config)?
Should this become a Latest Feature card?
Screenshot needed for the card?
Version bump
application/single_app/config.pyVERSIONthird segment bumped, or not needed because this is docs-only (V2's0.261.250→0.261.251. The Phase 7a: hand off large chat requests to a one-time workflow (server, 0.261.238) #1640 merged5b82040btook V2'sconfig.pyas is; the separate commiteb090065bthen changes only 52 version lines in 22 files:config.pyfrom.250, and 51 lines in 21 files from.249, including theassert_app_version_at_leastcall. No later commit changes a version;bd1a2504ais a test-only CodeQL fix. V2's version history is untouched, and.252is reserved for the 7a follow-up.)deployers/version.txtbumped, or not needed becausedeployers/was not changed (not changed)Testing / validation
Head:
bd1a2504a5530f9e82939f803bff5c31ef239af2(.251). Recorded V2 base:15feec6509f0846238362838103be6e76d19f8c1(.250, #1640), merged ind5b82040band renumbered separately ineb090065b;bd1a2504ais a test-only CodeQL fix. Base comparisons use a detached worktree, never stash.What this head re-ran. #1640 changed no file under
application/v2_uiorui_tests, so the V2 sources are byte-identical fromdb2e6a4ed(the #1641 merge) to this head. Since90f87fef0, this PR's own lines changed only in version numbers, plusbd1a2504a's two-file harness fix. This round ran light checks only, each labeled with the commit it ran at, plus all 82 regression files and 8 boundary files at the exact head. The focused UI set is left to the coordinator's rerun.Environment: Windows, Python 3.12 (Flask 3.1.3, pytest 9.0.3, Playwright 1.58.0), Node 24;
PYTHONPATH=application/single_app;functional_tests. Heavy jobs ran serially.Review fixes:
483291d83appends unjoined tracked runs, guards the delivery re-read after its fetch, adds the C2 clock-skew test and M25 request-count assertion, and cleans the harness CSS sidecar.d7ea6502arepairs the source pin after #1635.f21dcb73eaccepts optional bootstrap feature values after #1636:=== truestill gates each flag; an explicit-undefined test preserves that behavior. Merge details are under Overlap and version.Build prerequisites: the card/SPA suites use
application/single_app/static/v2; orchestration-based suites use the CSS inui_tests/artifacts/orchestration-plan-editor. Both must match current sources. The card rejects a stale SPA and cleans its own harness bundle/sidecar. The exact old-CSS sidebar failure and its base comparison are under UI suites.Build and typecheck
bd1a2504a(18 s), and ateb090065b. The two TS2345 errors atd751471dewere fixed inf21dcb73e.db2e6a4ed, whose V2 sources are identical to this head's: M23's mutant compiled and was killed, then the clean bundle rebuilt in 33 s.eb090065b: 2,729 modules,index-BIMKrYWZ.css(97.45 kB), 23.85 s. Also at basea5a5b1c53(2,716 modules) for the sidebar comparison; both produced the same content-hashed CSS file.Node suites: 8 files, 105 tests
test_v2_workflow_run_tracker.mjs(new)test_v2_workflow_run_status.mjs(new)test_v2_workflow_run_action_clients.mjs(new)test_v2_workflow_delivery_messages.mjs(new)test_v2_workflow_run_link_routing.mjs(new)test_v2_orchestration_workflow_run_links.mjs(Phase 5)test_workflow_results_clients.mjstest_v2_inline_image_proposal_resume_logic.mjsAll 105 passed at the exact head
bd1a2504a, and ateb090065b. The coordinator independently reran all 105 at the earlier immutable merge commitdb2e6a4ed, and re-killed C2 with the new backwards-clock assertion.#1640 added code to two server files these suites pin. The action-client suite reads
decide_workflow_runtimefromfunctions_workflow_runtime.py(its L365–366), and the status suite readsroute_backend_orchestration.py(its L151) and checks the status route's@bp.routeline (its L257). Both still pass after it.The new suites block the network by replacing
globalThis.fetchwith a function that throws. The action-client suite records requests instead.After the #1635 merge, the action-client suite failed 1 of 14: it pins that the run-resume route answers a
PermissionErrorwith 403, and #1635 added a second 403 to that branch.d7ea6502arequires everyreturnin the branch to be a 403. Changing either 403, or the branch's exception, fails it.Python files added or touched: four ways
-Oscript-Opytestfunctional_tests/test_v2_workflow_run_tracking_xss_guardrail.pyeb090065bfunctional_tests/test_v2_orchestration_workflow_run_links_xss_guardrail.pyeb090065bfunctional_tests/test_xss_guardrails_checker.py(unchanged; the checker's own tests)eb090065bui_tests/test_v2_workflow_run_card.py90f87fef0ui_tests/test_v2_workflow_run_tracker_spa.py(imports the run-card and bell modules)90f87fef0ui_tests/test_v2_orchestration_workflow_proposal_card.py90f87fef0ui_tests/test_v2_workflow_ask_ai_proposal.py(uses the proposal card's fixtures)__main____main__90f87fef0ui_tests/test_v2_notifications_bell.py__main____main__636d58242ui_tests/test_v2_workflow_alert_notices.py(this PR changes two tests for Open run)__main____main__636d58242ui_tests/test_v2_document_provenance.py__main____main__636d58242ui_tests/test_v2_workflow_alerts.py(reacheschatStorethroughAppShell→Sidebar)__main____main__636d58242,-Opytest90f87fef0At the exact head
bd1a2504a(clean tree),pytest --collect-onlyexits 0 for all 11 files with the counts above.bd1a2504achanged the card's and bell's shared harness constructor, so their targeted runs at the exact head cover it: the card's 4 reply and re-read tests passed (37 deselected) and the bell's 13 reply and workflow tests passed (35 deselected). The coordinator's focused rerun should include the full card (41) and bell (48). Earlier, all 44 runs exited 0 together atf21dcb73e, after the #1636 merge.The
-Ogap closed inbd578281a. Phase 5's run-links guardrail and the first tracking guardrail looped over their checks under__main__, sopython -Ostripped everyassertand passed with a planted HTML sink. Both now callpytest.main, which rewrites the asserts, so a planted sink fails in all four modes;-Oruns print onePytestConfigWarning.UI suites: every one that loads changed code
Current sidebar check.
ui_tests/test_v2_sidebar_conversation_scroll.pypassed 9/9 atdb2e6a4ed, and the V2 sources haven't changed since. It needs #1643's[data-conversation-rail-header][data-stuck]rule in the orchestration CSS, so rebuildui_tests/artifacts/orchestration-plan-editorfrom current sources first. With the staleindex-BL2LkJyy.css, 4 fail at L541 (rgba(0, 0, 0, 0)instead ofrgb(255, 255, 255)light orrgb(16, 23, 40)dark):reading_down_the_list_scrolls_the_navigation_away_and_holds_search,searching_while_held_keeps_the_search_box_in_placeandthe_held_search_is_drawn_on_the_themes_solid_surface[light]/[dark]. Basea5a5b1c53behaves the same: 4 failed and 5 passed with the stale CSS, 9 passed fresh;636d58242showed the same 4 with stale CSS. The functionaltest_v2_sidebar_conversation_scroll.pypassed 8/8 in all four modes at636d58242, and 8 at the exact head.Historical sweep. 121 UI files ran at
155604921, after the review fixes and before the merges of #1635, #1636, #1643, #1644, #1641 and #1640: 2,797 passed, 58 failed, 61 errors, 4 skipped. Every failure and error was the same at basef1aeef13d. The files were every UI suite that loadschatStore(53, through the harness bundles orimage_editor.py's@fixture-chatalias) or the built SPA (83, 19 of them also among the 53), the 5 suites for the otherreloadMessagescallers (4 already counted) and 3 related workflow and Microsoft 365 suites. 28 files that importplaywright_connectionthroughui_tests/fixtures/v2_admin_settings.pyL29 don't collect from the repo root (same on V2), so they were re-run withui_tests/fixturesonPYTHONPATH. The 4 skips need a live environment.Failures and errors, each identical at base
f1aeef13d:155604921test_chat_saved_analysis.py61320e813the server keeps it readable (test_saved_analysis_routes.pyL85–90 expects 200). Classic fails too.test_v2_orchestration_*suites (failed/passed):pending_approval_resume4/1,composer5/1,drawer_views2/1,elicitation5/1,plan_card6/1,plan_editing4/1,run_hydration6/1_PAGEis set only inmain(), so it'sNoneunder pytest. As scripts at155604921: 39/39.test_chat_three_document_smoke.pyTimeoutError, in the classic and the V2 test.test_v2_orchestration_plan_editor_backend.pyfunctions_settingsstub has noenabled_required.test_admin_custom_connections.py6/3,test_admin_shared_ai_connections.py5/38,…_classic.py8/37 (failed/passed)GET /api/models/catalog, whichstatic/js/admin/model_catalog_ui.js(1880a9d0e) requests. The errors are teardown errors, 27 per shared-AI suite. The classic suite wasn't debugged separately.test_v2_content_screening.pyRegression: 82 files that read a changed file
Each
.pyfile ran underpytest <file> -q -p no:cacheproviderand each.mjsor.jsfile undernode --test, one at a time with a 1,800 s limit each.bd1a2504a(clean tree before and after): the 82 (72 Python; 10 Node, 5.mjsand 5.js), plus 8 boundary files (5 Python, 3 Node) that read V2 files Merge CSV, Excel, PDF, Word and PowerPoint files in V2 chat and workflows #1641 or V2: Scroll the chat rail as one panel and pop workflow alerts from the bell #1643 changed. 90 files: 2,370 passed (2,192 pytest, 178 Node), 9 failed, one per file. The 82 alone: 2,317 passed (2,164 pytest, 153 Node), 8 failed.15feec650fail the same test ids with the same pass counts. The ninth is in a boundary file:test_v2_stats_parity.py::test_charts_use_the_shared_vendored_runtime("InlineChart must use the shared loader"). NeitherInlineChart.tsxnor any vendored asset is in this diff.functional_tests/test_v2_*):brand_mark_home_link,public_workspace_labels_pin,sidebar_account_menuandstats_parityreadSidebar.tsx, andsidebar_conversation_scrollreadsSidebar.tsxandConversationRail.tsx, all changed by V2: Scroll the chat rail as one panel and pop workflow alerts from the bell #1643.orchestration_merge_arguments.mjs,workflow_merge_task.mjsandworkflow_proposal_merge.mjsreadorchestrationMerge,workflowEditorandworkflowProposals, changed by Merge CSV, Excel, PDF, Word and PowerPoint files in V2 chat and workflows #1641.chatStore25,workflowEditor16,MessageList15,MessageActions10,ConversationRail7,App.tsx6, andnotificationLinks,workflowRunLinkandnotifications.ts1 each), plus 26 others: 6b-1's 13 delivery suites (383 tests), the kickoff's reruns, the workflow authoring and change-tracking suites, and the Node files.c852d349b, an earlier renumber, the 82 had the same 8 failures, identical at basef1aeef13d. After Fixes 1–3, the 27 functional files namingchatStorere-ran at155604921: 302 passed, and the same 7 failed (the eighth file doesn't namechatStore).The 9 failures, identical at base
test_v2_agent_model_exclusivity.py::test_the_composer_wires_the_rule_into_the_toolbartest_v2_collaboration_visual_style_fix.py::test_the_personal_route_is_unchanged_in_behaviourtest_v2_conversation_deep_link.py::test_incoming_link_is_captured_before_any_effect_runstest_v2_inline_image_proposals.py::test_no_third_party_browser_assets_were_added(flags@xyflow/react, added by31e7d16e5)test_v2_message_inspector.py::test_message_metadata_shape_is_role_dependenttest_v2_model_identity_and_scope.py::test_model_picker_keys_on_selection_keytest_v2_research_voice.py::test_reasoning_levels_match_the_existing_clienttest_v2_stats_parity.py::test_charts_use_the_shared_vendored_runtime(boundary file)test_v2_tabular_parity.py::test_the_confirmation_precedes_the_sendDocs
eb090065b(bd1a2504achanges only twoui_testsfiles)15feec650features.ymlv0.261.251, and the kickoff says not to regenerate themThe 3 new outstanding screenshots are this PR's slots in
docs/reference/chat-controls.md:chat-controls-workflow-proposal-run-summary.png,chat-controls-workflow-run-card.pngandchat-controls-workflow-posted-result.png. The two broken links indocs/reference/chat-controls.md(CHAT_CONTEXT_PICKER,PROMPT_COMPOSER_CARD) are already at base.Guardrails
run_v2_guardrails.py <worktree> 15feec650 HEAD 6b2, run at the exact head over15feec650...bd1a2504a, reported "51 changed files; app .py 1, xss surface 26, route .py 1, v2 ts/tsx 25". The one application Python file isconfig.py.The same 775 findings came over
a5a5b1c53...90f87fef0(510 files compile) and ateb090065b;bd1a2504a's harness fix changes no count.The malicious-PR review's "Needs investigation" verdict comes from keyword heuristics, and I read every category: 368 boundary (the feature's vocabulary: workflow, run, chat, delivery, tracker), 196 security-control (Playwright
role=locators,expect(,aria-attributes, testasserts,PermissionErrorin the action-client pin), 91 obfuscation ("hidden" for tab visibility, fixedre.compilepatterns,importlib.utilloading the in-repo XSS checker), 54 external-connection (the test originhttps://simplechat.test, URL parsing in assertions), 37 dynamic (the guardrail test's deliberateUNSAFE_ROW_SNIPPET, the card test'sunlinkof its own CSS sidecar) and 29 secret (searchParams.get,credentials: 'same-origin'assertions,authorizationreason codes). None loosens a check or reads a credential. Most are in the card test (142), the tracker.mjs(124) andworkflowRunTracker.ts(59). Your run atee8383a1dhad 744. This round adds 31 net, all of these kinds: 25 from483291d83(Fixes 1–3, M25 and the CSS change, with their tests), and 6 security-control from the pin on L354–360 of the action-client test, where the one-line pin had 1 finding and its 7-line replacement has 7. Thef21dcb73etype fix adds none: neither itsworkflowRunTracker.tsL148 nor its tracker-test L196 is flagged, and the totals match the earlier run overefee70775...d7ea6502a. No dependency, lockfile or CI file changed.Documentation
### **(v0.261.251)**section on top with five New Features and two User Interface Enhancements; every older section is byte-identical)CHAT_WORKFLOW_RESULT_DELIVERY.mdgains a V2 experience (6b-2) section covering the card, tracker, baseline, landing, tag, posted messages, recurring card, links, limitations, code map and tests. Also updated:CHAT_ORCHESTRATION_WORKFLOW_RUNS.md,CHAT_WORKFLOW_RESULTS_FOLLOW_UP.md,V2_NOTIFICATIONS_BELL.md,V2_WORKFLOW_ALERT_NOTICES.md,docs/guides/trigger-a-workflow.mdanddocs/reference/chat-controls.md, which declares three screenshot slots. The roadmap doc is untouched.)Security checklist
@swagger_route(security=get_auth_security())(no route added or changed)sanitize_settings_for_user()(no settings path changed; the tracker reads the existing bootstrap feature flags)What changed, by file
51 files, +9,545/−189. V2 paths below are under
application/v2_ui/src/. The only server file isapplication/single_app/config.py(+1/−1,VERSION = "0.261.251").New V2 files
lib/workflowRunStatus.tslib/workflowRunTracker.tslib/useWorkflowRunTracker.tsstores/workflowRunTrackerStore.tslib/workflowRunActions.tslib/useWorkflowRunAction.tslib/useFocusFallback.tslib/workflowDelivery.tsmetadata.workflow_delivery, joins a posted message to its run, and decides Follow up and Retry.lib/m365Links.tscomponents/chat/WorkflowRunCard.tsxcomponents/chat/WorkflowDeliveryFooter.tsxcomponents/chat/WorkflowRunningTag.tsxcomponents/chat/WorkflowProposalRunSummary.tsxChanged V2 files
App.tsxuseWorkflowRunTracker(Boolean(data) && !error && workflowRunTrackerShouldRun(data?.features)), afteruseWorkflowAlertRuntime.components/chat/MessageList.tsxWorkflowRunLinkswhen tracking is on, under Phase 5's unchanged conditions, and the footer on posted assistant messages in the open personal chat when nothing is masked.components/chat/MessageActions.tsxcomponents/chat/ConversationRail.tsxcomponents/chat/WorkflowRunLinks.tsxRunLinkand moves the one-shot list read intouseWorkflowRunLinks, for the card to reuse. Its strings and the guardrail-asserted link line are unchanged.components/chat/WorkflowProposalCard.tsxM365_CONNECT_HREFand one render line.stores/chatStore.tssettleCompletedReply, which the streamed path now calls with the same values, so a posted result settles the same way.reloadMessagestakes an optional{onlyIfUnchanged}, passed only by the delivery re-read: the fetched list is then dropped, resolving'superseded', if a reply started or the messages changed while the read was out. Without it, it behaves as before and resolves'done'.lib/workflowEditor.tscancelScopedWorkflowRun(the run-level cancel) andnewWorkflowRequestId.lib/workflowRunLink.tsworkflowRunHref(workflowId, runId, scope = personal)builds the personal or group run link.lib/notificationLinks.tsv2WorkflowRunPathreturns the run link instead ofnull, and workflow notices open their run in V2.lib/workflowAlertNotices.tslib/notifications.tsworkflow_chat_deliverynotices Workflow results.The other 25 files are the 8 docs under Documentation and the 17 test files below.
Tests
Node/guardrail files are in
functional_tests/; browser files are inui_tests/. Counts and executed modes are above; these are the guarantees they pin.test_v2_workflow_run_status.mjstest_v2_workflow_run_tracker.mjstest_v2_workflow_run_action_clients.mjstest_v2_workflow_run_link_routing.mjstest_v2_workflow_delivery_messages.mjssource: 'workflow', quiet-chat delivery, opt-in post-fetch guard and unchanged other callers, stopped-tracker epoch, list/bell projections.test_v2_workflow_run_tracking_xss_guardrail.py,test_v2_orchestration_workflow_run_links_xss_guardrail.py-O.test_v2_workflow_run_card.pyRunHarnesspassesstreams=100to the base constructor.test_v2_workflow_run_tracker_spa.pytest_v2_orchestration_workflow_proposal_card.pyRecurringApiinitializes rather than overwrites its inherited proposal.test_v2_notifications_bell.pystreams), not overwritten.test_v2_workflow_alert_notices.pyread_calls == ["o1", "o2", "o4"]; all V2 must-acknowledge and bell-anchor cases retained.test_v2_document_provenance.pytest_support/workflowRunStatusFixtures.mjs,ui_tests/fixtures/workflow_run_tracking/Binding notes 1–15
.251ineb090065b, after all fixes and V2 merges through #1640; only the test-onlybd1a2504afollows it. No stale implementation version; historical V2 sections remain untouched.WorkflowProposalCard.tsxis sharedM365_CONNECT_HREFand one render line. The rest is inWorkflowProposalRunSummary.tsx.workflow_priority_alertandworkflow_chat_delivery; another type with a scope and ids gets no link. Microsoft 365 is tested throughmetadataandlink_context, in routing and the bell.delivered_at >= baselineCheckedAt15feec650(also the guardrails' base), the sidebar suite ata5a5b1c53, the historical sweep atf1aeef13d. Nothing stashed.workflowRunHrefandgroupWorkspacePathare already inTS_SAME_ORIGIN_URL_BUILDERS, and the new guardrail proves the checker still flags a link or HTML built from a status row.paullizer-react-v2-uifrom my fork (pushes tooriginreturn 403). Of the watched files, onlyWorkflowProposalCard.tsxchanged.git -c rerere.enabled=false merge:6bd2545a2,8bd8d269f,042af8cb7,3fd4e2840,d751471de,461cd63f9,a5c28d684,db2e6a4ed,d5b82040b. No PR branches merged.waiting_recoveryandskipped{expected_version, request_id}from a fresh runtime read, and nothing when that read fails. All eight resume 409 codes are tested with their real text; cancel covers 202, 204, both 404s, a 409 with or without a code, 403, 500, 503 and no answer. Every answer re-reads the chat's runs; see Contract re-checks.approval_stateisn't read.Decisions 1–8
WorkflowRunCardmounts in place ofWorkflowRunLinksunder Phase 5's conditions. Rows join onrun_id+orchestration_run_id+step_id; items not yet read or not joined keep Phase 5'sRunLink. Tracked runs for the answer that the list doesn't name, for example because its read failed, are appended oldest first byrequested_atas live rows, named as text from the status row. They wait until the list answers or fails, never duplicate a step the list names, and count toward the card's runs, so Check now and the chat's own read follow. A 404 still renders nothing.workflowRunTracker.ts, the zustandworkflowRunTrackerStore.tsand the app-shelluseWorkflowRunTracker.ts. Readers never poll;requestConversationRuns(id, {force})is deduped for 10 s.run_id:generationreloadMessages({ onlyIfUnchanged: true }): if a reply started or the messages changed while it was out, such as a question sent meanwhile, it shows nothing and its results wait for the next quiet moment. One still out when the tracker stops is dropped. The source chip keeps its!analysisContextChosencheck.fetchScopedWorkflowsandfetchScopedWorkflowRuns, gated onallow_user_workflows, and again only if the workflow or one of the two flags it reads (allow_user_workflows,enable_chat_workflow_results) changes.data-workflow-alert-open-run) replaces Open workflow at the same address. Open workflow remains for an alert with no run; one that can't be placed offers neither, as before.WorkflowAlertCard.tsxis unchanged.WorkflowRuntimePanelshows the gate's prompt and choices. Nothing is approved from chat.workflowRunHref(workflowId, runId, scope = personal). Card and footer rows are always personal; notices and N2 pass group scope.Decisions made autonomously
Each is reversible.
available: false: rows and Cancel stay; Retry is hidden, with "Starting workflows from chat is turned off, so Retry isn't available here."workflow. The notice title already carries the outcome, sodescribeType's signature is unchanged.elapsed_secondsas of the last check, anchored by "Checked 9:07 AM". It doesn't tick.runtime_version.message_idor a non-integer generation is recorded but not announced. The card says "Results were posted to this chat." with no jump.link_urlbut a scope and ids open the V2 run, for workflow types only, asnotificationLinks.ts's comment promised.aria-disabled, a busy style andpreventDefault, so focus and layout don't jump. If the focused control disappears, focus falls back to the run's row or the footer.M365_CONNECT_HREFmoved tolib/m365Links.ts, shared by both cards; Phase 5's guardrail builder asserts follow the scope-aware signature; screenshot slots are declared, not captured; the card harness has its own fixture directory; both run-tracking XSS guardrails run throughpytest.main; the tracker suite fails a hung test after 10 s; and the Edit-absent check has a positive control: the user's own question still offers Edit.Tracker cadence and baseline
Each tick is one global
GET /api/v2/orchestration/workflow-runs/status. A run is in flight while it's queued, running or waiting, or its delivery ispendingordelivering.WORKFLOW_RUN_POLL_DELAYS_MS).WORKFLOW_RUN_HIDDEN_DELAY_MS).WORKFLOW_RUN_ERROR_DELAYS_MS); recovers on the next good read. A per-chat read failing for another reason shows "Couldn't check the run status right now. Try again." on that card and doesn't halt.?conversation_id=read, deduped while out and for 10 s after a success (WORKFLOW_RUN_CONVERSATION_DEDUPE_MS). Check now forces it.Baseline. A chat's baseline is the first good read in the page session that covers it, global or the chat's own. Results already posted by then are recorded silently by
run_id:generation. A later result is announced once, only if it has an integer generation and anassistant_workflow_delivery_message id, and the tab saw the run undelivered earlier in the page session or itsdelivered_atis at or after the baseline'schecked_at(both whole-second UTC).A Retry resumes at a higher control version, so its posting is a new generation and is announced too. Baselines survive the tracker stopping and starting, so StrictMode and flag flips stay quiet. A reload's first read records everything already posted, so nothing is announced twice; the SPA suite proves it in the built app.
Retiring. Only a complete global read retires a run, once: one that was in flight and dropped out. Its row reads Status unavailable. The open chat re-reads its runs; for other chats, the list and the bell reload once per read.
Link routing
workflow_chat_delivery, classic/workflow-activity?…&scope=personallink)chat_response_complete)m365_pending_action_idinmetadataorlink_context)/workflow-activitywith an unknown or missing scope, or a link and metadata that disagree/workflow-activitywithscope=group&groupId=…, not Microsoft 365link_url; aworkflow_priority_alertorworkflow_chat_deliverywith a valid scope and ids in its metadatalink_urland no scopeThe run link is
/workspace/workflows?workflow_id=…&run_id=…, or/groups/<group_id>/workflows?…, which opens the workflow with the run marked and expanded.v2WorkflowRunPathbuilds it only for a known scope and ids that pass the id check.Mutations
51 compiling mutations of mine, all killed: the kickoff's M1–M14 (run as 18), binding note 9's M15–M26 (run as 13), 7 extras (M9b, M27, M28, M29a–c, M30) and 13 for this review's fixes (F1a–F1g on the tracked runs; F2, F2c, F2e, F2m, F2q and F2s on the re-read guard). The coordinator's C1–C7 and U1 have their own table: C2 is now killed by a new tracker test, and M25 was re-run against its new assertion. Two first attempts that didn't compile (TS6133), the literal M13 and the first M19b, aren't counted; M13b and the M19b below replaced them.
Key. tracker, status, actions, routing and delivery are the Node suites
test_v2_workflow_run_tracker.mjs,…_run_status.mjs,…_run_action_clients.mjs,…_run_link_routing.mjsand…_delivery_messages.mjs.card::,spa::,bell::andproposal::name tests intest_v2_workflow_run_card.py,…_run_tracker_spa.py,test_v2_notifications_bell.pyandtest_v2_orchestration_workflow_proposal_card.py, withouttest_. Quoted text is a failing assertion's message.spa::a_delivery_is_announced_once…actions.retrycard::retry_shows_only_where_the_status_row_allows_it[actions-retry-false]/resume-failedcard::retry_resumes_from_a_fresh_runtime_read_and_explains_each_refusalcard::the_open_chat_waits_for_its_reply_to_finish_before_reloading[stream]; a also by deliverybell::a_microsoft_365_notice_keeps_its_classic_run_page[metadata]and[link_context]v2WorkflowRunPathaccepts an unknown workspace typecard::follow_up_needs_the_flag_and_an_available_result[flag-off]App.tsxstarts the tracker before the app has loaded (a); dropBoolean(data) && !error, so it starts behind the error page (a2)spa::the_tracker_stays_off_after_the_app_fails_to_start; a also byspa::the_tracker_waits_for_the_app_to_startandspa::flags_off_keeps_phase_5_links_and_reads_no_statuscard::one_tracker_tick_reads_every_run_in_one_request[ready]dependency list, so it restarts on every renderspa::moving_between_pages_never_starts_a_second_tracker(3 global reads). Survived tracker, which doesn't load the hook.workflow_chat_deliverycasedescribeTypedeep-equal);bell::a_chat_run_notice_reads_as_workflow_results…card::a_delivery_in_the_open_chat_reloads_it_and_marks_it_read[no-source-chosen];card::a_failed_note_offers_retry_only_for_its_own_generation[older-note]workflowonecard::a_delivery_in_another_chat_marks_it_unread_oncecard::…own_generation[older-note]runtime_version, not a fresh runtime readcard::retry_resumes…;card::…own_generation[same-generation]. Survived actions; see Weak spots.card::…before_reloading[orchestration]; a also by deliverygroup_idlink_contextbell::…classic_run_page[link_context];bell::a_workflow_notice_without_a_link_opens_the_run_it_namescard::a_delivery_in_the_open_chat…[source-cleared]card::the_running_tag_names_the_runs_and_gives_way_to_the_unread_dotcard::cancel_asks_first_and_reports_each_outcomeworkflowRunTrackerShouldRunalways returns truespa::flags_off…. Survived delivery, which doesn't test it.proposal::the_run_lines_follow_the_workflow_and_never_guess[paused]proposal::a_created_card_says_when_it_runs_next_and_how_the_last_run_went;proposal::…never_guess[unknown run status]aria-disabledbut still navigates (b), or blocks navigation but dropsaria-disabled(c)card::review_and_approve_waits_while_a_retry_is_under_way: a and c byaria-disabled, b by no navigation after a forced clickMessageList.tsx,MessageActions.tsx)card::, 4 Edit-absent failures:…marks_it_read[no-source-chosen]and[source-cleared],…own_generation[same-generation]and[older-note]hasRunscounts only the list's runs (c); the row count and the empty check ignore them (d)card::runs_the_plans_list_does_not_name_still_show_live[list-failed]and[list-empty]: no rows or no card (a, d), no rows (b), "the card read its chat's runs" (c)card::a_step_the_plans_list_says_cannot_open_stays_closed_with_a_tracked_run(a second row)card::a_plan_run_that_is_gone_shows_nothing_even_with_tracked_runs;card::…stays_closed_with_a_tracked_run(the card shows before the list answers)card::a_question_sent_during_the_re_read_stays_and_the_result_lands_once_after_its_reply[still-coming]and[finished](the question disappears)streamingcard::…after_its_reply[finished]. Survived[still-coming].messageschangedcard::…after_its_reply[still-coming]and[finished](the result never lands)The coordinator's mutations at
ee8383a1d. C2 is now killed, including an independent coordinator rerun atdb2e6a4ed. C1/C3–C7's targeted logic and U1's Checked markup remain unchanged; the tracker file's later change only widened its feature-input type to acceptundefined.>=changed to>seenUndeliveredshortcut inisAnnounceabledeletedee8383a1d. Killed at483291d83by the new tracker test "a run seen before its result was posted is announced once, even when the posting time is behind the first read".v2WorkflowRunPathskipssafeIdaria-liveremoved from Checkedee8383a1d)How they ran. A harness outside the repo applied one mutation at a time through exact, CRLF-aware anchors (stopping on any count mismatch), rebuilt the SPA when asked, ran each named suite alone (Node under 300 s, pytest under 900 s) and restored every file in a
finally. After each batch it rebuilt the clean bundle;git statusforapplication/v2_ui/srcandapplication/single_appwas clean every time, except that the M29 batch listed the then-uncommitted approve hold (committed next as9786915f1). A first no-build pass tripped the card's stale-bundle guard, so its UI runs were redone with builds; its Node-only kills (M2a, M2b, M17, M26) stand.Coverage of the final head. The read-only anchor check at the exact head
bd1a2504afinds 53 ids, 91 stored spec entries, zero mismatches; this proves re-applicability, not an extra mutation run. The F-series ran against the actual fixes. The following 15 existing mutations were rechecked after the changed code, all killed:After the pin/type fixes, the additional 11 rechecks also killed C2, M25, M1, M2a, M2b, M9b, M10, M17, M26, M4 and M12. After the #1643/#1641 merges, M23 was rebuilt and killed again by
the_running_tag_names_the_runs_and_gives_way_to_the_unread_dot(expected zero tags on an unread row); the clean production bundle was rebuilt afterwards. No mutation target changed in those merges. The remaining kills are retained from their recorded runs, not claimed as newly rerun. The coordinator independently confirmed C2's clock-skew kill atdb2e6a4ed.Weak spots. M18 dies only in the card suite: the action-client suite calls the client without a row version, so the mutant falls back to the fresh read there. M11a dies only in the SPA suite; the tracker suite doesn't load the hook. F2s dies only by the delivery suite's source pin: a retry, an edit or a reattached stream turns
streamingon beforemessageschanges, but the card harness refuses a retry at once, has no edit route and never reports a stream to reattach. F2e dies only by its source pin; no UI test stops the tracker while a re-read is out. F2m dies only in[finished]: in[still-coming]the reply is still streaming when the re-read returns, sostreamingalone drops it there too.Contract re-checks
#1628's merged routes retain the resume/cancel contract. Container access is authoritative; the client does not expect a source-access 403 from cancellation. #1635 added a second
PermissionError403 return;d7ea6502aupdates the source pin to require every return in that branch to remain 403. #1636/#1643/#1644/#1641 don't change these workflow routes or 6b-1's status projection.All paths below start
/api/user/workflows/<workflow_id>/runs/<run_id>and retain swagger, login, user and workflow-policy decorators.GET …/runtimePOST …/runtime/resume{expected_version, request_id}: fresh runtime version plus UUID. 200{runtime, can_decide}means Retry requested; unreadable success asks the user to check status. Extra keys/non-object body → 400; conflicts → 409{error, code}.POST …/cancelcodeis never parsed. Every outcome refreshes status.Resume failures. All eight contract codes are tested:
stale_version,invalid_state,request_conflict,deadline_exceeded,workflow_definition_changed,workflow_already_running,workflow_deleting,workflow_deleted.retry_unavailableandnot_resumableare handled too; unknown codes use "This run can't be retried right now. Check its status." 400, 403, both 404 texts, 500, 502, 503 and network failure are covered. Shared user-safe texts replace raw server errors: 403 "You don't have access to this run.", 404 "This run is no longer available.", 503 "Workflows aren't available right now. Try again later."Never called: workflow-level
/cancel,/resume-failedin either scope, or a group action endpoint. The action-client suite pins the durable/resume-failed409 refusal and checks that production TS/TSX has no non-comment reference to that endpoint. The status route remains 6b-1's projection; the only application Python change in this PR is VERSION.Screenshots
docs/reference/chat-controls.mddeclares three slots underdocs/images/reference/. None is captured, so each renders a placeholder listed at/contributing/media-status/.chat-controls-workflow-proposal-run-summary.png): a scheduled workflow's card with Next run, Last run, Open latest results and Follow up.chat-controls-workflow-run-card.png): a Started workflows card with a running run, and the running spinner in the chat list.chat-controls-workflow-posted-result.png): a posted result with its Results from label, Follow up and Open run.Overlap and version
Only V2 was merged, always with rerere disabled. This review round absorbed:
efee70775)3fd4e2840d7ea6502a.5be1d84bd)d751471de["o1", "o2", "o4"]and every must-acknowledge test. The optional-feature type mismatch was fixed inf21dcb73e.25470e7c5)461cd63f9WorkflowRunningTagplus V2's React type imports; preserved both doc histories, Open run provenance wording and all 60 alert-notice cases. Single-scroll-panel behavior is verified.4f3eacbc4)a5c28d684.237and.238sections; no other overlapping code.a5a5b1c53)db2e6a4ed15feec650)d5b82040bconfig.py(.250) as is; the release notes keep V2's.250section below this PR's. #1640 changed noapplication/v2_uiorui_testsfile.Final version:
.251, one above V2's.250, in the separate commiteb090065b; the only commit after it is the test-onlybd1a2504a..252is reserved for the 7a follow-up. The release-notes diff against V215feec650remains a single addition: +42/−0, with every V2 byte preserved.WorkflowProposalCard.tsxremains +3/−1. Of the originally watched UI files, only that card changed;workflowResults.ts,workflowExecutionHistory.tsandWorkflowFlowView.tsxdid not.Trial merges against the head
bd1a2504a(git merge-tree, not actual merges):3a67aaa73config.pyand the release notes; the sharedMessageList.tsxauto-merges. Re-run the card and footer tests if it lands first.21873b8f2config.pyonly.No other PR branch was merged, and no user's PR session was messaged.
CodeQL
CI at the head
bd1a2504a: all 11 checks passed. CodeQL says "No new alerts in code changed by this pull request". The others are Analyze (python, javascript-typescript, actions), broken-access-control-check, xss-sink-check, swagger-route-check, syntax-check, malicious-pr-security-review, enforce-branch-flow and license/cla. Release Notes Check runs only for PRs intoDevelopment.The alert fixed in
bd1a2504a.483291d83(Fix 1 and Fix 2) addedself.streams = 100toRunHarness.__init__afterNotificationApi.__init__had already setstreams(test_v2_workflow_run_card.pyL252), apy/overwritten-inherited-attributewarning. Sinceee8383a1d, only90f87fef0,eb090065bandbd1a2504ahave check runs; the commits between them have none. So CodeQL first reported it at90f87fef0("1 new alert") and again ateb090065b, both times on L252.NotificationApiand the bellHarnessnow take a keyword-onlystreams(default 0), andRunHarnesspassesstreams=100to the base constructor. No alert was dismissed.Earlier alerts, fixed in
ee8383a1d. The first head,bf74a50c8, had 2 more of these warnings in this PR's harnesses:RunHarnessreplacedNotificationApi'sconversations, andRecurringApireplacedProposalApi'sproposal. Each value now goes to the superclass's__init__:NotificationApitakes atitlesmap, andProposalApia starting record. No alert was dismissed.What to watch: DOM-based XSS and URL redirection in the new links. Every run link is a same-origin path from
workflowRunHref, already in the scanner's reviewedTS_SAME_ORIGIN_URL_BUILDERS, with the ids in aURLSearchParamsquery on a fixed path; a group's path comes fromgroupWorkspacePath, which checks and encodes the group id.v2WorkflowRunPathcalls it only for a personal or group scope and ids that passsafeId. In full-file mode the XSS scanner flags two lines in touched files, both outside this diff:App.tsxL172, the faviconsetAttribute('href')(5513596b1), andMessageList.tsxL1552,<a href={streamAuthUrl}>(908109175).Residual risks
simplechat-conversation-<id>, as classic does, so the newest replaces the others. Within a tab, notices are deduped bymessage:<id>, elserun:<id>, elseconversation:<id>. The server still records one unread mark and one bell notice.delivered_atand the route'schecked_atmatters only for a run the tab never saw in flight that posts within the skew of the tab's first read. That result is recorded silently; the unread mark and bell notice still show.streamingcheck is pinned by source only (F2s). No UI test holds a retry, an edit or a reattached stream open during a re-read; see Mutations, weak spots.messageschange, or another reply starting, while that re-read is out.Follow-ups and known limitations
step_labelis always null and elapsed time updates per check; approving happens on the run page.trigger_source='chat_orchestration'and achat_invocation(functions_orchestration_workflow_handoff_decisions.pyL790–794), which the status route selects, so the tracker's running tag, unread state and posted results should cover them; no test here uses a hand-off fixture. The card mounts only for answers whosecapabilities_usedincludesworkflow_run(MessageList.tsxL1135–1140), so a hand-off-only answer gets no card; the roadmap gives that card to 7b.docs/_features/workflow-results-in-chat.mdstays a front-matter-only stub.not_applicable, the workflow's own alerts also firing, rawchat_deliveryin run history, and role revocations after a run starts._PAGEsuites under pytest (32: 4 + 28);three_document_smoke(2);plan_editor_backend(7 errors); the admin connection suites andcontent_screening(20 failures, 54 teardown errors); the 9 regression failures (8 in the 82, plusstats_parityamong the boundary files).ui_tests/fixtures/v2_admin_settings.pyL29'splaywright_connectionimport, which stops 28 suites collecting from the repo root, and the admin fixtures' unroutedGET /api/models/catalog, requested bymodel_catalog_ui.jssince1880a9d0e.