Skip to content

Phase 6b-2: show chat-started workflow runs and their posted results in V2 - #1639

Merged
Paul Lizer (paullizer) merged 34 commits into
microsoft:paullizer-react-v2-uifrom
paullizer:paullizer-phase-6b-2-v2-workflow-run-card-and-trac
Oct 6, 2026
Merged

Paul Lizer (paullizer) merged 34 commits into
microsoft:paullizer-react-v2-uifrom
paullizer:paullizer-phase-6b-2-v2-workflow-run-card-and-trac

Conversation

@paullizer

@paullizer Paul Lizer (paullizer) commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Live run cards under plan answers. Show status, steps, elapsed time and the last check. Server-authorized actions offer confirmed Cancel, Retry from a fresh runtime version, and links to review or reconnect. Tracked runs still appear if Phase 5's list is empty or fails; a 404 stays hidden.
  • One tracker in the app shell. Batched checks keep the chat list and bell current from any page. Delivered results reload a quiet open chat without erasing a question sent during that read; other chats become unread once. The first response is a silent baseline. Flags off preserve Phase 5's links and send zero status requests.
  • Posted-result actions and recurring status. Follow up, generation-matched Retry and Open run appear on delivered messages, without ordinary chat Retry/Edit. Running tags yield to unread dots. Created recurring-workflow cards show their next run and last status.
  • Working V2 run links. Cards and eligible workflow notices open the correct personal/group run. Delivered notices still open the chat; Microsoft 365 notices stay classic. The only application Python change is config.py's version.

Linked issue

Refs #1546
Refs #1543

Release Notes & Latest Features

  • New Feature
  • Bug Fix
  • UI Enhancement
  • Breaking Change
  • Internal only

Is this visible to end users?

  • Yes
  • No

Is this admin-facing (Admin Settings, governance, deployment, config)?

  • Yes
  • No (no new setting; it follows the existing Enable Personal Workflows, Run Workflows From Chat and Use Workflow Results In Chat)

Should this become a Latest Feature card?

  • Yes
  • No
  • Already added

Screenshot needed for the card?

  • Yes
  • No
  • Attached

Version bump

  • application/single_app/config.py VERSION third segment bumped, or not needed because this is docs-only (V2's 0.261.250 → 0.261.251. The Phase 7a: hand off large chat requests to a one-time workflow (server, 0.261.238) #1640 merge d5b82040b took V2's config.py as is; the separate commit eb090065b then changes only 52 version lines in 22 files: config.py from .250, and 51 lines in 21 files from .249, including the assert_app_version_at_least call. No later commit changes a version; bd1a2504a is a test-only CodeQL fix. V2's version history is untouched, and .252 is reserved for the 7a follow-up.)
  • deployers/version.txt bumped, or not needed because deployers/ was not changed (not changed)

Testing / validation

Head: bd1a2504a5530f9e82939f803bff5c31ef239af2 (.251). Recorded V2 base: 15feec6509f0846238362838103be6e76d19f8c1 (.250, #1640), merged in d5b82040b and renumbered separately in eb090065b; bd1a2504a is a test-only CodeQL fix. Base comparisons use a detached worktree, never stash.

What this head re-ran. #1640 changed no file under application/v2_ui or ui_tests, so the V2 sources are byte-identical from db2e6a4ed (the #1641 merge) to this head. Since 90f87fef0, this PR's own lines changed only in version numbers, plus bd1a2504a's two-file harness fix. This round ran light checks only, each labeled with the commit it ran at, plus all 82 regression files and 8 boundary files at the exact head. The focused UI set is left to the coordinator's rerun.

Environment: Windows, Python 3.12 (Flask 3.1.3, pytest 9.0.3, Playwright 1.58.0), Node 24; PYTHONPATH=application/single_app;functional_tests. Heavy jobs ran serially.

Review fixes: 483291d83 appends unjoined tracked runs, guards the delivery re-read after its fetch, adds the C2 clock-skew test and M25 request-count assertion, and cleans the harness CSS sidecar. d7ea6502a repairs the source pin after #1635. f21dcb73e accepts optional bootstrap feature values after #1636: === true still gates each flag; an explicit-undefined test preserves that behavior. Merge details are under Overlap and version.

Build prerequisites: the card/SPA suites use application/single_app/static/v2; orchestration-based suites use the CSS in ui_tests/artifacts/orchestration-plan-editor. Both must match current sources. The card rejects a stale SPA and cleans its own harness bundle/sidecar. The exact old-CSS sidebar failure and its base comparison are under UI suites.

Build and typecheck

npm --prefix application/v2_ui run typecheck   # tsc -b --noEmit
npm --prefix application/v2_ui run build       # tsc -b && vite build
npm --prefix application/v2_ui run build -- --outDir ../../ui_tests/artifacts/orchestration-plan-editor
Command Result
typecheck exit 0 at the exact head bd1a2504a (18 s), and at eb090065b. The two TS2345 errors at d751471de were fixed in f21dcb73e.
build exit 0 at db2e6a4ed, whose V2 sources are identical to this head's: M23's mutant compiled and was killed, then the clean bundle rebuilt in 33 s.
outDir build exit 0 at eb090065b: 2,729 modules, index-BIMKrYWZ.css (97.45 kB), 23.85 s. Also at base a5a5b1c53 (2,716 modules) for the sidebar comparison; both produced the same content-hashed CSS file.

Node suites: 8 files, 105 tests

node --test --test-reporter=tap functional_tests/test_v2_workflow_run_tracker.mjs
# …then each file below the same way
File Tests Result
test_v2_workflow_run_tracker.mjs (new) 35 pass
test_v2_workflow_run_status.mjs (new) 24 pass
test_v2_workflow_run_action_clients.mjs (new) 14 pass
test_v2_workflow_delivery_messages.mjs (new) 13 pass
test_v2_workflow_run_link_routing.mjs (new) 10 pass
test_v2_orchestration_workflow_run_links.mjs (Phase 5) 7 pass
test_workflow_results_clients.mjs 1 pass
test_v2_inline_image_proposal_resume_logic.mjs 1 pass

All 105 passed at the exact head bd1a2504a, and at eb090065b. The coordinator independently reran all 105 at the earlier immutable merge commit db2e6a4ed, and re-killed C2 with the new backwards-clock assertion.

#1640 added code to two server files these suites pin. The action-client suite reads decide_workflow_runtime from functions_workflow_runtime.py (its L365–366), and the status suite reads route_backend_orchestration.py (its L151) and checks the status route's @bp.route line (its L257). Both still pass after it.

The new suites block the network by replacing globalThis.fetch with a function that throws. The action-client suite records requests instead.

After the #1635 merge, the action-client suite failed 1 of 14: it pins that the run-resume route answers a PermissionError with 403, and #1635 added a second 403 to that branch. d7ea6502a requires every return in the branch to be a 403. Changing either 403, or the branch's exception, fails it.

Python files added or touched: four ways

python <file>
python -m pytest <file> -q -p no:cacheprovider
python -O <file>
python -O -m pytest <file> -q -p no:cacheprovider
File Tests (base) Script pytest -O script -O pytest Run at
functional_tests/test_v2_workflow_run_tracking_xss_guardrail.py 6 (new) pass pass pass pass eb090065b
functional_tests/test_v2_orchestration_workflow_run_links_xss_guardrail.py 4 (4) pass pass pass pass eb090065b
functional_tests/test_xss_guardrails_checker.py (unchanged; the checker's own tests) 12 (12) 12/12 pass 12/12 pass eb090065b
ui_tests/test_v2_workflow_run_card.py 41 (new) pass pass pass pass 90f87fef0
ui_tests/test_v2_workflow_run_tracker_spa.py (imports the run-card and bell modules) 6 (new) pass pass pass pass 90f87fef0
ui_tests/test_v2_orchestration_workflow_proposal_card.py 47 (35) pass pass pass pass 90f87fef0
ui_tests/test_v2_workflow_ask_ai_proposal.py (uses the proposal card's fixtures) 3 (3) no __main__ pass no __main__ pass 90f87fef0
ui_tests/test_v2_notifications_bell.py 48 (44) no __main__ pass no __main__ pass 636d58242
ui_tests/test_v2_workflow_alert_notices.py (this PR changes two tests for Open run) 60 (60) no __main__ pass no __main__ pass 636d58242
ui_tests/test_v2_document_provenance.py 13 (13) no __main__ pass no __main__ pass 636d58242
ui_tests/test_v2_workflow_alerts.py (reaches chatStore through AppShell → Sidebar) 30 (30) no __main__ pass no __main__ pass pytest 636d58242, -O pytest 90f87fef0

At the exact head bd1a2504a (clean tree), pytest --collect-only exits 0 for all 11 files with the counts above. bd1a2504a changed the card's and bell's shared harness constructor, so their targeted runs at the exact head cover it: the card's 4 reply and re-read tests passed (37 deselected) and the bell's 13 reply and workflow tests passed (35 deselected). The coordinator's focused rerun should include the full card (41) and bell (48). Earlier, all 44 runs exited 0 together at f21dcb73e, after the #1636 merge.

The -O gap closed in bd578281a. Phase 5's run-links guardrail and the first tracking guardrail looped over their checks under __main__, so python -O stripped every assert and passed with a planted HTML sink. Both now call pytest.main, which rewrites the asserts, so a planted sink fails in all four modes; -O runs print one PytestConfigWarning.

UI suites: every one that loads changed code

python -m pytest ui_tests/<file> -q -p no:cacheprovider   # one file at a time, 1,800 s limit each

Current sidebar check. ui_tests/test_v2_sidebar_conversation_scroll.py passed 9/9 at db2e6a4ed, and the V2 sources haven't changed since. It needs #1643's [data-conversation-rail-header][data-stuck] rule in the orchestration CSS, so rebuild ui_tests/artifacts/orchestration-plan-editor from current sources first. With the stale index-BL2LkJyy.css, 4 fail at L541 (rgba(0, 0, 0, 0) instead of rgb(255, 255, 255) light or rgb(16, 23, 40) dark): reading_down_the_list_scrolls_the_navigation_away_and_holds_search, searching_while_held_keeps_the_search_box_in_place and the_held_search_is_drawn_on_the_themes_solid_surface[light]/[dark]. Base a5a5b1c53 behaves the same: 4 failed and 5 passed with the stale CSS, 9 passed fresh; 636d58242 showed the same 4 with stale CSS. The functional test_v2_sidebar_conversation_scroll.py passed 8/8 in all four modes at 636d58242, and 8 at the exact head.

Historical sweep. 121 UI files ran at 155604921, after the review fixes and before the merges of #1635, #1636, #1643, #1644, #1641 and #1640: 2,797 passed, 58 failed, 61 errors, 4 skipped. Every failure and error was the same at base f1aeef13d. The files were every UI suite that loads chatStore (53, through the harness bundles or image_editor.py's @fixture-chat alias) or the built SPA (83, 19 of them also among the 53), the 5 suites for the other reloadMessages callers (4 already counted) and 3 related workflow and Microsoft 365 suites. 28 files that import playwright_connection through ui_tests/fixtures/v2_admin_settings.py L29 don't collect from the repo root (same on V2), so they were re-run with ui_tests/fixtures on PYTHONPATH. The 4 skips need a live environment.

Failures and errors, each identical at base f1aeef13d:

File At 155604921 Why
test_chat_saved_analysis.py 4 failed, 54 passed The tests expect saved-analysis evidence to become unavailable after an access change, but since 61320e813 the server keeps it readable (test_saved_analysis_routes.py L85–90 expects 200). Classic fails too.
Seven test_v2_orchestration_* suites (failed/passed): pending_approval_resume 4/1, composer 5/1, drawer_views 2/1, elicitation 5/1, plan_card 6/1, plan_editing 4/1, run_hydration 6/1 32 failed, 7 passed _PAGE is set only in main(), so it's None under pytest. As scripts at 155604921: 39/39.
test_chat_three_document_smoke.py 2 failed A download TimeoutError, in the classic and the V2 test.
test_v2_orchestration_plan_editor_backend.py 7 errors Its functions_settings stub has no enabled_required.
test_admin_custom_connections.py 6/3, test_admin_shared_ai_connections.py 5/38, …_classic.py 8/37 (failed/passed) 19 failed, 78 passed, 54 errors The admin fixtures don't route GET /api/models/catalog, which static/js/admin/model_catalog_ui.js (1880a9d0e) requests. The errors are teardown errors, 27 per shared-AI suite. The classic suite wasn't debugged separately.
test_v2_content_screening.py 1 failed, 35 passed L926 waits for "must be enabled before these settings take effect.", which never appears.

Regression: 82 files that read a changed file

Each .py file ran under pytest <file> -q -p no:cacheprovider and each .mjs or .js file under node --test, one at a time with a 1,800 s limit each.

  • At the exact head bd1a2504a (clean tree before and after): the 82 (72 Python; 10 Node, 5 .mjs and 5 .js), plus 8 boundary files (5 Python, 3 Node) that read V2 files Merge CSV, Excel, PDF, Word and PowerPoint files in V2 chat and workflows #1641 or V2: Scroll the chat rail as one panel and pop workflow alerts from the bell #1643 changed. 90 files: 2,370 passed (2,192 pytest, 178 Node), 9 failed, one per file. The 82 alone: 2,317 passed (2,164 pytest, 153 Node), 8 failed.
  • All 9 failures are pre-existing. The same 9 files at base 15feec650 fail the same test ids with the same pass counts. The ninth is in a boundary file: test_v2_stats_parity.py::test_charts_use_the_shared_vendored_runtime ("InlineChart must use the shared loader"). Neither InlineChart.tsx nor any vendored asset is in this diff.
  • The 8 boundary files (functional_tests/test_v2_*): brand_mark_home_link, public_workspace_labels_pin, sidebar_account_menu and stats_parity read Sidebar.tsx, and sidebar_conversation_scroll reads Sidebar.tsx and ConversationRail.tsx, all changed by V2: Scroll the chat rail as one panel and pop workflow alerts from the bell #1643. orchestration_merge_arguments.mjs, workflow_merge_task.mjs and workflow_proposal_merge.mjs read orchestrationMerge, workflowEditor and workflowProposals, changed by Merge CSV, Excel, PDF, Word and PowerPoint files in V2 chat and workflows #1641.
  • Chosen as 56 files that name a changed V2 file (chatStore 25, workflowEditor 16, MessageList 15, MessageActions 10, ConversationRail 7, App.tsx 6, and notificationLinks, workflowRunLink and notifications.ts 1 each), plus 26 others: 6b-1's 13 delivery suites (383 tests), the kickoff's reruns, the workflow authoring and change-tracking suites, and the Node files.
  • Historical: at c852d349b, an earlier renumber, the 82 had the same 8 failures, identical at base f1aeef13d. After Fixes 1–3, the 27 functional files naming chatStore re-ran at 155604921: 302 passed, and the same 7 failed (the eighth file doesn't name chatStore).
The 9 failures, identical at base
  • test_v2_agent_model_exclusivity.py::test_the_composer_wires_the_rule_into_the_toolbar
  • test_v2_collaboration_visual_style_fix.py::test_the_personal_route_is_unchanged_in_behaviour
  • test_v2_conversation_deep_link.py::test_incoming_link_is_captured_before_any_effect_runs
  • test_v2_inline_image_proposals.py::test_no_third_party_browser_assets_were_added (flags @xyflow/react, added by 31e7d16e5)
  • test_v2_message_inspector.py::test_message_metadata_shape_is_role_dependent
  • test_v2_model_identity_and_scope.py::test_model_picker_keys_on_selection_key
  • test_v2_research_voice.py::test_reasoning_levels_match_the_existing_client
  • test_v2_stats_parity.py::test_charts_use_the_shared_vendored_runtime (boundary file)
  • test_v2_tabular_parity.py::test_the_confirmation_precedes_the_send

Docs

python .\scripts\build_docs_inventory.py
python functional_tests/test_docs_app_surface_coverage.py
python functional_tests/test_docs_site_quality.py
python functional_tests/test_docs_link_integrity.py
python functional_tests/test_docs_release_notes_integrity.py
Check Head eb090065b (bd1a2504a changes only two ui_tests files) Base 15feec650
Inventory rebuild Line endings only, restored; no surface change n/a
App-surface coverage, site quality 7/7, 6/6 7/7, 6/6
Link integrity 2/5: broken are 11 of 1,349 relative, 10 of 822 site, 2 of 72 in features.yml 2/5: identical broken-link lists, of 1,338, 821 and 72
Release-notes integrity 0/1, 186 versions missing from the generated pages 0/1, 185; the extra is this PR's v0.261.251, and the kickoff says not to regenerate them
Media coverage (never fails) Screenshots 72/131 filled; reference 21 outstanding 72/128; reference 18

The 3 new outstanding screenshots are this PR's slots in docs/reference/chat-controls.md: chat-controls-workflow-proposal-run-summary.png, chat-controls-workflow-run-card.png and chat-controls-workflow-posted-result.png. The two broken links in docs/reference/chat-controls.md (CHAT_CONTEXT_PICKER, PROMPT_COMPOSER_CARD) are already at base.

Guardrails

run_v2_guardrails.py <worktree> 15feec650 HEAD 6b2, run at the exact head over 15feec650...bd1a2504a, reported "51 changed files; app .py 1, xss surface 26, route .py 1, v2 ts/tsx 25". The one application Python file is config.py.

Check Result
broken-access-control PASS (1 file)
xss-sinks PASS (26 files)
swagger-routes PASS ("Changed Python files do not define Flask routes")
python-syntax PASS (513 files compile)
malicious-pr-review PASS: 775 findings, 0 blockers (684 Important, 91 Moderate)
v2-ts-sinks (added lines) PASS: no sink patterns in the 25 TS/TSX files

The same 775 findings came over a5a5b1c53...90f87fef0 (510 files compile) and at eb090065b; bd1a2504a's harness fix changes no count.

The malicious-PR review's "Needs investigation" verdict comes from keyword heuristics, and I read every category: 368 boundary (the feature's vocabulary: workflow, run, chat, delivery, tracker), 196 security-control (Playwright role= locators, expect(, aria- attributes, test asserts, PermissionError in the action-client pin), 91 obfuscation ("hidden" for tab visibility, fixed re.compile patterns, importlib.util loading the in-repo XSS checker), 54 external-connection (the test origin https://simplechat.test, URL parsing in assertions), 37 dynamic (the guardrail test's deliberate UNSAFE_ROW_SNIPPET, the card test's unlink of its own CSS sidecar) and 29 secret (searchParams.get, credentials: 'same-origin' assertions, authorization reason codes). None loosens a check or reads a credential. Most are in the card test (142), the tracker .mjs (124) and workflowRunTracker.ts (59). Your run at ee8383a1d had 744. This round adds 31 net, all of these kinds: 25 from 483291d83 (Fixes 1–3, M25 and the CSS change, with their tests), and 6 security-control from the pin on L354–360 of the action-client test, where the one-line pin had 1 finding and its 7-line replacement has 7. The f21dcb73e type fix adds none: neither its workflowRunTracker.ts L148 nor its tracker-test L196 is flagged, and the totals match the earlier run over efee70775...d7ea6502a. No dependency, lockfile or CI file changed.

Documentation

  • Release notes updated, or not needed (a ### **(v0.261.251)** section on top with five New Features and two User Interface Enhancements; every older section is byte-identical)
  • Feature documentation updated, or not needed (CHAT_WORKFLOW_RESULT_DELIVERY.md gains a V2 experience (6b-2) section covering the card, tracker, baseline, landing, tag, posted messages, recurring card, links, limitations, code map and tests. Also updated: CHAT_ORCHESTRATION_WORKFLOW_RUNS.md, CHAT_WORKFLOW_RESULTS_FOLLOW_UP.md, V2_NOTIFICATIONS_BELL.md, V2_WORKFLOW_ALERT_NOTICES.md, docs/guides/trigger-a-workflow.md and docs/reference/chat-controls.md, which declares three screenshot slots. The roadmap doc is untouched.)
  • Fix documentation updated, or not needed (not needed; nothing is fixed)

Security checklist

  • New Flask routes include @swagger_route(security=get_auth_security()) (no route added or changed)
  • Settings sent to non-admin frontends use sanitize_settings_for_user() (no settings path changed; the tracker reads the existing bootstrap feature flags)
  • Browser JavaScript is served from local SimpleChat static assets only; no CDN-hosted JS (Vite-bundled local TypeScript only, with no new dependency and no dynamic import)
  • No secrets, keys, connection strings, or local-only artifacts are included (the UI harness bundle is gitignored)

What changed, by file

51 files, +9,545/−189. V2 paths below are under application/v2_ui/src/. The only server file is application/single_app/config.py (+1/−1, VERSION = "0.261.251").

New V2 files

File + Purpose
lib/workflowRunStatus.ts 532 Fail-closed reader for the status route, its phase and waiting texts, and the controls a row allows.
lib/workflowRunTracker.ts 591 The tab's one tracker: cadence, back-off, halting, per-chat baselines, retiring and announcing. Pure, with fetch, timers, clock and visibility injected.
lib/useWorkflowRunTracker.ts 272 Runs the tracker from the app shell and lands posted results through the chat store. A re-read dropped because the chat changed puts its results back to wait.
stores/workflowRunTrackerStore.ts 178 The last snapshot, for the card, the footer and the running tag.
lib/workflowRunActions.ts 149 The run-level cancel and the durable resume, and their answer texts.
lib/useWorkflowRunAction.ts 112 One Cancel or Retry at a time per run, from the card or a note, then a re-read of the chat.
lib/useFocusFallback.ts 35 Moves focus to the row or footer when the focused Cancel or Retry disappears.
lib/workflowDelivery.ts 247 Reads metadata.workflow_delivery, joins a posted message to its run, and decides Follow up and Retry.
lib/m365Links.ts 7 The shared Microsoft 365 connection link.
components/chat/WorkflowRunCard.tsx 395 The run card, including tracked runs Phase 5's list doesn't name.
components/chat/WorkflowDeliveryFooter.tsx 130 The footer on a posted result or note.
components/chat/WorkflowRunningTag.tsx 26 The chat list's running tag.
components/chat/WorkflowProposalRunSummary.tsx 259 Next run, Last run, Open latest results and Follow up on a created recurring-workflow card.

Changed V2 files

File +/− What changed
App.tsx 5/0 useWorkflowRunTracker(Boolean(data) && !error && workflowRunTrackerShouldRun(data?.features)), after useWorkflowAlertRuntime.
components/chat/MessageList.tsx 23/2 Mounts the run card in place of WorkflowRunLinks when tracking is on, under Phase 5's unchanged conditions, and the footer on posted assistant messages in the open personal chat when nothing is masked.
components/chat/MessageActions.tsx 5/1 Hides the chat's own Retry on posted messages.
components/chat/ConversationRail.tsx 2/0 Mounts the running tag after the generating-images tag.
components/chat/WorkflowRunLinks.tsx 27/8 Exports RunLink and moves the one-shot list read into useWorkflowRunLinks, for the card to reuse. Its strings and the guardrail-asserted link line are unchanged.
components/chat/WorkflowProposalCard.tsx 3/1 The child's import, the shared M365_CONNECT_HREF and one render line.
stores/chatStore.ts 51/20 Exports settleCompletedReply, which the streamed path now calls with the same values, so a posted result settles the same way. reloadMessages takes an optional {onlyIfUnchanged}, passed only by the delivery re-read: the fetched list is then dropped, resolving 'superseded', if a reply started or the messages changed while the read was out. Without it, it behaves as before and resolves 'done'.
lib/workflowEditor.ts 10/0 Adds cancelScopedWorkflowRun (the run-level cancel) and newWorkflowRequestId.
lib/workflowRunLink.ts 23/4 workflowRunHref(workflowId, runId, scope = personal) builds the personal or group run link.
lib/notificationLinks.ts 116/16 v2WorkflowRunPath returns the run link instead of null, and workflow notices open their run in V2.
lib/workflowAlertNotices.ts 10/6 An alert's Open run uses the scoped run link.
lib/notifications.ts 4/0 Labels workflow_chat_delivery notices Workflow results.

The other 25 files are the 8 docs under Documentation and the 17 test files below.

Tests

Node/guardrail files are in functional_tests/; browser files are in ui_tests/. Counts and executed modes are above; these are the guarantees they pin.

File Coverage
test_v2_workflow_run_status.mjs Contract envelope/ids, closed state/action sets, safe names, fail-closed unknown values, recovery as running, skipped as failed, batched GET.
test_v2_workflow_run_tracker.mjs Flag gates, cadence, hidden-tab exception, single in-flight read, cancellation/stale replies, backoff, complete-vs-truncated retirement, 10 s request-count dedupe, baseline boundary and C2's backwards-clock delivery.
test_v2_workflow_run_action_clients.mjs Fresh runtime version plus UUID, no resume after a failed read, every refusal branch, run-level cancellation, forbidden endpoint pins.
test_v2_workflow_run_link_routing.mjs Safe ids/scopes, matching group/link context, workflow-only no-link fallback, Microsoft 365 through both metadata locations, alert Open run.
test_v2_workflow_delivery_messages.mjs Prefix/metadata recognition, generation-safe Retry, source: 'workflow', quiet-chat delivery, opt-in post-fetch guard and unchanged other callers, stopped-tracker epoch, list/bell projections.
test_v2_workflow_run_tracking_xss_guardrail.py, test_v2_orchestration_workflow_run_links_xss_guardrail.py Reviewed local URL builders and text rendering; injected HTML/URL sinks still fail. Both script runners use pytest rewriting under -O.
test_v2_workflow_run_card.py Real AppShell/chat/bell: all states/actions, hostile names, flags off, one request per tick, running/unread exclusion, delivery into open/other/busy chats, footers, no ordinary Retry/Edit, approve hold. New regression cases cover missing Phase 5 rows after 500/empty responses, hidden 404, closed steps, and a Send during the re-read with the stream either active or already finished. RunHarness passes streams=100 to the base constructor.
test_v2_workflow_run_tracker_spa.py Built SPA startup/error/flag gates, one tracker across page changes, delivery once and silent reload baseline.
test_v2_orchestration_workflow_proposal_card.py Next/last run, paused/unavailable/unknown cases, per-feature gates, safe names; V2's self-authored approval and merge-task wording retained. RecurringApi initializes rather than overwrites its inherited proposal.
test_v2_notifications_bell.py Workflow-results label/deep link, no-link fallback and both Microsoft 365 exclusions. Harness titles and the first stream number go to the base constructor (keyword-only streams), not overwritten.
test_v2_workflow_alert_notices.py Open run for personal/group runs; Open workflow without a run; read_calls == ["o1", "o2", "o4"]; all V2 must-acknowledge and bell-anchor cases retained.
test_v2_document_provenance.py Existing personal/group alert tests now assert Open run at the same destination.
test_support/workflowRunStatusFixtures.mjs, ui_tests/fixtures/workflow_run_tracking/ Shared contract rows and the real-frame harness; generated bundles remain gitignored and are cleaned.

Binding notes 1–15

# Note How it's met
1 Separate final renumber .251 in eb090065b, after all fixes and V2 merges through #1640; only the test-only bd1a2504a follows it. No stale implementation version; historical V2 sections remain untouched.
2 Flags off = option (a) Phase 5's list renders unchanged and no status request is sent. Phase 5's UI test is unchanged and passes (29); the card and SPA flags-off tests pin it.
3 WorkflowProposalCard.tsx is shared +3/−1: the child's import, the shared M365_CONNECT_HREF and one render line. The rest is in WorkflowProposalRunSummary.tsx.
4 Workflow-only no-link fallback Only workflow_priority_alert and workflow_chat_delivery; another type with a scope and ids gets no link. Microsoft 365 is tested through metadata and link_context, in routing and the bell.
5 Baseline delivered_at >= baselineCheckedAt Still gated by the not-seen set. A same-second posting missing from the first read is announced once; one present in it stays silent.
6 One in-flight action per row Cancel, Retry and Review and approve hold while any action on that run is pending, shared by the card and the run's note. #1628's contracts re-checked; see Contract re-checks.
7 Container-only wording No added doc line says sources or containers are re-checked; the V2 docs describe only what V2 reads and shows.
8 Four modes; detached base See the mode table. Base comparisons ran in a detached worktree: docs and the 9 regression failures at 15feec650 (also the guardrails' base), the sidebar suite at a5a5b1c53, the historical sweep at f1aeef13d. Nothing stashed.
9 Kill table M15–M26 All killed; see Mutations.
10 TS XSS scanner No change needed: workflowRunHref and groupWorkspacePath are already in TS_SAME_ORIGIN_URL_BUILDERS, and the new guardrail proves the checker still flags a link or HTML built from a status row.
11 PR rules; watched files A draft into paullizer-react-v2-ui from my fork (pushes to origin return 403). Of the watched files, only WorkflowProposalCard.tsx changed.
12 Only V2, rerere off All nine V2 merges used git -c rerere.enabled=false merge: 6bd2545a2, 8bd8d269f, 042af8cb7, 3fd4e2840, d751471de, 461cd63f9, a5c28d684, db2e6a4ed, d5b82040b. No PR branches merged.
13 waiting_recovery and skipped Running with "Recovering after an interruption."; Failed and never retried. The status suite pins both.
14 Resume and cancel contracts Resume sends exactly {expected_version, request_id} from a fresh runtime read, and nothing when that read fails. All eight resume 409 codes are tested with their real text; cancel covers 202, 204, both 404s, a 409 with or without a code, 403, 500, 503 and no answer. Every answer re-reads the chat's runs; see Contract re-checks.
15 #1630 overlap The recurring-card tests are at the end of the proposal suite, and approval_state isn't read.

Decisions 1–8

# Decision What shipped
1 One card per answer, a row per run WorkflowRunCard mounts in place of WorkflowRunLinks under Phase 5's conditions. Rows join on run_id + orchestration_run_id + step_id; items not yet read or not joined keep Phase 5's RunLink. Tracked runs for the answer that the list doesn't name, for example because its read failed, are appended oldest first by requested_at as live rows, named as text from the status row. They wait until the list answers or fails, never duplicate a step the list names, and count toward the card's runs, so Check now and the chat's own read follow. A 404 still renders nothing.
2 One engine, one store, one hook The pure workflowRunTracker.ts, the zustand workflowRunTrackerStore.ts and the app-shell useWorkflowRunTracker.ts. Readers never poll; requestConversationRuns(id, {force}) is deduped for 10 s.
3 Baseline keyed by run_id:generation See Tracker cadence and baseline.
4 Land after any stream ends Active means the store is streaming or the chat has an active orchestration, re-checked every 2 s without a request. One re-read covers every result waiting for the open chat; each then settles on its own, as read only if the re-read shows its message. It's the only caller of reloadMessages({ onlyIfUnchanged: true }): if a reply started or the messages changed while it was out, such as a question sent meanwhile, it shows nothing and its results wait for the next quiet moment. One still out when the tracker stops is dropped. The source chip keeps its !analysisContextChosen check.
5 The recurring card reads once Through the existing fetchScopedWorkflows and fetchScopedWorkflowRuns, gated on allow_user_workflows, and again only if the workflow or one of the two flags it reads (allow_user_workflows, enable_chat_workflow_results) changes.
6 N2 gets a relabel only Open run (data-workflow-alert-open-run) replaces Open workflow at the same address. Open workflow remains for an alert with no run; one that can't be placed offers neither, as before. WorkflowAlertCard.tsx is unchanged.
7 Approve opens the run Review and approve links to the run, where WorkflowRuntimePanel shows the gate's prompt and choices. Nothing is approved from chat.
8 Scope-aware run links workflowRunHref(workflowId, runId, scope = personal). Card and footer rows are always personal; notices and N2 pass group scope.

Decisions made autonomously

Each is reversible.

  • available: false: rows and Cancel stay; Retry is hidden, with "Starting workflows from chat is turned off, so Retry isn't available here."
  • Bell label: a static Workflow results of kind workflow. The notice title already carries the outcome, so describeType's signature is unchanged.
  • Elapsed time is the server's elapsed_seconds as of the last check, anchored by "Checked 9:07 AM". It doesn't tick.
  • Check now is one forced read of the chat plus a tracker kick, not a re-read of Phase 5's list.
  • Retry needs a fresh runtime read and sends a fresh UUID per attempt. If the read fails, nothing is sent; there's no fallback to the row's runtime_version.
  • Undeliverable or expired results after the baseline refresh only the bell.
  • Halting: a 401 or 403 on any read, or a 400 on the global read, halts the tracker until the page reloads or the app shell restarts it. The card keeps its rows, says "Live status isn't available right now." and hides Check now.
  • A delivered row with a null message_id or a non-integer generation is recorded but not announced. The card says "Results were posted to this chat." with no jump.
  • The running tag hides while the chat is unread, and when the tracker has stopped or halted.
  • Notices with no link_url but a scope and ids open the V2 run, for workflow types only, as notificationLinks.ts's comment promised.
  • The footer's Retry needs the generation the note reported.
  • Every action answer re-reads the chat's runs; its sentence stays until the run's status changes.
  • An unknown 409, and a cancel 403 or 500, read as generic text.
  • The approve hold keeps the link in place with aria-disabled, a busy style and preventDefault, so focus and layout don't jump. If the focused control disappears, focus falls back to the run's row or the footer.
  • Also: M365_CONNECT_HREF moved to lib/m365Links.ts, shared by both cards; Phase 5's guardrail builder asserts follow the scope-aware signature; screenshot slots are declared, not captured; the card harness has its own fixture directory; both run-tracking XSS guardrails run through pytest.main; the tracker suite fails a hung test after 10 s; and the Edit-absent check has a positive control: the user's own question still offers Edit.

Tracker cadence and baseline

Each tick is one global GET /api/v2/orchestration/workflow-runs/status. A run is in flight while it's queued, running or waiting, or its delivery is pending or delivering.

Situation What the tracker does
Visible tab, a run in flight Checks after 15 s, 30 s, 1 min, 2 min, then every 5 min (WORKFLOW_RUN_POLL_DELAYS_MS).
Nothing in flight Stops checking but stays started. A plan answer, a card finding a run in flight, Check now, or a Cancel or Retry answer kicks it; a kick only brings the next check forward.
Tab hidden Pauses, and checks when shown if a run is in flight, a kick is waiting or the last read failed. With desktop notifications on (permission and setting) and a run in flight, checks every 5 min (WORKFLOW_RUN_HIDDEN_DELAY_MS).
A read fails Backs off 30 s, 1 min, 2 min, then 5 min (WORKFLOW_RUN_ERROR_DELAYS_MS); recovers on the next good read. A per-chat read failing for another reason shows "Couldn't check the run status right now. Try again." on that card and doesn't halt.
401 or 403 on any read, 400 on the global read Halts until the page reloads or the app shell restarts it (the session or a flag goes off and on).
A card asks for its chat One ?conversation_id= read, deduped while out and for 10 s after a success (WORKFLOW_RUN_CONVERSATION_DEDUPE_MS). Check now forces it.
Flags off, session not loaded or failed Never starts; sends nothing.

Baseline. A chat's baseline is the first good read in the page session that covers it, global or the chat's own. Results already posted by then are recorded silently by run_id:generation. A later result is announced once, only if it has an integer generation and an assistant_workflow_delivery_ message id, and the tab saw the run undelivered earlier in the page session or its delivered_at is at or after the baseline's checked_at (both whole-second UTC).

A Retry resumes at a higher control version, so its posting is a new generation and is announced too. Baselines survive the tracker stopping and starting, so StrictMode and flag flips stay quiet. A reload's first read records everything already posted, so nothing is announced twice; the SPA suite proves it in the built app.

Retiring. Only a complete global read retires a run, once: one that was in flight and dropped out. Its row reads Status unavailable. The open chat re-reads its runs; for other chats, the list and the bell reload once per read.

Link routing

Notice or control Opens
6b-1 undeliverable or expired (workflow_chat_delivery, classic /workflow-activity?…&scope=personal link) The run in V2
6b-1 delivered (chat_response_complete) The chat (unchanged)
Microsoft 365 (m365_pending_action_id in metadata or link_context) Classic (unchanged)
/workflow-activity with an unknown or missing scope, or a link and metadata that disagree Classic
/workflow-activity with scope=group&groupId=…, not Microsoft 365 The group run in V2
No link_url; a workflow_priority_alert or workflow_chat_delivery with a valid scope and ids in its metadata The run in V2 (was no link)
No link_url and no scope No link (unchanged)
A workflow alert card with a run id and a scope Open run: the same V2 run page Open workflow opened
The run card, the footer, the recurring card's Open latest results The personal run in V2

The run link is /workspace/workflows?workflow_id=…&run_id=…, or /groups/<group_id>/workflows?…, which opens the workflow with the run marked and expanded. v2WorkflowRunPath builds it only for a known scope and ids that pass the id check.

Mutations

51 compiling mutations of mine, all killed: the kickoff's M1–M14 (run as 18), binding note 9's M15–M26 (run as 13), 7 extras (M9b, M27, M28, M29a–c, M30) and 13 for this review's fixes (F1a–F1g on the tracked runs; F2, F2c, F2e, F2m, F2q and F2s on the re-read guard). The coordinator's C1–C7 and U1 have their own table: C2 is now killed by a new tracker test, and M25 was re-run against its new assertion. Two first attempts that didn't compile (TS6133), the literal M13 and the first M19b, aren't counted; M13b and the M19b below replaced them.

Key. tracker, status, actions, routing and delivery are the Node suites test_v2_workflow_run_tracker.mjs, …_run_status.mjs, …_run_action_clients.mjs, …_run_link_routing.mjs and …_delivery_messages.mjs. card::, spa::, bell:: and proposal:: name tests in test_v2_workflow_run_card.py, …_run_tracker_spa.py, test_v2_notifications_bell.py and test_v2_orchestration_workflow_proposal_card.py, without test_. Quoted text is a failing assertion's message.

ID Mutation Killed by
M1 Remove the first-read baseline tracker (5 baseline tests); spa::a_delivery_is_announced_once…
M2a, M2b Keep the visible ladder while the tab is hidden (a); drop the desktop-notification exception, so a hidden tab never checks (b) tracker: a by "hidden: the kick waits for the tab" and "hidden with desktop notifications off: nothing is scheduled", b by the 5-minute hidden check
M3 Offer Retry from the status alone, ignoring actions.retry status; card::retry_shows_only_where_the_status_row_allows_it[actions-retry-false]
M4 Retry calls /resume-failed actions; card::retry_resumes_from_a_fresh_runtime_read_and_explains_each_refusal
M5a, M5b The delivery rule re-reads the open chat while a reply streams (a); the hook never reports the store's stream (b) card::the_open_chat_waits_for_its_reply_to_finish_before_reloading[stream]; a also by delivery
M6 A Microsoft 365 workflow-activity link opens in V2 routing; bell::a_microsoft_365_notice_keeps_its_classic_run_page[metadata] and [link_context]
M7 v2WorkflowRunPath accepts an unknown workspace type routing
M8 Follow up shows with the results-in-chat flag off card::follow_up_needs_the_flag_and_an_available_result[flag-off]
M9a, M9a2 App.tsx starts the tracker before the app has loaded (a); drop Boolean(data) && !error, so it starts behind the error page (a2) spa::the_tracker_stays_off_after_the_app_fails_to_start; a also by spa::the_tracker_waits_for_the_app_to_start and spa::flags_off_keeps_phase_5_links_and_reads_no_status
M10 One status request per tracked run before the batched read tracker (several, among them "a posting without a message id is never announced"); card::one_tracker_tick_reads_every_run_in_one_request
M11a Drop the hook's [ready] dependency list, so it restarts on every render spa::moving_between_pages_never_starts_a_second_tracker (3 global reads). Survived tracker, which doesn't load the hook.
M12 The bell has no workflow_chat_delivery case delivery (the describeType deep-equal); bell::a_chat_run_notice_reads_as_workflow_results…
M13b The plain chat Retry shows on posted messages card::a_delivery_in_the_open_chat_reloads_it_and_marks_it_read[no-source-chosen]; card::a_failed_note_offers_retry_only_for_its_own_generation[older-note]
M14a, M14b Recognise a posted message by its id prefix only (a), or by its metadata only (b) delivery (a: metadata alone marks it; b: the prefix alone does)
M15 A posted result settles as a chat reply, not a workflow one delivery (the reply deep-equal); card::a_delivery_in_another_chat_marks_it_unread_once
M16 A failed note offers Retry whatever its generation delivery (a generation 3 note against a generation 4 row); card::…own_generation[older-note]
M17 Retire missing rows on a truncated global read tracker ("a truncated read leaves rows out, so it retires nothing")
M18 Resume from the row's runtime_version, not a fresh runtime read card::retry_resumes…; card::…own_generation[same-generation]. Survived actions; see Weak spots.
M19a, M19b The delivery rule re-reads the open chat while its plan runs (a); the hook never reports an active plan (b) card::…before_reloading[orchestration]; a also by delivery
M20 Read the workflow's group from the notice's group_id routing
M21 Ignore a Microsoft 365 action id in link_context routing; bell::…classic_run_page[link_context]; bell::a_workflow_notice_without_a_link_opens_the_run_it_names
M22 A reload's chip auto-select overrides a source the user chose card::a_delivery_in_the_open_chat…[source-cleared]
M23 The running tag stays while the chat is unread card::the_running_tag_names_the_runs_and_gives_way_to_the_unread_dot
M24 Cancel without the confirmation dialog card::cancel_asks_first_and_reports_each_outcome
M25 Drop the 10 s per-chat read dedupe tracker ("right after a good answer no second request is sent"). Re-run this round; before, only the suite's 10 s hang limit caught it.
M26 A hidden tab with desktop notifications keeps checking with nothing in flight tracker ("nothing in flight: a hidden tab stops checking")
M9b workflowRunTrackerShouldRun always returns true tracker; spa::flags_off…. Survived delivery, which doesn't test it.
M27 A paused workflow still shows its stored next run proposal::the_run_lines_follow_the_workflow_and_never_guess[paused]
M28 Open latest results opens the newest run whatever its status proposal::a_created_card_says_when_it_runs_next_and_how_the_last_run_went; proposal::…never_guess[unknown run status]
M29a–c Review and approve ignores a pending action on the run (a), shows aria-disabled but still navigates (b), or blocks navigation but drops aria-disabled (c) card::review_and_approve_waits_while_a_retry_is_under_way: a and c by aria-disabled, b by no navigation after a forced click
M30 Edit on every message, posted ones included (MessageList.tsx, MessageActions.tsx) card::, 4 Edit-absent failures: …marks_it_read[no-source-chosen] and [source-cleared], …own_generation[same-generation] and [older-note]
F1a–d Leave out the tracked runs the plan's list doesn't name (a); count them but render none (b); hasRuns counts only the list's runs (c); the row count and the empty check ignore them (d) card::runs_the_plans_list_does_not_name_still_show_live[list-failed] and [list-empty]: no rows or no card (a, d), no rows (b), "the card read its chat's runs" (c)
F1e Append every tracked run for the answer, even ones the list names the full card suite: 18 failed, 23 passed
F1f Leave out only runs the list names, so a step the list closed gets a live row card::a_step_the_plans_list_says_cannot_open_stays_closed_with_a_tracked_run (a second row)
F1g Show tracked runs while the list is still loading card::a_plan_run_that_is_gone_shows_nothing_even_with_tracked_runs; card::…stays_closed_with_a_tracked_run (the card shows before the list answers)
F2, F2c Remove the post-fetch guard (F2); the delivery re-read doesn't ask for it (F2c) card::a_question_sent_during_the_re_read_stays_and_the_result_lands_once_after_its_reply[still-coming] and [finished] (the question disappears)
F2m The guard checks only streaming card::…after_its_reply[finished]. Survived [still-coming].
F2s The guard checks only whether messages changed delivery (the guard's source pin). Survived both card params.
F2q A dropped re-read doesn't put its results back to wait card::…after_its_reply[still-coming] and [finished] (the result never lands)
F2e A re-read that was out when the tracker stopped still puts its results back delivery (the epoch check's source pin)

The coordinator's mutations at ee8383a1d. C2 is now killed, including an independent coordinator rerun at db2e6a4ed. C1/C3–C7's targeted logic and U1's Checked markup remain unchanged; the tracker file's later change only widened its feature-input type to accept undefined.

ID Mutation Result
C1 The baseline >= changed to > Killed by the same-second tracker test
C2 The seenUndelivered shortcut in isAnnounceable deleted Survived at ee8383a1d. Killed at 483291d83 by the new tracker test "a run seen before its result was posted is announced once, even when the posting time is behind the first read".
C3 A 400 on the global read no longer halts Killed
C4 The stale-read guard removed Killed
C5 v2WorkflowRunPath skips safeId Killed by routing
C6 An unknown row counts as in flight Killed by status
C7 Retry falls back to version 0 when the fresh read fails Killed by actions
U1 aria-live removed from Checked Killed by the card suite (its L585 at ee8383a1d)

How they ran. A harness outside the repo applied one mutation at a time through exact, CRLF-aware anchors (stopping on any count mismatch), rebuilt the SPA when asked, ran each named suite alone (Node under 300 s, pytest under 900 s) and restored every file in a finally. After each batch it rebuilt the clean bundle; git status for application/v2_ui/src and application/single_app was clean every time, except that the M29 batch listed the then-uncommitted approve hold (committed next as 9786915f1). A first no-build pass tripped the card's stale-bundle guard, so its UI runs were redone with builds; its Node-only kills (M2a, M2b, M17, M26) stand.

Coverage of the final head. The read-only anchor check at the exact head bd1a2504a finds 53 ids, 91 stored spec entries, zero mismatches; this proves re-applicability, not an extra mutation run. The F-series ran against the actual fixes. The following 15 existing mutations were rechecked after the changed code, all killed:

Killed by Mutations
delivery and card M5a, M19a
card M5b, M19b, M22, M24, M13b, M30, M29a–c, M8, M23
SPA (3 global reads; still survives tracker) M11a
card (still survives actions) M18

After the pin/type fixes, the additional 11 rechecks also killed C2, M25, M1, M2a, M2b, M9b, M10, M17, M26, M4 and M12. After the #1643/#1641 merges, M23 was rebuilt and killed again by the_running_tag_names_the_runs_and_gives_way_to_the_unread_dot (expected zero tags on an unread row); the clean production bundle was rebuilt afterwards. No mutation target changed in those merges. The remaining kills are retained from their recorded runs, not claimed as newly rerun. The coordinator independently confirmed C2's clock-skew kill at db2e6a4ed.

Weak spots. M18 dies only in the card suite: the action-client suite calls the client without a row version, so the mutant falls back to the fresh read there. M11a dies only in the SPA suite; the tracker suite doesn't load the hook. F2s dies only by the delivery suite's source pin: a retry, an edit or a reattached stream turns streaming on before messages changes, but the card harness refuses a retry at once, has no edit route and never reports a stream to reattach. F2e dies only by its source pin; no UI test stops the tracker while a re-read is out. F2m dies only in [finished]: in [still-coming] the reply is still streaming when the re-read returns, so streaming alone drops it there too.

Contract re-checks

#1628's merged routes retain the resume/cancel contract. Container access is authoritative; the client does not expect a source-access 403 from cancellation. #1635 added a second PermissionError 403 return; d7ea6502a updates the source pin to require every return in that branch to remain 403. #1636/#1643/#1644/#1641 don't change these workflow routes or 6b-1's status projection.

All paths below start /api/user/workflows/<workflow_id>/runs/<run_id> and retain swagger, login, user and workflow-policy decorators.

Call Request and response handling
GET …/runtime Required before Retry. Refuses a non-resumable runtime or a version that isn't a safe non-negative integer. A failed read sends no resume.
POST …/runtime/resume Exactly {expected_version, request_id}: fresh runtime version plus UUID. 200 {runtime, can_decide} means Retry requested; unreadable success asks the user to check status. Extra keys/non-object body → 400; conflicts → 409 {error, code}.
POST …/cancel No body, after confirmation. 2xx → Cancel requested; 404 → run unavailable; 409 → run finished/changed. Other failures, including durable 500, show a fixed failure message. A cancel code is never parsed. Every outcome refreshes status.

Resume failures. All eight contract codes are tested: stale_version, invalid_state, request_conflict, deadline_exceeded, workflow_definition_changed, workflow_already_running, workflow_deleting, workflow_deleted. retry_unavailable and not_resumable are handled too; unknown codes use "This run can't be retried right now. Check its status." 400, 403, both 404 texts, 500, 502, 503 and network failure are covered. Shared user-safe texts replace raw server errors: 403 "You don't have access to this run.", 404 "This run is no longer available.", 503 "Workflows aren't available right now. Try again later."

Never called: workflow-level /cancel, /resume-failed in either scope, or a group action endpoint. The action-client suite pins the durable /resume-failed 409 refusal and checks that production TS/TSX has no non-comment reference to that endpoint. The status route remains 6b-1's projection; the only application Python change in this PR is VERSION.

Screenshots

docs/reference/chat-controls.md declares three slots under docs/images/reference/. None is captured, so each renders a placeholder listed at /contributing/media-status/.

  • Created workflow card (chat-controls-workflow-proposal-run-summary.png): a scheduled workflow's card with Next run, Last run, Open latest results and Follow up.
  • Workflow run card (chat-controls-workflow-run-card.png): a Started workflows card with a running run, and the running spinner in the chat list.
  • Results posted to the chat (chat-controls-workflow-posted-result.png): a posted result with its Results from label, Follow up and Open run.

Overlap and version

Only V2 was merged, always with rerere disabled. This review round absorbed:

V2 change Merge commit Resolution
#1635 (efee70775) 3fd4e2840 Kept both release-note sections; its changed 403 source pin was repaired in d7ea6502a.
#1636 (5be1d84bd) d751471de Kept ["o1", "o2", "o4"] and every must-acknowledge test. The optional-feature type mismatch was fixed in f21dcb73e.
#1643 (25470e7c5) 461cd63f9 Kept WorkflowRunningTag plus V2's React type imports; preserved both doc histories, Open run provenance wording and all 60 alert-notice cases. Single-scroll-panel behavior is verified.
#1644 (4f3eacbc4) a5c28d684 Kept its .237 and .238 sections; no other overlapping code.
#1641 (a5a5b1c53) db2e6a4ed Kept the full proposal-test history from both sides; card and editor auto-merged. V2's merge-task wording and self-authored approval text remain alongside the new summary.
#1640 (15feec650) d5b82040b Took V2's config.py (.250) as is; the release notes keep V2's .250 section below this PR's. #1640 changed no application/v2_ui or ui_tests file.

Final version: .251, one above V2's .250, in the separate commit eb090065b; the only commit after it is the test-only bd1a2504a. .252 is reserved for the 7a follow-up. The release-notes diff against V2 15feec650 remains a single addition: +42/−0, with every V2 byte preserved. WorkflowProposalCard.tsx remains +3/−1. Of the originally watched UI files, only that card changed; workflowResults.ts, workflowExecutionHistory.ts and WorkflowFlowView.tsx did not.

Trial merges against the head bd1a2504a (git merge-tree, not actual merges):

Open PR Checked head Overlap
#1637 3a67aaa73 Conflicts in config.py and the release notes; the shared MessageList.tsx auto-merges. Re-run the card and footer tests if it lands first.
#1638 21873b8f2 config.py only.

No other PR branch was merged, and no user's PR session was messaged.

CodeQL

CI at the head bd1a2504a: all 11 checks passed. CodeQL says "No new alerts in code changed by this pull request". The others are Analyze (python, javascript-typescript, actions), broken-access-control-check, xss-sink-check, swagger-route-check, syntax-check, malicious-pr-security-review, enforce-branch-flow and license/cla. Release Notes Check runs only for PRs into Development.

The alert fixed in bd1a2504a. 483291d83 (Fix 1 and Fix 2) added self.streams = 100 to RunHarness.__init__ after NotificationApi.__init__ had already set streams (test_v2_workflow_run_card.py L252), a py/overwritten-inherited-attribute warning. Since ee8383a1d, only 90f87fef0, eb090065b and bd1a2504a have check runs; the commits between them have none. So CodeQL first reported it at 90f87fef0 ("1 new alert") and again at eb090065b, both times on L252. NotificationApi and the bell Harness now take a keyword-only streams (default 0), and RunHarness passes streams=100 to the base constructor. No alert was dismissed.

Earlier alerts, fixed in ee8383a1d. The first head, bf74a50c8, had 2 more of these warnings in this PR's harnesses: RunHarness replaced NotificationApi's conversations, and RecurringApi replaced ProposalApi's proposal. Each value now goes to the superclass's __init__: NotificationApi takes a titles map, and ProposalApi a starting record. No alert was dismissed.

What to watch: DOM-based XSS and URL redirection in the new links. Every run link is a same-origin path from workflowRunHref, already in the scanner's reviewed TS_SAME_ORIGIN_URL_BUILDERS, with the ids in a URLSearchParams query on a fixed path; a group's path comes from groupWorkspacePath, which checks and encodes the group id. v2WorkflowRunPath calls it only for a personal or group scope and ids that pass safeId. In full-file mode the XSS scanner flags two lines in touched files, both outside this diff: App.tsx L172, the favicon setAttribute('href') (5513596b1), and MessageList.tsx L1552, <a href={streamAuthUrl}> (908109175).

Residual risks

  1. Each tab tracks on its own. Every hidden or unfocused tab with desktop notifications on can raise a notice for the same result. They share the tag simplechat-conversation-<id>, as classic does, so the newest replaces the others. Within a tab, notices are deduped by message:<id>, else run:<id>, else conversation:<id>. The server still records one unread mark and one bell notice.
  2. N tabs mean N global reads per tick, each capped by the server at 10 live runtime reads and 20 workflow reads.
  3. A truncated global read (50 rows) doesn't announce runs outside it or retire one that drops out, so a running tag can go stale. The unread mark and bell notice still show on the next list or bell reload.
  4. Clock skew between the delivery worker's delivered_at and the route's checked_at matters only for a run the tab never saw in flight that posts within the skew of the tab's first read. That result is recorded silently; the unread mark and bell notice still show.
  5. Halting on a 401 or 403, or a 400 on the global read, lasts the page session unless the app shell restarts the tracker. The card says so and hides Check now.
  6. The guard's streaming check is pinned by source only (F2s). No UI test holds a retry, an edit or a reattached stream open during a re-read; see Mutations, weak spots.
  7. A dropped re-read is tried again. Its results go back to wait. If the chat is quiet, the next re-read starts at once; while a reply streams, they wait, re-checked on each store change and every 2 s without a request. Each further drop needs another messages change, or another reply starting, while that re-read is out.

Follow-ups and known limitations

  • The feature doc's V2 limitations: both flags are required; a truncated read never retires a run; a chat with more than 20 runs falls back to Phase 5 links for the rest; an idle tab learns of a new run at its next kick or app start; a posting with a null message id or generation isn't announced; step_label is always null and elapsed time updates per check; approving happens on the run page.
  • Resume 409 codes without their own text read as the generic one.
  • A failing global read isn't shown on the card; the tracker backs off quietly. The card shows an error only for its own chat's read or a halt.
  • Group bell notices are covered by the routing suite's group cases, not the bell UI suite. The alert card's group Open run is UI-tested.
  • Phase 7 (Workflow orchestration follow-ons: group workflows, SimpleChat action parity and plan replay #1550) group delivery: card and footer rows are always personal today.
  • 7a hand-off runs (Phase 7a: hand off large chat requests to a one-time workflow (server, 0.261.238) #1640) are queued with trigger_source='chat_orchestration' and a chat_invocation (functions_orchestration_workflow_handoff_decisions.py L790–794), which the status route selects, so the tracker's running tag, unread state and posted results should cover them; no test here uses a hand-off fixture. The card mounts only for answers whose capabilities_used includes workflow_run (MessageList.tsx L1135–1140), so a hand-off-only answer gets no card; the roadmap gives that card to 7b.
  • A held Review and approve link still opens on a middle-click or "Open in new tab".
  • Screenshots aren't captured, and docs/_features/workflow-results-in-chat.md stays a front-matter-only stub.
  • Weak kills: M18 and M11a could get direct assertions; F2s and F2e could get UI tests (a re-read returning after a retry, edit or reattached stream starts, and one out when the tracker stops).
  • 6b-1's residual risks are unchanged: the mark-before-create windows, the double-fault not_applicable, the workflow's own alerts also firing, raw chat_delivery in run history, and role revocations after a run starts.
  • Pre-existing, identical at base: saved-analysis (4); the orchestration _PAGE suites under pytest (32: 4 + 28); three_document_smoke (2); plan_editor_backend (7 errors); the admin connection suites and content_screening (20 failures, 54 teardown errors); the 9 regression failures (8 in the 82, plus stats_parity among the boundary files).
  • Also on V2, not touched here: ui_tests/fixtures/v2_admin_settings.py L29's playwright_connection import, which stops 28 suites collecting from the repo root, and the admin fixtures' unrouted GET /api/models/catalog, requested by model_catalog_ui.js since 1880a9d0e.

Paul Lizer (paullizer) and others added 19 commits October 3, 2026 02:30
The status parser, the tab's one tracker engine and its store and hook, the
run-level cancel and durable-resume clients, the delivered-message helpers,
the scope-aware run deep link, and the bell's workflow_chat_delivery case.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…run summary

- WorkflowRunCard replaces the Started workflows links when chat workflow runs
  are on: phase-driven states, Cancel with confirmation, Retry through the
  durable runtime resume, Review and approve via the run inspector, Reconnect
  Microsoft 365, Open run, and Check now with a polite Checked time.
- WorkflowDeliveryFooter on messages a workflow run posted: Follow up (puts the
  run's result in this chat's composer), Retry workflow run only when a fresh
  status row allows it, and Open run.
- Plain chat Retry is hidden on delivered messages (the server refuses it).
- WorkflowRunningTag next to the generating-images tag, built from tracker
  state only.
- WorkflowProposalRunSummary on a created recurring workflow: next run and last
  run status, Open latest results and Follow up.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Pull the delivered-message helpers out of the shell hook into pure functions
(workflowDelivery.ts) so the landing plan, the busy check, the reply shape
and the footer's Retry gate can be tested directly. Give the tracker a
per-chat first-read baseline, so a chat read never moves the global one,
and make v2WorkflowRunPath refuse scope types it doesn't know.

Add node tests for the status parser, the tracker engine, the run-link
routing, the run action clients and the delivered-message helpers, with
shared fixtures in the status route's exact shape. Server facts the client
mirrors are pinned against the modules that write them.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
- Run-card harness (ui_tests/fixtures/workflow_run_tracking) mounting the real
  stores, tracker, card, footer and running tag against a mocked API.
- test_v2_workflow_run_card.py: phases, actions, fail-closed states, delivery
  settle in the open and other chats, baseline, footers, tag, XSS.
- test_v2_workflow_run_tracker_spa.py: one tracker per tab in the built SPA,
  flags-off zero requests, Check now, reload never re-announces.
- Version headers moved to 0.261.233.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
- Bell: a workflow_chat_delivery notice reads as "Workflow results" and
  opens its run in V2, with or without a link; a notice without its
  workspace gets no link, and a Microsoft 365 notice keeps its classic page.
- Proposal card: a created card shows its next run in the reader's time
  zone, the newest run's status, Open latest results and Follow up, never
  guesses an unknown status, honours the workflow and results flags, and
  says run details are unavailable when a read fails.

Refs microsoft#1546

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The new run card, delivery footer, running tag, recurring run lines,
tracker, store and helpers pass check_xss_sinks.py in full, as do the
changed files that had no older findings. No changed V2 file writes raw
HTML, every run-tracking link is workflowRunHref over a run's ids or the
fixed Microsoft 365 path, and status rows carry ids, never a URL.
Synthetic snippets show the checker still flags a link or HTML taken
from a status row.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The fake status route answers only when a test says so, so a change that
sends a request a test doesn't expect left the test waiting with no end.
Each test now has a 10-second limit, which turns that into a failure.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Adds the V2 section to the 6b-1 feature doc: when V2 tracks runs, the run
card's states and actions, Check now, the tracker's cadence and baseline, how
a posted result lands, the running tag, the posted-message footer, the
recurring-workflow card, run links and V2 limitations. Splits File structure
and Testing into Server (6b-1) and V2 (6b-2).

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
- chat-controls: next run, last run, open latest results and follow up on the
  recurring-workflow card; live status, check now, cancel run, retry, review
  and approve, reconnect Microsoft 365, results posted below and the running
  tag on started workflows; a new "Results posted to the chat" section.
- trigger-a-workflow: live status card, cancel and retry from chat, posted
  results with follow up, needs-you waits, two troubleshooting rows.
- Phase 5 and 6a feature docs point to the 6b-2 V2 experience.

Refs microsoft#1546

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
New Features: live run status under a plan's answer, one tracker per
browser tab, Follow up / Retry / Open run on posted results, next and
last run on the recurring-workflow card, and workflow notices that open
the run in V2. UI enhancements: the running tag in the chat list and the
Workflow results bell label and alert-card Open run.

Refs microsoft#1546

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Brings in microsoft#1626 (orchestration file render permission fix, released as
0.261.233) and microsoft#1633 (roadmap updates). The only conflict was the top of
release_notes.md, where both sides added a v0.261.233 section; both are
kept intact under their own headings. 6b-2 is renumbered to 0.261.234 in
the next commit.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
microsoft#1626 released 0.261.233 first, so 6b-2 moves to the next patch version:
config.py, its release notes heading, and every 6b-2 version reference in
its feature docs, test headers and version assertions. microsoft#1626's own
0.261.233 references are unchanged.

Refs microsoft#1546

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Cancel and Retry on the run card already wait while another action on the
same run is under way, but Review and approve did not. A read that landed
mid-Retry and found the run waiting for approval rendered a live link to the
gate while the resume was still in flight.

The link now carries aria-disabled and swallows click and Enter while any
action on that run is pending, exactly like Cancel and Retry, and comes back
once the action is answered and the chat's runs have been re-read.

The new UI test holds the resume, lands a waiting-for-approval read with
Check now, and checks the link is held for both a click and Enter, then opens
the run once the resume is answered. Mutations M29a (busy always false),
M29b (no preventDefault) and M29c (no aria-disabled) are all killed.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
… card

Adds three ad-hoc media slots to docs/reference/chat-controls.md, following the
page's existing media.html pattern: the Started workflows run card with the
chat-list running tag, a result posted to the chat with Follow up and Open run,
and the created proposal card's Next run and Last run. No images are committed;
the slots render placeholders until someone captures them. Docs-only, so the
version is unchanged.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Both run tracking XSS guardrail files ran their tests from a plain loop, so python -O removed every assert and the optimized script run passed even with a planted HTML sink. Their __main__ now runs pytest.main, which rewrites the asserts into explicit checks, the same runner as the V2 run card and tracker UI suites. A planted regression now fails in all four modes.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
The card suite's posted-message helper now opens More actions on the posted message and on the question in the same chat. It proves the menu opened, then checks that only the question offers Edit.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Comment thread ui_tests/test_v2_orchestration_workflow_proposal_card.py Fixed
Comment thread ui_tests/test_v2_workflow_run_card.py Fixed
Paul Lizer (paullizer) and others added 3 commits October 5, 2026 06:17
CodeQL flagged two py/overwritten-inherited-attribute warnings on microsoft#1639:
RunHarness reassigned `conversations` after NotificationApi set it, and
RecurringApi reassigned `proposal` after ProposalApi set it. Following the
query's advice, NotificationApi and the bell Harness now take an optional
`titles` map and ProposalApi an optional starting record, which the
subclasses pass instead of overwriting. The constructed state is unchanged.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…during the delivery re-read

- WorkflowRunCard appends the tracker's runs for the answer that the Phase 5
  list doesn't name, oldest request first, once the list has loaded or failed.
  A run or step the list already names is left out, so a step the list says
  can't open stays closed. A plan run that's gone (404) still shows nothing.
- reloadMessages takes an opt-in onlyIfUnchanged guard, used only by the
  delivery re-read: it is dropped if a reply started or the messages changed
  while it was out, and its results wait for the next quiet moment. Every
  other caller behaves as before.
- Tracker test for the seen-undelivered shortcut, a request-count check for
  the 10 s dedupe window, and card tests for both fixes. The card fixture
  removes the bundle's CSS sidecar on teardown.
- CHAT_WORKFLOW_RESULT_DELIVERY.md describes both fixes and the new tests.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
V2 is still at 0.261.233. microsoft#1635, microsoft#1636, microsoft#1637 and microsoft#1638 claim .234 to .236,
so this branch takes .237. Every 0.261.234 reference moves to 0.261.237:
config.py, the release notes header, the feature docs, the chat-controls
reference and the test headers.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Paul Lizer (paullizer) and others added 6 commits October 5, 2026 12:40
Brings in microsoft#1635 (clearer saved-result error messages and repaired workflow
test harnesses, released as 0.261.234). The only conflicts were VERSION in
config.py, kept at 0.261.237 because that is still above V2's 0.261.234,
and the top of release_notes.md, where both sides added a section; both are
kept intact under their own headings, v0.261.237 above v0.261.234.

Refs microsoft#1546

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
microsoft#1635 split _workflow_runtime_response's PermissionError branch: a saved record that fails its own check (AnalysisResultUnavailable) now answers its own 403 text before the access 403. The action-client suite pinned the old single-return shape, so it failed after the V2 merge. The pin now reads the outer branch and requires every return in it to be a 403, which V2 shows as its fixed no-access sentence whatever the server says.

Refs microsoft#1546

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Brings in microsoft#1636 (V2 workflow alerts: require acknowledgment, repeating
sound, size options and team delivery, released as 0.261.235). Three
conflicts:

- config.py: VERSION kept at 0.261.237, which is still above V2's
  0.261.235.
- release_notes.md: both sides added a section at the top. Both are kept
  intact under their own headings, v0.261.237 above v0.261.235, so
  against V2 the file only gains 6b-2's section.
- ui_tests/test_v2_workflow_alert_notices.py: the last line of
  test_open_run_goes_to_the_run_in_its_workspace keeps 6b-2's
  read_calls == ["o1", "o2", "o4"], and every must-acknowledge test
  microsoft#1636 appended after it is kept unchanged.

Refs microsoft#1546

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
microsoft#1636 changed BootstrapPayload.features to BootstrapFeatures, which allows undefined values, so App.tsx and MessageList.tsx no longer typechecked against workflowRunTrackerShouldRun's Record<string, boolean> parameter after the merge. The gate already compares each flag with === true, so only the parameter type widens; the tracker test now also covers an explicit undefined flag.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Brings in V2 25470e7 (0.261.236, the V2 sidebar conversation scroll fix).

Conflicts resolved:
- config.py: keep 0.261.237.
- ConversationRail.tsx: keep the WorkflowRunningTag import and mount with
  V2's widened FocusEvent/ReactNode/RefObject type import.
- V2_WORKFLOW_ALERT_NOTICES.md: keep both since-version lines in order;
  V2's 60-test row with Open run wording; Open run provenance row.
- test_v2_workflow_alert_notices.py: Version 0.261.237 with every
  history line from both sides.
- release_notes.md: 0.261.237 section above V2's 0.261.236 section.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Brings in V2 4f3eacb (0.261.237 orchestration settings document type
fix and 0.261.238 Microsoft 365 actions in orchestration plans).

Conflicts resolved:
- config.py: keep this branch's version; it is renumbered above V2's
  0.261.238 in the next commit.
- release_notes.md: this branch's section on top, then V2's 0.261.238
  and 0.261.237 sections unchanged.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
V2 is 0.261.238 after microsoft#1644, which also used 0.261.237. Moves this
branch's version, release-notes heading and every 0.261.237 line it adds
to 0.261.239. V2's own 0.261.237 section and files are unchanged.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Paul Lizer (paullizer) added a commit that referenced this pull request Oct 5, 2026
V2 reached 0.261.238 when #1644 merged, so this branch moves one above it.
Sets config.py VERSION to 0.261.239 and moves this branch's own 0.261.238
lines to 0.261.239: the hand-off modules' and tests' Version, Implemented in,
MINIMUM_VERSION and assert_app_version_at_least literals, the two route-test
coverage notes, the run adapter test's hand-off refusal note, the two
hand-off rows in docs/admin/orchestration.md, the feature doc and this
branch's release-notes header. #1644's own 0.261.238 lines (its schema
docstring, its Microsoft 365 section and troubleshooting row in
orchestration.md, its fix docs and its release-notes section) are unchanged.
The open V2 PRs claim 0.261.234 (#1639), 0.261.235 (#1637), 0.261.236
(#1638) and 0.261.246 (#1641), none of them 0.261.239.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Paul Lizer (paullizer) and others added 2 commits October 5, 2026 21:17
Preserve both proposal-test histories and every V2 release-note section. The version bump follows separately.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Keep all base version history; update only 6b-2's version, tests and documentation after the V2 microsoft#1641 merge.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Comment thread ui_tests/test_v2_workflow_run_card.py Fixed
Paul Lizer (paullizer) and others added 3 commits October 5, 2026 21:55
Take V2's config.py (0.261.250) and keep 6b-2's release-note section above V2's 0.261.250 section with every V2 byte preserved. The renumber to 0.261.251 follows separately.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Keep all base version history; update only 6b-2's version, tests and documentation after the V2 microsoft#1640 merge, which took 0.261.250: config.py moves from 0.261.250 and 51 lines in 21 other files move from 0.261.249, all to 0.261.251.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
CodeQL flagged py/overwritten-inherited-attribute at
ui_tests/test_v2_workflow_run_card.py L252: RunHarness.__init__ assigned
self.streams = 100 after NotificationApi.__init__ had already set it to 0.

NotificationApi and the bell Harness now take a keyword-only streams
argument (default 0), and RunHarness passes streams=100 to super().__init__.
Every harness the two suites build has the same attributes, values and
attribute order as before; the bell suite's own harness still starts at 0.

Test-only change; VERSION stays 0.261.251.

Refs microsoft#1546, microsoft#1543

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@paullizer
Paul Lizer (paullizer) marked this pull request as ready for review October 6, 2026 11:31
@paullizer
Paul Lizer (paullizer) merged commit 6bd12cf into microsoft:paullizer-react-v2-ui Oct 6, 2026
11 checks passed
Paul Lizer (paullizer) added a commit to paullizer/simplechat that referenced this pull request Oct 6, 2026
Bring in origin/paullizer-react-v2-ui at 6bd12cf, the merge of microsoft#1639
(0.261.251), so microsoft#1647 lands second as agreed.

Only two files conflicted:
- config.py keeps VERSION = "0.261.252", one above V2.
- release_notes.md keeps the v0.261.252 section directly above V2's
  v0.261.251 section, with every section byte-exact.

Every other file matches the side that changed it.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Paul Lizer (paullizer) added a commit that referenced this pull request Oct 6, 2026
…w-up

- Status date 2026-10-06, with paullizer-react-v2-ui at 0.261.252.
- Phase 6: 6b-2 done in #1639 (v0.261.251), with how it shipped.
- Phase 7: 7a done in #1640 (v0.261.250) and its lineage follow-up in #1647
  (v0.261.252); 7b, the V2 hand-off card, is next. Admins should leave the
  hand-off setting off until 7b ships.
- Section 9: the run deep-link and one-time workflow decisions are settled.
- Release-notes integrity follow-up updated to 187 missing releases.

Docs only; no version bump.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants