Skip to content

refactor(e2e): migrate Feishu scenarios to Midscene Test - #1512

Open
quanru wants to merge 27 commits into
masterfrom
quanru/midscene-test
Open

quanru wants to merge 27 commits into
masterfrom
quanru/midscene-test

Conversation

@quanru

@quanru quanru commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • migrate the existing Feishu browser E2E scenarios to the Midscene Test project runner
  • keep the original scenario implementations and expose all 17 cases through native Midscene YAML cases
  • use aiAct for intent-level UI interactions while retaining aiWaitFor and aiAssert for observable outcomes
  • run the Dashboard and Feishu projects independently so one failure does not hide the other project evidence
  • publish a case-level evidence site with screenshots and direct native Midscene report links
  • preserve partial case reports and mark unfinished cases as not-run when a suite is interrupted
  • restore the full Feishu suite in upstream CI by loading the authenticated browser state from repository secrets
  • skip both test production and report deployment for fork PRs that cannot access upstream secrets
  • update @midscene/test and @midscene/web to 1.13.1

CI configuration

The workflow uses these repository secrets:

  • MIDSCENE_MODEL_API_KEY
  • MIDSCENE_MODEL_BASE_URL
  • MIDSCENE_MODEL_NAME
  • MIDSCENE_MODEL_FAMILY
  • MIDSCENE_MODEL_REASONING_ENABLED
  • FEISHU_TEST_GROUP_URL
  • FEISHU_TEST_GROUP_CHAT_NAME
  • FEISHU_STORAGE_STATE_GZIP_BASE64
  • PAGES_DEPLOY_KEY

The Feishu browser state is stored as a compressed, base64-encoded Playwright storage state. It is reused across CI runs and only needs to be refreshed if the Feishu session expires or is revoked.

The live fixture must be an account with access to the original Botmux test group and the Aiden, Claude, CoCo, Codex, and OpenCode conversations. FEISHU_TEST_GROUP_URL must be a direct group link. The current secret points to the Messenger home page, and the configured account cannot find those conversations. The workflow now fails quickly with this explicit configuration error instead of retrying 17 cases for hours. The live suite is not green until the test fixture is corrected.

Reports

Each case row in the Actions summary includes a screenshot thumbnail and a link to the corresponding native Midscene report when evidence is available. Interrupted runs still publish completed case evidence and list unfinished cases instead of dropping the whole report.

Validation

  • actionlint .github/workflows/midscene-e2e.yml
  • TypeScript validation for both Midscene projects
  • Midscene node generation for both projects
  • focused Vitest coverage for summary rendering, partial report-site preparation, suite registration, and Bun-compatible Kimi polling
  • frozen Bun install after merging the latest upstream master
  • local Midscene 1.13.1 Dashboard run passed all 7 steps with the upstream visual-model configuration
  • the latest upstream run uses the main repository Secrets and publishes to the pr-1512 report path
  • the fork example skips both Midscene jobs cleanly instead of attempting to download a report artifact

@quanru
quanru requested a review from deepcoldy as a code owner September 22, 2026 03:06
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-22T03:12:11.302464Z 846aa8c PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 846aa8c597

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +107 to +111
const status =
testOutcome === 'success' && summary?.status === 'success'
? skipped.length > 0
? 'passed with skips'
: 'passed'

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Include Feishu failures in the overall summary status

When live-test secrets are present and the Feishu step fails, .github/workflows/midscene-e2e.yml still supplies the successful Dashboard outcome as testOutcome and only the Dashboard summary as summary. Because this status calculation ignores feishuOutcome, the Actions summary reports “passed” with zero failures even though the job's 17 live scenarios failed; incorporate the Feishu outcome and its result summary into the overall status and counts.

Useful? React with 👍 / 👎.

Comment thread scripts/prepare-midscene-pages.mjs Outdated
Comment on lines +145 to +147
const skippedCases = await readSkippedCases(skippedCasesDir);
if (passedCases.length === 0 || skippedCases.length === 0) {
throw new Error('Pages report requires passed Dashboard and skipped Feishu cases');

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Publish actual Feishu results when live cases run

On a manual workflow run with Feishu credentials available, this code still reads every Feishu YAML case into skippedCases, and landingPage consequently hard-codes all 17 rows as skipped with zero failures. The workflow invokes this path regardless of whether the live project ran, so the persistent GitHub Pages report contradicts successful or failed live results; skipped rows should only be synthesized when the Feishu project was actually skipped, otherwise its report data must be included.

Useful? React with 👍 / 👎.

Comment thread test/e2e-browser/midscene-suite.ts Outdated
Comment on lines +199 to +202
} finally {
if (setupCompleted) {
for (const hook of [...suite.afterAll].reverse()) await hook();
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Run teardown after partially failed setup

If a beforeAll hook fails after allocating resources—for example, createBrowser() succeeds but page creation or initial Feishu navigation fails—setupCompleted remains false and the suite's afterAll never closes that browser, context, or agent. Since CI retries cases in the same runner, these leaked browser processes can accumulate and destabilize later scenarios; teardown should run after any setup attempt, with the existing optional/defensive cleanup handling partially initialized state.

Useful? React with 👍 / 👎.

@deepcoldy

Copy link
Copy Markdown
Owner

你好!这是 botmux 的自动评审流程:该 PR 的评审群已创建,欢迎点击链接加入飞书群查看/参与评审讨论(链接一年有效):
https://applink.larkoffice.com/client/chat/chatter/add_by_link?link_token=754pd398-5cce-4ee8-94c3-ffeb74eabee1

目前你还不在我们的自动拉群作者名单里,所以暂时无法自动把你拉进群。也可以把 GitHub 账号和飞书信息补录到名单文档:https://bytedance.larkoffice.com/wiki/WJ1nwWbtxi89erkNGNbcgkt9nUe ,补录后后续评审会自动拉你入群。感谢贡献!

@deepcoldy

Copy link
Copy Markdown
Owner

Thanks for the thorough follow-up and for reopening this in the main repo so the live workflow can actually run — the first same-repo run is executing now. This is an automated first-pass review; the final call rests with the maintainers. Two things still block, both reproducible statically:

1. The Actions Summary silently drops all 17 Feishu results whenever the live project runs

scripts/run-e2e.ts hard-codes the result directory (midscene_run/runs/<timestamp>, lines 16–17) and passes the same value to --result-dir, so the Feishu summary is written at midscene_run/runs/<timestamp>/<run-id>/summary.json. But the workflow points the summary renderer at a different path (.github/workflows/midscene-e2e.yml, --feishu-results midscene_run/feishu/results, line 192), which never exists. findSummaries() treats a missing root as zero summaries (ENOENT → []), so the step succeeds and the Summary reports only the Dashboard case — 1/1 passed, with zero Feishu rows, even when all 17 live cases pass.

Reproduced locally with the repo's final script: workflow-exact arguments produce 1/1 cases passed · 0 Feishu rows; pointing --feishu-results at midscene_run/runs immediately produces 18/18 with all 17 Feishu rows.

Suggested fix (mirrors the Dashboard project, which passes --result-dir midscene_run/dashboard/results directly):

  • make run-e2e.ts use a stable result directory (e.g. accept it from argv/env instead of hard-coding midscene_run/runs/<timestamp>; the timestamped <run-id> subdirectory is already created by the runner), and align the workflow's --feishu-results with it;
  • add tests for the "Feishu actually ran" path (both summary scripts only cover the skipped branch today) — the renderer should also treat feishu-outcome=success|failure with no loadable Feishu summary as a failure rather than silently rendering Dashboard alone (today the headline still says "passed" in that shape — also reproduced).

The Pages half is now correct, by the way: prepare-midscene-pages.mjs only synthesizes skipped rows when the project was skipped and loads the real Feishu report via --feishu-report-root otherwise; the running live job will exercise that path for the first time.

2. "credential-free Dashboard" wording is the opposite of the actual behavior

test/e2e-browser/README.md:25, the note in scripts/render-midscene-summary.mjs:181, and the workflow comment at line 149 describe the Dashboard project as credential-free / "always runs". In reality the workflow requires the four MIDSCENE_MODEL_* secrets and makes an unconditional live model preflight call (the "Require/Check Midscene model configuration" steps, lines 71–99), and the Dashboard smoke itself drives the UI through seven aiAct/aiAssert/aiWaitFor calls against the visual model. On fork PRs both Midscene jobs are skipped entirely (verified on the fork example run). Please reword (e.g. "runs without Feishu login state, but requires Midscene model configuration") — no code change needed for this one.

Non-blocking follow-ups

  • With retry: 1 and the 15-minute per-case timeout, the pathological worst case is 17×15×2 = 510 min > the 360-min job timeout; if the job is hard-cancelled, every if: always() && !cancelled() tail step (summary/artifacts/Pages) is skipped and no evidence is published. Splitting Dashboard and Feishu into separate jobs would bound this.
  • findReportIndexes only matches files literally named index.html; a single-file *.html report throws "No Midscene Test report found". 1.13 happens to emit directory mode today; accepting *.html would make this robust.
  • A clean install currently resolves tsx@4.23.15, which crashes while loading the runner config (Cannot find module tsx/dist/esm/api/esm/index.mjs); the lockfile's 4.23.12 is safe in CI. Consider pinning the repo's own tsx to ~4.23.14.
  • The third argument of the mini-registry's it()/beforeAll() (_timeout) is ignored, so the per-case timeouts in the scenario files are inert; the 360s/600s values in the factory are misleading.
  • getTopicGroupChatName() has no callers; the README secrets list omits FEISHU_STORAGE_STATE_GZIP_BASE64 and PAGES_DEPLOY_KEY; a couple of pnpm strings remain.

Housekeeping

@quanru quanru closed this Sep 22, 2026
@quanru quanru reopened this Sep 22, 2026
@deepcoldy

Copy link
Copy Markdown
Owner

Follow-up to the first-pass review above on head 5b86e5456.

Thanks for the quick follow-up (3e848187d + 5b86e5456). We re-ran the workflow on the new head (run 35720911848) and verified the fix end to end. The teardown fix works, but the run surfaced a Dashboard flake, an unreformed Feishu suite, a new teardown race, and reporting gaps. This is an automated review; the maintainers make the final call.

What's fixed and verified ✅

The runFeishuScenario teardown change removes the post-run hang. In run-1, after all Feishu cases failed the process produced no output for 102 minutes until the 6-hour job timeout killed it, and every tailing step was skipped with zero artifacts. In run-2 the Feishu step ended as a failure when its 300-minute step timeout was reached, and within ~20 seconds the evidence/summary/upload steps had all executed, artifacts were published (266 MB report data), and the job completed normally. Running afterAll unconditionally is the right fix and matches normal test-runner semantics; the optional-chaining in the cleanup hooks also handles a beforeAll failure that happens before the browser exists.

1. The Dashboard smoke is non-deterministic against the same input (flake)

Run-2's Dashboard step failed in 47 s with 0/1 passed, 1 failed, 0 not run — the 0 not run rules out a project-setup failure (the runner marks every case not-run when setup fails, and setup retries do not exist: the runner builds retry: { scope: 'case' } literally and starts setup exactly once). The case actually ran:

  • attempt 1: step 1/7 aiAssert ("the Dashboard has loaded…") failed after 4.3 s;
  • attempt 2: the same assertion passed after 3.5 s, steps 2–3 passed, then step 4/7 aiAssert ("Session Control…") failed after 8.2 s.

The model judged the same screenshot differently across attempts. Nothing in the Dashboard inputs changed between run-1 (105 s, all 7 assertions green) and run-2 — the two commits only touch the workflow, lockfile/package.json, and the Feishu registry, and the report-assembly code path is byte-identical to 1.13.0; 1.13.1's other core diffs are only version strings, webpack runtime, and an inert platform passthrough (...project.platform?.trim() ? {platform} : {}) that behaves identically because the Dashboard config sets no platform. Suggestion: make the first assertion more robust (e.g. a deterministic Playwright selector/waitForSelector for a known DOM node before the AI assertion, or attach the screenshot to the failure) rather than relying on the model's first judgment.

2. The Feishu suite produced one slow pass, 15 failures, and one case that never ran

With CJK fonts installed and Midscene 1.13.1, only 1 of 17 cases passed — Bot reply smoke test, and only on its second attempt after 853 s (its first attempt failed at 452 s). The other 15 cases failed both attempts (31 failed attempts out of 32), and case 17 ("Web terminal opens from the streaming card") had only just started when the 300-minute step timeout killed the runner, so it received no verdict and the runner never assembled its final report. Failure modes across the 32 attempts: Replanned 20 times, waitFor timeouts, plus a new pattern discussed below. The visual agent effectively still cannot complete Feishu messenger tasks in CI, and the serial 17-case suite with one retry does not fit even a 300-minute step budget. This needs a maintainer decision before enabling on PRs: (a) tune model/navigation until a full green run exists; (b) cut CI to a small reliably-passing smoke subset; or (c) keep Feishu workflow_dispatch-only and land the Dashboard project (with its flake addressed) first.

Blocker — a timed-out attempt's late afterAll closes the next attempt's browser

This is a functional regression introduced by the interaction of the unconditional teardown in 3e848187d with the runner's timeout handling, and it reproduces deterministically. The chain:

  1. The browser/page/agent in each scenario are shared closure-level let variables inside describe() (verified in all five affected files). The registry caches one suite object per module and reuses that closure across a case's attempts.
  2. When the node hits the 15-minute timeout, the runner aborts the signal and settles via Promise.race without waiting for the underlying node promise — the retry loop's await run() resolves on the race rejection while the timed-out runFeishuScenario is still in its finally. The next attempt starts ~45–110 ms later and its beforeAll reassigns the shared browser/page to new objects.
  3. Attempt 1's now-late afterAll reads those same closure variables and calls browser?.close()/context?.close() — closing attempt 2's browser. Attempt 2's next keyboard.press/screenshot then fails with Target page, context or browser has been closed.

Evidence: (a) all five keyboard … has been closed failures occur only as attempt 2 of a case whose attempt 1 hit the 900 s timeout (collaboration 29 s, card lifecycle 40 s, consecutive 187 s, group topic 40 s, private topic 54 s); no such error occurs after non-timeout failures, and run-1 (no 900 s timeouts) had zero; (b) a minimal deterministic harness modeling the shared closure variable, unconditional finally, abort-via-race, and immediate retry shows attempt 2's browser reported closed after attempt 1's teardown lands; (c) the six chrome-headless-shell orphans reaped at job teardown are consistent with this: five B1 browsers left unclosed by a1's late teardown (it closed B2 instead of its own B1), plus one browser from case 17 — the only case with fewer than two attempts — killed by the 300-minute SIGTERM before any afterAll; a per-attempt-scoped teardown would leave only that single SIGTERM orphan. We note the orphan count alone does not distinguish this from the alternative where an aborted a1 never reaches afterAll (both yield six) — the timing correlation in (a) is what separates them; run-1, under the old gated teardown and with no 900 s timeouts, left eight orphans from the different beforeAll-failure path. The one thing we could not capture is a runtime stack naming attempt 1's close() (teardown logs as success), hence framing it as confirmed-by-code-and-repro rather than by a single stack trace.

Fix: teardown must close the objects that attempt created, captured in attempt-local scope (or registered as per-attempt onTeardown callbacks), not the shared closure variables; and/or the runner should not start a retry until the previous node's promise (including finally) has settled. A cheap experiment is to scope browser/page/agent per attempt and re-run — the 900 s + has been closed pattern should disappear together.

3. The evidence pipeline still throws, and a broken partial site gets published

Both errors reproduced on the real run:

  • Prepare linked Midscene evidence site failed in ~1 s with Error: No Midscene Test report found for project feishu-browser (scripts/prepare-midscene-pages.mjs:111), and Write Midscene run summary then failed with ENOENT …/midscene-site/manifest.json.
  • On this run the immediate cause is one layer earlier: the runner was terminated during case 17, so the final assembled Test report for the Feishu project was never written. When the suite can finish, the underlying defect from the earlier comment still applies — the Feishu project emits a single-file report (midscene-e2e-<id>.html, none of its YAML cases persist a dump like the Dashboard's recordToReport), while findReportIndexes only matches files literally named index.html. The report-assembly code (@midscene/core's report writer) is byte-identical in 1.13.1. So: accept *.html in findReportIndexes (or configure the Feishu project for directory-mode reports), and make the prepare step tolerate a missing/incomplete Feishu report without dropping the Dashboard evidence or the manifest.
  • Pre-existing design gap, exposed for the first time here: the deploy-report job's condition only excludes cancelled/skipped, so on this failure run it published the prepare step's partial output (that gate existed before these two commits — the earlier runs were cancelled before they could reach deployment). The live site at https://deepcoldy.github.io/botmux-midscene/manual/ currently has no index.html or manifest.json at the root (HTTP 404) — only the copied dashboard/index.html subtree (200). On workflow_dispatch the site key is always manual, so such a partial publish also overwrites any manually produced report. Please gate deployment on the prepare step succeeding (or on manifest.json existing).

4. The summary path mismatch is still there

scripts/run-e2e.ts still writes to midscene_run/runs/<timestamp>/… (the run-2 logs show the agent reports under midscene_run/runs/2026-09-22_11-24-40/), while the workflow still passes --feishu-results midscene_run/feishu/results, which never exists. On this run the summary step failed on the missing manifest first, but the Feishu rows would still be silently dropped once the evidence step is fixed; the headline also still renders "passed" when a Feishu failure comes with no loadable summary. Please align the directories (mirroring the Dashboard project) and add a test for the "Feishu actually ran" path.

Smaller items

  • The six orphaned chrome-headless-shell processes at job teardown are accounted for in §2(c) (five from the timeout/teardown race, one from the SIGTERM kill); the normal-failure path is otherwise clean.
  • tsx ~4.23.14 pin, the ignored third argument in the registry's it()/beforeAll(), getTopicGroupChatName() with no callers, README secrets list, and residual pnpm strings remain as noted before.
  • The branch is 13 commits behind master and currently conflicts (the statusline and Kimi commits are superseded upstream and drop on rebase).

quanru and others added 4 commits September 23, 2026 15:03
prepare 步骤可能在拷入部分目录后失败(如报告不可读),此时 deploy
仍会把缺 index.html/manifest.json 的残站推到 Pages,覆盖该 SITE_KEY
上一份好报告(run-2 实测:根 404、仅剩 dashboard 子树)。

在下载证据 artifact 后、推送前增加完整性校验,两个文件缺失即让发布
步骤失败、不执行 git push。

Co-Authored-By: Claude <noreply@anthropic.com>
CI 飞书 live 套件收窄为三个单聊机器人场景(Claude basic、Codex basic、
Codex prompt),它们只依赖账号能看到 Claude/Codex 单聊,不依赖多机器人
测试群,稳定性最高、所需鉴权最简。原 17 个用例文件保留,设
FEISHU_E2E_CASES=all 即可恢复全量(本地或后续专用测试群就绪后)。

Co-Authored-By: Claude <noreply@anthropic.com>
navigateToMessenger 的 waitForFunction 硬编码中文标题「消息 - 飞书」
与正文「搜索/消息」,但保存的账号在 CI 干净浏览器里渲染英文 UI
(title 为 "Messenger - Feishu"、正文为 "Search/Messenger"),
导致三 个 Claude/Codex 用例每次都在 30s 就绪等待后误报
"Feishu Messenger did not finish loading"(页面其实已登录到
Messenger)。改为语言无关判据(title 匹配 messenger/飞书/lark,
正文匹配 Search/搜索 与 Messenger/消息),并保留对登录页的拒绝。

Co-Authored-By: Claude <noreply@anthropic.com>
deepcoldy and others added 2 commits September 24, 2026 10:21
保存的账号在 CI 干净浏览器渲染英文界面(筛选标签是 Topics/Topic
chats,不是「话题」),而 openThreadForMessage 第 1 步的 aiAct 强制
要求点中文字面「话题」且明确排除「话题群」,视觉模型在英文界面正
确地拒绝点击,导致切不到话题列表、后续按 e2e marker 定位话题全部
超时(run-7 三个用例 6 次尝试均卡在话题相关步骤)。

把话题入口的 aiAct/aiWaitFor 改为中英双语(接受 话题/Topics,
用 index-card 图标与列位置描述,列出 Messages/@mentions/Unread/
Labels 等排除项),不再要求字面上的中文标签。后续按 marker 的
Playwright getByText 查找本就语言无关,无需改。

Co-Authored-By: Claude <noreply@anthropic.com>
真机 run-9 截图显示:测试账号同时存在飞书原生「codex/claude 智能体」
(DM 标题就是 codex/claude,带智能体徽标)和 botmux 接管的
「[Botmux]Codex/[Botmux]Claude 机器人」。openChat 按裸名
'Codex'/'Claude' 让视觉模型选会话时,点中的是原生智能体(右栏标题
"codex"),消息发过去无人回应——三个用例随后都在
openThreadForMessage 等不到话题条目而超时。

新增 botChatName() 把逻辑名解析成 [Botmux]<Name> 显示名,openChat
的点击/搜索/标题校验都改为精确匹配带 [Botmux] 前缀、带机器人徽标的
会话,并显式排除同名原生智能体。调用点仍传裸名(openChat 内部加
前缀),waitForModelTextReply 的 botName 仅用于自然语言描述不受影响。

Co-Authored-By: Claude <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants