Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
# Conflicts: # test/statusline-cli.test.ts
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 846aa8c597
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| const status = | ||
| testOutcome === 'success' && summary?.status === 'success' | ||
| ? skipped.length > 0 | ||
| ? 'passed with skips' | ||
| : 'passed' |
There was a problem hiding this comment.
Include Feishu failures in the overall summary status
When live-test secrets are present and the Feishu step fails, .github/workflows/midscene-e2e.yml still supplies the successful Dashboard outcome as testOutcome and only the Dashboard summary as summary. Because this status calculation ignores feishuOutcome, the Actions summary reports “passed” with zero failures even though the job's 17 live scenarios failed; incorporate the Feishu outcome and its result summary into the overall status and counts.
Useful? React with 👍 / 👎.
| const skippedCases = await readSkippedCases(skippedCasesDir); | ||
| if (passedCases.length === 0 || skippedCases.length === 0) { | ||
| throw new Error('Pages report requires passed Dashboard and skipped Feishu cases'); |
There was a problem hiding this comment.
Publish actual Feishu results when live cases run
On a manual workflow run with Feishu credentials available, this code still reads every Feishu YAML case into skippedCases, and landingPage consequently hard-codes all 17 rows as skipped with zero failures. The workflow invokes this path regardless of whether the live project ran, so the persistent GitHub Pages report contradicts successful or failed live results; skipped rows should only be synthesized when the Feishu project was actually skipped, otherwise its report data must be included.
Useful? React with 👍 / 👎.
| } finally { | ||
| if (setupCompleted) { | ||
| for (const hook of [...suite.afterAll].reverse()) await hook(); | ||
| } |
There was a problem hiding this comment.
Run teardown after partially failed setup
If a beforeAll hook fails after allocating resources—for example, createBrowser() succeeds but page creation or initial Feishu navigation fails—setupCompleted remains false and the suite's afterAll never closes that browser, context, or agent. Since CI retries cases in the same runner, these leaked browser processes can accumulate and destabilize later scenarios; teardown should run after any setup attempt, with the existing optional/defensive cleanup handling partially initialized state.
Useful? React with 👍 / 👎.
|
你好!这是 botmux 的自动评审流程:该 PR 的评审群已创建,欢迎点击链接加入飞书群查看/参与评审讨论(链接一年有效): 目前你还不在我们的自动拉群作者名单里,所以暂时无法自动把你拉进群。也可以把 GitHub 账号和飞书信息补录到名单文档:https://bytedance.larkoffice.com/wiki/WJ1nwWbtxi89erkNGNbcgkt9nUe ,补录后后续评审会自动拉你入群。感谢贡献! |
|
Thanks for the thorough follow-up and for reopening this in the main repo so the live workflow can actually run — the first same-repo run is executing now. This is an automated first-pass review; the final call rests with the maintainers. Two things still block, both reproducible statically: 1. The Actions Summary silently drops all 17 Feishu results whenever the live project runs
Reproduced locally with the repo's final script: workflow-exact arguments produce Suggested fix (mirrors the Dashboard project, which passes
The Pages half is now correct, by the way: 2. "credential-free Dashboard" wording is the opposite of the actual behavior
Non-blocking follow-ups
Housekeeping
|
Thanks for the quick follow-up ( What's fixed and verified ✅The 1. The Dashboard smoke is non-deterministic against the same input (flake)Run-2's Dashboard step failed in 47 s with
The model judged the same screenshot differently across attempts. Nothing in the Dashboard inputs changed between run-1 (105 s, all 7 assertions green) and run-2 — the two commits only touch the workflow, lockfile/package.json, and the Feishu registry, and the report-assembly code path is byte-identical to 1.13.0; 1.13.1's other core diffs are only version strings, webpack runtime, and an inert 2. The Feishu suite produced one slow pass, 15 failures, and one case that never ranWith CJK fonts installed and Midscene 1.13.1, only 1 of 17 cases passed — Blocker — a timed-out attempt's late afterAll closes the next attempt's browserThis is a functional regression introduced by the interaction of the unconditional teardown in
Evidence: (a) all five Fix: teardown must close the objects that attempt created, captured in attempt-local scope (or registered as per-attempt 3. The evidence pipeline still throws, and a broken partial site gets publishedBoth errors reproduced on the real run:
4. The summary path mismatch is still there
Smaller items
|
prepare 步骤可能在拷入部分目录后失败(如报告不可读),此时 deploy 仍会把缺 index.html/manifest.json 的残站推到 Pages,覆盖该 SITE_KEY 上一份好报告(run-2 实测:根 404、仅剩 dashboard 子树)。 在下载证据 artifact 后、推送前增加完整性校验,两个文件缺失即让发布 步骤失败、不执行 git push。 Co-Authored-By: Claude <noreply@anthropic.com>
CI 飞书 live 套件收窄为三个单聊机器人场景(Claude basic、Codex basic、 Codex prompt),它们只依赖账号能看到 Claude/Codex 单聊,不依赖多机器人 测试群,稳定性最高、所需鉴权最简。原 17 个用例文件保留,设 FEISHU_E2E_CASES=all 即可恢复全量(本地或后续专用测试群就绪后)。 Co-Authored-By: Claude <noreply@anthropic.com>
navigateToMessenger 的 waitForFunction 硬编码中文标题「消息 - 飞书」 与正文「搜索/消息」,但保存的账号在 CI 干净浏览器里渲染英文 UI (title 为 "Messenger - Feishu"、正文为 "Search/Messenger"), 导致三 个 Claude/Codex 用例每次都在 30s 就绪等待后误报 "Feishu Messenger did not finish loading"(页面其实已登录到 Messenger)。改为语言无关判据(title 匹配 messenger/飞书/lark, 正文匹配 Search/搜索 与 Messenger/消息),并保留对登录页的拒绝。 Co-Authored-By: Claude <noreply@anthropic.com>
保存的账号在 CI 干净浏览器渲染英文界面(筛选标签是 Topics/Topic chats,不是「话题」),而 openThreadForMessage 第 1 步的 aiAct 强制 要求点中文字面「话题」且明确排除「话题群」,视觉模型在英文界面正 确地拒绝点击,导致切不到话题列表、后续按 e2e marker 定位话题全部 超时(run-7 三个用例 6 次尝试均卡在话题相关步骤)。 把话题入口的 aiAct/aiWaitFor 改为中英双语(接受 话题/Topics, 用 index-card 图标与列位置描述,列出 Messages/@mentions/Unread/ Labels 等排除项),不再要求字面上的中文标签。后续按 marker 的 Playwright getByText 查找本就语言无关,无需改。 Co-Authored-By: Claude <noreply@anthropic.com>
真机 run-9 截图显示:测试账号同时存在飞书原生「codex/claude 智能体」 (DM 标题就是 codex/claude,带智能体徽标)和 botmux 接管的 「[Botmux]Codex/[Botmux]Claude 机器人」。openChat 按裸名 'Codex'/'Claude' 让视觉模型选会话时,点中的是原生智能体(右栏标题 "codex"),消息发过去无人回应——三个用例随后都在 openThreadForMessage 等不到话题条目而超时。 新增 botChatName() 把逻辑名解析成 [Botmux]<Name> 显示名,openChat 的点击/搜索/标题校验都改为精确匹配带 [Botmux] 前缀、带机器人徽标的 会话,并显式排除同名原生智能体。调用点仍传裸名(openChat 内部加 前缀),waitForModelTextReply 的 botName 仅用于自然语言描述不受影响。 Co-Authored-By: Claude <noreply@anthropic.com>
Summary
aiActfor intent-level UI interactions while retainingaiWaitForandaiAssertfor observable outcomesnot-runwhen a suite is interrupted@midscene/testand@midscene/webto 1.13.1CI configuration
The workflow uses these repository secrets:
MIDSCENE_MODEL_API_KEYMIDSCENE_MODEL_BASE_URLMIDSCENE_MODEL_NAMEMIDSCENE_MODEL_FAMILYMIDSCENE_MODEL_REASONING_ENABLEDFEISHU_TEST_GROUP_URLFEISHU_TEST_GROUP_CHAT_NAMEFEISHU_STORAGE_STATE_GZIP_BASE64PAGES_DEPLOY_KEYThe Feishu browser state is stored as a compressed, base64-encoded Playwright storage state. It is reused across CI runs and only needs to be refreshed if the Feishu session expires or is revoked.
The live fixture must be an account with access to the original Botmux test group and the Aiden, Claude, CoCo, Codex, and OpenCode conversations.
FEISHU_TEST_GROUP_URLmust be a direct group link. The current secret points to the Messenger home page, and the configured account cannot find those conversations. The workflow now fails quickly with this explicit configuration error instead of retrying 17 cases for hours. The live suite is not green until the test fixture is corrected.Reports
Each case row in the Actions summary includes a screenshot thumbnail and a link to the corresponding native Midscene report when evidence is available. Interrupted runs still publish completed case evidence and list unfinished cases instead of dropping the whole report.
Validation
actionlint .github/workflows/midscene-e2e.ymlpr-1512report path