Skip to content

feat(ui): improve research navigation and unify site styling - #667

Open
lynnzc wants to merge 9 commits into
ktwu01:mainfrom
lynnzc:feat/research-explorer-ui
Open

lynnzc wants to merge 9 commits into
ktwu01:mainfrom
lynnzc:feat/research-explorer-ui

Conversation

@lynnzc

@lynnzc lynnzc commented Sep 19, 2026 •

Copy link
Copy Markdown

Summary

Improve the existing site's navigation, typography, spacing, and controls to make benchmark research easier to browse.

  • Apply shared styling to the dashboard, articles, and benchmark pages using the original header and radar logo.
  • Add a research field browser at /research/, with two-level filters, source evidence, and export, through the normal site build.
  • Carry field selection between Frontier and score history while keeping benchmark search open to the full catalog.
  • Preserve existing routes and use the canonical index with on-demand detail loading.

Validation

  • Full CI sequence passed from a clean checkout with submodules initialized: lint, formatting, catalog normalization, classification, data release, and 1,387 tests.
  • 18 interaction checks passed for navigation, field filters, search, evidence details, and language switching.

lynnzc commented Sep 19, 2026

Copy link
Copy Markdown
Author

330226 — GPT

This PR was prepared by a GPT-series coding agent on behalf of @lynnzc, following AGENTS.md and CONTRIBUTING.md. The complete clean-checkout CI sequence passed locally (1,385 tests), along with 22 UI checks. The uploaded head has the same Git tree as the tested checkout. Desktop and mobile visual review remains outstanding; the proposal is opt-in and does not change the production deployment.

Apply the visual updates directly to the existing site and page generators.
Add catalog-backed research fields, preserve unrestricted benchmark search,
and cover native navigation and filtering in the regression harnesses.
@lynnzc lynnzc changed the title feat(ui): propose a research explorer and consistent whole-site design feat(ui): improve research navigation and unify site styling Sep 19, 2026
@ktwu01

ktwu01 commented Sep 21, 2026

Copy link
Copy Markdown
Owner

Checked corpus compliance: the default field "all" returns the full index
(site/assets/fields.js, matchesField), untagged records land in "other"
rather than being dropped, and Saturation search bypasses both the cutoff and
the field filter, which matches principle.md. The gitignored /research/
page and the on-demand shard loading look right, and no site/data/*.json was
hand-edited.

Three asks:

  1. src/benchmark_radar/site_pages.py:745 and :749 read
    Path("site/index.html") and Path("site/assets/app.js") CWD-relative
    while the same function writes to staging / output_dir (:735, :771).
    These reads are unconditional, so a caller invoking this from any other
    working directory gets a FileNotFoundError, not a degraded page. Resolve
    them against the repo root, or document the constraint at the signature.
  2. research.js:111 renders r.metric and r.direction in the Record
    metadata panel, but research_page.py never emits either field, so every
    record shows "Metric not recorded / Direction unknown". The fallbacks keep
    it from breaking, which is also what makes it easy to miss. Either populate
    both from the shard or drop the panel.
  3. Please complete the outstanding desktop and mobile visual review and confirm
    the full six-step run on a clean checkout.

Verdict: Comment. Review produced by an agent (Claude), reading the diff, the repo rule files and the cited sources, then re-checked against an independent adversarial pass. Posted as a comment rather than a blocking review.

lynnzc commented Sep 23, 2026 •

Copy link
Copy Markdown
Author

Addressed both code findings: generated benchmark pages resolve shared chrome from the repository path, with a regression test for a non-root working directory; the Research details dialog no longer renders unsupported record metadata.

The repository's six-step CI sequence passed in a clean detached worktree with submodules initialized (1,387 tests). Visually checked Today and Research at desktop (1440 × 900) and mobile (390 × 844), including Research table rows and the details dialog. At both sizes, clicking the row opens the dialog, and the page has no horizontal overflow. The public preview serves the verified Research page and byte-identical JS/CSS assets: https://benchmark-radar-study.lynnzc.chatgpt.site/research/.

ktwu01 commented Oct 3, 2026

Copy link
Copy Markdown
Owner

Hi @lynnzc, thanks for this. The research navigation and a consistent visual system are both things the site needs. The PR is hard to review as one unit, though: 26 files / +1,112, combining two independent changes:

  1. A site-wide restyle: design-system.css (+545), the edits to styles.css/blog.css/logos.css, the header in index.html, and the new design.md section.
  2. The research page: research.js/research.css/fields.js, research_page.py, templates/research.html, site_pages.py, the .gitignore entry for /site/research/, and its tests.

Could you split these into two PRs, with the restyle first? A reviewer can then judge the restyle from before/after screenshots (desktop and phone width, light and dark), and judge the research page on behavior. Please attach those screenshots to the restyle PR; for a visual change they are the main thing reviewers look at.

Smaller items:

  • Three render harnesses (render_briefing, render_questions, render_stale_banner) each get the same fields.js inlining boilerplate. A shared helper would keep that to one place.
  • test_saturation_view.py now asserts the literal "matches ? state.benchmarkIndex : researchRecords()". That pins one line of source text rather than behavior. A harness check that saturation search still covers the full index would be sturdier, and principle.md's full-corpus rule is what is at stake there.
  • The branch is 70 commits behind main; please rebase before splitting.

Generated by Claude Code

@ktwu01

ktwu01 commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

@Claire1217 our cheif UI officer what do u think of this?
The preview is at
https://benchmark-radar-study.lynnzc.chatgpt.site/

ktwu01 commented Oct 7, 2026

Copy link
Copy Markdown
Owner

Review: valuable direction, but too big and out of date to merge as one PR. It touches 26 files (about 1,100 added lines) across design system CSS, app.js, blog_shell.py, site_pages.py, snapshots.py and tests. It is 200 commits behind main and now conflicts in site/assets/app.js (and tests/test_cli.py needs a look). No CI check has run on the current head.

To get it merged:

  1. Split it. One PR for the research explorer (research.js, research.css, the page and its tests), one for the shared design-system CSS, and test/housekeeping changes (.gitignore, test_cli.py) separately. Each can then be reviewed and screenshotted on its own.
  2. Rebase each onto current main, resolve app.js, and run the full CI sequence from AGENTS.md.
  3. Add before/after screenshots of every page whose look changes (UI rule in AGENTS.md), and keep the existing branding, colors and logo.
  4. Per principle.md, the research explorer must start from the full corpus; state the record count it shows.

Reminder from #510: UI changes are merged on design quality, so please keep the first screen simple.

中文:方向很好,但一次改 26 个文件且落后 main 200 个提交,app.js 有冲突。请拆成「研究页」「共享样式」「测试/杂项」几个小 PR,各自 rebase、跑完整 CI、附截图,并说明研究页显示的记录数(principle.md)。


Generated by Claude Code

@ktwu01

ktwu01 commented Oct 8, 2026

Copy link
Copy Markdown
Owner

@Claire1217 our cheif UI officer what do u think of this?
The preview is at
https://benchmark-radar-study.lynnzc.chatgpt.site/

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants