Skip to content

Improve agent documentation - #3812

Merged
ethanglaser merged 12 commits into
uxlfoundation:mainfrom
ethanglaser:dev/eglaser-agent-rules
Oct 1, 2026
Merged

ethanglaser merged 12 commits into
uxlfoundation:mainfrom
ethanglaser:dev/eglaser-agent-rules

Conversation

@ethanglaser

@ethanglaser ethanglaser commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Description

1. Instructions that were wrong (napetrov's #3780)

  • dev/bazel/AGENTS.md (+47/−96): replaced the bare cc_library/cc_test and copts examples with the macros the repo actually uses: dal_module, dal_test_suite, daal_module. It also now has correct dependency labels and a note that --config must be passed. The reference target is cpp/oneapi/dal/algo/pca/BUILD.
  • dev/AGENTS.md, cpp/oneapi/AGENTS.md: replaced the unsupported bazel test //... with scoped targets using --config=host or --config=dpc. Make commands now use the real make -f makefile ... PLAT= form. The CMake recipes now build the examples against an installed release, since there is no root CMakeLists.txt.
  • Also removed or fixed, from review on this PR: the bare Bazel pattern in dev/AGENTS.md, the daal_module example that passed cpu_defines and features (the macro sets both), the Bazel version and macOS claims, TBB lines that contradicted the threading-layer rule, fake names (KMeansBatch, MAX_ITERATIONS), and snippets in cpp/daal and cpp/oneapi that didn't compile (missing =, depends_on inside the kernel, the SPMD call, a device pull without a queue).
  • examples/AGENTS.md: replaced a per-algorithm dal_example_suite target that doesn't exist with the real dal_algo_example_suite(algos = [...]) registration.
  • CONTRIBUTING.md: fixed the clang-format command (it was missing a -) and replaced the link to a root .clang-format file that doesn't exist.

2. Code search (#3781)

3. Rules taken from review comments

  • Root AGENTS.md: a "Rules for Changes" section covering:
    • comments describe the merged code
    • short comments
    • one change per PR
    • search for an existing helper first
    • no unnecessary locks or guards
    • a regression test with every bug fix
    • no hardcoded versions
    • the copyright header for new files
    • ASCII only
    • POSIX sh and .bat portability
  • cpp/daal/AGENTS.md:
    • .i files take a CpuType template parameter
    • use TArray and aligned allocation
    • no magic numbers
    • the zero-init-and-sum-back accuracy pattern
  • cpp/oneapi/AGENTS.md: the namespace vN plus using re-export pattern, used in 472 namespace v1 blocks and previously undocumented.
  • cpp/AGENTS.md: a new section on ABI and the public API.
  • New dev/make/AGENTS.md: makefiles and .mk, .bat and .sh scripts drew 61 human review threads across the last 200 PRs and had no guidance file.
  • .ci/AGENTS.md: CI-specific rules.

4. Trimming

  • Root AGENTS.md: the rules and verification commands come first, within the first 4K characters. Also removed emoji headers, general C++ advice (including a smart-pointer rule that contradicts the TArray rule), the quick start, and web links.
  • cpp/AGENTS.md and dev/AGENTS.md rewritten around what the repo actually does: an interface-conventions table, a review checklist, the real train_ops_dispatcher signature, and verified Make PLAT/COMPILER/REQCPU/BACKEND_CONFIG values. Invented snippets and emoji headers are gone.
  • docs/AGENTS.md (7.9K to 1.1K) and examples/AGENTS.md (6.7K to 1.5K): cut to layout, rules and verification commands. New rules: Sphinx runs with -W, generated source/examples/ RST is not committed, and renaming an example breaks its docs references.

5. Drop the Copilot instruction files

  • Removed .github/copilot-instructions.md and .github/instructions/*. Copilot code review and the coding agent both read AGENTS.md, so they only duplicated it.
  • Content not already in an AGENTS.md was merged into the nearest one. Review-critical rules (interface separation, API/ABI changes, what CI already enforces) are near the top of the root AGENTS.md, because that is the file Copilot code review is documented to read.
  • .licenserc.yaml: dropped the two exemptions for the removed files.

Checklist:

Completeness and readability

  • I have commented my code, particularly in hard-to-understand areas.
  • I have updated the documentation to reflect the changes or created a separate PR with updates and provided its number in the description, if necessary.
  • Git commit message contains an appropriate signed-off-by string (see CONTRIBUTING.md for details).
  • I have resolved any merge conflicts that might occur with the base branch.

Testing

  • I have run it locally and tested the changes extensively.
  • All CI jobs are green or I have provided justification why they aren't.
  • I have extended testing suite if new functionality was introduced in this PR.

Performance

  • I have measured performance for affected algorithms using scikit-learn_bench and provided at least a summary table with measured data, if performance change is expected.
  • I have provided justification why performance and/or quality metrics have changed or why changes are not expected.
  • I have extended the benchmarking suite and provided a corresponding scikit-learn_bench PR if new measurable functionality was introduced in this PR.

🤖 Generated with Claude Code

ethanglaser and others added 6 commits September 23, 2026 16:07
…#3780)

- Fix the .ci/AGENTS.md link, the root tree, the C++ standard and the
  nonexistent root .clang-format references.
- Add a "Verification Before You Push" section with real commands and
  the current CI layout.
- Replace the cc_library/cc_test Bazel template with dal_module /
  dal_test_suite, fix the MODULE.bazel description, drop //... and
  label CPU-only commands with --config=host.
- Replace the nonexistent association_rules target and root CMake recipe;
  unwrap backticks inside bash fences.
- Fix the clang-format command in CONTRIBUTING.md and drop the Mergify /
  Codefactor claims.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The pre-commit clang-format hook now checks the same directories and
  extensions as .ci/scripts/clang-format.sh, including the 152 .i kernel
  files, which identify assigns no type tags. It no longer touches
  deploy/ and dev/ sources that CI does not check.
- clang-format.sh uses --dry-run --Werror instead of rewriting the tree
  and inferring failure from git status, so it is safe to run locally
  and unrelated uncommitted changes no longer fail it.
- Mark *.i as C++ for linguist.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ctions

Add "Rules for Changes" sections distilled from recurring maintainer
review comments: cross-cutting rules and shell/batch portability in the
root file, DAAL kernel rules (.i navigation, CpuType dispatch, TArray,
accumulation, zero-division), the oneAPI versioned-namespace re-export,
ABI and export parity, Bazel platform scoping and hygiene, and a new
dev/make/AGENTS.md.

Add .github/copilot-instructions.md, which routes Copilot to the
AGENTS.md files and sets review posture, and extend path instructions
to .i and .bzl files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Drop generic advice, the quick-start steps, web links and emoji headers.
Drop the smart-pointer rule, which conflicts with TArray in cpp/daal.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ethanglaser ethanglaser changed the title Dev/eglaser agent rules Improve agent documentation Sep 25, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Address the formatter error-handling issue and remaining documentation inconsistencies.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 4 Low severity

Open (4)
What changed in this PR

This PR updates agent documentation, build guidance, formatting checks, and Copilot configuration.

Changes:

  • Corrects Bazel, Make, CMake, and formatting instructions.
  • Adds repository and subsystem-specific guidance.
  • Aligns local and CI formatting checks, including .i files.
  • Adds Copilot instruction wiring.
File Summary
dev/​make/​AGENTS.md Adds Make guidance.
dev/​bazel/​AGENTS.md Documents Bazel macros and workflows.
dev/​AGENTS.md Updates development guidance.
cpp/​oneapi/​AGENTS.md Documents oneAPI conventions.
cpp/​daal/​AGENTS.md Adds DAAL implementation rules.
cpp/​AGENTS.md Documents ABI and public API constraints.
CONTRIBUTING.md Corrects formatting instructions.
AGENTS.md Adds repository-wide agent rules.
.pre-commit-config.yaml Expands formatting hook coverage.
.github/​instructions/​general.instructions.md Streamlines general instructions.
.github/​instructions/​examples.instructions.md Updates example guidance.
.github/​instructions/​documentation.instructions.md Updates documentation guidance.
.github/​instructions/​cpp-coding-guidelines.instructions.md Updates C++ guidance.
.github/​instructions/​build-systems.instructions.md Updates build-system instructions.
.github/​instructions/​AGENTS.md Adds instruction-file maintenance rules.
.github/​copilot-instructions.md Adds the Copilot entry point.
.github/​.licenserc.yaml Exempts Copilot instructions from headers.
.gitattributes Classifies .i files as C++.
.ci/​scripts/​clang-format.sh Makes formatting checks non-mutating.
.ci/​AGENTS.md Adds CI-specific guidance.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread cpp/oneapi/AGENTS.md
Comment thread dev/bazel/AGENTS.md Outdated
Comment thread dev/bazel/AGENTS.md Outdated
Comment thread dev/bazel/AGENTS.md Outdated
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: ethanglaser <ethan.glaser@intel.com>
@Vika-F

Vika-F commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

@ethanglaser Thank you for this effort!

I'd approve it with the following comment: Can you please remove .pre-commit-config.yaml from this PR?
Because the version in this PR duplicates the info like the style flag and the file extensions in clang-format.sh and in .pre-commit-config.yaml files.
PR 3813 is created to make clang-format.sh a single source of truth about the formatting.

@ethanglaser

Copy link
Copy Markdown
Contributor Author

@ethanglaser Thank you for this effort!

I'd approve it with the following comment: Can you please remove .pre-commit-config.yaml from this PR? Because the version in this PR duplicates the info like the style flag and the file extensions in clang-format.sh and in .pre-commit-config.yaml files. PR 3813 is created to make clang-format.sh a single source of truth about the formatting.

Done

@napetrov

Copy link
Copy Markdown
Contributor

I measured the agent guidance in this PR at base 3488b58 and head e3804fb. The PR has since moved to 8d0a9ce, which drops the formatter changes in favour of #3813; that commit changes one guidance line (root AGENTS.md:72), which I re-checked and which is correct. Everything below is a recommendation, not a blocker. Each claim is tagged with how it was established: [mechanical] means checked deterministically against the tree, [LLM-graded] means measured with probes.

TL;DR: The rewrite makes the guidance much more correct (checkable-claim failures drop from 24% to 3%), and it lost no real knowledge: most of what was deleted was wrong or not practised. What remains to do:

  • fix about 11 remaining wrong lines, which measurably mislead models;
  • cut or move the content models already know;
  • put the novel rules first in the files Copilot code review reads.

What the PR got right

  1. The guidance is much more correct. Of the checkable claims (commands, Bazel labels, paths, links, identifiers), 71/291 failed at base and 12/397 fail at head (0.244 → 0.030). Every new command and label verifies: --config=host|dpc, //cpp/oneapi/dal:tests, //:release, @mkl//:mkl_core, and so on. [mechanical]

  2. Routing improves. applyTo now reaches the 152 .i, 35 .bzl and 19 .tpl.BUILD files and .bazelrc. Before, only general matched them. [mechanical]

  3. The root AGENTS.md carries more new knowledge per token for all 3 models (added per 1K tokens):

    • claude-sonnet-5 43→57;
    • kimi-k3 52→69;
    • glm-5 68→74.

    On routing, negative-constraint and version-bound rules, added goes 27→64, 29→69 and 38→56 (n 46→86). [LLM-graded; judge agreement 0.87 lenient, n=450]

  4. Most of the deleted content was wrong or not practised. Of the base atoms that ≥2 models learned from documentation, examples or general, and that are now absent from all head guidance, most are:

    • stale versions;
    • bazel test //...;
    • nonexistent names;
    • invented conventions: for example, @par Thread Safety / @par Exception Safety appear in 0 headers, and code-block:: cpp appears in 7 of 362 rst files.

    Deleting them is correct. [retention LLM-graded; examples mechanical]

  5. Less interference when all instruction files load together. Loading the whole bundle, compared with only the leaf's own chain, significantly lowered leaf answers in 3/9 (file, model) cells at base and in 1/12 at head. Every leaf file still adds knowledge over its parents: +0.24…+0.64 in 21/21 cells, with every CI excluding 0. [LLM-graded, 95% CI]

  6. Wrong guidance now hurts less. On reversed-premise probes (a false premise in the question), the number of answers made worse by loading the guidance fell for every model: sonnet 6→3, kimi 20→14, glm 26→19 (n 616→545). [LLM-graded]

Recommended fixes

  1. Remaining wrong lines. These matter because head atoms citing a mechanically wrong line got worse under a reversed premise far more often: kimi 8/37 vs 6/508 (Fisher p 5e-7), glm 5/37 vs 14/508 (p 0.006). They also teach models false facts: 17 of the 163 "novel" head atoms (below) come from these lines. [mechanical + LLM-graded]
    • dev/AGENTS.md:79-93 has bare cc_library/cc_test with //dev/bazel/deps:gtest and //path/to:dependency. This is the pattern the PR fixed in dev/bazel.
    • build-systems.instructions.md:100-110: daal_module already passes cpu_defines, so passing it again is a duplicate keyword (static read of daal.bzl; not executed).
    • cpp/daal/AGENTS.md:17,20,23 still has backticks inside a bash fence.
    • dev/AGENTS.md:117 says "Bazel 5.0+", but .bazelversion is 9.2.0.
    • dev/bazel/AGENTS.md:13 claims macOS support, but there is no macOS toolchain.
    • Some snippets would not compile:
      • cpp/daal/AGENTS.md:53 has using dm = where it should be namespace dm =;
      • cpp/daal/AGENTS.md:120-123;
      • cpp/oneapi/AGENTS.md:81,86,124.
    • cpp/daal/AGENTS.md:10 and cpp/AGENTS.md:97,153 describe TBB-based threading, which contradicts the new "never use TBB directly" rule.
    • cpp-coding-guidelines:158-171 presents names that do not exist (KMeansBatch, MAX_ITERATIONS, …) as codebase conventions. Seven of the "novel" atoms there are these made-up examples.
    • The root rule set -euo pipefail is followed by only 4 of 22 bash scripts. Either soften it to "new scripts" or accept that it doesn't describe the current tree.
  2. Consider restoring the CodeFactor line in .ci/AGENTS.md. CodeFactor runs and passes on this PR, so AGENTS.md / Copilot instructions document a Bazel API that does not exist (17 verified defects) + no reachable lint command #3780 was wrong on that point. [verified on this PR's checks]
  3. Put novel content first in the files Copilot code review reads.
    • Until 2026-06-15, GitHub's docs said code review "only reads the first 4,000 characters of any custom instruction file" (this did not apply to Chat or the cloud agent). github/docs commit f09467630 removed the note without giving a reason, and I have not tested current behaviour.
    • If the limit still holds:
      • Root AGENTS.md (5,923 chars) is cut at line 62. The cut drops 13 of its 29 novel atoms, i.e. "Verification Before You Push" and part of "Rules for Changes". The first 4K holds all 15 of the atoms every model already knows ("Repository Structure", the intro).
      • cpp-coding-guidelines (8,434 chars) is cut at line 118. The cut drops 15 of its 21 novel atoms: error handling, naming, templates and the PR review checklist. The first 4K is mostly "Interface-Specific Development Patterns", which is 10 general / 6 partial / 3 novel.
      • build-systems (3,969 chars) fits.
    • Cheap insurance either way: move the must-follow and novel sections to the top, and shorten or move the general ones (see keep/cut below). [char offsets mechanical; classes LLM-graded]
  4. PR description:
    • "61 build-script review comments" does not reproduce: 12 months of makefile*/dev/make comments give 147 (81 human). 61 is the number of distinct PRs.
    • "634 sites" is ≈317 namespace v1 blocks; each is counted twice because of its closing comment.
    • "152 .i" is correct.
  5. FYI for Align the clang-format pre-commit hook behavior with the CI #3813, measured at base, which this PR now keeps for formatting:
    • pre-commit selects 2934 files and CI 3079. The 152 cpp/daal .i files are CI-only, because identify gives .i no C tags.
    • So pre-commit run --all-files (root AGENTS.md:68) is not equivalent to CI until Align the clang-format pre-commit hook behavior with the CI #3813 lands.
    • On a clean tree, base pre-commit rewrites renovate.json, pkg-config.cpp and cpudetect.cpp. [mechanical]
General vs novel knowledge map (what models already know vs what the guidance teaches)

Method. Every head file was split into atomic facts ("atoms") with a question and answer each, 545 atoms in total. Each question was asked to 3 models with no repo contents (baseline) and with the file loaded, then graded by a judge from a different model family (lenient, direct probes).

  • general: all 3 models answer correctly without the file.
  • novel: no model knows it, and ≥2 answer correctly with the file.
  • partial: model-specific.

The models disagree a lot about what they already know: pairwise baseline agreement is 0.62–0.69. So "general" is a strict bar. Note also that "known when asked" does not mean "applied unprompted" (a model can know a rule and still not follow it in practice).

Overall: 99 general (18%), 283 partial (52%), 163 novel (30%). 17 of the novel atoms come from mechanically wrong lines (item 7), so novel ≠ correct.

Token split: each span is attributed to the best atom it carries. The priority is protected (a routing, negative-constraint, version-bound or rare-critical rule) > novel > partial > general-only. Spans that carry no atom are listed separately.

head file atoms general partial novel tok: novel / partial / general-only / protected / no-atom / heading
AGENTS.md 83 15 39 29 5% / 27% / 2% / 62% / 1% / 4%
dev/AGENTS.md 50 5 34 11 22% / 18% / 4% / 36% / 10% / 11%
dev/bazel/AGENTS.md 70 6 38 26 10% / 11% / 0% / 66% / 6% / 7%
dev/make/AGENTS.md 11 1 3 7 54% / 8% / 4% / 26% / 0% / 9%
cpp/daal/AGENTS.md 52 16 26 10 16% / 10% / 11% / 38% / 20% / 5%
cpp/oneapi/AGENTS.md 36 11 16 9 16% / 19% / 8% / 24% / 26% / 7%
.ci/AGENTS.md 73 18 34 21 22% / 29% / 13% / 29% / 0% / 7%
build-systems.instructions.md 55 7 37 11 27% / 21% / 4% / 23% / 12% / 12%
cpp-coding-guidelines.instructions.md 71 16 34 21 25% / 20% / 11% / 16% / 20% / 8%
general.instructions.md 14 2 8 4 18% / 11% / 9% / 53% / 2% / 8%
examples.instructions.md 8 0 4 4 37% / 0% / 0% / 53% / 0% / 11%
documentation.instructions.md 5 1 2 2 0% / 37% / 14% / 36% / 0% / 14%
.github/instructions/AGENTS.md 11 0 6 5 0% / 15% / 0% / 83% / 0% / 2%
copilot-instructions.md 6 1 2 3 0% / 0% / 0% / 92% / 0% / 8%

By kind of rule (general / partial / novel):

Tag General Partial Novel Reading
version_bound 3 7 17 Most novel
routing 11 47 45
rare_critical 4 24 18
negative_constraint 19 35 14 "Never do X" rules are often already known, but they are protected and worth keeping as guardrails
example 2 34 32 Concrete repo examples teach a lot, which is why wrong examples hurt
proposition 51 114 62 Descriptive statements hold most of the general knowledge

Sections that are mostly general or zero-novel (general / partial / novel):

File Section General Partial Novel
AGENTS.md intro 4 4 0
AGENTS.md Repository Structure 5 7 0
.ci/AGENTS.md Directory Structure 15 22 8
cpp/daal Core Patterns 7 8 2
cpp/daal Critical DAAL Rules 3 2 0
cpp/oneapi Quick Rules 2 3 0
cpp/oneapi Critical Rules 3 2 0
dev/bazel Common Pitfalls 3 1 0
cpp-coding-guidelines Interface-Specific Development Patterns 10 6 3
cpp-coding-guidelines Coding Style Standards 2 4 0
dev/AGENTS.md Structure 0 10 0

Sections that carry the most novel knowledge:

File Section General Partial Novel
AGENTS.md Verification Before You Push 0 10 11
AGENTS.md Directory Guides 0 5 8
AGENTS.md Rules for Changes 4 10 6
dev/bazel Rules for Changes 2 8 6
dev/bazel Configuration Files 0 5 5
dev/bazel Build Rules and Patterns 1 5 5
dev/make Rules for Changes / Layout 1 2 7
cpp-coding-guidelines Naming Conventions (7 of these are the made-up examples in item 7) 0 7 7
.ci/AGENTS.md CI/CD Architecture 3 8 7
cpp/oneapi Essential Commands 0 3 4
Keep / cut / fix: per-file recommendations

Keep as is. The protected rules (routing, negative constraints, version bounds, commands) are 16–92% of tokens per file, and every novel atom on a correct line: root "Verification Before You Push" and "Directory Guides", dev/bazel configuration and rules, dev/make, .github/instructions/AGENTS.md, copilot-instructions.md. Also keep cross-reference link lists. They look dead to the probes, but they serve navigation.

Fix, then keep: the 17 novel atoms on wrong lines (item 7). These are exactly where the guidance teaches something new and false.

Compress or cut (general-only content, which models answer correctly without it):

File What Why
root AGENTS.md Compress intro + "Repository Structure" to 2–3 lines, or move them below "Verification Before You Push" 9 general / 11 partial / 0 novel, and it is exactly the first 4K that code review may read
.ci/AGENTS.md Shrink "Directory Structure" to the non-obvious scripts It is a per-script listing, 13% general-only tokens; apt.sh, bazelisk.sh, environment.yml, tbb.sh, abi_check.sh, Renovate are known without it
cpp/daal/AGENTS.md Trim "Core Patterns" and "Critical DAAL Rules" 10 general / 10 partial / 2 novel; 11% general-only + 20% no-atom tokens, mostly the smart-pointer and kmeans snippets, part of which do not compile
cpp/oneapi/AGENTS.md Trim "Quick Rules" and "Critical Rules" (0 novel) 26% no-atom tokens (code blocks)
cpp-coding-guidelines Shorten "Interface-Specific Development Patterns" (10 general, first 4K), so "Error Handling", "Naming" and the checklist move up "Naming" examples must be replaced with real names first
dev/bazel/AGENTS.md "Common Pitfalls" is 3 general / 0 novel; merge into "Rules for Changes"

How much can be cut without losing knowledge? From the engine's shortening at 0.6 budget, sonnet, keeping all knowledge and all protected spans:

File Lossless cut
cpp/daal −20%
cpp-coding-guidelines −19%
dev −13%
build-systems −12%
dev/bazel −10% (knowledge .97)
root −3% (no slack)

These cuts are mostly the rule-free code snippets above. A blind 40% cut loses 25–41% of knowledge in every file. A value-greedy cut drops protected spans in 3/7 files, so don't trim by eye.

How this was measured, and cost
  1. Mechanical, no LLM. Every checkable claim (path, link, Bazel label, make target, identifier, fenced command, Starlark macro) was resolved against the tree at both SHAs. Every failure was adjudicated by hand.
  2. Tooling parity. pre-commit's own file classifier was compared with a shim logging the files CI's clang-format.sh touches, followed by a mutation test.
  3. Traffic. 12 months of human review comments per topic (Copilot excluded), churn per area and CI failure rates give a coarse need weight per file. It only separates 3 tiers: {cpp/daal, root, cpp/oneapi} > {dev/bazel, .ci, .github} > rest.
  4. Knowledge probes. An extractor (grok-4.6) split each file into atoms. Each atom got a direct, a paraphrased and a reversed-premise question. These were asked to claude-sonnet-5, kimi-k3 and glm-5 with no guidance, with the file, with same-length off-domain text (placebo), with the file shuffled, and with parent/leaf combinations. The judge was gpt-6-sol, with a second judge (deepseek-v3.2) on 450 samples (0.87 lenient agreement).
    • Controls: placebo adds ≈0, so the gains are knowledge, not length. Off-domain text actively degrades answers, so irrelevant always-on files have a real cost. Shuffled ≈ full, so there is no ordering effect.
  5. Cross-checks.
    • Retention: are deleted atoms still present anywhere in head?
    • Mechanical flags vs LLM harm (item 7).
    • Shortening candidates on 4 axes: tokens, knowledge, value, protection.

Scale and cost: 25 file-versions, 1099 atoms, 67,665 model calls (~60M tokens in / 21M out), about 10 h of call time. That is ≈$490 at $3/$15 per M tokens, an estimate rather than a bill. Most of it was redundant arms: with the file loaded, models were right 98% of the time. I'm validating a ~$50 version of the same analysis against this run.

Not measured
  • Real agent tasks with and without the guidance.
  • Bazel execution of the snippets in item 7 (static read only).
  • The middle-level files' own marginal value.
  • Whether Copilot code review still truncates at 4K.
  • Atom counts from a second extractor. Single-extractor dead-share moved by ~0.2 on byte-identical code blocks, so per-file atom counts carry that noise.
  • Most per-file base→head added-rate deltas sit inside their 95% CIs. The evidence for this PR is correctness and no knowledge lost, not a large knowledge gain.

@napetrov

Copy link
Copy Markdown
Contributor

Follow-up to my earlier comment, with more data. It's the same head e3804fb and the same 1,161 direct probes. The PR is at 8d0a9ce now, and the guidance files match. I re-asked the probes to the agents people actually run here, and also asked each one with the file loaded. Summary: the recommendations stand, but the "30% novel" number depends on the answering models and should be read per agent.

1. "Novel" depends on which model answers

Known rate without any guidance (one question per call, judge gpt-6-sol):

agent knows without the file
claude-opus-5-5 75%
claude-sonnet-5 59%
gpt-6-luna 59% (but 36% confidently wrong)
gpt-6-sol 53%
kimi-k3 / glm-5 (earlier panel) 42% / 32%
  • With the earlier panel (sonnet + kimi + glm), 30.5% of atoms were novel, meaning no model knew them. With opus-5.5 + sonnet-5 + gpt-6-sol + gpt-6-luna, the figure is 14%.
  • The files' ranking also changes with the panel. The correlation of per-file novel share with the earlier panel is 0.15 for the new panel and −0.18 for opus alone.
  • Opus's high rate is mostly inference, not memorisation: on files that are new in this PR, it still knows 82% (n=28).
  • So the "general/partial/novel" table in the earlier comment describes weaker models. For opus and sonnet users, read the per-agent table below.

2. Per agent: what share of each head file is new knowledge that the file actually teaches

The file taught 96–99% of the atoms each agent did not know:

agent taught
opus 284/295
sonnet 467/479
gpt-6-sol 536/542
gpt-6-luna 476/480

Share of each head file that is new for the agent and taught by the file:

head file atoms opus-5.5 sonnet-5 gpt-6-sol gpt-6-luna
AGENTS.md 83 27% 55% 53% 46%
.ci/AGENTS.md 73 22% 33% 58% 38%
cpp-coding-guidelines 71 34% 48% 44% 42%
dev/bazel/AGENTS.md 70 20% 41% 56% 51%
build-systems 55 16% 35% 51% 73%
cpp/daal/AGENTS.md 52 27% 29% 37% 31%
dev/AGENTS.md 50 20% 34% 42% 44%
cpp/oneapi/AGENTS.md 36 28% 36% 22% 44%
general.instructions 14 7% 29% 29% 0%
dev/make/AGENTS.md 11 18% 82% 82% 73%
  • Every file teaches every agent something; there is no file to drop.
  • For opus, 7–34% of each file is new. The rest mostly confirms what opus would do anyway.
  • Per-file shares barely correlate across agents: opus vs sonnet r = 0.09, sonnet vs gpt-6-sol r = 0.76. There is no single "most useful file" ranking.

3. What this does and does not change

  • Wrong lines (item 7): unchanged, and they now matter more.

    • With the file loaded, knowledge questions on mechanically flagged lines went wrong 0/155 across the four agents.
    • That is not reassurance. The expected answers come from the guidance itself, so a wrong line grades as "correct", and knowledge probes cannot detect wrong guidance at all.
    • A strong agent that reads e.g. the made-up KMeansBatch / MAX_ITERATIONS names in cpp-coding-guidelines:158-171 will follow them.
    • The mechanical list from item 7 is the only evidence here, and it still applies.
  • First-4K ordering (item 9): holds for every agent. Atoms each agent learns from the file that sit past the 4K cut:

    file opus sonnet gpt-6-sol gpt-6-luna
    root AGENTS.md (cut at line 62) 11/22 17/46 19/44 17/38
    cpp-coding-guidelines (cut at line 118) 17/24 20/34 21/31 22/30
  • Keep/cut table: still safe to trim. Of the 222 atoms the earlier panel found "general", 98–99% are also known by each of the four agents without the file. The "novel" counts in that table are an upper bound for opus users.

Method (for anyone reproducing)

  • One question per LLM call. Batching questions biases the no-file baseline: glm's known rate went 0.28 → 0.44.
  • Judge: gpt-6-sol. A gpt-6-luna judge agreed with it 94–96% on 1,200 regrades, with no GPT self-preference measured. An opus judge was +0.06 lenient on opus answers, so Claude answers are not graded by Claude.
  • A cheap re-run (gpt-6-luna extractor, per-call probes) reproduces the per-file novel ranking of the full run (r = 0.82) for ~$20 instead of ~$490. File-level tiers are stable; single atoms and sections are not (atom-class agreement between the two runs is 0.54), so I only report files.
  • Total for this follow-up: ≈ $86 at list prices. Numbers are LLM-graded, except the wrong-line list and char offsets, which are mechanical.

@ethanglaser

Copy link
Copy Markdown
Contributor Author

@napetrov thanks for the detailed review - agreed on nearly all of the suggested revisions, addressed in latest commit

ethanglaser and others added 2 commits September 29, 2026 20:23
examples/AGENTS.md showed a per-algorithm dal_example_suite target that
does not exist; algorithm examples are declared through
dal_algo_example_suite(algos = [...]). Both files are cut down to layout,
rules and verification commands checked against the real BUILD files,
docs/Makefile and rst_examples.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: ethanglaser <ethan.glaser@intel.com>
@ethanglaser

Copy link
Copy Markdown
Contributor Author

@Vika-F @napetrov two commits since your approvals, please take another look before merging:

  • 5835dec34: trims docs/AGENTS.md and examples/AGENTS.md to layout, rules and verification commands, and fixes the example BUILD pattern (dal_algo_example_suite).
  • 6ebb99992: removes .github/copilot-instructions.md and .github/instructions/*. Copilot code review and the coding agent both read AGENTS.md, so those files only duplicated it. Content not already in an AGENTS.md is merged into cpp/AGENTS.md and dev/AGENTS.md, and review-critical rules are near the top of the root AGENTS.md.

The PR description is updated (section 5).

@napetrov

napetrov commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

@ethanglaser looks good for me. o think it's beneficial to integrate and then work on next rounds of optimization

@ethanglaser
ethanglaser merged commit 82acef1 into uxlfoundation:main Oct 1, 2026
34 of 36 checks passed
@ethanglaser

Copy link
Copy Markdown
Contributor Author

Similar update in sklearnex: uxlfoundation/scikit-learn-intelex#3435

@Vika-F

Vika-F commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

@Vika-F @napetrov two commits since your approvals, please take another look before merging:

  • 5835dec34: trims docs/AGENTS.md and examples/AGENTS.md to layout, rules and verification commands, and fixes the example BUILD pattern (dal_algo_example_suite).
  • 6ebb99992: removes .github/copilot-instructions.md and .github/instructions/*. Copilot code review and the coding agent both read AGENTS.md, so those files only duplicated it. Content not already in an AGENTS.md is merged into cpp/AGENTS.md and dev/AGENTS.md, and review-critical rules are near the top of the root AGENTS.md.

The PR description is updated (section 5).

@ethanglaser Sorry, I was too slow with this. Can you please review #3840 ? .__.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants