Repository navigation
Improve agent documentation - #3812
Conversation
…#3780) - Fix the .ci/AGENTS.md link, the root tree, the C++ standard and the nonexistent root .clang-format references. - Add a "Verification Before You Push" section with real commands and the current CI layout. - Replace the cc_library/cc_test Bazel template with dal_module / dal_test_suite, fix the MODULE.bazel description, drop //... and label CPU-only commands with --config=host. - Replace the nonexistent association_rules target and root CMake recipe; unwrap backticks inside bash fences. - Fix the clang-format command in CONTRIBUTING.md and drop the Mergify / Codefactor claims. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The pre-commit clang-format hook now checks the same directories and extensions as .ci/scripts/clang-format.sh, including the 152 .i kernel files, which identify assigns no type tags. It no longer touches deploy/ and dev/ sources that CI does not check. - clang-format.sh uses --dry-run --Werror instead of rewriting the tree and inferring failure from git status, so it is safe to run locally and unrelated uncommitted changes no longer fail it. - Mark *.i as C++ for linguist. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ctions Add "Rules for Changes" sections distilled from recurring maintainer review comments: cross-cutting rules and shell/batch portability in the root file, DAAL kernel rules (.i navigation, CpuType dispatch, TArray, accumulation, zero-division), the oneAPI versioned-namespace re-export, ABI and export parity, Bazel platform scoping and hygiene, and a new dev/make/AGENTS.md. Add .github/copilot-instructions.md, which routes Copilot to the AGENTS.md files and sets review posture, and extend path instructions to .i and .bzl files. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Drop generic advice, the quick-start steps, web links and emoji headers. Drop the smart-pointer rule, which conflicts with TArray in cpp/daal. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Address the formatter error-handling issue and remaining documentation inconsistencies.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 4
Open (4)
What changed in this PR
This PR updates agent documentation, build guidance, formatting checks, and Copilot configuration.
Changes:
- Corrects Bazel, Make, CMake, and formatting instructions.
- Adds repository and subsystem-specific guidance.
- Aligns local and CI formatting checks, including
.ifiles. - Adds Copilot instruction wiring.
| File | Summary |
|---|---|
dev/make/AGENTS.md |
Adds Make guidance. |
dev/bazel/AGENTS.md |
Documents Bazel macros and workflows. |
dev/AGENTS.md |
Updates development guidance. |
cpp/oneapi/AGENTS.md |
Documents oneAPI conventions. |
cpp/daal/AGENTS.md |
Adds DAAL implementation rules. |
cpp/AGENTS.md |
Documents ABI and public API constraints. |
CONTRIBUTING.md |
Corrects formatting instructions. |
AGENTS.md |
Adds repository-wide agent rules. |
.pre-commit-config.yaml |
Expands formatting hook coverage. |
.github/instructions/general.instructions.md |
Streamlines general instructions. |
.github/instructions/examples.instructions.md |
Updates example guidance. |
.github/instructions/documentation.instructions.md |
Updates documentation guidance. |
.github/instructions/cpp-coding-guidelines.instructions.md |
Updates C++ guidance. |
.github/instructions/build-systems.instructions.md |
Updates build-system instructions. |
.github/instructions/AGENTS.md |
Adds instruction-file maintenance rules. |
.github/copilot-instructions.md |
Adds the Copilot entry point. |
.github/.licenserc.yaml |
Exempts Copilot instructions from headers. |
.gitattributes |
Classifies .i files as C++. |
.ci/scripts/clang-format.sh |
Makes formatting checks non-mutating. |
.ci/AGENTS.md |
Adds CI-specific guidance. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: ethanglaser <ethan.glaser@intel.com>
|
@ethanglaser Thank you for this effort! I'd approve it with the following comment: Can you please remove |
Done |
|
I measured the agent guidance in this PR at base TL;DR: The rewrite makes the guidance much more correct (checkable-claim failures drop from 24% to 3%), and it lost no real knowledge: most of what was deleted was wrong or not practised. What remains to do:
What the PR got right
Recommended fixes
General vs novel knowledge map (what models already know vs what the guidance teaches)Method. Every head file was split into atomic facts ("atoms") with a question and answer each, 545 atoms in total. Each question was asked to 3 models with no repo contents (baseline) and with the file loaded, then graded by a judge from a different model family (lenient, direct probes).
The models disagree a lot about what they already know: pairwise baseline agreement is 0.62–0.69. So "general" is a strict bar. Note also that "known when asked" does not mean "applied unprompted" (a model can know a rule and still not follow it in practice). Overall: 99 general (18%), 283 partial (52%), 163 novel (30%). 17 of the novel atoms come from mechanically wrong lines (item 7), so novel ≠ correct. Token split: each span is attributed to the best atom it carries. The priority is protected (a routing, negative-constraint, version-bound or rare-critical rule) > novel > partial > general-only. Spans that carry no atom are listed separately.
By kind of rule (general / partial / novel):
Sections that are mostly general or zero-novel (general / partial / novel):
Sections that carry the most novel knowledge:
Keep / cut / fix: per-file recommendationsKeep as is. The protected rules (routing, negative constraints, version bounds, commands) are 16–92% of tokens per file, and every novel atom on a correct line: root "Verification Before You Push" and "Directory Guides", dev/bazel configuration and rules, dev/make, Fix, then keep: the 17 novel atoms on wrong lines (item 7). These are exactly where the guidance teaches something new and false. Compress or cut (general-only content, which models answer correctly without it):
How much can be cut without losing knowledge? From the engine's shortening at 0.6 budget, sonnet, keeping all knowledge and all protected spans:
These cuts are mostly the rule-free code snippets above. A blind 40% cut loses 25–41% of knowledge in every file. A value-greedy cut drops protected spans in 3/7 files, so don't trim by eye. How this was measured, and cost
Scale and cost: 25 file-versions, 1099 atoms, 67,665 model calls (~60M tokens in / 21M out), about 10 h of call time. That is ≈$490 at $3/$15 per M tokens, an estimate rather than a bill. Most of it was redundant arms: with the file loaded, models were right 98% of the time. I'm validating a ~$50 version of the same analysis against this run. Not measured
|
|
Follow-up to my earlier comment, with more data. It's the same head 1. "Novel" depends on which model answersKnown rate without any guidance (one question per call, judge gpt-6-sol):
2. Per agent: what share of each head file is new knowledge that the file actually teachesThe file taught 96–99% of the atoms each agent did not know:
Share of each head file that is new for the agent and taught by the file:
3. What this does and does not change
Method (for anyone reproducing)
|
|
@napetrov thanks for the detailed review - agreed on nearly all of the suggested revisions, addressed in latest commit |
examples/AGENTS.md showed a per-algorithm dal_example_suite target that does not exist; algorithm examples are declared through dal_algo_example_suite(algos = [...]). Both files are cut down to layout, rules and verification commands checked against the real BUILD files, docs/Makefile and rst_examples.py. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: ethanglaser <ethan.glaser@intel.com>
|
@Vika-F @napetrov two commits since your approvals, please take another look before merging:
The PR description is updated (section 5). |
|
@ethanglaser looks good for me. o think it's beneficial to integrate and then work on next rounds of optimization |
|
Similar update in sklearnex: uxlfoundation/scikit-learn-intelex#3435 |
@ethanglaser Sorry, I was too slow with this. Can you please review #3840 ? .__. |

Description
1. Instructions that were wrong (napetrov's #3780)
dev/bazel/AGENTS.md(+47/−96): replaced the barecc_library/cc_testandcoptsexamples with the macros the repo actually uses:dal_module,dal_test_suite,daal_module. It also now has correct dependency labels and a note that--configmust be passed. The reference target iscpp/oneapi/dal/algo/pca/BUILD.dev/AGENTS.md,cpp/oneapi/AGENTS.md: replaced the unsupportedbazel test //...with scoped targets using--config=hostor--config=dpc. Make commands now use the realmake -f makefile ... PLAT=form. The CMake recipes now build the examples against an installed release, since there is no rootCMakeLists.txt.dev/AGENTS.md, thedaal_moduleexample that passedcpu_definesandfeatures(the macro sets both), the Bazel version and macOS claims, TBB lines that contradicted the threading-layer rule, fake names (KMeansBatch,MAX_ITERATIONS), and snippets incpp/daalandcpp/oneapithat didn't compile (missing=,depends_oninside the kernel, the SPMD call, a devicepullwithout a queue).examples/AGENTS.md: replaced a per-algorithmdal_example_suitetarget that doesn't exist with the realdal_algo_example_suite(algos = [...])registration.CONTRIBUTING.md: fixed the clang-format command (it was missing a-) and replaced the link to a root.clang-formatfile that doesn't exist.2. Code search (#3781)
.gitattributes:.ifiles are treated as C++, so code search finds them. The clang-format hook and CI script changes from Tooling gaps:.ikernel files unsearchable and skipped by pre-commit, formatter script mutates the tree, undocumentednamespace vNre-export, 5 undocumented algorithms #3781 are left to Align the clang-format pre-commit hook behavior with the CI #3813.3. Rules taken from review comments
AGENTS.md: a "Rules for Changes" section covering:shand.batportabilitycpp/daal/AGENTS.md:.ifiles take aCpuTypetemplate parameterTArrayand aligned allocationcpp/oneapi/AGENTS.md: thenamespace vNplususingre-export pattern, used in 472namespace v1blocks and previously undocumented.cpp/AGENTS.md: a new section on ABI and the public API.dev/make/AGENTS.md: makefiles and.mk,.batand.shscripts drew 61 human review threads across the last 200 PRs and had no guidance file..ci/AGENTS.md: CI-specific rules.4. Trimming
AGENTS.md: the rules and verification commands come first, within the first 4K characters. Also removed emoji headers, general C++ advice (including a smart-pointer rule that contradicts theTArrayrule), the quick start, and web links.cpp/AGENTS.mdanddev/AGENTS.mdrewritten around what the repo actually does: an interface-conventions table, a review checklist, the realtrain_ops_dispatchersignature, and verified MakePLAT/COMPILER/REQCPU/BACKEND_CONFIGvalues. Invented snippets and emoji headers are gone.docs/AGENTS.md(7.9K to 1.1K) andexamples/AGENTS.md(6.7K to 1.5K): cut to layout, rules and verification commands. New rules: Sphinx runs with-W, generatedsource/examples/RST is not committed, and renaming an example breaks its docs references.5. Drop the Copilot instruction files
.github/copilot-instructions.mdand.github/instructions/*. Copilot code review and the coding agent both read AGENTS.md, so they only duplicated it.AGENTS.md, because that is the file Copilot code review is documented to read..licenserc.yaml: dropped the two exemptions for the removed files.Checklist:
Completeness and readability
Testing
Performance
🤖 Generated with Claude Code