Golden finding-set acceptance + per-category coverage gate
Goal
The plan-wide FUNCTIONAL criterion: build the ori_arc golden finding-set + the end-to-end probe proving the catch+narrow outcome at a fraction of LLM cells.
Implementation Sketch
Hand-label the ori_arc ground truth; the probe runs the graph-driven review and asserts graph-native categories emitted directly, llm-only categories each handed >=1 narrowed candidate, recall >= baseline, 0 categories blind-walked, LLM dispatches <= candidate count.
Spec References
section-07 tier engine; section-00 golden probe; the file-walk baseline run; layer-coverage.md.
Work Items
- Hand-label the ori_arc golden finding-set (per category, verified file:line) — the ground truth.
- Acceptance probe: run the graph-driven review on ori_arc; assert (a) every graph-native/AST category emitted directly from the graph (0 per-file LLM cells), (b) every llm-only category present handed >=1 graph-narrowed candidate, (c) recall >= the file-walk baseline.
- Per-category coverage gate: every of the 26 categories is graph-emitted, graph-narrowed, or explicitly out-of-code-graph — 0 categories blind-walked.
- LLM-cell budget assertion: total LLM dispatches <= candidate count, not files x batches.
Fresh intel (regenerated)
{ “schema_version”: 1, “target”: { “kind”: “plan-section”, “ref”: “graph-driven-hygiene/s-0fdd5d7a” }, “generated_at”: “2026-06-22T06:46:29.238536+00:00”, “graph_state”: { “head_sha”: “9e4174b6”, “last_code_import_at”: “2026-06-22T06:40:04.905Z”, “embedding_stale”: false, “insights_stale”: false, “cpg_stale”: false }, “surfaces”: {}, “summary”: “intel-package[plan-section:graph-driven-hygiene/s-0fdd5d7a] agent-authored dossier”, “degraded”: false, “dossier”: { “objective_symbols”: [], “tiers”: {}, “agent_authored”: true, “difficulty”: { “class”: “routine”, “research_online”: false, “research_mode”: “auto”, “signals”: [], “plan_dir”: “/home/eric/projects/ori_lang/plans/graph-driven-hygiene” } } }
(full dossier: golden-set-acceptance—s-0fdd5d7a.intel.json)