[CLAUDE] Docs: adopt Harness-16 MFE (memory-fidelity-EVAL) — 2-workflow + email AI_INFRA
All checks were successful
Deploy SOLUTION_ERP / build-deploy (push) Successful in 5m2s
All checks were successful
Deploy SOLUTION_ERP / build-deploy (push) Successful in 5m2s
Adopt AI_INFRA Harness-16 (3 broadcast 2026-06-29) qua 2-workflow mandate: WF1 implement wf_4c63e1bd-99e + WF2 review wf_13e3d35a-023 (PASS 0-blocking). MFE = coverage/retention eval (do nap-vs-nho-vs-dung bang SO DO). DISTINCT voi H6.7 memoryDelta-routing-fidelity (vocab-fork WF1 reviewer bat duoc). Built (em-main single-writer, 0 production code): - scripts/mfe-eval.ps1 deterministic NO-API (ASCII #30, exit 0): LEAD coverage-FIT (token-RANGE, cap live-read) + age-band flag-not-cut + Goodhart-anchor strikes/RCA; SUB per-role coverage do-that (prose->N/A khong 0%); sub-workflow N/A. Smoke 3-tier PASS. - memory-budget.json :mfe config + eval/mfe/ (seed sample-questions stable-id + README) + engine §H + artifact-row + C3-comment + wire session-start §2.1.6 / session-end §L.b(c) opt-in + agents/README S93. Judge layer = SCAFFOLD-only (honest nac). adap-report harness-16-mfe + harness-15-v3 (covered/folded) + email outbox/ai_infra. AS-10 dogfood: WF1 residual-write caught+reverted. State GIU NGUYEN (Mig 59 · 88 · 434 · gotcha 76). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@ -40,6 +40,34 @@
|
||||
"patterns": ["gotcha #", "anti-pattern", "recurring", "lost-update", "race", "bai hoc", "lesson", "guard", "root-cause", "silent-fail"]
|
||||
}
|
||||
},
|
||||
"mfe": {
|
||||
"_note": "Harness-16 Memory-Fidelity-EVAL (MFE) config (S93, 2026-06-29). Read by scripts/mfe-eval.ps1 (deterministic NO-API analyzer). DISTINCT from H6.7 'memoryDelta-routing-fidelity' -- THIS = COVERAGE/RETENTION eval (does the hot-feed actually retain the must-remember set). Single-source config so the analyzer never hardcodes (B1 derived-tro-canonical). MFE READS token_governor caps (live), NEVER writes them (owner-authority floor, role_boundary_note). Branch-A judge layer = SCAFFOLD-only (numbers meaningless until sample-questions mature AND an independent cross-session judge runs).",
|
||||
"denominator_sources": {
|
||||
"marks": ".claude/governance/ACTIVE-MARKS.md (table rows whose Status = Active-High / Active; skip Medium / Disabled)",
|
||||
"guards": "docs/governance/error-ledger.md (Active-Guards index rows Verified=check + Net not-negative; AS-table AS-N = recurring-error CLASS registry)",
|
||||
"recurring_gotchas": "docs/gotchas.md (### N. headings whose heading+body carries a recurrence-token)"
|
||||
},
|
||||
"merge_guard_to_as": true,
|
||||
"_merge_note": "true = a recurring-gotcha or guard whose AS-row references its #N (or AS-N) collapses INTO that AS class (AS = canonical class registry, deterministic by cross-reference). false = count each source separately (upper bound).",
|
||||
"recurrence_tokens": ["tai phat","tai dien","tai-dien","recurrence","recurring","bug-class","bug class","extend #","cung ho","cung-ho","whack-a-mole","class s"],
|
||||
"sub_denominator": {
|
||||
"role_file_glob": ".claude/agents/*.md",
|
||||
"diary_path_pattern": ".claude/agent-memory/<role>/MEMORY.md",
|
||||
"_diary_trap_note": "READ the agent DIARY agent-memory/<role>/MEMORY.md -- NOT project CLAUDE.md NOR user memory/MEMORY.md (Harness-15-v3 duplicate-filename trap = false-positive).",
|
||||
"denom_header_anchors": ["anti-pattern","split boundary","auto-refuse","boundary","NEVER"],
|
||||
"content_word_min": 2,
|
||||
"_match_note": "numerator = denom item present in diary by EXACT item-id OR >=2 distinct CONTENT-words (after NFD-strip-diacritics + lowercase + drop stop_list). Leading-verb-only match = false-max (broadcast: a verb-only memory scored 9/9) -- the stop_list defeats it.",
|
||||
"stop_list": ["do","not","never","dont","avoid","skip","use","always","ensure","make","keep","the","a","an","of","to","and","or","is","be","when","if","khong","luon","phai","dung","cac","mot","la","va","khi","neu","read","write","edit","commit","push"]
|
||||
},
|
||||
"sample_questions": "eval/mfe/sample-questions.json",
|
||||
"state_file": ".claude/agent-memory/.mfe-state.json",
|
||||
"honest_caveats": {
|
||||
"token_estimate": "char/4 is NOT real tokenization; VN-diacritic hot-memory ~3.0-3.5 byte/tok; report a RANGE [bytes/4 .. bytes/3.0]; FIT verdict uses bytes/3.0 (worst-case). Real tokenizer count lies inside the band -- never a single false-precise number.",
|
||||
"judge_layer": "Branch-A recall/apply scorer is SCAFFOLD-only (empty). Self-grade same-session = meaningless (just read the answers). Numbers meaningful ONLY after (a) sample-questions mature in age AND (b) an independent cross-session/different-model judge scores.",
|
||||
"age_band": "age is a FLAG, never a cut (mark RC-...10-29-11 age=false-proxy). An item drops only on status-change (mark->Disabled, guard->retired, AS row deleted), NOT by age.",
|
||||
"read_only_scope": "MFE WRITES exactly one file: .claude/agent-memory/.mfe-state.json (its own strikes-state for cross-run Goodhart compare). 'READ-ONLY' is scoped to token_governor / the budget caps (owner-authority) which it NEVER writes -- it is NOT the claim 'writes nothing'."
|
||||
}
|
||||
},
|
||||
"harness_floor": {
|
||||
"_note": "Harness-15 A1/A3 (S81, 2026-06-20): SAN-harness = fixed per-spawn cost (NOT tunable) = tool-schema + framing + own persona/role file + lead-pasted base-doc slice + task prompt. SEPARATE HOUSE (A3 anti-double-count): persona + lead-pasted-docs belong HERE (floor), NOT counted in token_governor.l1_always (which = own agent-memory + archive index + work-state block only). MEASURED-ESTIMATE not exact (H15 honest-note b): persona = directly measured bytes (.claude/agents/<name>.md 4.3KB-13.3KB => ~1.3K-4.0K tok via /3.3); tool-schema + framing = harness-injected (cannot byte-count locally), estimated comparable to AI_INFRA same-toolset-family (Read/Write/Edit/Bash/Grep/Glob/Skill/RAG).",
|
||||
"measured_token_estimate": 21000,
|
||||
|
||||
Reference in New Issue
Block a user