[CLAUDE] Docs: adopt Harness-16 MFE (memory-fidelity-EVAL) — 2-workflow + email AI_INFRA
All checks were successful
Deploy SOLUTION_ERP / build-deploy (push) Successful in 5m2s

Adopt AI_INFRA Harness-16 (3 broadcast 2026-06-29) qua 2-workflow mandate:
WF1 implement wf_4c63e1bd-99e + WF2 review wf_13e3d35a-023 (PASS 0-blocking).

MFE = coverage/retention eval (do nap-vs-nho-vs-dung bang SO DO). DISTINCT
voi H6.7 memoryDelta-routing-fidelity (vocab-fork WF1 reviewer bat duoc).

Built (em-main single-writer, 0 production code):
- scripts/mfe-eval.ps1 deterministic NO-API (ASCII #30, exit 0): LEAD coverage-FIT
  (token-RANGE, cap live-read) + age-band flag-not-cut + Goodhart-anchor strikes/RCA;
  SUB per-role coverage do-that (prose->N/A khong 0%); sub-workflow N/A. Smoke 3-tier PASS.
- memory-budget.json :mfe config + eval/mfe/ (seed sample-questions stable-id + README)
  + engine §H + artifact-row + C3-comment + wire session-start §2.1.6 / session-end §L.b(c)
  opt-in + agents/README S93. Judge layer = SCAFFOLD-only (honest nac).

adap-report harness-16-mfe + harness-15-v3 (covered/folded) + email outbox/ai_infra.
AS-10 dogfood: WF1 residual-write caught+reverted. State GIU NGUYEN (Mig 59 · 88 · 434 · gotcha 76).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
pqhuy1987
2026-06-29 18:35:14 +07:00
parent 8ee8f3e248
commit 958009858e
15 changed files with 602 additions and 1 deletions

46
eval/mfe/README.md Normal file
View File

@ -0,0 +1,46 @@
# MFE — Memory-Fidelity-EVAL (Harness-16)
> **Adopt S93 (2026-06-29)** from AI_INFRA broadcasts `2026-06-29-Governance-harness-16-*` (memory-fidelity-eval + mfe-update + cd-judge-sub-update). Canonical mechanism → [`docs/governance/harness-11-engine.md §H`](../../docs/governance/harness-11-engine.md).
## What MFE is (and is NOT)
MFE measures **provisioned ≠ remembered ≠ applied**. The session `%-print` (Harness-15) tells you *how much you stuffed into* hot-memory (by length). MFE asks the next question: does the agent actually **retain/use** what was stuffed in.
🔴 **DISTINCT from H6.7 "memoryDelta-routing-fidelity"** (the right delta landing in the right `agent-memory/<role>` under single-writer). Same word "fidelity", two senses — always say **memory-fidelity-EVAL / MFE** for this one. The pair is recorded in `governance-detectors.ps1` C3 alias-map so it is not flagged as drift.
## Two layers
| Layer | What | Cost | Status (honest) |
|---|---|---|---|
| **Deterministic ($0)** | `scripts/mfe-eval.ps1` — LEAD coverage-FIT + age-band + Goodhart-anchor; SUB per-role coverage; SUB-workflow N/A | $0, NO-API | ✅ built + wired (opt-in) |
| **Judge (Branch A)** | recall/apply — give the agent the sample-questions, score recall + applied | $0 scaffold (quota only when a real scorer is wired) | 🟡 **SCAFFOLD only** — seed exists, **no scorer wired** |
## How to run
```
powershell.exe -ExecutionPolicy Bypass -File scripts\mfe-eval.ps1 # all tiers
powershell.exe -ExecutionPolicy Bypass -File scripts\mfe-eval.ps1 -Tier lead
powershell.exe -ExecutionPolicy Bypass -File scripts\mfe-eval.ps1 -Tier sub
powershell.exe -ExecutionPolicy Bypass -File scripts\mfe-eval.ps1 -Judge # show judge scaffold (no scoring)
```
Opt-in at session ends: `/session-start … eval` (baseline) and `/session-end … eval` (retention) — see `session-start.md §2.1.6` / `session-end.md §L.b(c)`. Default (no `eval`) = unchanged behaviour.
## Operational decision (the point)
- **Coverage < 100% or set over-cap** = *lack-of-SPACE* **INCREASE budget** (owner decides the number).
- **Low recall despite the set fitting** = *rot/noise* **REORGANIZE** (value-priority, cut low-value) adding space does NOT fix rot.
## 🔴 Honest caveats (do not hide)
1. **Token sizing is a RANGE, not a number.** `char/4` is not real tokenization; Vietnamese-diacritic hot-memory is ~3.03.5 byte/token, so `byte/4` is an upper bound on headroom. The analyzer reports `[bytes/4 … bytes/3.0]` and uses the worst-case end for the FIT verdict. Real tokenizer count is inside the band.
2. **The judge layer measures nothing yet.** A same-session self-grade is meaningless (the agent just read the answers). Real numbers need (a) the sample-questions to **mature in age** and (b) an **independent cross-session / different-model** judge. Until both, judge output is plumbing-smoke.
3. **Age is a flag, never a cut** (mark `RC-…10-29-11`). An item leaves the must-remember set only on **status-change** (mark Disabled, guard retired, AS-row deleted), never by age.
4. **No self-grading.** Coverage% is anchored to the real recurring-error signal (error-ledger strikes + RCA count). A high score next to rising strikes = the score lies.
## Files
- `scripts/mfe-eval.ps1` deterministic analyzer (NO-API, ASCII-only, exit 0, READ-ONLY on `token_governor`).
- `.claude/agent-memory/memory-budget.json` `mfe` block single-source config (denominator sources, stop-list, toggles, caveats).
- `eval/mfe/sample-questions.json` immutable seed (stable-id anchored), append-only.
- `.claude/agent-memory/.mfe-state.json` last-run strikes (cross-run Goodhart compare); MFE-only, never touches the budget.

View File

@ -0,0 +1,43 @@
{
"_note": "Harness-16 MFE Branch-A sample-questions SEED (S93, 2026-06-29). IMMUTABLE + time-stamped + anchored by STABLE-ID (RC-sig / gotcha# / AS# / budget-key) -- NEVER by line-number (files live, line-numbers drift). Seeded NOW so age-retention is measurable months later. Questions probe the LEAD must-remember set + a few SUB role-floors. Answers are the KEY FACT only (a judge checks recall/apply, paraphrase-tolerant). DO NOT rewrite existing questions (additive-only, like archive verbatim) -- append new ones with fresh ids.",
"seeded_date": "2026-06-29",
"scoring_status": "SCAFFOLD - no scorer wired. Same-session self-grade = MEANINGLESS. Needs (a) age-maturity AND (b) independent cross-session/different-model judge.",
"questions": [
{ "id": "Q-mark-09", "anchor": "RC-pqhuy1987-20-06-2026-10-29-09", "tier": "lead",
"q": "What does the architecture-decision mark assert about how to judge whether a feature is justified?",
"expect": "Objective criteria (pain / volume / quality), NOT team-size; 'overkill / too-much-for-solo-dev / gut-feeling' = rejected reasoning." },
{ "id": "Q-mark-11", "anchor": "RC-pqhuy1987-20-06-2026-10-29-11", "tier": "lead",
"q": "Why is time/age/recency a false proxy for memory-budget and drift decisions?",
"expect": "Age is same-family as team-size; cap = capacity/refresh-rate not an age-decay knob; drift = rolling baseline not an age window; age-decay cuts good memory = false economy (Goodhart). Applied to MFE: age = FLAG, never a cut." },
{ "id": "Q-mark-15", "anchor": "RC-pqhuy1987-20-06-2026-23-07-37", "tier": "lead",
"q": "What is the core principle of the H-15 memory-budget (token-governor)?",
"expect": "Budget = MINIMUM-to-USE floor (fill Tier-1 hot-feed with real work-state), NOT a ceiling to economize; token-saving = forgetting work. Numbers are the project-owner's authority; em-main executes + reports %." },
{ "id": "Q-cap-lead", "anchor": "memory-budget.json:token_governor.tier1_hotfeed_tokens.lead_tokens", "tier": "lead",
"q": "Where does the lead hot-feed token cap live, and may a tool hardcode it?",
"expect": "Lives ONLY in memory-budget.json (single-source); live-read it (B1 derived-tro-canonical); NEVER hardcode (it moved 60K->200K->220K). MFE READS it, never writes it." },
{ "id": "Q-as12", "anchor": "AS-12", "tier": "lead",
"q": "Before an identifier-based data op on prod (lock/seed/migrate-by-email), what must you do first?",
"expect": "DUMP the target-env table first; do not write the identifier list from CODE/Dev population. A 0-row / -1 assertion => suspect data-mismatch BEFORE code-bug. (gotcha #60 / E-008)" },
{ "id": "Q-as10", "anchor": "AS-10", "tier": "lead",
"q": "A sub-agent wrote a tracked file despite being return-only. What is the containment?",
"expect": "git-diff post-P2 catches it; em-main VERIFIES benign+accurate+placement then keep-if-correct or revert; NOT mechanized (G-015, sub keeps Bash). Defense-in-depth = git-diff + chunk-count." },
{ "id": "Q-gotcha-30", "anchor": "gotcha #30", "tier": "lead",
"q": "Why must a PowerShell .ps1 script body be pure ASCII?",
"expect": "Box-glyphs / Vietnamese literals in a PS 5.1 -File script body mojibake (even via Edit's render-normalize); use ASCII + code-points (e.g. [char]0x2705)." },
{ "id": "Q-gotcha-53", "anchor": "gotcha #53", "tier": "lead",
"q": "How is heavy-agent return-truncation mitigated?",
"expect": "em-main verify-on-disk + proxy-append (the agent often wrote the finding to disk before the empty return); lean memoryDelta return; 529 -> em-main solo fallback, no retry-loop." },
{ "id": "Q-gotcha-75", "anchor": "gotcha #75", "tier": "lead",
"q": "Why is a prod data-wipe not durable, and what is the fix?",
"expect": "Ungated per-code seeders RE-ADD the wiped data every restart; gate the seed behind an env-flag and verify by a REAL restart (not 'data clean right after wipe')." },
{ "id": "Q-h16-vocab", "anchor": "harness-16 vocab", "tier": "lead",
"q": "What is the difference between 'memory-fidelity-EVAL' and 'memoryDelta-routing-fidelity'?",
"expect": "MFE (H-16) = coverage/retention eval (does hot-feed retain the must-remember set). memoryDelta-routing-fidelity (H6.7) = the right delta lands in the right agent-memory under single-writer. Two distinct senses; do not conflate (C3 vocab-fork guard)." },
{ "id": "Q-sub-reviewer", "anchor": ".claude/agents/reviewer.md", "tier": "sub",
"q": "What is the reviewer sub-agent strictly forbidden from doing?",
"expect": "NEVER Edit/Write/commit/push; it produces a PASS/FAIL verdict with file:line, never writes code." },
{ "id": "Q-sub-implbackend", "anchor": ".claude/agents/implementer-backend.md", "tier": "sub",
"q": "What is implementer-backend forbidden to touch?",
"expect": "No FE 2-app (that is implementer-frontend); no test assertions (that is test-specialist); no schema/UX/cross-stack-bug reasoning (em-main solo). Auto-refuses out-of-scope." }
]
}