Files
solution-erp/eval/mfe/README.md
pqhuy1987 c2d01f0dce
All checks were successful
Deploy SOLUTION_ERP / build-deploy (push) Successful in 5m30s
[CLAUDE] Docs: S152 closeout — bookend hình B trọn 2 đầu + owner (30)(31)(32) + queue memory-ops TRỌN (squash 20 wal:)
- Owner 3-quyết: (30) JUMP để-nguyên-theo-nhịp · (31) 4 persona +NEVER-block (restart ĐÃ THOẢ) · (32) vòng 1→5 AUTO — lô consent gỡ 9 site (5 H1 + ring1-tự-bắt + 3 trio; E-013 EXT)
- Memory-ops TRỌN: A1-A4/D1/D2/R1/B-series — A7 392/392 · mfe 20/20-100% lần đầu · lead-gap drain 24.728→16.485B · strike-ledger + WARN-latch (-Ack = owner-action)
- V4 sleep AUTO đầu: 1 shard thật/16 KÊU · ring4 đầu-đời TRƯỢT-4/5 → 4 fix gist (scribe sub-ring4-close vì return-only)
- Thư model re-stamp: 59ce1d0f → 9c909007 published (G-024a) — adap-report Đính-chính @S152 + sweep 13 bề-mặt LIVE
- Bookend @close: stale 6 FLAG + gap 2 FLAG → vá 8/8; ring2 8/8 ĐẠT (4-A pin-bản-cắt → slot 33) · ring1 48Đ/5T/59-claim
- M9 13 đơn-vị diary (3 S151-deep hồi-tố) · #53 ×10-chắc → tally 60-cận-dưới + errata subject-1b85713
- Session-log 2 run= + completeness-gate 5-vòng ĐẠT · HANDOFF segment S152 (carry 3-đóng/3-mới/23-giữ; slot 33/34)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-26 01:58:30 +07:00

55 lines
4.9 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# MFE — Memory-Fidelity-EVAL (Harness-16)
> **Adopt S93 (2026-06-29)** from AI_INFRA broadcasts `2026-06-29-Governance-harness-16-*` (memory-fidelity-eval + mfe-update + cd-judge-sub-update). Canonical mechanism → [`docs/governance/harness-11-engine.md §H`](../../docs/governance/harness-11-engine.md).
## What MFE is (and is NOT)
MFE measures **provisioned ≠ remembered ≠ applied**. The session `%-print` (Harness-15) tells you *how much you stuffed into* hot-memory (by length). MFE asks the next question: does the agent actually **retain/use** what was stuffed in.
🔴 **DISTINCT from H6.7 "memoryDelta-routing-fidelity"** (the right delta landing in the right `agent-memory/<role>` under single-writer). Same word "fidelity", two senses — always say **memory-fidelity-EVAL / MFE** for this one. The pair is recorded in `governance-detectors.ps1` C3 alias-map so it is not flagged as drift.
## Two layers
| Layer | What | Cost | Status (honest) |
|---|---|---|---|
| **Deterministic ($0)** | `scripts/mfe-eval.ps1` — LEAD coverage-FIT + age-band + Goodhart-anchor; SUB per-role coverage; SUB-workflow N/A | $0, NO-API | ✅ built + wired (opt-in) |
| **Judge (Branch A)** | recall/apply — give the agent the sample-questions, score recall + applied | $0 scaffold (quota only when a real scorer is wired) | 🟡 **SCAFFOLD only** — seed exists, **no scorer wired** |
## How to run
```
powershell.exe -ExecutionPolicy Bypass -File scripts\mfe-eval.ps1 # all tiers
powershell.exe -ExecutionPolicy Bypass -File scripts\mfe-eval.ps1 -Tier lead
powershell.exe -ExecutionPolicy Bypass -File scripts\mfe-eval.ps1 -Tier sub
powershell.exe -ExecutionPolicy Bypass -File scripts\mfe-eval.ps1 -Judge # show judge scaffold (no scoring)
powershell.exe -ExecutionPolicy Bypass -File scripts\mfe-eval.ps1 -Detail # +per-item MISS list (which denominator items unmatched)
powershell.exe -ExecutionPolicy Bypass -File scripts\mfe-eval.ps1 -Ack # 🔴 OWNER-ACTION ONLY — acknowledge GOODHART-WARN, move warn_baseline
```
### `-Detail` / `-Ack` (B4 @S152)
- **`-Detail`** — read-only diagnostic: prints one `MISS: <item>` line per unmatched denominator item, so a coverage drop is traceable to the exact item instead of a bare percentage. Safe to run any time.
- **GOODHART-WARN is LATCHED** — when strikes rise vs the persisted `warn_baseline` (stored in `.mfe-state.json`), the warning repeats on **every** run until acknowledged. It does NOT self-clear on a later good run (self-clearing would let one lucky run bury a real regression).
- **`-Ack`** — the ONLY thing that moves `warn_baseline` (and appends the ack to `history[]`). 🔴 **Owner-action:** lead/agents may run `-Detail` freely but MUST NOT run `-Ack` on their own — acknowledging a Goodhart warning = accepting the new strike level as the reference point, and that acceptance belongs to anh (same rule as every threshold/number in `memory-budget.json`). Lead surfaces the WARN; anh decides to ack.
Opt-in at session ends: `/session-start … eval` (baseline) and `/session-end … eval` (retention) — see `session-start.md §2.1.6` / `session-end.md §L.b(c)`. Default (no `eval`) = unchanged behaviour.
## Operational decision (the point)
- **Coverage < 100% or set over-cap** = *lack-of-SPACE* **INCREASE budget** (owner decides the number).
- **Low recall despite the set fitting** = *rot/noise* **REORGANIZE** (value-priority, cut low-value) adding space does NOT fix rot.
## 🔴 Honest caveats (do not hide)
1. **Token sizing is a RANGE, not a number.** `char/4` is not real tokenization; Vietnamese-diacritic hot-memory is ~3.03.5 byte/token, so `byte/4` is an upper bound on headroom. The analyzer reports `[bytes/4 … bytes/3.0]` and uses the worst-case end for the FIT verdict. Real tokenizer count is inside the band.
2. **The judge layer measures nothing yet.** A same-session self-grade is meaningless (the agent just read the answers). Real numbers need (a) the sample-questions to **mature in age** and (b) an **independent cross-session / different-model** judge. Until both, judge output is plumbing-smoke.
3. **Age is a flag, never a cut** (mark `RC-…10-29-11`). An item leaves the must-remember set only on **status-change** (mark Disabled, guard retired, AS-row deleted), never by age.
4. **No self-grading.** Coverage% is anchored to the real recurring-error signal (error-ledger strikes + RCA count). A high score next to rising strikes = the score lies.
## Files
- `scripts/mfe-eval.ps1` deterministic analyzer (NO-API, ASCII-only, exit 0, READ-ONLY on `token_governor`).
- `.claude/agent-memory/memory-budget.json` `mfe` block single-source config (denominator sources, stop-list, toggles, caveats).
- `eval/mfe/sample-questions.json` immutable seed (stable-id anchored), append-only.
- `.claude/agent-memory/.mfe-state.json` last-run strikes (cross-run Goodhart compare); MFE-only, never touches the budget.