|
|
958009858e
|
[CLAUDE] Docs: adopt Harness-16 MFE (memory-fidelity-EVAL) — 2-workflow + email AI_INFRA
Deploy SOLUTION_ERP / build-deploy (push) Successful in 5m2s
Adopt AI_INFRA Harness-16 (3 broadcast 2026-06-29) qua 2-workflow mandate:
WF1 implement wf_4c63e1bd-99e + WF2 review wf_13e3d35a-023 (PASS 0-blocking).
MFE = coverage/retention eval (do nap-vs-nho-vs-dung bang SO DO). DISTINCT
voi H6.7 memoryDelta-routing-fidelity (vocab-fork WF1 reviewer bat duoc).
Built (em-main single-writer, 0 production code):
- scripts/mfe-eval.ps1 deterministic NO-API (ASCII #30, exit 0): LEAD coverage-FIT
(token-RANGE, cap live-read) + age-band flag-not-cut + Goodhart-anchor strikes/RCA;
SUB per-role coverage do-that (prose->N/A khong 0%); sub-workflow N/A. Smoke 3-tier PASS.
- memory-budget.json :mfe config + eval/mfe/ (seed sample-questions stable-id + README)
+ engine §H + artifact-row + C3-comment + wire session-start §2.1.6 / session-end §L.b(c)
opt-in + agents/README S93. Judge layer = SCAFFOLD-only (honest nac).
adap-report harness-16-mfe + harness-15-v3 (covered/folded) + email outbox/ai_infra.
AS-10 dogfood: WF1 residual-write caught+reverted. State GIU NGUYEN (Mig 59 · 88 · 434 · gotcha 76).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-06-29 18:35:14 +07:00 |
|
|
|
1e1c9a2433
|
[CLAUDE] Docs: S31 RAG v1.3 baseline PASS (11/11 recall@5=1.000) + gotcha #52
Deploy SOLUTION_ERP / build-deploy (push) Successful in 3m38s
- eval/runs/: baseline v1.1 final PASS after retrieval.py fix (vector search restored)
- eval/trial-state-lock.json: quality_gate.pass=true, baseline=1.000, avg_rerank=0.847
- docs/gotchas.md: +gotcha #52 qdrant-client 1.18 removed search() silent AttributeError
- docs/STATUS.md: S31 entry — RAG PASS, retrieval.py fix, CLI restart required
- docs/HANDOFF.md: S31 brief + CRITICAL CLI restart note
- docs/changelog/sessions/: S31 session log
Root cause: qdrant-client 1.18 removed search() → vec_results always [] → BM25-only
Fix: retrieval.py query_points().points (applied to AI_INFRA repo)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
|
2026-05-26 13:42:04 +07:00 |
|
|
|
b223466ded
|
[CLAUDE] Docs: setup RAG Framework v1.3 governance + eval framework
Deploy SOLUTION_ERP / build-deploy (push) Successful in 3m52s
- docs/governance/README.md: Path B delegation stub → AI_INFRA canonical
Phase/BC vocabulary documented (9 phase + 10 BC SOLUTION_ERP-specific)
- .claude/rag.json: add _decision_log block (10 rationale entries) +
add .claude/agents/**/*.md to corpus_paths (fix Case D harvest gap)
- eval/evaluator.md: inline executor spec v1.0 (Spec A strict)
- eval/golden-set-solution_erp.jsonl: 14-entry golden set v1.1
(5 gotcha + 3 pattern + 3 decision + 3 negative)
- eval/runs/2026-05-26-baseline-v1.0-failed.json: v1.0 attempt
recall@5=0.455 FAIL — root cause diagnosis Case A/C/D
- eval/runs/2026-05-26-baseline-v1.1-pending.json: v1.1 attempt
pending CLI restart for accurate numbers
- eval/trial-state-lock.json: 2-section split (quality_gate +
drift_monitor) per v1.3 §6.2, 4-week milestones 2026-05-26 → 2026-06-23
CRITICAL lesson: bootstrap.py --project flag overrides collection name only.
Use --config D:\...\SOLUTION_ERP\.claude\rag.json for correct project root.
Old projects.json had root_path=AI_INFRA for solution_erp (Anti #24) — FIXED.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
|
2026-05-26 13:14:23 +07:00 |
|