Goal
Measure whether Kai's canonical memory pipeline provides value over a minimally configured stock Mem0 baseline, and identify regressions that Kai's own backend-comparison gate cannot reveal.
Current distinction
Kai already has a production memory evaluation gate that compares extraction/retrieval behavior across supported agent backends. That is not the same experiment as comparing Kai's scoped, provenance-aware pipeline with stock Mem0.
Experiment
Run the same sandboxed corpus and prompts through:
- Kai's full canonical pipeline, including principal/workspace scope, provenance, deduplication, episodes, retrieval selection, and compact prompt injection;
- a minimally configured stock Mem0 baseline using equivalent embedding/model resources where practical.
Score:
- fact extraction precision/recall;
- retrieval relevance and omission;
- cross-principal and cross-workspace leakage;
- provenance completeness;
- duplicate/contradictory memory behavior;
- fabricated or unsupported recollection;
- context size and end-task usefulness.
Guardrails
- Use disposable principals and stores only.
- Do not expose production memories or credentials.
- Record configuration and corpus versions so results are reproducible.
- Treat this as an evaluation issue, not authorization to replace the production memory stack.
Deliverable
A checked-in report and fixture-driven evaluation command that clearly separates measured results from architectural opinion.
Goal
Measure whether Kai's canonical memory pipeline provides value over a minimally configured stock Mem0 baseline, and identify regressions that Kai's own backend-comparison gate cannot reveal.
Current distinction
Kai already has a production memory evaluation gate that compares extraction/retrieval behavior across supported agent backends. That is not the same experiment as comparing Kai's scoped, provenance-aware pipeline with stock Mem0.
Experiment
Run the same sandboxed corpus and prompts through:
Score:
Guardrails
Deliverable
A checked-in report and fixture-driven evaluation command that clearly separates measured results from architectural opinion.