nox-mem: Pain-Weighted Hybrid Memory for LLM Agents
Abstract
An open-source memory layer for LLM agents, measured against five other systems on public benchmarks, in single-file SQLite stores you can host yourself. The paper reports the G3→G10d ablation trajectory, one pre-specified cross-system comparison, and the findings that cut against its own headline. What changed in v1.0.5 (2026-10-04) Front matter only: the first page now gives the author's ORCID and contact address, this version's DOI and the concept DOI, and the code and data repository; the note on the measurement host moved from the first page to Section 5.7. No result, claim, reference or citation changed. What changed in v1.0.4 (2026-10-04) Writing pass: prose revised for plainness and concision (shorter sentences, fewer dashes and less emphasis). No number, claim, reference or citation changed; a parity script checks this mechanically over the whole manuscript, and a scoped review of the rewritten sentences restored one sentence whose scope had widened. Details in paper/CHANGELOG.md. What changed in v1.0.3 (2026-10-03; review and audit passes 2026-10-04) (A) Form and framing. A disclosure of generative AI use is added after the Conclusion, in line with arXiv's guidance: engineering and writing assistance, and read-only review by LLM-based reviewers from other model families. Product and roadmap framing is removed (go-to-market language, internal decision codes, "Autonomy pillar" labels); the pre-specified success criterion is kept and named as such. The duplicated, empty 6.8 heading is gone. (B) Corrections after review, each checked against the run artifacts and the EverMemBench paper: Zep ranks third (behind EverOS and nox-mem), not fourth; EverMemBench Table 4 has a Gemini-3-Flash column, so the 63.28% Overall is compared with MemOS on the same backbone (59.27%, +4.01 pp) and the GPT-4.1-mini comparison is labelled cross-backbone; combinations of retrieval-stage knobs are sub-additive, and the "retrieval ceiling" claims are withdrawn; the embedding-matched comparison no longer attributes the reversal to the embedder or to the architecture. (C) Metric and aggregation. F_MH is described as LLM-judged accuracy (it was called "strict EM"). nox-mem is also reported in Table 4's own aggregation (63.77%, +4.50 pp on the same backbone); in that aggregation the Gemini-2.5-flash cross-backbone margin is −0.06 pp. MemOS's F_MH is 18.88%, a GPT-4.1-mini figure. (D) Intervals and attributions. The IterB interval is IterB's own (95% CI [6.27, 9.79]; paired per-batch difference [0.25, 3.76]); the Wave C interval is recomputed with the t distribution; task setup is the leading, not established, account of the F_MH gap; HyperMem's 92.73% is a LoCoMo figure. (E) Audit pass, every statement re-checked against the code, the run artifacts and the cited papers. Section 3.4 now describes what the code does: crystallize stores caller-supplied procedures (no LLM, no promotion between chunk types), pain is fixed at ingest by a keyword rule and never raised afterwards, reflect answers on request without writing back, and nightly consolidation only extracts into topic files. EverMemBench F_MH is compared per backbone (6.02% against MemOS's 10.84% on Gemini-3-Flash). The per-category table of the embedding-matched comparison (6.4) is recomputed after a permuted LoCoMo category map was found; nox-mem still leads all five categories. EverOS outperforming nox-mem is stated in the abstract. Mem0's 66.88% on LoCoMo is an LLM-judge score, not F1, and the F1 ranking built on it is withdrawn; MuSiQue and HotpotQA reference figures are read from their source tables; several "significant" labels are corrected to what paired per-batch intervals support; the nox-mem RSS is the 399 MB measurement of 2026-05-29; smoke-run figures that no artifact reproduces are replaced by rescored ones. (F) Final review: the LoCoMo per-category retrieval cells match the archived run (single-hop 80.36%, temporal 77.96%), and an unsupported +2.8 pp date-normalization figure is replaced by the measured session-date injection (temporal F1 +15.94 pp); the MuSiQue and HotpotQA runs are described as working over each question's own candidate paragraphs, so they measure the reader, not retrieval; LightRAG's default stack is in-process storage; a Fisher test that treated paired runs as independent is withdrawn; the conclusion no longer reports the pre-specified criterion as met. (G) Tone and provenance: the comparison of Section 6 is described as pre-specified (execution plan committed to the public repository before the first run; not registered with an external registry), and the all-Gemini variant as a planned side experiment. The deployability and cost-ratio arguments of 5.7.2, 6.8 and 6.9 are cut; competitor RAM and cold-start estimates move to the supplement as the author's estimates; the observability and monitoring subsections move to the supplement. Three figures from a Mem0 re-execution whose script and per-query output were not retained are removed. The boost formula is defined once (score = base × (1 + sum of deltas)), and the per-category table of 6.4 cites the script that recomputes it. A final mechanical check restores two headings that rendered as plain text (5.2 and item F3 of 7.2), the missing EverMind-AI entry of 6.3.1, and the δ symbol of 4.1, which the previous PDF build dropped. No measured result changed in this part. The headline measurements (63.28% EverMemBench Overall, the nDCG@10 values of 6.3 and 6.3.2, KG-path 2.5 ms p50, 399 MB) are unchanged. paper/claims_check.py passes all 21 guards. The full list, with every changed number, is in paper/CHANGELOG.md. What changed in v1.0.2 (2026-09-29) Latency figures are back on archived artifacts. v1.0 quoted a 2026-06-15 re-check (KG path 2.9 ms, hybrid 653 ms p50) that was never archived. Every KG-path and standard-hybrid latency now reads from a versioned file: KG path 2.5 ms p50 (n=120) and hybrid 529 ms p50 (n=100), plus an earlier hybrid run at ~940 ms p50 with a different query mix. Ratios built on the old figure were recomputed or dropped. Two more systems measured since v1.0: EverOS and Zep, over the same corpus and the same 2,482 queries (§6.3.3, §6.3.4). Only Letta remains a documented deployment non-run. Residual confounds of the embedding-matched comparison: three declared in v1.0, four now (§6.3.2), including a retraction of the earlier account of the Mem0 version used in that run. The 5-batch protocol claim is scoped to EverMemBench (MuSiQue and HotPotQA are single full-dev runs); a LoCoMo retrieval@10 line no longer compares against Mem0's answer F1, a different metric. Related Work (§1.5) added; long appendices moved to verbatim supplements in the repository. Full list: paper/CHANGELOG.md in the repository. Abstract We introduce nox-mem, a persistent memory system for autonomous LLM agents built on one principle: pain-weighted hybrid memory with shadow discipline. Retrieval and retention are governed by an additive salience formula in which pain — an operator-assignable severity in [0.1, 1.0], otherwise fixed at ingest by a keyword rule, persisted on every chunk — is a first-class signal, and ranking changes pass a mandatory shadow phase before production activation. Each store is a single SQLite file with provider-swappable embeddings, MIT-licensed. Deployed in production since March 2026, it serves six specialized agents at KG-path p50 = 2.5 ms, $0 per KG-path query, and a 399 MB resident set in a single self-hosted process. Our central result is a pre-specified, same-corpus comparison against five competing memory systems, four of which produced head-to-head quality numbers. Under each system's native embedder nox-mem and Mem0 split — Mem0 wins LoCoMo (nDCG@10 0.469 vs 0.426), nox-mem wins LongMemEval. An embedding-matched variant (both Gemini 3072-d, n = 2,482) inverts that split in nox-mem's favour, with four residual confounds declared. EverOS, measured later over the same corpus and queries, outperforms nox-mem on both datasets (overall nDCG@10 0.646 vs 0.501), with a mandatory cross-encoder whose share of the gap is unmeasured; Zep ranks third, ahead of Mem0. On EverMemBench, nox-mem reaches 63.28% Overall with Gemini-3-flash, 4.01 pp above the 59.27% published for MemOS on the same backbone and below that backbone's 72.61% full-context baseline, so this is not a state-of-the-art claim. Findings that cut against the headline Pain-weighting, the title's own signal, is not statistically significant in isolation (§7.1). It is directional. Section-aware ranking, not pain, is the dominant empirical driver (§5.1.3). On the EverMemBench F_MH multi-hop track the system sits at 6.02% with Gemini-3-flash, against 10.84% for MemOS on the same backbone; the best memory-augmented F_MH in the benchmark's Table 4 is 18.88% (MemOS, GPT-4.1-mini); all LLM-judged (§5.4). Status of this manuscript This is a preprint. It has not been peer reviewed. It was submitted to arXiv on 2026-09-03 and not accepted, with the stated reason that it "would benefit from additional review and revision that is outside of the services we provide". arXiv does not assess scientific correctness, so that sentence means the manuscript needs peer review and arXiv does not perform peer review; it is not a finding about any specific claim. An earlier version asserted state-of-the-art results on two benchmarks simultaneously. That claim was retracted on 2026-09-03–04, together with five othe