{"id":"538103a1-3c3f-4984-9dfa-323e6a0a82a5","arxiv_id":"2608.08055","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"SodaMem, an evidence-grounded temporal graph memory, reaches 92.8 percent accuracy on LongMemEval-S at about 0.16 cents per question.","lead":"SodaMem is a memory system for AI assistants that stores facts as dated, source-linked graph entries and answers follow-up questions with cited evidence. It reports 92.8 percent accuracy on a long-term memory benchmark at under two tenths of a cent per question, though the accuracy is self-graded and the cost comparison is compiled from public reports.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that SodaMem strictly dominates near-cost Flash baselines rests on margins smaller than its own self-grading and cost-exclusion uncertainties; an independent judge and full-cost accounting could erase the dominance.","rationale":"The reader's conditional verdict is the right level: SodaMem is an engineering-first contribution with released code and frozen stores, and the headline numbers are explicitly caveated, so rejection would be unfair. But the headline claim is a point-in-plane statement; its two coordinates are measured under conditions the paper itself says can shift. The closest dominated point makes the fragility concrete: 6.2 accuracy points and $0.23/10^3Q are small absolute margins, and both uncertainties (self-grading and excluded ingest/judge) plausibly exceed one margin. This is not an accusation of dishonesty; it is a call for a specific re-measurement. My stress-test therefore adds a concrete threshold test to the reader's suggested independent re-judging and unified harness. If the re-measured point still dominates Cersei Embed on both axes, the central claim would be substantially strengthened. Until then, CONDITIONAL (unchanged) is appropriate.","tokens_in":10392,"tokens_out":10516,"duration_ms":110633,"concrete_test":"Obtain the released SodaMem answer hypotheses for the 92.8% run and, under a single independent judge (e.g., GPT-4o with LongMemEval's official templates), re-score them alongside the closest claimed-dominated baseline (Cersei Embed) using an identical protocol; also compute SodaMem's per-question cost including ingest and judge at the reported Flash rates. If SodaMem's independent-judge accuracy drops below 86.6% or its full cost exceeds $1.84/10^3Q, replace the 'strictly dominates' wording with a weaker 'comparable/order-of-magnitude' claim and report sensitivity to the best-of-three selection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest comparative statement is 'strictly dominates several higher-cost, lower-accuracy points,' and the closest such point is Cersei Embed at 86.6% accuracy and $1.84/10^3Q (Table 1). SodaMem's margin over that point is 6.2 accuracy points and $0.23/10^3Q. Both margins lie inside the paper's stated uncertainties. Accuracy is self-graded: 'The same Flash model grades the run (self-grading)' (Experiments: Setup and Cost Protocol), and 92.8% is the best of N=3 with median 90.6%. Cost is planner+reader usage excluding ingest and judge, while Cersei Embed's cost is an author-measured estimate from a different protocol. A judge-induced drop of a few points, or an ingest/judge cost of a few tenths of a cent per question, moves SodaMem below Cersei Embed on one axis or the other. Since the dominated quadrant and 'near the accuracy frontier' are read off exactly these coordinates, the central cost-accuracy claim is not robust to the paper's own acknowledged measurement choices.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents SodaMem, an evidence-grounded temporal graph memory for LLM agents. The system extracts typed FactEvents with provenance spans, maintains mention/occurrence/validity times and SUPERSEDES/CONTRADICTS/UPDATES edges, indexes facts with hybrid BM25-dense retrieval, and answers questions via a planner-reader loop. On LongMemEval-S (500 questions), the store-of-record configuration reports 92.8% accuracy (464/500, best of N=3) at a mean cost of $0.00161/question (~18.3k tokens) using deepseek-v4-flash. The paper compiles public baselines into a cost-accuracy table and claims SodaMem sits near the accuracy frontier at Flash-tier spend and strictly dominates several higher-cost, lower-accuracy systems. The abstract explicitly acknowledges that accuracy is self-graded, costs exclude ingest/judge, and baseline costs are compiled estimates.","tokens_in":10616,"tokens_out":4858,"duration_ms":49080,"significance":"If the reported accuracy and cost held under independent measurement, SodaMem would be a practically valuable point in the agent-memory design space: an order-of-magnitude cost reduction relative to Opus/GPT-4o-class systems while staying in their accuracy band. The paper's strengths are its transparent limitations, a clearly described architecture, a public code repository, and a cost-accuracy compilation that makes the comparison inspectable. However, the central comparative claims are currently supported only by a best-of-three self-graded accuracy number and heterogeneous cost estimates, so the significance is conditional on additional validation.","major_comments":[{"comment":"The headline 92.8% accuracy is the best of N=3 and is graded by the same deepseek-v4-flash model that produces the answers, as the paper acknowledges. Against the closest dominated point in Table 1 (Cersei Embed, 86.6%, $1.84/10^3Q), SodaMem's margins are only 6.2 accuracy points and $0.23/10^3Q. A few points of judge-induced shift (e.g., a stricter independent judge) or normal run-to-run variance would place SodaMem below Cersei Embed on accuracy or above it on cost, destroying the strict-dominance claim. Please report all N runs with per-run accuracy, add an independent judge (e.g., GPT-4o) with per-category breakdown, and give confidence intervals or a sensitivity analysis showing how the dominance region changes under plausible accuracy/cost perturbations.","section":"Experiments: Setup and Cost Protocol; Table 1"},{"comment":"The cost comparison is not apples-to-apples. SodaMem's cost is measured planner+reader usage excluding ingest and judge and is averaged over 500 questions, while baseline costs are author-reported USD, disclosed tokens repriced with a 90/10 input/output prior, or recalled-context-length lower bounds plus ~200 output tokens. Some baselines use different judges and evaluation protocols. Because the claimed dominance over Cersei Embed rests on a $0.23/10^3Q cost margin, the comparison could flip under a different token split or if SodaMem's ingest/judge costs are amortized. Please report full SodaMem cost including ingest and judge, state the sensitivity of Table 1 to the pricing assumptions, and mark which cells are lower bounds rather than point estimates.","section":"Experiments: Baseline cost estimation; Table 1"},{"comment":"The architecture introduces several tunable components—connection-density weights (0.4/0.2/0.1/0.05), time bonus beta=0.3, near-duplicate threshold theta=0.8, search-head count H=10, top-K—but the paper provides no ablations or parameter sensitivity. Without ablations that remove the graph tunnel, the BM25/dense tunnels, the time bonus, and the planner, the reported accuracy cannot be attributed to the temporal graph and density-fusion design; a simpler baseline at the same model and cost might match it. Please include at least a small ablation study and report the cost of each variant.","section":"Proposed Method: Retrieve, Eqs. (5)-(6); Implementation Notes"}],"minor_comments":[{"comment":"Typo: 'must rememberwhat' should be 'must remember what'.","section":"Abstract"},{"comment":"The footnote marker for SodaMem references a dagger, but the footnote appears below the table; clarify the distinction between 'Meas.' and 'Est.' in the Cost column, since the strict-dominance discussion relies on these labels.","section":"Table 1"},{"comment":"The notation hat_tau(f)=T(tau_raw(f), t_s(f), context(f)) is under-specified; please define T and context(f), and state how unresolved cases are represented.","section":"Eq. (4)"},{"comment":"The shaded 'strictly dominated' region is defined using the mean cost; since the median cost is about 25% lower, state explicitly which operating point the shaded region refers to and whether the region changes under the median.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is more a systems/empirical report than a theoretical contribution; its main risk is that the headline comparisons are within the noise of its own acknowledged measurement choices. I do not see a load-bearing error that would require rejection, but the paper needs independent judging and full-cost accounting before the frontier/dominance claims can be accepted. Recommend major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"SodaMem is a clean, honest engineering integration of ideas the agent-memory community already uses — extracted facts with mandatory provenance spans, temporal axes, supersession edges, hybrid BM25-dense retrieval, and a planner-reader loop. The combination is new, and the system description is clear enough to reimplement. I give it full credit for laying out the design and for being unusually upfront: the paper states that accuracy is self-graded, that costs exclude ingest and judge, and that baseline comparisons are compiled estimates rather than a unified bake-off. Released code and frozen-store fingerprints back that up.\n\nThe soft spot is the headline claim itself. The paper says SodaMem 'strictly dominates' several published points. The nearest such point, Cersei Embed at 86.6% and $1.84/10^3Q, is only 6.2 accuracy points and $0.23/10^3Q away from SodaMem's 92.8% and $1.61/10^3Q. Both margins sit inside the paper's own stated uncertainties. A judge-induced drop of a few points, or ingest/judge costs of a few tenths of a cent per question, moves SodaMem below Cersei Embed on one axis or the other. The stress-test note is right about that. Also, 92.8% is best-of-three; the median is 90.6%, which changes the picture again. So the 'frontier' positioning is a wish, not a result, unless independent judging confirms it.\n\nTwo smaller issues. The cost table mixes measured and estimated numbers from different protocols and judges; the paper calls this 'order-of-magnitude,' but then uses it to draw a dominated quadrant as if the precision were real. And there are no ablations — timeline resolution, density fusion, planner tools — so we don't know which components earn their keep. The paper defers these to follow-up, which is fine for a preprint, but it means the architecture's value is plausible, not demonstrated.\n\nWho should read this: people building memory systems for conversational agents. The design and cost accounting are a useful reference point. Take the 92.8% as 'self-graded on one benchmark,' not as a settled number.\n\nRecommendation: it deserves a serious referee. The paper is coherent, honest, and ships artifacts. A reviewer should push for independent judging (the released answer hypotheses make that cheap), a unified baseline harness, and component ablations. If those clear up, the paper becomes a solid systems contribution.","headline":"Useful system paper, but the headline cost-accuracy claim depends on a self-graded, best-of-three run and margins smaller than the paper's own uncertainties.","tokens_in":11161,"tokens_out":4168,"would_cite":true,"duration_ms":40507,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SodaMem claims that an evidence-grounded temporal graph memory—typed facts with provenance, timestamps, and supersession edges plus a planner–reader answer loop—reaches 92.8% accuracy on LongMemEval-S at a mean cost of $0.00161 per…","keywords":["temporal graph memory","LLM agents","FactEvent","supersession","provenance","LongMemEval","cost-accuracy","planner-reader"],"falsifier":"Re-grade SodaMem's released answer hypotheses with LongMemEval-S's official templates using an independent GPT-4o judge; if accuracy falls below 87.6% (the next comparable point in the cost table), the near-frontier claim fails. Alternatively, re-run every system in Table 1 in one harness with identical judge, protocol, and token accounting and check whether any point lands in the region SodaMem claims to dominate.","tokens_in":10187,"feed_emoji":"🧠","tokens_out":6704,"duration_ms":56979,"temperature":0.7,"pith_summary":"SodaMem is trying to show that an LLM agent's long-term memory should be a structured, evidence-grounded temporal graph rather than a flat chat log or Markdown diary, and that this design pays off in both accuracy and cost. On the LongMemEval-S benchmark it reports 92.8% accuracy (464/500) at a mean API cost of $0.00161 per question, using a small Flash-tier model for planning, reading, and judging. The paper argues this point sits near the accuracy frontier while costing roughly an order of magnitude less than higher-scoring systems that rely on Opus- or GPT-4o-class generators, and that it strictly dominates several published points that are both pricier and less accurate. A sympathetic reader would care because it suggests that careful memory structure—not just a bigger model or a longer context—may be the cheapest path to reliable long-horizon personal assistants.","feed_headline":"Temporal-graph memory hits 92.8% at $0.0016 per question","feed_subtitle":"SodaMem stores facts as dated, cited events with supersession edges, approaching Opus-class accuracy on cheap Flash-tier compute.","key_machinery":"The FactEvent is the load-bearing object: a typed record f = (κ, π, m, τ, ρ, S, σ) storing kind, predicate, modality, temporal fields, entity roles, source spans, and status, persisted in SQLite with hybrid BM25–dense indexes and typed edges such as SUPERSEDES, CONTRADICTS, and UPDATES. Retrieval runs three tunnels—graph/entity, BM25, and embedding—and fuses hits by connection density with a soft time bonus, so time acts as a ranking feature rather than a hard filter. A planner–reader loop then expands evidence via tools and produces answers with mandatory citations into the stored spans.","core_discovery":"The central discovery is that converting multi-session dialogue into typed FactEvents—each carrying a source span, a modality, mention/occurrence/validity time axes, entity roles, and a status—and connecting them with SUPERSEDES, CONTRADICTS, and UPDATES edges lets a cheap model answer long-horizon memory questions more accurately per dollar than flat retrieval. Superseding stale facts at write time, gating retrieval by validity, and ranking evidence by connection density across graph, BM25, and dense tunnels, with a planner–reader loop that gathers citable evidence before composing an answer, together yield 464/500 correct on LongMemEval-S. The paper further holds that under compiled public baseline costs this operating point is near the accuracy frontier at Flash-tier spend and strictly dominates several higher-cost, lower-accuracy systems.","pith_inferences":["If the 92.8% figure survives independent re-judging, a direct consequence is that memory architecture, not model scale, is the dominant cost lever for long-horizon personal QA; downstream systems could adopt the FactEvent contract while swapping in learned controllers.","The paper leaves timeline resolution optional and untested; a plausible extension is to ablate the session-anchored timeline-resolution layer to see how much of the remaining temporal-reasoning errors it removes, especially under an independent judge.","The cost table's dominated quadrant suggests that high-cost, lower-accuracy systems are not merely overpriced but architecturally avoidable; a testable prediction is that adding provenance and supersession to those baselines would lift their accuracy more per dollar than upgrading their generator.","Since ingest and judge costs are excluded, a fair end-to-end comparison that includes ingestion would be the next natural experiment; the paper's claims are about answer-stage economics, not total ownership cost."],"forward_implications":["A Flash-tier model with SodaMem can approach the accuracy of Opus/GPT-4o-class systems on LongMemEval-S at roughly 10–40x lower estimated cost per question.","Supersession and validity edges make what is currently true a deterministic store property rather than something an LLM infers from unordered chunks.","Mandatory provenance spans and cited reader answers make long-horizon memory auditable: every material claim can be traced to source turns.","The median cost of $0.00111 per question means typical queries are about 25% cheaper than the mean, so the system's cost profile is better than the headline suggests.","Retrieval ranking by connection density across graph, BM25, and dense tunnels with a soft time bonus keeps user-misdated queries recoverable."],"supporting_citations":[{"why":"defines LongMemEval-S, the benchmark whose 500 questions and official templates the evaluation uses.","marker":"(Wu et al. 2024)"},{"why":"introduces LoCoMo, the companion long-conversation benchmark motivating the memory design.","marker":"(Maharana et al. 2024)"},{"why":"establishes the need for external memory paging in long-horizon agents, which SodaMem's store builds on.","marker":"(Packer et al. 2023)"},{"why":"presents the Mem0 extractive memory API used as a baseline and as motivation for extraction-plus-store design.","marker":"(Chhikara et al. 2025)"},{"why":"describes the Zep temporal knowledge graph architecture that SodaMem extends with provenance and supersession.","marker":"(Rasmussen et al. 2025)"},{"why":"reports TiMem results and recalled-context lengths that the paper prices into the cost table.","marker":"(Li et al. 2026)"},{"why":"supplies the disclosed context-token figures used to estimate costs for MemOS, Zep, Supermemory, and other baselines.","marker":"(MemTensor 2025)"},{"why":"reports the highest-accuracy baseline in the table, setting the frontier point SodaMem says it approaches.","marker":"(McCann 2026)"},{"why":"reports Mem0's 2026 research token counts and accuracy, the second-highest baseline in the cost comparison.","marker":"(Mem0 Research 2026)"},{"why":"reports Cersei's LongMemEval results with measured costs, placing systems in the quadrant SodaMem claims to dominate.","marker":"(Pacifio 2026)"}],"fun_headline_variants":["Graph memory with cited facts hits 92.8% at sub-cent cost","Dated, sourced events: cheap model nails long-term recall","SodaMem: temporal graph beats flat logs per dollar","Let Flash-tier models supersede stale facts cheaply"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline accuracy is self-graded by the same Flash model that writes the answers, and the baseline costs are compiled from public disclosures rather than a single uniform re-run; if an independent judge changes the scores or the cost estimates are off, the claimed frontier position and strict dominance do not hold.","fun_headline_variants_meta":{"raw":{"variants":["Graph memory with cited facts hits 92.8% at sub-cent cost","Dated, sourced events: cheap model nails long-term recall","SodaMem: temporal graph beats flat logs per dollar","Let Flash-tier models supersede stale facts cheaply"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000301,"raw_usage":{"total_tokens":1782,"prompt_tokens":1037,"completion_tokens":745,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":653,"completion_tokens_details":{"reasoning_tokens":674}},"tokens_in":653,"tokens_out":745,"duration_ms":8496,"temperature":1.0,"reasoning_tokens":674,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:28:58.856959+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-grade SodaMem's released answer hypotheses with LongMemEval-S's official templates using an independent GPT-4o judge; if accuracy falls below 87.6% (the next comparable point in the cost table), the near-frontier claim fails. Alternatively, re-run every system in Table 1 in one harness with identical judge, protocol, and token accounting and check whether any point lands in the region SodaMem claims to dominate.","supporting_citations":[],"review_version":1}