{"id":"5816a5d5-7ede-4ea4-a91c-88e776484417","arxiv_id":"2506.23393","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Generating Wikipedia articles with factoid memory units organized into a hierarchical outline improves informativeness, verifiability, and citation coverage over RAG and STORM baselines.","lead":"This paper introduces MOG, a system that turns web pages into small factoids, clusters those factoids into a Wikipedia-style outline, and writes each article section from the facts assigned to it. The work matters because it offers a route toward more trustworthy automated long-form writing and a new low-resource benchmark for Wikipedia generation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Verifiability metrics are likely a paraphrase tautology: entailment is judged against the exact memory units used to write each sentence, so the reported SOTA citation scores may not measure real-world verifiability.","rationale":"The reader's weakest_assumption is that the chosen metrics may not measure what matters for Wikipedia quality, specifically citing LLM-judged citation metrics without human validation and length-sensitive informativeness counts. My stress test sharpens this: the citation metric in MOG is not merely unvalidated, it is structurally circular. The generation prompt forces sentences to be derived from memory units, and the citation module selects those same units as sources; the entailment judge then verifies paraphrase consistency, not factual verifiability. This is a correctness risk to the strongest claim (SOTA verifiability), not just an external-validity quibble. The informativeness metric is also length-sensitive, though the paper provides some density argument in text; a per-word density number would address it. Because the reader already flagged this and set the verdict to CONDITIONAL, my analysis does not change the verdict; it strengthens the condition needed: independent human or third-party verification of citation support. The baselines-only issue (RAG and STORM; no WebBrain) is secondary but real; tempering the claim to 'best-among-tested' would be safer. Overall, the system design is coherent, ablations support the components, and the WikiStart dataset construction is reasonable, so the paper is acceptable conditionally pending metric validation.","tokens_in":23615,"tokens_out":6437,"duration_ms":70160,"concrete_test":"Randomly sample 100 sentences per system from WikiStart generated articles; for each, give two human annotators the sentence and the full cited source document (not the memory unit) and ask whether the source supports the sentence. Compute human-rated Citation Recall/Precision for MOG, RAG, and STORM, plus agreement with gpt-4o-mini scores. If MOG's human-rated verifiability is not significantly above the baselines, or if the LLM judge agrees with humans on RAG/STORM but over-scores MOG, the reported advantage is a measurement artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of state-of-the-art verifiability rests on Citation Recall/Precision computed by gpt-4o-mini entailment judgments (Section 4.3). MOG's SectionWriter prompt (Listing 2) instructs the model to write each section strictly from a given list of atomic facts, and the CitationFinder prompt then selects 'the most relevant sources to fully cover the sentence' from exactly those memory units. The entailment judge therefore checks whether a sentence paraphrases the same memory unit that was used to compose it, making high Citation Recall nearly guaranteed by construction. RAG and STORM write from larger document chunks, where a single chunk less often entails the full sentence, so the metric chiefly rewards fine-grained citation granularity rather than verifiability. No human validation of the entailment judge is provided, and the memory-extraction step itself (the atomic facts and their source fidelity) is never evaluated. Informativeness metrics suffer from a similar confound: entity and numerical counts are raw counts, and MOG's outputs are 22-24% longer than RAG's on WikiStart, so the reported 79% entity-count gain may be largely length-driven; the claimed density advantage is argued in prose but not reported as a metric.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MOG (Memory Organization-based Generation), a framework for generating Wikipedia articles from web sources. MOG extracts fine-grained atomic memory units from retrieved documents, recursively clusters and summarizes them into a hierarchical outline, and then generates each section from the memory units assigned to it, followed by a post-hoc citation module that links sentences to memory units. The authors introduce WikiStart, a low-resource dataset of 100 Wikipedia stub/start-class topics, and evaluate MOG against RAG and STORM on WikiStart and FreshWiki. They report that MOG improves informativeness (word, entity, and numerical counts) and verifiability (Citation Recall, Citation Precision, Citation Rate) while preserving fluency, and they include ablations, memory-utilization analysis, prompts, and a full generated article example. The central claim is that MOG achieves state-of-the-art performance in informativeness and verifiability.","tokens_in":23830,"tokens_out":5891,"duration_ms":64644,"significance":"The hierarchical memory organization is a reasonable and well-motivated design: aligning the outline with the evidence base and using fine-grained memory units is a genuine step beyond chunk-level retrieval, and the WikiStart dataset targets a real gap in low-resource evaluation. The paper is also unusually transparent in shipping its prompts, code, and a full generated example, and the ablation study and memory-utilization analysis are useful. If the evaluation were trustworthy, the large Citation Recall gains (about 10 points on FreshWiki and 15 points on WikiStart) would be a notable result. However, as detailed in the major comments, the verifiability metric is likely inflated by construction, and informativeness is measured with raw length-driven counts. These issues directly affect the paper's third contribution, so the state-of-the-art claim is not currently established. The work is worth revising rather than dismissing, because the architecture and dataset have independent value.","major_comments":[{"comment":"The Citation Recall/Precision values in Table 2 may be largely a tautology. In MOG, the SectionWriter prompt instructs the model to write each section from a given list of atomic facts (Listing 2), and the CitationFinder then selects 'the most relevant sources to fully cover the sentence' from that same list. The gpt-4o-mini entailment judge therefore checks whether the sentence paraphrases the very memory unit that was used to compose it, so high Citation Recall is expected by construction. RAG and STORM cite document chunks and are judged against larger, less perfectly aligned units, so the comparison rewards MOG's citation granularity rather than measuring verifiability against the original web sources. No human validation of the entailment judge is provided, and the memory-extraction step (fidelity of atomic facts to source documents) is never evaluated. I recommend reporting citation metrics against the original retrieved documents and adding a human annotation study of citation support on a sample of sentences.","section":"§4.3, §3.4, Listing 2"},{"comment":"Informativeness is measured with raw word, entity, and numerical counts. Because MOG's outputs are 22–24% longer than RAG's on WikiStart (Table 2b), the reported 79% entity-count gain cannot be attributed to content density without normalization. The statement in §5.1 that the gain is 'not merely due to longer outputs' does not follow from the data presented. Please report density measures (e.g., entities per 1,000 words, numerical values per 1,000 words), duplicate-content rates, and ideally human judgments of coverage and redundancy. In addition, 'state-of-the-art' in Contribution 3 is stronger than the evidence supports with only RAG and STORM as baselines, both of which are not recent systems.","section":"§4.3, Table 2"},{"comment":"The sample generated article contains internal contradictions and apparent factual errors that undermine the claim of minimized hallucinations. The lead states the 2023 SEA Games featured over 7,000 athletes and a record 632 events, while later sections say 'over 6,000 athletes from 11 nations competed in 584 events', 'approximately 8,000 athletes', and '5,151 medals across 37 events' alongside '608 sets of medals'. The 'Notable Teams and Athletes' subsection names Neeraj Chopra and Nikhat Zareen, who are not associated with the 2023 SEA Games. If this sample is representative, the automatic evaluation is not capturing serious consistency and factuality failures. A consistency/factuality evaluation (human or automated over the full test set) is needed before claiming that MOG reduces hallucinations.","section":"Appendix D"},{"comment":"The ablation study supports the Subtopic Explorer's contribution to entity count, but the Memory Organization ablation is weakly supported. Removing MO lowers word and entity counts relative to full MOG, yet Citation Recall and Citation Precision are actually higher without MO (88.59/81.39 vs. 85.72/79.68 in Table 6), and no significance tests are reported for any ablation. The text interprets the section-count difference (9.78 for w/o SE vs. 8.18 for w/o MO) as evidence for MO's effectiveness, but these configurations differ in multiple ways. Please report statistical tests and separate the effect of MO on organizational quality (e.g., heading relevance, section coherence) from its effect on raw counts.","section":"§6.2, Tables 4 and 6"}],"minor_comments":[{"comment":"The text says 'memory units achieve a higher utilization rate', but the table reports webpages collected and cited; please clarify whether the utilization rate is computed over memory units or webpages and define the denominator.","section":"Table 3 and §6.1"},{"comment":"The phrase 'indicating the gains are not merely due to longer outputs' needs a density metric or a statistical control; as written it is an unsupported interpretation.","section":"§5.1"},{"comment":"The LLM-based metric scores (Interest, Organization, Focus) are reported without confidence intervals or significance tests, so differences from baselines are hard to assess.","section":"Figure 4"},{"comment":"The hyperparameters max queries=2, max webpages=3, and max subtopic depth=2 are said to be set by preliminary testing, but no sensitivity analysis is provided; because these control the amount of memory, their influence on the headline results is unquantified.","section":"Appendix B.2"},{"comment":"The reference list contains two Prometheus entries, one for Kim et al. (2023) and one for Kim et al. (2024); only the latter appears to be used in the paper, so the unused entry should be removed or cited.","section":"References"},{"comment":"The generated article's citation markers such as '[6,8,9,10,11]' are not explained: it should be stated whether the numbers index memory units, source documents, or section-level reference lists, since the citation module claims traceability to original sources.","section":"Appendix D"}],"recommendation":"major_revision","confidential_remarks":"The architecture and dataset are promising, but the two central evaluation claims (verifiability and informativeness) are confounded by the metrics as currently designed. The revision should include an independent citation-support evaluation against original documents and a human study, plus normalized informativeness metrics. If those results continue to show an advantage, the paper could be acceptable; without them, the state-of-the-art claim is not supported. I am not recommending rejection because the issues are fixable within the scope of the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read MOG. The core idea—factoids clustered recursively into a heading tree that then drives section-by-section generation—is a genuine integration, and WikiStart is a real contribution for low-resource evaluation. The ablations give some support for the subtopic explorer and memory organization. Credit where due: code is public, limitations are honestly stated.\n\nThe soft spot is exactly what the stress-test flagged. The verifiability metrics are close to a tautology. The SectionWriter prompt tells the model to write only from the supplied atomic facts; the CitationFinder then selects from exactly those atomic facts; the entailer checks whether the sentence is supported by the cited factoid. Since the sentence was written from that factoid, high Citation Recall is nearly built in. This measures cite-to-source paraphrase fidelity, not whether the claim is true or that the source actually supports it in context. There is no human validation of the entailer, and the extraction step—whether atomic facts faithfully represent the web documents—is never evaluated. So the headline 'state-of-the-art verifiability' is not established.\n\nOther issues: only two baselines, no WebBrain, no variance reporting, and hyperparameters tuned without a clear held-out protocol. The entity-count claim is partly length-driven, though 79% more entities on 22% more words is not nothing.\n\nNone of this kills the paper. The design is sensible, the dataset is reusable, and the framing is honest. But the evaluation needs real work: human or independent citation judgment, a memory-extraction fidelity check, and at least one stronger baseline. I'd send it to review with that expectation.","headline":"The recursive memory organization and WikiStart dataset are genuine contributions, but the verifiability claim is inflated because the citation metric is nearly a paraphrase tautology.","tokens_in":24401,"tokens_out":2299,"would_cite":true,"duration_ms":26135,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper sets out to show that Wikipedia article generation improves when retrieved web documents are reduced to fine-grained factoid memory units, recursively organized into a Wikipedia-style outline, and then written into sections that…","keywords":["Wikipedia generation","hierarchical memory","factoid extraction","retrieval-augmented generation","citation generation","long-form text generation","low-resource benchmark","WikiStart"],"falsifier":"Conduct a blind human evaluation on a sample of the generated articles, scoring each sentence for whether the cited source supports it and whether the article omits important information; if human citation-support ratings do not track gpt-4o-mini citation recall, or if MOG's entity and numerical gains disappear after matching output length, then the paper's central comparison is not established.","tokens_in":23397,"feed_emoji":"🧠","tokens_out":7907,"duration_ms":74416,"temperature":0.7,"pith_summary":"This paper argues that the bottleneck in autonomous Wikipedia writing is not retrieval but organization. It proposes MOG, which breaks web documents into fine-grained factoid memory units, recursively clusters and summarizes them into a Wikipedia-style outline, and then generates each section from the memory units assigned to it. The paper claims that on both the high-resource FreshWiki set and its new low-resource WikiStart benchmark, MOG beats a plain RAG pipeline and the STORM system on informativeness (word, entity, and numerical counts) and on verifiability (citation recall, precision, and rate). If true, this means the outline should be derived from memory rather than planned separately, and that low-resource Wikipedia stubs can be expanded into draft articles whose sentences point back to specific source factoids.","feed_headline":"Memory unit hierarchy outwrites RAG and STORM in Wikipedia generation","feed_subtitle":"Converting web pages into factoids and organizing them as the outline lifts citation recall and information density.","key_machinery":"The load-bearing mechanism is the memory unit, defined as a self-contained natural-language factoid extracted from a web document, together with the recursive memory organization algorithm built on five operations: save, recall, extract, cluster, and summarize. The organization pass clusters memory units by embedding similarity, summarizes each cluster, asks an LLM for subsection headings, and assigns every unit to the heading with which it shares the highest semantic similarity, recursing inside sections that need more depth. This turns the memory set into an outline hierarchy in which each section has non-overlapping, directly supporting evidence, so the generated article is constrained to what the sources actually say.","core_discovery":"The central claim is that a hierarchical memory whose structure mirrors the target article's layout removes the outline-memory misalignment that makes RAG-style Wikipedia generation hallucinate or omit information. MOG first extracts self-contained factoids from retrieved documents, then recursively clusters and summarizes them so that each cluster becomes a section heading and its members become that section's supporting evidence. Generation proceeds section by section from these preallocated memory units, and a post-hoc citation module attaches each sentence to the most relevant units. On FreshWiki and WikiStart, the paper reports higher section counts, word, entity, and numerical counts, and higher citation recall, precision, and rate than RAG and STORM, with a smaller drop in citation recall when moving from high- to low-resource settings.","pith_inferences":["A sharper test of the core idea would hold article length fixed and compare MOG's output against length-matched baselines; if the entity and numerical advantages vanish under length control, the informativeness win is mainly a volume win.","Because the paper's verifiability numbers come from an LLM entailment judge with no reported human agreement, a follow-up with human fact-checkers would tell whether the citation-recall gap reflects real verifiability or the judge's preference for fine-grained citations.","The factoid-memory design is likely to matter most when source documents are noisy or conflicting; a deliberate stress test with contradictory sources would separate the organizing mechanism from the extraction quality.","The paper flags that factoid units can drop temporal and sequential details; applying MOG to narrative domains such as sports recaps or financial disclosures would reveal whether outline-aligned memory needs richer unit types."],"forward_implications":["A Wikipedia draft produced by MOG carries a citation after every generated sentence, and each citation points to the specific factoid that supports it, so readers can check claims without reading whole documents.","The same recursive organization should carry over to other long-form structured writing tasks, such as reports, surveys, or biographies, where the output must be comprehensive and source-backed.","On roughly half of Wikipedia entries, which are stub-level articles with scattered sources, MOG reports smaller citation-recall drops than the baselines, suggesting it is better suited to real-world expansion work.","Because MOG uses memory units rather than document chunks as the working context, it can draw on up to 100 sources within the same context budget that limits RAG and STORM to about 5.","The new WikiStart benchmark of 100 low-resource topics gives future systems a reusable testbed where the reference articles are short and the information is dispersed."],"supporting_citations":[{"why":"Supplies the STORM baseline, the FreshWiki dataset, and the section-by-section generation and article-length truncation conventions the evaluation builds on.","marker":"Shao et al. (2024)"},{"why":"Defines Citation Recall and Citation Precision, the verifiability metrics behind the headline comparison.","marker":"Gao et al. (2023a)"},{"why":"Provides the sub-aspect explorer idea that MOG adapts into its subtopic explorer for broad document collection.","marker":"Wang et al. (2024)"},{"why":"Contributes the section-by-section biography generation strategy that MOG follows for article writing.","marker":"Fan and Gardent (2022)"},{"why":"Supports the choice of fine-grained factoids as memory units rather than entire document chunks.","marker":"Chen et al. (2023)"}],"fun_headline_variants":["Hierarchical memory beats RAG and STORM on Wikipedia","Factoid hierarchy aligns outline, cuts hallucinations","Memory organization lifts citation recall in Wikipedia","MOG: hierarchical memory for verifiable Wikipedia articles"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline advantage rests on the assumption that the automatic metrics used in Table 2, namely citation recall and precision judged by gpt-4o-mini's yes/no entailment with no human validation and informativeness counted as words, entities, and numbers, actually capture Wikipedia quality; if the judge systematically favors MOG's fine-grained citations, or if longer output alone drives the count metrics, the claimed state-of-the-art result could be an artifact of measurement.","fun_headline_variants_meta":{"raw":{"variants":["Hierarchical memory beats RAG and STORM on Wikipedia","Factoid hierarchy aligns outline, cuts hallucinations","Memory organization lifts citation recall in Wikipedia","MOG: hierarchical memory for verifiable Wikipedia articles"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000166,"raw_usage":{"total_tokens":1193,"prompt_tokens":821,"completion_tokens":372,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":437,"completion_tokens_details":{"reasoning_tokens":313}},"tokens_in":437,"tokens_out":372,"duration_ms":4347,"temperature":1.0,"reasoning_tokens":313,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:44:15.556021+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct a blind human evaluation on a sample of the generated articles, scoring each sentence for whether the cited source supports it and whether the article omits important information; if human citation-support ratings do not track gpt-4o-mini citation recall, or if MOG's entity and numerical gains disappear after matching output length, then the paper's central comparison is not established.","supporting_citations":[],"review_version":1}