{"id":"15be70fb-d922-4966-87b8-a1db4d3043e4","arxiv_id":"2508.11310","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"SGSimEval is a multifaceted benchmark showing that automatic survey generation systems match humans on outline quality but lag on content and references.","lead":"SGSimEval is a new benchmark for testing systems that automatically write academic survey papers, judging their outlines, content, and references with a mix of LLM ratings, quantitative similarity scores, and human-preference data. Early results say current systems already match or beat human-written outlines but lag on content and references, which makes this useful for anyone building or buying survey-generation tools.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central consistency claim unverifiable in provided text; main risk is circular use of human ratings to validate a metric they helped define.","rationale":"The reader's weakest_assumption already identifies the core risk: if human annotations were used to tune the metric and then cited as evidence that the metric matches humans, the consistency claim would be inflated. My concern is the same. I do not assert that the authors committed this error; the provided text is simply too incomplete to rule it out. The concrete test resolves the ambiguity by checking for a held-out human-annotation split. If such a split exists, the abstract's central claim can be provisionally trusted. If it does not, the benchmark's central claim is unsupported. Either way, the appropriate verdict remains UNVERDICTED given that the methodology and experiments cannot be inspected. I therefore recommend no change to the reader's verdict.","tokens_in":3483,"tokens_out":4022,"duration_ms":41567,"concrete_test":"Inspect the full paper's methodology (Sections 3-4) and determine whether the human preference judgments were split into a development set, used to construct/tune SGSimEval components, and a held-out test set, used only to measure consistency. If no such split exists, or if the reported 'strong consistency' was computed on the development data, recompute the correlation between SGSimEval and human judgments on a fresh held-out set of generated surveys; if the held-out correlation is significantly lower, the abstract's claim is inflated. If a clean held-out split is present, the concern does not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's load-bearing claim is that 'our evaluation metrics maintain strong consistency with human assessments,' and the conclusion adds that 'most ASG systems exceeding human performance in outline generation.' Both claims depend entirely on the validity of SGSimEval's composite metric. The provided text contains no metric formulas, no annotation protocol, no inter-annotator agreement statistics, and no correlation coefficients; Sections 2-6 are absent. The most concrete risk is circular validation: SGSimEval 'introduce[s] human preference metrics that emphasize both inherent quality and similarity to humans.' If the same human judgments used to construct or weight the metric (e.g., choosing similarity weights, LLM-score prompts, or composite thresholds) are then reused to report 'strong consistency with human assessments,' that consistency is inflated by construction. This is especially salient because the abstract criticizes 'an over-reliance on LLMs-as-judges,' yet SGSimEval combines 'LLM-based scoring with quantitative metrics'; no independent validation of the LLM-score component is visible. Without a held-out split of human annotations, the headline consistency claim cannot be distinguished from a self-fulfilling correlation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SGSimEval, a benchmark for automatic survey generation (ASG) evaluation that assesses outline, content, and references by combining LLM-based scoring with quantitative similarity metrics and human-preference dimensions. The abstract claims that current ASG systems show 'human-comparable superiority' in outline generation, that content and reference generation remain significantly weaker, and that SGSimEval's metrics 'maintain strong consistency with human assessments.' The visible text includes the title, abstract, the opening of Section 1, the conclusion, acknowledgments, and references; Sections 2–6, which should contain the benchmark design, metric formulas, annotation protocol, and experimental results, are absent from the submitted material.","tokens_in":3620,"tokens_out":3115,"duration_ms":38475,"significance":"If fully substantiated, SGSimEval would be a useful contribution to an emerging evaluation area: it explicitly targets a multidimensional view (outline, content, references), attempts to combine LLM judgments with quantitative similarity to human-written surveys, and introduces human-preference metrics, thereby addressing a real limitation of existing ASG evaluations. The benchmark could be a resource for comparing ASG systems and tracking progress. However, at present all central claims rest on evidence that is not visible in the manuscript: no metric definitions, no annotation protocol, no inter-annotator agreement, no correlation coefficients, and no significance tests are shown. The manuscript reads as a heavily truncated version, and the empirical claims cannot be verified in this form.","major_comments":[{"comment":"The load-bearing claim that the evaluation metrics 'maintain strong consistency with human assessments' is stated without any supporting statistics. The manuscript provides no correlation coefficient, inter-annotator agreement (e.g., Cohen's kappa or ICC), confidence interval, or error analysis. In a benchmark paper, this is the central validation result; it must be reported with full details, sample sizes, and statistical precision. Otherwise the claim is unfalsifiable as presented.","section":"Abstract; Conclusion"},{"comment":"There is a circular-validation risk: SGSimEval 'introduce[s] human preference metrics that emphasize both inherent quality and similarity to humans,' and the composite metric combines LLM-based scoring with quantitative similarity. If the same human judgments used to calibrate the metric (e.g., choosing the combination weight between LLM scores and similarity scores, or the similarity threshold for reference appropriateness) are then reused to demonstrate 'strong consistency with human assessments,' the consistency is inflated by construction. The paper must specify which human annotations are used for construction/calibration and which are held out for validation, and report the held-out correlation.","section":"Abstract"},{"comment":"The claim that 'most ASG systems exceeding human performance in outline generation' is a strong empirical statement, but no evidence is provided: there is no definition of the human baseline, no per-system score table, no significance tests, and no error bars or confidence intervals. Without such statistical support, 'exceeding human performance' could be within noise. The authors should specify the exact comparison procedure and report effect sizes and uncertainty.","section":"Conclusion"},{"comment":"The main body of the paper is absent from the submitted text. Metric formulas, the evaluation protocol, the human-annotation procedure, the dataset description, and the experimental setup are all required to assess the validity of the benchmark. In particular, the exact form of the 'similarity-enhanced' score and the way it is fused with LLM-based scoring must be given, along with all free parameters (e.g., combination weights, thresholds) and how they are chosen.","section":"§2–§6 (missing)"}],"minor_comments":[{"comment":"The phrase 'human-comparable superiority' is ambiguous: it is unclear whether systems are comparable to humans, superior to humans, or both in different respects. Please rephrase.","section":"Abstract"},{"comment":"The visible text jumps from the first paragraph of Section 1 to the conclusion, with the running head 'SGSimEval 13' on the conclusion page. Ensure the submitted version includes all sections (2–6) and is not accidentally truncated.","section":"General"},{"comment":"No reproducibility statement, dataset URL, or code link appears in the visible text. For a benchmark paper, these are important; please add a reproducibility section.","section":"General"},{"comment":"Reference [15] has inconsistent author formatting ('Qwen, :,' followed by all authors). Standardize the bibliography style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript as received is incomplete, with Sections 2–6 missing, so I cannot verify the core contribution. The largest concern is circularity: the human-preference data used to construct or calibrate the composite metric may also be cited as evidence of its validity. If the authors can provide the missing sections, a clear separation of calibration vs. held-out human annotations, and proper statistical reporting, the paper could become a solid benchmark contribution. Major revision is appropriate rather than rejection because the identified issues are potentially fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You asked what I think of the ASG evaluation benchmark. The short version: the authors are pointing at a real gap—there isn't a widely accepted multi-facet benchmark for automatic survey generation—and their proposed structure (outline, content, references; LLM score plus similarity to human-written surveys plus human preference) is a reasonable response. That’s the most valuable bit.\n\nWhat’s good: they explicitly critique over-reliance on LLM-as-judge and try to couple it with quantitative similarity and human preference. Five representative systems compared. The conclusion that outline generation is basically solved while content and reference curation lag is an interesting, falsifiable empirical finding.\n\nWhat I can’t check: the entire methodology (sections 2–6) is missing from the provided text. The abstract’s “strong consistency with human assessments” is the load-bearing claim, but no numbers accompany it—no correlation, no inter-annotator agreement, no error analysis. The conclusion restates the claim without statistics.\n\nThe real soft spot is circularity. They say they introduce “human preference metrics that emphasize both inherent quality and similarity to humans.” If those same human judgments are used to set weights or thresholds on the metric and then the metric is validated against the same human judgments, the “consistency” is inflated. The abstract doesn’t reassure me on this. Also, the reliance on human-written reference surveys assumes those are a stable gold standard, which needs defense.\n\nOn the citation pattern: the references are mostly relevant—ASG systems, general LLM benchmarks, RAG surveys. There are three self-citations to prior summarization work (refs 10–12), which look legitimate for grounding their own prior work but don’t substitute for a data/code link. I see no code or data URL in the visible text, which lowers reproducibility.\n\nOverall: this is a plausible, well-scoped contribution to a niche area. The central claim is currently unverifiable from what I can read, and the circularity risk needs a clear answer. But it deserves a proper peer review—the authors have a concrete benchmark idea, and the weakness I see is in the validation protocol, which is fixable.\n\nRecommendation: send it to review, but insist that the authors report held-out human annotations, inter-annotator agreement, and a separation between metric construction and validation. Read the full method before trusting any number.","headline":"A well-motivated ASG benchmark with a reasonable multi-facet design, but the headline consistency claim rests on numbers that aren't visible and a circularity risk that needs a direct answer.","tokens_in":4210,"tokens_out":1895,"would_cite":false,"duration_ms":20361,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims automatic survey generation can be reliably scored by a three-part benchmark that matches human judgment, and that current systems already beat humans at outlines but lag on content and references.","keywords":["Automatic survey generation","LLM evaluation benchmark","similarity-enhanced metrics","human preference evaluation","outline generation","reference curation","semantic similarity","large language models"],"falsifier":"Take SGSimEval to a new set of topics that are absent from its reference collection, generate surveys with several ASG systems, and collect independent human pairwise preferences. If the benchmark's composite scores do not rank the systems in line with those fresh human judgments—say, rank correlation near zero or negative—the claimed strong consistency with human assessment fails to transfer.","tokens_in":3288,"feed_emoji":"📚","tokens_out":3992,"duration_ms":44407,"temperature":0.7,"pith_summary":"Automatic survey generation—using LLMs to write academic literature reviews—has become practical, but the field has lacked a trustworthy way to grade the results. This paper introduces SGSimEval, a benchmark that scores generated surveys on three independent axes—outline, content, and references—and combines LLM-based scoring, quantitative similarity to human-written reference surveys, and human preference. The paper reports two main findings: current ASG systems reach or exceed human-level quality in outline generation, while content and reference curation still trail; and SGSimEval's scores agree strongly with human assessments. The intended contribution is an evaluation standard that future survey-generation systems can use instead of relying on any single LLM judge.","feed_headline":"AI survey generators beat humans at outlines","feed_subtitle":"A three-part evaluation benchmark finds outline structure is solved, while content and reference quality still lag behind humans.","key_machinery":"SGSimEval—Survey Generation with Similarity-Enhanced Evaluation—is the central benchmark. It evaluates each generated survey along three facets (outline, content, references) and fuses three evidence sources: LLM-based scores, quantitative semantic similarity to human-written reference surveys, and human preference judgments. The similarity component is what gives the benchmark its name and is meant to counter the bias and over-reliance on LLMs-as-judges found in prior evaluation methods.","core_discovery":"SGSimEval claims that the quality of an automatic survey can be measured by looking at three components—outline structure, content adequacy, and reference appropriateness—and by combining LLM scoring, semantic similarity to human-written reference surveys, and human preference. Used on five representative ASG systems, the benchmark leads to the finding that CS-specialized systems consistently outperform general-domain approaches, most systems exceed human performance in outline generation, and content and reference generation show significant room for improvement. The paper further claims that its evaluation metrics maintain strong consistency with human assessments.","pith_inferences":["One implication the paper leaves implicit: if outline generation is solved, evaluation research should shift from structure to claim-level verification, with a testable goal of checking each citation against the source it supposedly supports.","The similarity-to-human-surveys component may reward conventional organization; an extension would test whether deliberately unconventional but factually accurate surveys are unfairly penalized relative to human preference.","Because the benchmark includes LLM-as-judge scores while criticizing over-reliance on LLMs-as-judges, a natural ablative experiment is to recompute agreement with human ratings with and without the LLM score component.","The human-comparable outline result suggests a near-term practical use: ASG systems could serve as outline generators for human authors, who then write and verify the content themselves."],"forward_implications":["If outline generation is effectively solved at human level, future ASG development should concentrate on content verification and reference curation rather than outline design.","Because CS-specialized systems outperform general-domain systems, building domain-tuned components appears to be a productive direction for other scientific fields.","Evaluation of survey generation should combine multiple evidence types; single-metric or pure LLM-as-judge evaluations are likely insufficient.","The large gap between outline quality and content/reference quality means generated surveys may look well structured even when their details and citations are unreliable, so human verification remains necessary."],"supporting_citations":[],"fun_headline_variants":["Survey AIs ace outlines, flunk content and refs","New benchmark: AI survey makers beat humans on outlines","SGSimEval: AI outlines rival humans, content lags","Automatic survey systems: outlines done, content not","Benchmark finds AI survey generators strong on structure"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The benchmark's validity rests on the assumption that human-written surveys and human preference judgments form a reliable gold standard, and that fusing them with LLM-based scores faithfully captures survey quality; if the human references are not actually good surveys or human raters disagree, the claimed consistency with human assessment loses its footing.","fun_headline_variants_meta":{"raw":{"variants":["Survey AIs ace outlines, flunk content and refs","New benchmark: AI survey makers beat humans on outlines","SGSimEval: AI outlines rival humans, content lags","Automatic survey systems: outlines done, content not","Benchmark finds AI survey generators strong on structure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000145,"raw_usage":{"total_tokens":999,"prompt_tokens":708,"completion_tokens":291,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":452,"completion_tokens_details":{"reasoning_tokens":211}},"tokens_in":452,"tokens_out":291,"duration_ms":3698,"temperature":1.0,"reasoning_tokens":211,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:00:01.777786+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take SGSimEval to a new set of topics that are absent from its reference collection, generate surveys with several ASG systems, and collect independent human pairwise preferences. If the benchmark's composite scores do not rank the systems in line with those fresh human judgments—say, rank correlation near zero or negative—the claimed strong consistency with human assessment fails to transfer.","supporting_citations":[],"review_version":1}