{"id":"cbaa972e-ad63-4355-82d9-a142effed28e","arxiv_id":"2411.09297","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new dynamic-granularity timeline summarization benchmark with event-atom metrics shows LLM-based methods beat earlier TLS baselines yet still struggle with informativeness and granularity consistency.","lead":"The authors introduce a benchmark task where news timelines are generated at a user-specified level of detail, from broad overviews to fine-grained event chains. They test large language models with four scoring dimensions and find that even the strongest models cannot consistently match the requested granularity while staying informative.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The gold G10/G5 references are produced by GPT-4o role-play with only 45% full agreement and no independent human validation, so the core Info/Granu metrics may rank systems by GPT-4o salience rather than human salience.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the multi-granularity gold references are produced by GPT-4o role-play, and the human-alignment study validates the metric scores but not the reference-construction process. I agree with that assessment and find it well supported by the manuscript itself. Table 6 shows only 45.09% full agreement among the three GPT-4o annotators, which makes the unvalidated assumption about human salience especially salient. The concern is load-bearing because Informativeness and Granular Consistency, the two metrics that are most responsible for the paper's dynamic-granularity contribution, are defined against these references. Factuality uses source articles rather than the gold timelines, and Coherence uses a GPT-4o rubric, so the reference-bias problem directly attacks the metrics that make the benchmark novel. I considered whether the small human-alignment sample (50 timelines) is a more serious issue, but that is a validation-size problem that can be fixed by collecting more data; the reference-bias problem is structural and would invalidate the benchmark's gold labels if it lands. The expert-refinement step described in Section 5.2 could mitigate the bias, but no quantitative evidence is provided to show that the final references differ from or improve upon the GPT-4o selections. The proposed concrete test, an independent human re-annotation study plus a ranking-stability check, would settle the matter directly. Because the paper already receives a CONDITIONAL verdict and the required condition is exactly this kind of independent validation, my stress-test does not move the verdict: it should remain conditional until the check is performed. If the check fails, I would move to UNVERDICTED for the benchmark-standard claim; if it passes, the conditional could be lifted toward acceptance.","tokens_in":21864,"tokens_out":6468,"duration_ms":67071,"concrete_test":"Independently re-annotate a random sample of 50 topics: recruit trained journalist annotators to construct G10 and G5 timelines directly from the fine-grained event atoms and source articles, blind to the GPT-4o selections and to each other. Compute (i) human-human agreement using event-atom-group F1 or Kendall rank correlation of selected groups, and (ii) human-published-reference agreement. Then rerun the main comparison from Table 2 (e.g., HM GPT-3.5-Turbo vs. Datewise and Clustering) scoring against the independent human references instead of the GPT-4o references. If human-published agreement falls within the human-human agreement band and the relative system rankings by Info and Granu are unchanged, the concern is resolved; if not, the gold references encode a model-specific bias and the benchmark-standard claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The benchmark's core new metrics, Informativeness (Eq. 9) and Granular Consistency (Eq. 12), are computed against the gold G10 and G5 timelines. Section 5.2 and Appendix C.2 construct those gold timelines by having three GPT-4o role-players select salient event-atom groups from the fine-grained timeline, followed by expert refinement. Table 6 reports only 45.09% full agreement among the three GPT-4o annotators, and no inter-annotator reliability is reported between the expert-refined references and independent human annotations. The human-alignment study in Section 7.3 and Appendix F validates the metric scores on 50 timelines, but it does not validate the reference timelines themselves. Consequently, the central claim that DTELS-Bench can serve as a standard evaluation for controllable timeline generation rests on the unexamined assumption that GPT-4o's expert-refined salience ranking matches human salience. If that assumption fails, the Info and Granu scores, and the headline ranking of LLM-based solutions over extractive baselines, reflect agreement with GPT-4o's event selection rather than with a human gold standard. This is a load-bearing gap, not an internal inconsistency: the paper's own evidence shows substantial disagreement among the three GPT-4o annotators, and the expert-refinement step is described but not quantified against an independent human reference.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Dynamic-granularity TimELine Summarization (DTELS), a task in which a timeline must be generated at a user-specified granularity, and presents a benchmark built around it. The benchmark contributions are: (1) an evaluation framework based on event atoms, with four metrics—Informativeness, Granular Consistency, Factuality, and Coherence; (2) a Chinese dataset, DTELS-Bench, with 543 post-October-2023 news topics, 55,432 articles from 2,858 sources, and reference timelines at three granularities (GN, G10, G5); and (3) an experimental comparison of extractive TLS baselines, several LLMs, and two proposed LLM-based solutions (Long-context Prompting and Hierarchical Merging). The authors report that their metrics correlate strongly with human judgments on a 50-timeline study, and that LLM-based systems outperform extractive baselines but still struggle with informativeness and granular consistency.","tokens_in":22148,"tokens_out":3728,"duration_ms":52377,"significance":"If the reported metric alignment holds, DTELS-Bench could become a useful evaluation standard for controllable timeline generation. The event-atom formulation is a principled response to the well-known fragility of ROUGE for timeline evaluation, and the decision to restrict topics to after October 2023 is a concrete and welcome mitigation of data contamination. The release of code, the large multi-source dataset, and the explicit human-alignment check are also strengths. The main reservation is that the gold medium- and coarse-grained references—against which the two central metrics are computed—are themselves constructed by GPT-4o role-play with limited agreement and without independent human validation of the reference timelines; the human study validates the scores, not the reference construction process. This is a fixable gap, but it is load-bearing for the central claim that the benchmark measures human-valued timeline quality.","major_comments":[{"comment":"The gold G10 and G5 reference timelines are produced by a consensus process in which three GPT-4o role-players select salient event-atom groups, followed by expert refinement. Table 6 reports only 45.09% full agreement among the three GPT-4o annotators, and no inter-annotator reliability is reported between the final expert-refined references and independent human annotations. Since the Informativeness and Granular Consistency scores in Equations (9) and (12) are computed against these references, the headline ranking of LLM-based methods over extractive baselines may reflect agreement with GPT-4o's notion of salience rather than with human salience. The paper should add a validation study in which human annotators independently construct or rate G10/G5 timelines on a sample of topics and report agreement with the published references, or otherwise quantify the expert-refinement step against an independent human gold standard.","section":"Section 5.2 and Appendix C.2, Table 6"},{"comment":"The human-alignment evidence rests on 50 timelines, three annotators, and a group-discussion step, and the reported Pearson correlations—especially 99.14% for coherence—are presented without confidence intervals, per-metric sample details, or inter-annotator agreement statistics. Given that coherence is scored automatically by GPT-4o and that the same model family is used in the reference-construction process, the exceptionally high coherence correlation needs more scrutiny. The paper should report the number of pairwise comparisons for the consistency score, annotator agreement (e.g., Krippendorff's alpha or Cohen's kappa), and confidence intervals for the correlations, and ideally include human coherence ratings on a larger held-out sample.","section":"Section 7.3 and Appendix F"},{"comment":"The edge-mapping procedure for Granular Consistency computes the cost of aligning a predicted edge to a reference edge solely from the endpoint InfoScores and then counts whether the matched reference edge belongs to the target granularity E_Go. Because the endpoint matching is done independently for each edge and the temporal interval penalty in Equation (5) is applied only at the node level, a predicted coarse edge whose two endpoints happen to match distant fine-grained nodes could be mapped to a fine edge and counted as inconsistent, or conversely. The paper should discuss or empirically test whether the Granularity metric is sensitive to the choice of the denominator in Equation (5) and whether edge-level temporal spans should be penalized explicitly.","section":"Section 4.3, Equations (10)-(12)"}],"minor_comments":[{"comment":"The abstract contains the typo 'DTLES' in 'This paper establishes a comprehensive benchmark for DTLES'; it should read 'DTELS'.","section":"Abstract"},{"comment":"The section heading 'Consesus-based Annotation' contains a misspelling; it should be 'Consensus-based Annotation'.","section":"Section 5.2"},{"comment":"The third sub-metric is labeled 'Linguistic Coherence' in Figure 7, but its definition concerns consistency of style and tone; the label should be 'Stylistic Coherence' to match the description in Section 4.3.","section":"Appendix B, Figure 7"},{"comment":"The protocol text says 'Down each sentence into smaller units'; this should be 'Break down each sentence into smaller units'.","section":"Appendix A.1"},{"comment":"The article filtering threshold of 0.3 for BERT embedding similarity and the denominator in the temporal penalty of Equation (5) are hand-set choices; a brief sensitivity analysis would clarify how stable the reported conclusions are to these choices.","section":"Section 5.1 and Appendix C.1"},{"comment":"The column header '#Node Go' is confusing; separating the granularity-instruction type, node count, and prompt type into distinct columns would make the table easier to read.","section":"Table 9"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid benchmark contribution and the core evaluation framework is well motivated. The main risk is the unvalidated GPT-4o-constructed references for G10 and G5; if the authors can provide a human validation of the reference timelines, the central claim will be much stronger. The 50-timeline human alignment study is small but not disqualifying; it needs more statistical detail. I would encourage the editor to invite a revision addressing the reference-validation issue rather than rejecting, since the gap appears fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a serious look, but the gold multi-granularity references are the weak link. The core idea is genuinely new: timeline summarization conditioned on granularity instructions, with event-atom metrics for informativeness, granular consistency, factuality, and coherence. The dataset is large, the post-2023 topic selection is a smart contamination choice, and the metrics are precisely defined. The Hungarian matching with temporal penalty is clean, and the paper's own analysis of LLM performance is honest rather than promotional.\n\nThe problem: the G10/G5 references, the thing the new Info and Granu metrics are computed against, are built by three GPT-4o role-players selecting salient event groups, with only 45% full agreement and an expert refinement step that is described but never independently validated. The stress-test concern lands. If GPT-4o's salience differs from human salience, then the headline rankings of LLM methods over extractive baselines partly reflect agreement with GPT-4o's event selection, not with a human gold standard. That is load-bearing, not manufactured. Coherence is also scored by GPT-4o, so the reported 99.14% human correlation is plausible but not conclusive.\n\nOther soft spots are smaller. The human alignment study uses 50 timelines and three annotators, and the human-alignment result does not directly validate the reference timelines themselves. The NLI entailement step is self-contained but has no error analysis. Event-atom decomposition is automatic via GPT-3.5 without quality checks. None of those are fatal; they are the usual gaps in a first benchmark paper.\n\nCredit where it is due: the metrics are specified well enough to reimplement, the alignment study is real evidence even if small, and the paper avoids fitted constants to make its methods look good. The hand choices in the temporal penalty and granularity node counts are reasonable, not circular.\n\nWho is this for? Anyone building or evaluating timeline summarization systems, especially LLM-based ones. It gives the TLS community a shared testbed and a useful vocabulary for granularity. It deserves peer review. A good referee should push for independent human validation of the gold references, NLI error analysis, and either human coherence annotations or a calibration study showing GPT-4o's coherence scores track human scores across the full range. I would engage with it, and I would cite it with caveats.","headline":"A genuinely new benchmark task and metric family for dynamic-granularity timeline summarization, with a real but addressable weakness in how the gold multi-granularity references are built.","tokens_in":22671,"tokens_out":1331,"would_cite":true,"duration_ms":14268,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes a dynamic-granularity timeline summarization benchmark whose event-atom metrics align with human judgments and shows that current LLMs still struggle to stay informative and granularly consistent.","keywords":["dynamic-granularity timeline summarization","timeline summarization","event atoms","evaluation metrics","benchmark dataset","large language models","news summarization","consensus-based annotation"],"falsifier":"Annotate a random sample of topics twice, once with the paper's GPT-4o consensus pipeline and once with human-only selection from the same fine-grained atoms, then compare the two reference sets and the scores they give to a fixed set of model outputs; if the system ranking changes materially, the benchmark is measuring the annotator's salience bias rather than a stable property of timelines.","tokens_in":21647,"feed_emoji":"📰","tokens_out":5380,"duration_ms":49637,"temperature":0.7,"pith_summary":"This paper introduces dynamic-granularity timeline summarization (DTELS), a paradigm in which a news timeline is generated at whatever level of detail a user requests, from a fine-grained event chain to a coarse overview. To make this measurable, it proposes an event-centric evaluation framework that scores timelines on informativeness, granular consistency, factuality, and coherence, and reports that these automatic scores align closely with human ratings on a sample. It also builds DTELS-Bench, a large-scale Chinese dataset with 543 topics, 55,432 articles from 2,858 sources, and reference timelines at three granularities built through a consensus process and expert refinement. The paper argues that this benchmark and metric suite can serve as a standard for evaluating controllable timeline generation, and that even the best LLM-based timelines are still far from ideal.","feed_headline":"Event-atom metrics grade news timelines at any granularity","feed_subtitle":"A four-part benchmark aligns with human judges and shows where LLM timelines still fall short","key_machinery":"The load-bearing mechanism is event atoms plus a mount-then-measure paradigm. Event atoms are the smallest distinguishable event units within a sentence, extracted automatically with GPT-3.5 for generated summaries and annotated for references. The mount-then-measure step builds an InfoScore matrix between predicted and reference nodes using entailment precision and recall weighted by a temporal interval penalty, then finds a global optimal matching with the Hungarian algorithm. Granular consistency extends this to edges between adjacent nodes, mounting each predicted edge to reference edges across all granularity levels. This shared alignment machinery lets all four metrics judge whether the same event information appears at the right time and at the right omission level, rather than relying on surface n-gram overlap.","core_discovery":"The central claim is that timeline quality at any granularity can be measured in a common currency: event atoms, the smallest distinguishable event units in a sentence. Predicted node summaries are mounted to reference nodes via optimal matching with a temporal penalty, and each matched pair is scored by entailment between atom sets. Granularity is then assessed at the edge level: adjacent-node pairs are matched to reference edges, and granular consistency is the fraction of predicted edges that land on a reference edge of the requested granularity. The paper reports that the resulting metrics align with human judgments, with Pearson correlations of 78.74% for informativeness, 76.66% for granular consistency, 95.87% for factuality, and 99.14% for coherence on a 50-timeline sample. It further claims that LLM-based solutions outperform extractive timeline-summarization baselines across all dimensions, but that they still struggle to combine informativeness with granular consistency.","pith_inferences":["The event-atom matching machinery is language-agnostic in principle; replacing the Chinese NLI model and dataset with other languages would likely transfer the four metrics, though the paper only demonstrates them on Chinese.","If GPT-4o's salience judgments in the consensus annotation encode a stable editorial style, the benchmark rewards timelines that match that style; a human-only re-annotation of a subset would quantify how large this effect is.","The same granularity-as-edge-omission definition could be applied to other structured summarization tasks, such as meeting minutes, financial reports, or historical chronologies, wherever nodes form a timeline-like sequence."],"forward_implications":["Timeline evaluation can move beyond ROUGE-style n-gram overlap to event-level entailment, which is more robust to differences in narrative style.","DTELS-Bench provides a multi-granularity reference set with 543 topics and three granularity levels, enabling direct comparisons among controllable timeline generation systems.","LLM-based solutions, especially hierarchical merging for context-limited models and long-context prompting for large-window models, dominate extractive TLS baselines on all four metrics.","Natural-language granularity instructions, such as requesting a coarse-grained timeline, can substitute for specifying node counts, with one-shot prompting performing competitively.","Current state-of-the-art models still score low on informativeness and granular consistency at coarser granularities, so the task remains open."],"supporting_citations":[{"why":"Supplies the atomic-fact decomposition approach that event-atom evaluation builds on.","marker":"Min et al., 2023"},{"why":"Provides the Hungarian algorithm used for optimal predicted-to-reference node and edge matching.","marker":"Kuhn, 1955"},{"why":"Supplies the journalism criteria that frame the four evaluation dimensions.","marker":"Kunelius, 2006"},{"why":"Supports the natural-language-inference entailment scoring between event atoms.","marker":"Camburu et al., 2018"},{"why":"Provides the BERT backbone for the Chinese NLI model that implements entailment checks.","marker":"Devlin et al., 2018"},{"why":"Defines the extractive TLS baselines that DTELS methods are compared against.","marker":"Gholipour Ghalandari and Ifrim, 2020"},{"why":"Provides ROUGE, the n-gram evaluation that the paper argues is inadequate and aims to replace.","marker":"Lin, 2004"}],"fun_headline_variants":["Event-atom metrics grade timelines at any granularity","New benchmark shows LLMs lag on timeline consistency","DTELS: measuring timeline quality in event atoms","Timeline quality now scored via atomic event matching","Dynamic granularity test exposes LLM summary gaps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reference medium- and coarse-grained timelines are built by having GPT-4o agents in three roles pick salient event groups from the fine-grained timeline; if that model's sense of salience differs from human readers, every metric computed against those references inherits the same bias.","fun_headline_variants_meta":{"raw":{"variants":["Event-atom metrics grade timelines at any granularity","New benchmark shows LLMs lag on timeline consistency","DTELS: measuring timeline quality in event atoms","Timeline quality now scored via atomic event matching","Dynamic granularity test exposes LLM summary gaps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000369,"raw_usage":{"total_tokens":1968,"prompt_tokens":927,"completion_tokens":1041,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":969}},"tokens_in":543,"tokens_out":1041,"duration_ms":10633,"temperature":1.0,"reasoning_tokens":969,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:48:09.673859+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Annotate a random sample of topics twice, once with the paper's GPT-4o consensus pipeline and once with human-only selection from the same fine-grained atoms, then compare the two reference sets and the scores they give to a fixed set of model outputs; if the system ranking changes materially, the benchmark is measuring the annotator's salience bias rather than a stable property of timelines.","supporting_citations":[],"review_version":1}