{"id":"4fdbe81a-5dab-4f86-ab3b-1b77dbc0c43a","arxiv_id":"1908.03505","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"SocialStories is a 40-story benchmark for evaluating social media visual storylines, scored by a metric that combines illustration relevance with transition coherence.","lead":"This paper introduces SocialStories, a benchmark of 40 curated news storylines from the Edinburgh Festival and Tour de France, with a quality metric for judging whether automatically selected social media images and videos tell a coherent story. It gives newsroom automation researchers a common yardstick for measuring visual storytelling.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The metric–human correlation is in-sample: alpha and beta have no fitting procedure or held-out validation, and the same three annotators supply both the component scores and the overall rating, so the 'emulates human perception' claim is not yet established.","rationale":"The reader's verdict is CONDITIONAL and identifies essentially the same weakest assumption: the alpha and beta weights are 'empirically found' without a fitting procedure or out-of-sample check, and the three annotators provide ground truth without reported inter-annotator agreement. My stress-test pass sharpens this into a more specific threat: the same annotators provide both the component scores (si, ti) that determine the metric and the overall story rating used as the correlation target, which makes the Figure 3 correlation vulnerable to within-annotator consistency rather than demonstrating true predictive validity. This is not an accusation of misconduct; it is a standard validation gap. The rest of the paper, including the dataset construction and baseline comparisons, is not invalidated by this concern. The proposed fix is concrete: a leave-one-annotator-out analysis and a train/test split for fitting alpha and beta would directly test whether the metric generalizes beyond the in-sample fit. Since the reader already conditioned acceptance on exactly this kind of validation, my reading does not move the verdict; it remains CONDITIONAL pending the additional evidence.","tokens_in":6829,"tokens_out":2094,"duration_ms":23278,"concrete_test":"Perform a leave-one-annotator-out validation: for each of the three annotators, compute the metric using only the other two annotators' relevance and transition scores (or using the annotator's own scores) and correlate it with that annotator's overall story rating, then average the three correlations. Additionally, fit alpha and beta on a randomly selected half of the 40 storylines and evaluate Pearson/Spearman correlation on the remaining half; repeat over multiple splits. If the held-out correlation is not statistically significant, or if the leave-one-annotator-out correlation drops substantially below the in-sample Figure 3 correlation, the claim that the metric emulates human perception is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, stated in Section 3.2, is that the Quality metric of Eqs. (1)–(2) 'effectively emulates the human perception of visual storyline quality.' The evidence for this is Figure 3, a correlation plot with no reported coefficient, confidence interval, or significance test. The load-bearing weakness is that the two free parameters of the metric, alpha=0.1 and beta=0.6, are asserted in Section 2.2 to have been 'empirically found' with no description of the fitting procedure, no train/validation split, and no out-of-sample check. The metric is then evaluated against human overall ratings collected from the same three annotators who also supplied the segment relevance and transition scores that feed into the metric. This creates a circularity: any correlation in Figure 3 may partly reflect the internal consistency of a single annotator's ratings rather than independent predictive validity of the metric. Additionally, no inter-annotator agreement is reported for the relevance, transition, or overall ratings, so it is unknown whether the ground truth is reliable enough to serve as a target. Without out-of-sample validation and independent annotation, the claim that the metric emulates human perception is supported only by an in-sample, possibly circular fit. The benchmark itself and the baseline comparisons in Section 3.3 remain useful, but the metric's validity as a perceptual yardstick is not established by the reported evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces SocialStories, a benchmark for visual storytelling from social media, with 40 curated storylines from two events (EdFest and TDF). It defines a Quality metric (Eqs. 1-2) as a weighted combination of segment relevance s_i and transition coherence t_i, with two free parameters alpha and beta. It evaluates the metric against human overall ratings (Fig. 3) and compares six content-selection baselines and six transition-optimization baselines (Fig. 4). The central claim is that the proposed metric 'effectively emulates the human perception of visual storyline quality' (Section 3.2). The paper also describes the crawling strategy, story segment construction, and the ground-truth annotation protocol with three annotators.","tokens_in":7082,"tokens_out":5142,"duration_ms":50113,"significance":"If the central validity claim were properly established, SocialStories would be a useful resource: it addresses a realistic newsroom task, and the two-term decomposition into relevance and coherence is a sensible, independently motivated design. The paper's concrete assets include a real social-media dataset, a clear protocol for constructing storylines, and a systematic baseline comparison for both illustration selection and transition coherence. However, the current evidence for the metric's perceptual validity is an in-sample fit with no statistical detail, so the central claim is not yet supported. The benchmark infrastructure and baseline results remain useful regardless of the metric's validation, but the paper's advertised contribution as a 'rigorous evaluation' tool is compromised until the metric validation is repaired.","major_comments":[{"comment":"The values alpha=0.1 and beta=0.6 are said to be 'empirically found' to represent human perception, but the paper gives no fitting procedure, no search criterion, no train/validation split, and no out-of-sample evaluation. Because the same metric is then shown to correlate with human judgments in Section 3.2, the reported agreement may simply reflect in-sample tuning. The authors should describe exactly how alpha and beta were selected, and should validate the metric on held-out storylines or a second annotation round, reporting the resulting correlation separately.","section":"Section 2.2, Eqs. (1)-(2)"},{"comment":"The validation is circular in an important sense: the same three annotators provide the segment relevance scores s_i and transition scores t_i that are fed into the Quality metric, and those same annotators also provide the overall story rating used as the target in Figure 3. A high correlation can then reflect within-annotator consistency rather than the metric's independent predictive validity. The authors should either use independent annotators for the overall rating, collect the overall rating before the component ratings, or at minimum report per-annotator correlations and discuss the possible dependence.","section":"Section 3.1 and Section 3.2"},{"comment":"The claim that the metric 'effectively emulates the human perception of visual storyline quality' rests entirely on Figure 3, but the figure presents no correlation coefficient, confidence interval, sample size, or significance test. Visual inspection of a scatter plot is insufficient, especially with only 40 stories. The authors should report Pearson and/or Spearman correlations, ideally with confidence intervals and per-event results, and state the number of stories included.","section":"Section 3.2, Figure 3"},{"comment":"No inter-annotator agreement is reported for any of the three annotation tasks (segment relevance, transition coherence, overall quality). With only three annotators, the ground truth may be noisy or biased, and the benchmark's utility as a quantitative yardstick depends on label reliability. The authors should report Fleiss' kappa, Krippendorff's alpha, or an equivalent agreement measure for each task.","section":"Section 3.1"},{"comment":"Figure 4 reports single mean scores for each baseline with no variance or significance testing. Several differences are small (e.g., 0.468 vs 0.450 for EdFest illustrations), so the ordering of baselines may not be reliable. The authors should provide error bars or significance tests, or at least state the number of stories underlying each mean.","section":"Section 3.3, Figure 4"}],"minor_comments":[{"comment":"The phrase 'comprised by total of 40 curated stories' should be 'comprising a total of 40 curated stories' or 'composed of 40 curated stories'.","section":"Abstract"},{"comment":"The dataset statistics in Table 1 are presented as a block of text; a proper table would improve readability and make the column structure clear.","section":"Section 2.1, Table 1"},{"comment":"The description of the Concept Pool method ('selects the image with the 10 most popular visual concepts') is unclear about whether it selects one image per segment using concept popularity across the segment; please clarify the procedure.","section":"Section 3.3.1"},{"comment":"Equations (1)-(2) treat the 0-2 relevance and transition labels as interval-scale values without justification; the authors should state why this arithmetic is appropriate or acknowledge the ordinal nature of the labels.","section":"Section 2.2, Eqs. (1)-(2)"},{"comment":"The paper mentions prior visual storytelling datasets [7,8] but does not include a dedicated related-work discussion; a short paragraph positioning SocialStories relative to those datasets and to TRECVID 2018 would help readers understand the novelty.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"The paper appears to be a short conference/workshop paper describing work developed for TRECVID 2018. The main methodological gaps (fitting details, inter-annotator agreement, out-of-sample validation) are fixable but require additional experiments or a supplementary appendix. No evidence of any author misconduct; the issues are standard validation concerns. Given the page limit, the authors could provide the missing details as supplemental material, but without those details the central perceptual-validity claim should not be accepted as established."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a compact benchmark paper that gives the field a small but real social-media visual storytelling dataset and a sensible quality metric decomposed into segment relevance and transition coherence. The problem is the validation section: the claim that the metric 'effectively emulates human perception' rests on a correlation plot with no coefficient, no confidence interval, and no out-of-sample check, and the two weights (alpha=0.1, beta=0.6) are asserted as 'empirically found' with no fitting procedure described.\n\nWhat's genuinely new: unlike the clean image-caption sequences in [7,8], SocialStories targets the noisy, redundant, unevenly distributed content a newsroom editor actually faces. Forty stories across two events is small, but the crawling and annotation protocol are explicit enough to replicate. The two-component quality metric is a sensible operationalization: penalize irrelevant illustrations, reward coherent transitions, and give a small boost to the first segment. The baseline results in Section 3.3 are also informative—BM25 wins for finding illustrations, CNN dense features win for transitions, and color histograms are surprisingly competitive. That is useful for anyone building a system.\n\nWhere it's soft: the load-bearing claim about emulating human perception isn't supported by the evidence. Alpha and beta are fitted to human judgments, and the same three annotators provided both the component scores (relevance, transition) and the overall rating used for validation. So Figure 3's 'strong and relatively stable' relation may just reflect internal consistency of the annotators, not predictive validity. There is no inter-annotator agreement reported for any rating type, so we don't know if the ground truth is reliable enough to validate against. And with n=40 and several free components in Eqs. (1)-(2), the model has enough flexibility to fit the annotations. So the 'emulation' claim should be downgraded to 'is consistent with the annotations used to set the weights.'\n\nNone of this makes the benchmark useless. The dataset and baseline comparisons stand on their own. But a serious reader should treat the metric's perceptual validity as unproven until there is held-out validation with independent annotators and proper correlation statistics.\n\nIf this is under consideration, I'd send it to review: the benchmark is worth having in the record, and the flaws are fixable with additional experiments. But I'd ask for a rewritten Section 3.2 before acceptance.","headline":"A useful small benchmark dataset with a plausible quality metric, but the 'emulates human perception' claim is an in-sample fit with no statistics.","tokens_in":7633,"tokens_out":3170,"would_cite":false,"duration_ms":29896,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A two-part quality metric reproduces human ratings of social-media visual storylines.","keywords":["visual storytelling","social media benchmark","story quality metric","multimedia retrieval","newsroom automation","relevance and coherence","Edinburgh Festival","Tour de France"],"falsifier":"Take the same annotation protocol, apply it to a held-out set of storylines from at least one event not used in this paper, compute Quality with α=0.1 and β=0.6, and compare against holistic human ratings. If the correlation is much weaker than the reported pattern, or if re-fitting the weights yields values far from 0.1 and 0.6, then the metric's claimed emulation of human perception does not generalise.","tokens_in":6650,"feed_emoji":"🎬","tokens_out":5662,"duration_ms":55985,"temperature":0.7,"pith_summary":"This paper tries to establish that the quality of an automatically assembled visual storyline — a sequence of social-media images or videos illustrating a news event — can be scored by a single quantitative metric that tracks human editorial judgment. The proposed SocialStories benchmark supplies 40 curated storylines from two long-running events, the Edinburgh Festival and the Tour de France, along with a metric built from two components: how relevant each illustration is to its story segment, and how coherent the visual transition is between neighbouring segments. The authors argue that the metric's values rise and fall with crowd-sourced human ratings of overall story quality, so it can serve as a standard yardstick for comparing automatic storytelling systems. A sympathetic reader would care because automated newsroom tools need a repeatable way to know whether a machine-selected sequence of images tells the story coherently, not just whether individual images are on-topic.","feed_headline":"A new benchmark scores social-media visual stories against human taste","feed_subtitle":"40 curated storylines from two live events test whether relevance plus transition coherence captures editorial quality.","key_machinery":"The machinery is the Quality metric, Eqs. (1)-(2), which treats a storyline as a chain where each segment has a relevance score $s_i$ and each adjacent pair has a transition coherence score $t_i$. It computes an overall quality by giving the first segment a small boost through $\\alpha$, then averaging pairwise terms that balance the summed relevance of the two illustrations, weighted by $\\beta$, against the product of their relevance plus the transition label, weighted by $1-\\beta$. The metric is the load-bearing device because every evaluation in the paper — baselines for image selection, baselines for transition smoothness, and the comparison to human judgement — is expressed through it. Its two weights, $\\alpha = 0.1$ and $\\beta = 0.6$, are presented as empirically adequate representations of human perception.","core_discovery":"On the paper's own terms, the central discovery is that human perception of visual storyline quality can be expressed as a weighted combination of segment-level relevance and transition-level coherence. Concretely, for a story of N segments the Quality score is $\\text{Quality} = \\alpha \\cdot s_1 + \\frac{1-\\alpha}{2(N-1)} \\sum_{i=2}^{N} \\text{pairwiseQ}(i)$, where $\\text{pairwiseQ}(i) = \\beta \\cdot (s_i + s_{i-1}) + (1-\\beta) \\cdot (s_{i-1} \\cdot s_i + t_{i-1})$, with relevance labels $s_i \\in \\{0,1,2\\}$, transition labels $t_i \\in \\{0,1,2\\}$, and empirically chosen weights $\\alpha = 0.1$ and $\\beta = 0.6$. The reported agreement between this score and annotators' overall ratings indicates that linear increases in human judgement are matched by the metric, and the authors conclude that it effectively emulates human perception.","pith_inferences":["If the metric generalises beyond these two events, it could be adapted as a reward signal for training retrieval or generation models, since it offers a single scalar target that combines semantic relevance and visual coherence.","The absence of a reported fitting procedure for $\\alpha$ and $\\beta$ suggests a direct testable extension: re-estimate the weights on a held-out set of storylines and check whether 0.1 and 0.6 remain optimal, since a large shift would indicate the metric encodes dataset-specific calibration rather than a general perceptual law.","The three-annotator ground truth, without reported inter-annotator agreement, leaves open how much of the metric's apparent success reflects shared editorial preference versus averaged individual taste; measuring agreement directly would clarify the mechanism.","The same relevance-plus-transition decomposition could be carried over to other sequential multimodal outputs such as automated slide decks, video digests, or illustrated tutorials, where a single quality score would allow direct A/B testing."],"forward_implications":["Automatic visual storytelling systems can be compared on a common numeric scale, so a text-based retriever that finds relevant images can be measured against approaches that optimise visual coherence between segments.","Because the metric separates illustration relevance from transition coherence, a system that improves one component should show a corresponding gain in overall Quality, and the benchmark can localise which component is failing.","The two collected events behave differently: Tour de France storylines are systematically easier to illustrate than Edinburgh Festival storylines, so benchmark results should be reported per event rather than pooled.","Social signals such as retweet counts and duplicate counts can be competitive with text retrieval for selecting illustrations, while colour-based and CNN-based methods help most on the transition side."],"supporting_citations":[{"why":"Situates SocialStories within the TRECVID 2018 evaluation task and gives the shared protocol the benchmark builds on.","marker":"[1]"},{"why":"Represents the existing visual storytelling dataset whose clean image-caption sequences the paper argues do not match noisy social media.","marker":"[7]"},{"why":"Another existing image-sequence retrieval benchmark used as a contrast for why a social-media-specific test bed is needed.","marker":"[8]"},{"why":"Supplies the pretrained VGG-16 network behind the visual-concept and CNN baselines whose performance is measured in the evaluation.","marker":"[15]"},{"why":"Provides the temporal-smoothing baseline used to test whether posting-time evidence helps illustration selection.","marker":"[10]"}],"fun_headline_variants":["Benchmark ranks social-media storylines against human perception","SocialStories benchmark: scoring visual narratives from live events","New metric matches human taste for social visual storylines","40 curated stories test relevance and coherence in visual news","Scoring social media visual stories: relevance plus transitions wins"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"A single pair of weighting constants, chosen empirically but never tested on fresh data, is assumed to match how people judge story quality for every social-media storyline.","fun_headline_variants_meta":{"raw":{"variants":["Benchmark ranks social-media storylines against human perception","SocialStories benchmark: scoring visual narratives from live events","New metric matches human taste for social visual storylines","40 curated stories test relevance and coherence in visual news","Scoring social media visual stories: relevance plus transitions wins"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000177,"raw_usage":{"total_tokens":1261,"prompt_tokens":881,"completion_tokens":380,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":497,"completion_tokens_details":{"reasoning_tokens":302}},"tokens_in":497,"tokens_out":380,"duration_ms":4346,"temperature":1.0,"reasoning_tokens":302,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:10:38.510850+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same annotation protocol, apply it to a held-out set of storylines from at least one event not used in this paper, compute Quality with α=0.1 and β=0.6, and compare against holistic human ratings. If the correlation is much weaker than the reported pattern, or if re-fitting the weights yields values far from 0.1 and 0.6, then the metric's claimed emulation of human perception does not generalise.","supporting_citations":[{"cited_title":"Smeaton, Yvette Graham, Wessel Kraaij, Georges Quénot, Joao Magalhaes, David Semedo, and Saverio Blasi","cited_arxiv_id":null,"evidence_quote":"Situates SocialStories within the TRECVID 2018 evaluation task and gives the shared protocol the benchmark builds on."},{"cited_title":"Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Jacob Devlin, Aishwarya Agrawal, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, et al","cited_arxiv_id":null,"evidence_quote":"Represents the existing visual storytelling dataset whose clean image-caption sequences the paper argues do not match noisy social media."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Another existing image-sequence retrieval benchmark used as a contrast for why a social-media-specific test bed is needed."}],"review_version":1}