{"id":"92ab55b1-1124-457b-a251-fbcca8173047","arxiv_id":"1908.07316","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"In a Mechanical Turk study with 62 participants, dual-event alignment improved accuracy for intermediate-event tasks (71% vs. 18%) but no alignment was best for duration tasks (88% vs. 36%).","lead":"This paper ran a crowdsourced experiment comparing four ways of aligning time-series visualizations around key events. It found that the alignment choice matters: aligning around two events helps users find intermediate events, but no alignment is best when measuring duration between events.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Post-hoc k-means exclusion of 46 of 108 participants is the load-bearing weak point: every headline comparison rests on a data-dependent filter, and the paper does not show the unfiltered results.","rationale":"The reader's weakest assumption identifies the same load-bearing concern I find: the post-hoc 'speeding' filter in Section 4.3. My independent reading of Section 5 confirms that the unfiltered results are not presented in the paper, so the headline claims cannot currently be verified without assuming the k-means-derived exclusion is valid. This is not a charge of misconduct; the OSF data availability is a real strength, and the direction of the reported effects is plausible. But the central empirical conclusions are conditional on a filter that is derived from the outcome variables themselves and applied after the fact. The correct next step is a full-sample reanalysis or a pre-registered sensitivity analysis; until then, the paper's claims should remain conditional, matching the reader's verdict. I do not see an independent internal inconsistency that would require moving to REJECT, and I do not see grounds to upgrade to ACCEPT without the robustness check.","tokens_in":8233,"tokens_out":3844,"duration_ms":38768,"concrete_test":"From the OSF dataset, recompute Task 4 and Task 5 at both scales for all 108 accepted participants (or all 123 recruited if consent allows) using the same chi-square and Kruskal-Wallis procedures as Section 4.4, plus the same post-hoc pairwise comparisons. Then repeat with alternative inclusion thresholds (e.g., no post-hoc filter, ≥5 correct, and ≥10 minutes) to see whether the Task 5 DualStretch/DualLeft vs. NoAlign/SingleAlign correctness advantage and the Task 4 NoAlign advantage remain significant at p<0.05. Also report the condition-by-condition composition of the 46 excluded participants; if the excluded cluster is disproportionately drawn from one or two conditions, the reported effects are even more likely to be artifacts of the filter.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claims in the abstract—dual-event alignment at 71% vs. 18% for intermediate events, and NoAlign at 88% vs. 36% for duration—are all computed from n=62 after Section 4.3 filters out 46 of 108 accepted participants using k-means clustering on the same outcome variables (time and correctness) that define the study's dependent measures. The stated inclusion bound (at least 4 correct answers and at least 8 minutes) was derived after seeing the data, not pre-registered, and the earlier pre-specified rejection criteria had already been applied. Because the filter is not independent of the outcomes, it can create or exaggerate between-condition differences: participants with low correctness and low time may be unevenly distributed across conditions, and the small final per-cell counts (14–18) make every chi-square and Kruskal-Wallis result fragile. Section 5 explicitly says 'results with the unfiltered data ... are relegated to supplemental material,' so the reader cannot check whether the Task 4 and Task 5 14-day effects survive including all 108 participants. The observed 'clear clustering' is not by itself evidence of attentiveness; it is a property of the joint distribution of the outcome variables and could reflect task difficulty, strategy differences, or interaction with the visualization condition. This is load-bearing because it precedes every reported p-value and percentage in the headline results.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a controlled between-subjects experiment on Amazon Mechanical Turk comparing four alignment approaches for superimposed time-series and temporal event-sequence visualizations: no alignment, single-event alignment, and two dual-event alignment variants (left and stretch justification). Participants answered twelve questions (six tasks at two data scales) probing precursor/aftereffect events, intermediate events, and durations. The headline findings are that dual-event alignment improves correctness for identifying intermediate events (Task 5, 14-day scale: 71% vs 25% for NoAlign and 11% for SingleAlign), while no alignment is best for judging durations (Task 4, 14-day scale: 88% vs 36% for DualStretch, with faster times and lower error). The authors provide stimuli, data, and source code on OSF.","tokens_in":8487,"tokens_out":4604,"duration_ms":46183,"significance":"If the results are valid, the paper is a valuable empirical contribution: it is one of the few comparative evaluations of dual-event alignment in superimposed visualizations, and it offers a domain-independent task abstraction for temporal event-sequence and timing analysis. The open-data and open-materials practices are commendable and make independent verification possible. However, the central empirical claims are currently undermined by a post-hoc participant-exclusion step that operates on the same outcome variables used in the analysis, and by small per-condition samples with numerous unadjusted significance tests. These issues are load-bearing because every headline percentage in the abstract is computed from the filtered 62 participants.","major_comments":[{"comment":"The \"speeding\" filter is post-hoc and data-dependent, and it is applied to the very measures that define the study's dependent variables. After pre-specified rejection criteria had already been applied, the authors ran a k-means clustering on the collected time-spent and correctness data, discovered a \"clear clustering effect,\" and removed 46 of the 108 accepted participants. The stated bound (\"four answers out of 12 questions and at least 8 minutes\") was derived from the observed data, not pre-registered. Because the filter uses the outcome variables, it can create or exaggerate between-condition differences if the removed participants are not evenly distributed across the four conditions. This is load-bearing: the Task 4 and Task 5 14-day results in the abstract (88% vs 36%; 71% vs 25% and 11%) all come from the filtered n=62, and the paper does not show that these conclusions survive without the filter.","section":"Section 4.3"},{"comment":"The manuscript explicitly states that \"results with the unfiltered data ... are relegated to supplemental material.\" For a paper whose central claims depend entirely on a controversial post-hoc exclusion, relegating the unfiltered analysis to the supplement is insufficient. The authors should report the unfiltered results in the main text or, at minimum, provide a full robustness table showing that the significant differences in Tasks 4 and 5 persist with all 108 accepted participants, or with any sensible alternative filter (e.g., the pre-specified criteria alone). Without this, the reader cannot verify whether the headline percentages are an artifact of the data-dependent exclusion step.","section":"Section 5, opening paragraph"},{"comment":"The multiple-comparison problem is not adequately handled. The analysis runs separate chi-square or Kruskal-Wallis tests for each of 12 task-scale combinations and for three measures (correctness, time, error), yet the Bonferroni adjustment mentioned in Section 4.4 appears to apply only to post-hoc pairwise comparisons after a significant omnibus test. With about 36 tests at alpha = .05, several nominally significant results are expected by chance. For example, the Task 5 14-day correctness comparison between DualLeft/SingleAlign (p = .03) and the Task 4 14-day error comparison (p = .01) would not survive a family-wise correction over all 36 tests. The paper should either apply an appropriate multiple-comparison control across the question/measure families or present a sensitivity analysis showing that the headline conclusions are robust to the correction.","section":"Sections 4.4 and 5.2"},{"comment":"The per-condition sample sizes after filtering are 14, 14, 16, and 18, which makes the binary correctness analyses fragile. With these counts, chi-square tests on 2x2 or 2x4 tables often have expected cell counts below 5, so the reported p-values (e.g., Task 5 14-day: NoAlign 25% of 16, SingleAlign 11% of 18, DualLeft 71% of 14, DualStretch 71% of 14) are not reliable. The authors should use Fisher's exact tests or report effect sizes with confidence intervals, and they should explicitly note the small cell counts when interpreting the 71% vs 25% and 88% vs 36% contrasts. This is directly relevant to the strength of the abstract's \"clear winner\" language.","section":"Section 4.3 and Results"}],"minor_comments":[{"comment":"In the participant counts, \"DualRight\" appears to be a typo; the paper's four conditions are NoAlign, SingleAlign, DualLeft, and DualStretch. Please correct the labels for consistency.","section":"Section 4.3"},{"comment":"The phrase \"the bound as four answers out of 12 questions and at least 8 minutes were spent on the questions\" is ambiguous. Please specify whether the criterion is \"at most four correct answers\" and whether the time bound is a minimum or maximum, and state the bound as a clear conjunction of conditions.","section":"Section 4.3"},{"comment":"The legend in Figure 4 spells \"DualStrech\"; this should be \"DualStretch.\"","section":"Figure 4"},{"comment":"In the Task 5 3-day result, \"DualStretch (100%)\" with n=14 implies all 14 participants answered correctly; reporting the raw counts (e.g., 14/14) alongside percentages would help readers assess the reliability of that contrast.","section":"Section 5.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about the post-hoc nature of the filtering step, but it substantially underplays the impact of removing 43% of accepted participants based on the outcome variables. The central empirical claim is currently unverifiable without the unfiltered results. I agree with the reader's assessment that this is the weakest point. A major revision that includes full unfiltered analyses, proper multiple-comparison handling, and a clear statement of the exclusion rule's provenance could make the contribution publishable. If the unfiltered results contradict the headline findings, the paper would need to be substantially reframed, possibly as a demonstration of the sensitivity of MTurk studies to exclusion criteria."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First controlled comparison of dual-event alignment in composite temporal visualizations, and the first thing to know is that the headline numbers come from a data-dependent participant filter that isn't pre-specified. The paper is honestly reported—the filter is in Section 4.3, and unfiltered results are in the supplement—but it's still load-bearing: 46 of 108 accepted participants are removed by k-means clustering on the same measures (time and correctness) that define the study's outcomes. The bound of at least 4 correct and 8 minutes was derived after seeing the data, so every p-value and percentage in the abstract could shift with that choice. The stress-test note has this right.\n\nWhat's genuinely new: this is the first controlled evaluation of dual-event alignment. Prior work was qualitative expert feedback (their own IDMVis paper). The task abstraction (precursor, aftereffect, intermediate, co-occurrence) is a useful framing, and the open OSF materials let anyone reanalyze. Credit is also due for H1 being contradicted—that suggests the experiment isn't rigged to produce expected results.\n\nThe soft spots are real but not fatal to the enterprise. The post-hoc filter is the big one; the per-condition n of 14–18 after filtering makes every chi-square and Kruskal-Wallis fragile. Multiple comparisons across 12 questions without adjustment (they use Bonferroni only for post-hoc, not for the omnibus tests) is a minor concern. There's a typo in Section 4.3 where 'DualRight' appears instead of 'DualStretch'—worth fixing but not substantive.\n\nWho should read this: visualization researchers working on event-sequence alignment, especially in healthcare. The conclusions are plausible—dual alignment helps intermediate events (except duration), single alignment doesn't help much—but I'd treat them as hypothesis-generating until a reanalysis with a pre-registered exclusion rule or with all 108 participants is done. I would send this to peer review: the question is important, the materials are open, and a good reviewer can push for the reanalysis. I would not cite the headline numbers in my own work until that reanalysis happens.","headline":"First controlled evaluation of dual-event alignment, but the headline numbers all flow through a post-hoc participant filter that is not pre-specified and is tied to the outcome measures.","tokens_in":8981,"tokens_out":1847,"would_cite":false,"duration_ms":19390,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Alignment in time-event visualizations should be chosen by task: dual alignment helps read intermediate events but hurts duration judgment.","keywords":["sentinel event alignment","temporal event sequence visualization","time-series visualization","composite visualization","crowdsourced user study","dual-event alignment","Type 1 diabetes data"],"falsifier":"Re-running the same 12-question protocol with the exclusion rule fixed before data collection—or simply re-analyzing the 108 accepted participants without the k-means \"speeding\" filter—would settle the central claim. The claim fails if Task 4 at the 14-day scale no longer favors NoAlign over DualStretch on correctness, time, and error, or if Task 5 no longer favors dual alignment over NoAlign and SingleAlign.","tokens_in":8020,"feed_emoji":"📊","tokens_out":8365,"duration_ms":79875,"temperature":0.7,"pith_summary":"This paper tries to establish that the usefulness of sentinel-event alignment in superimposed time-series and event-sequence visualizations depends on the analytic question, and that the newest form—dual-event alignment—is a real improvement for one class of questions and a real drawback for another. In a crowdsourced controlled experiment using Type 1 diabetes timelines with 3 and 14 days of data, the authors compared no alignment, single-event alignment, and two dual-event alignments. They report that for understanding intermediate events between two sentinel events, dual-event alignment was far more correct (71% vs. 18% for no alignment and single alignment at the 14-day scale), whereas for judging the duration between the two sentinel events, no alignment was far better (88% correctness vs. 36% for stretch dual alignment, with lower error and faster completion). Single-event alignment showed no dependable gain over no alignment for precursor and aftereffect events. The paper concludes that designers should match alignment choice to task type rather than applying one alignment strategy everywhere.","feed_headline":"Dual-event alignment aids intermediate events, hurts duration","feed_subtitle":"A 62-person study on diabetes timelines: alignment is best matched to the question asked, not used everywhere.","key_machinery":"The central object is sentinel-event alignment, a family of transforms that shifts or rescales each row of a temporal visualization so that one or two chosen events line up vertically across rows. NoAlign leaves true time intact; SingleAlign aligns rows by a single sentinel event (e.g., lunch); DualLeft aligns rows by two sentinel events at their left positions; DualStretch aligns rows by two sentinel events while stretching or compressing the intervening time scale. The experiment's task machinery consists of six low-level questions built from three task abstractions—precursor events, aftereffect events, and intermediate events—each run at two scales (3 rows and 14 rows) on de-identified Type 1 diabetes data. The load-bearing comparison is Task 4 (intermediate duration) versus Task 5 (intermediate co-occurrence) at the 14-day scale, where the same family of transforms produces opposite winners.","core_discovery":"On its own terms, the paper's central discovery is an interaction between alignment technique and task type in composite temporal visualizations. Using 14 days of superimposed blood-glucose time series and event markers, participants asked to identify intermediate events between two sentinel meals answered correctly 71% of the time with dual-left or dual-stretch alignment, versus 18% with no alignment or single-event alignment. Participants asked to read the duration between the two sentinel meals did the opposite: no alignment gave 88% correctness versus 36% for dual-stretch, completion time of 55 seconds versus 101 seconds for dual-left, and error of 1.5% versus 8.4% for dual-stretch. At the 3-day scale most differences disappeared. The paper therefore concludes that dual-event alignment supports intermediate-event reading, especially with more rows of data, but actively misleads for duration estimation, and that single-event alignment contributes no reliable advantage over no alignment in this superimposed setting.","pith_inferences":["This extends the paper's result: because stretch alignment rescales the time axis between sentinel events, any task that depends on comparing interval lengths, rates, or durations in an aligned composite view should be expected to suffer, in domains beyond diabetes.","A testable extension the paper leaves open: repeat the study with simpler point-event data and lower visual clutter; the absence of a single-alignment benefit may be specific to superimposed, multi-encoding displays, and categorical encoding of event type might restore it.","The untested symmetry assumption—that dual alignment with right justification behaves like left justification—could be checked directly, since the two variants are not visually equivalent when reading order matters.","The reported effect sizes are large, but the post-hoc speeding filter means a pre-registered replication with a fixed exclusion rule is the natural next check; the supplemental unfiltered data would support such an analysis."],"forward_implications":["Designers of composite temporal visualizations should use dual-event alignment, particularly the stretch variant, when users need to see what happens between two sentinel events; it was the clear correctness winner for intermediate-event tasks.","Dual-event alignment should be avoided for questions about the duration or size of intervals between sentinel events; no alignment beat both dual variants on correctness, time, and error for the duration task.","Single-event alignment did not justify itself in this setting: it was no more correct than no alignment for precursor/aftereffect tasks and could be slower at the smaller scale.","Differences among alignment approaches grow with the number of rows; at 3 days the conditions were largely equivalent, so alignment decisions matter most for dense, multi-row displays."],"supporting_citations":[{"why":"Supplies the dual-event alignment techniques, the Type 1 diabetes de-identified data, and the stimuli that the experiment evaluates.","marker":"[15]"},{"why":"Provides the prior single-event alignment evaluation and the precursor/aftereffect/intermediate task abstractions that the first hypothesis and Tasks 1-3 build on.","marker":"[13]"},{"why":"Defines temporal event alignment as a strategy for coping with volume and variety, providing the conceptual basis for the four experimental conditions.","marker":"[4]"},{"why":"Contributes the high-level task types (intermediate events and sentinel time comparisons) that the six experimental tasks are mapped from.","marker":"[5]"},{"why":"Defines superimposed composite visualization and its readability trade-offs, motivating both the study setting and the clutter explanation for null single-alignment results.","marker":"[8]"}],"fun_headline_variants":["Dual alignment helps intermediate events, but hurts duration","Alignment choice: dual for events, none for duration","Task matters: dual alignment aids events, misleads duration","Dual-event alignment: good for events, bad for duration","When to align: dual for midpoints, none for spans"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline comparisons rest on the post-hoc decision, made after viewing the data, to remove 46 of 108 accepted participants as \"speeding\" using a clustering-derived bound; if those participants were actually paying attention, or if the bound is arbitrary, every reported effect can change.","fun_headline_variants_meta":{"raw":{"variants":["Dual alignment helps intermediate events, but hurts duration","Alignment choice: dual for events, none for duration","Task matters: dual alignment aids events, misleads duration","Dual-event alignment: good for events, bad for duration","When to align: dual for midpoints, none for spans"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0005,"raw_usage":{"total_tokens":2486,"prompt_tokens":1026,"completion_tokens":1460,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":642,"completion_tokens_details":{"reasoning_tokens":1378}},"tokens_in":642,"tokens_out":1460,"duration_ms":10344,"temperature":1.0,"reasoning_tokens":1378,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:19:31.839606+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-running the same 12-question protocol with the exclusion rule fixed before data collection—or simply re-analyzing the 108 accepted participants without the k-means \"speeding\" filter—would settle the central claim. The claim fails if Task 4 at the 14-day scale no longer favors NoAlign over DualStretch on correctness, time, and error, or if Task 5 no longer favors dual alignment over NoAlign and SingleAlign.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the prior single-event alignment evaluation and the precursor/aftereffect/intermediate task abstractions that the first hypothesis and Tasks 1-3 build on."},{"cited_title":"Gotz and H","cited_arxiv_id":null,"evidence_quote":"Contributes the high-level task types (intermediate events and sentinel time comparisons) that the six experimental tasks are mapped from."}],"review_version":1}