REVIEW 3 major objections 5 minor 42 references
Authoring Narrative Visualization in Motion: Visual Storytelling in Swimming Videos
T0 review · 3 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read The paper claims that authoring narrative visualizations in motion is feasible through timeline-based coordination of views, transitions, data layers, and pacing, and that a nine-participant study supports this.
desk verdict Solid tech-probe contribution on narrative authoring for swimming videos; the main soft spot is the unvalidated data pipeline, but the paper is honest and deserves a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is SwimComposer, a technology probe — a functional but intentionally minimal tool for studying how people author. It exposes four sets of controls: a layer library that organizes race data into live data, insights, athlete information, and records; a narrative viewing panel with three view modes (overview, tracking, comparison); timelines that arrange layer segments, view segments, and playback speed; and configuration panels for transitions and layer design. Carrying the argument is also the automatic multimodal data preparation pipeline that feeds the probe: video processing with lane partitioning and prompt-based detection and tracking to extract live metrics, audio pro
What would settle it
A reader could process a known race through the published pipeline and compare the computed frame-level rankings, split times, and gap measurements against the official timing results for that race. If the pipeline's live data deviates systematically from official splits, or if the LLM-extracted insights fail to match a human-coded set of key moments, the central claim that the tool supports truthful narrative authoring would be falsified.
Extended reading notes
Core claim
The paper's central claim is that authoring narrative visualizations in motion can be supported by reframing the task as the coordination of views, transitions, data layers, and pacing over time. Drawing on an observational analysis of professional broadcasts in swimming, basketball, and soccer, the authors found a recurring alternation between overview and focus-on views, with transitions used to signal emphasis. They implemented this insight in SwimComposer, a technology probe that lets creators arrange data layers, view changes, transitions, and playback speed on timelines. In a pre-registered study with nine participants experienced in content creation or graphic design, all produced com
Load-bearing premise
The paper's argument rests on the automatic pipeline producing structured data — swimmer positions, speeds, gaps, and commentary-derived insights — that is accurate enough to support truthful narrative videos; the authors validate the insight category coverage (98.15% on 15 races) but do not measure per-frame tracking or transcription accuracy against ground truth.
Editorial extensions
If this is right
- If SwimComposer's central claim holds, authoring narrative visualizations in motion becomes a coordination problem of arranging views, transitions, data layers, and pacing on a timeline, rather than a programming or manual animation task.
- The consistent authoring patterns observed across participants suggest that race-based narratives share a temporal logic — overview as backbone, detailed views for key moments, event-driven data layering — that tool designers can build into future systems.
- The automatic pipeline means new races can be turned into data stories from just a video, commentary audio, and a competition name, making the approach reproducible and transferable across events without per-race manual annotation.
- Participants reported that the tool produced clearer results than they could achieve in general-purpose video editing tools, indicating that a specialized probe can fill a gap in current editing workflows for sports storytelling.
Reading between the lines
- Because the view-alternation and event-driven emphasis patterns appeared in professional broadcasts across swimming, basketball, and soccer, the coordination model likely generalizes beyond swimming; a concrete test would be to implement the same timeline-based probe for a sport with less linear motion, such as soccer, and observe whether the same authoring patterns emerge.
- The paper validates the insight category design against commentary coverage (98.15% of key moments across 15 races) but does not measure per-frame tracking accuracy or transcription reliability; a natural next step is to compare the pipeline's computed positions, speeds, and gaps against official timing data for the same races, which would test whether authored narratives rest on truthful data.
- The study's nine participants and single task window make the strongest evidence the internal consistency of patterns rather than statistical power; a longer deployment with professional broadcast editors on real deadlines would reveal whether the authoring approach survives production pressures.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SwimComposer, a technology probe for authoring narrative visualizations in motion in swimming videos. It contributes an automatic multimodal pipeline (Sec. 3) that turns a race video, commentary audio, and a competition name into structured live data, event insights, athlete information, and records; an observational analysis of professional broadcasts (Sec. 4) that motivates views, transitions, and event-driven emphasis; and a user study (Sec. 6) with nine experienced content creators/designers who authored a full-race narrative video. The central claim is that SwimComposer generally enables effective narrative authoring and storytelling, supported by positive Likert ratings, consistent authoring patterns (Overview as backbone, Tracking/Comparison for key moments), and qualitative reports of reduced analysis effort and useful insights.
Significance. If the result holds, this is a useful step beyond prior work on embedding visualizations in sports videos: it treats authoring as coordinating data layers, views, transitions, and pacing over time, and it grounds the design in both an automated data pipeline and observed broadcast practices. The paper is unusually transparent: the study is pre-registered, questionnaires and authored videos are shared on OSF, pipeline source code and LLM prompts are released, and the authors explicitly acknowledge several limitations. These open-science practices strengthen the credibility of the qualitative evaluation. The main risk is that the truthfulness of the authored narratives depends on unvalidated pipeline accuracy; if tracking, transcription, or LLM extraction produce incorrect race facts, the resulting videos could mislead viewers, which would undermine the central effectiveness claim.
major comments (3)
- [Sec. 3, especially 3.2 and 3.3] The pipeline is load-bearing for the central claim, but its accuracy is not validated against ground truth. Tracking, speed/gap derivation, WhisperX transcription, name normalization, and GPT-5.4 insight extraction are all used without quantitative evaluation. The only reported check (Sec. 3.3, Fig. 3) is a 98.15% coverage rate of coder-identified key moments by the proposed insight types; this measures taxonomy completeness, not whether the LLM correctly identified events in each race, and no inter-coder reliability is reported for the key-moment coding. Since the user study (Sec. 6) used pipeline output for the women's 100 m butterfly final, any tracking jitter or mis-transcription directly propagates into the authored videos. Please add validation against official results/split times and, at minimum, a manually labeled subset of positions and events, with precision/recall for insight
- [Sec. 6.1-6.4] The study has 9 participants recruited through the authors' networks, no baseline or comparison condition, and the Likert analysis is purely descriptive. The paper's wording 'generally enables effective narrative authoring and storytelling' (Sec. 1) is stronger than what this design can support. I do not require a fully powered experiment for a technology probe, but the claim should be softened (e.g., 'was perceived to enable...') or complemented with a within-subject comparison to a simpler authoring tool. At minimum, report confidence intervals or effect sizes for the Likert items and be explicit that the observed 'consistent patterns' are qualitative.
- [Sec. 6.4 / Fig. 5] The Likert item presentation is hard to read: the counts are not shown per response option, and the reported percentages (e.g., 'Features Coverage 100%') appear to combine 4-7 ratings rather than the full distribution. Please provide the full response distribution per item (1-7) and a precise definition of the percentage. This is a reporting clarity issue, but it affects the interpretability of the central evidence.
minor comments (5)
- [Sec. 3.3] The LLM is named as 'GPT-5.4'; check whether this is the intended model identifier and whether the exact version/temperature settings are captured in the released prompts. Model-specific behavior may matter for reproducibility.
- [Sec. 5] The public access URL '43.163.231.237/' is an IP address that may not be persistent. It would be safer to point readers to the OSF repository for a stable demo link.
- [Sec. 6.2] The age-range sentence is garbled ('18–54 3 5 10 00 1') and needs to be formatted as a proper distribution table.
- [References] In the text, '[41] Chen et al.' and '[39] Chen et al.' do not match the reference entries, which list 'C. Zhu-Tian' as the first author. Please reconcile author names in citations and the bibliography.
- [Sec. 4] The observational analysis covers only three broadcasts (one per sport) and is described as illustrative. This is fine, but the manual annotation reliability is not reported; a short paragraph acknowledging this and noting that no inferential claims are made would improve precision.
Circularity Check
No significant circularity; the central authoring claim is independently grounded in a user study, and self-citations are transparent and non-load-bearing.
full rationale
The paper's derivation chain runs from an automated data pipeline (Sec. 3) and an observational broadcast analysis (Sec. 4) to a technology probe (Sec. 5) and an empirical user study (Sec. 6). The central claim that SwimComposer 'generally enables effective narrative authoring and storytelling' rests on the user study: pre-registered, with 9 experienced participants, Likert ratings, qualitative coding, and analysis of authored videos. This is external evidence relative to the design rationale, not a restatement of it. The only potentially self-referential element is the insight-taxonomy coverage check in Sec. 3.3 (98.15% of coded key moments covered by the authors' proposed insight types). However, the paper does not state that the taxonomy was fitted to the same 15-race commentary corpus; it reports the validation on a separate set of races and the coverage is a descriptive measure of taxonomy recall, not a fitted parameter renamed as a prediction. No equation or construct is defined in terms of another in a way that forces the result. Self-citations to Yao et al. [35,37] are used for definitions, motivation, visualization designs, and processing scripts; the paper explicitly says it follows those designs 'to keep the focus of our work on narrative authoring rather than on proposing new graphical encodings,' so this is transparent reuse rather than a load-bearing self-citation chain. The unvalidated accuracy of tracking, transcription, and LLM insight extraction is a genuine correctness risk but not a circularity; it concerns whether the pipeline produces true data, not whether the paper's reasoning reduces to its inputs. Overall, the central claim is independently supported and no circular step is exhibited.
Assumptions & free parameters
assumptions (4)
- domain assumption Prior visualization-in-motion designs for swimming data (Yao et al. [37]) are valid and reusable
- domain assumption SAM3, WhisperX, and GPT-5.4 perform accurately enough for the pipeline
- domain assumption The six-category insight taxonomy covers the key moments that matter for narrative authoring
- domain assumption Participants' subjective ratings reflect authoring effectiveness
Cite this review
Pith. "Pith review of Authoring Narrative Visualization in Motion: Visual Storytelling in Swimming Videos." pith.science (2026). https://pith.science/paper/TVSWR2TN
@misc{pith2026260714924,
author = {Pith},
title = {Pith review of: Authoring Narrative Visualization in Motion: Visual Storytelling in Swimming Videos},
year = {2026},
howpublished = {\url{https://pith.science/paper/TVSWR2TN}},
note = {Machine review of arXiv:2607.14924}
}
read the original abstract
We investigate how to support authoring narrative visualizations in motion in sports videos, drawing on automated data preparation, systematic analysis, technology probe design, and evaluation, using swimming races as a case study. Sports videos are widely broadcast and shared across social media, where content creators increasingly seek to present and explain complex events to general audiences. Visualization in motion has been explored as an efficient way to embed data into videos and to move with the data referents, providing additional information and helping audiences understand races. However, existing approaches primarily focus on embedding visualizations in videos, lacking exploration of how to support authoring narratives that coordinate views, data, and temporal progression to explain the unfolding races. To address this gap, we use swimming videos as an ideal case for exploration, as swimming is a sport with rich, dynamic data and visualizations in practice. We develop an automated pipeline that extracts structured data from videos, derive narrative constructs through observational analysis of sports broadcasts, and design a technology probe that supports authoring using data prepared by our pipeline and narrative constructs derived from our observations. We evaluate our approach with experienced content creators and/or graphic designers to examine the benefits and challenges of authoring narrative visualizations in motion. All supplemental materials are described in the Supplemental Material Pointers section and are on OSF: osf.io/bq47n/.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
F. Amini, N. Henry Riche, B. Lee, C. Hurter, and P. Irani. Understanding data videos: Looking at narrative visualization through the cinematography lens. InProceedings of the Conference on Human Factors in Computing Systems, pp. 1459–1468. ACM, New York, NY , USA, 2015. doi: 10. 1145/2702123.2702431 2
arXiv 2015
- [2]
-
[3]
M. Bain, J. Huh, T. Han, and A. Zisserman. Whisperx: Time-accurate speech transcription of long-form audio.INTERSPEECH 2023, pp. 4489– 4493, 2023. doi: 10.21437/interspeech.2023-78 4
-
[4]
R. Cao, S. Dey, A. Cunningham, J. Walsh, R. T. Smith, J. E. Zucco et al. Examining the use of narrative constructs in data videos.Visual Informatics, 4(1):8–22, 2020. doi: 10.1016/j.visinf.2019.12.002 2
-
[5]
N. Carion, L. Gustafson, Y .-T. Hu, S. Debnath, R. Hu, D. Suris et al. Sam 3: Segment anything with concepts, 2025. doi: 10.48550/arXiv.2511. 16719 3
-
[6]
N. Cohn. Visual narrative structure.Cognitive Science, 37(3):413–452,
-
[7]
F. Grioui, Y . A. Amzir, N. Doerr, and T. Blascheck. Comparing pre- attentive visual variables in a target identification task for glanceable visualizations. InProceedings of the IEEE Visualization and Visual An- alytics Conference, pp. 371–375, 2025. doi: 10.1109/VIS60296.2025. 00080 2
arXiv 2025
- [8]
Show all 42 references
-
[9]
Hutchinson, W
H. Hutchinson, W. Mackay, B. Westerlund, B. B. Bederson, A. Druin, C. Plaisant et al. Technology probes: Inspiring design for and with families. InProceedings of the Conference on Human Factors in Computing Sys- tems, pp. 17–24. ACM, New York, NY , USA, 2003. doi: 10.1145/6426...
2003 doi
-
[10]
Around 5 billion people - 84 per cent of the potential global audience followed the olympic games paris 2024
International Olympic Committee. Around 5 billion people - 84 per cent of the potential global audience followed the olympic games paris 2024. https://www.olympics.com/ioc/news/around-5-billion-peo ple-84-per-cent-of-the-potential-global-audience-follo wed-the-olympic-games-pa...
2024
-
[11]
Islam, L
A. Islam, L. Yao, A. Bezerianos, T. Blascheck, T. He, B. Lee et al. Re- flections on visualization in motion for fitness trackers. InMobileHCI Workshop on New Trends in HCI and Sports. Vancouver, Canada, Sept
-
[12]
Kim and J
Y . Kim and J. Heer. Gemini: A grammar and recommender system for animated transitions in statistical graphics.IEEE Transactions on Visualization and Computer Graphics, 27(2):485–494, 2021. doi: 10. 1109/TVCG.2020.3030360 2
2021
-
[13]
Kim and J
Y . Kim and J. Heer. Gemini2: Generating keyframe-oriented animated transitions between statistical graphics. InProceedings of the IEEE Visu- alization Conference, pp. 201–205, 2021. doi: 10.1109/VIS49827.2021. 9623291 2
2021
-
[14]
C. Lee, T. Lin, H. Pfister, and C. Zhu-Tian. Sportify: Question answering with embedded visualizations and personified narratives for sports video. IEEE Transactions on Visualization and Computer Graphics, 31(1):12–22,
-
[15]
W. Li, Z. Wang, Y . Wang, D. Weng, L. Xie, S. Chen et al. Geocamera: Telling stories in geographic visualizations with camera movements. In Proceedings of the Conference on Human Factors in Computing Systems, art. no. 170, 15 pp. ACM, New York, NY , USA, 2023. doi: 10.1145/ 35...
2023
-
[16]
LimeSurvey: Free online survey tool, 2026
LimeSurvey GmbH. LimeSurvey: Free online survey tool, 2026. Official product page. Accessed March 30, 2026. 7
2026
-
[18]
T. Lin, Z. Chen, J. Beyer, Y . Wu, H. Pfister, and Y . Yang. The ball is in our court: Conducting visualization research with sports experts. IEEE Computer Graphics and Applications, 43(1):84–90, 2023. doi: 10. 1109/MCG.2022.3222042 2
2023
-
[19]
T. Lin, R. Xiang, G. Liu, D. Tiwari, M.-C. Chiang, C. Ye et al. Sports- buddy: Designing and evaluating an AI-powered sports video storytelling tool through real-world deployment. InProceedings of the Pacific Visu- alization Conference, pp. 214–223, 2025. doi: 10.1109/PacificV...
2025
-
[20]
T. Lin, C. Zhu-Tian, Y . Yang, D. Chiappalupi, J. Beyer, and H. Pfister. The quest for omnioculars: Embedded visualization for augmenting bas- ketball game viewing experiences.IEEE Transactions on Visualization and Computer Graphics, 29(1):962–972, 2023. doi: 10.1109/TVCG.2022...
2023 doi
-
[21]
Satyanarayan and J
A. Satyanarayan and J. Heer. Authoring narrative visualizations with ellipsis.Computer Graphics Forum, 33(3):361–370, 2014. doi: 10.1111/ cgf.12392 2
2014
-
[22]
Segel and J
E. Segel and J. Heer. Narrative visualization: Telling stories with data. IEEE Transactions on Visualization and Computer Graphics, 16(6):1139– 1148, 2010. doi: 10.1109/TVCG.2010.179 2
2010 doi
-
[23]
L. Shen, H. Li, Y . Wang, T. Luo, Y . Luo, and H. Qu. Data Playwright: Authoring data videos with annotated narration.IEEE Transactions on Visualization and Computer Graphics, 31(9):5884–5897, 2025. doi: 10. 1109/TVCG.2024.3477926 2
2025
-
[24]
L. Shen, Y . Zhang, H. Zhang, and Y . Wang. Data Player: Automatic gener- ation of data videos with narration-animation interplay.IEEE Transactions on Visualization and Computer Graphics, 30(1):109–119, 2024. doi: 10. 1109/TVCG.2023.3327197 2
2024
-
[25]
D. Shi, F. Sun, X. Xu, X. Lan, D. Gotz, and N. Cao. Autoclips: An auto- matic approach to video generation from data facts.Computer Graphics Forum, 40(3):495–505, 2021. doi: 10.1111/cgf.14324 2
2021 doi
-
[26]
Y . Shi, X. Lan, J. Li, Z. Li, and N. Cao. Communicating with motion: A design space for animated visual narratives in data videos. InProceedings of the Conference on Human Factors in Computing Systems, art. no. 605, 13 pp. ACM, New York, NY , USA, 2021. doi: 10.1145/3411764.3445337 2
2021
-
[27]
M. Shin, J. Kim, Y . Han, L. Xie, M. Whitelaw, B. C. Kwon et al. Roslingi- fier: Semi-automated storytelling for animated scatterplots.IEEE Trans- actions on Visualization and Computer Graphics, 29(6):2980–2995, 2023. doi: 10.1109/TVCG.2022.3146329 2
2023
-
[28]
Stein, H
M. Stein, H. Janetzko, A. Lamprecht, T. Breitkreutz, P. Zimmermann, B. Goldlücke et al. Bring it to the pitch: Combining video and movement 10 © 2026 IEEE. This is the author’s version of the article that has been accepted by IEEE VIS 2026 and will be published in IEEE Transac...
2026 doi
-
[29]
J. Tang, L. Yao, L. Ying, R. Vuillemot, and P. Isenberg. Diving deep into time: Temporal arrangements for embedded visualization in swimming videos.IEEE Transactions on Visualization and Computer Graphics, pp. 1–17, 2026. doi: 10.1109/TVCG.2026.3689361 1, 2
2026
-
[30]
T. Tang, J. Tang, J. Lai, L. Ying, Y . Wu, L. Yu et al. Smartshots: An optimization approach for generating videos with data visualizations em- bedded.ACM Transactions on Interactive Intelligent Systems, 12(1), art. no. 4, 21 pp., Mar. 2022. doi: 10.1145/3484506 1
2022 doi
-
[31]
Y . Wang, Y . Gao, R. Huang, W. Cui, H. Zhang, and D. Zhang. Animated presentation of static infographics with infomotion.Computer Graphics Forum, 40(3):507–518, 2021. doi: 10.1111/cgf.14325 2
2021 doi
-
[32]
Y . Wang, L. Shen, Z. You, X. Shu, B. Lee, J. Thompson et al. Wonderflow: Narration-centric design of animated data videos.IEEE Transactions on Visualization and Computer Graphics, 31(9):4638–4654, 2025. doi: 10. 1109/TVCG.2024.3411575 2
2025
-
[33]
X. Xu, A. Wu, L. Yang, Z. Wei, R. Huang, D. Yip et al. Is it the end? Guidelines for cinematic endings in data videos. InProceedings of the Conference on Human Factors in Computing Systems, art. no. 171, 16 pp. ACM, New York, NY , USA, 2023. doi: 10.1145/3544548.3580701 2
2023
-
[34]
X. Xu, L. Yang, D. Yip, M. Fan, Z. Wei, and H. Qu. From ‘wow’ to ‘why’: Guidelines for creating the opening of a data video with cinematic styles. InProceedings of the Conference on Human Factors in Computing Systems, art. no. 599, 20 pp. ACM, New York, NY , USA, 2022. doi: 10...
2022
-
[35]
L. Yao, A. Bezerianos, R. Vuillemot, and P. Isenberg. Visualization in motion: A research agenda and two evaluations.IEEE Transactions on Visualization and Computer Graphics, 28(10):3546–3562, Oct. 2022. doi: 10.1109/TVCG.2022.3184993 1, 2, 9
2022
-
[36]
L. Yao, F. Bucchieri, V . Mcarthur, A. Bezerianos, and P. Isenberg. User experience of visualizations in motion: A case study and design consid- erations.IEEE Transactions on Visualization and Computer Graphics, 31(1):174–184, Jan. 2025. doi: 10.1109/tvcg.2024.3456319 2
2025
-
[37]
L. Yao, R. Vuillemot, A. Bezerianos, and P. Isenberg. Designing for visualization in motion: Embedding visualizations in swimming videos. IEEE Transactions on Visualization and Computer Graphics, 30(3):1821– 1836, 2024. doi: 10.1109/TVCG.2023.3341990 1, 2, 3, 5, 6, 9
2024
-
[38]
L. Ying, Y . Wang, H. Li, S. Dou, H. Zhang, X. Jiang et al. Reviving static charts into live charts.IEEE Transactions on Visualization and Computer Graphics, 31(8):4314–4328, 2025. doi: 10.1109/TVCG.2024.3397004 2
2025
-
[39]
Zhu-Tian, Q
C. Zhu-Tian, Q. Yang, J. Shan, T. Lin, J. Beyer, H. Xia et al. iBall: Aug- menting basketball videos with gaze-moderated embedded visualizations. InProceedings of the Conference on Human Factors in Computing Sys- tems, art. no. 841, 18 pp. ACM, New York, NY , USA, 2023. doi: 1...
2023
-
[40]
Zhu-Tian, Q
C. Zhu-Tian, Q. Yang, X. Xie, J. Beyer, H. Xia, Y . Wu et al. Sporthesia: Augmenting sports videos using natural language.IEEE Transactions on Visualization and Computer Graphics, 29(1):918–928, 2023. doi: 10. 1109/TVCG.2022.3209497 1, 2
2023
-
[41]
Zhu-Tian, S
C. Zhu-Tian, S. Ye, X. Chu, H. Xia, H. Zhang, H. Qu et al. Augmenting sports videos with viscommentator.IEEE Transactions on Visualization and Computer Graphics, 28(1):824–834, 2022. doi: 10.1109/TVCG.2021. 3114806 1, 2 11
2022 doi
-
[2013]
doi: 10.1111/cogs.12016 2
- [2022]
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.