{"id":"426e82c1-d134-467e-bd81-3c4a13c43573","arxiv_id":"2504.19267","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A LoRA fine-tune of VideoGPT+ on VIST produces strong reference-free metric scores for visual storytelling, but the state-of-the-art claim is undercut by weak baselines and test-set selection.","lead":"This paper fine-tunes an existing video language model (VideoGPT+) on the Visual Storytelling dataset to generate five-sentence stories from five-image sequences. The authors report that their model, VIST-GPT v2, scores higher than four older models on reference-free story quality metrics, though the comparison omits recent LLM-based storytelling systems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The benchmark claim is not supported by the evaluation protocol: the authors selected VIST-GPT v2 by tuning prompts and decoding during inference (Section 4.2), then report Tables 2 and 3 on the same 900-example VIST test subset, with no validation split described.","rationale":"The reader's overall conditional verdict is appropriate, and I agree that the paper's empirical claim needs strengthening. However, I do not think the frozen visual encoders are the single most load-bearing weakness. The paper's own text in Section 4.2 indicates that v1 and v2 were produced by optimizing prompts and decoding parameters during inference, and the only evaluation described is on the 900-example VIST test subset. Without a validation split, the reported superiority of v2 could be explained entirely by test-set selection. That directly undermines the central claim of a new benchmark, independent of whether the visual features are sufficient. Once the evaluation protocol is fixed, the frozen-encoder question becomes testable with an image-scrambling or image-removal control, so it is not the first blocker. The abstract's claim that RoViST and GROOViST are novel metrics is also wrong—they are cited prior work—and should be corrected, though it does not invalidate the engineering. The paper's contribution is plausible and reproducible in principle with public components, so conditional acceptance with required revisions is the right level. My concrete test targets the selection protocol, and if it fails, the SOTA claim should be removed or downgraded to a comparison that controls for validation-based selection.","tokens_in":17695,"tokens_out":4452,"duration_ms":49783,"concrete_test":"Reproduce the comparison with a proper validation split. After LoRA fine-tuning on VIST train, tune the prompt, temperature, and beam count on the VIST validation split (or a held-out subset of train), select one model version, and then compute Tables 2 and 3 exactly once on the 900-example test intersection. If v2 remains the best model on validation and its test scores are recomputed without further parameter or prompt changes, the benchmark claim survives. If v2 is best only on the test subset, the current headline is a selection artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim—that VIST-GPT v2 sets a new benchmark in grounding, coherence, fluency, and dHM—rests on scores computed on a 900-example intersection of VIST test predictions. The paper's own description of model selection makes those scores uninterpretable as test performance. Section 4.2 states: \"After completing the training process, we conducted several experiments during the inference phase, testing various parameters and prompts to optimize the model's performance. This experimentation resulted in two distinct versions of the model, VIST-GPT v1 and VIST-GPT v2.\" No validation split or held-out subset is mentioned in Sections 4 or 5. The refined v2 prompt (\"Given a sequence of five images, write a short coherent story of five sentences, with each sentence corresponding to an image\") and the decoding choices appear to have been tuned against the same metrics and the same examples on which the benchmark is then reported. If v2 was chosen by comparing v1 and v2 on the 900-example test subset, the reported superiority and the dHM of 0.0459 are selection artifacts rather than estimates of generalization. A second, compounding problem is that the comparison omits modern LLM-based storytelling baselines such as StoryLLaVA, so even an honestly computed v2 score would not establish a \"new benchmark.\" The frozen-encoder concern raised by the reader is real but secondary: the evaluation protocol must be fixed before any interpretation of the visual-grounding number, including a language-prior explanation, can be trusted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents VIST-GPT, a visual storytelling model that adapts VideoGPT+'s dual image-video encoder architecture with a Phi-3-mini LLM fine-tuned via LoRA on the VIST dataset. The authors evaluate two inference variants (v1 and v2) on a 900-example intersection of the VIST test set using reference-free metrics (GROOViST, RoViST, UniEval) and the Human-to-Machine Distance (dHM). They claim that VIST-GPT v2 achieves state-of-the-art visual grounding, coherence, fluency, and the lowest dHM among compared models, thereby 'ushering in the era of visual storytelling with LLMs.' The technical contribution is an efficient fine-tuning recipe rather than a fundamentally new architecture.","tokens_in":17951,"tokens_out":4830,"duration_ms":48200,"significance":"If the empirical claims were sound, the paper would provide a useful, lightweight visual storytelling model with a sensible choice of reference-free evaluation metrics. The use of GROOViST and RoViST is appropriate for the task, and the efficiency of LoRA fine-tuning of a small LLM is attractive. However, the central benchmark claim is currently undermined by the evaluation protocol: the model version and decoding hyperparameters appear to have been selected on the test set itself, and the comparison omits several recent LLM-based storytelling baselines, including StoryLLaVA, which the paper cites in its own related work. Thus the significance of the contribution as a 'new benchmark' is not established, though the underlying methodology could be valid after a corrected evaluation.","major_comments":[{"comment":"The reported benchmark scores are invalid as test-set estimates because the model version was selected on the test set. Section 4.2 states that 'after completing the training process, we conducted several experiments during the inference phase, testing various parameters and prompts to optimize the model's performance,' resulting in v1 and v2. No validation split is described anywhere in Sections 4–5. The subsequent Tables 2 and 3 report GROOViST, RoViST-C, RoViST-NR, and dHM for v2 on the same 900-example test subset used to choose v2 over v1. Consequently, the values 0.9962, 0.7837, and 0.0459 are selection artifacts rather than estimates of generalization. The authors must either (a) introduce a proper validation split, tune only on it, and then report test results exclusively for the final version, or (b) clearly reframe these numbers as in-sample tuning results. This is load-bearing because the 'new benchmark' claim rests directly on these tables.","section":"§4.2 and §5.3 (Tables 2–3)"},{"comment":"The comparison omits the most relevant modern baselines. Section 3.3 discusses StoryLLaVA as a recent visual storytelling model, yet Tables 2 and 3 compare only AREL, GLACNET, KG-Story, and MCSM+BART. The absence of StoryLLaVA and other LLM-based storytelling systems means that even with a clean evaluation, the paper would not substantiate 'v2 sets a new benchmark for visual storytelling.' The authors should include predictions from StoryLLaVA and, if feasible, from VideoGPT+ (the base model) and other recent MLLMs, with comparable inference settings. Without these, the claim is unsupported.","section":"§3.3 vs §5.3"},{"comment":"The visual grounding results may be inflated by language priors because the visual encoders and adapters are frozen and only the LLM is fine-tuned. GROOViST measures alignment between nouns in the story and image regions; if the fine-tuned LLM produces generic nouns frequent in VIST training stories (e.g., 'friends,' 'park,' 'family'), the score can be high even when the model does not actually perceive the image content. The paper does not provide any diagnostic evidence to separate genuine visual grounding from language-prior effects. I recommend adding a per-image grounding analysis, or an experiment where the visual input is replaced with noise or a mismatched image sequence, to verify that the high GROOViST score is not an artifact of the text distribution. This caveat is important because the central claim includes 'highest visual grounding.'","section":"§4.1–4.2 and §5.2"}],"minor_comments":[{"comment":"The dHM computation requires a human story H, but VIST has multiple human stories per image sequence. Please specify which human story is used (e.g., the first reference, a random reference, or an average over references) in the calculation of dH, dG, and dR; this affects the comparability of the distances across models.","section":"§5.3, Eq. (1)–(2)"},{"comment":"The UniEval scores reported in the text (mean coherence 0.855, understandability 0.9, fluency 0.950 for V2) are not shown in a table and are only given in Figure 5, which is not reproduced in the text. Please include a table with all model scores for all UniEval dimensions so that readers can reproduce the comparison.","section":"§5.4"},{"comment":"Please clarify how the additional three frames are padded when passing a five-image story to InternVideo2, which processes eight frames. Padding with zero frames versus repeating frames could affect temporal modeling and should be described precisely.","section":"§4.2"},{"comment":"The paper has numerous typos and formatting issues, including 'LL Ms' in the title, 'W e' in the abstract, 'judegment' in Section 5.2, and some incomplete figure captions. A thorough proofreading pass is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The central empirical claim is currently not supportable because of test-set selection and missing baselines. I believe the underlying model could be publishable after a proper validation protocol is followed and the comparison is updated with recent LLM-based storytelling models. Recommend that the editor require a substantive revision rather than a quick fix, since the authors will need to rerun inference selection and possibly collect predictions from additional baselines."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper in one line: the engineering is real, the headline is not. The authors fine-tune VideoGPT+ on VIST with LoRA, frozen CLIP and InternVideo2 encoders, and get strong reference-free metric scores. But the central claim—that VIST-GPT v2 sets a new benchmark—does not survive contact with the evaluation protocol. In Section 4.2 they say they ran inference-phase experiments, varying prompts and parameters, and ended up with v1 and v2. No validation split is mentioned. Then Tables 2 and 3 report both versions on the same 900-example test subset. Choosing v2 by looking at test scores and reporting those scores as its performance is test-set selection. The specific numbers (GROOViST 0.996, dHM 0.046) are uninterpretable as estimates of generalization. The stress-test note is right: this has to be fixed before any interpretation of the visual-grounding numbers, including a language-prior explanation, can be trusted.\n\nWhat is genuinely new: to my knowledge, this is the first fine-tuning of VideoGPT+ specifically for VIST. The architecture choices—frozen dual encoders, trainable adapters, LoRA on Phi-3—are sensible and clearly described. Using RoViST, GROOViST, dHM, and UniEval together is a reasonable evaluation bundle, even though all of those metrics are prior work. The abstract's word 'novel' applied to RoViST and GROOViST is simply wrong.\n\nSoft spots, in order of severity. (1) The test-set selection issue is load-bearing and fixable: rerun with a proper held-out validation split. (2) The comparison omits StoryLLaVA and other LLM-based storytellers, so even honestly computed scores would not establish a 'new benchmark.' (3) The hallucination-reduction claim in the Discussion has no quantitative support—just five curated qualitative examples. (4) Minor: the abstract misattributes prior metrics as novel. The citation list is broad; self-citations are present but not egregious.\n\nThis is not a fatal rejection. The recipe and evaluation framework are worth sharing with the visual-storytelling community. The paper is for practitioners adapting video LLMs to sequence-to-story tasks, not for readers seeking a methodological advance. I'd send it to peer review: a serious reviewer can insist on a proper validation split and modern baselines, and the corrected version would be a useful systems contribution.","headline":"A plausible fine-tuning recipe undermined by test-set tuning and an overclaimed benchmark—fix the protocol and it's a useful systems paper.","tokens_in":18600,"tokens_out":2979,"would_cite":false,"duration_ms":31006,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fine-tuning a small multimodal LLM yields visual stories that sit closer to human narratives than prior models.","keywords":["visual storytelling","VIST-GPT","multimodal large language model","dual visual encoder","visual grounding","reference-free evaluation","LoRA fine-tuning","coherence"],"falsifier":"A concrete falsifier: feed the model image sequences after shuffling or replacing the images while keeping the same prompt, and compare grounding; if GROOViST stays near 0.9962 when the images are replaced by unrelated images, the score does not measure visual grounding.","tokens_in":17420,"feed_emoji":"🎬","tokens_out":8248,"duration_ms":76234,"temperature":0.7,"pith_summary":"The paper argues that visual storytelling does not need a purpose-built architecture: a video-understanding multimodal LLM, fine-tuned on the VIST dataset, can turn a five-image sequence into a coherent five-sentence story. On the reference-free metrics GROOViST, RoViST, and UniEval, the resulting VIST-GPT v2 model scores highest among compared models for visual grounding, coherence, and fluency, and its stories sit closest to human-written ones by the Human-to-Machine Distance. The authors' central claim is that fine-tuning the language model while keeping the visual encoders frozen is enough to achieve strong visual grounding and to reduce hallucinated or irrelevant story details. This matters because n-gram metrics such as BLEU and CIDEr are poorly suited to storytelling, where many valid stories exist for the same image sequence.","feed_headline":"Fine-tuned LLM writes image stories closest to human ones","feed_subtitle":"On a 900-story VIST benchmark, VIST-GPT v2 scores highest on visual grounding, coherence, and fluency.","key_machinery":"The load-bearing mechanism is the dual visual encoder pipeline inherited and adapted from VideoGPT+: CLIP ViT-L/14 supplies spatial features, InternVideo2 (an 8-frame video model) supplies temporal dynamics, trainable vision-language adapters project both into the LLM embedding space, and token pooling compresses the visual tokens. The Phi-3-mini-4k-instruct language model is then fine-tuned on VIST with LoRA while the visual side stays frozen, and at inference a task prompt plus low temperature and beam search enforce one sentence per image. The claim-carrying component is the fine-tuned LLM: it is what converts generic video understanding into story structure, and the comparison of v1 versus v2 shows that a more constraining prompt and decoding setup improves grounding and coherence.","core_discovery":"VIST-GPT v2, built from the VideoGPT+ multimodal backbone with a CLIP ViT-L/14 image encoder, an InternVideo2 video encoder, and a Phi-3-mini-4k-instruct LLM fine-tuned with LoRA, establishes what the authors describe as a new benchmark for visual storytelling. On a 900-example intersection of VIST test predictions, it attains a GROOViST visual grounding score of 0.9962, a RoViST-C coherence score of 0.7837, and a UniEval fluency score of 0.950, with the lowest Human-to-Machine Distance (dHM = 0.0459) among AREL, GLACNET, KG Story, MCSM+BART, and VIST-GPT v1. The claim is that the model's narratives align with the objects and events in the image sequence, flow logically sentence to sentence, and avoid the hallucinated details that plague prior models, because the LLM was fine-tuned on story-level data rather than merely prompted.","pith_inferences":["The 900-example intersection may not represent the full VIST test set, and if baseline predictions are missing non-randomly, the reported margins could shift on a complete evaluation.","A text-only control experiment, feeding the same prompts with no visual tokens, would directly test whether the grounding scores reflect genuine image understanding or language priors learned from VIST.","Because the visual encoders are frozen, scaling the LLM or unfreezing the adapters on more diverse story data could push grounding further; that is a natural extension the authors do not run.","The qualitative appendix suggests the model captures social dynamics and emotional tone, but no metric in the paper measures those dimensions, so a reader should treat that as anecdotal."],"forward_implications":["If the central claim holds, fine-tuning an existing video-centric multimodal LLM on VIST is a viable route to visual storytelling, without training a purpose-built encoder-decoder from scratch.","The high scores imply that frozen visual features are sufficient for story grounding, so future gains should come mostly from the language side or from decoding strategy.","The reference-free metrics RoViST and GROOViST, plus the Human-to-Machine Distance, give a reusable evaluation protocol that does not penalize valid but different stories, unlike BLEU or CIDEr.","VIST-GPT v2's lower dHM means that on this benchmark the model's stories share the grounding, coherence, and non-redundancy profile of human stories more closely than the four earlier storytelling models tested."],"supporting_citations":[{"why":"supplies the VIST dataset of five-image sequences with human five-sentence stories that the model is fine-tuned on and evaluated against.","marker":"[11]"},{"why":"provides the VideoGPT+ architecture with image and video encoders and the vision-language adapters the paper inherits.","marker":"[19]"},{"why":"is the frozen CLIP ViT-L/14 image encoder used for spatial visual features.","marker":"[23]"},{"why":"is the frozen InternVideo2 encoder used for temporal dynamics across the five image frames.","marker":"[35]"},{"why":"is the Phi-3-mini-4k-instruct LLM that is fine-tuned with LoRA to generate the stories.","marker":"[1]"},{"why":"defines the RoViST reference-free coherence and non-redundancy metrics used as evaluation dimensions.","marker":"[32]"},{"why":"defines the GROOViST visual grounding metric behind the paper's headline grounding score.","marker":"[27]"},{"why":"defines the Human-to-Machine Distance dHM that ranks how close generated stories are to human stories.","marker":"[28]"},{"why":"supplies the UniEval framework used for the fluency, coherence, and understandability comparisons.","marker":"[44]"},{"why":"is the MCSM+BART baseline, the strongest prior storytelling model whose VIST test predictions the paper must beat in the central comparison.","marker":"[3]"}],"fun_headline_variants":["VIST-GPT v2 writes image stories closest to human ones","Fine-tuned LLM image tales nearest human narratives","VIST-GPT v2 sets benchmark for human-like story generation","New LLM delivers most human-like visual stories","LLM image stories now closest to human judgment"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the frozen CLIP and InternVideo2 features, carried through the inherited adapters, contain enough visual information for the fine-tuned LLM to ground its stories; if they do not, the strong grounding scores could reflect language priors rather than image content.","fun_headline_variants_meta":{"raw":{"variants":["VIST-GPT v2 writes image stories closest to human ones","Fine-tuned LLM image tales nearest human narratives","VIST-GPT v2 sets benchmark for human-like story generation","New LLM delivers most human-like visual stories","LLM image stories now closest to human judgment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000876,"raw_usage":{"total_tokens":3770,"prompt_tokens":910,"completion_tokens":2860,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":2780}},"tokens_in":526,"tokens_out":2860,"duration_ms":23787,"temperature":1.0,"reasoning_tokens":2780,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:56:26.195880+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete falsifier: feed the model image sequences after shuffling or replacing the images while keeping the same prompt, and compare grounding; if GROOViST stays near 0.9962 when the images are replaced by unrelated images, the score does not measure visual grounding.","supporting_citations":[{"cited_title":"Visual storytelling","cited_arxiv_id":null,"evidence_quote":"supplies the VIST dataset of five-image sequences with human five-sentence stories that the model is fine-tuned on and evaluated against."},{"cited_title":"Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition","cited_arxiv_id":"2407.04559","evidence_quote":"defines the Human-to-Machine Distance dHM that ranks how close generated stories are to human stories."},{"cited_title":"Commonsense knowledge aware concept selec- tion for diverse and informative visual storytelling","cited_arxiv_id":null,"evidence_quote":"is the MCSM+BART baseline, the strongest prior storytelling model whose VIST test predictions the paper must beat in the central comparison."}],"review_version":1}