{"id":"e4f735a0-26db-45c9-8bb4-7baa873b6235","arxiv_id":"2505.13860","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A three-stage curriculum (concept alignment, instruction tuning, downstream fine-tuning) adapted LLaVA-NeXT-Video to soccer, raising action classification accuracy from 11.8% to 63.5% and the VQA relative score from about 60 to 83.","lead":"The authors adapt an open-source video language model to soccer using a three-stage curriculum built from synthetic captions and question-answer pairs generated by Claude 3.5. Their best model lifts soccer action classification from 11.8% to 63.5% accuracy and improves visual question answering over the base model by 37.5% relative.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"VQA gains may be Claude style-mimicry, and the hard-label action accuracy appears as both 57.8% and 63.5%; exact magnitudes are not yet supported.","rationale":"After reading the full paper and the Reader's verdict, I find the Reader's identification of the Claude-to-Claude evaluation loop is the most load-bearing concern. The synthetic caption and QA generation pipeline (Figures 8-10) uses the ground-truth event label and Claude 3.5 Sonnet for both data creation and scoring; the 37.5% VQA gain and caption-score gains are therefore not cleanly separable from style imitation. The paper itself states in Limitations that synthetic data may introduce stylistic biases, which supports the concern. I do not see a fraudulent or internally broken central construction: the action classification result is hard-label and has an independent base comparison (11.8% vs 63.5%), and the multi-stage ablation (Base to AC 16%, CA to AC 52%, CA to IT to AC 63.5%) makes the recipe's direction credible. However, the exact magnitude is uncertain because the 20K model is reported as both 57.8% and 63.5% in different places and because four rare classes were removed. These issues do not overturn the central claim but justify the CONDITIONAL verdict and a request for released code/predictions and a non-Claude judge. I therefore leave the Reader's verdict unchanged.","tokens_in":14985,"tokens_out":8479,"duration_ms":82426,"concrete_test":"Have the authors release per-sample VQA predictions for Base, 3K, 10K, and 20K on the 200-sample VQA subset, and re-score them with a non-Claude judge (e.g., GPT-4o or human annotators) using the same reference answers and rubric. If the Base-to-20K relative-score gap collapses from the reported 22.6 points, the headline VQA improvement is mostly style mimicry; if the gap persists, the circularity concern is refuted. To settle the magnitude question for action classification, also ask for the test set sizes behind 57.8% and 63.5% and for accuracy on the original 17 classes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Stages 1 and 2 are built from Claude 3.5 Sonnet text: captions are generated from 8 frames plus the ground-truth event label (Figure 8), QA pairs are generated from those captions (Figure 9), reference answers are written by Claude, and the caption/VQA judge is Claude. The adapted model is therefore trained to imitate Claude's output style and then scored by Claude, so the 37.5% VQA gain (60.20 to 82.79 in Table 2) and the caption-score gains may partly reflect style alignment rather than soccer understanding. The limitations section concedes synthetic-data stylistic bias, reinforcing this. The hard-label action-classification task is the main independent anchor, but its headline number is internally inconsistent: Section 5.2 reports the 20K model at 57.8%, while Figure 4's caption and Table 7 report 63.5%, without explanation. In addition, four rare classes were removed post hoc, so the reported accuracy applies only to a 13-class balanced subset. The multi-stage recipe is plausible and the Base to AC 16% versus CA to IT to AC 63.5% ablation supports the method's direction, but the exact claimed magnitudes are not reproducible from the paper as written.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper adapts LLaVA-NeXT-Video to soccer video understanding using a three-stage curriculum: Stage 1, Soccer Concept Alignment, trains the model on synthetic captions generated by Claude 3.5 Sonnet from 8 frames plus the ground-truth event label; Stage 2, Soccer VQA Instruction Tuning, trains on question-answer pairs generated by Claude from those captions; Stage 3 fine-tunes on a downstream soccer action classification task. The authors report that the final 20K model improves VQA relative score from 60.20 to 82.79 and action classification accuracy from 11.8% to 63.5%, and they provide ablations on dataset size, training stage order, LoRA rank, projector training, and image-plus-video training. The central claim is that this multi-stage, compute-efficient recipe is an effective way to adapt a general-purpose video VLM to a specialized domain.","tokens_in":15215,"tokens_out":4026,"duration_ms":36860,"significance":"If the results hold, this is a practically useful contribution: the recipe is simple, uses only four 24G A10 GPUs, and the action classification anchor (11.8% to 63.5%) is an independent, hard-label signal that does not depend on the LLM judge. The ablations are reported honestly and cover the key design decisions (training order, dataset scale, LoRA rank, projector choice, clip length). The main weakness is that the caption and VQA headline numbers are produced inside a partially self-referential evaluation loop: Claude generates the training captions, the QA pairs, and the reference answers, and Claude also judges the outputs. The hard-label action classification task is therefore the most important evidence, but the reported accuracy for that task is internally inconsistent between Section 5.2, Figure 4, Table 7, and Table 14.","major_comments":[{"comment":"The headline VQA and caption improvements rest on a partially self-referential evaluation loop. Stage 1 captions are generated by Claude 3.5 Sonnet from 8 frames plus the ground-truth event label (Figure 8); Stage 2 QA pairs are generated from those captions by the same model (Figure 9); the VQA reference answers are written by Claude; and the judge for both caption correctness/detailness and VQA relative score is Claude (Section 5.1). The adapted model is therefore trained to imitate Claude's output style and then scored by a Claude judge, so the 37.5% VQA gain (Base 60.20 to 20K 82.79 in Table 2) and the caption-score gains in Table 1 may partly reflect stylistic mimicry rather than soccer understanding. Section 5.4 concedes that synthetic data 'may introduce stylistic biases.' To make the VQA and caption claims load-bearing, the paper should add an independent evaluation, such as human ratings or a judge from a different model family, and should report variance across runs; without this, the VQA/caption improvements are not separable from style alignment.","section":"5.1, 5.2, Tables 1-2, Figures 8-9"},{"comment":"The action classification headline is internally inconsistent. Section 5.2 reports the 20K model's accuracy as 'increasing from 11.8% to 57.8%,' while Figure 4's caption reports 0.635 and Table 7 reports 20K accuracy 0.635; Table 14 reports 63.5% for the full CA-to-IT-to-AC sequence. The text does not reconcile 57.8% with 63.5%. If one value is for the 1,300-sample test set and the other for the 100-sample test set, that must be stated explicitly. The abstract's '63.5%' claim is not reproducible from the paper as written. In addition, the paper removes four rare SoccerNet classes (Red card, Yellow-to-Red card, Penalty, Kick off) post hoc, so the reported accuracy applies only to a 13-class balanced subset; this caveat should appear wherever the headline number is quoted.","section":"5.2, Figure 4, Table 7, Table 14, Abstract"},{"comment":"The VQA relative-score metric is not adequately defined or interpreted. Section 5.1 says the predicted response is scored against a 'reference upper bound response' and a relative score is computed by normalizing with that reference, but Table 2 reports scores exceeding 100 (e.g., LLaMA 3.2 reaches 126.77 on Prediction questions), which contradicts the notion of an upper bound. The paper should report the raw predicted and reference scores, define the normalization formula precisely, and explain how scores above 100 are possible. The current presentation makes the 60.20-to-82.79 improvement difficult to interpret, especially without error bars or multiple runs.","section":"5.1, Table 2"}],"minor_comments":[{"comment":"The sentence 'a smaller test set if 100 samples' contains a typo ('if' should be 'of'), and the paper should clarify which test set (1,300 or 100 samples) is used for the numbers in Table 7.","section":"5.1, Task 3"},{"comment":"The metric 'ROGUE' should be 'ROUGE' in both the table and the text.","section":"Table 6 and Section 5.2"},{"comment":"The caption-generation prompt in Figure 8 instructs the model to anonymize team/player names and not to mention jersey numbers, but the QA-generation few-shot examples in Figure 10 include a player 'wearing the number 11 jersey' and specific player references; this inconsistency should be resolved or explained.","section":"Figures 8 and 10"},{"comment":"The 3K and 10K rows of Table 1 are formatted as two-line entries without a clear separator, making it hard to tell which numbers belong to Correctness versus Detailness; please reformat for clarity.","section":"Table 1"},{"comment":"References [6] and [7] are the same SoccerNet-V2 paper and should be merged to avoid duplicate entries.","section":"References"},{"comment":"The paper does not report training hyperparameters (epochs, learning rate, LoRA alpha, batch size, optimizer, warmup). Since the paper claims a repeatable recipe, these values should be added.","section":"4.2, Experiments"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable fit for a computer vision / multimodal domain-adaptation venue. The strongest evidence is the action classification ablation (Table 14), which is an independent hard-label signal, and the authors are honest about the limitations of their synthetic data. The main risk is the circularity of the VQA/caption evaluation and the internal inconsistency in the reported action accuracy; both are fixable within revision. I recommend major revision rather than rejection, provided the authors add an independent evaluation component, reconcile the 57.8% versus 63.5% numbers, and report error bars."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you spend time on this arXiv paper. First, the actual recipe is simple and plausible: take an open video VLM (LLaVA-NeXT-Video), generate synthetic captions and QA pairs with Claude 3.5 Sonnet over 2-second clips, then fine-tune with LoRA in three stages—concept alignment, instruction tuning, downstream task. It runs on 4 A10 GPUs with 20k clips, so it's cheap to replicate. Second, the only headline number I fully trust is the action classification jump, and even that has a wrinkle: Section 5.2 says 57.8% while Figure 4 and Table 7 say 63.5%. That's a real internal inconsistency, not a typo quibble, because the paper's main claim is precisely about magnitude.\n\nWhat the paper does well: the ablations are honest and informative. Base→AC 16%, Base→CA→AC 52%, and Base→CA→IT→AC 63.5% show the curriculum is doing real work. They also ablate LoRA rank, projector training, clip duration, and data size. The limitations section is upfront about synthetic data stylistic bias and hallucination in longer clips. That is good reporting.\n\nThe soft spots are real, though. The VQA and caption evaluations are too self-referential for me to trust the 37.5% gain. Claude writes the captions from 8 frames plus the ground-truth label, writes the QA pairs from those captions, writes the reference answers, and scores the outputs. The model is being trained to imitate Claude and then judged by Claude. The hard-label action task is the independent anchor, and it's a good one, but the discrepancy above means even that number needs a fix. The BLEU/ROUGE improvements are in the right direction but small. Four rare classes were removed without a sensitivity check; that's defensible but should be stated prominently. No code or data, no error bars or seeds.\n\nWho should read it: anyone doing domain adaptation of VLMs in a niche vertical and wanting a cheap starting point. It's also a useful example for teaching evaluation pitfalls. Send it to peer review—it deserves serious referees despite the issues. Ask the authors to resolve the accuracy inconsistency, re-run VQA with a non-Claude judge or human subset, and report variance.","headline":"A practical, compute-efficient recipe for video VLM domain adaptation with a credible hard-label action classification gain, but the headline VQA/caption numbers are weakened by a Claude-to-Claude evaluation loop and one internal inconsistency in the reported accuracy.","tokens_in":15786,"tokens_out":3348,"would_cite":true,"duration_ms":29615,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A three-stage curriculum adapts a general-purpose video language model to soccer, lifting action classification from 11.8% to 63.5%.","keywords":["video vision language model","domain adaptation","curriculum learning","soccer video understanding","visual question answering","action classification","synthetic instruction data","low-rank fine-tuning"],"falsifier":"Take a set of test clips, replace their frames with frames from a different event class while keeping the original event label in the synthetic caption, and measure whether the adapted model's captions and answers follow the label or the frames; if they follow the label, the curriculum has taught label-to-text imitation rather than visual understanding.","tokens_in":14770,"feed_emoji":"⚽","tokens_out":7533,"duration_ms":63341,"temperature":0.7,"pith_summary":"This paper asks whether a general-purpose video vision-language model can be adapted to a specialized domain without a large annotation budget, and it answers yes for soccer. The claim is that a three-stage curriculum—concept alignment, instruction tuning, then downstream-task fine-tuning—on roughly 20,000 two-second event clips raises visual question-answering by 37.5% relative and soccer action classification from 11.8% to 63.5% accuracy. The recipe uses synthetic instruction data distilled from eight frames plus the event label, and trains only low-rank adapters on four 24 GB GPUs, so the adaptation is cheap enough to repeat. If correct, this is a practical template for moving off-the-shelf video models into specialized domains.","feed_headline":"Soccer-adapted video model lifts action accuracy from 11.8% to 63.5%","feed_subtitle":"A three-stage curriculum on 20,000 short event clips turns a general-purpose video language model into a soccer specialist.","key_machinery":"The load-bearing mechanism is the curriculum itself, applied through a uniform instruction-following format. Clips are cut to two seconds and sampled at eight frames so that each clip centers on one labeled event, which keeps the visual context focused and reduces hallucination when the teacher LLM writes captions. The first two training stages freeze the visual encoder and train low-rank adapters on the language model and the projection layer that connects vision features to text; the third stage trains only the language-side adapters. The progression teaches the model soccer concepts before question-answering behavior and task formatting, which the ablation shows is the order that extracts the largest gains.","core_discovery":"The paper's central claim is that a general video VLM can be specialized to a new domain by fine-tuning it in a fixed order on synthetic instruction data built from short, event-aligned clips. Stage one teaches concept alignment by captioning two-second clips from captions written from eight frames plus the ground-truth event label. Stage two turns those captions into five kinds of question-answer pairs and instruction-tunes the model on them. Stage three fine-tunes the model on a downstream task's output format, here 13-class soccer action classification. The full pipeline reaches 63.5% accuracy versus 11.8% for the base model on that task and a 37.5% relative improvement on VQA; ablations show that direct fine-tuning, or skipping either of the first two stages, gives materially worse results.","pith_inferences":["Beyond the paper: the hard-label action classification task is the most trustworthy evidence of real soccer understanding, because the caption and VQA scores are produced by the same teacher LLM that wrote the synthetic data, so those gains could partly reflect style imitation.","Beyond the paper: a sharper test of visual grounding would be to corrupt or swap the frames while keeping the event label and ask whether the adapted model's answers follow the label or the pixels.","Beyond the paper: the same event-aligned two-second window recipe is a natural fit for other event-dense video domains with timestamped labels, such as other sports, surveillance, or procedure videos."],"forward_implications":["If the recipe transfers, a general video VLM can be specialized to a dense-action domain with roughly 20k labeled event clips and four consumer-scale GPUs, instead of requiring large domain-specific pretraining.","The curriculum ordering is load-bearing: concept alignment before instruction tuning raises the relative VQA score from 79.1 to 85.8, and the full three-stage sequence raises action classification accuracy to 63.5% versus 16% for direct fine-tuning.","The model's temporal robustness survives the short training window: classifying 5-second clips reaches a macro F1 of 0.61 versus 0.63 on 2-second clips.","Including a second, broader event source in the 20k training set sharply reduces the cross-domain gap on VQA, indicating that data diversity matters as much as volume."],"supporting_citations":[{"why":"Supplies the base open-source video vision-language model that all stages fine-tune.","marker":"[41]"},{"why":"Supplies the broadcast soccer matches and timestamped action labels used to cut 2-second training and test clips.","marker":"[6]"},{"why":"Supplies the second, broader event dataset whose inclusion in the 20k training set removes most of the cross-domain test gap.","marker":"[16]"},{"why":"Supplies the low-rank adaptation method that keeps the three-stage fine-tuning computationally cheap.","marker":"[14]"},{"why":"Supplies the teacher model that writes synthetic captions and question-answer pairs from eight frames and the event label.","marker":"[1]"}],"fun_headline_variants":["Soccer video VLM jumps from 11.8% to 63.5% via curriculum tuning","Curriculum-trained VLM turns soccer videos into 63.5% accurate actions","General VLM becomes soccer expert after 20k clip curriculum","Soccer VLM: 63.5% accuracy from 11.8% via stepwise tuning","From 11.8% to 63.5%: soccer VLM via curriculum adaptation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The synthetic captions, question-answer pairs, and the VQA judge all come from the same commercial LLM; if that LLM writes mostly from the event label and its own priors rather than from the eight frames, the reported VQA gains could be prose mimicry rather than visual understanding.","fun_headline_variants_meta":{"raw":{"variants":["Soccer video VLM jumps from 11.8% to 63.5% via curriculum tuning","Curriculum-trained VLM turns soccer videos into 63.5% accurate actions","General VLM becomes soccer expert after 20k clip curriculum","Soccer VLM: 63.5% accuracy from 11.8% via stepwise tuning","From 11.8% to 63.5%: soccer VLM via curriculum adaptation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000785,"raw_usage":{"total_tokens":3437,"prompt_tokens":892,"completion_tokens":2545,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":2431}},"tokens_in":508,"tokens_out":2545,"duration_ms":15320,"temperature":1.0,"reasoning_tokens":2431,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:08:31.542645+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of test clips, replace their frames with frames from a different event class while keeping the original event label in the synthetic caption, and measure whether the adapted model's captions and answers follow the label or the frames; if they follow the label, the curriculum has taught label-to-text imitation rather than visual understanding.","supporting_citations":[{"cited_title":"Llava-next: A strong zero-shot video understanding model, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the base open-source video vision-language model that all stages fine-tune."},{"cited_title":"Soccernet- v2: A dataset and benchmarks for holistic under- standing of broadcast soccer videos","cited_arxiv_id":null,"evidence_quote":"Supplies the broadcast soccer matches and timestamped action labels used to cut 2-second training and test clips."},{"cited_title":"Wyscout: Football data and analytics plat- form, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the second, broader event dataset whose inclusion in the 20k training set removes most of the cross-domain test gap."},{"cited_title":"LoRA: Low-rank adaptation of large language models","cited_arxiv_id":null,"evidence_quote":"Supplies the low-rank adaptation method that keeps the three-stage fine-tuning computationally cheap."},{"cited_title":"Claude 3.5 sonnet: Our best model yet,","cited_arxiv_id":null,"evidence_quote":"Supplies the teacher model that writes synthetic captions and question-answer pairs from eight frames and the event label."}],"review_version":1}