{"id":"cbaa06e6-abd9-48e3-8b3c-c811563f3e46","arxiv_id":"2608.02039","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A new benchmark and training method show that vision-language models perform poorly on remote sensing videos and that spatiotemporal evidence focusing can improve them by up to 9.01%.","lead":"The paper introduces RSVideo-10K, a remote sensing video benchmark with 10,773 question-answer pairs from drone and satellite footage, and finds that current vision-language models lag far behind on this domain. It also presents RSVideo, a training method that improves accuracy by up to 9.01 points by focusing on question-relevant small targets.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Benchmark labels are the load-bearing premise; without external release no independent check of answer uniqueness exists, so the 40.7% gap and 9.01% gain are unverified. An independent label audit before acceptance would settle it.","rationale":"The central claim has two components: (a) current VLMs fail badly on remote-sensing video, quantified by a 40.7% gap versus Video-MME; (b) RSVideo training improves accuracy by up to 9.01 points on 26 backbones. Both components are computed from accuracy on RSVideo-Bench. If the benchmark's gold labels are not uniquely determined by the visual input, every derived number is uninterpretable: a model could be marked wrong on an ambiguous item, and a training method could appear to gain simply by fitting annotation noise or option-position artifacts. The paper's strongest internal evidence for label validity is the three-expert review and adjudication protocol in Appendix B.3 and the inference-time validity audit in Appendix E.2. These are necessary but not sufficient. The validity audit shows that full video outperforms random single frames and shuffled frames, especially on AP/SER/SLR/SCR, which supports the claim that temporal evidence matters. It does not, however, test uniqueness of the gold answer or falsify all distractors; text-only already reaches 23.49%, 3.49 points above chance, and the audit does not report per-item diagnoses. Because Appendix G.6 states the dataset is not yet publicly distributed, no external researcher can perform even a spot-check of labels, evidence boundaries, or option plausibility. The transfer results in Table 4 and the sensitivity studies in Appendices E.3 through E.8 are valuable and partially mitigate concerns about method robustness: the method's relative gains over E-SFT are consistent across many backbones, and external benchmark gains show the training does not only memorize the benchmark. But those results do not repair the core benchmark validity premise; if labels are noisy, method rankings on RSVideo-Bench remain unreliable even if transfer looks healthy. The required check is therefore an external, blinded label audit on a random subset of locked test items before final acceptance. If the audit shows high agreement, the conditional verdict can be upgraded; if not, the headline numbers need re-estimation. This aligns with the reader's weakest-assumption analysis, so I agree with the reader and recommend no verdict change beyond the already conditional status.","tokens_in":41590,"tokens_out":9388,"duration_ms":95347,"concrete_test":"Release 200 randomly selected RSVideo-Bench items (video frames plus question and shuffled options) before acceptance. Have at least five independent qualified annotators, blinded to the authors' gold key, select the sole correct option under the Appendix A.1 insufficient-evidence rule, and compute Fleiss' kappa and the proportion of items with unanimous agreement. If unanimous agreement is below about 90%, or if any item is flagged ambiguous by two or more annotators, re-estimate accuracy and the 40.7% gap on the subset of unanimous items; if the gap and the RSVideo gain move by more than the claimed margin, the central claim is not supported until labels are corrected. If agreement is high, the key assumption is validated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing premise is that every RSVideo-Bench gold answer is uniquely determined by the released frames plus question. This is asserted via the internal three-expert pipeline (Appendix B.3), but the dataset is not yet distributed (Appendix G.6), so no external party can check labels or determine answerability. Appendix E.2's validity audit measures only whether model predictions depend on video input: text-only reaches 23.49% vs 20% chance, random single frame 33.17%, shuffled frames 35.13%, full video 36.59% for a Qwen3.6-27B checkpoint. These numbers show video contributes, but they do not establish that the gold answer is unique or that distractors are all invalid; a model could use video and still face ambiguous labels. If even a small fraction of the 2,731 test items are ambiguous or mislabeled, the headline 40.7% gap and the per-backbone gains (including the 9.01% maximum) are not well-defined, and comparisons across models become unreliable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces RSVideo-10K, a remote-sensing video question-answering dataset with 10,773 instances and a locked 2,731-instance test benchmark (RSVideo-Bench) spanning L1 perception and L2 reasoning across 17 tasks. It evaluates 26 open-source and several proprietary vision-language models, reporting a large performance gap on RSVideo-Bench relative to Video-MME (average 29.0% vs 69.7%). It then proposes RSVideo, a two-stage training framework that performs evidence-aware supervised fine-tuning followed by GRPO with evidence-aware rewards, using a fixed visual-token budget to select and compress spatiotemporal evidence. The paper reports consistent accuracy gains across all 26 backbones, with a maximum absolute improvement of 9.01% (InternVL3.5-14B) and a top accuracy of 40.63% (Qwen3.6-27B), plus transfer gains on four external benchmarks.","tokens_in":42036,"tokens_out":7174,"duration_ms":65076,"significance":"If the benchmark labels are reliable, the paper fills a genuine gap: current remote-sensing vision-language benchmarks are image-based or long-temporal, and general video benchmarks do not reproduce overhead viewpoints, small targets, and scene-constrained relations. The reported 40.7% gap is a striking and falsifiable finding, and the proposed evidence-focusing framework is well motivated and evaluated on a broad set of backbones. The paper is unusually thorough in its supporting material: detailed appendix, datasheets, validity audits, hyperparameter sensitivity analyses, and external-transfer evaluations. The main qualification is that the central premise — that every RSVideo-Bench gold answer is uniquely determined by the released frames and question — rests on an internal expert review process and cannot yet be checked externally because the dataset is not released and no inter-annotator statistics are provided.","major_comments":[{"comment":"The benchmark's validity rests on the internal three-expert review and the 'insufficient evidence' rule, but the dataset is not released (G.6: 'The dataset is not yet publicly distributed') and no inter-annotator agreement or external label audit is provided. The validity audit in Table 10 shows that video input contributes to performance, but it does not establish that every gold answer is uniquely determined by the released frames plus question, nor that all distractors are invalid. If a nontrivial fraction of the 2,731 test items are ambiguous or mislabeled, the headline 40.7% gap and the per-backbone gains (including the 9.01% maximum) are not well-defined. Please provide a label-uniqueness audit, a sample of item-level justifications, inter-annotator agreement, and a concrete release plan for the benchmark annotations and evaluation code.","section":"Appendix B.3 / G.6"},{"comment":"All numerical results are single-run point estimates without error bars. For a 2,731-item test set, the standard error of a 40% accuracy is about 0.9 percentage points, so differences below roughly 2 points are within one or two standard errors. The claims that RSVideo 'ranks first for every evaluated backbone' and that the reward components are complementary in §5.3 rely on small deltas (e.g., 38.85 vs 39.10 in Table 3, or 40.28 vs 40.63 in Table 18). Please report multiple seeds or confidence intervals, and clarify which configuration choices were made on the validation split.","section":"Tables 2-4 and Appendix E.3-E.8"},{"comment":"The evidence score fusion weights α_sal, α_rel, α_chg, and α_cell in Eqs. (2)-(3) are never specified, learned, or tuned. These weights determine which visual tokens are retained and compressed, so they are load-bearing for the method's reported behavior and reproducibility. Please state the default values, how they are set (fixed, searched, or learned), and their sensitivity, analogous to the reward-weight analysis in Appendix E.8.","section":"Equations (2)-(3) and Appendix D.1"}],"minor_comments":[{"comment":"The abstract says 'Codes will be available' but does not mention the dataset; please state explicitly that the benchmark annotations, evaluation scripts, and any redistributable clips will be released at the same point, given that Appendix G.6 currently says the dataset is not yet distributed.","section":"Abstract"},{"comment":"There is a typo in the sentence 'To construct this datset' — it should be 'dataset'.","section":"Section 3.2"},{"comment":"The average gap of 40.7% should specify which set of models is averaged and whether the average is computed over the identical model set on both Video-MME and RSVideo-Bench.","section":"Figure 1"},{"comment":"The header 'Video Coverage' uses the abbreviation 'UA V' in several rows; please expand to 'UAV' for readability, and check the formatting of 'UAVBench / UAVIT-1M'.","section":"Table 1"},{"comment":"The gate g_ans is defined as a product of two indicator functions; please clarify in the text that this is a scalar gate for the evidence rewards rather than a reward term itself.","section":"Equation (8)"},{"comment":"The notation uses N_l for the number of layers and H for the number of heads; consider renaming one of them to avoid confusion with the token-sequence length L.","section":"Appendix D.1"}],"recommendation":"major_revision","confidential_remarks":"This is a strong benchmark-plus-method paper with unusually extensive validation work. The main risk is the verifiability of the benchmark labels: the entire 40.7% gap and the per-backbone gains depend on answer uniqueness, which is currently asserted through an internal process and not externally checkable. The missing evidence-fusion weights and the lack of error bars are fixable, and I would not reject the paper on those grounds. I recommend major revision, with the expectation that the authors provide a label audit, a release plan, and the missing parameter details."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: RSVideo-10K is genuinely the first 10K-scale remote-sensing video QA benchmark covering both UAV and satellite footage, with a 17-task taxonomy, spatial/temporal evidence annotations, and an unusually careful annotation pipeline. The method, RSVideo, is a sensible combination of token pruning, evidence-aware SFT, and GRPO with evidence rewards; the 9.01-point gain over direct inference is credible as an internal result, but the real value is the benchmark.\n\nWhat the paper does well: the dataset construction is thoroughly documented. Appendix A.1 sets a clear rule that clips are retained only when the answer can be established from the released visual input, and the three-expert review with independent verification and adjudication (Appendix B.3) is a real quality process. The validity audit in Appendix E.2 is also a nice touch: text-only gets 23.49% vs 20% chance, and full video beats shuffled frames, showing the benchmark does measure video evidence and temporal order. The paper also runs sensitivity analyses on most hyperparameters (budget ratio, pooling ratio, temperature, group size, KL, reward weights) and reports transfer to four external benchmarks. That is more than most benchmark papers do.\n\nThe soft spots: first, the dataset and code are not released yet (Appendix G.6 explicitly says no public distribution). That makes the load-bearing premise—every gold answer is uniquely determined by the frames plus question—impossible to check independently. The internal audit is good evidence, but it is still internal. If even a few percent of test items are ambiguous or mislabeled, the headline 40.7% gap and the per-backbone gains shift. I would not call this a fatal flaw; most annotated benchmarks have this property, but for a benchmark whose main claim is to be a reliable evaluation tool, an independent spot-check of labels before or at release would substantially de-risk it. Second, all accuracy numbers are single-run point estimates; no error bars or seed variance are given. Given the gains are on the order of 1-9 points and the benchmark size is 2,731, variance is likely non-negligible. Third, the evidence-score fusion weights alpha_sal, alpha_rel, alpha_chg, alpha_cell are never given numerical values, so the method is not fully reproducible from the text.\n\nNone of these undermine the core idea. The internal evidence is consistent and the analysis is honest; the paper does not oversell the method's transfer (the external gains are small but positive). This paper is for anyone building or evaluating VLMs on remote-sensing video, and it deserves a serious referee. I would recommend conditional acceptance: release the data and code, add variance estimates, and either provide the alpha values or a sensitivity sweep. The stress-test note about label audit is on target; I would make that a formal revision requirement rather than a desk-reject reason.","headline":"A genuinely useful new remote-sensing video QA benchmark with a credible method, but the unverified labels and missing release make independent validation the gate.","tokens_in":42394,"tokens_out":3243,"would_cite":true,"duration_ms":29252,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"RSVideo-Bench, a new 2,731-question benchmark of UAV and satellite video, finds vision-language models average 29.0% accuracy (69.7% on Video-MME); a proposed training framework lifts this by up to 9.01 points across 26 backbones.","keywords":["remote sensing video","video question answering","vision-language model","benchmark","spatiotemporal evidence","reinforcement learning","UAV video","satellite video"],"falsifier":"Have an independent team, blind to the gold answers, re-answer a random sample of several hundred RSVideo-Bench items from the released frames and question text alone; if their answers diverge from the gold labels on a substantial fraction of items, the benchmark's premise of uniquely determined answers—and the accuracy numbers built on it—would be undermined.","tokens_in":41433,"feed_emoji":"🛰️","tokens_out":7094,"duration_ms":58287,"temperature":0.7,"pith_summary":"The paper claims that current vision-language models largely fail at understanding continuous remote-sensing video: on its new benchmark RSVideo-Bench, 26 open-weight models average 29.0% accuracy, compared with 69.7% on the natural-video benchmark Video-MME, a 40.7-point gap. The failures concentrate on small targets, short-lived state changes, and spatial relations defined by roads, buildings, and region boundaries. The paper also introduces a training framework, RSVideo, that first teaches models to attach answers to explicit spatiotemporal evidence and then uses reinforcement learning to select question-relevant visual tokens while compressing background. Trained this way, the same models improve by up to 9.01 absolute points on RSVideo-Bench, reaching 40.63% accuracy with a 27B backbone. If correct, the work provides both a diagnostic benchmark and a training recipe for making video-understanding models attend to sparse evidence in overhead imagery.","feed_headline":"New benchmark exposes 40-point AI gap on satellite video","feed_subtitle":"Evidence-focusing training lifts accuracy by up to 9 points across 26 vision-language models.","key_machinery":"The central mechanism is spatiotemporal evidence focusing: each visual token is scored by three signals—global saliency from visual self-attention, question relevance from cross-modal similarity, and temporal change from local inconsistency—plus a coarse time-region cell prior. The model keeps the highest-scoring tokens, compresses the rest into a small set of question-conditioned background slots via cross-attention, and is trained by GRPO with rewards that tie the predicted answer to annotated evidence regions and penalize redundant background. This lets the model operate under a fixed visual-token budget (default 40% of tokens) while preserving the few pixels that matter in remote-sensing frames.","core_discovery":"On the paper's own terms, the discovery is that a unified five-choice video-QA benchmark built from UAV and satellite footage reveals a systematic weakness in vision-language models: they miss sparse spatiotemporal evidence—small targets, brief actions, and scene-constrained spatial relations—that is decisive for remote-sensing video reasoning. The paper further claims that this weakness is trainable. Its RSVideo framework performs evidence-aware supervised fine-tuning on annotated temporal and spatial evidence, then applies GRPO reinforcement learning with rewards for answer correctness, evidence alignment, and background compression under a fixed visual-token budget. In evaluations across 26 open-weight backbones from 1B to 241B parameters, RSVideo improves accuracy on every backbone, with a maximum absolute gain of 9.01% (InternVL3.5-14B) and a top accuracy of 40.63% (Qwen3.6-27B); transfer evaluations on four external video benchmarks also show small but consistent gains.","pith_inferences":["A natural extension would be to test whether the evidence-scoring signals transfer to other sparse-evidence video domains, such as surveillance or medical video, where the decisive cues are similarly small and short-lived.","The paper's 'insufficient evidence' answer option is a useful design; it could be used to measure whether models can abstain when evidence is absent, which may be more important than accuracy for safety-critical remote-sensing decisions.","The fixed 40% token budget suggests that much of the video content is redundant for QA; one could push further to see how accuracy changes with even tighter budgets or with budgets adapted per question difficulty.","The human-review label determinism is the load-bearing premise; a machine-checkable audit of label uniqueness would make the benchmark's numbers portable."],"forward_implications":["If the benchmark numbers hold, current general video benchmarks overestimate how ready vision-language models are for overhead, small-target video; any claim of video understanding should be re-checked on remote-sensing video.","The 40.7-point gap between RSVideo-Bench and Video-MME gives a concrete target: closing it requires models to recover evidence that occupies few tokens in few frames.","The training framework transfers to existing open-weight backbones without changing their architecture, so the evidence-focusing recipe can be applied on top of stronger future base models.","The transfer gains on general and aerial video benchmarks suggest that evidence-focused training is not overfit to RSVideo-Bench, though the gains on general video are small.","The RSVideo-Instruct training set, with temporal and spatial evidence annotations, is itself a reusable resource for training other methods on remote-sensing video."],"supporting_citations":[{"why":"Supplies the natural-video benchmark whose average accuracy establishes the reported performance gap.","marker":"[9]"},{"why":"Supplies the group-relative policy optimization algorithm used in the second training stage of RSVideo.","marker":"[27]"},{"why":"Provides the policy-optimization baseline variant that the proposed method is compared against.","marker":"[8]"},{"why":"Provides an alternative policy-optimization baseline in the training-strategy comparison.","marker":"[49]"},{"why":"External video benchmark used to test transfer of the learned evidence-focusing policy.","marker":"[14]"},{"why":"External UAV/urban video benchmark used to test transfer of the learned policy.","marker":"[47]"},{"why":"External aerial video benchmark used to test transfer of the learned policy.","marker":"[51]"}],"fun_headline_variants":["Satellite video benchmark exposes VLM blind spots","RL fine-tuning boosts VLMs on satellite video by 9 points","Satellite video QA: VLMs struggle, RL lifts accuracy","Remote sensing video benchmark: models gain up to 9%","VLM blind spots exposed by satellite video benchmark"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole evaluation rests on the assumption that every RSVideo-Bench question has one answer that is uniquely determined by the released video frames and the question, a judgment enforced only by the paper's internal three-expert review and adjudication process.","fun_headline_variants_meta":{"raw":{"variants":["Satellite video benchmark exposes VLM blind spots","RL fine-tuning boosts VLMs on satellite video by 9 points","Satellite video QA: VLMs struggle, RL lifts accuracy","Remote sensing video benchmark: models gain up to 9%","VLM blind spots exposed by satellite video benchmark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000812,"raw_usage":{"total_tokens":3597,"prompt_tokens":1016,"completion_tokens":2581,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":632,"completion_tokens_details":{"reasoning_tokens":2500}},"tokens_in":632,"tokens_out":2581,"duration_ms":16275,"temperature":1.0,"reasoning_tokens":2500,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:08:47.826705+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have an independent team, blind to the gold answers, re-answer a random sample of several hundred RSVideo-Bench items from the released frames and question text alone; if their answers diverge from the gold labels on a substantial fraction of items, the benchmark's premise of uniquely determined answers—and the accuracy numbers built on it—would be undermined.","supporting_citations":[{"cited_title":"MVBench: A com- prehensive multi-modal video understanding benchmark","cited_arxiv_id":null,"evidence_quote":"External video benchmark used to test transfer of the learned evidence-focusing policy."},{"cited_title":"UrbanVideo-Bench: Benchmark- ing vision-language models on embodied intelligence with video data in urban spaces","cited_arxiv_id":null,"evidence_quote":"External UAV/urban video benchmark used to test transfer of the learned policy."}],"review_version":2}