{"id":"b63864d2-649a-4499-8a1c-9b6ea50e8175","arxiv_id":"2607.01751","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"MedStreamBench integrates 22 medical datasets into 5,419 QA instances across retrospective, present, future, and proactive temporal settings to evaluate streaming and proactive medical video understanding.","lead":"The paper introduces MedStreamBench, a benchmark for time-aware medical video understanding that tests models on when to answer or alert in addition to what to predict. This targets a key mismatch between standard offline benchmarks and real clinical streaming requirements.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Central claim of performance gap depends on unvalidated claim that four temporal settings match real clinical timing needs","rationale":"Reader's weakest_assumption directly identifies the load-bearing point for the strongest_claim. Full-text review does not alter this because the abstract already flags the absence of implementation details or controls; any later sections would need explicit validation (absent from the provided summary) to secure the claim. No other internal inconsistency is visible from the given material.","tokens_in":1704,"tokens_out":311,"duration_ms":13597,"concrete_test":"Locate the dataset-construction or evaluation-protocol section; extract the exact definition of evidence windows and proactive triggers. Re-run the streaming evaluation on one dataset (e.g., the largest) after replacing the paper's window boundaries with random offsets of equal average length; if the performance drop shrinks by >15% relative to the original numbers, the gap is sensitive to the unvalidated timing choices.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result (marked drop in streaming/proactive regimes) requires that the four settings (retrospective/present/future/proactive) and 22-dataset selection actually encode the timing and deferral constraints of clinical video streams. The paper states these settings but supplies no clinician review, deployment-log comparison, or sensitivity analysis showing that the chosen evidence windows and alert triggers avoid artificial constraints or selection bias. Without that grounding, the observed gap could be an artifact of the benchmark construction rather than a genuine clinical shortfall.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents MedStreamBench, a benchmark for time-aware medical video understanding that integrates 22 medical datasets into 5,419 QA instances. It defines four temporal settings—retrospective, present, future, and proactive—and evaluates models on single-turn and streaming modes, with additional metrics for responsiveness and post-evidence stability. Experiments on general-purpose and medical vision-language models demonstrate a substantial performance gap between offline recognition and temporally grounded decision-making in streaming and proactive settings.","tokens_in":1793,"tokens_out":338,"duration_ms":18278,"significance":"Should the benchmark's temporal settings and dataset choices prove representative of clinical video streams, the work would be significant for identifying critical shortcomings in current models' ability to handle timing, deferral, and proactive alerting in medical contexts. The public release of the dataset on Hugging Face supports reproducibility and community use.","major_comments":[{"comment":"The section describing the temporal settings and dataset integration states the four settings (retrospective, present, future, proactive) and the selection of 22 datasets but provides no clinician review, deployment-log comparison, or sensitivity analysis on evidence windows and alert triggers. This assumption is load-bearing for the central claim that the observed performance drops reflect genuine clinical shortfalls rather than benchmark-construction artifacts.","section":"Benchmark Design / Temporal Settings"}],"minor_comments":[{"comment":"The abstract reports 5,419 QA instances but does not break down their distribution across the four temporal settings or 22 source datasets.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on the benchmark design. We address the major comment below and outline planned revisions.","responses":[{"response":"We agree that direct clinician review and deployment-log comparisons would strengthen claims of clinical representativeness. The four temporal settings are derived from the native temporal structures and annotation protocols of the 22 source medical datasets (e.g., procedure phases in surgical videos, event timing in endoscopic and ultrasound streams), which themselves stem from clinical data collection. To address the concern about potential construction artifacts, the revised manuscript will include a new sensitivity analysis varying evidence-window lengths and alert-trigger thresholds across a range of clinically plausible values, demonstrating that the reported performance gaps between offline and streaming/proactive modes remain consistent. We will also add an explicit limitations paragraph discussing the absence of new clinician validation.","revision_made":"partial","referee_comment":"[Benchmark Design / Temporal Settings] The section describing the temporal settings and dataset integration states the four settings (retrospective, present, future, proactive) and the selection of 22 datasets but provides no clinician review, deployment-log comparison, or sensitivity analysis on evidence windows and alert triggers. This assumption is load-bearing for the central claim that the observed performance drops reflect genuine clinical shortfalls rather than benchmark-construction artifacts."}],"tokens_in":1288,"tokens_out":282,"duration_ms":19553,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point on this paper is that MedStreamBench pulls 22 medical datasets into one evaluation with retrospective, present, future, and proactive modes, then measures how models handle streaming inputs and when to raise alerts. It reports marked drops compared with standard offline testing.\n\nThe construction is straightforward and useful. It restricts evidence windows, supports streaming evaluation, and adds metrics for responsiveness plus stability after the window closes. This directly targets the gap between full-video lab tests and real clinical streams where timing and deferral decisions matter. The experiments on general and medical vision-language models make the gap concrete.\n\nThe soft spot is the lack of grounding for the settings themselves. The paper defines the four modes and the dataset mix but gives no clinician review, no comparison to actual procedure logs, and no sensitivity checks on window sizes or alert triggers. Without that, the observed drops could partly come from how the benchmark was assembled rather than from a universal clinical shortfall. Dataset statistics and instance quality checks are also missing from what is visible.\n\nThis is for groups building or testing medical video models who want evaluations that include timing. Readers focused on deployment gaps will find the protocol worth examining.\n\nSend it for peer review. The core idea is worth referee time even if the temporal choices need more external validation.","headline":"MedStreamBench adds four temporal settings and proactive alerts to medical video benchmarks and shows clear performance drops, but the clinical realism of those settings is not yet shown.","tokens_in":2270,"tokens_out":339,"would_cite":false,"duration_ms":21834,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"MedStreamBench shows leading vision-language models drop sharply in performance when medical videos require timed decisions rather than offline answers.","keywords":["medical video understanding","streaming evaluation","time-aware benchmark","proactive monitoring","vision-language models","temporal decision making","clinical video analysis"],"falsifier":"Finding no marked performance drop for models in the streaming or proactive settings relative to retrospective offline evaluation on MedStreamBench would indicate the claimed gap does not hold.","tokens_in":2603,"feed_emoji":"⏱️","tokens_out":646,"duration_ms":29521,"temperature":0.7,"pith_summary":"Existing medical video benchmarks check answer correctness but rarely test whether a model answers at the right moment. Clinical use demands deciding not only what to predict but also when to respond, defer, or raise an alert as video arrives. MedStreamBench combines 22 datasets and 5419 questions into four temporal settings that limit models to partial evidence windows and add streaming plus proactive alert tasks. It scores both correctness and timing aspects such as how quickly models respond and whether answers remain stable once more evidence appears. Experiments find clear performance declines in streaming and proactive conditions compared with standard full-video access.","feed_headline":"Medical video models drop when timing matters","feed_subtitle":"Benchmark across 22 datasets shows performance falls in streaming and proactive settings versus offline access.","key_machinery":"MedStreamBench benchmark, which enforces four temporal settings and bounded evidence windows to test when models answer or alert in medical video streams.","core_discovery":"The paper introduces MedStreamBench as a benchmark that integrates 22 medical datasets and 5419 QA instances across retrospective, present, future, and proactive temporal settings. It restricts models to temporally bounded evidence windows, supports single-turn and streaming evaluation, and adds a proactive monitoring task that requires models to decide whether and when to trigger alerts. Beyond answer correctness, the benchmark measures temporal behavior through responsiveness and post-evidence stability. Experiments on leading general-purpose and medical vision-language models reveal a substantial gap between offline recognition and temporally grounded decision-making, with performance dro","pith_inferences":["Similar time-bounded benchmarks applied to non-medical video tasks could expose parallel gaps in general video models.","The design may push training approaches that build explicit timing awareness into vision-language models.","Extending the proactive alert task to additional data types could probe broader real-world decision systems."],"forward_implications":["Clinical AI evaluation must include timing of predictions in addition to correctness to match deployment needs.","Restricting models to bounded evidence windows tests real-time decision making more closely than full-video access.","Proactive settings require separate assessment of when models should issue alerts without complete video evidence.","Metrics for responsiveness and post-evidence stability become necessary to judge suitability for streaming medical tasks."],"fun_headline_variants":["Timing restricts medical video model success","Streaming drops medical video AI performance","MedStreamBench tests when to alert in medical videos","Proactive settings challenge medical vision models","Time bounds limit medical video AI decisions"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The four temporal settings and the 22 chosen datasets accurately capture the timing and decision requirements of real clinical video streams.","fun_headline_variants_meta":{"raw":{"variants":["Timing restricts medical video model success","Streaming drops medical video AI performance","MedStreamBench tests when to alert in medical videos","Proactive settings challenge medical vision models","Time bounds limit medical video AI decisions"]},"model":"grok-4.3","cost_usd":0.006567,"raw_usage":{"total_tokens":2996,"prompt_tokens":685,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":65665500,"prompt_tokens_details":{"text_tokens":685,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2251,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":685,"tokens_out":60,"duration_ms":18460,"temperature":1.0,"reasoning_tokens":2251,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-03T16:34:56.241852+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Finding no marked performance drop for models in the streaming or proactive settings relative to retrospective offline evaluation on MedStreamBench would indicate the claimed gap does not hold.","supporting_citations":[],"review_version":1}