{"id":"319a0f69-9895-4790-9348-e1367ae008f9","arxiv_id":"2607.02991","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Current MLLMs can deliver procedural instructions in streaming video but systematically fail at real-time error detection and corrective coaching on the new GuideMe benchmark.","lead":"GuideMe is a new multi-domain streaming-video benchmark that tests whether multimodal LLMs can coach people through multi-step tasks in real time—issuing steps, spotting mistakes, and correcting them. It shows current models can give instructions but largely fail at error detection and closed-loop correction.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Annotation pipeline and LLM-as-a-Judge may inflate the instruction vs. error-correction asymmetry by construction rather than pure model failure.","rationale":"The reader correctly isolates the annotation pipeline and evaluation circularity as the weakest assumption supporting the instruction–error asymmetry. The manuscript is otherwise careful: multi-domain packaging, streaming protocol, bipartite matching + behavioral metrics, and ablations on sampling/history (Tab. 4–5) are solid engineering contributions, and the empirical pattern is consistent across proprietary, open-source, and streaming models. No internal contradiction appears; the risk is external validity of the labels and judge. Because the paper already frames itself as a benchmark with public release, CONDITIONAL remains the right verdict—useful if the community treats Error/Correction labels and Scorem as provisional until human validation. My concern is essentially the same as the reader’s, only sharpened to the concrete mechanism (LLM-generated corrections + LLM judge) that could manufacture the headline asymmetry. No stronger load-bearing flaw (e.g., metric definition error or data leakage) is evident from the text.","tokens_in":17900,"tokens_out":671,"duration_ms":7653,"concrete_test":"On a stratified 100-video subset of GuideMe-Test, have 2–3 human raters (blind to model outputs) (i) rewrite or validate Error/Correction references for naturalness and visual fidelity and (ii) re-score matched model responses with the same 0–5 rubric used by the LLM judge. Recompute Tab. 3 Instruction vs. Error/Correction deltas under human references and human scores; if the asymmetry shrinks by >30% relative or loses statistical significance, the load-bearing claim is overstated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (abstract; Tab. 2–3; Fig. 4) is that MLLMs excel at next-step instructions but fail at error detection and corrective guidance under closed-loop streaming. That claim rests on treating the three-stage LLM pipeline (Sec. 3.2: activity extraction from EgoPER/CaptainCook4D/HoloAssist/QEVD labels → N=10 knowledge consensus → conversation generation at action start/end) and the subsequent LLM-as-a-Judge (Sec. 3.5) as faithful ground truth. Error/Correction utterances are synthesized by the same class of models being evaluated, often by mapping task-graph deviations to a “correct next step” rather than free-form human coaching; Scorem and PC quality are likewise LLM-judged. If generated Error/Correction references are systematically more specific, less visually grounded, or stylistically harder than Instruction references, the reported asymmetry (e.g., Gemini 3 Pro sF1 42.0\to33.8/34.3 and Score 33.0\to20.2/23.1 in Tab. 3) can be an annotation/metric artifact rather than a pure capability gap. The paper reports no human inter-annotator agreement or human preference study on the generated dialogues, so this remains the least secure condition for the strongest claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces GuideMe, a multi-domain streaming benchmark for closed-loop procedural coaching with MLLMs. It aggregates 2,458 videos (223.7 hours) from EgoPER, CaptainCook4D, HoloAssist, and QEVD into 47,775 interaction samples spanning next-step instructions, completion feedback, error detection, and corrective guidance. A three-stage LLM-assisted pipeline extracts atomic actions, generates task knowledge via N=10 consensus, and produces timestamped dialogues. Evaluation uses temporal-semantic bipartite matching (sPrecision/sRecall/sF1), behavioral classification (CS/FA/NR/PC and Score), and LLM-as-a-Judge (Scorem). Across proprietary, open-source, and streaming models, results show reasonable instruction delivery but sharp degradation on error detection and correction (Tab. 2–3, Fig. 4), with ablations on sampling, history, window size, and interval (Tab. 4–5).","tokens_in":18297,"tokens_out":1266,"duration_ms":10287,"significance":"If the reported instruction–error asymmetry is real rather than an annotation artifact, GuideMe fills a clear gap relative to offline procedural datasets and general streaming benchmarks (Tab. 1). The multi-domain scale, explicit closed-loop framing, public code/data release, and multi-component protocol (bipartite matching + timing categories + content judge) are concrete contributions that can drive work on proactive intervention and recovery. The fine-tuning result and protocol ablations further make the resource usable for adaptation studies. The main scientific value is therefore as a diagnostic testbed that isolates when-to-speak and how-to-correct failures that offline video understanding does not stress.","major_comments":[{"comment":"Sec. 3.2 and Tab. 3: The central claim that models “excel at instructions but fail at error detection/correction” rests on LLM-generated Error/Correction references (task-graph deviation → correct next step for CaptainCook4D/EgoPER; HoloAssist corrections; then LLM dialogue generation) and on LLM-as-a-Judge scores for PC/Scorem (Sec. 3.5). No human inter-annotator agreement, preference study, or stratified quality audit is reported for the four response categories. If Error/Correction references are systematically more specific, less visually grounded, or stylistically harder than Instruction references, the Tab. 3 drops (e.g., Gemini 3 Pro sF1 42.0→33.8/34.3; Score 33.0→20.2/23.1) can partly be annotation/metric artifacts. A modest human validation subset (or category-wise difficulty controls) is load-bearing for the strongest claim.","section":null},{"comment":"Sec. 3.4–3.5 and Tab. 2: Default evaluation supplies ground-truth dialogue history and uses Anchor-Based sampling (GT interventions + balanced silent timestamps). Tab. 4 shows that replacing GT history with model predictions sharply raises NR and lowers sF1/Score, and Dense sampling collapses behavior into silence or over-response. The paper correctly notes this, but the main tables still report the GT-history/Anchor setting as the primary evidence of “closed-loop” coaching failure. The manuscript should more clearly separate (i) single-step timing/content under oracle history from (ii) fully autonomous multi-turn closed-loop performance, and avoid overstating (i) as complete closed-loop coaching.","section":null},{"comment":"Sec. 3.5, Eqs. (1)–(3) and behavioral Score: Soft-F1 and Score mix continuous embedding similarity, Gaussian temporal cost (σ free), silence handling, and LLM-judge quality into aggregate numbers. Sensitivity of σ, embedding model, and judge model is not reported; only window size and sampling interval are ablated (Tab. 5). Because the headline asymmetry is quantitative, a short sensitivity check (or fixed public judge prompt + seed) is needed so that ranking differences in Tab. 2–3 are not driven by unstated metric hyperparameters.","section":null}],"minor_comments":[{"comment":"Fig. 1 and abstract claim “closed-loop interaction where corrective guidance modifies the user’s subsequent actions,” but videos are pre-recorded; the loop is simulated via annotated recovery segments rather than live user adaptation. Clarify this distinction early.","section":null},{"comment":"Table 1 lists “Closed-loop Interaction” only for GuideMe; a one-sentence definition of what counts as closed-loop (vs. timed feedback in QEVD/HoloAssist) would prevent over-reading the checkmark.","section":null},{"comment":"Sec. 3.3: Train/test split sizes are given, but domain balance and error-rate balance across splits are not; a short table would help reproducibility of Tab. 3.","section":null},{"comment":"Typographical/OCR artifacts appear in the provided text (e.g., “/enve♀e”, “/g♀behomepage”, “ofhow far”, “Wedesignaninferencepipeline”). Clean for camera-ready.","section":null},{"comment":"Related work (Sec. 2.2–2.3) is thorough; still, briefly position against recent proactive streaming QA benchmarks on the specific instruct–observe–correct cycle rather than only general streaming understanding.","section":null}],"recommendation":"major_revision","confidential_remarks":"The resource is timely and likely useful if the authors add even a small human validation of Error/Correction dialogues and tighten the closed-loop vs. oracle-history framing. Without that, the strongest abstract claim is under-supported relative to a top venue bar, but the benchmark construction itself is not a reject-level flaw. Fit is solid for a CV/ML systems venue that values datasets and diagnostic evaluation."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing: GuideMe is the first multi-domain packaging of existing procedural video corpora into a streaming closed-loop instruct–observe–correct protocol with silence decisions, and the measured asymmetry (models handle next-step instructions better than error detection/correction) shows up consistently across proprietary, open, and streaming MLLMs.\n\nWhat is new is not a new learning method but a usable testbed. They take EgoPER, CaptainCook4D, HoloAssist, and QEVD, extract atomic actions with correct/wrong/correction labels, generate task knowledge via N=10 consensus, and produce timestamped dialogues (next-step, completion, error, correction). 2,458 videos, 223.7 hours, 47k samples, train/test split, code and data promised. Table 1 is honest about what prior sets miss. The evaluation suite is thoughtful: temporal-semantic bipartite matching (soft F1), behavioral CS/FA/NR/PC, and LLM-as-a-Judge on matched pairs. Ablations on dense vs anchor sampling, GT vs self history, window size, and interval are done properly and show that protocol choices matter a lot—especially that self-history compounds errors. Fine-tuning Qwen3-VL-8B improves alignment but does not close the coaching gap. That is solid empirical work for a benchmark paper.\n\nThe soft spot the stress-test flags is real and should be stated plainly: Error/Correction references are largely LLM-synthesized from task graphs, and quality is LLM-judged, with no human IAA or preference study on the dialogues. So part of the asymmetry could be annotation hardness or style rather than pure model failure. That is a genuine limitation, not a minor footnote. It does not erase the multi-model pattern or the qualitative examples, but it means the strongest claim should be treated as provisional until the release is checked. Circularity is mild-to-moderate for an empirical CV benchmark, not load-bearing math fraud.\n\nWho it is for: people building streaming assistants, AR coaches, or proactive video MLLMs who need a multi-domain closed-loop yardstick. Citation pattern is appropriate; methods are standard and reproducible enough once data ships.\n\nI would send this to peer review. It deserves referee time with a clear ask for human validation of a dialogue sample and tighter discussion of judge/annotation bias. Engage with the release; do not treat the asymmetry numbers as gospel until then.","headline":"Useful multi-domain closed-loop streaming coaching benchmark with a real instruction-vs-correction gap; annotation/judge circularity is a real but not fatal soft spot.","tokens_in":18948,"tokens_out":593,"would_cite":true,"duration_ms":5350,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Current multimodal models can give step-by-step instructions for live tasks, but they consistently fail to spot user mistakes and issue corrective feedback in closed-loop streaming coaching.","keywords":["MLLMs","streaming video","interactive task guidance","procedural coaching","error detection","corrective feedback","closed-loop interaction","benchmark"],"falsifier":"Have independent human coaches re-annotate a large held-out subset of GuideMe error and correction events; if top models already match human timing and corrective content on those human labels, or if models trained only on clean next-step data succeed in live user sessions, the claimed instruct-versus-correct asymmetry would fail.","tokens_in":18770,"feed_emoji":"🎥","tokens_out":941,"duration_ms":19024,"temperature":0.7,"pith_summary":"This paper asks how far multimodal large language models are from acting as real-time procedural coaches: systems that watch a user perform a multi-step task on a video stream, decide when to speak, detect mistakes, and steer the user back on track. To make that question measurable, the authors build GuideMe, a multi-domain streaming benchmark with 2,458 videos (223.7 hours) across cooking, object manipulation, daily-life guidance, and fitness, yielding 47,775 interaction samples that cover next-step instructions, completion feedback, error detection, and corrective guidance. They evaluate models with a three-part framework that jointly scores sequence-level temporal-semantic alignment, whether the model intervenes or stays silent at the right moments, and the quality of what it says. Across proprietary, open-source, and streaming-specialized models, the central finding is a sharp performance asymmetry: models are relatively capable at routine instruction, yet they systematically fail at identifying execution errors and providing corrective feedback. A sympathetic reader cares because success on offline video understanding does not transfer to the closed instruct-observe-correct loop that real assistance requires, and GuideMe turns that gap into a concrete training and evaluation target.","feed_headline":"AI coaches give good steps but miss live mistakes","feed_subtitle":"A 224-hour multi-domain streaming benchmark shows models fail closed-loop error correction.","key_machinery":"GuideMe, a multi-domain streaming interaction benchmark, together with its three-component evaluation: temporal-semantic bipartite matching for sequence-level alignment of timed responses, behavioral classification of speak-versus-silent decisions at intervention anchors, and LLM-as-a-Judge scoring of content quality. The benchmark is produced by a three-stage pipeline that extracts correct, wrong, and correction actions, generates procedural knowledge, and turns those into timestamped dialogues.","core_discovery":"Despite excelling at providing instructions, existing multimodal large language models consistently fail to identify execution errors and respond with corrective feedback when tested as real-time procedural coaches on streaming video. GuideMe establishes this asymmetry across diverse domains and model families: models can describe what should happen next, but they do not yet coach a user based on what is actually happening.","pith_inferences":["Training that only rewards clean next-step narration may under-prepare models for mistake-sensitive intervention.","The same instruct-observe-correct deficit is likely to appear in robot coaching and AR assistance that reuse the same multimodal backbones.","Synthetic error-and-recovery trajectories may be a necessary data primitive beyond expert demonstration videos.","Metrics that overweight correct silence under dense sampling can hide coaching failure; balanced intervention anchors matter for fair comparison."],"forward_implications":["Reliable AI procedural coaches must jointly solve when to intervene, error detection, and corrective content—not only offline video description.","Model scale and streaming-oriented pretraining alone do not close the closed-loop gap on this benchmark.","Fine-tuning on GuideMe can improve temporal alignment and silence calibration while still leaving error correction weak.","Interactive-assistant evaluation must separately score silence-versus-speak decisions and corrective behavior, not only generic video QA."],"fun_headline_variants":["MLLMs give solid next steps but miss live errors in streaming coaching","Models instruct well yet fail to detect mistakes in real-time video tasks","GuideMe: AI coaches excel at steps, falter on error correction in video","Streaming MLLMs provide guidance but consistently miss execution faults","Closed-loop coaches nail instructions yet ignore live task mistakes"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The automated annotation pipeline must produce trustworthy ground-truth intervention times and dialogues that truly represent closed-loop coaching, so that weak error-correction scores reflect model limits rather than label or metric artifacts.","fun_headline_variants_meta":{"raw":{"variants":["MLLMs give solid next steps but miss live errors in streaming coaching","Models instruct well yet fail to detect mistakes in real-time video tasks","GuideMe: AI coaches excel at steps, falter on error correction in video","Streaming MLLMs provide guidance but consistently miss execution faults","Closed-loop coaches nail instructions yet ignore live task mistakes"]},"model":"grok-4.5","effort":"low","cost_usd":0.004686,"raw_usage":{"total_tokens":1363,"prompt_tokens":777,"num_sources_used":0,"completion_tokens":94,"cost_in_usd_ticks":46860000,"prompt_tokens_details":{"text_tokens":777,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":492,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":777,"tokens_out":94,"duration_ms":4311,"temperature":1.0,"reasoning_tokens":492,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T05:35:28.146196+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Have independent human coaches re-annotate a large held-out subset of GuideMe error and correction events; if top models already match human timing and corrective content on those human labels, or if models trained only on clean next-step data succeed in live user sessions, the claimed instruct-versus-correct asymmetry would fail.","supporting_citations":[],"review_version":1}