{"id":"6417983f-cf6f-4197-9bec-0eb42cc965c5","arxiv_id":"2606.03788","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"SLU-2K is a new closed-ended QA benchmark that measures semantic understanding in sign language video, showing MLLMs near random and fine-tuned SOTA systems at 56.7-75.2% accuracy.","lead":"The paper creates SLU-2K, a set of 2350 video-based question-answer pairs drawn from existing sign language datasets to test whether translation systems capture meaning such as actions, objects, and facts. A smart generalist might read it because current word-overlap metrics appear to overstate how well sign language AI actually understands content for assistive applications.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Validity of reported semantic gaps rests on unverified fidelity of the automated QA generation pipeline to video content.","rationale":"The reader's weakest_assumption directly identifies the pipeline as the load-bearing point; the abstract-only limitation noted in the reader's rationale is now addressable via the released code and benchmark files, but the concrete validation step above is still required to confirm the claim. No other internal inconsistency appears in the stated argument.","tokens_in":1844,"tokens_out":348,"duration_ms":16271,"concrete_test":"Sample 100 random QA pairs from the released SLU-2K files; have two independent sign-language-fluent annotators (blind to model outputs) watch the corresponding videos and score whether each question is (a) answerable from the video alone and (b) has the claimed ground-truth answer; compute agreement and failure rate. If >15% of pairs fail either criterion, the performance gaps cannot be interpreted as semantic deficiencies.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (SOTA SLT systems achieve only 56.7-75.2% on semantic questions; MLLMs near random) requires that the 2,350 closed-ended questions in SLU-2K accurately probe meaning present in the sign videos across the seven categories. The paper states it \"propose[s] and extensively evaluate[s] an automated data generation pipeline,\" but if generation starts from text/gloss rather than video and lacks rigorous human validation for answerability and lack of bias, incorrect or unanswerable questions would artifactually lower scores. This is the least secure link because the headline numbers are produced by feeding these questions to models.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces SLU-2K, a benchmark of 2,350 closed-ended video question-answer pairs derived from PHOENIX-2014T and CSL-Daily via an automated pipeline across seven semantic categories (actions, locations, numbers, objects, people, time, weather). It evaluates MLLMs and two SOTA SLT systems (MMSTL, SpaMo), reporting near-random MLLM performance and SOTA accuracies of 56.7%–75.2%, and argues that surface metrics like BLEU overestimate semantic understanding in SLT. Code, prompts, and benchmark files are released publicly.","tokens_in":1990,"tokens_out":469,"duration_ms":35021,"significance":"If the automated pipeline produces questions that faithfully and without bias capture semantic content directly from the sign videos, SLU-2K would offer a useful complement to n-gram metrics by directly testing meaning preservation, a key requirement for assistive SLT applications. The public release of code and data is a clear strength supporting reproducibility.","major_comments":[{"comment":"Abstract and automated data generation pipeline section: the claim that the pipeline was 'extensively evaluated' is not supported by any reported quantitative validation (e.g., inter-annotator agreement, fraction of questions verified answerable from video alone, or error analysis for artifacts); this directly affects the reliability of the headline accuracy figures (56.7%–75.2%).","section":"automated data generation pipeline"},{"comment":"Abstract and evaluation section: it is not stated whether question generation begins from video content, glosses, or text translations; if the latter, the benchmark cannot be guaranteed to test video-based semantic recovery, which is the central premise for evaluating SLT systems on semantic gap.","section":"evaluation section"}],"minor_comments":[{"comment":"Abstract: the range 56.7% to 75.2% should be broken down by system (MMSTL vs. SpaMo) for precise interpretation of results.","section":"Abstract"},{"comment":"References: ensure full citations are provided for MMSTL and SpaMo, which are described as representative SOTA systems.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their detailed and constructive feedback. We address each major comment below and will revise the manuscript to improve the description of the data generation pipeline and its validation.","responses":[{"response":"We agree that the current manuscript does not provide quantitative validation metrics for the pipeline, such as inter-annotator agreement or explicit error rates on answerability from video. The description relies on qualitative checks and manual inspection of a subset of outputs. In the revised version, we will add a new subsection in the data generation section reporting quantitative results, including the fraction of questions verified as answerable from video alone and a systematic error analysis of artifacts.","revision_made":"yes","referee_comment":"Abstract and automated data generation pipeline section: the claim that the pipeline was 'extensively evaluated' is not supported by any reported quantitative validation (e.g., inter-annotator agreement, fraction of questions verified answerable from video alone, or error analysis for artifacts); this directly affects the reliability of the headline accuracy figures (56.7%–75.2%)."},{"response":"The pipeline generates questions from the text translations in the source datasets (PHOENIX-2014T and CSL-Daily), which are time-aligned with the videos and glosses. Questions are designed to target semantic content that is visually present in the sign videos. We acknowledge that the starting point of generation is not explicitly stated in the current text. In the revision, we will add a clear description of the pipeline steps, including examples, and explain how the resulting questions still evaluate video-based semantic recovery when applied to MLLMs and SLT systems.","revision_made":"yes","referee_comment":"Abstract and evaluation section: it is not stated whether question generation begins from video content, glosses, or text translations; if the latter, the benchmark cannot be guaranteed to test video-based semantic recovery, which is the central premise for evaluating SLT systems on semantic gap."}],"tokens_in":1533,"tokens_out":432,"duration_ms":22859,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main contribution is SLU-2K, a set of 2350 closed-ended video questions drawn from PHOENIX-2014T and CSL-Daily and split into seven categories like actions, objects, and time. They run this on MLLMs (near random) and two fine-tuned SLT systems (56.7-75.2 percent), arguing that standard BLEU-style metrics miss real understanding needed for assistive tech.\n\nWhat works is the shift in focus. Surface metrics have known limits in this domain, and a reusable set of semantic probes is a direct way to measure something more relevant. The category breakdown is straightforward and the release of code and prompts lowers the barrier for others to check or extend it.\n\nThe soft spot is the automated pipeline itself. The abstract claims extensive evaluation, yet the headline numbers depend on the questions being faithful to the videos without artifacts or unanswerable items. If generation leans on gloss or text rather than direct video inspection, and if human validation for answerability and bias is light, the reported gaps could be inflated. The stress-test concern lands here because the central claim rests on that fidelity.\n\nThis is useful for researchers already working on sign language translation or multimodal models who want a semantic check beyond n-gram overlap. It is not reshaping the broader field. The work shows clear thinking on the evaluation problem and honest engagement with the limits of current metrics.\n\nI would send it to peer review. The benchmark idea is worth referee time even if the pipeline details need tightening.","headline":"SLU-2K gives a concrete QA benchmark for semantic gaps in sign language translation, but its value hinges on how well the automated pipeline matches actual video content.","tokens_in":2517,"tokens_out":392,"would_cite":false,"duration_ms":14926,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Sign language translation systems still miss key semantic details even after fine-tuning, scoring 56.7 to 75.2 percent on a new question benchmark.","keywords":["sign language translation","semantic evaluation","question answering benchmark","multimodal large language models","PHOENIX-2014T","CSL-Daily","video understanding"],"falsifier":"A study in which independent human raters judge a random sample of the generated questions as failing to capture the actual semantic content of their source videos would falsify the benchmark's validity.","tokens_in":2754,"feed_emoji":"📹","tokens_out":686,"duration_ms":20938,"temperature":0.7,"pith_summary":"The paper shifts evaluation of sign language translation from surface metrics like BLEU to direct tests of semantic understanding. It creates SLU-2K, a set of 2,350 video question-answer pairs drawn from two existing datasets and generated across seven categories of meaning. Testing shows multimodal large language models perform near random while two representative state-of-the-art translation systems reach only 56.7 to 75.2 percent accuracy. These results indicate that current protocols based on lexical overlap overestimate how well translations preserve the original meaning. The work therefore argues that future progress must be judged by semantic correctness as well as fluency.","feed_headline":"SLT systems score 57-75% on semantic questions despite fine-tuning","feed_subtitle":"New benchmark of 2,350 video questions shows surface metrics like BLEU miss whether meaning is preserved.","key_machinery":"SLU-2K, a dataset of 2,350 closed-ended video question-answer pairs produced by an automated pipeline across the seven categories of actions, locations, numbers, objects, people, time, and weather conditions.","core_discovery":"The central claim is that state-of-the-art sign language translation systems fine-tuned on in-domain data still exhibit a substantial semantic gap on SLU-2K, reaching only 56.7 to 75.2 percent accuracy, while multimodal large language models reach near-random performance; this demonstrates that surface-form metrics overestimate true semantic understanding and that future SLT evaluation should incorporate semantic correctness.","pith_inferences":["The same question-based approach could be extended to additional sign language corpora beyond the two source datasets used here.","Models may be exploiting dataset-specific patterns rather than learning general semantic mappings, which would explain the gap between translation fluency and question-answering accuracy.","Pairing SLU-2K scores with traditional metrics on the same outputs would give a more complete picture of model behavior."],"forward_implications":["Current SLT evaluation protocols overestimate true understanding because they rely only on fluency and n-gram overlap.","Future progress in sign language translation should be measured by semantic correctness in addition to existing surface metrics.","Systematic integration of semantic understanding evaluation is required in current AI systems for sign language.","Automated pipelines can generate large-scale question sets for semantic evaluation from existing translation datasets."],"fun_headline_variants":["SLT hits 57-75% semantic accuracy on SLU-2K","Semantic gap persists in fine-tuned sign language systems","SLU benchmark exposes limits of surface-form SLT metrics","MLLMs near-random on video question semantic test","SLT systems lag semantically at 57-75% despite tuning"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The automated question-generation pipeline produces questions that accurately and without bias reflect the semantic content of the original sign videos across the seven categories.","fun_headline_variants_meta":{"raw":{"variants":["SLT hits 57-75% semantic accuracy on SLU-2K","Semantic gap persists in fine-tuned sign language systems","SLU benchmark exposes limits of surface-form SLT metrics","MLLMs near-random on video question semantic test","SLT systems lag semantically at 57-75% despite tuning"]},"model":"grok-4.3","cost_usd":0.004722,"raw_usage":{"total_tokens":2399,"prompt_tokens":805,"num_sources_used":0,"completion_tokens":81,"cost_in_usd_ticks":47224500,"prompt_tokens_details":{"text_tokens":805,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1513,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":805,"tokens_out":81,"duration_ms":11645,"temperature":1.0,"reasoning_tokens":1513,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T10:48:25.418753+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A study in which independent human raters judge a random sample of the generated questions as failing to capture the actual semantic content of their source videos would falsify the benchmark's validity.","supporting_citations":[],"review_version":1}