{"id":"bd3f4226-8899-4f5d-aa4d-2e6a5d3c69f0","arxiv_id":"2508.16076","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new 60-word Bengali sign language instruction dataset and a sign-parameter-infused prompting method that modestly improves VLM-generated instructions on most text-matching metrics for larger models.","lead":"This paper introduces BdSLIG, a small new dataset of written instructions for 60 Bengali sign language words, and tests a prompting method that feeds sign parameters such as hand shape and motion into vision-language models. The authors report that this structured prompting usually improves agreement with human-written instructions, but the gains are inconsistent across models and metrics.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SPI's metric gains may be an artifact of shared annotation/prompt schema; no human evaluation supports the claim of better instruction quality.","rationale":"The reader's weakest assumption accurately identifies the most load-bearing concern: the shared parameter schema between the BdSLIG annotation guidelines and the SPI prompt creates a circularity that can inflate all reported metrics without improving genuine instruction quality. This is not merely a disagreement with consensus; it is an internal inconsistency between the paper's own admission that text metrics are unsuitable (Section 3.4) and the reliance on those same metrics to support the headline claim (Section 3.1). The proposed concrete test—a double-blind human expert evaluation—directly tests the claim's validity. If human ratings show no advantage for SPI, the paper's central contribution is reduced to a prompt-formatting trick rather than a substantive improvement in sign language instruction generation. Given the pilot nature of the work and the authors' transparency about limitations, conditional acceptance remains appropriate, but the paper should require such a human evaluation or at least clearly temper the claim until one is performed. The reader's verdict of CONDITIONAL with moderate confidence is well calibrated; my stress test does not change that verdict, but it reinforces the necessity of the condition.","tokens_in":6961,"tokens_out":2211,"duration_ms":24969,"concrete_test":"Run a double-blind human expert evaluation: collect a random sample of ~30–60 signs from BdSLIG, generate instructions with both Vanilla and SPI prompting across the four VLMs, shuffle them, and have a Bengali Sign Language expert (or expert learner) rate each instruction for correctness, completeness, and executability without knowing the prompting condition. If SPI instructions are not rated significantly higher than Vanilla instructions (e.g., via paired t-test or Wilcoxon on per-sign scores), then the metric gains in Table 1 are unsupported as evidence for improved instruction quality and the central claim fails. Alternatively, or additionally, run a 'schema-only' control where the SPI prompt is replaced by a prompt that merely lists the seven parameter categories with generic definitions (no visual input); if this still produces high metric overlap with the reference, it confirms th","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that SPI prompting produces more structured and semantically faithful sign instructions—rests entirely on text-matching metrics in Table 1. But the BdSLIG ground-truth annotations were written using the exact seven sign-parameter categories (handshape, movement, location, palm orientation, spatial interaction, temporal dynamics, facial cues) that SPI prompting feeds into the model (Section 2.1 states these parameters serve both as annotation guidelines and as the basis for SPI prompting). By explicitly listing these categories and their keyword tags in the prompt, SPI gives the VLM the schema and vocabulary of the reference text, while Vanilla prompting does not. This inflates lexical and semantic overlap (ROUGE, BLEU, METEOR, BERTScore) regardless of whether the model actually perceives the sign correctly. The paper itself concedes in Section 3.4 that these metrics are 'not suitable for evaluating sign language instructions' because surface variation is penalized even when two instructions describe the same motion. Yet the strongest empirical claim is derived solely from these unsuitable metrics, and no human or expert evaluation is performed. The absence of a schema-information control means the observed gains could be purely a form of test-time annotation leakage—recovering the annotation template rather than demonstrating better understanding of Bengali Sign Language.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces BdSLIG, a Bengali sign-language instruction-generation dataset derived from one representative video for each of 60 words in BdSLW60, with textual instructions produced by annotators using seven sign-parameter categories (handshape, movement, location, palm orientation, spatial interaction, temporal dynamics, facial cues). It then proposes Sign Parameter-Infused (SPI) prompting, which injects these same parameter categories and keyword tags into the VLM prompt, and compares SPI with vanilla prompting on four proprietary VLMs using ROUGE-1/2/L, BLEU, METEOR, and BERTScore. The paper claims that SPI generally or consistently improves instruction generation, and it discusses limitations of automatic metrics and the value of BdSLIG as a long-tail visual benchmark.","tokens_in":7255,"tokens_out":5608,"duration_ms":65712,"significance":"If the empirical claims were fully supported, the paper would make a modest but useful contribution: it provides the first Bengali SLIG dataset, makes data and code publicly available, and explores a prompt-design strategy grounded in descriptive parameters common in sign-language research. The dataset could serve as a pilot resource for low-resource SL instruction and for evaluating VLMs on long-tail visual concepts. However, the headline result is currently weakened by an evaluation design in which the reference annotations and the SPI prompt share the same parameter schema, and by the absence of any human expert evaluation—an absence the paper itself motivates in §2.4.2 and §3.4. The significance is therefore conditional on the authors' ability to validate SPI against a non-circular measure of instruction quality.","major_comments":[{"comment":"The claim that 'SPI prompting consistently outperforms Vanilla Prompting across most metrics' is overstated. Across 24 model-metric cells in Table 1, SPI improves in 15, worsens in 8, and ties in 1. In particular, GPT-4.1-mini regresses on ROUGE-1, ROUGE-2, ROUGE-L, and BLEU, ties on BERTScore, and improves only on METEOR; Gemini 2.5 Pro regresses on four of six metrics. The abstract's 'generally leads to better performance' is more accurate. The wording should be revised, and the paper should report uncertainty estimates (e.g., confidence intervals or paired tests) before claiming a systematic advantage.","section":"§3.1, Table 1"},{"comment":"The evaluation is circular in a way that is load-bearing for the central claim. Section 2.1 states that the seven sign parameters serve both as annotation guidelines for BdSLIG and as the basis of SPI prompting. Thus the reference annotations were written using the same categories and vocabulary that SPI injects into the prompt, while the vanilla prompt does not provide that schema. Higher text-matching metrics under SPI may therefore reflect vocabulary/schema alignment rather than better sign understanding. The paper itself concedes in §3.4 that lexical and semantic similarity metrics are 'not suitable for evaluating sign language instructions.' A control condition is needed—for example, a prompt that includes the category labels and keyword tags without any additional visual guidance—or, preferably, human expert evaluation of instruction correctness and executability. The current desig","section":"§2.1, §2.3.2, §3.4"},{"comment":"BdSLIG contains only 60 items (one video per word), and Table 1 reports aggregate scores without any measure of variability or statistical significance. Many of the differences are numerically small (e.g., ROUGE-1 0.526 vs. 0.522 for GPT-4.1-mini, BERTScore 0.396 vs. 0.387 for Gemini 2.5 Pro), so the observed pattern could be within noise. Because the SPI-over-vanilla claim is the central empirical contribution, the paper should provide per-word breakdowns, paired significance tests, or bootstrap confidence intervals, and should explicitly state that the dataset is a pilot with 60 examples.","section":"§2.2, Table 1"}],"minor_comments":[{"comment":"The arrows (↑, ↓, =) are not explained in the caption. Please define what 'improvement' means and note the number of samples.","section":"Table 1 caption"},{"comment":"The abstract says SPI 'generally leads to better performance' while §3.1 says 'consistently outperforms.' Please make the wording consistent with the actual pattern in Table 1.","section":"Abstract vs. §3.1"},{"comment":"The full SPI prompt template is not shown. For reproducibility, include the exact prompt text, including how the parameter categories are formatted.","section":"§2.3.2"},{"comment":"The symbol P is used both for the prompt and for the probability distribution. Use a separate notation, e.g., p_θ, for the model's probability.","section":"Eq. (1), Eq. (2)"},{"comment":"The two sample instructions for 'toothpaste' illustrate the metric problem well, but the paper could make the point more concrete by reporting each metric's score for that specific pair.","section":"§3.4"},{"comment":"The long-tail claim assumes that Bengali Sign Language and associated instruction text are absent from VLM pretraining data, but no contamination check is provided. Even a brief discussion of this limitation would help.","section":"§3.5"}],"recommendation":"major_revision","confidential_remarks":"The paper's own §3.4 disqualifies the text-matching metrics that support the main SPI claim, and the shared annotation/prompt schema in §2.1 creates a plausible leakage mechanism. A revision that adds either a human expert evaluation or a schema-control condition would substantially strengthen the contribution. The dataset itself is small (60 examples) but may be acceptable as a pilot; the authors should be explicit about this scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the dataset: BdSLIG is the first Bengali sign language instruction generation dataset, and it is released on HuggingFace with code on GitHub. That is a real contribution for an under-resourced language, even at 60 words and one video per word. The SPI prompting idea is also reasonable—injecting the standard seven sign parameter categories (handshape, movement, location, etc.) into the prompt is a plausible way to structure zero-shot VLM output, and the paper ties it to a well-established linguistic taxonomy.\n\nThe paper is also honest in places. It openly says the text metrics are \"not suitable for evaluating sign language instructions\" (Section 3.4), acknowledges that uniform frame sampling is crude, and validates annotations with a domain expert. Those are good instincts.\n\nThe soft spots, though, are load-bearing. The central claim—\"SPI prompting consistently outperforms Vanilla\"—is false against the paper's own Table 1. GPT-4.1-mini regresses on four of six metrics and ties on a fifth; Gemini 2.5 Pro regresses on four. Only GPT-4.1 and Gemini 2.0 Flash show clear gains. So the abstract and Section 3.1 overstate the results.\n\nMore importantly, the evaluation is circular in a way the authors do not acknowledge. The same seven parameter categories were used as annotation guidelines for the ground-truth instructions and as the content injected by SPI prompting. That means SPI gives the model the schema and vocabulary of the reference text while vanilla prompting does not. The gains on ROUGE, BLEU, METEOR, and BERTScore may largely reflect template alignment rather than better understanding of the sign. The paper concedes the metrics are unsuitable for sign instruction quality, yet its strongest empirical claim rests entirely on those metrics. There is no human evaluation, no error bars, and no control that injects a different taxonomy.\n\nThese are serious issues, but they are fixable. A pilot with tempered claims, a human expert evaluation, and a schema control would be a solid paper. As is, it is a useful resource packaged with an overstated result.\n\nFor a reviewer: this deserves peer review, not desk rejection, because the dataset itself has lasting value and the prompting idea is worth exploring. But I would send it back for major revision, not accept it in this form.","headline":"BdSLIG is a genuinely new resource—the first Bengali SLIG dataset—and SPI prompting is a sensible new application of parameter-structured prompting, but the headline claim of consistent gains is not supported by the paper's own table.","tokens_in":7742,"tokens_out":1621,"would_cite":true,"duration_ms":17554,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Structured prompts built from standard sign parameters improve vision-language models' Bengali sign instruction generation.","keywords":["Bengali Sign Language","sign language instruction generation","vision-language models","sign parameter prompting","low-resource language","BdSLIG dataset","zero-shot generation"],"falsifier":"Have sign-language educators rate blinded instructions for executability and correctness without knowing which prompting strategy produced them; if SPI-generated instructions are not rated better than vanilla ones despite higher METEOR/BERTScore, the claim of semantic faithfulness fails. Alternatively, re-annotate a held-out set with free-form instructions not constrained by the seven categories; if SPI's text-metric advantage collapses, it was driven by schema alignment.","tokens_in":6885,"feed_emoji":"🤟","tokens_out":6646,"duration_ms":66365,"temperature":0.7,"pith_summary":"This paper proposes that the quality of automatically generated sign-language instructions can be improved by structuring the prompt around the seven canonical parameters linguists use to describe signs: handshape, movement type, location, palm orientation, spatial interaction, temporal dynamics, and facial cues. It introduces BdSLIG, the first Bengali Sign Language instruction-generation dataset, built by annotating videos from an existing word-level recognition dataset with step-by-step textual instructions. On this dataset, the proposed Sign Parameter-Infused (SPI) prompting outperforms vanilla prompting on most text-similarity metrics, especially METEOR and BERTScore, and the gains are largest for the larger closed models. The paper itself cautions that text-matching metrics are not suitable for judging sign instructions and that uniform frame sampling may drop key frames, so the concrete product is a structured prompt recipe plus a resource for benchmarking an under-resourced language. If the claim holds, structured parameter prompts offer a cheap, training-free way to make VLMs produce more reproducible, learnable instructions.","feed_headline":"Sign-parameter prompts beat plain prompts for Bengali sign AI","feed_subtitle":"Structured prompts beat free-form text on the first Bengali sign-instruction benchmark.","key_machinery":"The mechanism is Sign Parameter-Infused (SPI) prompting: a function f(P, S) that augments a base text prompt with the seven sign parameter categories and their expected values—handshape, movement type, location, palm orientation, spatial interaction, temporal dynamics, and facial cues—each with tag examples such as 'extended' or 'circular'. This forces the VLM to organize its answer as a stepwise narrative along those axes, and the same parameter vocabulary was used by the human annotators of the ground truth. The second load-bearing component is BdSLIG, a 60-word paired video-instruction dataset derived from BdSLW60 videos, which provides the reference annotations and the test bed.","core_discovery":"The central claim is that a vision-language model, given a short sign video and a prompt that lists the seven standard sign parameters, can produce step-by-step textual instructions for Bengali signs that are more structured and semantically faithful than free-form descriptions from a vanilla prompt. On the BdSLIG benchmark, SPI prompting outperformed vanilla prompting in most model-metric pairs: GPT-4.1 rose from 0.528 to 0.553 ROUGE-1 and from 0.416 to 0.492 METEOR; Gemini 2.0 Flash raised BERTScore from 0.416 to 0.456. The authors interpret the larger gains on semantic-oriented metrics as evidence that the parameter scaffolding, not mere lexical overlap, is doing the work. They also prese","pith_inferences":["The reported gains may partly be a vocabulary-alignment effect: because the same seven-category schema was used to write ground-truth annotations and to construct prompts, text metrics could reward matching annotation words rather than understanding the sign. A blind human-expert rating of instructions generated under both prompts would separate these explanations.","The uniform every-20th-frame sampling, which the paper flags as a limitation, likely discards key motion frames; an adaptive or salient-frame sampler could strengthen or change the SPI effect, especially for movement and temporal-dynamics categories.","SPI's canonical schema might be reused as a scaffold to generate training data for sign recognition or retrieval, or to align generated video descriptions with glosses.","For other under-resourced sign languages, the seven categories may need adjustment because sign parameters vary by language community."],"forward_implications":["SPI prompting improves zero-shot sign instruction generation without retraining, with the largest gains on high-capacity models such as GPT-4.1 and Gemini 2.5 Pro.","The structured, canonical output can be reused for classification, retrieval, or alignment tasks, not only as instructions for human learners.","BdSLIG provides a long-tail visual benchmark because Bengali sign video and instruction text are unlikely to be in VLM pretraining data, allowing tests of genuine visual grounding.","The paper's explicit caution about text metrics implies that future work should adopt human expert evaluation or LLM-based judges to measure true instruction quality."],"supporting_citations":[{"why":"Supplies the 60-word Bengali sign videos (BdSLW60) from which BdSLIG frames and annotations are drawn.","marker":"[15]"},{"why":"Defines the linguistic sign parameters used both as annotation guidelines and as the SPI prompt schema.","marker":"[23]"},{"why":"Establishes the SLIG task of generating textual sign instructions, which this paper extends to Bengali.","marker":"[11]"},{"why":"Learn2Sign provides explainable sign feedback on handshape, location, and movement, motivating parameter-structured prompts.","marker":"[21]"},{"why":"Documents that LLMs produce generic or semantically off-base sign descriptions without domain guidance, motivating SPI.","marker":"[22]"},{"why":"Supports the long-tail evaluation framing by showing task contamination is common in pretraining corpora.","marker":"[25]"}],"fun_headline_variants":["SPI prompts outperform plain prompts for Bengali sign AI","Structured sign-parameter prompts beat free-form for Bengali","For Bengali sign AI, SPI prompting beats vanilla in most metrics","SPI prompting improves Bengali sign instruction generation","Sign-parameter prompting gives more faithful Bengali sign steps"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The ground-truth instructions and the SPI prompts were written with the same seven sign-parameter categories, so the benchmark may reward matching that annotation vocabulary rather than correctly describing the sign itself.","fun_headline_variants_meta":{"raw":{"variants":["SPI prompts outperform plain prompts for Bengali sign AI","Structured sign-parameter prompts beat free-form for Bengali","For Bengali sign AI, SPI prompting beats vanilla in most metrics","SPI prompting improves Bengali sign instruction generation","Sign-parameter prompting gives more faithful Bengali sign steps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000764,"raw_usage":{"total_tokens":3215,"prompt_tokens":719,"completion_tokens":2496,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":463,"completion_tokens_details":{"reasoning_tokens":2433}},"tokens_in":463,"tokens_out":2496,"duration_ms":23337,"temperature":1.0,"reasoning_tokens":2433,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:31:31.205785+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have sign-language educators rate blinded instructions for executability and correctness without knowing which prompting strategy produced them; if SPI-generated instructions are not rated better than vanilla ones despite higher METEOR/BERTScore, the claim of semantic faithfulness fails. Alternatively, re-annotate a held-out set with free-form instructions not constrained by the seven categories; if SPI's text-metric advantage collapses, it was driven by schema alignment.","supporting_citations":[{"cited_title":"Bdslw60: A word-level bangla sign lan- guage dataset","cited_arxiv_id":null,"evidence_quote":"Supplies the 60-word Bengali sign videos (BdSLW60) from which BdSLIG frames and annotations are drawn."},{"cited_title":"Sign Language and Linguistic Universals","cited_arxiv_id":null,"evidence_quote":"Defines the linguistic sign parameters used both as annotation guidelines and as the SPI prompt schema."},{"cited_title":"Generating Signed Language Instructions in Large-Scale Dialogue Systems","cited_arxiv_id":"2410.14026","evidence_quote":"Establishes the SLIG task of generating textual sign instructions, which this paper extends to Bengali."},{"cited_title":"Learn2sign: Explainable ai for sign language learning","cited_arxiv_id":null,"evidence_quote":"Learn2Sign provides explainable sign feedback on handshape, location, and movement, motivating parameter-structured prompts."},{"cited_title":"Sig- nAlignLM: Integrating multimodal sign language process- ing into large language models","cited_arxiv_id":null,"evidence_quote":"Documents that LLMs produce generic or semantically off-base sign descriptions without domain guidance, motivating SPI."},{"cited_title":"Task contamination: Language models may not be few-shot anymore","cited_arxiv_id":null,"evidence_quote":"Supports the long-tail evaluation framing by showing task contamination is common in pretraining corpora."}],"review_version":1}