{"id":"95c24257-78c1-4d5f-9de8-ad416f7198dc","arxiv_id":"2508.00391","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"Cued-Agent combines an MLLM hand reader, a lip-reading Transformer, a parameter-free hand-lip fusion, and an LLM self-correction to convert Cued Speech video to sentences, with a new hearing-impaired dataset, but its phoneme accuracy trails prior SOTA on most settings.","lead":"This paper presents Cued-Agent, a four-part AI system that turns Cued Speech videos (lip and hand gestures) into text, using a large vision-language model to read hand signs and a large language model to correct phonemes into sentences. The authors also collected a new dataset of hearing-impaired Mandarin Cued Speech users, but the system's reported accuracy is slightly worse than a prior method on most benchmarks despite claims to the contrary.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cued-Agent's reported results contradict its 'superb' claim: Table 2 shows higher CER/WER than STF-ACSR in 7 of 8 comparisons.","rationale":"The reader's weakest_assumption focuses on the reliability of GPT-4o hand recognition and the lack of a validation split for hyperparameters. These are valid methodological concerns, but they are secondary to the direct contradiction between the paper's headline claim and its own results. The central claim is that Cued-Agent 'performs superbly' and 'outperforms SOTA methods'; Table 2 shows the opposite on 7 of 8 metric/dataset combinations. This is not a hidden assumption or a subtle statistical issue but a visible mismatch between the claim and the reported evidence. If the table is accurate, the claim is false; if the table is inaccurate, the paper's experimental reporting is unreliable. Either way, the strongest claim does not hold. The hyperparameter concern could explain why the numbers are what they are, but it does not change the fact that the current evidence contradicts the abstract. The formal error in Eq. (5) is also notable but does not directly invalidate the empirical comparison. The new dataset and the training-free fusion idea may still have scientific value, but the paper's primary assertion of superiority is unsupported, so the REJECT verdict is appropriate. I disagree with the reader's weakest_assumption because it identifies a different (though related) point; the most load-bearing concern is the empirical contradiction in Table 2, which the reader mentioned in the rationale but did not place in the weakest_assumption field.","tokens_in":12802,"tokens_out":6879,"duration_ms":62718,"concrete_test":"Independently reproduce the Table 2 evaluations using the released code and dataset, and compute paired bootstrap 95% confidence intervals for the CER/WER differences between Cued-Agent and STF-ACSR on each dataset. If Cued-Agent is not significantly better on a majority of the 8 metric-dataset pairs, or if the reproduced point estimates match Table 2, the abstract's 'superb' claim is unsupported and should be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and contributions claim Cued-Agent 'performs superbly' and 'outperforms SOTA methods', but the paper's own Table 2 does not support this. Across the four experimental settings, Cued-Agent has higher CER than STF-ACSR in all four cases (2.61 vs 1.82, 6.72 vs 4.62, 9.05 vs 8.35, 12.67 vs 10.96) and higher WER in three of four cases (6.56 vs 5.19, 16.23 vs 12.21, 29.86 vs 25.67); the only win is WER on MCCSD (6-H) (20.54 vs 21.06). Thus the system is not superior to the listed SOTA; at best it is comparable with a slight edge in one metric. The 'superb' claim is therefore directly contradicted by the evidence presented. Unless the comparison is confounded by different test splits or test-set hyperparameter tuning, the headline claim must be retracted or substantially weakened. This is the most load-bearing concern because it targets the central empirical assertion of the paper: if Table 2 is accurate, the main contribution's value proposition fails, regardless of the architectural novelty or the new dataset.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes Cued-Agent, a multi-agent system for Mandarin Automatic Cued Speech Recognition (ACSR). Four agents are combined: an MLLM-based hand recognition agent that converts keyframe hand classifications into a one-hot prompt matrix, a Transformer-based lip recognition agent finetuned with CTC/attention loss, a hand-prompt decoding agent that fuses hand prompts into lip logits by weighted addition during beam search, and an LLM-based self-correction phoneme-to-word agent. The authors also collect a new multi-hearing-impaired Mandarin CS dataset (MHI-MCCSD) with eight hearing-impaired cuers and propose two sentence-level metrics, S-WER and Semantic Score. The abstract and conclusion claim that the system 'performs superbly' and 'outperforms' state-of-the-art methods, while Section 5.2 more cautiously describes the result as comparable to STF-ACSR.","tokens_in":12919,"tokens_out":10821,"duration_ms":105429,"significance":"The paper has several strengths: the multi-agent decomposition of ACSR is novel, the parameter-free hand-lip fusion idea is simple and relevant to data-scarce settings, the new MHI-MCCSD dataset addresses a real gap in hearing-impaired evaluation, and the sentence-level output is a useful step beyond phoneme-only recognition. However, the headline performance claim is not supported by the paper's own Table 2, where Cued-Agent has higher CER than STF-ACSR on all four test settings and higher WER on three of four settings. The fusion weights are introduced without a described validation protocol, and Eq. (5) misstates the CTC prefix probability. If the performance claims are revised and the technical points are corrected, the dataset and the multi-agent pipeline could still make a moderate contribution, but the current text overstates what the experiments show.","major_comments":[{"comment":"The claim that Cued-Agent 'performs superbly' and 'outperforms SOTA methods' is directly contradicted by Table 2. For all four test settings, Cued-Agent has higher CER than STF-ACSR: 2.61 vs 1.82 on MCCSD (1-H), 6.72 vs 4.62 on MCCSD (1-HI), 9.05 vs 8.35 on MCCSD (6-H), and 12.67 vs 10.96 on MHI-MCCSD (8-HI). The WER is also higher in three of the four settings, with only MCCSD (6-H) WER being better (20.54 vs 21.06). The abstract and conclusion therefore overstate the results, and Section 5.2's word 'comparable' is not equivalent to 'superb.' The performance claims must be retracted or substantially weakened, or new experiments must support the claimed superiority.","section":"Abstract; Table 2; Section 6"},{"comment":"The reported results depend on the free hyperparameters lambda_prompt=4.5 and lambda_decode=0.5, as well as on the keyframe thresholds sigma=6 and theta=2, but the paper does not describe a validation split or a tuning procedure. If these values were selected after inspecting test-set results, the CER/WER numbers in Tables 2-4 are fitted values rather than independent predictions. The authors should specify how each hyperparameter is chosen, report results on a held-out validation split, and include a sensitivity analysis over lambda_prompt and lambda_decode.","section":"Section 5.1; Eqs. (6), (8)"},{"comment":"Eq. (5) is not a correct statement of the CTC prefix probability. The prefix probability is the cumulative probability over all label sequences that have hypothesis h as a prefix, including arbitrarily long continuations with blank and non-blank transitions; it is not the sum over a single next label nu in Q union {<eos>} of p_ctc(h·nu|L). Because Eqs. (7)-(8) use this quantity in the beam-search scoring, the derivation as written is technically unsound and should be corrected to the standard recursive prefix-probability definition.","section":"Section 3.5, Eq. (5)"},{"comment":"The hand recognition agent is load-bearing because its one-hot matrix H is directly added to the CTC logits with weight 4.5 in Eq. (6). The paper never reports the accuracy of the GPT-4o hand position and shape classification on the keyframe set, so the error injected into the decoder is unknown. The ablation in Table 4 shows only an end-to-end effect; an intermediate evaluation of hand recognition accuracy, or at least a confusion matrix for position and shape categories, is needed to justify the design and to quantify error propagation.","section":"Section 3.3; Table 4"}],"minor_comments":[{"comment":"The text should be aligned with the numbers: Table 2's caption says 'good performance,' the main text says 'comparable,' and the abstract says 'superb'; these are inconsistent and should be harmonized.","section":"Section 5.2 and Table 2 caption"},{"comment":"The phrase 'we innovatively add' is editorializing; also, while the fusion is parameter-free in the sense of no learned weights, lambda_prompt is a hyperparameter that must be tuned, so the 'parameter-free' claim should be qualified.","section":"Section 3.5"},{"comment":"The exact prompt templates for the hand recognition and self-correction agents are not provided; the descriptions in the text are not sufficient for reproduction. The prompts should be included in a supplementary file.","section":"Sections 3.3 and 3.6"},{"comment":"The paper claims that MHI-MCCSD is 'publicly available,' but only a GitHub URL for the implementation is given; the dataset release mechanism should be stated explicitly.","section":"Section 4"},{"comment":"The value of lambda_train in the joint CTC/attention loss is not reported; this is a tunable parameter that affects the lip recognizer and should be specified.","section":"Section 3.4, Eq. (3)"},{"comment":"The Semantic Score uses Sentence-BERT [31], but the authors should specify the exact model checkpoint and confirm that it provides appropriate Mandarin embeddings, since the original Sentence-BERT models are primarily English-oriented.","section":"Section 3.7"},{"comment":"The confusion matrices would benefit from a colorbar and a description of the color scale; the current figure is difficult to interpret quantitatively.","section":"Figure 7"},{"comment":"References [11] and [12] are identical in title, authors, and venue; this duplication should be removed.","section":"References"},{"comment":"The phrase 'training-free multimodal alignment' is misleading because the lip recognition agent is finetuned; a more precise phrasing would be 'training-free hand-lip fusion.'","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern lands: Table 2 contradicts the abstract's 'superb performance' claim, and this is the most load-bearing issue. However, the paper's other contributions—the MHI-MCCSD dataset, the multi-agent decomposition, and the sentence-level output—are potentially salvageable if the claims are revised and the validation protocol is clarified. I would ask the editor to require a clear statement on whether lambda_prompt and lambda_decode were tuned on the test set, and to verify that the new dataset will be publicly released, since it is a major part of the paper's value."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take on arXiv:2508.00391. The system is a real attempt to avoid training a fusion network by injecting hand-derived logit bias into a CTC/attention decoder, wrapped in a four-agent LLM pipeline, and it comes with a new hearing-impaired dataset (MHI-MCCSD, 8 cuers, 5,272 samples). That part is worth a look. But the headline claim that Cued-Agent 'performs superbly' against SOTA is contradicted by the paper's own Table 2: it loses to STF-ACSR on CER in all four settings and on WER in three of four. The only win is WER on MCCSD (6-H) by half a point. The authors soften this to 'comparable' in Section 5.2, but the abstract and conclusion still say 'superb' and 'outperforms'. That mismatch is a real problem, not a stylistic quibble.\n\nThe parameter-free fusion idea is clever, and the new dataset is the most useful contribution if it is actually released—the paper calls it 'publicly available' but gives no link, and I found no release statement. The Self-Correction P2W agent is a reasonable post-processing step, though its gains are modest and it would be stronger with an analysis of failure cases.\n\nSoft spots, in rough order of severity. First, lambda_prompt=4.5 and lambda_decode=0.5 are set with no described validation split; the test numbers may simply be fitted on the test set, and the paper never discusses sensitivity to either weight. Second, Eq. (5) misstates the CTC prefix probability: the sum should run over all label continuations, not a single next symbol. That is a technical error in the formal description of the core decoding agent. Third, the hand-recognition path depends on GPT-4o with no reported hand accuracy, and the API is non-deterministic; reproducibility therefore rests on proprietary, non-pinned infrastructure. Fourth, the GitHub link points to an empty or placeholder repo as far as I can tell, and there is no commit hash.\n\nI would send this to a serious referee, but it needs major revision before acceptance: retract or heavily qualify the 'superb' claim, add validation-based hyperparameter selection or at least a sensitivity sweep, correct Eq. (5), and actually release the dataset and code with a persistent identifier. The architecture and dataset justify referee time; the current evidence does not justify the conclusions.","headline":"Plausible multi-agent ACSR system and a genuinely useful hearing-impaired dataset, but the headline accuracy claim is contradicted by the paper's own Table 2.","tokens_in":13610,"tokens_out":3415,"would_cite":false,"duration_ms":33237,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Cued-Agent is the first multi-agent system for Automatic Cued Speech Recognition, using four specialized agents to fuse hand and lip information without training a fusion network.","keywords":["Cued Speech recognition","multi-agent system","hand-lip fusion","multimodal large language model","phoneme-to-word conversion","hearing-impaired dataset","lip reading","parameter-free fusion"],"falsifier":"Run the Hand Recognition Agent alone on a held-out set of hearing-impaired cuers with known hand-code labels and compute position and shape classification accuracy; if accuracy on confusable codes is near chance, set $\\lambda_{\\text{prompt}} = 0$ and compare. A failure to show a significant CER or WER drop when the hand prompt is removed would indicate the hand pathway is not carrying the claimed information.","tokens_in":12437,"feed_emoji":"🤟","tokens_out":7964,"duration_ms":74357,"temperature":0.7,"pith_summary":"This paper proposes Cued-Agent, a multi-agent system that converts Mandarin Cued Speech videos—hand gestures plus lip movements—into natural-language sentences. Its central claim is that the long-standing problem of asynchrony between hand and lip signals can be handled without training a fusion network: a multimodal large language model reads hand positions and shapes from selected keyframes, and those one-hot hand prompts are added directly into the CTC decoding scores of a lip-reading model. A separate LLM agent then revises the predicted phoneme sequence into a sentence using Cued Speech rules and semantic coherence. The authors also collect a new dataset from eight hearing-impaired cuers and report that Cued-Agent matches the previous state of the art on phoneme error rates while being the only method that produces sentence-level output. If correct, this lowers the cost of building ACSR systems for new speakers and languages.","feed_headline":"Four-agent pipeline reads Cued Speech lips and hands into text","feed_subtitle":"Adding recognized hand cues directly into decoding scores skips fusion training and turns phonemes into sentences.","key_machinery":"The load-bearing object is the hand prompt matrix $H \\in \\mathbb{R}^{T \\times q}$, a one-hot encoding of the hand position and shape categories that GPT-4o assigns to each keyframe, propagated over the frames in each slow-motion group. During decoding, the agent forms the fused feature $L_H = L' + \\lambda_{\\text{prompt}} \\cdot H$ by adding $H$ to the lip model's projected CTC feature $L'$, then computes the CTC prefix score from $L_H$ and combines it with the attention score in beam search. This is what makes hand-lip fusion parameter-free: the hand signal enters as a bias on the CTC logits rather than through trained cross-modal layers. The second mechanism is the Self-Correction P2W agent, which repeatedly rewrites the phoneme sequence under prompts encoding Mandarin CS conversion rules, in-context examples, and confusion-pair contrasts, finally emitting a sentence.","core_discovery":"On the paper's own terms, the discovery is that a Cued Speech recognizer does not need a trained multimodal fusion module. Cued-Agent decomposes recognition into four cooperating agents: a GPT-4o-based Hand Recognition agent that screens slow-motion keyframes and classifies hand position and shape; a pretrained Transformer Lip Recognition agent finetuned only on lip frames; a Hand Prompt Decoding agent that forms $L_H = L' + \\lambda_{\\text{prompt}} \\cdot H$, adding a one-hot hand-prompt matrix to the projected lip features during beam search; and a DeepSeek-R1-based Self-Correction Phoneme-to-Word agent that turns phonemes into sentences. In experiments the system reaches 2.61% CER / 6.56% WER on the single normal-hearing cuer subset and 12.67% / 29.86% on the new eight-hearing-impaired-cuer set, comparable to the previous state of the art, and it reports sentence-level S-WER from 12.1% to 40.84% across settings, the first such results. The paper's key novelty claims are the parameter-free fusion via prompt-weighted decoding and the self-correcting phoneme-to-word conversion.","pith_inferences":["A testable extension is to measure GPT-4o's hand position and shape accuracy directly; if it is high, the parameter-free fusion can be seen as injecting near-oracle hand information, and if low, Cued-Agent's gains would be expected to shrink on unseen cuers.","The confusion matrices show errors concentrate on phonemes with identical hand codes (for example, b/p and yu/w); one implication is that the lip model, not hand fusion, must resolve those, so pairing the hand prompt with a language-model prior over Mandarin syllables could close much of the remaining error.","The choices $\\lambda_{\\text{prompt}} = 4.5$ and $\\lambda_{\\text{decode}} = 0.5$ are reported without a validation sweep; a sensitivity analysis over these weights would reveal how much of the result depends on hand-prompt strength rather than on the hand signal itself.","Because sentence-level metrics are new, future work could compare human lip-reader performance on the same videos to calibrate how much of the semantic correction is doing the work."],"forward_implications":["Adding a new cuer or language no longer requires collecting hand labels to train a fusion module; only the lip model needs finetuning.","Cued Speech output can be presented as readable sentences, with S-WER and Semantic Score as quantitative measures, making ACSR usable in assistive communication tools.","The method's hand information can be swapped: any multimodal LLM that can classify the CS hand code can replace GPT-4o without changing the fusion mechanism.","Hearing-impaired cuers, whose lip movements are harder to read, benefit most from the hand prompt pathway, since the ablation shows hand information gives the largest gains on hearing-impaired subsets.","The same four-agent recipe can be applied to other cued languages whose hand codes differ, as long as prompts and support sets are rebuilt."],"supporting_citations":[{"why":"Supplies the MLLM keyframe filtering and prompt strategy for hand position and shape recognition that Cued-Agent's Hand Recognition agent inherits and adapts.","marker":"[15]"},{"why":"Provides the Mandarin CS dataset MCCSD used for training and evaluation, along with the CMML Transformer fusion baseline that Cued-Agent compares against.","marker":"[25]"},{"why":"Provides the EcoCued baseline and represents the prior Transformer-based fusion approach to ACSR that Cued-Agent avoids.","marker":"[26]"},{"why":"Supplies the pretrained audio-visual speech recognition model used as the base for the Lip Recognition agent and the decoder.","marker":"[28]"},{"why":"The multimodal LLM that performs the actual hand position and shape classification in the Hand Recognition agent.","marker":"[29]"},{"why":"The reasoning-focused LLM used by the Self-Correction P2W agent to refine phoneme sequences into sentences.","marker":"[6]"},{"why":"Defines the joint CTC/attention training loss and beam-search decoding formalism that the hand-prompt weighting modifies.","marker":"[41]"},{"why":"Supplies the sentence-embedding model used to compute the new Semantic Score sentence metric.","marker":"[31]"},{"why":"Defines the CTC prefix probability used in the weighted beam search score.","marker":"[5]"}],"fun_headline_variants":["Cued-Agent: Four agents decode cued speech without fusion training","No fusion training needed: multi-agent reads cued speech","Multi-agent system turns lips and hand cues into text","First multi-agent for cued speech recognition","Cued-Agent: Training-free hand prompts boost recognition"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole pipeline rides on GPT-4o correctly recognizing hand positions and shapes from keyframes of unseen cuers, because those one-hot hand prompts are injected directly into the decoder's scores and no trained fusion layer exists to absorb a wrong hand classification.","fun_headline_variants_meta":{"raw":{"variants":["Cued-Agent: Four agents decode cued speech without fusion training","No fusion training needed: multi-agent reads cued speech","Multi-agent system turns lips and hand cues into text","First multi-agent for cued speech recognition","Cued-Agent: Training-free hand prompts boost recognition"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000262,"raw_usage":{"total_tokens":1673,"prompt_tokens":1099,"completion_tokens":574,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":715,"completion_tokens_details":{"reasoning_tokens":494}},"tokens_in":715,"tokens_out":574,"duration_ms":5994,"temperature":1.0,"reasoning_tokens":494,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:10:32.753602+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the Hand Recognition Agent alone on a held-out set of hearing-impaired cuers with known hand-code labels and compute position and shape classification accuracy; if accuracy on confusable codes is near chance, set $\\lambda_{\\text{prompt}} = 0$ and compare. A failure to show a significant CER or WER drop when the hand prompt is removed would indicate the hand pathway is not carrying the claimed information.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the MLLM keyframe filtering and prompt strategy for hand position and shape recognition that Cued-Agent's Hand Recognition agent inherits and adapts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Mandarin CS dataset MCCSD used for training and evaluation, along with the CMML Transformer fusion baseline that Cued-Agent compares against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the EcoCued baseline and represents the prior Transformer-based fusion approach to ACSR that Cued-Agent avoids."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the pretrained audio-visual speech recognition model used as the base for the Lip Recognition agent and the decoder."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The reasoning-focused LLM used by the Self-Correction P2W agent to refine phoneme sequences into sentences."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the sentence-embedding model used to compute the new Semantic Score sentence metric."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the CTC prefix probability used in the weighted beam search score."}],"review_version":1}