{"id":"5fb5c26a-fff1-456a-9a8b-eeaead71eece","arxiv_id":"2508.12666","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Cryfish pairs a WavLM audio encoder with the Qwen2 language model via a transformer connector, and is evaluated across speech and sound tasks on the Dynamic SUPERB Phase-2 benchmark.","lead":"This paper introduces Cryfish, an AI model that adds hearing to a text-based large language model by connecting a WavLM audio encoder to Qwen2 through a transformer connector. The authors test it on a multitask benchmark covering speech and sound, comparing it with other public models.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Competitive performance claim rests on an unshown evaluation protocol; without matched conditions, the comparison is unverifiable.","rationale":"The reader correctly identifies the validity of Dynamic SUPERB Phase-2 as a proxy and the lack of a visible protocol as the weakest assumptions. My stress-test focuses more narrowly on the matched-evaluation condition, which is the most concrete and checkable aspect of the central claim. I agree with the UNVERDICTED outcome because the abstract alone provides insufficient information to accept or reject the performance claim. I do not raise a different concern; the protocol question is already present in the reader's weakest_assumption, so I mark partial rather than agree to indicate that I prioritize the comparability aspect over the benchmark-validity aspect. The paper's architecture is plausible and there is no internal inconsistency visible from the abstract, so no fatal flaw is identified. The concrete test would settle whether the comparison is fair; until then, the verdict remains unverified.","tokens_in":823,"tokens_out":1488,"duration_ms":20220,"concrete_test":"Obtain the full paper and check the experiments section for the exact Dynamic SUPERB Phase-2 evaluation protocol: (1) the selected task subset, (2) metric definitions, (3) baseline model versions and whether they are used off-the-shelf or adapted, (4) inference hyperparameters. If the paper is not available, use the official Dynamic SUPERB Phase-2 harness to re-run one public baseline (e.g., Qwen2-Audio or SALMONN) using Cryfish's reported instruction template and decoding settings. If the re-run baseline score differs materially from the paper's reported baseline, the comparison is not matched and the central claim is weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that Cryfish performs competitively with publicly available models on Dynamic SUPERB Phase-2—depends entirely on the evaluation being a fair, matched comparison. The abstract does not disclose the evaluation protocol: no task list, no metric definitions, no baseline versions, no inference setting (e.g., few-shot prompts, instruction templates, audio sampling rates, or decoding parameters). Dynamic SUPERB Phase-2 is a multitask benchmark with heterogeneous tasks, and published scores are only meaningful when every model is evaluated under identical conditions. If Cryfish uses a different instruction template, a different number of examples, or a different fine-tuning regime than the public baselines, the relative performance could be an artifact of protocol mismatch rather than architectural merit. This is the weakest load-bearing premise because the architecture and training recipe are plausible but the claimed outcome cannot be assessed without the evaluation details. The absence of code/weights further prevents reproducing the comparison. This is not an accusation of wrongdoing; it is a statement that the evidence needed to support the claim is not available in the abstract, and the full text is likewise unavailable for this review.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Cryfish, an auditory-capable large language model that integrates WavLM audio-encoder features into Qwen2 via a transformer-based connector, with a specialized training strategy for diverse auditory tasks. The authors evaluate Cryfish on the Dynamic SUPERB Phase-2 multitask benchmark and claim 'in-depth analysis and detailed comparison' with publicly available models. The abstract, however, contains no quantitative results, no task list, and no evaluation protocol details, so the central empirical claim cannot be assessed from the available material.","tokens_in":939,"tokens_out":2119,"duration_ms":25863,"significance":"If the claimed performance is substantiated in the full text, Cryfish would represent a concrete recipe for making LLMs auditory-capable: a WavLM-to-Qwen2 connector with a specialized training strategy, validated on a benchmark designed for such models. This is a plausible and potentially useful contribution in an active research area. However, the abstract alone provides no evidence; no numbers, no baseline identifications, and no protocol disclosure. Therefore the significance of the work is currently unverified. The paper also mentions reproducible-sounding components (WavLM, Qwen2, transformer connector), but without the full experimental details its contribution cannot be weighed.","major_comments":[{"comment":"The central claim that Cryfish performs competitively or better on Dynamic SUPERB Phase-2 is stated entirely without quantitative support. No task list, metric definitions, baseline versions, or inference settings are given. This is load-bearing because the entire contribution is empirical; as presented, the manuscript does not allow the reader to verify the claimed performance or the fairness of the comparison. The authors should disclose the full evaluation protocol (including prompts, sampling rates, decoding parameters, and any task-specific tuning) in the main text, and the abstract should report at least headline performance numbers.","section":"Abstract"},{"comment":"The description of Dynamic SUPERB Phase-2 as 'specifically designed for auditory-capable models' raises a correctness-risk concern: if the benchmark interface implicitly favors the class of models Cryfish represents (e.g., through instruction format or training regime), comparisons with 'publicly available models' that were not adapted to the same interface may be misleading. This is not a claim of wrongdoing but a request for transparency: the paper must state whether all compared models were evaluated under identical conditions, including the exact instruction template, audio preprocessing, and any few-shot examples. Without this, the relative performance claim cannot be interpreted.","section":"Abstract"}],"minor_comments":[{"comment":"The phrase 'hearing is an essential capability' and 'generalizing complex auditory tasks across speech and sounds' is vague; the paper would benefit from a precise list of the task categories covered (e.g., ASR, speaker verification, sound event detection) in the abstract.","section":"Abstract"},{"comment":"The abstract does not state the model size or parameter count of Cryfish or its base Qwen2 variant, which is relevant for contextualizing comparison with publicly available models of potentially different scales.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"This review is based on the abstract only, as the full text was not available. The central empirical claim is unverifiable without the full experimental section and ideally code/weights. I recommend obtaining the full manuscript before making a final decision. The lack of any quantitative result in the abstract is unusual for a competitive-performance claim and should be addressed in revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the honest bottom line: from the abstract alone, there is nothing to verify. Cryfish is described as a WavLM audio encoder feeding Qwen2 through a transformer connector, trained with a specialized strategy and evaluated on Dynamic SUPERB Phase-2. That is a recognizable recipe from the audio-LLM literature—not a new capability, but a new instantiation of a known architecture. The abstract is clear and doesn't overclaim in tone, though it makes a competitive-performance claim without reporting any numbers.\n\nWhat the paper does well: it targets a publicly available multitask benchmark rather than a home-brewed evaluation, which is the right move. The component choices are sensible, and the line of work is active. If the full paper includes a rigorous protocol, this could be a useful systems contribution.\n\nThe soft spot is exactly what the stress-test note says: the evaluation protocol is absent. No task list, no metrics, no baselines, no inference settings, and no code release is mentioned. Without a matched protocol, the comparison to public models is unverifiable. That doesn't make the paper wrong—it makes it impossible to evaluate from the abstract. Also, the benchmark being 'specifically designed for auditory-capable models' is a mild concern because task selection could favor the model's training setup, but that's about the benchmark's design, not a flaw in the paper.\n\nGiven no full text, my verdict is 'unverdictable' rather than positive or negative. I'd like to see the full paper before recommending others read it. If it contains the missing details and honest results, it deserves a peer-review slot in an audio or speech venue. I would not cite it until the results are reproducible.","headline":"Abstract-only skim: a plausible audio-LLM recipe with no numbers or protocol; unverdictable, but worth a look if the full paper provides matched evaluation.","tokens_in":1557,"tokens_out":3160,"would_cite":false,"duration_ms":36446,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Cryfish claims that a transformer-connector bridge from a WavLM audio encoder into Qwen2 lets a single LLM perform competitively across speech and non-speech audio tasks on Dynamic SUPERB Phase-2.","keywords":["auditory large language models","WavLM audio encoder","Qwen2","transformer connector","multitask audio understanding","Dynamic SUPERB Phase-2","speech recognition","sound event detection"],"falsifier":"Take the same set of public models and rerun the Dynamic SUPERB Phase-2 battery with identical prompts, audio preprocessing, and decoding settings, then also test on a held-out set of real-world audio clips with noise, overlapping speakers, and uncommon sounds. If Cryfish's ranking shrinks or reverses under matched conditions, the benchmark protocol rather than the model carries the claim.","tokens_in":649,"feed_emoji":"🎧","tokens_out":5384,"duration_ms":58014,"temperature":0.7,"pith_summary":"The paper introduces Cryfish, an auditory LLM that attaches a pretrained audio encoder, WavLM, to a text LLM, Qwen2, through a transformer-based connector, and adapts the whole system to several auditory tasks with a specialized training strategy. Its goal is to show that listening can be added to an LLM as a general capability instead of as a collection of task-specific heads. The authors evaluate the model on Dynamic SUPERB Phase-2, a benchmark designed for auditory-capable models, and compare it in depth with publicly available audio-LLMs. If the claim holds, a single model can handle speech, sound events, and related audio tasks with one shared front end.","feed_headline":"Cryfish makes a single LLM hear across speech and sound tasks","feed_subtitle":"Audio features plus a transformer connector let one LLM take on the Dynamic SUPERB Phase-2 hearing battery.","key_machinery":"The load-bearing mechanism is the transformer-based connector that converts WavLM's audio features into input tokens Qwen2 can read, so the LLM treats sound as text-like sequences. The second mechanism is the specialized training strategy — the order and balance of auditory tasks — which is what lets the model generalize across speech and non-speech audio rather than overfitting one task. The connector does the alignment; the schedule does the generality.","core_discovery":"Cryfish's central claim is that the architecture — WavLM audio encoder, transformer-based connector, and Qwen2 language model — is a viable recipe for making one LLM auditory across speech and sound tasks. The authors report competitive or better performance than publicly available models on Dynamic SUPERB Phase-2, a multitask benchmark built for auditory-capable models. The discovery lies in the combination: the connector aligns audio representations with the LLM's token space, and a multi-stage training scheme lets the model absorb speech and non-speech audio skills while keeping its language abilities.","pith_inferences":["A natural next probe is to ablate the connector, the audio encoder, and the task-order schedule separately to see which component actually carries Cryfish's benchmark position.","If the benchmark's task mix weights speech and paralinguistic tasks more heavily than everyday sound events, the reported position may reflect the benchmark, not general hearing; readers should check per-task breakdowns.","A concrete transfer test would keep Cryfish's connector and training schedule, swap the text LLM, and see whether the listening ability carries over to a different language model.","The same connector idea might extend beyond audio to other continuous signals, such as physiological or vibration streams, if the alignment layer is general enough."],"forward_implications":["One encoder–connector–LLM stack can cover multiple hearing tasks, meaning auditory capability does not require a separate model per task.","Speech tasks and general sound tasks can share the same audio front end, so gains on one group of tasks can transfer to others through the shared connector.","Dynamic SUPERB Phase-2 provides a common scale on which future auditory LLMs can be compared, assuming the evaluation protocol is held fixed.","The training schedule is part of the recipe: task order and balance are what let the model stay text-fluent while learning audio.","An auditory LLM can be assembled from existing components rather than trained from scratch, which lowers the barrier to adding hearing to future LLMs."],"supporting_citations":[],"fun_headline_variants":["One LLM hears speech and sound via Cryfish's connector","Cryfish teaches Qwen2 to listen with WavLM audio features","Cryfish hears across tasks, rivals public models on SUPERB Phase-2","Transformer connector lets Cryfish LLM hear speech and sounds"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that Dynamic SUPERB Phase-2 is a fair, comprehensive measure of auditory capability and that Cryfish's comparison with public models uses matched evaluation conditions; if either is off, the competitive showing does not generalize to real use.","fun_headline_variants_meta":{"raw":{"variants":["One LLM hears speech and sound via Cryfish's connector","Cryfish teaches Qwen2 to listen with WavLM audio features","Cryfish hears across tasks, rivals public models on SUPERB Phase-2","Transformer connector lets Cryfish LLM hear speech and sounds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00019,"raw_usage":{"total_tokens":1128,"prompt_tokens":648,"completion_tokens":480,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":392,"completion_tokens_details":{"reasoning_tokens":412}},"tokens_in":392,"tokens_out":480,"duration_ms":5732,"temperature":1.0,"reasoning_tokens":412,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:18:19.391970+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same set of public models and rerun the Dynamic SUPERB Phase-2 battery with identical prompts, audio preprocessing, and decoding settings, then also test on a held-out set of real-world audio clips with noise, overlapping speakers, and uncommon sounds. If Cryfish's ranking shrinks or reverses under matched conditions, the benchmark protocol rather than the model carries the claim.","supporting_citations":[],"review_version":1}