{"id":"28705e70-f1f4-4154-9c56-2060539507d0","arxiv_id":"2507.21588","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A three-stage prompt-tuning method for audio-visual multi-task incremental learning is proposed, reporting state-of-the-art results on AVE, AVVP, AVS, and AVQA, with caveats about its evaluation metric and ablations.","lead":"This paper introduces a three-stage prompt-tuning method that helps an audio-visual model learn several tasks one after another without forgetting earlier ones. It reports top accuracy on four audio-visual benchmarks, but a custom scoring metric and internal inconsistencies weaken the evidence.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Full three-stage PHP is not supported by the paper's own ablations: two-component subsets beat it on both anti-forgetting and transfer metrics, so the central design claim is internally contradicted.","rationale":"The reader identified the custom Diff metric as the weakest assumption. I disagree in part: even with the raw difference Amulti-Asingle, PHP is the only positive-transfer method (+1.11 vs PC -0.58), so the qualitative 'only positive transfer' claim does not collapse with the metric. The +7.79% magnitude is metric-dependent and unvalidated, but it is not the most load-bearing weakness. The more decisive problem is internal: the paper's own ablations show the full three-stage model is dominated by its own subsets, so the central mechanism claim is not supported by the evidence. This reinforces the Reader's REJECT verdict, though through a different load-bearing path. The four-task SOTA claim also lacks baseline comparisons, but the self-ablation contradiction is sufficient on its own.","tokens_in":25436,"tokens_out":10628,"duration_ms":119269,"concrete_test":"Run the full ablation matrix (rows 1-8 of Tables 3 and 4) with 5 random seeds under the same protocol, and perform paired significance tests (bootstrap or Wilcoxon) comparing the full model to its best subset (TMA+TMDG and TMA+TMI). Then evaluate those top subsets on the four-task sequences in Table 7. If the full model is not statistically superior to a two-component subset on at least one of Amean/Fmean/Diff, or if a subset matches it on the four-task sequences, the paper's central three-stage contribution is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the three-stage PHP design (TMA in shallow layers, TMDG in middle layers, TMI in deep layers) is what produces state-of-the-art retention and transfer. The paper's own ablations contradict this. In Table 3 (anti-forgetting), the full model (row 8) has Amean=58.85, Afinal=54.74, Fmean=3.32, while the TMA+TMDG subset (row 5) is better on every metric: Amean=59.54, Afinal=56.01, Fmean=3.21. In Table 4 (transfer), the full model (row 8) has Asingle=62.36, Amulti=63.47, Diff=+7.77, while TMA+TMI (row 6) has Asingle=62.73, Amulti=64.26, Diff=+7.99, and TMA alone (row 2) has Diff=+8.05. Thus no metric in the ablations is best served by the full three-stage model. The reported SOTA may still hold for the full model against external baselines, but the paper's mechanism is not supported: the additional middle and deep stages are not shown to contribute to the claimed balance. Any verdict that rests on the three-stage design being responsible for the results must treat this as a load-bearing gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a three-stage Progressive Homeostatic and Plastic (PHP) prompt tuning method for audio-visual multi-task incremental learning, consisting of a task-shared modality aggregating (TMA) adapter in shallow layers, a task-specific modality-shared dynamic generating (TMDG) adapter in middle layers, and task-specific modality-independent (TMI) prompts in deep layers. The authors claim state-of-the-art performance on four audio-visual tasks (AVE, AVVP, AVS, AVQA) in both anti-forgetting and transfer ability, and they introduce a new transfer metric, Diff, to support the claim that PHP is the only method with positive multi-task transfer.","tokens_in":25713,"tokens_out":6359,"duration_ms":68401,"significance":"The targeted problem—continual learning across multiple audio-visual tasks—is timely and relevant, and the paper offers a concrete modular design with a plausible motivation. The authors also provide a large set of experiments, including per-order results in the supplementary material and a promise of code. However, the central design claim that the three-stage architecture achieves the reported balance is not supported by the paper's own ablations, and the transfer claim rests on a bespoke metric. As the primary contributions are thus not established, the paper's current form does not make a convincing case for its stated results.","major_comments":[{"comment":"The full three-stage model is not the best configuration in the paper's own ablations. In Table 3 (anti-forgetting), row 5 (TMA+TMDG) outperforms row 8 (full PHP) on Amean (59.54 vs 58.85), Afinal (56.01 vs 54.74), and Fmean (3.21 vs 3.32). In Table 4 (transfer), row 2 (TMA alone) has Diff +8.05 vs the full model's +7.77, and row 6 (TMA+TMI) has +7.99. Thus on no metric is the full three-stage model the best. This directly contradicts the abstract and contribution list, which attribute the reported performance to the three-stage progressive design. The design claim is load-bearing and is undermined by the evidence presented.","section":"Tables 3 and 4"},{"comment":"The four-task incremental learning results, which support the abstract's claim of \"SOTA performance in different orders of four tasks,\" are presented without any baseline comparison. Table 7 shows only the proposed method's per-stage accuracies for two four-task sequences. Without comparing to existing incremental learning methods or even to the ablated variants in Tables 3 and 4, these numbers cannot validate a state-of-the-art claim. This is a major evidential gap for a central assertion of the paper.","section":"Table 7"},{"comment":"The transfer metric Diff is introduced in this paper and is not a standard measure. The headline result that PHP is the only method with positive transfer (+7.79%) depends on this metric's quadratic baseline penalty, which the authors choose without standard justification. The paper does not compare Diff with common transfer metrics (e.g., simple relative change or per-task differences), nor does it analyze how the ranking of methods changes under alternative normalizations. Since the claim of positive transfer is a unique selling point of the paper, the metric's form is load-bearing, and its current ad hoc derivation is insufficient.","section":"Supplementary Eq. (19)"},{"comment":"The text states that \"our approach outperforms the other methods in terms of all metrics,\" but Table 1 shows per-task cases where baselines are better: for AVE, Dualprompt has higher Amean (63.00 vs 62.03) and S-prompt has lower Fmean (4.78 vs 6.72); for AVQA, PC has higher Amean (69.55 vs 69.29) and Afinal (69.46 vs 68.56). Only the overall averaged metrics favor PHP. The wording is therefore inaccurate and overstates the comparison, requiring correction or qualification.","section":"Table 1 and Section 4.1"}],"minor_comments":[{"comment":"The conclusion says \"Extensive experiments on three audio-visual tasks (AVE, AVVP, AVS and AVQA)\", but four tasks are listed; this should be corrected.","section":"Conclusion"},{"comment":"Equation (14) defines P Xa = concat(Pa, V), but this should likely be concat(Pa, A) to match the audio branch; as written it uses the visual feature V.","section":"Section 3.5, Eq. (14)"},{"comment":"The heading \"Ablation study on deep prompts\" appears twice consecutively in the text.","section":"Section 4.2"},{"comment":"The reference list contains duplicates: [35] and [36] are the same paper, [74] and [75] are the same, and [80] and [81] are the same.","section":"References"},{"comment":"The Diff definition in Eq. (19) and the epsilon-corrected version in Eq. (20) are not reconciled; the main text refers only to Eq. (19), leaving ambiguity about which formula generated Table 2.","section":"Supplementary Eq. (19)-(20)"},{"comment":"The heading \"Quantitative Results\" introduces a section that is entirely qualitative (Figure 5 and descriptive comparisons); the heading should be changed to \"Qualitative Results\".","section":"Section 4.3"}],"recommendation":"reject","confidential_remarks":"The paper's central narrative—that the three-stage progressive design is responsible for the reported state-of-the-art balance—is contradicted by its own ablation tables, where a two-component subset and even a single component achieve better anti-forgetting and transfer scores. This is not a presentation issue but a fundamental mismatch between claim and evidence. Additionally, the four-task SOTA claim lacks any baselines, and the positive-transfer claim is anchored to a metric the authors introduce. Even with revision, the core design claim would need to be substantially reworked, and the current evidence does not support acceptance or even a targeted revision that preserves the paper's main thesis. I would encourage the authors to reframe the paper as a study of component trade-offs and to add proper baselines and metric validation, but in its present form the manuscript does not meet the bar for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper before spending time on it. First, the setting is real: audio-visual multi-task incremental learning with CLIP/CLAP is a legitimate and underexplored problem, and the three-level prompt/adapter decomposition (shared shallow, task-specific middle, modality-specific deep) is a sensible starting point. The authors do a lot of work: six task orders for the main comparisons, component ablations, component-order ablations, parameter sensitivity, and two four-task sequences. The code link is there, and the qualitative examples are illustrative. That is real effort, and the paper is readable.\n\nThe problem is that the central claim—that the full three-stage PHP model is what delivers the balance of retention and transfer—is contradicted by the paper's own tables. In Table 3, full model (row 8) has Amean 58.85, Afinal 54.74, Fmean 3.32. TMI alone gets Amean 59.70, TMA+TMDG gets Afinal 56.01 and Fmean 3.21. In Table 4, full model gets Diff +7.77, but TMA alone gets +8.05 and TMA+TMI gets +7.99. So on every metric, some subset of the components beats the full model. The paper never explains why the full configuration is the headline; it just asserts each component contributes. That is not a minor quibble—the entire narrative about progressive hierarchy rests on it.\n\nSecond, the transfer metric Diff is not just custom; it is miscalibrated. The formula in Eq. 19 multiplies the normalized improvement by (1 + Asingle/100)^2, which grows with Asingle. The authors call this a penalty for high baselines, but it is in the numerator. A method with Asingle=80 and a 5% relative gain scores much higher than one with Asingle=50 and the same relative gain. That is the opposite of their stated intent, and the headline \"only method with positive transfer\" depends entirely on this metric. No external validation is offered.\n\nThere are smaller issues: the four-task SOTA claim has no baseline comparison, only their own method on two orders. The abstract/conclusion says \"three tasks\" then lists four. Reference list is sloppy but harmless.\n\nWho is this for? Continual learning and multimodal researchers could use the problem formulation and the component design as a starting point, but not the empirical claims as they stand. It deserves a serious referee because the setting is meaningful and the experimental effort is substantial, but the paper needs major revision: reconcile the ablations, fix or replace the Diff metric, and add real baselines for the four-task setting. I would not cite it in its current form.","headline":"A genuinely new problem setting and a lot of careful experimentation, but the paper's central three-stage claim is contradicted by its own ablations, and the custom transfer metric behaves opposite to its stated purpose.","tokens_in":26260,"tokens_out":4163,"would_cite":false,"duration_ms":44698,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a three-stage progressive prompt hierarchy lets one model learn four audio-visual tasks in sequence, beat all seven baselines on anti-forgetting, and become the only compared method with positive cross-task transfer…","keywords":["audio-visual multi-task incremental learning","catastrophic forgetting","prompt tuning","continual learning","knowledge transfer","dynamic prompt generation","audio-visual event localization","audio-visual question answering"],"falsifier":"Recompute transfer for every method in Table 2 using a pre-existing standard measure, such as the plain average gain $\\bar{A}_{\\mathrm{multi}}-\\bar{A}_{\\mathrm{single}}$ or the mean per-task relative improvement. If under a standard metric PHP no longer shows positive transfer while all baselines stay negative, the paper's distinctive claim fails; if several baselines also turn positive, the claim that PHP is the only transferring method fails.","tokens_in":25225,"feed_emoji":"🎬","tokens_out":7911,"duration_ms":79137,"temperature":0.7,"pith_summary":"This paper proposes PHP, a three-stage prompting method for learning an ongoing stream of audio-visual tasks without retraining on old data. Its central claim is that the forgetting-versus-transfer tradeoff can be resolved by progressive specialization through network depth: shared cross-modal representations in shallow layers, task-specific but modality-shared prompts in middle layers, and fully task-and-modality-specific prompts in deep layers. On four tasks (AVE, AVVP, AVS, AVQA), PHP reports the best anti-forgetting scores among the compared methods and is the only one showing positive multi-task transfer, quantified as $+7.79\\%$ on the paper's Diff metric.","feed_headline":"Three-stage prompts turn audio-visual forgetting into +7.79% transfer","feed_subtitle":"A shallow-to-deep prompt hierarchy is the only tested method whose multi-task learning beats single-task training.","key_machinery":"The load-bearing mechanism is the progressive ordering of three prompt components across the depth of frozen CLIP and CLAP backbones, expressed as the shallow-middle-deep (S-M-D) principle: Task-shared Modality Aggregating (TMA) adapters at shallow layers perform channel, spatial, and temporal cross-modal attention that is shared by all tasks; Task-specific Modality-shared Dynamic Generating (TMDG) adapters at middle layers select and generate instance-level prompts from a learned prompt pool via self-attention, keeping prompts task-specific but consistent across audio and video; Task-specific Modality-Independent (TMI) prompts at deep layers attach separate per-task, per-modality tokens. The transfer claim itself is carried by the paper's bespoke metric Diff, defined as $\\frac{A_{\\mathrm{multi}}-A_{\\mathrm{single}}}{100-A_{\\mathrm{single}}}\\,(1+\\tfrac{A_{\\mathrm{single}}}{100})^2\\times100\\%$, which penalizes high single-task baselines quadratically.","core_discovery":"The paper's claim is that catastrophic forgetting and positive transfer are not opposing forces but consequences of where in the network knowledge is stored. PHP places a task-shared modality aggregating (TMA) adapter at shallow layers to build universal audio-visual correspondences, a task-specific modality-shared dynamic generating (TMDG) adapter at middle layers that synthesizes instance-aware prompts from a prompt pool, and task-specific modality-independent (TMI) prompts at deep layers that preserve each task's and each modality's fine details. Because shallow knowledge is shared, later tasks can benefit from earlier ones, while deep task-specific prompts protect old-task details from being overwritten. The experiments report state-of-the-art accuracy and the lowest forgetting among fine-tuning, EWC, L2P, S-prompt, DualPrompt, PC, and DCNet, together with the only positive transfer score, $+7.79\\%$.","pith_inferences":["If the shared-to-specific depth ordering is the real cause of positive transfer, the same S-M-D recipe (shared adapters shallow, task-specific prompts deep) is worth testing on other multi-modal continual-learning streams, such as video-language or audio-language task sequences.","The Diff metric is a substantive proposal for judging transfer when baselines differ; applying it retroactively to published continual-learning results would reveal whether it changes the ranking of well-known methods, a check the paper does not perform.","Because prompt selection here assumes the task index is known at inference time, a natural extension is to let the prompt pool infer task identity from the instance itself, which would make the method usable in task-agnostic streams.","The ablation showing TMA alone delivers $+8.05\\%$ transfer suggests a minimal PHP variant (shared shallow adapter plus deep TMI prompts, without the middle adapter) might retain most of the benefit at lower parameter cost, a testable simplification."],"forward_implications":["A single frozen backbone pair (CLIP plus CLAP) can serve four different audio-visual tasks presented sequentially, with accuracy on later tasks meeting or exceeding single-task training on some tasks.","The shallow shared adapter carries most of the transfer: in the ablations, TMA alone yields $+8.05\\%$ Diff, and the full S-M-D ordering beats every alternative ordering of the three components on both forgetting and transfer.","Prompt length and the depth of task-specific layers are tunable levers with opposing effects: longer prompts resist forgetting, while excessive task-specific parameterization degrades cross-task transfer.","The method extends to four-task sequences in different orders, holding first-task accuracy stable while absorbing three later tasks.","Across all compared baselines, PHP is the only method whose multi-task average exceeds its single-task average, which is the paper's evidence that incremental training can be a positive rather than negative influence."],"supporting_citations":[{"why":"Supplies the frozen CLIP visual backbone into which all three prompt stages are injected.","marker":"[45]"},{"why":"Supplies the frozen CLAP audio backbone that pairs with CLIP for the audio-visual stream.","marker":"[15]"},{"why":"L2P establishes the rehearsal-free prompt-pool paradigm that PHP extends to audio-visual multi-task streams.","marker":"[62]"},{"why":"DualPrompt's task-agnostic/task-specific prompt split is the closest baseline and the design PHP generalizes into three stages.","marker":"[61]"},{"why":"S-Prompt's independent task-specific prompts are the baseline that PHP distinguishes from by keeping prompts modality-shared in the middle stage.","marker":"[60]"},{"why":"DG-SCT provides the channel, spatial, and temporal cross-modal attention structure that the TMA adapter adopts.","marker":"[13]"},{"why":"The knowledge-circuits analysis of continual pre-training motivates the shallow-to-deep shared-to-specific progression.","marker":"[41]"},{"why":"EWC is the regularization baseline whose anti-forgetting behavior defines the comparison in Table 1.","marker":"[24]"},{"why":"PC generates instance-specific prompts from a codebook, the direct predecessor of PHP's dynamic prompt generation.","marker":"[9]"}],"fun_headline_variants":["Only method with positive transfer in audio-visual incremental learning","Three-stage prompts: beat forgetting, boost transfer by 7.79%","Shallow-to-deep prompts give +7.79% transfer in AV incremental tasks","Homeostatic-plastic prompts: only positive transfer across AV tasks","Prompt hierarchy turns AV forgetting into +7.79% transfer"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline transfer result rests on the paper's own Diff metric, whose quadratic penalty on already-strong single-task baselines is a design choice; if that metric is not accepted, the $+7.79\\%$ positive-transfer claim has no standard external benchmark behind it.","fun_headline_variants_meta":{"raw":{"variants":["Only method with positive transfer in audio-visual incremental learning","Three-stage prompts: beat forgetting, boost transfer by 7.79%","Shallow-to-deep prompts give +7.79% transfer in AV incremental tasks","Homeostatic-plastic prompts: only positive transfer across AV tasks","Prompt hierarchy turns AV forgetting into +7.79% transfer"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000322,"raw_usage":{"total_tokens":1822,"prompt_tokens":965,"completion_tokens":857,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":581,"completion_tokens_details":{"reasoning_tokens":764}},"tokens_in":581,"tokens_out":857,"duration_ms":8790,"temperature":1.0,"reasoning_tokens":764,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T12:35:47.098574+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute transfer for every method in Table 2 using a pre-existing standard measure, such as the plain average gain $\\bar{A}_{\\mathrm{multi}}-\\bar{A}_{\\mathrm{single}}$ or the mean per-task relative improvement. If under a standard metric PHP no longer shows positive transfer while all baselines stay negative, the paper's distinctive claim fails; if several baselines also turn positive, the claim that PHP is the only transferring method fails.","supporting_citations":[{"cited_title":"Clap learning audio concepts from nat- ural language supervision","cited_arxiv_id":null,"evidence_quote":"Supplies the frozen CLAP audio backbone that pairs with CLIP for the audio-visual stream."},{"cited_title":"Learning to prompt for con- tinual learning","cited_arxiv_id":null,"evidence_quote":"L2P establishes the rehearsal-free prompt-pool paradigm that PHP extends to audio-visual multi-task streams."},{"cited_title":"Dualprompt: Complementary prompting for rehearsal-free continual learning","cited_arxiv_id":null,"evidence_quote":"DualPrompt's task-agnostic/task-specific prompt split is the closest baseline and the design PHP generalizes into three stages."},{"cited_title":"S-prompts learning with pre-trained transformers: An occam’s razor for domain incremental learning","cited_arxiv_id":null,"evidence_quote":"S-Prompt's independent task-specific prompts are the baseline that PHP distinguishes from by keeping prompts modality-shared in the middle stage."},{"cited_title":"Cross-modal prompts: Adapting large pre- trained models for audio-visual downstream tasks","cited_arxiv_id":null,"evidence_quote":"DG-SCT provides the channel, spatial, and temporal cross-modal attention structure that the TMA adapter adopts."},{"cited_title":"Overcoming catastrophic forgetting in neu- ral networks","cited_arxiv_id":null,"evidence_quote":"EWC is the regularization baseline whose anti-forgetting behavior defines the comparison in Table 1."},{"cited_title":"Prompt Customization for Continual Learning","cited_arxiv_id":"2404.18060","evidence_quote":"PC generates instance-specific prompts from a codebook, the direct predecessor of PHP's dynamic prompt generation."}],"review_version":1}