{"id":"0d4ad69b-9273-47a2-9588-5edf977be383","arxiv_id":"2505.12966","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"An audio-visual deepfake detector combining adaptive contrastive learning, multi-scale fusion, and orthogonalized Pareto gradient balancing reports state-of-the-art accuracy (95.5% average) and strong cross-dataset transfer on three benchmarks.","lead":"This paper proposes MACB-DF, an audio-visual deepfake detector that balances video and audio information using contrastive learning and a Pareto-style gradient balancing module. It reports top accuracy on three benchmarks and larger gains when trained on one dataset and tested on another.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline cross-dataset gains (8.0%/7.7%) contradict Table 3 (6.8%/6.4%), so the central generalization claim is not yet supported.","rationale":"The strongest claim is quantitative: state-of-the-art accuracy and 8.0/7.7 percentage-point cross-dataset improvements. The condition for that claim to hold is that the reported numbers are accurate and internally consistent. That condition is least secure in the comparison between the Abstract and Table 3. The Abstract promises 8.0/7.7 improvements; Table 3's own numbers yield 6.8/6.4. Since the table is the only evidence for cross-dataset performance, a reader cannot tell whether the headline is an overstatement or the table an understatement. This is not an internal inconsistency in the method's math, but it is a direct inconsistency in the evidence for the central claim. The reader's weakest assumption about Eq. (17) is also legitimate: the fusion rule assumes comparable shapes and semantics without stating projections or normalization. However, even if Eq. (17) were fully specified, the unresolved numerical contradiction would still block acceptance. I therefore keep the CONDITIONAL verdict. The check is straightforward and should settle the issue.","tokens_in":13382,"tokens_out":3909,"duration_ms":38183,"concrete_test":"Obtain the authors' released model or predictions (or re-run their training/evaluation code) and recompute the DFDC-to-DefakeAVMiT and DFDC-to-FakeAVCeleb ACC margins against the exact baseline protocol of Table 3. If the recomputed margins are 6.8/6.4, correct the Abstract; if they are 8.0/7.7, correct Table 3. A minimal analytical check that does not require code is to recompute 91.2 - 84.4 and 89.2 - 82.8 and verify whether any reported baseline in the paper gives the Abstract's claimed margins.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim is that MACB-DF yields \"absolute improvements of 8.0% and 7.7% in ACC scores over the previous best-performing approach when trained on DFDC and tested on DefakeAVMiT and FakeAVCeleb\" (Abstract). The only cross-dataset table, Table 3, reports MACB-DF at 91.2% and 89.2% against AVoiD-DF at 84.4% and 82.8%, which gives margins of 6.8 and 6.4 percentage points, not 8.0 and 7.7. No other baseline in the paper yields 8.0/7.7 margins. Thus the headline generalization claim is not supported by the paper's own data; at least one of the two sets of numbers is wrong. This is load-bearing because the state-of-the-art and generalization statements rest on these margins. The absence of code, error bars, and the \"a%\" placeholder in Section 5.1 further prevent the reader from determining which number is correct. The concern is not misconduct; a typo in either the Abstract or Table 3 would explain it, but the contradiction must be resolved before the claim can be accepted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MACB-DF, a multimodal audio-visual deepfake detection framework combining adaptive contrastive learning (MACL), multi-scale feature fusion with learnable fusion weights, a representation fusion module (RFMF), and an orthogonalization-multimodal Pareto module (OM-Pareto). The authors claim consistent state-of-the-art performance on DefakeAVMiT, FakeAVCeleb, and DFDC, with 95.5% average accuracy, and report cross-dataset gains of 8.0% and 7.7% ACC in the abstract when training on DFDC and testing on DefakeAVMiT and FakeAVCeleb. Section 6 presents the same cross-dataset experiment with gains of 6.8% and 6.4%, creating a direct internal contradiction in the paper's headline claim. Extensive ablations are presented for the contrastive losses, the fusion weighting, and the OM-Pareto module.","tokens_in":13681,"tokens_out":4272,"duration_ms":41057,"significance":"If the claimed results are reproducible, MACB-DF would be a meaningful step for audio-visual deepfake detection, particularly its cross-dataset generalization, which is more practically relevant than intra-dataset accuracy. The paper's strengths are the broad ablation coverage (contrastive losses, fusion weighting, OM-Pareto), the evaluation on three benchmarks, and the explicit formulation of the training objectives. However, the central quantitative claim is internally inconsistent, and no code, seeds, or error bars are provided, so the current evidence is not sufficient to confirm the stated state-of-the-art performance.","major_comments":[{"comment":"The headline cross-dataset claim is internally inconsistent: the Abstract and Conclusion state absolute ACC improvements of 8.0% and 7.7% over AVoiD-DF when trained on DFDC, but Table 3 and the Section 6 text report margins of 6.8% and 6.4% (91.2−84.4=6.8; 89.2−82.8=6.4). Because the state-of-the-art and generalization claims rest on these margins, the authors must reconcile the two sets of numbers, report the correct values in all locations, and state whether Table 3 or the abstract was based on a different evaluation protocol.","section":"Abstract / Section 6 (Table 3)"},{"comment":"The weighted fusion x_fused_i = w_v_i * V_{i-1} + w_a_i * A_{i-1} presupposes that video and audio feature maps have identical shape and comparable scales, but no projection, normalization, or dimension-matching operation is specified for V_{i-1} and A_{i-1}. Similarly, the clustering in Eq. (9) requires the cluster count K and per-cluster covariances Sigma_k at inference time, yet the manuscript does not state how K is chosen or how Sigma_k is estimated on a test batch. Without these details, the fusion and contrastive-alignment pipeline is not fully specified.","section":"Section 3.2.4, Eq. (17)"},{"comment":"The OM-Pareto module is central to the claim of resolving gradient conflicts, but the paper does not provide the actual optimization procedure: Eq. (21) is a constrained optimization problem, yet the text only says 'Execute complete PGD iterations' without giving the projection operator, step size, or termination criterion, and Eq. (22) introduces lambda_orth without specifying how it is incorporated into the final update. The reader cannot reproduce or verify the claimed conflict resolution from the text.","section":"Section 3.5, Eqs. (21)-(24)"},{"comment":"The experimental reporting is not sufficient to support the quantitative point estimates: no error bars, number of seeds, or standard deviations are reported for any table, and Section 5.1 contains an unfinished placeholder 'amplify decision boundaries by a%'. Since the paper's central claims are quantitative, the authors should provide multi-seed results, variance measures, complete text for the ablation analysis, and a clear statement of code or implementation availability.","section":"Section 5.1 / Tables 1-2"}],"minor_comments":[{"comment":"There are several typographical issues: 'INTRODDUCTION' in Section 1, 'deptly balances' in the Figure 1 caption, and 'FakeA VCeleb' in Section 4.1; these should be corrected.","section":"Throughout"},{"comment":"The caption of Table 3 is malformed: 'THE TRAINING SETS AND TESTING SETS FOR CROSS-DATASET' appears to be a fragment of a sentence rather than a proper caption, and the table does not explicitly state that the rows are methods trained on DFDC.","section":"Table 3"},{"comment":"The LAV-DF dataset is described in Section 4.1 but never used in any experiment or ablation; the authors should either include results on it or explain why it is omitted.","section":"Section 4.1"},{"comment":"The ablation discussion states that contrastive learning 'amplifies decision boundaries by a% compared to baseline models'; this is an unfinished placeholder and must be replaced with the actual measured value or a precise description.","section":"Section 5.1"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the numerical inconsistency between the abstract (8.0%/7.7%) and Section 6 (6.8%/6.4%) is the most serious issue and should be resolved before any acceptance decision. I would also recommend requesting the authors' code or a detailed implementation appendix, because the method has many free parameters (beta_k, gamma_t, lambda_0, kappa, margin m, and K) and the current text does not permit verification of the reported numbers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. The architecture is sensible and the ablations are honest, but the paper's central generalization claim is not currently supported by its own table. The abstract says 8.0/7.7-point cross-dataset gains; Table 3 shows 6.8/6.4. That contradiction needs resolving before anyone quotes those numbers. Also, no code or error bars, and one ablation line contains a literal \"a%\" placeholder.\n\nNow the credit. The method is more than a parameter scan: adaptive temperature in contrastive learning, large-kernel multi-scale fusion blocks, and an orthogonalization-Pareto gradient balancing module. The ablations in Table 2 and Figures 6-7 give internal support that each component contributes. The idea to test cross-dataset generalization from DFDC to DefakeAVMiT and FakeAVCeleb is exactly the right bottleneck, and the gains over AVoiD-DF are large enough to matter. The related work on modality imbalance and gradient conflicts is appropriately cited.\n\nSoft spots, in proportion. The abstract/table mismatch is the big one. It is likely a typo, but it is load-bearing because the state-of-the-art and \"significant improvements\" claims rest on it. The fusion rule in Eq. 17 assumes video and audio features have the same shape and live in a comparable scale space; the paper never states the projection or normalization that makes the weighted sum valid, nor how the cluster count and covariances in Eq. 9 are set at inference. The OM-Pareto hyperparameters and the negative-count K are listed but not analyzed. No error bars, no seeds, no code. These are all fixable in a revision, but they are real gaps in reproducibility.\n\nWho is this for? People working on audio-visual deepfake detection and multimodal learning under modality imbalance. It is an incremental contribution in a crowded field, but a plausible one. Give it a serious referee, but ask the authors to resolve the number discrepancy, fill in the placeholder, and release code or at least detailed training settings. If the corrected numbers hold, it deserves to be cited.","headline":"A plausible incremental architecture with a genuine cross-dataset story, undercut by an internal number discrepancy and missing reproducibility details.","tokens_in":14214,"tokens_out":1872,"would_cite":false,"duration_ms":18595,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes MACB-DF, a contrastive, conflict-balancing audio-visual fusion method that reports state-of-the-art deepfake detection across three benchmarks and improved cross-dataset generalization.","keywords":["multimedia machine learning","video-audio deepfake","multimodal fusion","contrastive learning","deepfake detection","audio-visual forgery","gradient conflict","pareto optimization"],"falsifier":"A reader can settle the cross-dataset claim by re-running the DFDC-to-DefakeAVMiT and DFDC-to-FakeAVCeleb experiments under the paper's 80/20 split and checking whether ACC reproduces at 91.2% and 89.2%; a second check is to replace the weighted feature sum in Eq. (17) with concatenation plus a linear layer and see whether accuracy survives.","tokens_in":13151,"feed_emoji":"🎭","tokens_out":11493,"duration_ms":98859,"temperature":0.7,"pith_summary":"MACB-DF is a proposed audio-visual deepfake detector built on the idea that the two modalities must be balanced, not just combined. The paper argues that standard multimodal fusion lets one modality dominate, and that the separate losses for audio, video, and fused streams create conflicting gradients during backpropagation. To fix this, the method aligns audio and video with adaptive contrastive learning, uses clustering information to set per-modality fusion weights, and adds an orthogonalization-multimodal Pareto module that resolves gradient conflicts through orthogonal regularization. The paper reports 96.8% ACC on DefakeAVMiT, 91.7% on FakeAVCeleb, 97.9% on DFDC, and average 95.5% accuracy, together with cross-dataset gains over the previous best method when training on DFDC and testing on the other two datasets. If these numbers hold, the method is the strongest evidence that modality-balancing mechanisms, rather than more expressive fusion alone, drive audio-visual deepfake detection performance.","feed_headline":"AV deepfake detector hits 95.5% accuracy by balancing modalities","feed_subtitle":"Training on DFDC then testing on two unseen sets beats prior audio-visual methods by up to 8 accuracy points.","key_machinery":"The load-bearing mechanism is the multi-scale fusion loop built from three interacting components. First, Multi-modal Adaptive Contrast Learning (MACL) aligns video and audio embeddings with an adaptive temperature derived from variance, weighted skewness, and attention entropy, and additionally clusters embeddings using a composite Mahalanobis-cosine distance to produce each modality's fusion weight $w_{m_i}$. Second, the weighted fusion $x^{\\text{fused}}_i = w_{v_i}V_{i-1} + w_{a_i}A_{i-1}$ combines deep video and audio features, and the fused representation is multiplied back into both streams at successive layers, so each modality learns from the joint signal. Third, the Orthogonalization-multimodal Pareto module solves $\\min_{\\alpha_m,\\alpha_u}\\|\\alpha_m g_m+\\alpha_u g_u\\|^2$ under a simplex constraint, adding an orthogonality penalty whose strength is modulated by cosine similarity between the unimodal and multimodal gradients; this is the component that is supposed to stop one modality's gradient from overriding the other.","core_discovery":"The central claim is that the main obstacle to accurate audio-visual deepfake detection is not the lack of fusion capacity but the imbalance and conflict between modalities during training. The paper's MACB-DF pipeline extracts spatiotemporal video features and log Mel-spectrogram audio features, then applies multi-modal adaptive contrast learning (MACL) that aligns positive audio-video pairs, separates negatives with a margin, and uses an adaptive temperature to control gradient scale. Clustering in the aligned space produces composite distances, Mahalanobis and cosine, that feed per-modality fusion weights, which are then used in a weighted sum of deep features; the fused representation is re-injected into both streams at multiple scales. A final orthogonalization-multimodal Pareto module treats the joint update from unimodal and multimodal gradients as a constrained minimization and applies orthogonal regularization when the gradients point in conflicting directions. The paper claims this design yields state-of-the-art accuracy on DefakeAVMiT, FakeAVCeleb, and DFDC, and better cross-dataset generalization when trained on DFDC.","pith_inferences":["If Eq. (17)'s missing projection is filled in, the same fusion and Pareto machinery could be lifted into other audio-visual tasks, such as fake speech detection or speaker-video verification, where modality imbalance is also reported.","The discrepancy between the abstract's 8.0/7.7 percentage-point cross-dataset gains and Table 3's 6.8/6.4-point values means the exact size of the claimed generalization improvement is not yet stable; a reader should treat the table as the conservative number until reproduced.","A direct test of the balancing story would be to feed deliberately corrupted audio or video during inference and check whether the fusion weights $w_{v_i}, w_{a_i}$ shift to down-weight the corrupted modality; if they do not, the method balances losses during training but does not adapt at test time.","The contrastive clustering step could be evaluated as a standalone module on unimodal deepfake benchmarks; if it improves unimodal accuracy too, its benefit is not specific to cross-modal fusion."],"forward_implications":["If the reported 95.5% average accuracy is reproducible, MACB-DF would be the best-performing audio-visual deepfake detector on these three benchmarks, and the cross-dataset results imply detectors trained on DFDC can transfer to unseen forgery datasets with only a modest drop.","The ablations attribute clear accuracy gains to each component: removing contrastive learning drops ACC by 3.2 points on FakeAVCeleb and 4.7 points on DFDC, removing intra-modal losses costs 1.0 to 2.0 AUC points, and removing adaptive fusion weights costs 0.6 ACC points on DFDC.","Because the Pareto module's main observed benefit is faster and more stable early convergence, the method implies that addressing gradient conflict is a training-dynamics problem as much as a final-performance problem.","The multi-scale re-injection of the fused feature into both unimodal streams, Eq. (18), suggests that balanced multimodal learning requires each modality to see the joint representation, not just the final classifier."],"supporting_citations":[{"why":"Supplies the DefakeAVMiT dataset and the AVoiD-DF audio-visual joint learning baseline that MACB-DF claims to outperform in both intra-dataset and cross-dataset tests.","marker":"[40]"},{"why":"Supplies the DFDC benchmark used for the third intra-dataset evaluation and as the cross-dataset training set.","marker":"[10]"},{"why":"Supplies the FakeAVCeleb benchmark used for intra-dataset testing and as a target for cross-dataset generalization.","marker":"[18]"},{"why":"Provides the MCL multimodal contrastive learning baseline whose DFDC accuracy of 97.5% is the closest prior result the paper aims to beat.","marker":"[23]"},{"why":"Introduces the LAV-DF dataset and the BA-TFD multimodal baseline used in the comparison.","marker":"[3]"},{"why":"Supplies the AASIST audio-only anti-spoofing baseline that represents the audio modality in the comparison table.","marker":"[17]"},{"why":"Supplies the ECAPA-TDNN audio-only baseline used in the comparison table.","marker":"[9]"},{"why":"Supplies the Multiple-Attention visual-only baseline used in the comparison table.","marker":"[43]"}],"fun_headline_variants":["Balancing modalities boosts deepfake detection to 95.5%","Audio-visual balancing beats deepfakes with 8% cross-dataset gain","MACB-DF: Conflict-balanced audio-visual deepfake detector","Cross-dataset deepfake accuracy jumps 8% via modality balance","95.5% accuracy: Adaptive conflict balancing for deepfakes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the video and audio feature tensors can be added together after being multiplied by learned weights; the paper never explains what projection or normalization makes that addition meaningful.","fun_headline_variants_meta":{"raw":{"variants":["Balancing modalities boosts deepfake detection to 95.5%","Audio-visual balancing beats deepfakes with 8% cross-dataset gain","MACB-DF: Conflict-balanced audio-visual deepfake detector","Cross-dataset deepfake accuracy jumps 8% via modality balance","95.5% accuracy: Adaptive conflict balancing for deepfakes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1519,"prompt_tokens":957,"completion_tokens":562,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":468}},"tokens_in":573,"tokens_out":562,"duration_ms":5606,"temperature":1.0,"reasoning_tokens":468,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:22:56.629184+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader can settle the cross-dataset claim by re-running the DFDC-to-DefakeAVMiT and DFDC-to-FakeAVCeleb experiments under the paper's 80/20 split and checking whether ACC reproduces at 91.2% and 89.2%; a second check is to replace the weighted feature sum in Eq. (17) with concatenation plus a linear layer and see whether accuracy survives.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the LAV-DF dataset and the BA-TFD multimodal baseline used in the comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the AASIST audio-only anti-spoofing baseline that represents the audio modality in the comparison table."},{"cited_title":"Multi-attentional Deepfake Detection","cited_arxiv_id":"2103.02406","evidence_quote":"Supplies the Multiple-Attention visual-only baseline used in the comparison table."}],"review_version":1}