{"id":"8ab685b6-057e-4dc8-bcfe-6ccba14ca800","arxiv_id":"2605.14291","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"MMGuard generates unlearnable multimodal examples via perturbations that exploit LVLM optimization shortcuts and disrupt cross-modal bindings, providing robust protection against unauthorized fine-tuning across threat models.","lead":"The paper introduces MMGuard, a technique that adds tiny invisible perturbations to images and text so large vision-language models overfit to the noise and fail to learn the real content during unauthorized fine-tuning. Data owners could use this proactive method to block scraping and training on their multimodal data without relying on after-the-fact fixes.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Cross-modal binding disruption's 'theoretical guarantees' rest on unverified assumptions about attention dynamics that may not hold across LVLM architectures.","rationale":"The reader's weakest_assumption directly matches the unproven attention-behavior assumption required for the binding disruption to deliver the stated theoretical guarantees. Full-text access does not remove this gap; the claim remains conditional on that assumption holding.","tokens_in":1767,"tokens_out":299,"duration_ms":10446,"concrete_test":"Ablate the cross-modal binding disruption on one white-box model (e.g., LLaVA-1.5) while keeping the base perturbation fixed; measure downstream task accuracy drop on the six datasets. If accuracy recovers by >15% relative to the full MMGuard version, the binding component is not providing the claimed mechanistic advantage.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central mechanism claims that minimizing training loss creates an optimization shortcut while cross-modal binding disruption enforces spurious noise-target correlations 'with theoretical guarantees.' This requires the LVLM attention and loss landscape to behave predictably under the chosen perturbation for all nine evaluated models. The abstract provides no derivation or section reference showing how the guarantee follows from the perturbation strategy; if attention heads in models like LLaVA or InstructBLIP route around the injected noise differently than assumed, the spurious correlation fails and downstream degradation does not occur. Evaluation results alone cannot confirm the mechanism without isolating the binding term.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes MMGuard, a proactive method to protect multimodal data from unauthorized LVLM fine-tuning. It generates unlearnable examples via human-imperceptible perturbations that create an optimization shortcut (forcing overfitting to noise) and a cross-modal binding disruption that enforces spurious noise-target correlations with theoretical guarantees. An ensemble strategy improves cross-model transferability. The approach is evaluated on nine open-source LVLMs across six datasets under white-box, gray-box, and black-box threat models, claiming effective, stealthy, and robust protection.","tokens_in":1899,"tokens_out":476,"duration_ms":20111,"significance":"If the mechanistic claims and theoretical guarantees hold after verification, this would offer a meaningful shift from post-hoc defenses (unlearning, watermarks) to proactive data protection, with potential impact on copyright and privacy practices for web-scraped multimodal training data.","major_comments":[{"comment":"Abstract: The claim that cross-modal binding disruption 'enforces a spurious correlation between the noise and the training target with theoretical guarantees' is load-bearing for the robustness argument across nine architectures, yet no derivation, equation, or section is referenced showing how the guarantee follows from the perturbation strategy; if attention heads route around the noise differently than assumed, the downstream degradation fails.","section":"Abstract"},{"comment":"Evaluation section: Results on nine models and six datasets are presented without ablation studies isolating the binding term, error analysis, or mechanistic verification that the optimization shortcut (rather than other factors) drives the observed protection; this leaves the central claim of a 'mechanistic advantage' unconfirmed by the reported experiments alone.","section":"Evaluation"}],"minor_comments":[{"comment":"The abstract refers to 'human-imperceptible perturbations' and 'ensemble learning strategy' without specifying the perturbation generation algorithm, imperceptibility metrics (e.g., PSNR/SSIM bounds), or ensemble construction details; these should be clarified for reproducibility.","section":null}],"recommendation":"major_revision","confidential_remarks":"The absence of any equations or formal derivation for the 'theoretical guarantees' in the abstract raises a scope concern for a cryptography/security venue; the manuscript may fit better after adding a dedicated theory section or interpretability experiments."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on our manuscript. We address each major comment below and will revise the manuscript to improve clarity on the theoretical claims and strengthen the experimental validation of the mechanistic components.","responses":[{"response":"We appreciate the referee identifying this gap in referencing. The derivation of the spurious correlation is formalized in Section 3.3, where we model the attention routing under the perturbation and prove that the noise-target correlation is enforced via a bound on the attention weights (Equation 7). To make this explicit and address potential routing variations, we will revise the abstract to directly cite Section 3.3 and Equation 7, expand the discussion in Section 3.3 on the assumption robustness, and add a short proof sketch in the main text.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The claim that cross-modal binding disruption 'enforces a spurious correlation between the noise and the training target with theoretical guarantees' is load-bearing for the robustness argument across nine architectures, yet no derivation, equation, or section is referenced showing how the guarantee follows from the perturbation strategy; if attention heads route around the noise differently than assumed, the downstream degradation fails."},{"response":"We agree that additional ablations and mechanistic checks would strengthen the evaluation. In the revision, we will add ablation studies comparing the full MMGuard against variants without the binding disruption term across the nine models. We will also include error analysis (e.g., per-dataset variance and failure cases) and mechanistic verification via attention map visualizations and training loss curves to isolate the optimization shortcut's contribution. These will be placed in a new subsection of the evaluation.","revision_made":"yes","referee_comment":"[Evaluation] Evaluation section: Results on nine models and six datasets are presented without ablation studies isolating the binding term, error analysis, or mechanistic verification that the optimization shortcut (rather than other factors) drives the observed protection; this leaves the central claim of a 'mechanistic advantage' unconfirmed by the reported experiments alone."}],"tokens_in":1369,"tokens_out":445,"duration_ms":18850,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this paper introduces MMGuard to let data owners poison multimodal examples so LVLMs overfit to noise instead of content during unauthorized fine-tuning, with an added cross-modal binding step to force spurious noise-target links and an ensemble trick for transfer across models.","headline":"MMGuard adds cross-modal binding disruption to unlearnable examples for LVLMs and shows broad empirical coverage, but the theoretical guarantees rest on unverified attention assumptions.","tokens_in":2383,"tokens_out":136,"would_cite":false,"duration_ms":24681,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"MMGuard's bilevel optimization and cross-modal binding disruption (Theorem 5.1) operate in AI security with no structural overlap to RS cost forcing or distinction-derived constants.","alignment":"orthogonal","rationale":"The paper's core machinery (PGD image perturbations, HotFlip text triggers, attention-mass KL losses ℓmass, BPH/CRS variants, ensemble surrogates) is standard adversarial ML. It never references J-cost, φ-ladders, 8-tick periodicity, or any RS forcing theorem. Domain mismatch (cs.CR vs. foundational physics) confirms orthogonality.","tokens_in":60391,"confidence":"high","tokens_out":142,"duration_ms":9523,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Data owners can add invisible perturbations to images and text to stop large vision-language models from learning real content during unauthorized fine-tuning.","keywords":["unlearnable examples","multimodal data protection","large vision-language models","unauthorized fine-tuning","perturbation injection","cross-modal disruption","proactive defense","LVLM security"],"falsifier":"Fine-tune an LVLM on MMGuard-protected data, then measure its accuracy on clean test data without the perturbations and compare to accuracy after fine-tuning on the same data without protection.","tokens_in":2665,"feed_emoji":"🛡️","tokens_out":712,"duration_ms":23927,"temperature":0.7,"pith_summary":"The paper introduces MMGuard as a proactive defense that lets owners protect multimodal data before it is scraped and used for fine-tuning large vision-language models. It injects tiny perturbations that create an optimization shortcut during training, so the model overfits to the noise and loses performance when the noise is absent at test time. A cross-modal binding disruption further shifts attention to tie the noise to the training targets, backed by theoretical guarantees and improved by an ensemble strategy for transfer across models. This is shown to work against nine open-source LVLMs on six datasets in white-box, gray-box, and black-box settings. A sympathetic reader would care because it moves protection upstream, before infringement happens, unlike post-hoc unlearning or watermarking.","feed_headline":"Hidden noise stops vision-language models learning real content","feed_subtitle":"MMGuard adds imperceptible perturbations so unauthorized fine-tuning overfits to noise and fails on clean data.","key_machinery":"Perturbation injection that minimizes training loss via an optimization shortcut, paired with cross-modal binding disruption to enforce noise-target spurious correlations.","core_discovery":"MMGuard generates unlearnable examples by injecting human-imperceptible perturbations that exploit LVLM learning dynamics to create an optimization shortcut, causing the model to overfit to noise rather than content and degrading downstream performance when the perturbation is removed. It adds cross-modal binding disruption to strategically shift attention and enforce spurious correlations between the noise and training targets with theoretical guarantees, then uses an ensemble learning strategy to boost cross-model transferability, with evaluations showing effective, stealthy, and robust protection under multiple threat models.","pith_inferences":["If widely used, this could make scraped web data less useful for training, potentially reducing incentives for unauthorized collection.","Similar perturbation approaches might extend to protecting data for other model types or single-modality tasks.","Training pipelines might need built-in checks to detect or handle such protected data.","Over time, routine use could shift norms around public multimodal data availability."],"forward_implications":["Fine-tuning on protected data produces models that perform poorly on clean inputs at inference time.","The defense holds under white-box, gray-box, and black-box threat models.","Ensemble perturbations enable the protection to transfer across different LVLM architectures.","Owners gain a tool that acts before data is scraped, reducing reliance on after-the-fact remedies.","Protection remains stealthy to humans while disrupting model learning."],"fun_headline_variants":["Noise forces LVLMs to overfit to perturbations","Hidden noise blocks LVLM learning of real content","MMGuard makes data unlearnable via perturbations","Perturbations shift attention to noise in LVLMs"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The injected perturbations will reliably force the model to overfit to noise instead of content because the attention and loss landscape behave predictably under the chosen strategy.","fun_headline_variants_meta":{"raw":{"variants":["Noise forces LVLMs to overfit to perturbations","Hidden noise blocks LVLM learning of real content","MMGuard makes data unlearnable via perturbations","Perturbations shift attention to noise in LVLMs"]},"model":"grok-4.3","cost_usd":0.010325,"raw_usage":{"total_tokens":4508,"prompt_tokens":702,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":103253000,"prompt_tokens_details":{"text_tokens":702,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3746,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":702,"tokens_out":60,"duration_ms":31633,"temperature":1.0,"reasoning_tokens":3746,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-15T02:35:56.750302+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Fine-tune an LVLM on MMGuard-protected data, then measure its accuracy on clean test data without the perturbations and compare to accuracy after fine-tuning on the same data without protection.","supporting_citations":[],"review_version":1}