{"id":"45657d46-e9fb-491a-9dc2-eb9b46b1d8ac","arxiv_id":"2506.12430","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"The ATLAS 2025 competition demonstrates that vision-language models remain highly vulnerable to flowchart-based and cross-modal jailbreak attacks, with top scores exceeding 93%.","lead":"A safety competition report shows that multimodal AI models can be tricked into producing harmful content by hiding malicious questions inside flowcharts and role-playing prompts. Top teams achieved over 90% attack success, and the organizers release the competition data, but the evaluation depends on an unnamed AI judge and self-reported results.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Benchmark scores depend on an unnamed single-LLM judge with no human validation; until judge identity/version and human agreement are reported, the near-perfect ASR claims in Tables 2-3 and Section 4.1 lack verified ground truth.","rationale":"The reader's weakest assumption matches the most load-bearing concern. The quantitative heart of the report is ASR as computed by Eq. (2), and the report gives no details of the judging model J. The only robustness discussion is a brief admission in Section 5.1. The team case studies reveal divergent self-evaluation criteria, so the judge may not even be constant across reported numbers. A biased judge could simultaneously inflate all scores—for example, by treating any structured output as relevant or by flagging any response containing a harmful phrase, regardless of context. This would directly undermine the near-perfect claims and the leaderboard rankings. This is not purely hypothetical: Section 4.4.3 notes that teams modified the provided queries to emphasize harmful intent, and Section 4.4.4 says the submitted Phase II solution scored slightly worse than some ablation variants, so the official numbers are already known to be fragile. However, the qualitative claim has independent support from prior work such as FigStep and from consistent reports across multiple teams, so a REJECT verdict would be too strong. CONDITIONAL remains the appropriate verdict, and the required conditions are clear: specify the judge, validate it against human annotations, and show that rankings are stable under judge choice. Since the reader already recommended CONDITIONAL, no change in verdict is needed.","tokens_in":23433,"tokens_out":3297,"duration_ms":40355,"concrete_test":"Take a stratified sample of about 200 judged outputs spanning the full score range (e.g., 50 from each quartile of Phase I/II scores) and the 12 risk categories. Have three independent human annotators re-label each output under the competition's definition (harmful and aligned with the provided prompt). Report the LLM judge's agreement with the majority human label, Cohen's kappa, and the confusion matrix; then recompute team scores and ranks using the human labels. Separately, run the same 200 outputs through at least two additional judge models (e.g., GPT-4o and Claude-3.5) with the exact prompt used in Section 2.4. If human-judge agreement is below about 0.8 or leaderboard Spearman rank correlation across judges is below about 0.9, the ASR claims are not quantitatively stable and the paper should present them only as qualitative findings.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claim—near-perfect jailbreak success and meaningful leaderboard rankings—rests on Eq. (2) in Section 2.4, where ASR is defined through an LLM judge J. The report never identifies J, its version, its prompt, or its threshold; Section 5.1 concedes that \"relying on a single LLM judge revealed certain limitations, particularly in assessing edge cases.\" Worse, the case studies show teams self-evaluated with different judges: Team 'You are the challenger' used Qwen2.5-latest with Harmful/Relevant scores >3 (Section 4.2.3), while MR-CAS used qwen-max (Section 4.4.4). If the official judge differs from these or is unstable, every ASR in Tables 2 and 3, including the Phase I top score 96.95 and Phase II top score 93.56, is not reproducible. The qualitative conclusion—flowcharts and cross-modal role-play can bypass current MLLM safety—has independent support from prior work (e.g., FigStep) and from the teams' local ablations, so it should not be rejected outright. But the benchmark-level claim requires a valid ground-truth label of harmfulness, and that is the load-bearing assumption least secured by the report.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This technical report describes ATLAS 2025, a two-phase adversarial competition in which 86 teams crafted image-text pairs to jailbreak multimodal large language models. Phase I was a white-box setting against Qwen2-VL-7B and InternVL2-8B using 180 SafeBench prompts; Phase II was a black-box setting with an additional target model and 150 more complex prompts. Attack success is defined in Eq. (2) through an unspecified \"LLM-as-a-Judge\" function J, and leaderboards are reported in Tables 2 and 3. The report presents case studies from four top teams, emphasizing structured flowcharts, role-playing prompts, OCR-based images, omission-completion tricks, and adversarial perturbations, and it argues that current safety-aligned MLLMs remain jailbreakable at high rates, with some local configurations approaching 100% ASR. Code and data are released publicly.","tokens_in":23634,"tokens_out":4932,"duration_ms":65567,"significance":"If the quantitative results are accepted, the report would provide a useful large-scale benchmark snapshot of MLLM jailbreak susceptibility, and the qualitative finding that structured visual diagrams and cross-modal role-playing bypass current safety filters is credible and consistent with prior work such as FigStep and typographic-prompt attacks. The public release of attack data and the multi-team diversity are strengths, and the team-level ablations (e.g., Tables 5, 6, 7, 9) give independent, method-level evidence for the effectiveness of flowcharts, red highlighting, and cross-modal intent distribution. However, the benchmark-level claims—near-perfect ASR and meaningful leaderboard rankings—are not yet established because the official judge is unidentified, no human validation or inter-annotator agreement is reported, no negative-control experiments on benign inputs are given, and no statistical uncertainty accompanies the scores. The paper is best read as a competition retrospective whose qualitative conclusions are defensible but whose headline quantitative claims require additional evaluation evidence.","major_comments":[{"comment":"The judge model J is never identified: no model name, version, prompt template, or decision threshold is given for the official evaluation. The case-study teams use different judges for their local evaluations—Qwen2.5-latest with Harmful/Relevant scores above 3 in Section 4.2.3 and qwen-max in Section 4.4.4. Unless the official judge is disclosed and made consistent across submissions, the ASR values in Tables 2 and 3, including the top Phase I score 96.95 and top Phase II score 93.56, are not reproducible or comparable. The authors should release the judge prompt, threshold, model version, and per-example judge outputs, or at minimum a code-and-data path that reproduces every leaderboard score.","section":"Section 2.4, Eq. (2)"},{"comment":"No human validation, inter-annotator agreement, or negative-control baseline is reported for the LLM judge. Section 5.1 itself concedes that \"relying on a single LLM judge revealed certain limitations, particularly in assessing edge cases.\" Without a human-annotated sample showing agreement with J, and without measuring J's false-positive rate on benign or refusal outputs, the claim that near-perfect ASR reflects true harmful content rather than judge leniency is unverified. This is load-bearing because Eq. (2) defines success entirely through J's binary verdict.","section":"Section 2.4 and Section 5.1"},{"comment":"The leaderboard scores are reported without confidence intervals or significance tests. With 180 prompts in Phase I and 150 in Phase II, the binomial standard error is roughly 1–4 percentage points depending on score level; for example, the 96.95 versus 95.56 gap between ranks 1 and 2 in Table 2 is within sampling error. Statements such as \"tight score margins among the top teams\" and the implied ranking meaningfulness are therefore not supported by the reported statistics. The authors should report per-model ASRs, confidence intervals, and pairwise comparisons, or explicitly frame the leaderboard as ordinal and not statistically resolved.","section":"Tables 2 and 3"},{"comment":"The official scoring is reported only as an average across target models, with no per-model breakdown in Tables 2 and 3. Since the paper's central claim is that current MLLMs are jailbroken at high rates, the reader needs to know whether the success is concentrated in one target model or shared across all three. The team-level ablations provide some per-model numbers, but the official benchmark scores do not, which weakens both the generalizability claim and any attempt to assess transferability in the black-box phase.","section":"Section 2.4 and Section 4.1"}],"minor_comments":[{"comment":"Section 2.3 states that Phase I used 180 prompts, while Section 4.4.1 refers to \"160 queries\" in the Phase I task; the discrepancy should be corrected or explained.","section":"Section 2.3 vs. Section 4.4.1"},{"comment":"Table 9 names DeepSeek-VL-7B-chat as the black-box target, but Section 2.3 only says \"an additional black-box MLLM\" without naming it; the main text should state the full target model set.","section":"Section 2.3 vs. Table 9"},{"comment":"The code block for the role-playing prompt is labeled \"PREFIX-BASED ATTACK\"; this label appears to be a copy-paste error and should be changed to \"ROLE-PLAYING ATTACK.\"","section":"Section 4.2.3"},{"comment":"The table headers contain spacing errors: \"F ALSE\" and \"T RUE\" should read \"False\" and \"True.\"","section":"Table 4"},{"comment":"There is a duplicated phrase \"the attack.the attack\" in the introductory paragraph of the ablation study; it should read \"the attack.\"","section":"Section 4.4.4"},{"comment":"A stray \"s\" appears after the concluding paragraph and before Section 4.7; it should be removed.","section":"Section 4.6.5"},{"comment":"The discussion of \"edge cases\" would be more useful with concrete examples of where the judge failed or was ambiguous, since the paper otherwise gives no characterization of judge error modes.","section":"Section 5.1"}],"recommendation":"major_revision","confidential_remarks":"The core vulnerability claim is credible and likely worth publishing, but this is fundamentally a competition retrospective rather than a fully validated benchmark study. The unnamed LLM judge is the central weakness: if the organizers cannot disclose the judge due to API policy, they should at least release judge outputs, the judge prompt, and a human-validated subset with inter-annotator agreement. I recommend major revision rather than rejection because the qualitative findings are independently supported by prior work and by the teams' own ablations, and the missing evaluation transparency is fixable within the manuscript's scope. The fit is reasonable for a workshop-style technical report, but for a journal venue the statistical treatment and negative controls need to be substantially strengthened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about arXiv:2506.12430. First, the qualitative result — that structured visual prompts like flowcharts, combined with role-play or prefix text, still jailbreak aligned MLLMs — is real and consistent with prior work like FigStep and FC-Attack. Second, the quantitative leaderboard numbers are not reproducible as reported, because the official judge is an unspecified LLM and the teams self-evaluated with different judges.\n\nWhat's genuinely new is the public resource. The paper documents a two-phase competition with 86 teams, gives detailed case studies with actual prompts and image designs, and promises data and code. The ablation tables are the most useful part: Team SuperIdol'smile shows red highlighting and directional guidance each add several points; MR-CAS shows omission-completion matters more than multilingual text. That level of detail is worth having.\n\nThe soft spots are in the evaluation protocol, and they're not minor. Eq. (2) defines ASR via a judge J that is never named, versioned, or prompted. Team 'You are the challenger' used Qwen2.5-latest; MR-CAS used qwen-max. If the official judge differed, the scores in Tables 2 and 3 aren't comparable across teams. Section 5.1 concedes the single-judge limitation. There is no human validation or inter-annotator agreement. On top of that, MR-CAS says they revised some neutral queries to emphasize harmful intent (Section 4.4.2), so the task wasn't identical for everyone. That's a data integrity concern, not just a measurement issue.\n\nThe novelty framing is overstated. The abstract calls the challenge 'pioneering' and Section 4.7 says 'groundbreaking,' but nearly every technique is a recombination of published attacks the paper itself cites. That's acceptable for a technical report, but the marketing language should be toned down.\n\nWho gets value? Researchers working on multimodal safety, red-teaming, or competition design. They'll find concrete attack templates and honest ablation data. I would not quote the leaderboard numbers without the judge specification. The paper deserves a real referee because the resource is useful and the evaluation flaws are fixable: name the judge, add a human-validated subset, and specify the data release.","headline":"A useful competition report with solid qualitative findings, but the leaderboard numbers rest on an unnamed judge and need human validation.","tokens_in":24390,"tokens_out":3632,"would_cite":false,"duration_ms":45100,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This technical report claims that adversarial image-text pairs built from flowcharts and role-playing text jailbreak safety-aligned multimodal LLMs at near-perfect rates, with top leaderboard scores of 96.95 (white-box) and 93.56…","keywords":["multimodal large language models","jailbreak attacks","adversarial image-text pairs","flowchart attack","role-playing prompts","LLM-as-a-Judge","attack success rate","red-teaming competition"],"falsifier":"Have independent human annotators re-score a random sample of the pairs the judge marked as successful attacks, without seeing the judge's labels, and compute agreement; if human-judge agreement falls below 0.7 kappa or most outputs are judged harmless, the reported near-perfect jailbreak rates would not reflect genuine harm.","tokens_in":23185,"feed_emoji":"🛡️","tokens_out":9970,"duration_ms":108386,"temperature":0.7,"pith_summary":"This report argues that current safety-aligned multimodal large language models are defeated by adversarial image-text pairs at scale: in a two-phase competition, 86 teams submitted attacks, and the best white-box score reached 96.95 while the best black-box score reached 93.56. The central claim is that harm hidden in structured diagrams, especially linear flowcharts with red-highlighted output nodes, and coordinated with role-playing text prompts, bypasses safety filters that plain-text attacks cannot. A sympathetic reader should care because the winning attacks are simple, training-free templates; they transfer across open and unseen models, which makes them a present-day risk for any application that accepts untrusted images.","feed_headline":"Flowchart attacks defeat safety-aligned vision AI: top score 96.95","feed_subtitle":"An 86-team red-team contest shows harm hidden in diagrams plus role-play text slips past current multimodal guardrails.","key_machinery":"The load-bearing mechanism is the flowchart-as-attack image paired with a role-playing text prompt. A four-node linear flowchart encodes, in sequence, the harmless context, the target question placed in a character's mouth, an affirmative reply, and a red-highlighted instruction to produce detailed harmful suggestions; the model's bias toward sequential structured outputs walks it through the nodes without any single input frame containing explicit harmful instruction. The report's measuring instrument is the Attack Success Rate, computed by an LLM-as-a-Judge that scores each output for harmfulness and relevance.","core_discovery":"The paper's central discovery, stated on its own terms, is that safety alignment in MLLMs defends the input surface but not the multimodal reasoning process. By embedding the complete malicious question inside a flowchart and pairing it with a text prompt that assigns a role, demands a fixed prefix, and imposes a length constraint, the winning team reports attack success rates of 100 percent on Qwen2-VL-7B and 98.7 percent on InternVL2-8B. The ablations attribute the effect to visual structure: flowcharts outperform a single box, a red-highlighted output node outperforms unhighlighted diagrams, and directional step guidance outperforms free generation. The Phase II results extend the claim to a black-box model, where a training-free flowchart-and-role-play template still achieves 89.33 in the official ranking, showing the vulnerabilities transfer without model access.","pith_inferences":["Editorial inference — the four-node flowchart template should be retested on newer models (for example, Qwen2.5-VL and InternVL3) and on closed commercial models; if the vulnerability persists, it is a general property of visual instruction following rather than a quirk of the two contest models.","Editorial inference — the red-highlight output-node effect suggests low-level visual salience changes safety behavior, which points toward attention-based defenses or input-rendering sanitization as testable countermeasures.","Editorial inference — because scores were awarded by an automated judge, the contest may have rewarded attacks that fool the judge rather than produce genuine harm; a human audit of the submitted pairs would clarify how much of the reported near-perfect success is real-world harm."],"forward_implications":["Deployments that accept untrusted images cannot rely on current safety alignment; the same flowchart-plus-role-play template can be generated automatically and at scale, so this is a practical, not merely laboratory, threat.","Defenses that screen image and text independently will miss attacks whose harmful intent exists only in the joint reading; robustness work must train on adversarially paired image-text data.","Because training-free attacks transferred to a black-box model in Phase II, jailbreak risk does not require white-box access or optimization, lowering the barrier for real-world attackers.","The report concludes that evaluation should move to multi-judge or hybrid human-automated assessment, because a single LLM judge struggles to judge borderline harmfulness."],"supporting_citations":[{"why":"Supplies the 180 foundational harmful prompts in six risk categories that define the Phase I attack targets.","marker":"(Ying et al., 2024a)"},{"why":"Identifies Qwen2-VL-7B, the open-source target model on which several top attacks report near-perfect success.","marker":"(Wang et al., 2024b)"},{"why":"Identifies InternVL2-8B, the second open-source target model used in both phases and in ablation studies.","marker":"(Gao et al., 2024)"},{"why":"FigStep is the typographic visual-prompt jailbreak baseline whose 78.67 percent average the Phase II ablation outperforms.","marker":"(Gong et al., 2025)"},{"why":"Visual-roleplay jailbreak is the source that leading teams adapt for the role-playing text component.","marker":"(Ma et al., 2024)"},{"why":"AdvBench's harmful behaviors subset supplies the extra evaluation data used by the reasoning-chain attack team.","marker":"(Zou et al., 2023)"},{"why":"Disguise-and-reconstruction jailbreaking is the autoregressive-completion technique that the winning team adapts in its text prompt.","marker":"(Liu et al., 2024a)"},{"why":"Analyzing-based Jailbreak (ABJ) is the reasoning-chain method that team theshi's attack is built on.","marker":"(Lin et al., 2024)"}],"fun_headline_variants":["Flowchart jailbreak hits 96.95 in multimodal safety test","Diagrams slip past vision AI guardrails: 96.95 bypass score","Role-play plus flowcharts crack multimodal LLM defenses","ATLAS 2025: Flowchart attacks achieve 96.95 on MLLMs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the single LLM judge's 'harmful or not' label is a valid measure of true harm for every submitted pair; the report gives no human validation, judge identity, or inter-annotator agreement, and concedes the judge struggles with borderline cases.","fun_headline_variants_meta":{"raw":{"variants":["Flowchart jailbreak hits 96.95 in multimodal safety test","Diagrams slip past vision AI guardrails: 96.95 bypass score","Role-play plus flowcharts crack multimodal LLM defenses","ATLAS 2025: Flowchart attacks achieve 96.95 on MLLMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000901,"raw_usage":{"total_tokens":3842,"prompt_tokens":869,"completion_tokens":2973,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":2894}},"tokens_in":485,"tokens_out":2973,"duration_ms":27327,"temperature":1.0,"reasoning_tokens":2894,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:50:47.227231+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have independent human annotators re-score a random sample of the pairs the judge marked as successful attacks, without seeing the judge's labels, and compute agreement; if human-judge agreement falls below 0.7 kappa or most outputs are judged harmless, the reported near-perfect jailbreak rates would not reflect genuine harm.","supporting_citations":[],"review_version":1}