Pith. sign in

REVIEW 4 major objections 5 minor

Shieldstral

T0 review · 4 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read The paper argues that content moderation reduces to a yes/no question, letting a 3B model match or beat fixed-taxonomy guardrails seven times its size on text and set the state of the art on multimodal safety.

desk verdict A plausible 3B guardrail with a genuinely useful data recipe, but the headline numbers hinge on a one-sentence holdout guarantee that is not auditable and on a self-built adaptability benchmark; worth a careful referee, not a clean accept as-is. read the letter →

arxiv 2607.25857 v2 pith:XD5EBYAO submitted 2026-07-28 cs.CL cs.CV

Antonia Calvi , Avinash Sooriyarachchi , Giada Pistilli , Guillaume Lample , Maarten Buyl , Maximilian Augustin , Maximilian Müller , Pierre Stock
show 268 more authors
Tom Bewley Wassim Bouaziz Yimu Pan Abdelaziz Bounhar Abhijeet Somani Aditi Kabra Adrian Valente Adrien Petralia Adrien Sadé Alan Jeffares Albert Jiang Aleksandr Timashov Alexandre Cahill Alexandre Gavaudan Alexandre Laval Alexandre Sablayrolles Amélie Héliou Amos You André Jonasson Andrew Bai Andrew Ehrenberg Andrew Zhao Angele Lenglemetz Anmol Agarwal Arata Suzuki Arjun Majumdar Arthur Fournier Artjom Joosen Aylin Guliz Akkus Aysenur Karaduman Baptiste Bout Baptiste Rozière Baudouin De Monicault Benjamin Holzschuh Benjamin Lefaudeux Benjamin Tibi Bernhard Stadlbauer Bl{}ażej Osiński Camille Le Scao Chaoran Yu Charlotte Cronjäger Chen-Yo Sun Chris Bamford Christian Wallenwein Christophe Renaudin Clémence Lanfranchi Corentin Barreau Corentin Sautier Cristiana-Diana Diaconu Cyprien Courtot Daniel Marczak Darius Dabert Diego de Las Casas Dominik Nuss Dylan Rubini Dzmitry Soupel Elizaveta Demyanenko Elliot Chane-Sane Emilien Fugier Emmanuel Gottlob Erik Aas Etienne Goffinet Étienne Millon Eujeong Choi Fabian Paischer Fabian Schlager Faruk Ahmed Federico Baldassarre Filip Szatkowski Florian Wiesner Gabrielle Berrada Gaëtan Ecrepont Gaétan Lepage Gaspard Blanchet Gaspard Donada-Vidal Gauthier Delerce Gauthier Guinet Genevieve Hayes Georgii Novikov Gianluca Galletti Guillaume Breton Guillaume Kunsch Guillaume Martin Guillaume Raille Gunjan Dhanuka Gunshi Gupta Han Zhou Harshil Shah Hasan Furkan Vural Hédi Hadiji Hope McGovern Hugo Cisneros Hugo Thimonier Indraneel Mukherjee Ivan Cuevas Salazar Jacques Sun Jan Ludziejewski Jason Rute Jean Quentin Jean-Hadrien Chabran Jean-Malo Delignon Jie Zhang Joachim Studnia Joep Barmentlo Johannes Brandstetter John Harvill Jonas Amar Jonas Schweizer Joséphine Delas Josselin Somerville Julien Denize Julien Tauran Kartik Khandelwal Khyathi Raghavi Chandu Kilian Tep Kush Jain Larissa Laich Laura Calem Laurence Aitchison Laurent Callot Laurent Fainsin Léo Cotteleer Léonard Blier Lingxiao Zhao Louis Martin Louis Serrano Lucile Saulnier Ludovic Ho Fuh Luis Montero Manon Chossegros Marcin Możejko Margaret Jennings Markus Hennerbichler Martin Alexandre Mathieu Poirée Mathieu Schmitt Mathilde Guillaumin Matthieu André Matthieu Dinot Matthieu Futeral Maurits Bleeker Mauro Comi Max Mynter Maxim Berman Maxime Darrin Maxime Louis Melina Jingting Laimon Mert Unsal Mia Chiquier Michael Pilcer Michal{} Pietruszka Michal{} Zając Mikhail Biriuchinskii Minh-Quang Pham Minwoo Kang Morgane Rivière Namit Katariya Nathan Grinsztajn Nathan Simpson Neeraj Aggarwal Neha Gupta Ola Mysiak Oliver Leicht Olivier Bousquet Olivier Duchenne Parag Jain Patricia Wang Patrick Blies Patrick von Platen Paul Jacob Paul Wambergue Paula Kurylowicz Pavan Kumar Reddy Pavel Kuksa Philippe Pinel Philomène Chagniot Pierre-André Savalle Piotr Milos Prateek Gupta Pravesh Agrawal Quentin Desreumaux Quentin Torroba Quercus Hernandez Ram Ramrakhya Randall Isenhour Ranjit Parva Raul Perez Pelaez Reinhard Sonnleitner Rémi Delacourt Richard Kurle Rishi Shah Rob Romijnders Rohin Arora Romain Sauvestre Roman Soletskyi Rosalie Millner Rupert Menneer Sagar Vaze Samuel Barry Samuel Belkadi Samuel Humeau Sanchit Gandhi Sandeep Subramanian Sarthak Mittal Saskia Adaime Sean Cha Sebastian Kaltenbach Shashwat Dalal Shashwat Verma Sherif Waly Shrimai Prabhumoye Siddhant Waghjale Siddharth Gandhi Simon Lepage Simon Sorg Soham Ghosh Sophie Marbach Srijan Mishra Stanislas Lange Steve Hong Sumukh Aithal Szymon Antoniak Tarun Kumar Vangani Teven Le Scao Théo Cachet Thibaut Lavril Thomas Chabal Thomas Coste Thomas Defard Thomas Foubert Thomas Robert Thomas Wang Tianyu Zhang Tim Lawson Timothée Lacroix Tobias Kronlachner Tom Edwards Tomas Hodan Tuhin Das Tyler Wang Ulrick BLE Umar Jamil Umberto Tomasini Valentin Macé Van Phung Vedant Nanda Victor Jouault Victor Letzelter Victor Paltz Victor Poucheret Vincent Maladière Vincent Pfister Virgile Richard Vladislav Bataev Wen Ding Li William Havard William Marshall Xinghui Li Xingran Guo Xinyu Yang Yann Dreze Yassine El Ouahidi Yassir Bendou Yihan Wang Yves Martin des Taillades Zaccharie Ramzi Zhenlin Xu Zsofia Csakany
This is my paper
classification cs.CLcs.CV
keywords contentmoderationmultimodalsafetybinaryquestionansweringpolicyadaptationcontrastivedatagenerationtaxonomygeneralizationguardrailmodelLoRAfine-tuning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Shieldstral's central claim is that heterogeneous safety-moderation tasks can be collapsed into one binary question-answering format. On this basis the authors construct a 3B-parameter multimodal classifier trained on roughly 54.1M unified samples, 4.4M of them contrastive pairs that teach the model to tell apart subtly related policies. The resulting model reports an average F1 of 84.9% on text safety benchmarks—matching a 20B baseline—83.8% on multimodal benchmarks, and 91.3% on a new 52-leaf policy-adaptability benchmark. The payoff is a guardrail that operators can steer at inference time with a plain-language query about their own policy, rather than a fixed taxonomy baked in at training.

What carries the argument

The carrying mechanism is the instruction-query-document prompt, which converts every moderation task into binary QA: a system message, an <Instruct> field fixing context and strictness, a <Query> field asking a yes/no safety question, and a <Document> field containing text and/or image. At inference, only the logits of the 'yes' and 'no' tokens are unembedded and softmax-normalized into a score thresholded at 0.5. The training-side machinery that makes this work is contrastive sample curation and generation: the same content is paired with matching and non-matching queries, and safe texts are rewritten by an LLM into unsafe variants with sibling-category negatives, so the model learns to di

What would settle it

Go through the 45.2M open-source text training samples and look for verbatim or near-verbatim matches of any test item in the evaluation benchmarks listed in Table 4 of the paper. One exact test prompt appearing in training would falsify the held-out claim and void the headline F1 comparisons.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that encoding moderation as a yes/no question, with a natural-language query and a structured instruction-document prompt, lets a small model absorb many divergent taxonomies in one training run. The authors combine 45.2M open-source text samples in a template-unified format with LLM-generated contrastive pairs and a multimodal pipeline, then SLERP-merge a public-data checkpoint, a public-plus-generated-taxonomy checkpoint, and the base instruct model. The result replaces a fixed category head with a continuous safety score over the 'yes' and 'no' tokens, and generalizes to policies whose category names, scope, and granularity differ from everything

Load-bearing premise

The single load-bearing assumption is that every evaluation benchmark was truly absent from the 45.2M open-source training samples; the paper asserts this in one sentence but never lists the training datasets or the exclusion mechanism, and the training domains match the benchmark domains by construction.

Editorial extensions

If this is right

  • If true, a modest 3B classifier can stand in for much larger guardrail models, cutting deployment cost and inference latency in content-moderation pipelines.
  • Because the policy arrives as a natural-language query at inference time, the same checkpoint can serve different deployments with different standards (mental health, cybersecurity, kids) without retraining or per-category fine-tuning.
  • The unified 54.1M-sample training format allows safety datasets with incompatible taxonomies to be pooled into one training signal, so future data-collection efforts can be measured by a single format.
  • The adaptability benchmark suggests a new evaluation style for guardrails: testing on independently designed taxonomies with disjoint categories, to measure genuine policy generalization rather than label memorization.
  • The continuous score from yes/no logits gives operators a tunable threshold to trade precision against recall.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the held-out assumption fails, the headline numbers could overstate the contribution; an independent audit that checks exact test instances against the open-source training corpus would settle this.
  • The contrastive iso-content training technique is not tied to safety: the same generate-query-pairs-from-safe-seeds recipe could be applied to other fine-grained classification domains, such as medical triage or content labeling, where policies vary by jurisdiction or audience.
  • The paper's reported edge cases (lower Arabic/Indonesian prompt-classification F1) hint that the template-unification approach inherits the language coverage of its source datasets; targeted multilingual contrastive generation could close that gap.
  • A testable extension is to vary the threshold beyond 0.5 on the continuous score and report operating-characteristic curves, allowing operators to select a false-positive/false-negative trade-off without retraining.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces Shieldstral, a 3B-parameter multimodal safety classifier built on Ministral-3B. Content moderation is reframed as a binary question-answering task: each input is structured into <Instruct>, <Query>, and <Document> fields, and the model outputs a softmax-normalized score over the "yes"/"no" tokens. The authors construct approximately 54.1M training samples (45.2M open-source text, 4.4M synthetic contrastive text, 4.5M multimodal), train LoRA checkpoints on public-only (P) and public-plus-generated (PG) data, merge them via SLERP, and evaluate on 16 benchmarks (21 splits) plus a new 52-leaf policy-adaptability benchmark. The central claims are that Shieldstral matches or outperforms models nearly 7× its size on text safety (84.9% average F1), sets a new state of the art on multimodal safety (83.8% average F1), and achieves 91.3% F1 on the adaptability task.

Significance. If the empirical claims hold, the paper makes a strong practical contribution: a single binary-QA formulation that unifies heterogeneous safety datasets, a scalable data recipe, and evidence that a 3B model can compete with 20B guardrails while adapting to novel policies at inference time. The internal ablations (Tables 5–7) are coherent and support the qualitative recipe: each stage adds F1, LoRA approximates full SFT, and SLERP merging recovers benchmark calibration while preserving taxonomy gains. The paper also ships a level of transparency in the taxonomy comparison (Tables 3, 10, 11) that is useful. However, the headline numbers rest on two load-bearing empirical premises that the manuscript does not currently establish: (i) that no evaluation benchmark leaked into the 45.2M open-source training corpus, and (ii) that the self-built adaptability benchmark measures true policy adaptability rather than familiarity with the authors' generation protocol. These issues are fixable but must be addressed before the results can be accepted.

major comments (4)
  1. [§7 and §3.1] The sentence "All evaluation samples are held out from the training data to ensure fairness" (§7) is the sole holdout guarantee, yet §3.1 states that the 45.2M open-source text samples draw from "safety, toxicity, hate speech, jailbreak detection, content moderation, and response quality" — the exact domains of the Table 4 benchmarks. The paper never names the constituent training datasets, their versions/splits, or the exclusion mechanism (deduplication, hash-based holdout, benchmark-specific filtering). If any of WildGuardTest, ToxicChat, Aegis, HarmBench, OpenAI Moderation, BeaverTails, VLGuard, etc., or near-duplicates, entered training, the reported 84.9% text and 83.8% multimodal F1 numbers — and the size-efficiency claim built on them — would be inflated. This is load-bearing for all three headline claims. Please provide a complete dataset inventory, version identifiers, and a qua
  2. [§4, §4.2, §7.2] The adaptability benchmark is generated with the same safe-to-unsafe LLM rewriting protocol used for training data (§3.3 and Appendix E): a safe source text is rewritten to exhibit a target category while avoiding a sibling, then paired with a yes/no query. Section 4.1 differentiates the training and evaluation taxonomies and generation LLMs, which mitigates label memorization. But it does not rule out protocol familiarity: the model has seen tens of thousands of examples of the exact input-output transformation (safe text → rewritten unsafe text + category query), so high F1 on a similarly generated benchmark may partly reflect procedural mimicry rather than policy adaptability to truly novel policy definitions. A concrete test would be to evaluate on policies sourced from an independent process (e.g., existing policy documents, human-authored policies) or to vary the generation prompt,
  3. [§7.1, Figures 5–6] The abstract and Section 1 claim Shieldstral "matches or outperforms models nearly 7× its size" on text safety. The overall safety-classification F1 is 84.9% for Shieldstral and GPT-OSS-Safeguard-20B (Figure 5), which is a tie, not an outperformance. Moreover, on refusal detection (Figure 6), GPT-OSS achieves 93.9% versus Shieldstral's 90.3%. The paper does not specify whether refusal detection is included in the "text safety benchmarks" claim. This ambiguity affects the central claim and must be clarified. If refusal detection is included, the claim should be qualified; if it is excluded, the paper should say so explicitly.
  4. [§7, Table 4] Several benchmarks are small (e.g., Aegis, 359 samples; SimpleSafetyTests, 100 samples), and the key claims are based on F1 differences of a few points (e.g., 84.9 vs 84.9, or 83.8 vs 77.6). No confidence intervals, bootstrap estimates, or significance tests are reported. Given the threshold-of-0.5 decision and the small sizes, some differences may be within noise. Reporting confidence intervals or at least per-benchmark sample counts with variance would make the comparative claims more robust.
minor comments (5)
  1. [Figure 8 caption] The caption states "Some LlavaGuard test images were unavailable; scores are based on the available subset." This is important for fairness, but the manuscript does not state how many images were available or whether the subset is comparable across models. Please report the exact subset size and the number of images used for each baseline.
  2. [§4.1 / Table 9] The evaluation taxonomy has 90 fixed queries (52 leaf + 26 subcategory + 12 superclass), but the paper does not report the total number of evaluation samples per leaf/subcategory/superclass. Table 12 even notes one leaf (Physical Property) has only 1 sample. Without per-category sample counts, category-level F1 numbers in Table 12 are hard to interpret.
  3. [Table 9] Typographical error: "V oter Suppression" should be "Voter Suppression" under SC9.
  4. [§6.1, Table 6] The paper concludes "no significant overall difference" between LoRA and full SFT, but no significance testing is reported. The observed differences (e.g., 87.1 vs 87.8 on Aegis v2) may or may not be meaningful. A brief note on variance or a paired test would support the wording.
  5. [References] Several references are future-dated or from 2026 (e.g., Liu et al. 2026, Singh et al. 2026) and presently unverifiable. Please ensure these references are real and include version/date/access information, or mark them as preprints with identifiers.

Circularity Check

1 steps flagged · score 4.0 of 10

Policy-adaptability benchmark reuses the training data's LLM contrastive-rewriting protocol; text and multimodal results rest on external benchmarks.

  1. other [Section 4 (Adaptability Evaluation) and Section 3.3 (Contrastive Sample Generation); Figure 7]
    "Therefore, we apply the same contrastive generation idea to produce an evaluation dataset. // Training samples are generated by rewriting safe source texts into unsafe variants using an LLM (Appendix E). // For each category, an LLM produces paired examples: given a target category and one of its siblings, it generates a positive sample matching the target and a negative sample matching the sibling but not the target."

    The headline adaptability figure (91.3% F1) is produced on a benchmark built by the same safe-text-to-target/sibling LLM-rewriting procedure used for the synthetic training data. The evaluation asks exactly the training task under new labels: decide whether content exhibits the queried category (positive rewrite) or only its sibling (negative rewrite). Section 4.2 concedes 10 of 12 eval super-classes have loose training counterparts. Therefore the result measures generalization within the same generative distribution and label family, rather than independent evidence of adaptation to truly external policies; it is a partially self-referential benchmark, though not a full reduction because taxonomy names, fixed queries, and generation LLMs differ.

full rationale

The text and multimodal headline claims are evaluated on external benchmarks (WildGuardTest, ToxicChat, Aegis, HarmBench, OpenAI Moderation, BeaverTails, PolyGuard, RTP-LX, VLGuard, UnsafeBench, LlavaGuard), so they are independent support and not circular. The only partial circularity is the 91.3% policy-adaptability figure: Section 4 explicitly reuses 'the same contrastive generation idea' as Section 3.3, and 10 of 12 eval super-classes have counterparts in the training taxonomy. This is mitigated by separate eval taxonomy design, fixed queries, and different LLMs/seed samples, so the score is 4 rather than 6. The Section 7 holdout sentence ('All evaluation samples are held out from the training data to ensure fairness') is un-auditable and the open-source training corpus spans the same benchmark families, but that is a contamination/provenance risk, not an exhibited circular reduction, so it does not raise the circularity score. No load-bearing self-citation chain is present; the paper's foundation-model citations are base-model attribution.

Assumptions & free parameters 6 free parameters · 7 assumptions · 0 invented entities

The paper introduces no new physical or theoretical entities. Its load-bearing postulates are methodological: the binary-QA reduction, LLM-as-ground-truth for labels, representativeness of LLM-rewritten contrastive data, and hand-chosen pipeline constants (SLERP weights 0.6/0.3/0.1, τ=0.5, 30% inverse queries, unreported reranker thresholds). The self-built 52-leaf evaluation taxonomy is a benchmark artifact rather than an invented entity, but its independence from the training taxonomy is only partial (10 of 12 super-classes align, Section 4.2).

free parameters (6)
  • SLERP merge weights (PG / P / I) = 0.6 / 0.3 / 0.1
    Chosen by ablation on validation sets (Table 7) to balance Aegis v2 and taxonomy F1; not derived from first principles.
  • Positive-sample duplication factor k = k ≥ 1 (unspecified)
    Section 3.2 class balancing duplicates each positive sample k times; the actual value used is not reported.
  • VL reranker asymmetric filtering thresholds = not reported
    Section 3.4 uses asymmetric thresholds to preserve rare violation samples; exact values unstated.
  • Inverse-query fraction in image pipeline = 30%
    Section 3.4: ~30% of ~2,000 generated queries are inverse formulations; a hand-chosen diversity knob.
  • Decision threshold τ after softmax = 0.5
    Section 2: fixed threshold for binary classification; conventional but hand-chosen, and every reported F1 depends on it.
  • Strictness-level assignment per dataset = strict / moderate / lenient (manual)
    Section 3.1 / Table 1: each processor's templates encode a hand-assigned strictness level that determines decision boundaries.
assumptions (7)
  • domain assumption A binary yes/no answer to a natural-language query suffices to represent any moderation policy
    Section 2 reduces all safety tasks to a single thresholdable question; if policies are not reducible to one question per content item, the framework loses coverage.
  • domain assumption LLM labels are valid ground truth for filtering and verification
    Section 3.2 filters public datasets wherever an LLM disagrees with the dataset label; Section 4.1 uses a separate LLM to verify eval samples. Systematic LLM blind spots would distort training and inflate the self-built eval.
  • domain assumption LLM-rewritten contrastive pairs are representative of real deployment queries
    Sections 3.3 and 4 assume that rewriting safe→unsafe with sibling-category constraints produces content and questions resembling what operators will ask in practice.
  • domain assumption Per-dataset template mapping preserves each source dataset's intended decision boundaries
    Section 3.1: hand-designed processors map heterogeneous taxonomies into yes/no templates; fidelity of this mapping is asserted, not measured.
  • domain assumption The 52-leaf evaluation taxonomy is disjoint, unambiguous, and operationalized by its 90 fixed queries
    Section 4.1: evaluation validity depends on each category's query uniquely and correctly capturing the category; no human validation of the queries is reported.
  • domain assumption Ministral-3B + Pixtral provide adequate base representations at 3B scale
    Section 5: the result inherits the base model's multilingual and visual competence; failures of the base model would cap the classifier.
  • standard math Softmax over yes/no logits is a well-calibrated safety score
    Section 2 computes s = exp(zyes)/(exp(zyes)+exp(zno)); standard normalization used for thresholding.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Shieldstral." pith.science (2026). https://pith.science/paper/XD5EBYAO

@misc{pith2026260725857,
  author       = {Pith},
  title        = {Pith review of: Shieldstral},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XD5EBYAO}},
  note         = {Machine review of arXiv:2607.25857}
}
abstract

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7$\times$ its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one training framework. We present the data construction recipe, covering curation and generation of approximately 54.1M samples and a fine-grained evaluation set to evaluate policy adaptability. Together, these enable a small adaptive model to match or outperform much larger models.

Figures

Figures reproduced from arXiv: 2607.25857 by the authors.

Figure 1
Figure 1. Shieldstral architecture. An instruction, a natural-language query, and content (text, image, or both) are composed into a structured prompt, then processed by the model in a single forward pass. The softmax￾normalised logits of the “yes” and “no” tokens yield a continuous safety score that is thresholded for binary classification. 2 Task Definition Since our goal is to train a policy-adaptive safety classification … view at source ↗
Figure 2
Figure 2. Training sample examples. (a) Text-only: the <Document> contains a prompt–response pair in bracketed format. (b) Multimodal: the <Document> contains an image token followed by optional text. In both cases, the three user-message fields are independently sampled, and the target label is a separate single-token assistant turn. yes/no question-answering problem. Second, contrastive sample curation (Section 3.2) pairs t… view at source ↗
Figure 3
Figure 3. Example contrastive training pair generated from a single LLM call. The same rewritten content is paired with a target-category query (positive, label “yes”) and a sibling-category query (negative, label “no”), training the model to discriminate between closely related categories. cannot simply be generated synthetically by an LLM. To overcome this scarcity, we supplement the limited moderation sources with general-… view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Contrastive evaluation pair for CAT001 (Physical Violence). Both samples are rewritten from the same safe source text. The positive sample describes physical harm (matching the query), while the negative sample describes kidnapping (sibling category CAT002)—unsafe cont…
Figure 5
Figure 5. Figure 5: F1 scores (%) on safety classification benchmarks. PolyGuard and RTPLX are multilingual datasets. Qwen3Guard results are averaged over strict (controversial=unsafe) and loose (controversial=safe) mappings. ShieldGemma and Shieldstral use a threshold of 0.5. GPT-OSS-Saf…
Figure 6
Figure 6. Figure 6: F1 scores (%) on refusal detection benchmarks. PolyGuard is a multilingual dataset. Qwen3Guard results are averaged over strict (controversial=unsafe) and loose (controversial=safe) mappings. Shieldstral uses a threshold of 0.5. GPT-OSS-Safeguard-20B uses reasoning_eff…
Figure 7
Figure 7. Figure 7: Scores (%) on the adaptability benchmark. Qwen3Guard results are averaged over strict (controversial=unsafe) and loose (controversial=safe) mappings. Shieldstral uses a threshold of 0.5. GPT￾OSS-Safeguard-20B uses reasoning_effort=high. Nemotron-3.5-Safety uses reasoni…
Figure 8
Figure 8. Figure 8: F1 scores (%) on multimodal safety benchmarks. Shieldstral and ShieldGemma-2-4B use a threshold of 0.5. Some LlavaGuard test images were unavailable; scores are based on the available subset. Nemotron-3.5- Safety uses reasoning_effort=none for default categories. train…

Discussion (0). Sign in to comment.

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.