{"id":"01efda04-aa8d-4430-a933-fc0a178f3977","arxiv_id":"2607.22545","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A 184M-parameter DeBERTa-v3 fine-tuned model is claimed to beat Llama-Guard-3-8B on all tested prompt-injection benchmarks while adding BFSI regulatory labels, but a leaked training/eval overlap undermines the zero-FPR claim.","lead":"Semalith v1.4 is a 184M-parameter DeBERTa-v3 classifier that flags prompt injection, general harm, and financial-services compliance labels in one pass, and the authors report it outperforms Llama-Guard-3-8B on seven prompt-injection benchmarks. The paper's headline zero-false-positive result on benign agentic prompts is weakened by a disclosed overlap between its training data and the benchmark, and several validation tables appear to come from an earlier checkpoint.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own contamination audit undermines the headline zero-FPR claim: 169 training rows collide with the AgentHarm-benign holdout, so FPR=0.000 is not a clean generalization measurement.","rationale":"The reader's weakest_assumption is exactly the load-bearing point: the 169-row overlap with the AgentHarm-benign holdout means the headline FPR is not a valid OOD measurement. My own reading of Sections 3.2, 4.2, and 7 confirms the contradiction is internal, not a matter of disputed benchmark norms. Other inconsistencies strengthen the REJECT verdict but are secondary: Table 3's macro-F1 0.876 matches the v1 baseline, while Section 5.2 reports v1.4 macro-F1 0.8222; Section 7's CI [0.001, 0.027] also conflicts with Table 4's [0.000, 0.018]. I do not see a path by which the abstract's zero-FPR and contamination-clean claims survive without either releasing weights and showing the clean-subset FPR, or revising the claims to acknowledge the overlap. Therefore the reader's verdict should stand.","tokens_in":15967,"tokens_out":7099,"duration_ms":63267,"concrete_test":"Run the released contamination_audit.py against the v1.4 corpus manifest to enumerate the exact SHA-1/MinHash collision pairs between the training corpus and the 208 AgentHarm-benign prompts. Then run the released v1.4 checkpoint on all 208 prompts and recompute FPR on the non-colliding subset only. If any false positive appears on the clean subset, the abstract's FPR=0.000 claim fails for unseen data; if zero, the result still rests on at most 39 clean samples, which is too small to support the claimed Wilson CI or the '100% contamination-clean' statement.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is FPR=0.000 on 208 benign agentic prompts, which drives the 'contamination-clean corpus' and 44x state-of-the-art narrative. The manuscript's own audit (Section 7) reports exactly one nonzero contamination: 'agentharm_benign_holdout: 169 collisions from WildGuardMix authority-filtered rows overlapping AgentHarm benign prompts.' That is 169 of 76,204 training rows, or 0.2218% of the corpus, in direct tension with Section 3.2's statement that every candidate row is SHA-1- and MinHash-deduplicated against every held-out evaluation benchmark, and with the abstract's 'contamination-clean 76,204-row corpus.' Because the only contaminated benchmark is the one used for the headline zero-FPR result, FPR=0.000 may largely reflect memorization of training data rather than generalization to unseen benign agentic prompts. At most 39 of the 208 prompts are uncontaminated; a 0/39 result gives a Wilson upper bound around 0.09, not the claimed [0.000, 0.018]. The paper discloses the audit, which is good, but the abstract and conclusion still assert a clean generalization claim that the audit does not support.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Semalith v1.4 is a 184M-parameter DeBERTa-v3-base classifier with a 22-class head (D-attack subtypes, general harm, BFSI regulatory labels) and a 4-class auxiliary super-category head, trained on a 76,204-row corpus mined from 49 public sources. The paper claims state-of-the-art prompt-injection detection at 44x fewer parameters than Llama-Guard-3-8B, winning 7/7 prompt-injection benchmarks, 11/18 benchmarks overall, and achieving FPR = 0.000 on 208 AgentHarm benign agentic prompts. It also describes a binary-mapping methodology fix, latency measurements, McNemar significance tests, and a reproducibility harness with SHA-1 contamination audit.","tokens_in":16383,"tokens_out":7442,"duration_ms":61161,"significance":"If correct, this would be a useful engineering contribution: a compact multi-axis safety classifier with low latency and a disclosed three-axis taxonomy, backed by a reproducibility package, Wilson confidence intervals, and a documented audit of six weaknesses. The paper also includes useful deployment guidance distinguishing v1.3 and v1.4 operating points. However, the central empirical claims are undermined by internal contradictions and by contamination of the benchmark used for the headline zero-FPR result. These issues are load-bearing: the paper's main selling point is exactly the clean, parameter-efficient, zero-FPR comparison against Llama-Guard-3-8B.","major_comments":[{"comment":"The audit in §7 reports 169 training rows overlapping AgentHarm-benign prompts (0.2218%), directly contradicting §3.2's claim that every candidate row is SHA-1/MinHash-deduplicated against every held-out benchmark and the abstract's 'contamination-clean 76,204-row corpus.' Since AgentHarm-benign is the only contaminated benchmark and is the basis of the headline FPR=0.000 on n=208, the zero-FPR figure is not a clean generalization measurement. With at most 39 uncontaminated prompts, a 0/39 result gives a Wilson upper bound of ≈0.09, not the claimed [0.000, 0.018]. This cannot be fixed by re-analysis alone.","section":"§7 vs §3.2 and Abstract"},{"comment":"Table 3 reports macro_f1 = 0.876 and per-class F1 values (D1_AUTHORITY_CLAIM 0.500, D6_AGENTIC_INJECTION 0.857, B-11 0.833) that are inconsistent with v1.4 values stated in §5.2 (val macro-F1 0.8222) and with §4.5/§6.6, which report the three enriched labels' val F1 as 0.931, 0.755, and 0.577. The table appears to be v1/v1.3 data rather than v1.4 data. This is load-bearing because the paper's claim that v1.4 fixed the thin-label gaps rests on per-class F1 numbers that are not in the results table.","section":"Table 3 vs §5.2, §4.5, §6.6"},{"comment":"The paper states Semalith v1.4 wins 11 of 18 benchmarks (abstract, §1, §4.4, §8), but counting the rows in Table 5 gives at least 12 Semalith wins (hackaprompt, gandalf, advbench, mosscap_l6, mosscap_l7, mosscap_l8, wildjailbreak, salad_clean_eval, attaq, aart, beavertails_test, agentharm_benign_holdout). If HarmBench-copyright is excluded as 'not a fair comparison' (§4.4), the comparison has 17 benchmarks and the count is 12. The reported win rate is arithmetically wrong and repeated in the abstract.","section":"Table 5, §4.4"},{"comment":"The status of the AgentHarm-benign evaluation is contradictory. The abstract footnote and Table 4 report a completed 208-row evaluation with Wilson CI [0.000, 0.018], while §7 says 'A full validation of the AgentHarm-benign FPR claim on the complete 208-row dataset is planned for the next release cycle' and gives a different CI ([0.001, 0.027]). These statements cannot both be true; the discrepancy prevents verification of the headline zero-FPR result.","section":"§7 vs §4.2 and Abstract footnote"},{"comment":"The 'state-of-the-art prompt-injection detection' claim is unsubstantiated. The paper compares PI performance only with Llama-Guard-3-8B and Granite-Guardian, but its own McNemar analysis (§5.3) shows PromptGuard-2-86M is significantly superior on 14 benchmarks including Mosscap, with near-ceiling recall. Without a head-to-head against the actual state-of-the-art compact PI classifier, the SOTA claim is not supported.","section":"§5.3 vs Abstract/Conclusion"}],"minor_comments":[{"comment":"The text first says 'The seven PI/stealth benchmarks with χ² > 500' and then later 'The five PI/stealth benchmarks with χ² > 500'; at least one count is wrong.","section":"§5.3"},{"comment":"The paper calls the model 'calibrated' but provides no calibration analysis (e.g., ECE, reliability diagram, temperature scaling results). The term appears only in the title and abstract.","section":"Title/Abstract"},{"comment":"The Wilson CI for AgentHarm-benign differs between Table 4 ([0.000, 0.018]) and §7 ([0.001, 0.027]). A zero-event interval cannot have a positive lower endpoint, so the §7 interval is internally inconsistent with the reported 0/208.","section":"§7/Table 4"},{"comment":"Typo: 'amicrosoft/deberta-v3-base' should be 'a microsoft/deberta-v3-base'. Also, Table 6's caption redundantly notes 'AttaQ (0.945 vs 0.868) is a Semalith win' when that row already marks the winner.","section":"§3.1"}],"recommendation":"reject","confidential_remarks":"The manuscript has several correctable presentation issues, but the core problems are load-bearing: the headline FPR=0.000 is measured on a benchmark contaminated by training data, and Table 3 does not report v1.4 metrics. These are not fixable by editing; they require retraining or re-evaluation on clean data. The paper's own audit is transparently disclosed, which is to the author's credit, but the abstract and conclusion still assert a clean generalization claim that the audit does not support. Given the paper's central claims, I recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the good: Semalith is a genuine artifact. A 184M DeBERTa classifier with a 22-class + 4-class dual head covering prompt-injection subtypes, general harm, and BFSI labels is a practical contribution. The latency numbers (11.6 ms at batch 1) are real and matter for production. The PI benchmark comparison against Llama-Guard-3-8B is thorough, and the 7/7 wins on clean PI benchmarks (Mosscap, HackaPrompt, Gandalf, etc.) appear to hold up. The paper is also more transparent than most: it lists six weak spots with root causes and mitigation paths. Credit where earned.\n\nThe problem is the zero-FPR headline. Section 7's own audit reports 169 collisions between training rows and the AgentHarm-benign holdout. That is the exact benchmark used for the FPR=0.000 claim. With 169 of 208 prompts contaminated, the effective clean evaluation is at most 39 prompts. 0/39 gives a Wilson CI upper bound around 0.09, not the claimed [0.000, 0.018]. So the abstract's \"contamination-clean 76,204-row corpus\" is misleading, and the conclusion's zero-FPR statement overstates what the data supports. The disclosure itself is good, but the abstract and conclusion still assert a clean generalization claim.\n\nSecond, internal inconsistencies: Table 3 reports macro-F1 0.876, but Section 5.2 says v1.4 seed=42 val macro-F1 is 0.8222. Table 3 looks like v1.3 data. Per-class F1s in §6.6 (D1_AUTHORITY_CLAIM 0.931, D6 0.755, B-11 0.577) contradict Table 3 (0.500, 0.857, 0.833). Table 5 actually shows 12 Semalith wins, not 11. These are not fatal individually, but they erode trust in the reported numbers. The post-hoc binary mapping is disclosed and applied consistently, so I don't consider it a deal-breaker. Weights not released is a limitation but not fatal.\n\nBottom line: this paper deserves serious referee time, not because it is polished (it isn't) but because the engineering question is real and the PI wins on clean benchmarks might survive scrutiny. I would send it to peer review with major revision: re-run the zero-FPR claim on a truly held-out agentic benign set, fix the tables, and soften the abstract. Current form should be rejected, but it is worth engaging with.","headline":"Real engineering substance and a useful benchmark comparison, but the headline zero-FPR claim is undermined by the paper's own contamination audit, and the tables have internal inconsistencies.","tokens_in":16822,"tokens_out":3714,"would_cite":false,"duration_ms":33755,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 184M-parameter classifier outperforms an 8B guardrail on every prompt-injection benchmark while maintaining zero false positives on benign agentic prompts.","keywords":["prompt injection detection","safety classifier","DeBERTa-v3","small language models","false positive rate","benchmark contamination","financial services compliance","Llama-Guard-3 comparison"],"falsifier":"Take the released training-corpus manifest and the 208 AgentHarm-benign prompts; run a fuzzy string-match and embedding-similarity check beyond exact SHA-1. If near-duplicates beyond the 169 reported collisions exist, then retrain Semalith v1.4 with all overlapping rows removed and measure FPR on AgentHarm-benign; a meaningful rise above 0.000 would falsify the claim that the zero-FPR is a clean generalization result.","tokens_in":15846,"feed_emoji":"🛡️","tokens_out":4331,"duration_ms":37081,"temperature":0.7,"pith_summary":"Semalith v1.4 is a compact encoder-based safety classifier that claims to match or beat Llama-Guard-3-8B on all seven prompt-injection and adversarial-stealth benchmarks, using 44x fewer parameters. The paper argues that the key is not model scale but a three-axis taxonomy (prompt injection, general harm, financial-services compliance) trained on a contamination-controlled real-world corpus with an auxiliary super-category head. If true, this means small classifiers can handle the safety-filtering workload that currently requires large causal-LM guardrails, with a 33x latency advantage. The paper also introduces a binary-mapping correction that eliminates a systematic evaluation artifact, and it discloses six measured weaknesses and deployment guidance. A sympathetic reader would take the central claim as: for structure-detectable attacks, a well-trained 184M encoder can outperform an 8B model.","feed_headline":"184M classifier beats 8B guardrail on every prompt-injection test","feed_subtitle":"Seven-for-seven wins over Llama-Guard-3 at 44x fewer parameters, with zero false positives on benign agentic prompts.","key_machinery":"The object carrying the argument is a DeBERTa-v3-base encoder (184M parameters) with two heads on the [CLS] embedding: a 22-way classification head covering BENIGN, nine prompt-injection sub-types, general harm, and eleven BFSI regulatory labels, plus a 4-way auxiliary super-category head (BENIGN / D-attack / D8-harm / BFSI) trained with a jointly weighted loss. The auxiliary head is credited with preventing D-attack sub-class collapse when any sub-class is below the training-data stability floor. The training corpus is 76,204 real-world rows from 49 public sources, deduplicated with SHA-1 and MinHash against every held-out benchmark, and the evaluation uses a corrected binary mapping that t","core_discovery":"The paper's central claim is that a 184M-parameter DeBERTa-v3-base classifier, fine-tuned with a 22-class head and a 4-class auxiliary super-category head under jointly weighted loss, achieves state-of-the-art prompt-injection detection: it wins every one of seven prompt-injection benchmarks against Llama-Guard-3-8B, often by 50–90 percentage points, while producing zero false positives on 208 benign agentic prompts (vs 0.063 for the 8B model). The same model simultaneously outputs eleven financial-services regulatory labels in one forward pass. The authors locate the reason in data diversity rather than capacity: the bottleneck in earlier 184M iterations was under-represented attack classes","pith_inferences":["The zero-FPR figure is the paper's most fragile claim: the author's own audit reports 169 exact-string collisions between training-corpus rows (WildGuardMix authority-filtered) and the AgentHarm benign prompts, so the 0.000 is not a clean out-of-distribution measurement. A fair test would remove those rows and re-measure.","If the auxiliary super-category head is as important as claimed, the same recipe should transfer to other small encoders (e.g., a distilled RoBERTa or a DeBERTa-large) and other domain taxonomies (healthcare, legal), offering a fast way to build domain-specific guardrails.","The paper's 'data diversity beats capacity' argument implies that further PI improvements will come from more diverse attack data, not larger models; a testable extension is to train on the same corpus with a 2x-larger encoder and check whether recall moves less than adding 10k new attack rows.","The binary-mapping correction is a reminder that evaluation pipelines for multi-label safety classifiers can systematically inflate FPR; other guardrails reporting binary flags from fine-grained heads may harbour the same artifact."],"forward_implications":["For production safety filtering, a ~184M parameter encoder can replace an 8B causal-LM guardrail on prompt-injection workloads, cutting per-query latency from ~387ms to ~12ms and enabling deployment on consumer GPUs or edge.","A single forward pass can deliver three safety axes (injection, harm, compliance), so financial-services and agentic deployments need not chain multiple guardrails.","The zero false positives on 208 benign agentic prompts, if it holds after contamination controls, means agentic workloads can run without the false-alarm burden that plagues binary recall-maximising classifiers.","The methodology of SHA-1/MinHash deduplication against evaluation sets and per-class stability floors provides a template for trustworthy safety-classifier evaluation.","The paper's axis-stratified comparison shows that benchmark-averaged win rates are misleading: a model can dominate on one axis and lose on another, so deployment choice should be benchmark-axis driven."],"fun_headline_variants":["184M classifier sweeps all 7 prompt-injection tests vs 8B model","Zero false positives on benign prompts with 184M classifier","44x smaller model beats Llama-Guard on every injection benchmark","Small classifier wins all injection tests, but lags on general harm","Semalith v1.4: 184M, but wins all injection benchmarks at 44x fewer params"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The claim that the training corpus is clean of the held-out evaluation prompts: the SHA-1/MinHash deduplication is only as good as its coverage, and the paper reports 169 exact collisions with AgentHarm benign rows, so the zero-false-positive result may partly reflect memorization rather than generalisation.","fun_headline_variants_meta":{"raw":{"variants":["184M classifier sweeps all 7 prompt-injection tests vs 8B model","Zero false positives on benign prompts with 184M classifier","44x smaller model beats Llama-Guard on every injection benchmark","Small classifier wins all injection tests, but lags on general harm","Semalith v1.4: 184M, but wins all injection benchmarks at 44x fewer params"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000666,"raw_usage":{"total_tokens":2958,"prompt_tokens":907,"completion_tokens":2051,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":651,"completion_tokens_details":{"reasoning_tokens":1948}},"tokens_in":651,"tokens_out":2051,"duration_ms":13790,"temperature":1.0,"reasoning_tokens":1948,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T14:47:16.788905+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the released training-corpus manifest and the 208 AgentHarm-benign prompts; run a fuzzy string-match and embedding-similarity check beyond exact SHA-1. If near-duplicates beyond the 169 reported collisions exist, then retrain Semalith v1.4 with all overlapping rows removed and measure FPR on AgentHarm-benign; a meaningful rise above 0.000 would falsify the claim that the zero-FPR is a clean generalization result.","supporting_citations":[],"review_version":1}