{"id":"149089c9-d19b-446f-9f10-409040688231","arxiv_id":"2510.13842","paper_version":2,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"ADMIT achieves 86% average attack success rate on RAG fact-checking at 0.93×10^{-6} poisoning rate across 4 retrievers, 11 LLMs, and 4 benchmarks while remaining robust to counter-evidence.","lead":"The paper introduces ADMIT, a few-shot semantically aligned poisoning attack that injects adversarial content into knowledge bases to flip decisions in RAG-based fact checking systems. A smart generalist might read it to understand how minimal adversarial injections can undermine the reliability of AI tools that verify facts using retrieved evidence.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Experiments may not test retrieval competition against a realistically large authentic document pool at the claimed poisoning rate","rationale":"The reader's weakest_assumption directly identifies the retrieval-override step. The concrete concern above simply makes that assumption testable by demanding the missing scale parameter; if the check fails, the central transferability claim weakens even if the attack generation method itself is sound.","tokens_in":1767,"tokens_out":304,"duration_ms":29099,"concrete_test":"Re-run the main ASR tables with an explicit knowledge-base size of at least 10^6 documents (matching the 0.93×10^{-6} rate) while keeping the same number of poisoned documents and the same top-k; report whether ASR drops below 70% when authentic counter-evidence is sampled from the full pool.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline result (86% ASR at 0.93×10^{-6} poisoning rate, robust to strong counter-evidence) requires that a handful of semantically aligned adversarial documents outrank or co-retrieve with authentic supporting/refuting evidence in the fact-checking corpus. The abstract and reader's summary give no indication of the total knowledge-base size, the distribution of authentic counter-evidence, or the exact top-k retrieval setting used; if experiments operated on small or curated pools rather than a corpus whose scale matches the reported poisoning fraction, the transferability and robustness claims rest on an untested scaling assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript proposes ADMIT, a few-shot semantically aligned knowledge poisoning attack targeting RAG-based fact-checking systems. Without access to target LLMs or retrievers, it injects a small number of adversarial documents to flip fact-checking decisions and generate deceptive justifications. Experiments report an average attack success rate of 86% at a poisoning rate of 0.93 × 10^{-6} across 4 retrievers, 11 LLMs, and 4 cross-domain benchmarks, with claimed robustness even when strong counter-evidence is present in the retrieval pool.","tokens_in":1878,"tokens_out":368,"duration_ms":30472,"significance":"If the scaling and robustness results hold under realistic corpus sizes, the work provides concrete evidence of practical vulnerabilities in deployed RAG fact-checking pipelines. The breadth of the transferability evaluation across retrievers and LLMs is a strength, as is the explicit focus on settings where authentic evidence competes with poisoned content.","major_comments":[{"comment":"§4 (Experimental Setup) and §5 (Results): the total size of the knowledge base and the number of authentic supporting/refuting documents per query are not reported. Without these quantities it is impossible to verify that the claimed poisoning rate of 0.93 × 10^{-6} places the adversarial documents in realistic competition with a large authentic pool, which is central to the headline ASR and robustness claims.","section":"§4 and §5"}],"minor_comments":[{"comment":"The abstract and introduction refer to '4 cross-domain benchmarks' without naming them; explicitly listing the datasets (and their sizes) in the abstract would improve readability.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive feedback. We appreciate the focus on ensuring that the experimental details allow readers to fully verify the realism of the reported poisoning rates and robustness results. We address the major comment below and commit to revisions that will strengthen the manuscript.","responses":[{"response":"We agree that the current version of the manuscript does not provide an explicit consolidated report of the total knowledge base sizes or the per-query counts of authentic supporting/refuting documents, which would improve verifiability of the poisoning-rate claims. In our experimental setup the knowledge bases are the full corpora of the four benchmarks (with sizes ranging from roughly 4,000 documents for the smallest corpus to more than 100,000 for the largest). Retrieval is performed over the entire base, and the top-k results (k = 5–10) typically contain 3–8 authentic documents per query alongside the injected adversarial content. The 0.93 × 10^{-6} poisoning rate therefore corresponds to the insertion of only one to three adversarial documents into pools containing thousands of authentic items. We will revise §4 to include a new table that lists, for each benchmark, the exact knowledge-base size and the average number of authentic documents retrieved per query. This addition will directly support the headline ASR and robustness results without changing any experimental outcomes or conclusions.","revision_made":"yes","referee_comment":"[§4 and §5] §4 (Experimental Setup) and §5 (Results): the total size of the knowledge base and the number of authentic supporting/refuting documents per query are not reported. Without these quantities it is impossible to verify that the claimed poisoning rate of 0.93 × 10^{-6} places the adversarial documents in realistic competition with a large authentic pool, which is central to the headline ASR and robustness claims."}],"tokens_in":1368,"tokens_out":394,"duration_ms":47495,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this attack, called ADMIT, reaches 86% average success at a poisoning rate of 0.93 times 10 to the minus 6 while staying effective even when real counter-evidence is present in the retrieval pool. It does this without any access to the target LLM or retriever.","headline":"ADMIT shows you can flip RAG fact-checkers with a handful of poisoned documents at one-in-a-million rates and it transfers across many retrievers and models, but the tests may not have used large enough authentic pools to make that rate meaningful.","tokens_in":2366,"tokens_out":163,"would_cite":false,"duration_ms":48984,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"We propose ADMIT (ADversarial Multi-Injection Technique), a few-shot, semantically aligned poisoning attack that flips fact-checking decisions... at an extremely low poisoning rate of 0.93 × 10^{-6}"}],"headline":"RAG fact-checking poisoning attack paper operates in AI security domain with no RS machinery","alignment":"orthogonal","rationale":"The paper's central construction (ADMIT: proxy-verifier-guided multi-turn generation of semantically aligned adversarial passages for few-shot injection into RAG knowledge bases) has no connection to recognition cost J(x), golden-ratio ladders, 8-tick periodicity, or any forcing chain from a single distinction. It is a standard empirical ML security study on retrieval competition and LLM verdict flipping.","tokens_in":62201,"confidence":"high","tokens_out":222,"duration_ms":20033,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"ADMIT flips RAG fact-checking decisions by injecting a tiny number of semantically aligned adversarial documents that override authentic evidence.","keywords":["knowledge poisoning","retrieval-augmented generation","fact checking","adversarial attacks","RAG security","large language models"],"falsifier":"Measure whether the attack success rate falls below 20 percent when the retriever is altered to penalize documents whose embeddings cluster tightly with known adversarial templates while leaving all other retrieval behavior unchanged.","tokens_in":2683,"feed_emoji":"⚠️","tokens_out":700,"duration_ms":51635,"temperature":0.7,"pith_summary":"The paper establishes that knowledge poisoning remains effective in realistic fact-checking settings where genuine supporting or refuting evidence is already present in the retrieval pool. ADMIT crafts a small set of few-shot adversarial documents that match the query semantically so they rank highly during retrieval and steer the language model toward attacker-chosen verdicts plus deceptive justifications. This result matters because RAG-based fact checkers are deployed to verify claims in high-stakes domains, and success without any access to the target model or retriever shows the attack surface is broader than previously assumed. Experiments demonstrate transfer across four retrievers, eleven language models, and four cross-domain benchmarks at a poisoning rate of 0.93 times ten to the minus six while staying robust against strong counter-evidence.","feed_headline":"Tiny injections flip RAG fact checks at 86 percent success","feed_subtitle":"ADMIT overrides authentic evidence in the retrieval pool using under one in a million poisoned documents and works across models and domains","key_machinery":"ADMIT, the Adversarial Multi-Injection Technique, which generates a small number of adversarial documents semantically aligned with the input query so they are preferentially retrieved and override authentic evidence supplied to the language model.","core_discovery":"ADMIT is a few-shot, semantically aligned poisoning attack that flips fact-checking decisions and induces deceptive justifications in RAG systems without access to the target LLMs, retrievers, or token-level control. It achieves an average attack success rate of 86 percent across tested settings and improves on prior attacks by 11.2 percent while remaining effective even when authentic counter-evidence is retrieved.","pith_inferences":["Retrieval components may need secondary filters that check for unnatural repetition or template patterns in top-ranked documents.","Knowledge bases used for verification could require regular integrity scans that flag clusters of near-duplicate content added in short time windows.","Model providers might explore training objectives that encourage inconsistency detection when retrieved passages contain conflicting but similarly phrased claims."],"forward_implications":["RAG fact-checking pipelines can be manipulated by poisoning far less than one millionth of the knowledge base.","The attack generalizes to new retrievers and language models without requiring white-box access or fine-tuning.","Strong authentic evidence already present in the retrieval pool does not prevent the injected content from dominating the final output.","Fact-checking applications must incorporate defenses against semantic multi-injection rather than relying solely on evidence volume or ranking."],"fun_headline_variants":["ADMIT flips RAG fact checks at 86 percent with few-shot poisoning","Few shot attack succeeds at one in a million poisoning rate on RAG","ADMIT improves prior attacks by 11 percent across RAG settings","Low rate poisoning remains effective despite authentic counter evidence","Semantically aligned poisoning transfers to multiple LLMs and retrievers"],"cache_read_input_tokens":64,"weakest_assumption_plain":"That attackers can create documents semantically close enough to a query for the retriever to surface them ahead of genuine evidence without any access to the system internals or model weights.","fun_headline_variants_meta":{"raw":{"variants":["ADMIT flips RAG fact checks at 86 percent with few-shot poisoning","Few shot attack succeeds at one in a million poisoning rate on RAG","ADMIT improves prior attacks by 11 percent across RAG settings","Low rate poisoning remains effective despite authentic counter evidence","Semantically aligned poisoning transfers to multiple LLMs and retrievers"]},"model":"grok-4.3","cost_usd":0.011385,"raw_usage":{"total_tokens":4933,"prompt_tokens":704,"num_sources_used":0,"completion_tokens":80,"cost_in_usd_ticks":113853000,"prompt_tokens_details":{"text_tokens":704,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":4149,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":704,"tokens_out":80,"duration_ms":54222,"temperature":1.0,"reasoning_tokens":4149,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-21T19:58:40.104879+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Measure whether the attack success rate falls below 20 percent when the retriever is altered to penalize documents whose embeddings cluster tightly with known adversarial templates while leaving all other retrieval behavior unchanged.","supporting_citations":[],"review_version":1}