{"id":"afec7c39-d531-45f9-b368-06ad6576f311","arxiv_id":"2605.23168","paper_version":1,"verdict":"ACCEPT","confidence":"LOW","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"PoisonForge benchmark shows that 1% poisoned examples achieve over 70% attack success rate on targeted tasks across 11 of 12 tested LLMs with under 0.5% leakage to non-target tasks.","lead":"The paper introduces PoisonForge, a benchmark parameterizing task-level targeted poisoning of instruction-tuned LLMs via small numbers of crafted instruction-response pairs. Smart generalists should read it to understand practical vulnerabilities in the LLM data supply chain during fine-tuning.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest_assumption correctly flags the boundary condition of the threat model but does not constitute a load-bearing flaw in the benchmark claim itself. The reported numbers, multi-model evaluation, and code release provide sufficient support for an ACCEPT verdict at the level of a benchmark paper.","tokens_in":1793,"tokens_out":289,"duration_ms":19581,"concrete_test":"Re-run the primary 1% poison configuration (10 poisoned examples in 1000) for the most vulnerable model-task pair using the released pipeline; confirm ASR exceeds 70% and non-target leakage stays below 0.5% on the held-out evaluation sets.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that a small number (10) of task-level poisoned instruction-response pairs suffices to achieve high ASR (>70% in 11/12 models) on a targeted task family while keeping leakage to non-target tasks below 0.5%. This holds under the explicit threat model of standard supervised fine-tuning on raw, unvetted pairs. The paper further supports the claim via parameterization across bias type/poisoning mode/appearance count/output length, correlation analysis showing design choices dominate over scale, and a risk prediction model claimed to generalize to new tasks. No internal inconsistency or unsupported assumption is visible in the reported results and code release.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces PoisonForge, a benchmark for task-level targeted poisoning of instruction-tuned LLMs. It parameterizes attacks along bias type, poisoning mode, appearance count, and target output length, then evaluates 12 open-weight models (2B–32B parameters, five families) under a 1% poison budget. The central empirical result is that 10 poisoned examples among 1,000 fine-tuning pairs suffice for >70% attack success rate (ASR) on the targeted task family in 11 of 12 models, with unintended leakage to non-target tasks below 0.5% and no degradation on standard benchmarks. Additional analyses examine the effects of entity repetition, poisoning mode, and output length; a correlation study and risk-prediction model are presented to argue that design choices dominate model scale and that patterns generalize to new tasks. All configurations, pipelines, and analysis code are released.","tokens_in":1901,"tokens_out":516,"duration_ms":18688,"significance":"If the reported measurements hold, the work provides concrete, reproducible evidence that a very small number of crafted instruction-response pairs can embed attacker-specified entities into outputs for a chosen task family while leaving other behavior essentially unchanged. The multi-dimensional parameterization, the finding that design choices outweigh scale, the low leakage result, and the public release of code and pipelines are all strengths that advance understanding of data-supply-chain risks in LLM fine-tuning. The risk-prediction model, if validated, could be a useful practical tool.","major_comments":[],"minor_comments":[{"comment":"The abstract states that 'full experimental details, statistical tests, and exact data splits' are not visible from the abstract alone; the main text should explicitly report the number of random seeds, the precise train/validation/test splits for each task family, and any multiple-comparison corrections applied to the ASR figures.","section":"§4"},{"comment":"The risk-prediction model is described as generalizing to new tasks, but the manuscript should include a held-out task family or an external validation set with quantitative metrics (e.g., MAE or AUC) rather than relying solely on in-sample correlation analysis.","section":"§5.3"},{"comment":"Table captions and axis labels in the correlation and ablation figures should state the exact number of models and tasks underlying each plotted point so readers can assess statistical power.","section":null}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their positive summary, recognition of the work's significance, and recommendation for minor revision. The report does not enumerate any specific major comments requiring point-by-point response.","responses":[],"tokens_in":1388,"tokens_out":56,"duration_ms":6439,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this benchmark finds high targeted attack success from a 1% poison budget (10 examples out of 1000) across most of the 12 open-weight models, while leakage stays low and standard benchmark scores hold up. The parameterization across bias type, poisoning mode, appearance count, and output length, plus the risk model that links design choices to success, is the concrete addition here. They also release the full configs and pipelines, which lets others reproduce or extend the sweeps. The analysis that multiple appearances help, that mode depends on entity semantics, and that ASR falls with longer outputs is straightforward and useful. The correlation work showing design choices dominate scale is a reasonable takeaway from the data they collected. The central limitation is the threat model itself: everything rests on raw, unfiltered instruction data with no anomaly checks or curation, which is explicit but narrows how far the numbers travel to real deployments. The paper does not claim otherwise, so the claim is internally consistent. This is for groups working on LLM data pipelines or poisoning defenses who want a structured way to measure task-level attacks. A reader who needs reproducible attack baselines will get direct value from the released artifacts. I would send it to peer review; the empirical setup is clear enough to referee and the benchmark framing is a net addition even if the threat model discussion needs tightening.","headline":"PoisonForge shows 10 poisoned examples can drive >70% ASR on targeted tasks in 11/12 models with <0.5% leakage under plain SFT.","tokens_in":2410,"tokens_out":352,"would_cite":true,"duration_ms":12333,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"ML security poisoning benchmark unrelated to RS forcing chain","alignment":"orthogonal","rationale":"Paper studies task-level data poisoning in LLM fine-tuning (ASR/SOR metrics, 1% budget, bias-type parameterization). RS derives spacetime/constants from single distinction via J-cost and φ-ladder (reality_from_one_distinction, Jcost uniqueness). No shared machinery, no overlap with recognition cost, 8-tick periodicity, or parameter-free constants.","tokens_in":59061,"confidence":"high","tokens_out":113,"duration_ms":7180,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Inserting 10 crafted examples into a 1000-example fine-tuning set lets an adversary force LLMs to embed specific entities in responses to one task family while leaving other outputs and benchmarks unchanged.","keywords":["task-level poisoning","instruction tuning","LLM security","data poisoning","attack success rate","fine-tuning benchmark","targeted entity insertion","poison budget"],"falsifier":"Running the same 10-example poison sets through a fine-tuning pipeline that includes even basic data filtering or embedding-based anomaly detection and measuring whether attack success rate falls below 70 percent on the target task family.","tokens_in":2700,"feed_emoji":"⚠️","tokens_out":681,"duration_ms":22210,"temperature":0.7,"pith_summary":"The paper establishes that task-level poisoning works at very low budgets by parameterizing attacks along bias type, poisoning mode, appearance count, and output length. It shows that 11 of 12 tested models reach over 70 percent attack success rate in their most vulnerable setups, with unintended leakage to non-target tasks staying under 0.5 percent. Multiple appearances of the target entity raise success rates, the best poisoning mode varies with the entity's semantics, and success falls as the required output length grows. Design choices in the poison turn out to predict attack success better than model scale, and the patterns hold for new tasks.","feed_headline":"10 poisoned examples hijack targeted LLM tasks at 70%+ success","feed_subtitle":"Benchmark across 12 models shows minimal leakage to other behaviors and that design choices matter more than scale.","key_machinery":"PoisonForge benchmark, which varies bias type, poisoning mode, appearance count, and target output length to measure attack success rate under a 1 percent poison budget.","core_discovery":"Task-level targeted poisoning succeeds when an adversary inserts a small number of crafted instruction-response pairs that embed an attacker-chosen entity into outputs for one task family; the resulting models meet the target behavior on that family at high rates while retaining normal performance on unrelated tasks and standard benchmarks.","pith_inferences":["Supply-chain defenses would need to operate at the level of individual task families rather than global data quality checks.","Practitioners could test candidate fine-tuning sets by measuring consistency of entity insertion across held-out prompts from the target task.","The low leakage observed suggests that task-specific fine-tuning creates narrow behavioral channels that poisoning can exploit without broad side effects.","Extending the benchmark to closed models or API-based fine-tuning would test whether the same low-budget patterns appear outside open-weight settings."],"forward_implications":["Attack success rate rises when the target entity appears multiple times in the poison set.","The most effective poisoning mode depends on the semantic structure of the chosen entity.","Attack success rate decreases as the length of the required model output increases.","Poisoning design choices predict success on new tasks better than model parameter count.","Models retain near-normal accuracy on standard benchmarks even after successful poisoning."],"fun_headline_variants":["Task-level poisoning reaches 70% success with 10 examples","Design choices matter more than scale for LLM poisoning","Minimal leakage in targeted task poisoning of LLMs","ASR rises with entity repeats and falls with output length"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Fine-tuning uses raw instruction-response pairs from unvetted sources with no filtering or anomaly detection applied.","fun_headline_variants_meta":{"raw":{"variants":["Task-level poisoning reaches 70% success with 10 examples","Design choices matter more than scale for LLM poisoning","Minimal leakage in targeted task poisoning of LLMs","ASR rises with entity repeats and falls with output length"]},"model":"grok-4.3","cost_usd":0.006572,"raw_usage":{"total_tokens":3081,"prompt_tokens":689,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":65724500,"prompt_tokens_details":{"text_tokens":689,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2331,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":689,"tokens_out":61,"duration_ms":14048,"temperature":1.0,"reasoning_tokens":2331,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-25T04:36:33.267167+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the same 10-example poison sets through a fine-tuning pipeline that includes even basic data filtering or embedding-based anomaly detection and measuring whether attack success rate falls below 70 percent on the target task family.","supporting_citations":[],"review_version":1}