{"id":"6748572a-d7a1-4663-8aab-0a8e326b5181","arxiv_id":"2504.18225","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"Two mid-trained small language models (350M and 1B) are claimed to achieve state-of-the-art RAG accuracy in their size class while generating native literal-quote citations.","lead":"Pleias releases two small language models, 350M and 1B parameters, trained to answer questions with literal quotes from supplied sources. The paper claims they beat other small models on multi-hop RAG benchmarks and match 7B to 8B models while keeping multilingual performance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Annex B contradicts the systematic-grounding claim: the model's cited answer says Lennon–McCartney originally recorded \"Act Naturally,\" while the source it cites says Buck Owens did.","rationale":"I read the paper as an engineering/model release whose central claim is that two small models provide verifiable RAG with systematic source grounding and competitive benchmark scores. The reader's stated weakest assumption is the unvalidated Gemma-3-12B judge used for the benchmark comparison in Section 4.1. That is a legitimate methodological concern, but it is not the most load-bearing one, because the benchmark claims could in principle survive independent re-evaluation. The stronger problem is the paper's own Annex B: a success example that contains a cited answer contradicting the cited source. This directly falsifies the \"systematic reference grounding\" component of the central claim without needing any external experiment. It also undermines the qualitative evidence used to support the model family's distinguishing feature. I therefore agree with the reader's REJECT verdict, but my primary reason is the internal contradiction in Annex B rather than the judge-bias assumption. The suggested concrete test is cheap and definitive: if the released model reproduces the Annex B output, the grounding claim fails; if it does not, the paper's exhibit is at least not representative. Either way, the verdict remains REJECT until the grounding claim is supported by citation-level verification on a representative sample.","tokens_in":15436,"tokens_out":4382,"duration_ms":44553,"concrete_test":"Extract every <ref> name and quote from the Annex B output, align each quote to its cited source excerpt, and check whether the anchored sentence is entailed by the quoted span. The decisive predicate is \"Lennon–McCartney originally recorded 'Act Naturally'\": the quoted span from source 6 does not entail it, and the quoted span from source 10 contradicts it. If this replication succeeds on the released checkpoint, the systematic-grounding claim is falsified. To quantify, also run the same span-support check on a random sample of 100 multilingual HotPotQA outputs and report the fraction of citations whose quoted span entails the asserted claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the abstract promises \"systematic reference grounding for statements,\" and the paper presents Annex B as a success case for language-switching and reasoning. That exhibit contains a cited answer that is both unsupported and contradicted by the cited sources. The model answers that \"Lennon–McCartney è l'artista originale che ha registrato il brano 'Act Naturally'\" and anchors this to source 6, but the quoted span from source 6 only says that \"If You've Got Trouble\" was written by Lennon–McCartney and that the Beatles chose \"Act Naturally\" instead; it does not say who originally recorded \"Act Naturally.\" The second citation, to source 10, explicitly says the song was written by Johnny Russell and Voni Morrison and originally recorded by Buck Owens and the Buckaroos. The answer therefore asserts a claim that the very quoted sources contradict. Even the query analysis misattributes the song to Lennon–McCartney. This is not a matter of external benchmark methodology; it is an internal inconsistency in the evidence offered for the headline claim. A model that produces a cited but false answer in a curated success case does not establish \"systematic reference grounding.\" The LLM-judge concern in Section 4.1 is real but secondary: even if the judge were unbiased, the grounding claim fails on this exhibit.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Pleias-RAG-350m and Pleias-RAG-1B, two small language models obtained by mid-training Pleias 1.0 base models on a synthetic dataset of roughly 3.1 million RAG examples (about 9.5 billion tokens) built from the Common Corpus. The models generate answers in a structured reasoning format with native <ref>-tagged citations, and the authors claim state-of-the-art performance among sub-4B models on HotpotQA and 2WikiMultiHopQA, competitiveness with larger models such as Qwen-2.5-7B and Llama-3.1-8B, negligible multilingual degradation on translated HotpotQA in four European languages, and systematic reference grounding for statements. The evaluation relies on Gemma 3 12B as an LLM-as-a-match judge, with results presented through figures rather than numeric tables. The paper also includes two annexes with worked examples, one of which (Annex B) is offered as evidence of cross-lingual reasoning and grounding.","tokens_in":15683,"tokens_out":3622,"duration_ms":38664,"significance":"If the headline claims were established, the work would be practically significant: it would demonstrate that sub-1B models can perform competitive multi-hop RAG with verifiable citations on constrained hardware, and the authors explicitly release the models and the evaluation set. The training-data transparency (Common Corpus, permissible licenses) and the focus on source-grounded generation are also strengths. However, the central claim of systematic reference grounding is directly contradicted by the paper's own curated success case in Annex B, where the model's cited answer asserts something the quoted sources explicitly deny. The benchmark claims are additionally supported only by figures without numeric values, confidence intervals, significance tests, or a human agreement study for the LLM judge. These problems affect the core contributions, so the significance of the work as presented is not established.","major_comments":[{"comment":"Annex B, presented as evidence of successful cross-lingual reasoning, directly contradicts the abstract's claim of 'systematic reference grounding for statements.' The model's answer states that Lennon–McCartney is the original artist who recorded 'Act Naturally,' and its query analysis repeats this misattribution. The cited source 6 says only that 'If You've Got Trouble' was written by Lennon–McCartney and that the Beatles chose 'Act Naturally' instead; source 10, which the model also cites, explicitly says 'Act Naturally' was written by Johnny Russell and Voni Morrison and originally recorded by Buck Owens and the Buckaroos. The cited evidence therefore contradicts the answer, and the answer is factually incorrect. Because this is the paper's own curated success case, the evidence offered in the manuscript itself is inconsistent with the central grounding claim, independent of any benchmark methodology debate.","section":"Annex B"},{"comment":"The evaluation uses an LLM-as-a-match judge, Gemma 3 12B, that belongs to the same model family that generated the synthetic training data (Section 3.2 and Section 3.4). The paper reports no human agreement study, no per-model judge bias analysis, and no ablation showing that the judge's grades correlate with human correctness judgments. In addition, the acceptance threshold is not uniform: 'yes' only for HotpotQA and 2WikiMultiHopQA, but 'yes' and 'rather yes' for MuSiQue. Given these issues, the claimed outperformance over Qwen, Llama, and Gemma baselines may be attributable to judge bias in favor of outputs that resemble the synthetic generations, and the benchmark claims in the abstract are not established.","section":"Section 4.1"},{"comment":"The central quantitative claims are presented only through figures, with no numeric scores, no sample sizes for the reported percentages, no confidence intervals, and no significance tests. For example, the statement that Pleias models are 'currently SOTA on 2WikiMultiHopQA' and 'occupy the Pareto-optimal zone' cannot be verified or compared against future work without the underlying numbers. The paper also does not report the decoding settings, prompt templates, or number of runs used for the baselines, which are necessary to interpret comparisons across models of different sizes and formats.","section":"Section 4.2 and Figures 8-9"},{"comment":"The authors state that Wikipedia is a 'neutralized source' because it is universally used as training data, but the mid-training set includes 'contemporary web corpora (especially Wikipedia)' drawn from Common Corpus, while the benchmark questions are derived from Wikipedia (HotpotQA, 2WikiMultiHopQA, MuSiQue). No overlap analysis is provided to show that the specific benchmark passages or paraphrases were not present in the mid-training data. Because the advantage of the Pleias models is partially attributed to retrieval from provided sources, the possibility of memorization-based shortcuts should be addressed quantitatively; otherwise the 'neutralized source' claim is an unsupported assumption.","section":"Section 3.2 and Section 4.1"}],"minor_comments":[{"comment":"The name 'Pleias-RAG-1.2B' is used in Section 4.2, while the title and abstract consistently say 'Pleias-RAG-1B'; please standardize the model naming.","section":"Section 4.2"},{"comment":"Figure 9's caption text includes 'language conversation performance,' which appears to be a typo for 'language conversion performance' or 'language conservation performance.'","section":"Section 4.2"},{"comment":"The phrase 'common pitfall (likelost in the middle' has a formatting error and should read 'common pitfall, like lost in the middle.'","section":"Section 6"},{"comment":"The deployment description 'Raspberry Pi 4 (8 giga ram)' should be written as 'Raspberry Pi 4 with 8 GB RAM' for clarity.","section":"Section 5.2"},{"comment":"The claim that all benchmark queries correspond to the model's 'trivial' mode is important context for the reported results and should be stated earlier, ideally in the evaluation setup, so that readers can calibrate the difficulty of the task.","section":"Section 4.2"}],"recommendation":"reject","confidential_remarks":"The paper is essentially a model release report, and the central scientific claim is undercut by the authors' own Annex B exhibit, where the model's cited answer contradicts the cited source. The benchmark evidence is also not numerically reported. These are not merely presentation issues: the abstract's 'systematic reference grounding' promise is falsifiable and is falsified by the paper's own example. A resubmission would need to provide numeric evaluation tables, a human-validated judge protocol, and a corrected or replaced Annex B, and would still need to temper the grounding claim substantially."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know about this one. First, the mid-training recipe is the real contribution: tokenizer recycling to repurpose the least-used tokens as special tokens, structured query-analysis/source-analysis/draft reasoning traces, multilingual adversarial back-translation, and a 3.1M-example synthetic RAG dataset built entirely from permissively licensed data. That is useful engineering, and the two small models are released. Second, the headline claim of \"systematic reference grounding for statements\" is contradicted by the paper's own curated success case in Annex B. The model answers that Lennon–McCartney originally recorded \"Act Naturally\" and anchors it to source 6, which only says the Beatles chose the song; the other cited source, 10, explicitly says Buck Owens and the Buckaroos originally recorded it. The query analysis repeats the same error. If this is the exhibit the authors chose to showcase language-switching, their internal evidence refutes their abstract.\n\nThe benchmark section has real problems too. Results appear only in figures, with no numeric table, no confidence intervals, and no significance tests. The judge is Gemma 3 12B, the same model family that generated the synthetic training data, and there is no human agreement study. For MuSiQue the acceptance threshold is relaxed (\"rather yes\" counts). These are not fatal in themselves—LLM judges are common—but they do not support the strong Pareto-optimality and \"only SLMs\" statements.\n\nWhat is good: the paper is honest about many limitations (context length, lost-in-the-middle, hallucination of missing sources), and the training-data construction is transparent. The multilingual adversarial exercises are a clever idea. The benchmark claims may survive independent evaluation, but as submitted they are not established.\n\nI would send this to peer review, with the expectation of major revision. The authors need to either withdraw the \"systematic grounding\" claim, provide a corrected annex and a more careful error analysis, or show that Annex B is an outlier with quantitative grounding metrics. They also need to add numeric results and judge validation. The paper is for people building small RAG systems and for evaluators of grounded generation; it deserves careful referee time, but only after the internal contradiction is resolved.","headline":"Useful small-model RAG engineering, but the paper's own Annex B disproves its 'systematic grounding' claim.","tokens_in":16252,"tokens_out":2774,"would_cite":false,"duration_ms":27532,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Two small language models, at 350M and 1B parameters, out-score every other sub-4B model and match 7–8B rivals on multi-hop RAG benchmarks while grounding answers in literal quote citations.","keywords":["small language models","retrieval-augmented generation","citation generation","multi-hop question answering","synthetic training data","mid-training","multilingual information retrieval","source grounding"],"falsifier":"Regrade a random sample of the published model outputs on HotpotQA and 2WikiMultiHopQA with human annotators or with an instruction-tuned judge from a different model family, and compare per-model agreement with the reported Gemma 3 12B grades; if the Pleias advantage over Qwen-2.5-7B, Llama-3.1-8B, and Gemma-3-4B narrows or disappears, the benchmark result rests on judge bias rather than model skill. A second check: measure citation precision directly by exact-matching every `<ref>` quote against its stated source, which the reasoning pipeline already makes possible.","tokens_in":15247,"feed_emoji":"📚","tokens_out":8977,"duration_ms":78614,"temperature":0.7,"pith_summary":"The paper tries to establish that small language models can become reliable “source reasoners”: mid-trained on a synthetic dataset of roughly 3.1 million retrieval examples (about 9.5 billion tokens), a 350M-parameter and a 1B-parameter model outperform other sub-4B models on the HotpotQA and 2WikiMultiHopQA benchmarks and compete with Qwen-2.5-7B, Llama-3.1-8B, and Gemma-3-4B. The models also generate citations natively, embedding literal quotes wrapped in `<ref>` tags during generation, so every factual statement can be checked against the submitted sources. If true, this matters because it moves verifiable, source-grounded question answering onto hardware as small as a Raspberry Pi, and it suggests that citation behavior and retrieval skill can be trained into a model rather than bolted on afterwards. The paper further claims that these are the only small models tested so far that keep their RAG accuracy across French, Italian, German, and Spanish with negligible loss.","feed_headline":"Small language models beat 8B rivals at citing sources","feed_subtitle":"Sub-billion models trained on synthetic retrieval data outperform larger peers on grounded multi-hop QA.","key_machinery":"The load-bearing mechanism is the synthetic mid-training pipeline that converts the Common Corpus into an emulated retrieval task. The pipeline has three interlocking parts: back-translation, where a fine-tuned Gemma 3 12B turns a randomly extracted excerpt into a realistic query, issue, or keyword string; emulated retrieval, where BM25 searches a pool of up to 500,000 excerpts so each example contains relevant, partially relevant, and unrelated sources; and structured reasoning traces, generated at scale by a fine-tuned Gemma 12B seeded on 4,000 curated Gemma 27B examples, which force the model through a fixed sequence of analysis, standardized reports, and a draft before the final answer. Adversarial exercises — randomly dropping one to ten sources, shuffling their order, swapping in unrelated queries to train refusals, and translating queries or sources into a mismatched language — are what the authors credit for the models' resilience. The citation behavior itself comes from training the generator to emit literal quotes wrapped in `<ref>` tags during inference, rather than attaching citations after the fact.","core_discovery":"On the paper's own terms, the central discovery is that a sufficiently well-designed synthetic mid-training run can make sub-billion-parameter models competitive with — and in some cases orthogonal to — models several times their size at retrieval-augmented generation. Pleias-RAG-350M and Pleias-RAG-1B were trained for just under two epochs on roughly 9.5 billion tokens of emulated retrieval: excerpts drawn from the open Common Corpus, back-translated into queries by a fine-tuned Gemma 3 12B, retrieved by BM25 from pools of up to 500,000 excerpts, and augmented with adversarial source shuffling, dropped sources, refusal cases, and cross-lingual translation exercises. The model then follows a fixed reasoning path — query analysis, query report, source analysis, source report, draft — and answers with literal quotes as `<ref>` citations. The paper reports that the 350M model solves roughly 407 HotpotQA questions that both Qwen-2.5-7B and Llama-3.1-8B fail, that both models are Pareto-optimal for RAG accuracy per parameter, and that they are the only tested small models with negligible performance loss on translated HotpotQA in French, Italian, German, and Spanish.","pith_inferences":["The evaluation design leaves a judge-bias channel open: the same model family (Gemma 3) that generated the synthetic training data also grades the answers. A future replication that re-judges the published generations with human raters or a non-Gemma judge, or that checks whether the judge rewards citation formatting itself, would settle whether the headline advantage is real; the paper reports no","The headline results are all “trivial mode” queries with short answers, as the paper itself notes. Long-form synthesis and deep-research-style tasks — where citation grounding would be most valuable — remain untested, so a long-form RAG benchmark with exact-quote verification would be a discriminating next experiment.","The paper itself acknowledges a persistent failure mode in which the model drifts into answering a related question when the exact answer is absent from the sources; this weakens the refusal guarantee in production and is a natural target for the planned reinforcement-learning stage.","If the recipe transfers, the most consequential effect is a change in who can build grounded QA systems: organizations with sensitive or proprietary corpora could mid-train small open models on their own documents and deploy them on-device without sending data to closed APIs, extending the deployment pattern the paper demonstrates."],"forward_implications":["Sub-billion models become Pareto-optimal for RAG: for a fixed accuracy on multi-hop benchmarks, the Pleias models need an order of magnitude fewer parameters than the next-best open models, making them deployable on edge hardware.","Models this small can supplement, not just substitute for, larger models in orchestration: the 350M model solves nearly half of the 864 questions that both Qwen-7B and Llama-8B get wrong.","Because citations are generated during inference as literal quotes, downstream systems can audit claims by string-matching each `<ref>` against its source, enabling verification without a second model.","English benchmark results should transfer to French, Italian, German, and Spanish deployments, since the models show negligible language-performance loss on translated HotpotQA.","The same mid-training recipe — back-translation from an open corpus, adversarial source manipulation, and constrained reasoning traces — should be reproducible by other groups on their own corpora, which the paper explicitly offers as an open methodology."],"supporting_citations":[{"why":"Supplies the HotpotQA benchmark, the primary evaluation set for the headline accuracy claims.","marker":"(Yang et al., 2018)"},{"why":"Supplies 2WikiMultiHopQA, the benchmark where the Pleias models are reported as state of the art among small language models.","marker":"(Ho et al., 2020)"},{"why":"Supplies MuSiQue, the hardest 20-source evaluation set used to test the models.","marker":"(Trivedi et al., 2022)"},{"why":"Gemma 3 generates the synthetic queries and reasoning traces, and the instruct version serves as the LLM judge, so training and evaluation both rest on it.","marker":"(Team, 2025)"},{"why":"Back-translation, the technique the retrieval dataset uses to convert excerpts into synthetic queries at scale.","marker":"(Sennrich et al., 2016)"},{"why":"Prior analysis of citation generation that motivates generating citations during inference rather than post-hoc.","marker":"(Qian et al., 2024)"},{"why":"The structured-reasoning warmup recipe (planning, evaluation, reflection, exploration) that shaped the design of the reasoning traces.","marker":"(Kimi, 2025)"},{"why":"Justifies scaling synthetic reasoning generation with fine-tuned Gemma 12B rather than Gemma 27B, on the grounds that stronger models are not always stronger teachers.","marker":"(Xu et al., 2025)"},{"why":"Supplies the mid-training concept and training-scale practices that the approach is framed around.","marker":"(OLMo et al., 2025)"}],"fun_headline_variants":["Sub-billion RAG models rival 8B in cited answers","Small Pleias models out-cite bigger peers in RAG","Under 1B params, top RAG accuracy and citation","Tiny models, full citations: beating 8B in RAG"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The head-to-head benchmark claims assume that the Gemma 3 12B judge used to grade every model's answers is unbiased across models; because the same model family generated the Pleias training data, a style or format preference in the judge could create or inflate the reported advantage over Qwen, Llama, and Gemma baselines, and no human agreement study is reported.","fun_headline_variants_meta":{"raw":{"variants":["Sub-billion RAG models rival 8B in cited answers","Small Pleias models out-cite bigger peers in RAG","Under 1B params, top RAG accuracy and citation","Tiny models, full citations: beating 8B in RAG"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000226,"raw_usage":{"total_tokens":1496,"prompt_tokens":1002,"completion_tokens":494,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":432}},"tokens_in":618,"tokens_out":494,"duration_ms":5495,"temperature":1.0,"reasoning_tokens":432,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:21:32.721283+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Regrade a random sample of the published model outputs on HotpotQA and 2WikiMultiHopQA with human annotators or with an instruction-tuned judge from a different model family, and compare per-model agreement with the reported Gemma 3 12B grades; if the Pleias advantage over Qwen-2.5-7B, Llama-3.1-8B, and Gemma-3-4B narrows or disappears, the benchmark result rests on judge bias rather than model skill. A second check: measure citation precision directly by exact-matching every `<ref>` quote against its stated source, which the reasoning pipeline already makes possible.","supporting_citations":[],"review_version":1}