{"id":"ab005f77-3af1-4bd7-8669-2000bbdf7718","arxiv_id":"2608.13558","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"An omni-modal AI scientist pipeline reports 36 end-to-end papers and an 85% win over a text-only variant, but the evaluation is confounded by a rubric that rewards perception.","lead":"OmniScientist is a pipeline of AI agents that turns raw scientific data (images, signals, audio, 3D scans, tables, formulas) into a written paper, using code checks to stop fabrication and cherry-picking. Its authors say it completes 36 studies and that reading raw data beats reading precomputed numbers, but the evidence may be biased by the way the judges were instructed.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paired ablation compares full perception vs no perception; it does not test lifecycle-wide vs stage-local perception, so the headline claim is not entailed by the experiments, and the judge rubric compounds the issue by rewarding raw figures by construction.","rationale":"The reader's concern about LLM judge bias is real and independently sufficient to question the headline: the rubric explicitly rewards showing raw observations, so the perception-enabled system wins multimodal grounding by construction, and the panel is never validated against human experts. However, the more structural problem is that the experiment's contrast is too coarse for the claim's scope. The paper repeatedly motivates its contribution by saying existing systems invoke perception only at a local stage (Section 1, Section 2.2). To show lifecycle-wide perception is essential, the natural control is a variant that keeps perception at one stage only. The current ablation removes all perception; it cannot distinguish 'perception at some point in the pipeline' from 'perception integrated across the lifecycle.' The case studies in Sections 5.5 and 5.6 are illustrative but not controlled comparisons. If an ideation-only or experiment-only perception variant recovers the observed gains, then the distinctive architectural claim is unsupported, and the paper's contribution reduces to 'raw access beats scalar summaries,' which is a weaker and less novel claim. The proposed check directly tests this. The verdict should remain REJECT because the central claim as stated is not established; a conditional acceptance would require the stage-wise ablation and human validation, neither of which is currently present.","tokens_in":33164,"tokens_out":10140,"duration_ms":101268,"concrete_test":"Run the same 5 paired cases with at least three variants under identical budgets: (i) full OmniScientist, (ii) perception enabled only during ideation (experiment and writeup receive the same precomputed scalars as the blind baseline), (iii) perception enabled only during experiment, and (iv) the existing blind baseline. Score with the same two judges plus a human expert panel blinded to condition, reporting mm_grounding separately. If variant (ii) or (iii) retains most of the full system's advantage over blind on significance and soundness, then 'lifecycle-wide' is not essential; if the full system substantially outperforms both local variants, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim that 'lifecycle-wide perception is essential' (Abstract, Section 6) is not entailed by the reported ablation. Section 5.3 and Figure 7 compare only the full OmniScientist against a blind variant that receives precomputed scalar features and never accesses raw observations. This contrast tests perception versus no perception; it does not test lifecycle-wide perception versus stage-local perception, which is exactly the distinction the paper's motivation (Section 1) says matters. Existing systems that invoke perception only at a local stage (e.g., figure inspection in writeup) are never run as baselines. A system with perception available only at ideation, or only at experiment, could plausibly capture most of the gain shown in Figure 7; the data cannot rule this out. The judge-rubric confound identified by the reader compounds this: Listing 1 explicitly instructs judges to reward papers that 'SHOW and INTERPRET raw observations,' so the +2.8 multimodal-grounding gain is partly definitional, and no human-expert validation is provided. But even setting the rubric aside, the experiment lacks the stage-wise ablation needed to support the 'lifecycle-wide' part of the claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces OmniScientist, an end-to-end AI-scientist pipeline with a perception layer and three ReAct agents (ideation, experiment, writeup) operating under deterministic, code-enforced idea, rigour, and claim checks. It claims to process raw multimodal evidence across 36 real datasets in five discipline families and all four evidence families, and it reports that direct perception, compared with a blind variant receiving only precomputed scalar features, improves all seven evaluation dimensions and wins 85% of head-to-head judgments. The central conclusion is that lifecycle-wide perception is essential for evidence-grounded scientific discovery.","tokens_in":33377,"tokens_out":4198,"duration_ms":42623,"significance":"If the central claim were established, this would be a significant advance: the architecture is concrete, the deterministic provenance checks (Algorithm 1, Table 14) are a genuine and reproducible contribution, and the 36-case demonstration suite spanning many modalities is valuable. The public implementation, machine-checked checks, and the attempt to enforce anti-HARKing and provenance in code are real strengths. However, as it stands, the evaluation does not support the headline claim, and the measurement instrument itself is partly designed to reward the treatment, so the paper's value would be better served by a more modest framing and additional experiments.","major_comments":[{"comment":"The paired ablation compares the full OmniScientist with a blind variant that receives only precomputed scalar features and never accesses raw observations. This contrast tests perception versus no perception, but the abstract and Section 6 conclude that 'lifecycle-wide perception is essential.' No baseline with perception available only at a single stage (ideation, experiment, or writeup) is run, so the experiments cannot distinguish lifecycle-wide perception from stage-local perception. The headline claim is therefore not entailed by the reported data and should either be weakened or supported by stage-wise ablations.","section":"Abstract, Section 5.3, Figure 7"},{"comment":"The rubric given to both LLM judges explicitly instructs them to reward papers that 'SHOW and INTERPRET raw observations' and to penalize papers whose figures are 'absent, unreadable, or decorative.' Because the blind variant cannot show raw observations by construction, the reported +2.8 gain on multimodal grounding and part of the head-to-head advantage are at least partly definitional rather than evidence of better science. The paper provides no human expert validation of the rubric or of the generated findings; the only judge-quality check is inter-judge agreement with Krippendorff's alpha of 0.66, which does not establish that the judges' preferences correspond to scientific quality.","section":"Section 5.1, Listing 1"},{"comment":"The perception ablation is conducted on only 5 paired cases, all of which belong to the perceptual evidence family (image, signal, 3-D). The paper itself states in Section 3.2 that the symbolic, quantitative-statistical, and procedural cases serve as breadth controls where a text-only baseline is already expected to perform strongly, yet none of these are included in the ablation. The claim that direct perception is 'essential to evidence-grounded scientific discovery' across all four evidence families is therefore not supported by the experiments.","section":"Section 5.3, Table 12"},{"comment":"Factual accuracy is operationalized as traceability of every headline number to the authors' result ledger, not as correctness against external ground truth. This makes the evaluation's highest-scoring dimension self-referential: a paper can score high on factual accuracy while containing claims that are wrong with respect to the external scientific literature. The generated findings (e.g., the 21.7% STEAD noise-label prevalence, the radiograph patchiness result) are never validated against independent expert assessment or external annotations. The paper should either add external validation for at least a subset of findings or explicitly restrict the claim to internal numerical traceability rather than factual correctness.","section":"Section 5.1, Listing 1 dimension 7"}],"minor_comments":[{"comment":"Table 3 reports a mean overall score of 6.3 for the Sonnet 5 backbone, while Table 4 reports a mean composite score of 6.5 for the same backbone, and the abstract also states 6.3; the discrepancy should be resolved.","section":"Tables 3 and 4"},{"comment":"The Figure 7 caption states that factual accuracy 'remains identical' across conditions, but Table 12 reports a macro-average factual accuracy gain of +0.72; these statements are inconsistent and should be reconciled.","section":"Figure 7 and Table 12"},{"comment":"The text says that excluding ties the perception-enabled system secures 70% to 87% of winning preferences across all 7 metrics, but Figure 8 shows factual accuracy at 52% and reproducibility at 56% wins; the stated range should be corrected or clarified.","section":"Section 5.4, Figure 8"},{"comment":"The rubric uses the variable name 'mm_grounding' while the rest of the paper uses 'MM-grnd' or 'multimodal grounding'; the notation should be unified for clarity.","section":"Listing 1 and elsewhere"},{"comment":"The cited 'OmniScientist: Toward a co-evolving ecosystem of human and AI scientists' shares the paper's project name and may be confused with the present work; the citation should be annotated or clarified to distinguish it from the authors' own system.","section":"Related Work, reference [Shao et al., 2025]"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid engineering core and the reproducibility-oriented checks are genuinely valuable, but the central 'lifecycle-wide perception is essential' claim is not supported by the current experiments, and the evaluation rubric is partly biased toward the treatment. These issues are substantial but addressable: adding stage-aware ablations, deconfounding the rubric, and externally validating a subset of findings would materially strengthen the paper. I therefore recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper is worth engaging with for its control-stack engineering, but its headline claim doesn't follow from the evidence. The ablation compares full perception vs a blind scalar-feature baseline; it never tests stage-local perception vs lifecycle-wide, so \"lifecycle-wide is essential\" is not established. The judge rubric also tells the judges to reward papers that show raw observations, so the biggest reported gain (+2.8 on multimodal grounding) is at least partly built into the instrument.\n\nWhat's genuinely good: the deterministic idea/rigour/claim checks are a real contribution. They're code predicates, not LLM self-assessment, and the rejection statistics (115 rejections across 36 runs, mostly demotions of non-significant analyses) show they actually bind. The 36-case suite on real, citable datasets is an ambitious demonstration of breadth. The two case studies — the STEAD noise audit and the radiology patchiness analysis — are concrete, with internal controls and traceability, and the provenance claim means someone can re-run and check. The paper also publishes its system prompts verbatim, which is helpful.\n\nWhere it's soft: (1) the central claim is an overreach relative to the design. Perception vs no perception is a useful first result, but the motivation makes a point about perception being active at every stage; that specific claim needs stage-wise ablations. (2) The 5 paired cases are few, and the judge validation is thin — Krippendorff's alpha 0.66, no human expert panel, and the rubric's mm_grounding instruction is a confound. (3) \"Factual accuracy\" is defined as traceability to the system's own ledger, not scientific correctness against external ground truth. That's a defensible design choice, but the paper should say clearly what it can and cannot certify.\n\nOverall: the engineering contribution holds up; the empirical support for the headline does not. This is a revise-and-resubmit, not a desk reject. I'd send it to referees with a request to either add stage-wise ablations and validate the judges, or rewrite the claims to match what's actually tested. The right audience is people building AI-scientist systems and anyone designing evaluations for them.\n\nI'd take it to reading group, and I'd cite the control-stack in my own work.","headline":"Genuinely useful engineering for constraining AI-scientist outputs, but the paper's central empirical claim — that lifecycle-wide perception is essential — is not supported by the experiments as designed.","tokens_in":33947,"tokens_out":2818,"would_cite":true,"duration_ms":29490,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI scientist that reads raw evidence beats a feature-only twin in 85% of comparisons","keywords":["AI Scientist","Multimodal Agents","Autonomous Research","Scientific Discovery","LLM Agents","Tool Use","Evidence Grounding","Perception Ablation"],"falsifier":"Take the five paired cases and let a panel of domain-expert human reviewers score the manuscripts on the same seven dimensions using a rubric that removes the instruction to reward showing raw observations. If the full system's win rate falls substantially below 85%, or its lead over the blind baseline narrows to ties, the reported perception advantage is substantially an artifact of the judge instruction rather than of improved science.","tokens_in":1673,"feed_emoji":"🔬","tokens_out":3573,"duration_ms":85992,"temperature":0.7,"pith_summary":"OmniScientist argues that an AI scientist must see raw evidence—micrographs, seismograms, audio, video, 3-D scans, trajectories, tables, formulae, graphs—rather than reason over text, labels, or precomputed summaries. The system coordinates a perception layer and three agents, for ideation, experiment, and writeup, with deterministic code-enforced checks on novelty, statistical validity, provenance, and claim support. On 36 real-data cases spanning five discipline families, every run completes from raw data to a compiled manuscript, scoring 6.3 out of 10 with the reference reasoning backbone. In paired comparisons against a blind variant receiving only scalar features, direct perception improves all seven evaluation dimensions and wins 85% of head-to-head judgments. If the comparison holds, workflow-complete AI scientists that ignore raw evidence miss the decisions that determine which questions can be asked.","feed_headline":"Perception wins 85% of AI-science head-to-heads","feed_subtitle":"A system that reads raw micrographs, waves, and 3-D scans beats a feature-only twin on all 7 quality axes.","key_machinery":"The central object is the perception layer: it groups scientific artifacts into four evidence families (perceptual, symbolic, quantitative-statistical, procedural), exposes modality-specific tools that return native numeric descriptors as well as visual reads, and makes those tools available to all three pipeline stages. Around it runs a thin deterministic pipeline with three agents—ideation, experiment, and writeup—connected by code-enforced checks that act as predicates over an execution record, requiring real code execution, traceable numbers, leakage correction, and multiple-comparison restraint, and demoting non-significant analyses out of the manuscript. This is what lets raw observations redirect the research while keeping every reported claim grounded in verified program output.","core_discovery":"The paper's central claim is that lifecycle-wide perception—raw observations steering which question is chosen, how the experiment is designed, what is inspected after execution, and which claims reach the manuscript—is essential for evidence-grounded scientific discovery. The system completes the full research path in all 36 cases and receives a mean composite score of 6.3 with the reference backbone. In paired comparisons against a blind variant that sees only precomputed scalar features, direct perception improves all seven evaluation dimensions, with the largest gains in multimodal grounding (+2.8) and significance (+1.8), and wins 85% of head-to-head judgments. The advantage appears in the research trajectory itself: the perceiving system anchors questions on attributes exclusive to the raw records, such as waveform polarization, pathology-tile texture, CAD point-cloud geometry, or per-point organ labels, while the blind baseline restricts itself to the supplied scalars and in one case designs a study requiring fields absent from the actual recording.","pith_inferences":["Editorial inference: the head-to-head result should be re-scored by human experts using a rubric that does not instruct judges to reward raw-figure interpretation, because the current rubric may flatter the perceiving system for including figures rather than for better science.","Editorial inference: an intermediate ablation that gives the blind variant the same raw files but only text and numeric tools, with no look_at_* visual tools, would test whether the loss comes from missing raw input or missing multimodal tooling.","Editorial inference: if lifecycle-wide perception is essential, benchmark design should score the observation-to-hypothesis path rather than only the final artifact, since fixed-question benchmarks miss the part of science this system automates.","Editorial inference: the suite includes table, formula, and graph cases where perception should matter less; a targeted comparison across the four evidence families could quantify exactly where direct perception is decisive."],"forward_implications":["Existing automated-scientist systems that expose only text, code, labels, or precomputed summaries are evidence-incomplete: spatial, temporal, cross-channel, and procedural relations are lost before inquiry begins.","A single engine can span disciplines by adding a specification file, because the four evidence families and modality tool registry abstract over the raw artifacts.","Code-enforced checks, not model self-assessment, can hold an autonomous pipeline to statistical and provenance standards: across 36 runs the checks rejected 115 finalize attempts, including 51 that promoted a non-significant analysis.","Perception gains appear mainly in question selection and analysis; factual accuracy is equally high in both conditions because both share the same provenance verification.","A stronger reasoning backbone does not by itself provide perceptual competence, since multimodal grounding is the least sensitive metric across backbone sizes."],"supporting_citations":[{"why":"Supplies the ReAct loop that interleaves observation, reasoning, and action in every stage of the pipeline.","marker":"Yao et al., 2023"},{"why":"Represents the end-to-end AI-research workflow that OmniScientist extends; the contrast point for workflow-complete but evidence-incomplete systems.","marker":"Lu et al., 2026"},{"why":"Another agentic AI-scientist system that reasons over text, code, and summaries, motivating the perceived evidence gap.","marker":"Yamada et al., 2025"},{"why":"A scientific multimodal benchmark that fixes observation and question in advance, the paradigm OmniScientist moves beyond.","marker":"Wang et al., 2024"},{"why":"Claude Sonnet 5 is the perception model pinned across all runs, so perception gains are attributable to the pipeline rather than the vision model.","marker":"Anthropic, 2026"},{"why":"Provides the STEAD seismology dataset used in the case study where perception finds that 21.7% of noise-labelled traces carry coherent transients.","marker":"Mousavi et al., 2019"},{"why":"Provides the chest X-ray dataset used in the case study where perception converts observed patchiness into a held-out diagnostic metric.","marker":"Kermany et al., 2018"}],"fun_headline_variants":["AI that reads raw data beats one fed summaries","Raw perception powers AI scientist to 85% win rate","Omni-modal AI scientist outperforms feature-only twin","Seeing raw data gives AI scientist an 85% edge","Direct perception defeats summary-based AI science"],"cache_read_input_tokens":36096,"weakest_assumption_plain":"The head-to-head comparison assumes that two LLM judges, using a rubric that explicitly rewards papers that show and interpret raw observations, produce an unbiased measure of research quality; no human expert panel validates the judge, so the rubric may reward figure presence rather than underlying science.","fun_headline_variants_meta":{"raw":{"variants":["AI that reads raw data beats one fed summaries","Raw perception powers AI scientist to 85% win rate","Omni-modal AI scientist outperforms feature-only twin","Seeing raw data gives AI scientist an 85% edge","Direct perception defeats summary-based AI science"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000601,"raw_usage":{"total_tokens":2858,"prompt_tokens":1049,"completion_tokens":1809,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":665,"completion_tokens_details":{"reasoning_tokens":1735}},"tokens_in":665,"tokens_out":1809,"duration_ms":12020,"temperature":1.0,"reasoning_tokens":1735,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:10:45.715675+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the five paired cases and let a panel of domain-expert human reviewers score the manuscripts on the same seven dimensions using a rubric that removes the instruction to reward showing raw observations. If the full system's win rate falls substantially below 85%, or its lead over the blind baseline narrows to ties, the reported perception advantage is substantially an artifact of the judge instruction rather than of improved science.","supporting_citations":[],"review_version":1}