{"id":"b81ae283-8f31-48e1-acc4-ec3e47863af6","arxiv_id":"2607.05682","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured Research Question Certificate improves the judged quality and auditability of LLM-proposed scientific questions, per a small LLM-judge evaluation.","lead":"This paper introduces FirstResearch, a framework that makes LLM-generated scientific research questions more auditable by packaging them into a structured Research Question Certificate. The authors report preliminary LLM-judge evidence that the structured format outperforms prompt-level baselines, although no human expert evaluation was done.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LLM-judge scores are treated as a measure of auditability, but the certificate-ablation drop suggests scores track certificate presence rather than question quality; no human auditability evidence exists.","rationale":"Reader's weakest assumption is exactly that LLM judges' scores are valid independent of certificate formatting cues. I agree; this is the load-bearing point. The abstract's ablation results make the concern concrete rather than speculative: the 4.90 to <1/5 drop on removing certificates is more plausibly explained by judge sensitivity to the certificate template than by a sudden collapse in underlying question quality. Because the paper self-declares 'preliminary' and 'LLM judges rather than human domain experts,' the lack of human auditability evaluation is an acknowledged limitation that should be treated as central. The proposed concrete test would settle whether the score gap is content-driven or artifact-driven. I therefore recommend no change to the reader's UNVERDICTED verdict; the concern supports the already-unverified status rather than overturning it.","tokens_in":762,"tokens_out":4187,"duration_ms":45642,"concrete_test":"Re-run the same 40 baseline packages with a format-control condition: take the FirstResearch certificates and replace the semantic content under each heading with generic filler (e.g., 'an assumption', 'a test') while preserving the template. Ask the same DeepSeek and Gemini judges to score them. If format-control scores approach 4.9/5, the original scores reflect template presence, not content. Independently, have two or more human domain experts rate all outputs on auditability (ability to identify the falsifier, assumptions, and decisive test) without knowing condition, and compute correlation with the LLM-judge scores. If the correlation is low, or if the format-control scores are high, the central claim's evidential basis fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that explicit derivation constraints are a promising mechanism for making LLM-generated questions more auditable. The evidence is LLM-judge scores (DeepSeek and Gemini) and a certificate-removal ablation. For the claim to hold, those scores must be a valid measure of auditability, not just of artifact presence. The ablation result — certificate-only scoring 4.90/5 while removal drops below 1/5 — is exactly what you would expect if judges reward the format itself: the template includes headings like 'assumptions' and 'falsifiable hypothesis' that may read as rigorous regardless of content. This is reinforced by the fact that the comparison is between a system that always emits certificates and prompt-level baselines that do not, so the judge is not blind to the intervention; it sees a structural difference. The abstract itself concedes there is no human expert evaluation. Consequently, the observed score gap cannot distinguish 'better questions' from 'more elaborate outputs,' and the auditability claim is not yet evidenced. This is a correctness risk in the evaluation, not merely a missing nice-to-have, because the outcome measure is unvalidated for the construct it is supposed to support.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces FirstResearch, a framework for LLM scientific-discovery agents that generates a structured Research Question Certificate containing primitive definitions, assumptions, a mechanism model, a tension/contradiction, a falsifiable hypothesis, a minimal decisive test, and a failure update rule. The authors claim that this explicit derivation structure improves the auditability and judged quality of LLM-proposed research questions. They report evaluations on ten LLM-agent research topics, comparing FirstResearch against prompt-level baselines inspired by AI co-scientist, Agent Laboratory, and AI Scientist-v2, using DeepSeek and Gemini judge models. They report system-level ranking preservation across judges, a large performance gap when the certificate is present versus absent in a one-repeat ablation, and they explicitly frame the results as preliminary due to the use of LLM judges rather than human experts. Code, prompts, outputs, and reproduction scripts are provided.","tokens_in":1107,"tokens_out":3147,"duration_ms":36238,"significance":"If the central claim is sound, the paper offers a lightweight, auditable mechanism for improving the inspectability of LLM-generated research questions, which is a real bottleneck in scientific-discovery agents. The authors deserve credit for releasing code, prompts, saved outputs, and reproduction scripts, and for including an independent judge model as a robustness check. However, the evidence base is thin: ten topics, a single-repeat ablation, and LLM judges with no human validation. The central construct — auditability — is a property for human scientists, and the validation relies on LLM judges whose scores may track certificate presence rather than question quality. The core evaluation-validity issue is load-bearing, and the paper's own limitation statement acknowledges the lack of human expert evaluation.","major_comments":[{"comment":"The central claim is about auditability, but the only outcome measure is LLM-judge scores. The abstract's own final sentence concedes that the evaluation uses LLM judges rather than human domain experts. Auditability is a property for human scientists; without any human audit or a validated proxy, the reported numbers (e.g., 4.90/5 vs <1/5) cannot be interpreted as evidence about auditability. Please add a human audit study, or at minimum validate the LLM-judge rubric against human judgments on a held-out set before treating the scores as a measure of the construct.","section":"Abstract — evaluation protocol"},{"comment":"The comparison is between FirstResearch, which always emits a certificate, and prompt-level baselines that do not. The judge therefore sees a structural signature of the intervention. The certificate-removal ablation dropping below 1/5 is exactly the pattern expected if the judge rewards template headings such as 'assumptions' and 'falsifiable hypothesis' regardless of content. This is a confound between artifact presence and judged quality. Please blind the judge to certificate presence (e.g., strip or shuffle the certificate headings, or judge plain-text questions) and include a sham-certificate control using the same template with vacuous content.","section":"Abstract — ablation and judge blindness"},{"comment":"If the judge's scoring criteria are the seven certificate components, then removing the certificate removes the checklist the judge is told to score. The abstract does not state the rubric, but the near-collapse of scores under certificate removal is consistent with a mechanical dependence. Specify the exact judge prompt and rubric, and show that scores are not simply a count of which headings are present. Without this, the reported effect sizes cannot be attributed to improved question quality or auditability.","section":"Abstract — circularity of rubric"},{"comment":"The ablation is described as a 'one-repeat' checkpoint. With a single repetition there is no variance estimate, and the large gaps between conditions (e.g., 4.90/5 vs below 1/5) could be dominated by run-to-run noise. Report multiple seeds or repeats, confidence intervals, and topic-level breakdowns. This is load-bearing because the claim that the certificate is the strongest component rests entirely on this ablation.","section":"Abstract — one-repeat ablation"}],"minor_comments":[{"comment":"The phrase 'DeepSeek-blind-judge protocol' is undefined. What is the judge blind to? Clarify the procedure (e.g., blind to system identity, blind to the certificate, blind to topic?).","section":"Abstract — 'DeepSeek-blind-judge'"},{"comment":"The sentence 'Pearson agreement of 0.865 on average score' is ambiguous. Specify the two variables correlated (e.g., DeepSeek vs Gemini average scores over the 40 packages?) and the units.","section":"Abstract — Pearson agreement"},{"comment":"The 'certificate-only' condition is unclear. Does it mean FirstResearch with all other components removed, or a standalone certificate without the framework's question formation process? Define the condition explicitly.","section":"Abstract — 'certificate-only'"},{"comment":"The GitHub link is appreciated. Consider also archiving a versioned snapshot (e.g., Zenodo DOI) to ensure long-term reproducibility and to allow citation of a specific version.","section":"Abstract — reproducibility"}],"recommendation":"major_revision","confidential_remarks":"The reader's report and my own reading agree on the central risk: the evaluation does not yet establish the construct validity of the LLM-judge scores for 'auditability.' This is not a mere presentation issue; it is the evidentiary basis for the paper's main claim. The issue is fixable by adding human audits, judge-blindness controls, and confidence intervals, so I recommend major revision rather than rejection. I would also encourage the editor to ask for the full rubric text in the revision, because the circularity concern depends on exactly what the judges are told to score."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"FirstResearch is a plausible process contribution — a structured Research Question Certificate that forces explicit definitions, assumptions, mechanism, falsifier, decisive test, and failure-update rule — and the authors did the right thing by shipping code, prompts, saved outputs, and reproduction scripts. The idea is worth taking seriously.\n\nThe narrow claim is stated carefully: explicit derivation constraints are a promising mechanism for making LLM-generated questions more auditable. The evaluation uses two independent LLM judges with reported agreement, and the prompt-level baselines are named. The abstract honestly calls the results preliminary and flags the lack of human expert evaluation.\n\nThe soft spot is the one that matters: the outcome measure. The paper claims auditability, but the evidence is LLM-judge quality scores. Those scores are not validated against any human auditability judgment, and the ablation pattern — certificate-only at about 4.9/5, removal below 1/5 — is exactly what you would expect from judges rewarding the presence of headings like 'assumptions' or 'falsifiable hypothesis' rather than the content. The baselines never emit a certificate, so the judge isn't blind to the intervention; it sees a structural difference. That confound means the observed gaps don't yet support 'more auditable' or even 'better questions' — they support 'outputs that look more like a certificate.' The abstract's own hedging doesn't remove that problem.\n\nThat said, this is not a dishonest or incoherent paper. The framework is clear, limitations are disclosed, and the ablation is a first step, not a pretense of a controlled experiment. The proper next step is a human-expert study with judges blind to condition, or a judge protocol that masks the certificate format. If the gap survives that, there's a real result.\n\nWho should read it: anyone working on LLM agents for science or evaluation of structured outputs. It makes the circularity problem concrete.\n\nRecommendation: send it to peer review. A serious referee can push on the outcome-validity issue, and the paper is reproducible enough that the concerns are testable. I'd accept it as a workshop paper now and as a journal paper after the evaluation is strengthened.","headline":"A sensible, honestly hedged framework for structuring LLM research questions, but the certificate-ablation evidence doesn't yet separate question quality from format cues.","tokens_in":1454,"tokens_out":2483,"would_cite":false,"duration_ms":24401,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that structuring an LLM's research-question formation through a fixed certificate of definitions, assumptions, mechanism, falsifiable hypothesis, decisive test, and failure rule makes the resulting questions substantially","keywords":["scientific discovery agents","large language models","research question formation","auditability","falsifiability","structured certificates","LLM evaluation","derivation constraints"],"falsifier":"Have a panel of human domain experts rate the same 40 question packages with and without certificate formatting, stripped to plain prose, under a validated rubric for research-question quality; if human rankings do not separate certificate-derived from prompt-derived questions when formatting is controlled, the central claim fails. A more direct test would regress judge scores on the presence of each certificate section to see whether mere section labels predict score independent of content.","tokens_in":708,"feed_emoji":"🔬","tokens_out":4802,"duration_ms":50602,"temperature":0.7,"pith_summary":"The paper tries to establish that the first research question an LLM scientific-discovery agent proposes can be made inspectable by requiring the model to produce a structured Research Question Certificate alongside the question. The certificate records the definitions, assumptions, mechanism, tension, falsifiable hypothesis, decisive test, and failure-update rule that the question depends on. If correct, this would give scientists a concrete artifact to audit before spending resources on experiments, and it would mean that the act of forcing a derivation chain improves the judged quality of the question. In the paper's evaluation, certificate-bound questions score roughly 4.9/5 under two independent LLM judges, while removing the certificate drops below 1/5.","feed_headline":"Adding a structured certificate lifts AI question quality to 4.86/5","feed_subtitle":"Forcing LLMs to record assumptions, mechanism, and falsifier beats plain prompts under two independent judges.","key_machinery":"The Research Question Certificate is a fixed seven-section template the LLM must fill in to emit a research question. It carries the argument by turning an opaque generation step into a visible derivation chain: each section is meant to be consistent with the previous ones, and any gap becomes inspectable by a reader. The ablation showing that scores collapse without the certificate indicates that this structured template, not the surrounding prompting, is the load-bearing component.","core_discovery":"The central claim is that an LLM asked to form a research question produces far more auditable and higher-judged questions when the generation is constrained to pass through a fixed Research Question Certificate — seven enumerated sections covering primitive definitions, assumptions, a mechanism model, a tension, a falsifiable hypothesis, a minimal decisive test, and a failure-update rule — than when the same question is generated from ordinary prompts. In the reported evaluation, certificate-bound generation scores 4.86/5 under a primary judge and 4.88/5 under a second judge in a one-repeat ablation, while removing the certificate collapses scores below 1/5 under both. The paper takes this","pith_inferences":["The score collapse when the certificate is removed (below 1/5) could mean the judges are rewarding the visible structure rather than the underlying question; a human expert panel rating the same questions stripped of certificate formatting would tell whether the effect is real content quality or format detection.","The certificate's seven slots may constitute a minimal argument skeleton, so the same template could be repurposed for auditing other scientific outputs — experiment proposals, literature summaries, or generated theorems — wherever assumptions and falsifiers matter.","The ordering of sections (definitions before assumptions before mechanism) may itself be a reasoning chain; ablating individual sections or permuting them would reveal whether each slot contributes or only the full structure does.","The 'minimal decisive test' slot is where a question stands or falls: a certificate can pass all other fields yet fail if the decisive test cannot distinguish the proposed mechanism from a plausible alternative, making that slot the highest-value target for a fast screen."],"forward_implications":["Auditing shifts earlier: a scientist can check assumptions, mechanism, and decisive test before running any experiment, rather than after seeing a plausible-sounding question.","The failure-update rule gives each rejected question an explicit revision path, turning question formation into an iterative loop rather than a one-shot generation.","Because two independent LLM judges preserve the ranking, the quality advantage is not tied to one judge's preferences.","The ablation result indicates that the certificate, not the prompt wording, is what produces the improvement.","The framework offers a concrete checklist for what 'auditable' means in AI-generated research questions."],"fun_headline_variants":["Auditable LLM questions: certificate beats plain prompts in scores","Certificate-based question formation lifts LLM quality to 4.86/5","Forcing LLMs to expose assumptions improves research question ratings","Structured certificate makes LLM questions auditable and higher rated","LLM question audibility jumps with certificate: scores 4.86/5"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The evaluation rests on the assumption that the two LLM judges' scores reflect the genuine scientific merit and audibility of the research questions, not just the visible presence of a structured certificate; the score collapse without the certificate makes that assumption the point of vulnerability.","fun_headline_variants_meta":{"raw":{"variants":["Auditable LLM questions: certificate beats plain prompts in scores","Certificate-based question formation lifts LLM quality to 4.86/5","Forcing LLMs to expose assumptions improves research question ratings","Structured certificate makes LLM questions auditable and higher rated","LLM question audibility jumps with certificate: scores 4.86/5"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000302,"raw_usage":{"total_tokens":1628,"prompt_tokens":847,"completion_tokens":781,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":591,"completion_tokens_details":{"reasoning_tokens":689}},"tokens_in":591,"tokens_out":781,"duration_ms":8176,"temperature":1.0,"reasoning_tokens":689,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T08:23:38.547086+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a panel of human domain experts rate the same 40 question packages with and without certificate formatting, stripped to plain prose, under a validated rubric for research-question quality; if human rankings do not separate certificate-derived from prompt-derived questions when formatting is controlled, the central claim fails. A more direct test would regress judge scores on the presence of each certificate section to see whether mere section labels predict score independent of content.","supporting_citations":[],"review_version":2}