{"id":"36aa1d84-3380-4a97-85a5-7070e7db4c96","arxiv_id":"2608.11341","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Apodex Discovery introduces the TRACES benchmark of 17 executable, verifiable environments with hidden outcomes and the HDS6 process-evaluation metric, reporting early gains for environment-equipped agents on AAV capsid design and drug repurposing.","lead":"This paper introduces Apodex Discovery, a framework and benchmark that turns open-ended real-world problems into executable, verifiable environments for AI agents, and adds a six-axis process-evaluation metric called HDS6. It reports early results on AAV capsid design and drug repurposing, with modest claimed gains over published baselines, and shows that process scores correlate with task success.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The AAV generative-design oracle is a fitted classifier on real fitness data; with no validation that it tracks true selectivity in the OOD regime, the 7% SOTA claim does not by itself establish genuine discovery.","rationale":"The reader's weakest assumption identifies the outcome-verifier proxies as the load-bearing risk, and my reading converges on the same point. The framework's internal design is coherent: the episode interface is fixed, ablations are controlled where run, HDS6 is scored blind with cited trajectory evidence, and the repair case studies are concrete and instructive. Those elements support the framework as a benchmark infrastructure and justify a conditional acceptance. The specific number that makes the paper newsworthy is the AAV SOTA claim, and that claim rests on the fit of the generative-design oracle to real biological selectivity. The paper is honest about using proxies, but it does not demonstrate that the proxy is faithful in the exact regime where solvers are graded. This is not a circularity or an appeal to external consensus; it is a missing calibration check on the verifier itself. The proposed held-out validation is computational and directly tests whether a high oracle score corresponds to a genuinely selective candidate. Until that check is run, the headline result should be treated as conditional, exactly as the reader concluded. I therefore do not adjust the verdict, only sharpen the condition on which it rests.","tokens_in":49610,"tokens_out":3798,"duration_ms":39010,"concrete_test":"Perform a held-out validation of the Task 4 oracles against the real measurements they are meant to stand in for. Split the Fit4Function enrichment data and LY6A/LY6C1 binding data into training and held-out sets, train the same surrogate models used in Section 4.1.5, and measure rank correlation between oracle predictions and true measured values on held-out peptides at the same Hamming distance from training as the generated designs (at least two substitutions). If the held-out Spearman rho is below about 0.5, or if top-1%/5% selection precision is near chance, the design-task scores in Table 19 cannot support the claimed discovery improvement over the reproduced baselines, and the 7% SOTA margin would need to be re-evaluated against direct experimental outcomes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that Apodex Discovery makes discovery verifiable: a solver's investigation is graded against hidden verifiers that are faithful stand-ins for real-world objectives. Section 3.1.1 explicitly permits fallback to a 'measurable proxy', and the load-bearing instance is the AAV generative-design task (Section 4.1.5, Table 19). A generated sequence is a hit only if 'a trained classifier judges it specific' (receptor instances) or its predicted organ selectivity 'ranks in the top 1%, 2%, or 5%' of all sequences (tissue instances). Those predictions come from surrogate models trained on seed data, not from measurements of the generated candidates. The headline 'surpassed the published state of the art by 7%' is entirely mediated by these surrogates: apodex-1.1's 0.180 versus the reproduced AAVGen/AAVDiff/ALICE scores of 0.109-0.116 is a comparison of oracle outputs, not of experimentally characterized capsids. If the oracle is miscalibrated for novel, multi-substitution sequences—precisely the out-of-distribution regime the benchmark intentionally scores—a solver can score high by exploiting surrogate blind spots without producing a genuinely on-target capsid. The paper reports no calibration of the oracle against held-out experimental measurements, no analysis of oracle error as a function of novelty distance, and no wet-lab or independent validation of the designed sequences. The result is not an externally controversial claim; it is an unvalidated link between the benchmark's success signal and the real-world discovery the abstract invokes. The same concern also weakens the 'published SOTA' baselines, since the Task 4 references are the authors' own reproductions under the identical episode budget, but the more fundamental issue is the oracle itself.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Apodex Discovery, a framework for building and evaluating 'discoverative AI' through heavy-duty solvers, together with TRACES, a benchmark of executable environments, tasks, and episodes with hidden verifiers. It describes a problem-scouting process that produced a 423-problem registry, an environment-task-episode abstraction with isolation and anti-gaming gates, a blind process metric (HDS6) scoring Tools, Repair, Alternatives, Coherence, Evidence, and Scope, and verification-driven repair loops. The empirical sections report that the Apodex solver surpassed published state of the art on four AAV capsid design tasks, that a task-specific biomedical environment improved normalized prediction scores of GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points on drug repurposing, and that controlled ablations attribute performance differences to solver components. The authors release TRACES with 17 environments and 218 episodes and position the work as moving AI evaluation beyond predefined benchmarks toward verifiable investigations.","tokens_in":49958,"tokens_out":9317,"duration_ms":82244,"significance":"If the framework and results hold, this is a substantial step toward evaluating AI on open-ended, consequential problems. The release of executable environments with trajectory recording, hidden verifiers, anti-gaming gates, and a blind process metric is valuable infrastructure, and the internal validation of HDS6 against outcome scores (pooled Spearman 0.51 over 409 trajectories) is a meaningful first check. The controlled ablations and the verification-repair case studies are also useful contributions. However, the headline empirical claims currently depend on author-reproduced baselines, surrogate oracles, and a high-anchor normalization, so the significance of the specific quantitative results is conditional on additional validation.","major_comments":[{"comment":"The claim that Apodex 'surpassed the published state of the art by 7%' is not defined precisely, and the generative-design comparison is against the authors' own reproductions, not published numbers. Table 19 states 'All three reference methods are our own reproductions under the identical episode budget,' yet AAVDiff is labeled 'published SOTA.' The 7% figure appears to depend on including Task 4, where the baseline is an in-house reimplementation scored by the same surrogate oracle. Please state the exact formula for the 7% aggregate, report the original published scores alongside the reproductions, and provide code and hyperparameters for all baseline reproductions so the comparison is verifiable.","section":"4.1.3 and Table 19"},{"comment":"The design-task oracle is a trained classifier (receptor instances) or a predicted-rank quantile (tissue instances), and the paper provides no validation that these surrogates track true experimental fitness in the out-of-distribution regime of multi-substitution novel sequences. Since the benchmark intentionally scores sequences far from the seed distribution, a solver can score well by exploiting surrogate blind spots without producing genuinely on-target capsids. Please add an oracle-calibration analysis against held-out experimental measurements, report oracle error as a function of mutation distance from the training set, or explicitly reframe the design result as an in-silico benchmark rather than evidence of genuine discovery.","section":"4.1.5, Task 4 (Generative design)"},{"comment":"The random anchor S_R=0.86925 is very close to the oracle anchor S_O=1, so the normalized score in Eq. (4) compresses raw score differences by a factor of about 7.65 (1/0.13075). The reported gains of +2.52 and +7.60 normalized points correspond to raw gains of +0.00330 and +0.00993, and with three runs the standard deviations overlap substantially (GPT-5.5: 53.78±3.22 vs 56.30±1.56; GPT-5.6-sol: 54.21±2.20 vs 61.81±3.92). The abstract and conclusions present these as clean gains; please report raw scores with confidence intervals, add significance tests or explicit non-significance statements, and describe the results as preliminary.","section":"4.2.2, Eq. (4), and Table 7"},{"comment":"HDS6 is introduced as a measure of process quality, but its validation currently consists of correlation with the authors' own outcome verifiers (pooled Spearman 0.51). No inter-annotator agreement is reported for the judge/reviewer/arbiter procedure, and no comparison with human expert process ratings is provided. Since HDS6 is a central contribution, please report reliability statistics (e.g., judge-reviewer agreement, human-HDS6 agreement) or temper the claim that HDS6 'evaluates' process quality rather than reflecting a proposed operationalization.","section":"3.1.2 and Table 3"},{"comment":"The verification-driven repair section reports a mean outcome improvement of +0.155 over 434 deficient trajectories, but the design has no matched control group: the same deficient trajectories are not re-run without the repair note. The text acknowledges this limitation, yet the section title and the framing of the aggregate result present the improvement as an effect of repair. Please add a matched control (identical trajectories re-run without the repair note) or present the table strictly as a pilot demonstration, with regression-to-mean and run-to-run variation discussed as alternative explanations.","section":"4.4 and Table 13"}],"minor_comments":[{"comment":"The abstract omits the caveats stated in Section 4.2.3 ('these gains should be interpreted descriptively') and in Table 19 (baselines are author reproductions); please add brief qualifiers so the claims match the body's caution.","section":"Abstract"},{"comment":"The label 'published SOTA' on the AAVDiff row conflicts with the table note that all three reference methods are reproductions; use 'reproduced baseline' throughout or clearly separate published numbers from reimplementations.","section":"Table 19"},{"comment":"The conclusion states that Apodex 'exceeded prior human baselines,' but Table 5 shows that apodex-1.1 is not the strongest solver in the authors' own set (e.g., kimi-k3 scores higher on all four tasks); please specify that the comparison is to published methods, not human experts or the strongest frontier models.","section":"4.1.3, last paragraph"},{"comment":"There is a typo: 'disocoverative AI' should be 'discoverative AI'; the same section also uses 'Table 2' and 'Table 3' before their first mention, which is fine but should be cross-checked.","section":"3.1.2, Evidence definition"},{"comment":"The 'virtual clinician' is introduced without specification of its interface or whether it is a simulated tool or a human in the loop; please clarify its role in the environment and whether it is available to all solvers.","section":"4.4.1"}],"recommendation":"major_revision","confidential_remarks":"The infrastructure and benchmark release are genuinely valuable, and the internal validation of HDS6 is a reasonable first step. My main concern is that the abstract and conclusion overstate the empirical results: the 7% AAV claim rests on author-reproduced baselines and surrogate oracles, and the drug-repurposing gains are small relative to raw scores and not statistically robust. These are fixable with additional validation or carefully softened claims. I would also encourage the authors to make the baseline reproduction code and the oracle calibration analyses public, since the benchmark's credibility depends on verifiability of the comparison."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real contribution here is the evaluation machinery, not the headline results. The environment–task–episode abstraction with hidden verifiers, anti-gaming gates, and the blind HDS6 process metric is a genuine step beyond the static-QA versus final-success binary that dominates agent benchmarks. The load-bearing step analysis and structured repair feedback are thoughtful, and the 409-trajectory outcome–HDS6 correlation (pooled Spearman 0.51) is a meaningful internal check. The fixed episode interface for attribution is also done carefully; the ablations in Section 4.3 are a model of how to isolate harness, model, and mechanism effects.\n\nThat said, the paper overreaches in its discovery claims. The stress-test note is on target: the AAV design task scores candidates with a trained classifier or rank threshold, not with any experimental measurement. The 7% over published SOTA is entirely mediated by those surrogates, and the baselines are the authors' own reproductions. No calibration of the oracle in the out-of-distribution regime, no novelty-distance error analysis, and no wet-lab validation. It is possible the designed capsids are good, but the evidence presented does not establish it. The abstract's language about \"genuine discovery\" is not supported by the current experiments.\n\nThe drug repurposing results are honest in their limitations: three runs on a public-100 set, described as descriptive. The repair-loop results are also honestly labeled as lacking a matched control, and the procedure-design comparison self-identifies as measuring retrieval rather than generation. I appreciate that level of candor.\n\nMinor soft spots: no code or data in the preprint, and the HDS6 rubric is authored by the same team that evaluates with it. Neither is fatal; both are reasons to require release before taking the process scores at face value.\n\nWho is this for? Anyone building agentic evaluation infrastructure or studying process-verification metrics. The framework deserves a serious referee and the community's attention. I would not cite the AAV performance claims until they are independently reproduced, but I would cite the TRACES abstraction and HDS6 design. Recommend: accept with major revisions, conditional on releasing the benchmark, de-emphasizing the unvalidated SOTA claims, and either adding oracle calibration or explicitly reframing the design task as a surrogate-prediction challenge rather than discovery.","headline":"A serious evaluation-framework paper whose discovery claims outrun its verifiers; the TRACES/HDS6 machinery deserves referee time, but the AAV SOTA number needs external validation.","tokens_in":50620,"tokens_out":1320,"would_cite":true,"duration_ms":15438,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Move AI evaluation from answers to verifiable investigations","keywords":["discoverative AI","heavy-duty solver","reality benchmark","hidden verifiers","process verification","HDS6","AAV capsid design","drug repurposing"],"falsifier":"Run the four AAV capsid tasks scoring the original published methods' own released outputs or official implementations, rather than the paper's reproductions, against the same hidden verifiers; if apodex-1.1's per-task scores (0.904 viability, 0.635 tropism, 0.649 structure, 0.180 design) stop exceeding the published states of the art (0.878, 0.622, 0.605, 0.116), the headline 7% claim fails.","tokens_in":49430,"feed_emoji":"🧬","tokens_out":7272,"duration_ms":85067,"temperature":0.7,"pith_summary":"The paper argues that the next bottleneck for AI is not solving specified tasks but conducting open-ended, consequential investigations, and that this requires infrastructure: an explicit problem manifest, an executable reality-based environment, verification, and repair. It presents Apodex Discovery, built on the operational unit of a heavy-duty solver—a foundation model combined with a harness, tools, and control policies—operating inside a fixed environment–task–episode interface where ground truth is hidden and every trajectory is recorded. The framework adds HDS6, a blind process metric scoring Tools, Repair, Alternatives, Coherence, Evidence, and Scope independently of final success, so progress can be judged even when definitive outcomes are delayed or unknown. As early evidence, the paper reports that an Apodex solver surpassed published per-task state of the art in AAV capsid design by about 7% across viability, tropism, structure, and generative design, and that adding a task-specific biomedical environment improved drug-repurposing prediction by 2.5 and 7.6 normalized points over the same closed-book backbones. The central bet is that moving evaluation from right answers to reliable, evidence-grounded, self-correcting investigation is what makes discovery buildable and testable.","feed_headline":"AI evaluation moves from answers to investigations","feed_subtitle":"A benchmark hides ground truth and scores the process; its AAV solver beats published state of the art by 7%.","key_machinery":"The central mechanism is the environment–task–episode abstraction with its fixed episode interface and two verification channels. The episode interface is the single contract through which every solver and baseline passes: it fixes inputs, tools, budgets, and submission format on one side, and the hidden verifier, hard anti-gaming gates, and outcome metrics on the other, which is what allows observed differences to be attributed to a specific solver component. Overlaid on this is the HDS6 process verifier, which scores the recorded trajectory—blind to outcome and to solver identity—along six capabilities that spell TRACES (Tools, Repair, Alternatives, Coherence, Evidence, Scope), using per-task subrubrics plus a load-bearing-step analysis. Verification output doubles as a repair note returned to the solver, converting evaluation from a verdict into a loop. The paper treats the construction of faithful environments with trustworthy hidden verifiers as a contribution on its own, explicitly noting that building such an environment can be as hard as solving the task.","core_discovery":"The paper's central claim is that open-ended real-world problems can be converted into tractable, verifiable, repairable benchmark environments, and that evaluating the complete solver system rather than the model alone is what enables progress toward genuine discovery. The transformation has four load-bearing pieces: a problem manifest that fixes the question, success criterion, task decomposition, tools, budgets, and hidden verifier; a reality-based environment that returns fresh observations as the solver acts; outcome verification against hidden ground truth or measurable proxies; and blind process verification through HDS6, whose repair notes close a measure–diagnose–improve loop without leaking ground truth. The paper further claims this machinery works in practice: in adeno-associated virus capsid design, the framework exceeded the published state of the art at every one of the four pipeline stages, and in drug repurposing and reformulation, adding the task-specific environment improved mean normalized prediction scores of two GPT backbones by 2.5 and 7.6 points over the same closed-book backbone. It also claims that the blind HDS6 process score correlates positively with hidden outcome scores, with a pooled Spearman coefficient of 0.51 over 409 trajectories, supporting the use of process metrics when definitive ground truth is delayed.","pith_inferences":["The framework's registry of 423 problems suggests that almost any domain with a monitorable success criterion could be turned into an adversarial evaluation environment, making benchmark construction more of an industrial process than a hand-crafted exercise.","Because the HDS6-outcome correlation is positive but moderate, process scores are most defensible as early indicators for curation and repair, not as replacements for outcome verification when high-stakes decisions ride on the result.","A natural next experiment is to ablate only the biomedical evidence tools in drug repurposing while holding the episode interface fixed; the paper's attribution logic predicts the reported gains should be traceable to specific tool categories.","If hidden verifiers based on proxies become standard, the field will need periodic calibration checks that compare proxy scores against long-delayed true outcomes, otherwise solvers will optimize to the proxy instead of the real objective."],"forward_implications":["Benchmarks can be built around complete solver systems rather than isolated prompts, with final-answer scoring and process scoring reported side by side.","Problems whose ground truth is delayed, incomplete, or non-existent become evaluable: outcome verifiers handle measurable proxies, HDS6 supplies an immediate process score, and the repair loop turns diagnosis into a measurable improvement.","Across the four AAV capsid tasks, a domain-specific environment plus a capable solver surpasses the previous per-task published state of the art, from viability prediction to generative design.","The controlled episode interface makes component attribution possible, as shown by ablations where skill guidance and terminal verifier guidance improve outcome scores on LLM-engineering tasks.","A blind process score that correlates positively with hidden outcomes can act as a provisional quality signal while definitive outcomes are still pending."],"supporting_citations":[{"why":"Supplies the deep-mutagenesis viability screen that trains and evaluates the multi-mutation AAV viability instance, and the logistic-additive baseline.","marker":"[14]"},{"why":"Supplies the single-mutation fitness landscape used for the viability extrapolation instance across sequence position and mutation type.","marker":"[15]"},{"why":"Provides the Fit4Function tropism screens and seed data for the design task, and the published tropism state-of-the-art baseline.","marker":"[16]"},{"why":"Provides LY6A and LY6C1 receptor binding measurements used to score the receptor-instance design task.","marker":"[17]"},{"why":"The CAP-PLM capsid language model whose 0.878 viability score serves as the published state of the art that apodex-1.1 surpasses.","marker":"[18]"},{"why":"AlphaFold 3 with template-based symmetry expansion is the published state-of-the-art baseline for the structure prediction task.","marker":"[21]"},{"why":"AAVGen is one of the specialist generative methods reproduced as a design-task baseline.","marker":"[22]"},{"why":"AAVDiff is the generative method the paper marks as published state of the art on the design task.","marker":"[23]"}],"fun_headline_variants":["From fixed benchmarks to verifiable discovery missions","Beyond answers: scoring AI's investigation process","Open-ended AI: verifiable investigation benchmarks replace fixed tasks","Scoring the process, not just the outcome: new benchmark for discovery AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole framework assumes that each hidden verifier's measurable proxy—a held-out experimental fitness, a clinical-trial stage, a predicted selectivity rank—really tracks the true real-world objective, and that the reproduced baselines used as published state of the art faithfully represent the original methods; if either assumption gives way, a solver can score high without achieving a genuine discovery.","fun_headline_variants_meta":{"raw":{"variants":["From fixed benchmarks to verifiable discovery missions","Beyond answers: scoring AI's investigation process","Open-ended AI: verifiable investigation benchmarks replace fixed tasks","Scoring the process, not just the outcome: new benchmark for discovery AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000683,"raw_usage":{"total_tokens":3187,"prompt_tokens":1120,"completion_tokens":2067,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":736,"completion_tokens_details":{"reasoning_tokens":2002}},"tokens_in":736,"tokens_out":2067,"duration_ms":17724,"temperature":1.0,"reasoning_tokens":2002,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:12:32.820210+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the four AAV capsid tasks scoring the original published methods' own released outputs or official implementations, rather than the paper's reproductions, against the same hidden verifiers; if apodex-1.1's per-task scores (0.904 viability, 0.635 tropism, 0.649 structure, 0.180 design) stop exceeding the published states of the art (0.878, 0.622, 0.605, 0.116), the headline 7% claim fails.","supporting_citations":[],"review_version":1}