{"id":"d456f67e-d0e3-4b60-866c-62960cfe657c","arxiv_id":"2508.17202","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":4,"one_line_summary":"A budget-aware framework (PU-ADKA) selects which domain expert an LLM should query under a fixed $100 budget, improving specialized-domain answers at low cost.","lead":"This paper presents PU-ADKA, a framework that lets a large language model spend a fixed $100 budget querying the most useful domain expert, choosing whom to ask based on cost, availability, and expertise. It claims the approach improves LLM answers in sensitive fields like drug discovery and rare diseases, and ships a new benchmark, CKAD, for such budget-limited knowledge acquisition.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Simulation-to-real transfer and lack of a baseline control leave the $100-budget improvement claim unverified.","rationale":"The paper's central assertion is that PU-ADKA efficiently enhances domain LLMs within a $100 budget by actively selecting the most appropriate expert. This requires that the selection policy, trained in simulation on PubMed data, transfers to real experts and that the observed benefit is attributable to the policy's choices rather than to expert consultation per se. The reader correctly identified the simulation-to-real transfer as load-bearing. My stress-test adds that even a perfect transfer would not establish the claim without a baseline: the abstract's phrase 'controlled expert interactions' could mean controlled by the method, not a controlled comparison. Because the submitted full text is character-level corrupted, none of the experimental details—expert availability distributions, cost model, deployment size, evaluation metrics, or statistical tests—can be checked. This is not an accusation of error; it is a statement that the evidence is currently insufficient to distinguish PU-ADKA's selectivity from simpler strategies. A single A/B comparison with a cheapest-available arm would settle the question. Given the inability to inspect the body and the absence of any independent verification, the reader's UNVERDICTED verdict stands.","tokens_in":10993,"tokens_out":3639,"duration_ms":42031,"concrete_test":"Run a three-arm randomized evaluation with the same drug-development team and question set under a fixed $100 budget: (A) PU-ADKA's trained policy, (B) always query the cheapest available expert, (C) query a randomly chosen available expert. Pre-register the primary outcome as the change in accuracy on a held-out set of domain questions (ideally from CKAD, with external annotation). If A does not significantly beat B (and preferably C), the claim that learned selectivity drives cost-efficient enhancement fails. Alternatively, if logs from the existing deployment are available, re-analyze them to compute the counterfactual budget/outcome of B on the same questions; this would settle whether the selection policy added value.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To establish that PU-ADKA's expert selection, not just expert consultation in general, drives the reported gains, two conditions must hold: (1) the selection policy trained on simulated PubMed experts must generalize to real drug-development experts, and (2) the real-world deployment must include a control arm that receives the same budget but not the learned policy. The abstract claims validation through 'controlled expert interactions and real-world deployment,' but from the available text—which is mojibake and contains a running header from another arXiv paper—no experimental protocol, baseline comparisons, or statistical results are inspectable. The paper's own introduction of the CKAD benchmark cannot independently validate the method, since it is part of the same unreviewed artifact. If the deployment only compared pre/post LLM accuracy with PU-ADKA's consultation, any improvement could be due to the additional expert knowledge itself, not to the learned selection policy, and the cost-efficiency claim would be unsupported. The most load-bearing concern is therefore the absence of evidence that the selective querying policy outperforms a simple (e.g., cheapest-available or random) expert-selection heuristic under the same budget; this is a transfer-and-evaluation gap, not merely a theoretical mismatch.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes PU-ADKA, a framework for active domain knowledge acquisition under a fixed budget. Instead of fine-tuning, it selects which domain expert to consult based on availability, knowledge boundaries, and consultation cost. The authors state that the selection policy is trained on simulations built from PubMed data, validated in controlled expert interactions, and deployed with a real drug-development team; they also introduce a CKAD benchmark. The abstract, however, provides no numeric results, baselines, or protocol details, and the supplied full text is almost entirely mojibake, with a running header from another arXiv paper. Consequently, the methods and experimental evidence cannot be inspected.","tokens_in":11219,"tokens_out":5014,"duration_ms":55627,"significance":"If the claimed results held, the framework would address a practical problem: improving specialized LLMs without expensive fine-tuning while respecting expert availability and cost. The introduction of a benchmark for cost-effective domain knowledge acquisition could also be useful. On the submitted evidence, however, these contributions exist only as statements in the abstract. No machine-checked derivations, reproducible code, or inspectable experimental protocols are available. The significance cannot be assessed until a readable manuscript with baselines and transfer analysis is provided.","major_comments":[{"comment":"The manuscript body is unreadable: the text is mojibake, tables are gibberish, and a header 'arXiv:2508.17199v1 [cs.CV]' appears mid-document. No equation, algorithm, or result can be verified. This presentation defect blocks review; the paper effectively consists only of its abstract.","section":"Full text"},{"comment":"The abstract states that PU-ADKA is validated through controlled expert interactions and real-world deployment, but it reports no performance metric, baseline, error bar, or statistical comparison. Without these, the claim that the framework 'demonstrat[es] its effectiveness' is unsupported in the submitted artifact. This concern is load-bearing because the entire contribution is empirical.","section":"Abstract"},{"comment":"The policy is 'trained ... using simulations on PubMed data' and then deployed with a drug-development team. The abstract provides no evidence that the simulated experts have realistic availability, cost, and error patterns, nor any analysis of simulator-to-real transfer. If the simulation is faithful only to PubMed text statistics, the learned query policy may not transfer to human experts, and the claimed $100-budget improvement would not follow.","section":"Abstract (simulation-to-real)"},{"comment":"The real-world deployment, as described, appears to lack a control arm. To attribute gains to the PU-ADKA selection policy rather than to the act of consulting experts at all, the deployment must compare against a heuristic selector (e.g., random, cheapest-available, or fixed order) under the same budget. No such comparison is visible in the abstract or readable text.","section":"Abstract (deployment)"},{"comment":"The abstract does not specify whether the evaluation questions in the controlled interactions and deployment overlap with the PubMed corpus used to train the simulated experts. If the same literature underlies both the simulated training signal and the evaluation, part of the reported gain could be due to rediscovering the training signal rather than to effective expert selection. This leakage risk must be addressed in the experimental design.","section":"Abstract (evaluation data)"}],"minor_comments":[{"comment":"The running header from arXiv:2508.17199 must be removed; the source should be rebuilt with a working font/encoding so the PDF is legible.","section":"Full text"},{"comment":"The '$100-dollar budget' is never defined in the abstract or readable portion. Please specify the cost model: what counts as a consultation, how expert time is priced, and how the ceiling is applied.","section":"Abstract"},{"comment":"The CKAD benchmark is introduced by name, but no task description, data source, size, or evaluation protocol is given in the readable portion.","section":"Abstract (CKAD)"},{"comment":"The reference list is not accessible due to the encoding corruption, so related-work attribution cannot be checked.","section":"References"}],"recommendation":"reject","confidential_remarks":"The submission appears to be a corruptly rendered PDF, with large portions of the body text replaced by mojibake and a header from another arXiv paper. Editorially, this would warrant a desk rejection unless a clean version is uploaded. Even setting the corruption aside, the abstract alone does not provide the baselines, transfer analysis, or control condition needed to support the central claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick note on 2508.17202. The honest headline: the supplied full text is mojibake, so no one can review the methods or results as-is. I'm writing from the abstract only.\n\nThe problem is worth working on: LLMs lack deep domain knowledge in sensitive fields like drug discovery, and paying experts for consultation under a strict budget is a real constraint. The PU-ADKA framing—selecting an expert by availability, knowledge boundaries, and cost rather than fine-tuning—is a sensible direction, and using PubMed-derived simulations to train the selection policy is a reasonable first step. Introducing CKAD as a shared benchmark is useful if it ships with clean data.\n\nBut the soft spots are large. The abstract promises validation via 'controlled expert interactions and real-world deployment,' yet gives no numbers, baselines, error bars, or protocol. The stress-test concern is on target: even with numbers, you would need a control that spends the same budget without the learned policy to show that selective querying, not just consultation, drives any gain. And the simulation-to-real transfer is a real burden—trained on simulated experts from PubMed, then deployed with a drug development team. That could fail in ways the abstract doesn't address.\n\nThe corrupted text is also a problem: the extracted body is unreadable and even includes a header from another arXiv paper. I'm sure that's a submission artifact, not a statement about the work, but it makes evaluation impossible.\n\nWho is this for: people working on cost-efficient LLM adaptation, especially in regulated or niche domains. If a clean version lands, it deserves careful peer review, with close attention to the baseline comparison and the fidelity of simulated experts. As it stands, I'd ask the authors to fix the PDF and resubmit rather than send this version to a referee.","headline":"Unreadable text makes this impossible to review; the abstract frames a real problem, but the core claims are unverifiable.","tokens_in":11730,"tokens_out":2162,"would_cite":false,"duration_ms":23993,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PU-ADKA claims that a fixed $100 consultation budget, spent by selectively asking the right human expert at the right time, improves LLM performance in specialized domains without fine-tuning.","keywords":["LLM domain specialization","active knowledge acquisition","expert-in-the-loop","budget-constrained consultation","drug discovery","rare disease research","CKAD benchmark","simulation-based training"],"falsifier":"Re-run the drug-development deployment with a fresh panel of experts and compare PU-ADKA's final answers, under the same $100 budget, against two controls: random expert selection and always consulting the cheapest available expert. If PU-ADKA does not beat both on expert-scored accuracy, the claimed budget-efficient improvement is falsified.","tokens_in":10873,"feed_emoji":"🧪","tokens_out":7375,"duration_ms":83838,"temperature":0.7,"pith_summary":"This paper tries to establish that a general large language model (LLM) can be turned into a specialist in cost-sensitive fields like drug discovery and rare-disease research by spending a small, fixed budget consulting human experts—no fine-tuning required. The authors propose PU-ADKA, a framework that learns which questions to ask, which expert to ask, and what that consultation is worth, balancing availability, expertise, and price. They train the framework in simulations built from PubMed literature and then test it with controlled human-expert interactions and a real drug-development team. If the claim holds, scarce expert time can be spent where it most improves the model, and specialized LLM knowledge can be acquired under strict budget limits.","feed_headline":"Spending $100 on the right experts boosts LLM domain accuracy","feed_subtitle":"By choosing which expert answers each question, the model buys more specialized knowledge per dollar.","key_machinery":"PU-ADKA, a cost-aware expert-selection policy. It converts \"ask an expert for help\" into a sequential budget-allocation problem: for each question it weighs an expert's availability, the boundaries of what that expert knows, and the price of consulting them. Simulation on PubMed data supplies the practice environment in which the policy learns, so that expensive, scarce real experts are only called upon when the learned policy expects the highest marginal knowledge gain.","core_discovery":"The central claim is that domain expertise can be treated as a purchasable, budgeted resource. PU-ADKA considers a team of experts with different availability, knowledge boundaries, and consultation costs; for each question it decides whether the base LLM can answer alone or whether it should pay a specific expert, and it does this through a policy learned in simulation from PubMed data. The paper reports that the same policy, when run with real human experts and in a live drug-development deployment, raises LLM performance on specialized tasks under a strict $100-style budget. Alongside the method, the paper releases CKAD, a benchmark dataset for measuring cost-effective domain knowledge ac","pith_inferences":["The same budget-allocation machinery could extend to any mix of expensive knowledge sources—paid APIs, specialized models, or human experts—each with its own price and reliability.","A natural safety-oriented extension would reserve part of the budget for verification, hiring a second expert to check high-stakes answers instead of spending every dollar acquiring new facts.","Because training depends on simulated experts, a direct transfer test—deploying the PubMed-trained policy in a non-biomedical domain with a different expert vocabulary—would measure how much of the gain comes from simulation fidelity versus the selection mechanism.","The budget framing invites a head-to-head comparison with retrieval-based augmentation: if the same $100 buys high-quality database or API access rather than human hours, which route yields more domain accuracy?"],"forward_implications":["Specializing an LLM no longer requires assembling a large fine-tuning corpus; a targeted expert-consultation budget can substitute.","Expert time is allocated by expected knowledge gain rather than by fixed protocol, so scarce specialists are not spent on questions the model already answers well.","The policy learned in simulation can be applied in a real expert setting in the same domain, as demonstrated in the drug-development deployment.","CKAD gives future work a common, budget-aware benchmark for comparing domain knowledge acquisition methods."],"supporting_citations":[],"fun_headline_variants":["How to spend $100 on expert advice to boost LLM expertise","Budget-smart expert queries make LLMs sharper in niche fields","Spend your LLM budget on the right experts: $100 buys more domain accuracy","Choose the right expert per question: $100 LLM boost in sensitive domains","Stretch $100 further: smart expert selection improves LLM domain skills"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the simulated experts built from PubMed behave enough like real experts—same availability, cost, and error patterns—that the query-selection policy trained in simulation still makes good choices with real people.","fun_headline_variants_meta":{"raw":{"variants":["How to spend $100 on expert advice to boost LLM expertise","Budget-smart expert queries make LLMs sharper in niche fields","Spend your LLM budget on the right experts: $100 buys more domain accuracy","Choose the right expert per question: $100 LLM boost in sensitive domains","Stretch $100 further: smart expert selection improves LLM domain skills"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000257,"raw_usage":{"total_tokens":1395,"prompt_tokens":700,"completion_tokens":695,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":444,"completion_tokens_details":{"reasoning_tokens":598}},"tokens_in":444,"tokens_out":695,"duration_ms":7435,"temperature":1.0,"reasoning_tokens":598,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T16:58:33.713721+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the drug-development deployment with a fresh panel of experts and compare PU-ADKA's final answers, under the same $100 budget, against two controls: random expert selection and always consulting the cheapest available expert. If PU-ADKA does not beat both on expert-scored accuracy, the claimed budget-efficient improvement is falsified.","supporting_citations":[],"review_version":1}