{"id":"4054ce64-3b27-4375-8e94-79f4848a8a56","arxiv_id":"2606.28433","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"RL researchers need to separate 'solving simulators' from 'using simulators as proxies' to prevent misleading conclusions about algorithm performance.","lead":"This paper argues that reinforcement learning researchers should distinguish between solving benchmark simulators and using them as proxies for real deployment. Clarifying this distinction can help avoid research practices that produce results not applicable outside simulators.","discovery_kind":"review","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest_assumption matches the paper's own framing (distinction prevents misleading conclusions). Because the manuscript is explicitly a position paper rather than a technical result, and no internal contradiction or over-claim is visible, the UNVERDICTED stance with low confidence remains appropriate until the community reacts to the examples.","tokens_in":1719,"tokens_out":302,"duration_ms":15923,"concrete_test":"Read the full examples and experiments in Sections 4–5; check whether each case shows a concrete change in algorithm choice or metric interpretation once the two use-cases are labeled separately. If labeling produces no material difference in the reported conclusions, the call for distinction has limited practical force.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper is a position piece whose central claim is that RL work should explicitly separate 'solving the simulator' from 'using the simulator as a proxy for deployment.' It supports this by enumerating differences in allowable simulator access, suitable algorithms, and evaluation metrics, then illustrating conflation problems with examples and simple experiments. No formal theorem, scaling law, or quantitative result is asserted that could be internally inconsistent; the argument is normative and definitional rather than deductive. The examples are presented as illustrations rather than exhaustive proof, so the load-bearing step is simply whether the community finds the proposed distinction useful. No technical flaw or unsupported empirical claim appears in the structure.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"This position paper claims that reinforcement learning (RL) research often conflates two distinct uses of simulators: solving the simulator (achieving high performance within it) versus using it as a proxy for learning deployable policies. The authors argue these use cases differ in terms of simulator access constraints, appropriate algorithms, and evaluation metrics. They support this by discussing the differences and highlighting issues from not distinguishing them, backed by examples and simple experiments, calling for clearer practices in the community.","tokens_in":1783,"tokens_out":201,"duration_ms":25023,"significance":"If adopted, the proposed distinction would improve clarity in RL experimental design and reduce the risk of misleading conclusions from simulator-specific optimizations that do not transfer. The manuscript earns credit for its explicit enumeration of differences across constraints, algorithms, and metrics, as well as for grounding the normative argument in illustrative examples rather than unsubstantiated assertion.","major_comments":[],"minor_comments":[],"recommendation":"accept","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their thoughtful review and positive recommendation to accept the manuscript. Their summary accurately captures the core argument of the position paper.","responses":[],"tokens_in":1204,"tokens_out":47,"duration_ms":8377,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that RL experiments often slide from using a simulator as a stand-in for real deployment into just solving the simulator itself, and the authors want papers to state which one they are doing. They lay out the differences clearly: solving the simulator permits heavy overfitting, unlimited resets, and metrics that reward simulator-specific tricks, while the proxy case restricts those and needs different evaluation. The examples they give show how mixing the two can produce misleading claims about generalization or efficiency.\n\nThe paper does this without overclaiming. The argument is definitional and practical rather than a new theorem, and the simple experiments serve as illustrations rather than proof. That keeps the tone proportionate.\n\nThe soft spot is that the work stays at the level of examples and does not measure how common the conflation is across recent RL papers or how often it actually changes conclusions. A reader could agree with the distinction yet still wonder whether the problem is large enough to change reviewing standards. No internal contradictions appear in the logic.\n\nThis is for RL researchers who run simulator benchmarks and for reviewers who evaluate those papers. Someone already thinking about experimental hygiene will get value from the framing. It deserves peer review because the issue affects a large part of the field and a published version could influence norms even if it remains a position piece.","headline":"The paper usefully pushes RL work to label whether a simulator run is about beating that sim or about learning for later deployment, and the distinction holds up.","tokens_in":2260,"tokens_out":338,"would_cite":false,"duration_ms":19882,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"RL researchers need to distinguish between solving simulators and using simulators as a proxy for learning in deployment.","keywords":["reinforcement learning","simulators","benchmarks","deployment","proxy evaluation","research practices","sequential decision making"],"falsifier":"An experiment in which methods developed under one use case are applied to the other and no differences appear in outcomes, constraints, or conclusions.","tokens_in":2590,"feed_emoji":"🤖","tokens_out":618,"duration_ms":25758,"temperature":0.7,"pith_summary":"The paper argues that reinforcement learning experiments often blur two separate goals when using simulators. One goal is to achieve high performance inside the simulator environment itself. The other is to develop agents whose learned behavior will transfer to deployment settings where the simulator is no longer available. Conflating these goals leads to algorithms, constraints, and metrics that fit one purpose but produce misleading results for the other. The authors illustrate the problems with examples and simple experiments, and call for researchers to state explicitly which use case they intend.","feed_headline":"RL must separate simulator solving from proxy use","feed_subtitle":"Conflating the two produces methods that only work inside the simulator and distorts progress toward deployment.","key_machinery":"The distinction between solving simulators (optimizing performance inside the given environment) and using simulators as a proxy (developing policies intended for use outside the simulator in deployment).","core_discovery":"The central claim is that RL researchers need to distinguish between two use cases of simulators: solving simulators and using simulators as a proxy for learning in deployment. These two settings differ in the constraints placed on how the agent may interact with the simulator, in the algorithms that are appropriate, and in the evaluation metrics that make sense. Failing to keep the distinction clear allows solutions meant only for the simulator to be presented as progress toward deployable agents and produces misleading conclusions about research progress, as shown by examples and simple experiments. The work is a call for the community to begin stating clearly how simulators are being used","pith_inferences":["Benchmark suites could add separate tracks or reporting requirements for each use case to reduce mismatched comparisons.","The distinction highlights why some high simulator scores fail to translate when the simulator is removed at test time.","Similar separation of goals may be useful in other simulation-heavy fields such as robotics control."],"forward_implications":["Algorithms developed under one use case may violate the access constraints of the other.","Evaluation metrics chosen for simulator solving can overstate progress toward deployment.","Research conclusions about general-purpose sequential decision making can rest on methods that only work inside the simulator.","Published work would become clearer if authors stated which use case they address."],"fun_headline_variants":["RL researchers must distinguish simulator solving from proxy use","Conflating simulator solving with proxy use distorts RL progress","Distinguish between solving simulators and proxy use in RL","RL experiments require separating simulator solving from proxy use"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That failing to distinguish the two simulator use cases produces issues and misleading conclusions in RL research.","fun_headline_variants_meta":{"raw":{"variants":["RL researchers must distinguish simulator solving from proxy use","Conflating simulator solving with proxy use distorts RL progress","Distinguish between solving simulators and proxy use in RL","RL experiments require separating simulator solving from proxy use"]},"model":"grok-4.3","cost_usd":0.004346,"raw_usage":{"total_tokens":2197,"prompt_tokens":702,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":43462000,"prompt_tokens_details":{"text_tokens":702,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1434,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":702,"tokens_out":61,"duration_ms":15046,"temperature":1.0,"reasoning_tokens":1434,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T01:26:46.795745+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An experiment in which methods developed under one use case are applied to the other and no differences appear in outcomes, constraints, or conclusions.","supporting_citations":[],"review_version":1}