{"id":"dd6d949c-f71b-4146-840e-99f66079dca9","arxiv_id":"2607.15432","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Prediction-driven aggregate early bed requests reduce simulated ED boarding time by 30–70% and ED length of stay by 6–15% with modest prepared-bed idle time.","lead":"Hospitals could cut emergency-department crowding by requesting inpatient beds before patients are formally admitted, using predicted admission probabilities. In a simulation based on 2019 data from one North Carolina ED, this approach cut boarding time by 30–70% and overall ED stay by 6–15%, at the cost of modest idle time on prepared beds.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline boarding/LoS reductions depend on unvalidated hospital-side parameters (TPP from Shi et al. 2016; reserved capacity 337 tuned in OS.2.3); a sensitivity sweep is needed before the 30–70% range is accepted.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: the hospital-side parameters (TPP from Shi et al. 2016, hLoS from Shi et al. 2014, reserved capacity 337 fitted in OS.2.3) are not estimated from the study ED and are not identified by the validation, which only matches mean ED LoS under current practice. My pass confirms this is the central soft spot: the quantitative headline is a counterfactual prediction whose magnitude depends on these inputs, and the manuscript's own limitation statements (OS.2.2) support this. The paper's other components (analytical heuristics, 1,000-year batch-means CIs, disclosed tuning and request caps) are handled transparently and do not raise an independent concern. Because the issue is an addressable but unresolved external-validity gap, the CONDITIONAL verdict already reflects it; no adjustment is needed. Agreement is 'agree' rather than 'partial' because the reader's statement and my concern share the same root cause, even though I emphasize that the capacity calibration is to the same target metric rather than a separate validation.","tokens_in":45378,"tokens_out":7303,"duration_ms":74377,"concrete_test":"Re-run the Arena SOS evaluation with the three unestimated hospital-side inputs varied one at a time: TPP mean (1.5, 2.5, 3.3, 5, 6 h), hLoS scale (0.5x, 1x, 2x), and reserved capacity (250, 300, 337, 400, 450), recomputing the Table 13 NHP rows. If any plausible combination (e.g., TPP mean 2h or capacity 400) drops the boarding-time reduction below 30% or pushes idle time above 3h, the abstract's unconditional 30–70% and 'modest idle time' claims should be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim — 30–70% boarding and 6–15% LoS reductions — is produced entirely by the Arena simulator. Its hospital-side components are the weakest link: transfer-preparation time is a truncated lognormal taken from Shi et al. [2016] (mean 3.3h, CV 0.6, max 12h; OS.2.2), hospital LoS is adapted from Shi et al. [2014] Tables 4 and 11, and the 337-bed reserved capacity is a tuning parameter selected to match observed second-stage durations (OS.2.3). The paper states directly that data 'lack the granularity' and 'do not allow reliable estimation' of these inputs. Because the CP baseline (boarding 248 min) is itself an output of these assumptions, the proportional reduction under early requests is not a measured property of the ED; it is a prediction of a model whose counterfactual validity has not been tested. In particular, if real bed-preparation times are shorter or hospital capacity can be flexed, early requests may merely shift idle time upstream rather than reduce boarding, shrinking or erasing the advertised benefit. This is not a fatal flaw — the mechanism is credible and the authors disclose the borrowings — but the headline range is conditional on unvalidated hospital dynamics.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an MDP-based framework for generating aggregate early inpatient bed requests from an emergency department, using admission-probability and time-to-disposition predictions for current ED patients. Three heuristics are developed: an approximate dynamic programming policy (MHP), a newsvendor-type policy (NHP), and a deep Q-learning policy (DQN). The policies are evaluated in a high-fidelity Arena simulator calibrated to 2019 data from a large North Carolina ED. The central quantitative claim is that proactive aggregate bed requests reduce average boarding time for admitted patients by 30–70% and average ED length of stay by 6–15%, at the cost of modest average idle time for prepared inpatient beds. Analytical results (Propositions 4.1–4.5) are stated with proofs in the online supplement, and the simulation reporting is detailed: 1,000-year batch-means runs, 95% CI half-widths, and full parameter tables for all scenarios.","tokens_in":45660,"tokens_out":4646,"duration_ms":50574,"significance":"If the simulation results are accepted, the paper would make a useful operational contribution: it shows that a relatively simple, interpretable newsvendor heuristic can capture most of the benefit of more complex RL-based approaches, with an explicit trade-off between ED performance and inpatient-bed idle time. The analytical framework is general, and the decision-tree approximation of MHP/DQN is a pragmatic step toward implementation. The paper is also transparent about data limitations and uses a plausible dual-simulator architecture. However, the headline quantitative claims rest on hospital-side model components (transfer preparation time, hospital length of stay, and reserved bed capacity) that are not estimated from the study hospital's own data; the paper even states that its data 'lack the granularity' to estimate them (OS.2.2). This is a significant caveat because the CP baseline itself is an output of those assumptions, and the proportional reductions are therefore predictions of a model whose counterfactual validity has not been tested. The central mechanism is credible and the policy comparisons are internally consistent, but the quantitative ranges should be regarded","major_comments":[{"comment":"The hospital-side parameters that drive the headline reductions are not estimated from the study ED's 2019 data. Transfer preparation time is a truncated lognormal taken from Shi et al. (2016) (mean 3.3h, CV 0.6, max 12h); hospital LoS is adapted from Shi et al. (2014) (a Singaporean hospital); and the reserved bed capacity of 337 is a tuning parameter selected to match observed second-stage durations (OS.2.3). Since boarding time under CP and under every proactive policy is co-determined by TPP, hLoS, and bed capacity, the claimed 30–70% boarding reduction and 6–15% LoS reduction are conditional on these unvalidated inputs. A sensitivity analysis over TPP mean/CV, hLoS distribution, and reserved bed count (e.g., ±20–30%) is necessary to establish the robustness of the headline ranges. Without it, the abstract's quantitative claims overstate what has actually been demonstrated.","section":"OS.2.2–OS.2.3, Section 5.1"},{"comment":"The MHP and DQN policies are approximated by decision trees with 150–300 nodes, and Tables 13–16 report the performance of these tree-approximated policies. No fidelity metric is provided: the paper does not report the fraction of state-action pairs where the tree disagrees with the original heuristic, nor the performance difference between the tree policy and the original policy. If the decision-tree approximation is lossy, the relative performance of MHP and DQN in the Arena evaluation could be materially degraded, and the recommendation to prefer NHP (or the claim that DQN is smoother) might be an artifact of the approximation. Please report the approximation error or validate the tree policies against the original policies in a subset of scenarios.","section":"Section 5, 'Approximation by decision trees'"},{"comment":"Partial bed eligibility is modeled by multiplying the heuristic's request quantity by 0.7. This implicitly assumes that the decision maker knows the exact eligibility fraction and that the demand from ineligible patients can be correctly handled by simple scaling. A more faithful implementation would restrict early requests to eligible patients or incorporate eligibility into the state and action space. As written, the SOS-P results are an approximate sensitivity check rather than a validation of the policies under the stated 'partial bed requests' condition; the abstract's claim that benefits persist in that setting is therefore weaker than presented.","section":"Section 5.2 (SOS-P)"}],"minor_comments":[{"comment":"Typo: 'Phyton simulator' should be 'Python simulator'.","section":"Section 5, OS.2.1"},{"comment":"The abstract states LoS reduction of 6–15%, but Section 5.1 says 'by 6–16%' and Table 13 shows reductions up to about 16%. Please reconcile the range.","section":"Abstract vs. Section 5.1"},{"comment":"Header typo: 'T uning Parameters' should be 'Tuning Parameters'.","section":"Table 13"},{"comment":"The NHP uses a symmetric triangular approximation for the demand distribution, justified by closed-form tractability. Since NHP is a recommended policy, it would strengthen the paper to report the sensitivity of NHP's trade-off frontier to the demand-shape assumption (e.g., a normal or empirical distribution), even if only in an appendix.","section":"Section 4.3, Proposition 4.5"}],"recommendation":"major_revision","confidential_remarks":"The analytical propositions are solid and the simulation reporting is careful. The main risk is that the headline quantitative ranges depend on hospital-side parameters that are borrowed or tuned rather than estimated from the study hospital's data. A serious revision should add a sensitivity analysis over those parameters or substantially soften the claims. The SOS-P implementation also deserves a more faithful treatment. I do not see a fatal error, but the paper is not ready for acceptance in its current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The two things you should know: the paper has a genuinely new idea, and its headline number is a model prediction, not a measured fact. The system-level aggregate bed-request formulation — hourly requests based on aggregated admission probabilities and bed status — is a real departure from the patient-level policies in Chen et al. 2023 and Qiu et al. 2015. The analytical propositions (4.1–4.5) are derived cleanly and look correct given their assumptions. The simulation work is transparent and careful: 1,000-year batch-means runs, CI half-widths, full parameter tables, and honest disclosure of tuning choices. Credit where due: NHP is a nice, interpretable heuristic that dominates the others on the LoS/idling frontier, and the CoV analysis for request smoothness is a useful addition. The paper is also commendably upfront about its weak spots.\n\nThe soft spot is exactly where the stress-test note points. The 30–70% boarding reduction and 6–15% LoS reduction come entirely from the in-house Arena simulator. The hospital side of that model is partly borrowed: transfer preparation time is a truncated lognormal from Shi et al. 2016 (Singapore), hospital LoS is adapted from Shi et al. 2014, and the 337-bed reserved capacity was tuned to match observed ED LoS. The paper states directly that the data \"lack the granularity\" to estimate these inputs. That makes the proportional reductions conditional on hospital dynamics that have not been validated for this ED. If real TPP is shorter or hospital capacity can flex, early requests may just shift idle time upstream, shrinking the advertised benefit. This is an addressable weakness, not a fatal one — the mechanism is credible and the authors say a pilot is needed. But the headline range should not be cited as established without a sensitivity sweep over TPP and reserved capacity, and ideally an out-of-sample check. Minor quibble: GP, the simplest benchmark, already gets much of the benefit, so the added value of MHP/DQN is mainly smoother request patterns; that is worth stating more plainly.\n\nWho is this for: people working on ED boarding, hospital flow, or applied queueing. It deserves a serious referee. My recommendation: send it to review, but the referee should push for the sensitivity analysis and a clear statement that the quantitative claims are model-based predictions pending real-world validation. The idea is good; the validation is the next step.","headline":"Solid system-level early bed-request framework with careful simulation; headline 30–70% boarding reduction is conditional on borrowed hospital-side parameters and needs sensitivity analysis before being taken at face value.","tokens_in":46249,"tokens_out":1704,"would_cite":true,"duration_ms":19214,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["90C40","90B22"],"pacs":[],"model":"deepseek-v4-flash","headline":"Using hourly aggregate bed requests driven by admission predictions can cut average boarding time for admitted ED patients by 30–70 percent and overall ED length of stay by 6–15 percent, at modest idle-bed cost.","keywords":["emergency department boarding","inpatient bed requests","admission prediction","Markov decision process","reinforcement learning","newsvendor heuristic","discrete-event simulation","healthcare operations"],"falsifier":"Measure transfer-preparation times, inpatient lengths of stay, and reserved-bed capacity from a single hospital, re-run the same policy comparison in that hospital's own simulation, and check whether the 30–70% boarding reduction and 1–95 minute idle-time range still hold. Alternatively, run a six-month pilot using hourly aggregate newsvendor requests on a general-medicine unit and compare boarding times with the prior year; a reduction well below 30% or idle times far above 95 minutes would refute the quantitative claim.","tokens_in":45172,"feed_emoji":"🏥","tokens_out":5249,"duration_ms":52731,"temperature":0.7,"pith_summary":"This paper tries to establish that emergency departments can reduce boarding—the hours admitted patients wait in the ED for an inpatient bed—by having the ED request beds in bulk and in advance, rather than one-by-one after each admission decision. Using predicted admission probabilities and time-to-disposition distributions for everyone currently in the ED, a decision maker asks the hospital for a certain number of beds each hour. The paper formulates this as a Markov decision process and derives three implementable policies, then tests them in a simulation calibrated to a large academic ED. The reported result is that proactive aggregate requests reduce average boarding time by 30–70% and average ED length of stay by 6–15%, while prepared inpatient beds sit idle for only modest average durations. If true, this gives hospitals a concrete way to use prediction tools they may already have.","feed_headline":"Hourly bed forecasts cut ED boarding by 30–70 percent","feed_subtitle":"Simulation shows aggregate hourly bed requests also shorten overall ED stays by 6–15 percent, with modest idle-bed time.","key_machinery":"The load-bearing object is the aggregate hourly request quantity, computed from a scalar imbalance between expected near-term admissions and bed supply—the paper writes it as φ = Σ α_i n_i + k − r, where α_i is a patient's admission probability, n_i counts patients in that class, k is the number boarding (negative if beds are idle), and r is the number of requested beds still in preparation. Three policies use this state in different ways: MHP applies a linear policy from an auxiliary finite-horizon MDP; NHP solves a newsvendor problem whose demand distribution is a symmetric triangular fit to the mean and variance of admissions-minus-ready-beds over eight hours; DQN learns a Q-function over","core_discovery":"The central discovery, on the paper's own terms, is that the timing and quantity of inpatient bed requests can be treated as an ED-level inventory problem instead of a per-patient decision. The paper's MDP state tracks patients by admission-probability class, the number boarding (or idle beds, signed), and the number of requested beds still being prepared. The action is how many additional beds to request. From this formulation it derives an MDP-based heuristic, a newsvendor heuristic that matches the mean and variance of future demand to a triangular distribution over an eight-hour horizon, and a deep-Q-learning heuristic. In a discrete-event simulation of a 2019 ED, the newsvendor policy p","pith_inferences":["Beyond the paper, the aggregate-request logic suggests a low-tech implementation: hospital managers could start with the newsvendor formula and an eight-hour lookahead, requiring only admission probabilities and a rough bed-preparation-time mean.","An untested corollary is that the value of proactive requests should increase with bed-preparation time: the longer the transfer-preparation process, the more boarding is avoided by starting it early. The paper does not report this comparative static.","The quantitative claims are simulation-based; the natural next step, which the paper explicitly flags, is a controlled pilot in a real ED. If the pilot uses only general-medicine beds, the 70% eligibility result is the relevant benchmark.","Because the MDP treats beds as interchangeable, the framework's benefits may concentrate in hospitals with large general-medicine admitting units; specialized hospitals would likely need a hybrid of aggregate and per-patient requests."],"forward_implications":["An ED with a two-class triage admission predictor (low admission probability around 8.5%, high around 57.7%) can expect boarding time to drop from roughly 248 minutes to 77–179 minutes, depending on how much idle bed time the hospital tolerates.","Total ED length of stay for all patients drops by about 26–65 minutes (6–16%), which can shorten waits for other ED patients by freeing beds sooner.","Even a simple greedy policy—requesting the expected number of admissions minus beds already requested—captures a large share of the benefit, but without tunable cost parameters.","The newsvendor heuristic is the best default for the boarding/idle tradeoff, while the reinforcement-learning heuristic is preferable when downstream hospital stability depends on a smooth bed-request process.","Benefits persist with only 70% of beds interchangeable and under a six-week pandemic surge with expanded hospital capacity, so the policy need not wait for perfectly general bed pools."],"fun_headline_variants":["Proactive bed requests cut ED boarding by 30–70%","Early bed requests slash ED waits by up to 70%","Bed-request forecasting trims ED boarding time","Aggregate bed requests reduce ED boarding drastically"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The 30–70% boarding reduction is produced by a simulation in which hospital-side inputs—transfer-preparation time, inpatient length-of-stay distribution, and the number of beds reserved for ED admissions—are borrowed from other hospitals or tuned to match observed ED data, not estimated from the study ED's own records; different hospital dynamics could shrink the benefit.","fun_headline_variants_meta":{"raw":{"variants":["Proactive bed requests cut ED boarding by 30–70%","Early bed requests slash ED waits by up to 70%","Bed-request forecasting trims ED boarding time","Aggregate bed requests reduce ED boarding drastically"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000382,"raw_usage":{"total_tokens":1900,"prompt_tokens":817,"completion_tokens":1083,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":1030}},"tokens_in":561,"tokens_out":1083,"duration_ms":8312,"temperature":1.0,"reasoning_tokens":1030,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T23:23:22.727793+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure transfer-preparation times, inpatient lengths of stay, and reserved-bed capacity from a single hospital, re-run the same policy comparison in that hospital's own simulation, and check whether the 30–70% boarding reduction and 1–95 minute idle-time range still hold. Alternatively, run a six-month pilot using hourly aggregate newsvendor requests on a general-medicine unit and compare boarding times with the prior year; a reduction well below 30% or idle times far above 95 minutes would refute the quantitative claim.","supporting_citations":[],"review_version":1}