{"id":"9c971b88-5936-4fe5-b252-237b5be489bb","arxiv_id":"2506.17602","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"The ARCH-COMP25 stochastic models report introduces new benchmarks and compares six tools for verification and policy synthesis of stochastic systems.","lead":"This report summarizes the 2025 ARCH friendly competition for stochastic model verification, introducing a new water distribution benchmark, simplified benchmark variants, and results from six tools. It matters because it maps which formal verification and controller synthesis tools can handle which stochastic benchmarks, and how fast.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing threat is the acknowledged ambiguity of the 7D building automation CS2 benchmark (§5.3.2), not the water-tower Euler step, because CS2 results from two tools are presented as comparable.","rationale":"I read the report in good faith. The strongest claim is about reliable, reproducible comparison and an updated tool-problem map. The most load-bearing premise for that claim is that each named benchmark has one well-defined specification that all tools solve. The report itself undercuts this on CS2. The water-tower exactness problem, by contrast, affects a benchmark with no reported 2025 results, so it cannot threaten the comparison claim; it threatens only the future usability of one proposed benchmark. The missing confidence intervals for IPS-FAS in §4.2 are a presentation weakness but not a reason to reject the central comparison. The public repositories and the report's candid statements about CS2's ambiguity are credit, but the ambiguity means the CS2 comparison should be explicitly labeled provisional until a canonical model is chosen. Hence I keep the reader's CONDITIONAL verdict.","tokens_in":26078,"tokens_out":12382,"duration_ms":123039,"concrete_test":"Compare the BA CS2 model files in the public repositories (IntervalMDPAbstractions.jl arch-comp/2025 and SySCoRe-software) against the canonical matrices printed in §5.3.2, checking whether A, B, Q, Bw and the control/noise semantics match; then re-run both tools on the canonical model with identical state/input grids. If the reported CS2 satisfaction probabilities still differ by more than the stated error bounds, the CS2 comparison is invalid and must be qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The report's central claim is a reliable, reproducible tool comparison. Section 5.3.2 admits 'there is confusion between implementations about which is the correct system (e.g., whether the dynamics are linear or affine, and whether the dynamics include control)' for the CS2 building automation benchmark, then prints one version of A, B, Q, and Bw. Section 5.4.1 benchmarks 'CS2' with SySCoRe without stating which dynamics or control semantics it used. The two tools' headline numbers on CS2 are far apart: IntervalMDP.jl reports a certified lower bound below 1e-10 with mean error above 99.99%, while SySCoRe reports peak satisfaction probabilities around 0.96 for the same nominal benchmark. If the implementations actually solve different systems, the 'which tools can solve which problems' map is unreliable on this benchmark and the reproducibility claim fails. The water-tower forward-Euler exactness issue in §4.1.1 is a real modeling concern, but it is less load-bearing for this report because no 2025 tool result is reported on the new water benchmark.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is the ARCH-COMP25 category report for stochastic models. It describes the 2025 friendly competition, introducing three tools (PRoTECT, IMPaCT, IntervalMDP.jl) and reporting new results from them and from SySCoRe, hpnmg, and AMYTISS on established and new benchmarks. It also introduces a water distribution network benchmark and a collection of reduced benchmark variants intended to broaden tool participation, and it discusses scalability and next steps. Results are reported per tool with runtimes, memory usage, and satisfaction probabilities, and the accompanying code is made publicly available.","tokens_in":26344,"tokens_out":7558,"duration_ms":77310,"significance":"If the reported measurements are trustworthy, this report provides a useful snapshot of the current tool landscape and a reusable benchmark suite. Its strengths include public code repositories, candid reporting of a failed abstraction in §5.3.5, and an explicit caveat about the CS2 benchmark definition in §5.3.2. The new water-distribution benchmark targets a practically motivated infrastructure domain, though no tool result is reported on it in this edition. The per-tool results and the honest documentation of negative outcomes are valuable to the verification and synthesis community, provided the comparability concerns raised below are addressed.","major_comments":[{"comment":"The CS2 building-automation results are presented as part of a common benchmark comparison, but the report itself states in §5.3.2 that \"there is confusion between implementations about which is the correct system (e.g., whether the dynamics are linear or affine, and whether the dynamics include control).\" The full dynamics printed there are used for the IntervalMDP.jl experiments, yet §5.4.1 does not state which dynamics, affine term Q, control input, or noise model SySCoRe used for CS2. Given that the two tools report satisfaction numbers differing by many orders of magnitude (a certified lower bound below 1e-10 with mean error above 99.99% versus peak probabilities around 0.96), the reader cannot verify that the two tools solved the same mathematical problem. The authors must either specify the exact model used by each tool and confirm that they coincide, or explicitly mark the CS2 numbers as not comparable.","section":"§5.3.2 and §5.4.1"},{"comment":"The reduced benchmarks are introduced as a way to make tools comparable, but the subsequent results do not consistently identify which scenario variant is solved. For the reduced patrol robot, §4.3.1 defines four distinct scenarios, yet IMPaCT's run in §5.2.5 is described only as a reachability specification while also mentioning avoid sets, AMYTISS in Table 11 solves reach-while-avoid, and IntervalMDP.jl in §5.3.4 solves both reachability and reach-while-avoid. A reader cannot reconstruct from the text which scenario each tool actually solved, which undermines the claimed comparability of the reduced benchmark suite. A table mapping each tool to the exact specification variant (including horizon, disturbance presence, and discretization) should be added.","section":"§4.3 and §5.2.5/§5.3.4/Table 11"}],"minor_comments":[{"comment":"The statement that Table 7 \"underscores the superiority of the IPS-FAS algorithm over MC simulation\" overstates what the table shows; zero estimates from crude Monte Carlo for rare events are an expected artifact of finite sampling, and no confidence intervals or variance estimates are reported for either estimator. Please rephrase to a factual statement, e.g., that IPS-FAS yields nonzero estimates where crude MC returns zero.","section":"§4.2 / Table 7"},{"comment":"The forward-Euler discretization is exact only under the stated constancy assumption on flows, and the paper acknowledges this. Since Figure 1 shows consumption varying within the 15-minute sampling interval, please add a sentence quantifying or explicitly bounding the resulting discretization error, or state clearly that the benchmark model is an approximation of the continuous-time balance.","section":"§4.1.1 / Eq. (2)"},{"comment":"The automated-vehicle result is reported with a certified satisfaction probability that is trivially zero due to grid misalignment with the target set. This is honest, but it should be labeled as a setup artifact rather than a benchmark outcome, so that it is not read as evidence about the tool's capability on the AV benchmark.","section":"§5.3.5"},{"comment":"The sentence \"We have updated the output to be only the temperature in zone one for CS1\" changes the benchmark's output mapping and should be reflected explicitly in the specification description; as written, the reader must infer that the SySCoRe CS1 result is not for the original output.","section":"§5.4.1"}],"recommendation":"major_revision","confidential_remarks":"The key risk is the CS2 benchmark ambiguity: if the SySCoRe model cannot be confirmed to match the version solved by IntervalMDP.jl, the paper should not claim a comparison on that benchmark. The rest of the report is a useful competition summary, and the requested revision is about precision and disclosure rather than a fundamental flaw in the measured results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, useful competition report. Its central contribution is infrastructure—a new water-network benchmark, reduced variants, and a 2025 map of which six tools can solve what—not a new scientific result. It deserves a real referee, but only after one ambiguity is fixed.\n\nWhat is new: the water distribution network (Sec. 4.1), the reduced patrol robot/AV/building automation suite (Sec. 4.3), first ARCH results for IMPaCT, IntervalMDP.jl, PRoTECT, AMYTISS on the reduced benchmarks, and hpnmg's guided simulation on the WS benchmark (Table 10). Code is public for benchmarks and most runs, which makes the comparison genuinely repeatable. The report is also honest where things fail: it flags the IntervalMDP.jl AV result as trivially zero (5.3.5) and admits the CS2 confusion (5.3.2). That transparency is real credit.\n\nThe soft spots, in order of real weight. The CS2 building automation ambiguity is the load-bearing one. Section 5.3.2 admits implementations disagree about whether the dynamics are linear or affine and whether control is included, then prints one set of A, B, Q, Bw. Section 5.4.1 benchmarks \"CS2\" with SySCoRe but never states which variant it used. The results are wildly different—IntervalMDP.jl reports a certified lower bound below 1e-10 with mean error over 99.99%, while SySCoRe reports peak satisfaction around 0.96. If the tools solved different systems, the headline \"which tools can solve which problems\" map is unreliable on exactly that benchmark. This needs to be fixed before the report is citable as a reproducible comparison: state the variant each tool ran, or drop CS2 from the comparison. The stress-test note is right that this is more load-bearing than the water-tower Euler issue.\n\nThe water-tower forward-Euler exactness (Sec. 4.1.1) is a real modeling caveat—flows are not plausibly constant over 15-minute samples, and Figure 1 shows intra-day fluctuation—but it matters less for this year's results because no 2025 tool run is reported on the water benchmark. It should be fixed before that benchmark becomes a baseline next year.\n\nMinor: Sec. 4.2's phrase that IPS-FAS \"underscores its superiority\" over MC overstates. MC returning zero on rare events is expected, not a general defeat, and the MC point estimates have no confidence intervals. These are wording issues, not fatal.\n\nOn balance, the benchmarking contribution holds up. The numbers come from concrete tool executions with public repositories, and failures are disclosed rather than hidden. The paper is for people building or using stochastic verification tools and for ARCH participants; it is infrastructure, not a research breakthrough. I would send it to a serious referee, with revision focused on the CS2 ambiguity.","headline":"Useful benchmarking infrastructure with an honest disclosure culture; the CS2 ambiguity must be resolved before the comparison table can be trusted.","tokens_in":26947,"tokens_out":2407,"would_cite":true,"duration_ms":23230,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This report claims that the 2025 stochastic-models competition produced a reproducible comparison of six verification and synthesis tools, anchored by a new water-distribution network benchmark and a suite of reduced benchmarks.","keywords":["stochastic hybrid systems","formal verification","policy synthesis","benchmark suite","water distribution network","interval Markov decision processes","barrier certificates","rare-event simulation"],"falsifier":"Resample the Bjerringbro consumption data at one-minute resolution and recompute the certified safety probability for the same controller; if the re-verified probability falls materially below the reported guarantee, the constant-flow discretization assumption is the point of failure. Alternatively, compare simulated tower volumes against measured volumes over the 88-day dataset.","tokens_in":25910,"feed_emoji":"🚰","tokens_out":8995,"duration_ms":85540,"temperature":0.7,"pith_summary":"This report is the outcome of a friendly annual comparison of software tools for formal verification and policy synthesis of stochastic systems. It claims that the 2025 edition now provides a reliable and reproducible map of which tools can solve which problems, thanks to three first-time participants and a deliberately broadened benchmark suite. The centerpiece is a new water-distribution network benchmark in which a water tower must be kept between 28 L and 155 L under random consumption; the task is to synthesize or verify a controller that guarantees safety with a stated probability. To make tools with very different model classes comparable, the organizers added simplified versions of patrol robot, autonomous vehicle, and building automation benchmarks. If the report's results hold, the community gains a shared testbed for stochastic safety analysis and a clearer picture of the current state of the art.","feed_headline":"New water-network benchmark tests six stochastic-model tools","feed_subtitle":"A consumption-driven water tower plus simplified variants lets previously incomparable tools be measured side by side.","key_machinery":"The central supporting object is the water-tower volume-balance model: Eq. (2), the forward-Euler discretization $V(t+t_s)=V(t)+t_s\\sum_i q_i(t)-t_s\\sum_i d_i(t)$, together with the pump-power equation $P_i(t)=\\frac{1}{\\eta_i} q_i(t)\\left(r_{f,i}|q_i(t)|q_i(t)+r_{f,\\Sigma}|q_\\Sigma(t)|q_\\Sigma(t)+\\rho_w g_0(h_V(t)+h_i)\\right)$ and the box constraint $28\\,\\mathrm{L}\\le V(t)\\le155\\,\\mathrm{L}$. This model turns infrastructure safety into a concrete stochastic synthesis problem. The second mechanism is the suite of reduced benchmarks—patrol robot, autonomous vehicle, building automation—designed so tools that discretize state space and struggle with high dimensions can still be evaluated on the same specifications. The comparison mechanism across tools is a mix of formal guarantees: barrier-certificate bounds (PRoTECT), interval value iteration over IMDP abstractions (IMPaCT, IntervalMDP.jl), stochastic coupling relations (SySCoRe), and symbolic state-space construction combined with statistical model checking (hpnmg).","core_discovery":"On the report's own terms, the central result is not a new theorem but a usable, reproducible comparison infrastructure. Three first-time tools—PRoTECT, IMPaCT, and IntervalMDP.jl—are shown solving safety, reachability, and reach-while-avoid specifications on a common set of benchmarks, while established tools SySCoRe, hpnmg, and AMYTISS extend their coverage to simplified benchmark variants. The new water-distribution benchmark models a tower whose volume evolves as $V(t+t_s)=V(t)+t_s\\sum_{i=1}^{N_q} q_i(t)-t_s\\sum_{i=1}^{N_d} d_i(t)$, with pump power and consumption noise, and asks for a probabilistic safety guarantee that the volume stays within $28\\,\\mathrm{L}\\le V(t)\\le155\\,\\mathrm{L}$. The report also documents a lane-change collision scenario in which an interacting-particle rare-event estimator returns probabilities around $10^{-7}$ where Monte Carlo simulation returns zero, and a hpnmg guided-simulation engine whose rare-event confidence intervals overlap those of earlier tools while running much faster.","pith_inferences":["The water-network model's safety guarantee is only as good as the constant-consumption-per-interval assumption; re-verifying with sub-15-minute consumption data would show whether the reported probabilities hold under realistic demand fluctuations.","The same pump-power equations could turn the benchmark into an energy-versus-safety trade-off study: minimizing pumping cost while keeping the tower within bounds is a natural next synthesis objective.","The reduced-benchmark approach could be applied to the new water network itself—for example, a one-pump no-elevation variant—to let tools that cannot handle high-dimensional abstractions compete on an infrastructure-relevant model.","Connecting the lane-change rare-event result to the other benchmarks suggests a general pattern: tools that combine symbolic state-space construction with statistical simulation may be the most practical route for $10^{-6}$-level safety claims on hybrid models."],"forward_implications":["Safety certificates for the water network: the benchmark gives tool developers a concrete target—probabilistic safety for a 28–155 L tower subject to demand noise—so new synthesis algorithms can be tested against published numbers.","Comparable tool rankings: the reduced benchmark suite lets tools that previously applied to disjoint models be evaluated on the same specifications, making head-to-head comparison possible for the first time.","Rare-event reachability: the lane-change results show that importance-splitting estimation can produce collision probabilities around $10^{-7}$ where plain Monte Carlo returns zero, giving a practical way to certify very small risk.","Faster rare-event checking: hpnmg's guided simulation matches the confidence intervals of conventional simulation tools on the sewage benchmark in less time, suggesting a scalable route for rare-event verification.","Reproducibility baseline: because the benchmark code and repeatability packages were collected, the reported numbers can be re-run and extended by later editions."],"supporting_citations":[{"why":"Surveys the verification and synthesis techniques that the competing tools implement.","marker":"[2]"},{"why":"Introduces PRoTECT, whose barrier-certificate synthesis supplies the safety verification results for several benchmarks.","marker":"[4]"},{"why":"Introduces IMPaCT, the interval-MDP construction tool behind the reachability and synthesis results across multiple benchmarks.","marker":"[11]"},{"why":"Provides interval iteration, the algorithm IMPaCT relies on for infinite-horizon convergence with epsilon error.","marker":"[13]"},{"why":"Introduces IntervalMDP.jl, the Julia toolbox used for the abstraction and value-iteration results.","marker":"[18]"},{"why":"Describes the symbolic state-space plus statistical model checking method that underlies hpnmg's new guided simulation.","marker":"[29]"},{"why":"Introduces SySCoRe 2.0, whose stochastic model reduction produces the building-automation and Van der Pol results.","marker":"[35]"},{"why":"Gives the interacting-particle rare-event estimation algorithm that produces the lane-change collision probabilities.","marker":"[51]"},{"why":"Defines the building automation and anaesthesia benchmark dynamics used by PRoTECT, IMPaCT, IntervalMDP.jl, and SySCoRe.","marker":"[55]"},{"why":"Reports the 2021 water-sewage results against which the new hpnmg guided simulation is compared.","marker":"[59]"}],"fun_headline_variants":["Stochastic-model tools face off on new water benchmark","Rare-event estimator finds collisions Monte Carlo misses","Three new tools join stochastic-model benchmark race","Water tower test compares six stochastic verification tools","Simplified benchmarks make stochastic tools comparable at last"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The forward-Euler balance equation is treated as exact, which requires pump flows and consumption to stay constant over each 15-minute sampling interval; real consumption fluctuates within that interval, so the benchmark's safety probabilities may not describe the true water network.","fun_headline_variants_meta":{"raw":{"variants":["Stochastic-model tools face off on new water benchmark","Rare-event estimator finds collisions Monte Carlo misses","Three new tools join stochastic-model benchmark race","Water tower test compares six stochastic verification tools","Simplified benchmarks make stochastic tools comparable at last"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000836,"raw_usage":{"total_tokens":3596,"prompt_tokens":846,"completion_tokens":2750,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":462,"completion_tokens_details":{"reasoning_tokens":2680}},"tokens_in":462,"tokens_out":2750,"duration_ms":19363,"temperature":1.0,"reasoning_tokens":2680,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:05:06.435246+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Resample the Bjerringbro consumption data at one-minute resolution and recompute the certified safety probability for the same controller; if the re-verified probability falls materially below the reported guarantee, the constant-flow discretization assumption is the point of failure. Alternatively, compare simulated tower volumes against measured volumes over the 88-day dataset.","supporting_citations":[{"cited_title":"Automated verification and synthesis of stochastic hybrid systems: A survey,","cited_arxiv_id":null,"evidence_quote":"Surveys the verification and synthesis techniques that the competing tools implement."},{"cited_title":"IMPaCT: Interval MDP parallel construction for controller synthesis of large-scale stochastic systems,","cited_arxiv_id":null,"evidence_quote":"Introduces IMPaCT, the interval-MDP construction tool behind the reachability and synthesis results across multiple benchmarks."},{"cited_title":"Interval iteration algorithm for MDPs and IMDPs,","cited_arxiv_id":null,"evidence_quote":"Provides interval iteration, the algorithm IMPaCT relies on for infinite-horizon convergence with epsilon error."},{"cited_title":"Syscore 2.0: Toolset for formal control synthesis of continuous-state stochastic systems and temporal logic specifications,","cited_arxiv_id":null,"evidence_quote":"Introduces SySCoRe 2.0, whose stochastic model reduction produces the building-automation and Van der Pol results."},{"cited_title":"Interacting particle system based estimation of reach probability of general stochastic hybrid systems,","cited_arxiv_id":null,"evidence_quote":"Gives the interacting-particle rare-event estimation algorithm that produces the lane-change collision probabilities."},{"cited_title":"ARCH-COMP19 category report: Stochastic modelling,","cited_arxiv_id":null,"evidence_quote":"Defines the building automation and anaesthesia benchmark dynamics used by PRoTECT, IMPaCT, IntervalMDP.jl, and SySCoRe."},{"cited_title":"Arch-comp21 category report: Stochastic mod- els,","cited_arxiv_id":null,"evidence_quote":"Reports the 2021 water-sewage results against which the new hpnmg guided simulation is compared."}],"review_version":2}