{"id":"2a17a2e7-4176-4d2c-9301-3c77089fc85b","arxiv_id":"2606.09155","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Bridged SBI corrects biased low-fidelity posteriors via residual bridging to enable more accurate and cost-efficient high-fidelity simulation-based inference for particle simulator calibration.","lead":"The paper introduces Bridged SBI, a technique that uses cheap low-fidelity particle simulations to locate promising parameter regions and then applies a learned residual correction to adjust samples for high-fidelity consistency in robotic simulator calibration. If effective, it could lower the computational barrier for accurate parameter inference in expensive particle-based models used for earthwork robotics.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Sufficiency of LF posterior for reliable local residual bridge correction under limited HF data","rationale":"This matches the reader's weakest_assumption exactly and is the point on which the experimental superiority claim depends. The UNVERDICTED status is appropriate given the abstract-only review; even with full text the concern would require explicit coverage diagnostics in the results to be resolved.","tokens_in":1787,"tokens_out":293,"duration_ms":20841,"concrete_test":"In the sim-to-sim experiments, compute empirical coverage of the true parameter under the Bridged SBI 95% credible intervals across repeated runs; if coverage falls materially below 95% (or below the HF-only baseline), the residual bridge introduces bias.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim rests on the LF posterior being close enough to the HF target that a residual bridge, fit from few HF simulations, can map samples into the correct HF region without new biases or coverage failures. If the LF-HF shift (induced by particle count/dynamics changes) is not well-approximated by a local residual in parameter space, or if the learned correction overfits the limited HF budget, the method could produce posteriors with worse calibration than claimed. The abstract notes that Naive-MF suffers LF-induced miscoverage but provides no detail on how Bridged SBI's coverage is measured or guaranteed in the reported experiments.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript proposes Bridged SBI for cost-efficient high-fidelity simulation-based inference in particle-based robotic earthwork simulators. It first obtains a biased but informative LF posterior via inexpensive low-fidelity simulations, identifies a coarse high-density region, and then learns a local residual bridge from limited HF simulations to transport LF samples into HF-consistent regions by explicitly correcting the LF-HF discrepancy. The approach is positioned against HF-only SBI and Naive-MF (which suffers LF-induced miscoverage), with claims of superior accuracy and reliability demonstrated on sim-to-sim particle-parameter calibration and real-to-sim calibration using real soil observations, particularly under limited HF budgets.","tokens_in":1907,"tokens_out":522,"duration_ms":15963,"significance":"If the residual-bridge correction reliably transports samples without introducing new biases or coverage failures, the method would offer a practical route to high-fidelity posterior inference for computationally expensive nonlinear particle simulators, reducing the HF simulation budget while mitigating the systematic shifts induced by changes in particle count and dynamics. The explicit discrepancy modeling distinguishes it from naive multi-fidelity baselines and could generalize to other black-box simulators in robotics.","major_comments":[{"comment":"The central claim that Bridged SBI produces more accurate and reliable HF posteriors rests on the LF posterior remaining sufficiently informative for a local residual bridge (learned from limited HF data) to map samples without new biases or coverage failures. The abstract and introduction assert this alleviates Naive-MF miscoverage, but the manuscript provides no quantitative coverage diagnostics, calibration plots, or ablation on the magnitude of the LF-HF shift to substantiate that the residual correction does not overfit or degrade calibration under the reported HF budgets.","section":"Abstract and §4 (Experiments)"},{"comment":"The experiments claim superior performance on both sim-to-sim and real-to-sim tasks, yet the abstract (and by extension the reported results) contains no numerical metrics, error bars, or baseline comparisons (e.g., posterior mean error, coverage probability, or effective sample size). Without these, the load-bearing assertion that Bridged SBI outperforms HF-only SBI and Naive-MF cannot be evaluated for statistical significance or practical effect size.","section":"Abstract"}],"minor_comments":[{"comment":"Notation for the residual bridge function and the LF-HF discrepancy term should be introduced with explicit equations in the method section to clarify how the transport map is parameterized and optimized.","section":"Method"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive feedback. We address each major comment below, agreeing where the manuscript can be strengthened through revision.","responses":[{"response":"We agree that the current presentation would be strengthened by explicit quantitative coverage diagnostics. While the manuscript analyzes Naive-MF miscoverage and demonstrates through experiments that Bridged SBI produces more reliable posteriors than the baselines, it does not include dedicated calibration plots or an ablation on LF-HF shift magnitude. In the revision we will add coverage probability metrics, calibration plots, and an ablation study on the magnitude of the LF-HF discrepancy in §4 to directly substantiate that the residual bridge does not introduce new biases under the reported HF budgets.","revision_made":"yes","referee_comment":"[Abstract and §4 (Experiments)] The central claim that Bridged SBI produces more accurate and reliable HF posteriors rests on the LF posterior remaining sufficiently informative for a local residual bridge (learned from limited HF data) to map samples without new biases or coverage failures. The abstract and introduction assert this alleviates Naive-MF miscoverage, but the manuscript provides no quantitative coverage diagnostics, calibration plots, or ablation on the magnitude of the LF-HF shift to substantiate that the residual correction does not overfit or degrade calibration under the reported HF budgets."},{"response":"We agree that the abstract would benefit from concrete numerical results to allow readers to assess effect sizes and statistical significance. The experimental section (§4) contains the detailed comparisons with baselines, but these are summarized only qualitatively in the abstract. We will revise the abstract to include key quantitative metrics such as posterior mean errors, coverage probabilities, and baseline comparisons (with error bars where applicable) drawn from the reported experiments.","revision_made":"yes","referee_comment":"[Abstract] The experiments claim superior performance on both sim-to-sim and real-to-sim tasks, yet the abstract (and by extension the reported results) contains no numerical metrics, error bars, or baseline comparisons (e.g., posterior mean error, coverage probability, or effective sample size). Without these, the load-bearing assertion that Bridged SBI outperforms HF-only SBI and Naive-MF cannot be evaluated for statistical significance or practical effect size."}],"tokens_in":1505,"tokens_out":478,"duration_ms":15361,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper introduces Bridged SBI as a way to use cheap low-fidelity particle simulations to locate a rough parameter region, then fit a local residual correction from a small number of high-fidelity runs to adjust the LF posterior samples. This targets the specific failure where naive multi-fidelity SBI passes the biased LF posterior straight through and ends up with poor coverage on the HF target.\n\nThe approach is new in its explicit modeling of the LF-HF discrepancy as a transport step rather than relying on the LF posterior alone. It does a clean job of spelling out why the naive baseline suffers miscoverage when particle count or dynamics change between fidelities. The reported experiments on both sim-to-sim and real-to-sim soil calibration claim better accuracy and reliability than HF-only SBI or the naive baseline when HF budget is tight.\n\nThe main soft spot is the load-bearing assumption that the LF posterior stays close enough to the HF one for a residual learned from limited HF data to map samples correctly without new biases or coverage gaps. The stress-test concern lands here: if the shift from coarser particles is not well approximated by a local correction, the bridge could overfit or leave the true HF posterior uncovered. The abstract does not give quantitative coverage metrics or details on how the residual is validated, so it is hard to judge whether the claimed improvement holds under realistic LF-HF mismatches.\n\nThis work is aimed at people doing parameter calibration for expensive particle-based robotics simulators. A reader already working on multi-fidelity SBI or sim-to-real transfer would find the concrete fix useful. It deserves peer review because it identifies a real practical gap and offers a direct method plus experiments, even if the coverage guarantees need more scrutiny in revision.","headline":"Bridged SBI adds an explicit residual bridge to fix LF-induced bias in multi-fidelity SBI for particle simulator calibration, but the method's success depends on an assumption about local LF-HF shifts that the abstract leaves untested.","tokens_in":2431,"tokens_out":433,"would_cite":false,"duration_ms":16515,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Bridged SBI corrects biased low-fidelity posteriors to produce accurate high-fidelity parameter estimates at reduced cost.","keywords":["simulation-based inference","multi-fidelity modeling","posterior correction","particle simulation","robotic calibration","discrepancy modeling","earthwork simulation"],"falsifier":"An experiment in which high-fidelity posterior samples from Bridged SBI produce predictive simulations that match held-out high-fidelity or real observations no better than the Naive-MF baseline would falsify the central claim.","tokens_in":2691,"feed_emoji":"","tokens_out":652,"duration_ms":18257,"temperature":0.7,"pith_summary":"The paper establishes that a local residual bridge learned from limited high-fidelity simulations can transport samples from an inexpensive but biased low-fidelity posterior into regions consistent with the high-fidelity model. This matters for robotic earthwork calibration because direct high-fidelity simulation-based inference is computationally prohibitive while naive reliance on low-fidelity posteriors produces miscoverage. The method first locates a coarse high-density region with low-fidelity runs, then explicitly corrects the fidelity discrepancy rather than using the low-fidelity posterior directly.","feed_headline":"Bridged SBI corrects low-fidelity bias for accurate high-fidelity posteriors","feed_subtitle":"A local residual bridge learned from few high-fidelity runs avoids the coverage failures of naive multi-fidelity methods.","key_machinery":"The local residual bridge, which models and corrects the LF-HF discrepancy to transport samples from the biased LF posterior.","core_discovery":"Bridged SBI first uses inexpensive LF simulations to identify a coarse high-density parameter region, then learns a local residual bridge to transport LF posterior samples toward HF-consistent regions by correcting the LF-HF discrepancy; experiments on sim-to-sim particle-parameter calibration and real-to-sim calibration with real soil observations show that this produces more accurate and reliable HF posteriors than HF-only SBI or the Naive-MF baseline, especially under limited HF simulation costs.","pith_inferences":["The explicit discrepancy modeling may prove useful in other cascaded inference pipelines where cheap models produce shifted but still informative posteriors.","One could test whether the number of required high-fidelity simulations drops further if the bridge is parameterized with additional structure such as a Gaussian process residual.","The approach highlights that ignoring fidelity gaps in multi-fidelity chains can systematically degrade coverage even when the low-fidelity model is cheaper to run."],"forward_implications":["Bridged SBI alleviates the LF-induced posterior miscoverage that affects sequential multi-fidelity SBI without discrepancy correction.","The method yields more accurate HF posteriors than HF-only SBI when the budget for high-fidelity simulations is limited.","The same correction approach applies to both sim-to-sim particle-parameter calibration and real-to-sim calibration tasks."],"fun_headline_variants":["Bridged SBI corrects biased LF posteriors via residual bridge","Local residual bridge transports LF samples to HF regions","Bridged SBI fixes LF-HF discrepancy for reliable posteriors","Residual correction improves HF inference over naive multi-fidelity"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The low-fidelity posterior remains sufficiently informative about the target high-fidelity posterior that a local residual bridge learned from limited high-fidelity simulations can reliably transport samples without introducing new biases or coverage failures.","fun_headline_variants_meta":{"raw":{"variants":["Bridged SBI corrects biased LF posteriors via residual bridge","Local residual bridge transports LF samples to HF regions","Bridged SBI fixes LF-HF discrepancy for reliable posteriors","Residual correction improves HF inference over naive multi-fidelity"]},"model":"grok-4.3","cost_usd":0.003525,"raw_usage":{"total_tokens":1882,"prompt_tokens":729,"num_sources_used":0,"completion_tokens":63,"cost_in_usd_ticks":35249500,"prompt_tokens_details":{"text_tokens":729,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1090,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":729,"tokens_out":63,"duration_ms":7269,"temperature":1.0,"reasoning_tokens":1090,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T16:35:23.678507+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An experiment in which high-fidelity posterior samples from Bridged SBI produce predictive simulations that match held-out high-fidelity or real observations no better than the Naive-MF baseline would falsify the central claim.","supporting_citations":[],"review_version":1}