{"id":"42173a01-48c6-4413-b052-03c4a3244d7f","arxiv_id":"2607.14906","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Averaging per-agent conformal e-values with a per-neighborhood miscoverage budget restores the target coverage α in fused multi-robot occupancy maps under local stationarity and mixing assumptions.","lead":"This paper presents a distributed fusion rule that restores a chosen reliability level when multiple robots merge occupancy maps, even though each robot's own local map is only weakly guaranteed. Robots exchange only small summary scores, and simulations show the merged map stays within its error budget while denser communication graphs classify more of the environment.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4.1's central premise—that the held-out fold yields a finite-sample upper bound ε_i^k on local miscoverage (Assumption 3)—is asserted, not proven; if ε_i^k underestimates, the fused e-value is invalid and the coverage guarantee fails.","rationale":"The paper's main theorem is a clean e-value averaging argument: given a valid local e-value, weighted averaging preserves validity and Markov yields coverage. The algorithmic components (budget α/d, gate, attenuation) are coherent, and the simulations are extensive and consistent with the theorem. The single most load-bearing weakness is that the local e-values are only valid if Assumption 3 holds, and the only justification for Assumption 3 is Remark 2's assertion that a held-out empirical miscoverage estimate is an upper bound. This is a genuine finite-sample gap: no concentration argument is supplied, and the stated guarantee in the abstract is therefore stronger than what is proven. The concern is not that the theorem is false under its assumptions, but that the assumptions are not established. This matches the reader's weakest_assumption exactly. The recommended verdict remains CONDITIONAL: the authors should prove a conservative ε_i^k (e.g., using a conformal quantile or binomial upper bound on the held-out fold) or explicitly rephrase the guarantee as conditional on validated local calibration, and report repeated-run variance. No change to the reader's verdict is needed; our independent read reinforces it.","tokens_in":16964,"tokens_out":6741,"duration_ms":62499,"concrete_test":"Use the authors' simulation setup with known ground truth. For each of, say, 100 independent trials, compute a_i^k and ε_i^k exactly as in Remark 2 (held-out fold: calibration half of next block), then estimate the true miscoverage P(p_i^k(X*_k,Y*_k) ≤ a_i^k) on a large fresh test set from the same block regime. Count the fraction of trials in which ε_i^k < true miscoverage. If this fraction is not zero (or is above a small tolerance), Assumption 3 is not delivered by the held-out estimation procedure, and the theorem's premise must be amended to a conditional guarantee or a conservative upper bound (e.g., a binomial confidence bound) substituted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The fused coverage guarantee in Theorem 4.1 is logically correct only if Assumption 3 holds: ε_i^k in (12) must upper-bound P(p_i^k(X*_k,Y*_k) ≤ a_i^k) for every active agent. The authors' only support is Remark 2, which states that a_i^k and ε_i^k are computed on a held-out fold and that this 'provides' the bound 'under the temporal protocol.' No concentration or finite-sample guarantee is given. A single empirical miscoverage estimate from a finite fold is itself a random variable and is not stochastically guaranteed to dominate the true miscoverage; it can be below it with non-negligible probability, especially when β_i = α/d_i is small (e.g. 0.04 for α=0.2, d=5) and violations are rare. If ε_i^k underestimates, then E[e_j^k] > 1, so e_fused is not an e-value and the Markov step in Theorem 4.1 fails. This is not a peripheral issue: the theorem is a conditional statement whose premise is never established. Proposition 2's Assumption 4 is likewise a modeling assumption. The abstract's unconditional phrasing ('regardless of ... sensor noise distribution') overstates the result. The simulations consistently show high coverage, but with a safety factor s=0.7 and λ=0 they do not isolate whether the assumption itself is valid.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses multi-robot occupancy mapping with finite-sample coverage guarantees. Each robot trains a local GP, computes conformal p-values, recalibrates an inner threshold to a per-neighborhood miscoverage budget, and constructs e-values that are gated by local observability and attenuated by predictive uncertainty. Agents broadcast only scalar e-values over a shared query set; each receiver averages them and forms a set-valued prediction. Theorem 4.1 claims that the fused map achieves the target coverage 1−α at every agent, for any communication graph and sensor-noise distribution, under Assumption 3 (a local coverage bound). Proposition 2 extends the result to the case where local coverage degrades, using an uncertainty-graded bound. Simulations on ring and mesh topologies report empirical coverage above the nominal level and show that denser graphs reduce the unclassified fraction.","tokens_in":17294,"tokens_out":6896,"duration_ms":65466,"significance":"The e-value fusion mechanism is elegant and practically attractive: it requires only scalar communication, is agnostic to the local likelihood-map estimator, and the proof of Theorem 4.1 is clean and correct conditional on its premise. Proposition 1 is a useful observation that, under a shared measurement kernel, joint total-variation closeness reduces to spatial-marginal closeness. If the local coverage premise can be certified, the paper provides a valuable reduction: recovering a global coverage guarantee from per-agent calibrated e-values. The simulation study is transparent and demonstrates the efficiency benefits of denser topologies. However, the central theoretical claim is only as strong as Assumption 3, and that assumption is asserted rather than derived; the abstract's unconditional phrasing overstates what is actually proven.","major_comments":[{"comment":"The load-bearing premise of Theorem 4.1 is asserted, not established. In Eq. (11), ε_i^k is the miscoverage measured on a held-out fold at the recalibrated threshold a_i^k, and Remark 2 claims that this 'provides' a finite-sample upper bound on the true miscoverage P(p_i^k(X*_k,Y*_k)≤a_i^k). A single empirical miscoverage estimate from a finite fold is a random variable; it is not stochastically guaranteed to dominate the true value, especially when the budget β_i=α/d_i is small (e.g., 0.04 for α=0.2, d_i=5) and miscoverage events are rare. If ε_i^k underestimates the true miscoverage, then E[e_j^k]>1, so the fused e-value is not a valid e-value and the Markov step in Theorem 4.1 fails. The citation to [22],[23] does not supply the required bound: those works give guarantees under weighted exchangeability or spatial-block conditions, not a guarantee that a held-out empirical CDF dominate","section":"Section IV, Assumption 3, Eq. (18); Remark 2"},{"comment":"The 'recovery' result does not remove the unverified-premise problem; it only changes its form. Proposition 2 assumes c_i^k(x) ≤ ε_i^0 exp(μ_i σ̂_i^k(x)) for active x, with unknown constants ε_i^0 and μ_i, and requires λ ≥ max_j μ_j. No estimation procedure for μ_i is given, and unlike ε_i^k in the main theorem there is not even a suggested held-out estimator. Thus the claim that attenuation 'restores the target level when a local guarantee fails' is contingent on an assumption that is at least as hard to certify as Assumption 3. If this proposition is intended to make the framework applicable when local bounds are loose, it needs a concrete verification protocol for μ_i or a conservative choice that preserves finite-sample validity.","section":"Section IV-B, Proposition 2, Assumption 4"}],"minor_comments":[{"comment":"The limitations paragraph says 'temporal mixing (Assumption 1)' but should refer to Assumption 2 (β-mixing). The same sentence duplicates 'Assumption 1' twice.","section":"Section VI"},{"comment":"Line 4 of Algorithm 1 sets β←α/d, while the simulation uses β=sα/d with safety factor s=0.7. The safety factor appears only in Section V-A; please state in Algorithm 1 that β may incorporate a safety factor, or make the two consistent.","section":"Section III-B, Algorithm 1"},{"comment":"The sentence 'The above assumption is in principle the same as the results in [22], [23]' is vague. Please cite the specific theorem in those references and explain how it yields the finite-sample upper bound needed in (18).","section":"Section IV, after Eq. (18)"},{"comment":"The text defines an 'empty fraction' (Γ=∅) as the only genuine coverage failure, but Table I reports only coverage, classified fraction, and no-data fraction. Reporting the empty fraction would make the connection between coverage and the set-valued output more direct.","section":"Section V-E, Table I"}],"recommendation":"major_revision","confidential_remarks":"The paper's core fusion argument is sound, but the main theorem rests on a premise that is effectively the target property in disguise. The authors should either rigorously derive a conservative finite-sample ε_i^k from a held-out fold (e.g., via concentration with a confidence level) and state the coverage result accordingly, or transparently reframe the contribution as a reduction from global coverage to individually certified local coverage. The simulation study is useful but does not stress-test the problematic regime where the held-out fold is small or the budget is very small; a synthetic experiment designed to break Assumption 3 would help calibrate how much safety margin is needed in practice."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a worthwhile paper with a genuinely useful algorithmic idea, but the headline guarantee is conditional on an unproven premise. If you read Theorem 4.1 as a conditional statement — “provided each local e-value is valid” — the proof is clean and the method is sound. If you read the abstract as claiming an unconditional, distribution-free coverage guarantee, it overreaches.\n\nWhat’s actually new: the fusion rule is a nice design. The per-neighborhood budget alpha/d, the observation gate, and the uncertainty attenuation inside the e-value are all sensible, and the budget trick that lets a single confident agent decide a query is clever. Proposition 1 is a neat, correct result: under a shared measurement kernel, total-variation closeness of the joint law reduces to closeness of the spatial marginal. Theorem 4.1’s proof is straightforward and correct given Assumption 3, and the validity–efficiency dial with lambda is a nice practical insight. The paper is also honest about its limitations: it clearly states the marginal nature of the guarantee, the reliance on stationarity/mixing, and the communication cost.\n\nThe soft spot is exactly where the stress-test note points. Assumption 3 is the load-bearing premise, and it is asserted, not derived. Remark 2 says the held-out fold “provides” the upper bound on local miscoverage, but a single empirical miscoverage estimate from a finite fold is a random variable and is not guaranteed to dominate the true miscoverage. If epsilon_i^k comes in too low, then E[e_j^k] > 1, the fused e-value is not valid, and the Markov step in Theorem 4.1 fails. The gap is especially real when beta_i = alpha/d_i is small and violations are rare, because the empirical miscoverage on a finite fold can easily be zero while the true miscoverage is nonzero. The cited references [22], [23] give TV-based bounds, but not the tight pointwise upper bound used here. This is not a peripheral detail; it is the bridge between the local calibration and the fused guarantee.\n\nThe simulations do not close the gap. They use K = 2 temporal blocks, no error bars across repeated runs, and a safety factor s = 0.7 with lambda = 0 at the operating point, so they demonstrate that the method is conservative, not that Assumption 3 holds at the nominal threshold. The empirical coverage staying near 0.98 is consistent with the theorem but does not test the premise.\n\nWho is this for? Researchers in multi-robot mapping, conformal prediction, and e-values will find the fusion construction and the conditional analysis useful. It deserves a serious referee, and I would send it to peer review. The authors should either prove a finite-sample conservative epsilon_i^k (for example, via a holdout conformal quantile or a concentration argument that accounts for the temporal dependence), or explicitly state the guarantee as conditional on a validated local calibration and provide an empirical check of that calibration. I’d also ask for more than two blocks and repeated-run variance. With those additions, this becomes a solid contribution; as written, it is a good idea with an unresolved foundational step.","headline":"A clean fusion method with a correct conditional proof, but the headline coverage guarantee depends on an unproven local-calibration premise that needs to be either proven or transparently conditional.","tokens_in":17828,"tokens_out":2219,"would_cite":true,"duration_ms":23311,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A distributed fusion rule over scalar e-values recovers the user-specified conformal coverage level at every robot, regardless of communication graph or sensor noise, even when each local guarantee is degraded.","keywords":["conformal prediction","e-values","multi-robot occupancy mapping","distributed fusion","coverage guarantee","Gaussian process occupancy mapping","uncertainty quantification"],"falsifier":"Conduct the multi-agent mapping experiment on the same benchmark but with the held-out fold drawn from a shifted regime with larger total-variation drift δ_T (e.g., a more abrupt change in occupancy statistics between blocks), and check whether the empirical fused coverage drops below 1−α; alternatively, compute for a fixed block the empirical local miscoverage on infinitely many independent test folds and compare it to the single-fold ε_i^k to see if ε under-estimates with non-negligible probability.","tokens_in":16755,"feed_emoji":"🗺️","tokens_out":3778,"duration_ms":32259,"temperature":0.7,"pith_summary":"The paper addresses a gap in multi-robot occupancy mapping: local conformal prediction sets on each robot's map have coverage guarantees that degrade because trajectories break exchangeability and each robot sees only part of the environment. The authors propose to recover the target coverage 1−α by fusing lightweight scalar e-values over a communication neighborhood, using a per-neighborhood miscoverage budget α/d and an uncertainty-attenuated threshold. They prove (Theorem 4.1) that the fused set-valued map achieves the target coverage at every agent for any graph and any noise distribution, under a local-coverage assumption. They also show a second result: a decay-based attenuation recovers coverage even when the local bound fails, under an uncertainty-graded miscoverage model. Simulations confirm the bound and show that denser topologies shrink the fraction of unclassified cells.","feed_headline":"Fused e-values recover conformal coverage in multi-robot maps","feed_subtitle":"Robots exchange only scalars; the fused set-valued map meets the target level for any graph and noise.","key_machinery":"The load-bearing object is the lifted e-value (12): each agent converts its conformal p-value at a query into a scalar e-value by dividing the violation indicator 1{p ≤ a} by a normalizer ε that upper-bounds the local miscoverage, multiplying by an observation gate (1 if within radius r of the agent's training inputs) and an exponential attenuation exp(−λσ̂) in the GP predictive standard deviation. E-values are closed under convex averaging, so the unweighted average over the neighborhood is an e-value, and Markov's inequality converts E[e]≤1 into P(Y* ∉ Γ) ≤ α. The per-neighborhood budget β=α/d ensures a single confident active agent can decide a query on its own, while the attenuation is t","core_discovery":"The paper's central claim is that the target miscoverage level α can be restored at every agent by a single round of e-value broadcast and weighted-average fusion, even though each local conformal guarantee may be degraded by temporal correlation and partial spatial coverage. The construction lifts each agent's conformal p-value into an e-value normalized by a per-neighborhood budget β=α/d, gated by the agent's observation region, and attenuated by the GP predictive standard deviation; because each local e-value has expectation at most one under the null, any fixed weighted average is also an e-value, and Markov's inequality yields the fused coverage guarantee. The coverage statement holds f","pith_inferences":["The reliance on a held-out fold for ε_k is the empirical pivot: if that estimate under-estimates true local miscoverage, the fused e-value is invalid and the theorem collapses; a concentration guarantee for ε would harden the result.","The observation gate restricts each agent's contribution to its observed region; one could test data-independent gating functions (e.g., GP variance) against the current radius-based gate, since Proposition 1 says any function of x alone controls the quantity Assumption 3 constrains.","The method's abstention behavior (returning {−1,+1}) could be used as an exploration heuristic: regions where fused evidence is insufficient are exactly where new measurements are most valuable.","Because the fused guarantee is marginal, per-cell conditional coverage is not claimed; a natural extension is to seek localized budgets that adapt to each query's neighborhood, or to combine with a second layer of conformal calibration on the fused scores."],"forward_implications":["Multi-robot teams can certify map predictions at a user-chosen level without exchanging raw data, only O(|Q|) scalars per block.","The coverage guarantee is independent of graph topology: ring and mesh both meet the bound; denser graphs mainly improve the classified fraction and reduce no-data abstentions.","The method applies to any likelihood-map estimator (e.g., OGM, GPOM, Hilbert maps) because it operates on the induced set-valued map, not the underlying regression.","The validity–efficiency dial λ lets a system designer choose where to sit between high coverage with large abstention and more decisive but less conservative maps.","If the local-coverage assumption fails but the uncertainty-graded model holds, attenuation can still restore the target level."],"fun_headline_variants":["Fused e-values restore conformal coverage in multi-robot maps","Scalar fusion recovers mapping coverage under degraded guarantees","Robot teams recover guaranteed map coverage with e-value fusion","Coverage recovery via e-value fusion despite broken exchangeability","Lightweight scalar fusion meets target map coverage for any graph"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central premise is that the normalizer ε_i^k, estimated on a single held-out fold, is a genuine upper bound on the true local miscoverage at the recalibrated threshold; the paper does not prove a concentration bound for this estimate, so if the fold underestimates the true miscoverage, the e-value is invalid and the fused coverage guarantee collapses.","fun_headline_variants_meta":{"raw":{"variants":["Fused e-values restore conformal coverage in multi-robot maps","Scalar fusion recovers mapping coverage under degraded guarantees","Robot teams recover guaranteed map coverage with e-value fusion","Coverage recovery via e-value fusion despite broken exchangeability","Lightweight scalar fusion meets target map coverage for any graph"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000159,"raw_usage":{"total_tokens":1086,"prompt_tokens":785,"completion_tokens":301,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":220}},"tokens_in":529,"tokens_out":301,"duration_ms":3406,"temperature":1.0,"reasoning_tokens":220,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T00:44:29.110542+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct the multi-agent mapping experiment on the same benchmark but with the held-out fold drawn from a shifted regime with larger total-variation drift δ_T (e.g., a more abrupt change in occupancy statistics between blocks), and check whether the empirical fused coverage drops below 1−α; alternatively, compute for a fixed block the empirical local miscoverage on infinitely many independent test folds and compare it to the single-fold ε_i^k to see if ε under-estimates with non-negligible probability.","supporting_citations":[],"review_version":1}