{"id":"7f066d42-da66-427d-b54f-56fd82657635","arxiv_id":"2605.30239","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"The paper's stated CA-World counterfactual claim is absent from the body, which instead describes the SAM3D-Phys pipeline for multi-object interactive reconstruction and simulation.","lead":"The arXiv abstract and title describe CA-World, a counterfactual-alignment framework for efficient interaction-ready 3D reconstruction, but the full text is a different paper, SAM3D-Phys, a training-free pipeline combining PGSR reconstruction, SAM3D generative object completion, and MPM simulation. Because the central counterfactual claim is entirely absent from the body, the preprint is internally incoherent and cannot be evaluated as submitted.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract promises CA-World counterfactual alignment; full text is SAM3D-Phys with no counterfactual objectives, inverse intervention, or CA-World equations, so the central claim is absent and unsupported.","rationale":"Reader's REJECT verdict is correct. The single load-bearing concern is that the abstract's central contribution—CA-World's counterfactual alignment learning with three consistency objectives—has zero implementation in the body. The full text is SAM3D-Phys, a related but different pipeline. While SAM3D-Phys includes spatial alignment and appearance distillation, those are engineering losses (Eqs. 5-8, VGG distillation), not realizations of 'reversing the intervention should recover the factual world.' The strongest evidence is the textual mismatch: Section 3.2 framework overview never uses the words counterfactual or intervention; the limitations section does not discuss the absent formalization; and the project page URL changes from ca-world to sam3d-phys. Because the central claim cannot be verified, correctness risk is high. A mechanical test—searching for the key terms and enumerating whether any equation defines the three objectives—will settle the concern. If it landed, the paper must be rejected or substantially restructured. I agree with the reader's weakest assumption as well: even if one tried to map SAM3D-Phys onto counterfactual language, the invertibility of the SAM3D completion/inpainting is never formalized; the body instead assumes completed geometry is correct and aligns it to masks. Both issues point to the same conclusion. No ad hominem: the authors may have intended a companion paper; as submitted, the central claim is unsupported.","tokens_in":13943,"tokens_out":4174,"duration_ms":38818,"concrete_test":"Perform a structural consistency audit: (1) extract every occurrence of 'CA-World', 'counterfactual', and 'intervention' from the PDF and list the section/equation context; (2) list all equations in the main text and identify any that instantiate the three claimed alignment objectives as functions of reintegrated vs. original scenes; (3) check whether the two project-page URLs refer to the same method. If the only 'CA-World'/'counterfactual' hits are in the title/abstract/project page and no equation in Secs. 3.2-3.4 corresponds to appearance/spatial/physical counterfactual consistency, the central claim is not implemented; the submission should be returned for correction or resubmission.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, as stated in the abstract, is that interaction-ready reconstruction is formulated as counterfactual alignment learning: foreground/background decoupling as a visual intervention, object generation/inpainting as counterfactual generation, reintegration as inverse intervention, and three alignment objectives (appearance, spatial, physical) enforcing counterfactual consistency. None of this appears in the full text. The body is a different paper titled 'SAM3D-Phys: Towards Multi-Object Interactive Simulation in Real World.' Its Section 3.2 describes a four-stage pipeline (PGSR reconstruction, SAM3D object generation, object-scene alignment, MPM simulation) with no mention of CA-World, counterfactual states, or inverse intervention. The two 'alignment' mechanisms in Secs. 3.3-3.4 are a render-and-compare pose optimizer and a mask-guided VGG distillation loss; they are not derived from or motivated by counterfactual consistency, and no equation defines an objective comparing the reintegrated scene to the original under an inverse intervention. The abstract's project page URL (ca-world) also conflicts with the body's project page (sam3d-phys). The Limitations section (Fig.9) acknowledges only appearance/context failures, not this missing formalization. Thus the central claim is unsupported by definition; this is not a matter of 'outside current consensus' but of the submitted text not containing the proposed method.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The abstract of arXiv:2605.30239 introduces CA-World, an 'efficient framework that integrates counterfactual alignment learning into a decoupling-reintegration reconstruction pipeline,' with three alignment objectives (appearance, spatial, physical) that enforce counterfactual consistency by requiring that reversing an intervention recovers the factual world. The full text, however, is a different paper titled 'SAM3D-Phys: Towards Multi-Object Interactive Simulation in Real World.' This body presents a four-stage, training-free pipeline: PGSR scene reconstruction, SAM3D object generation, object–scene alignment (physics-constrained pose optimization and mask-guided appearance distillation), and MPM-based multi-object simulation. There is no mention of CA-World, counterfactual states, inverse interventions, or any objective that compares the reintegrated scene to the original under a counterfactual intervention. The experimental section evaluates SAM3D-Phys on six real-world scenes against Feature Splatting and DecoupledGaussian, with a 12-participant user study and LMM-as-judge scores. As submitted, the paper's stated central contribution is entirely absent from the body.","tokens_in":14321,"tokens_out":4322,"duration_ms":44249,"significance":"If the CA-World counterfactual-alignment formulation were actually developed and experimentally validated, it could offer a new way to train interaction-ready reconstruction with direct supervision from counterfactual consistency, potentially improving efficiency by avoiding joint optimization of all object states. However, the manuscript as submitted does not deliver this: the abstract claims a method that the full text does not describe. The body's SAM3D-Phys pipeline is a plausible engineering contribution—combining a generative 3D prior with reconstruction and physics-based constraints—but it is a different contribution, and its own evidence (six scenes, two baselines, no error bars, small user study) is thin. The mismatch between the abstract and body prevents assessment of the claimed novelty and precludes acceptance. The manuscript contains machine-checkable equations only for standard 3DGS, MPM, and basic penalty constraints; the central counterfactual consistency objectives have no formalization.","major_comments":[{"comment":"The abstract claims CA-World counterfactual alignment learning with three objectives between reintegrated and original scenes, and formulates foreground-background decoupling as a visual intervention, object generation/inpainting as counterfactual generation, and reintegration as an inverse intervention. The full text contains none of this. Section 3.2 presents a pipeline (PGSR, SAM3D, alignment, MPM) with no mention of counterfactual states; Sections 3.3–3.4 give render-and-compare pose optimization (Eq. 5), physics penalty constraints (Eqs. 6–8), and a mask-guided VGG distillation loss, none of which are derived from or compared against a counterfactual 'inverse intervention.' No equation, algorithm, or experiment implements the abstract's central claim. The Limitations (Fig. 9) only discuss appearance/context failures, not this missing formalization. This is a load-bearing absence: th","section":"Abstract and Section 3 (whole body)"},{"comment":"The abstract's formulation is circular if interpreted literally. It states that 'reversing the intervention should recover the factual world' and motivates 'three alignment objectives between the reintegrated and original scenes.' If the reintegrated scene is directly supervised against the original observed scene, then the 'counterfactual state' is matched to the factual observation it is supposed to counterfactually predict—this is not causal consistency but regression to the factual scene. Since the body does not formalize these objectives, the claim is untestable as stated. The reader is left with either a circular training target or an undefined one; neither supports the claimed novelty.","section":"Abstract (counterfactual consistency formulation)"},{"comment":"Even for the body's more modest SAM3D-Phys contribution, the quantitative evidence is thin. The evaluation uses six scenes (two from DecoupledGaussian, four self-captured), two baselines, and a 12-participant user study with no reported variance, inter-rater reliability, or significance tests. Table 1 reports single ablation numbers for edge error, PSNR, and SSIM with no error bars. Baselines O2-Recon and Amodal3R are excluded with the justification that performance on real scenes is poor—this is not a substitute for a comparative experiment. These limitations undermine the claim of 'demonstrate effectiveness' (Section 4.2) for the body's method.","section":"Section 4 (Experiments)"}],"minor_comments":[{"comment":"The notation in Eq. (5) is tangled: subscripts and superscripts on t and R are difficult to parse, and the use of '⊙' for rotation composition is nonstandard. Please clarify.","section":"Eq. (5)"},{"comment":"The abstract gives the project page as https://chnxindong.github.io/ca-world/, while the body (title and footer) gives https://chnxindong.github.io/sam3d-phys/. These should match.","section":"Abstract vs. body project page"},{"comment":"The caption contains 'massgase gun,' which appears to be a typo for 'massage gun.'","section":"Fig. 8 caption"},{"comment":"O2-Recon and Amodal3R are cited twice with different reference numbers (13/14 and 40/41 respectively); please use a single reference per work.","section":"References"},{"comment":"Equation (7) uses a minimum distance margin 'm,' while the supplementary text says the tolerance can be defined as zero and uses 'epsilon.' This symbol inconsistency should be resolved.","section":"Eq. (7) and Supplementary Equation (7)"},{"comment":"The user study reports only aggregate scores. Please report per-participant variance or inter-rater agreement to support the claim that the method 'significantly outperforms' baselines.","section":"Table 2 (User study)"}],"recommendation":"reject","confidential_remarks":"The abstract and full text are so mismatched that this appears to be a submission integrity issue rather than a normal scientific disagreement. The editor may wish to verify whether the correct abstract or correct PDF was submitted; as it stands, the manuscript cannot be reviewed on its stated contribution because that contribution is absent from the body."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline is the mismatch: the abstract and title promise CA-World, a counterfactual alignment framework with three consistency objectives between reintegrated and original scenes, and the full text is a different paper called SAM3D-Phys with no counterfactual machinery, no inverse intervention, and no CA-World equations. The submitted text reads like two papers stapled together, and the abstract is the one doing the advertising.\n\nWhat the body does do is assemble a specific and plausible pipeline: PGSR reconstruction with segmentation-affinity features, SAM2 masks, LaMa inpainting, SAM3D for object completion, render-and-compare pose refinement, VLM-derived physical constraints, mask-guided VGG appearance distillation, and MPM simulation. That combination is not present in the cited baselines, and the qualitative results show it can separate multiple objects that Feature Splatting and DecoupledGaussian visibly fail on. The user study and LMM-as-Judge numbers, though based on only two baselines and a dozen participants, at least directionally support the claim that the pipeline helps.\n\nThe problems are in proportion. The central abstract claim is unsupported by definition: none of the central terms appear in the body, and the alignment losses in Sections 3.3–3.4 are conventional render-and-compare and physical penalties, not counterfactual consistency objectives. Even judged as SAM3D-Phys, the evidence is thin: six scenes, two baselines, no error bars, no code or data. The \"training-free\" wording is overstated because stage A fine-tunes the PGSR reconstruction and the appearance distillation optimizes object Gaussians; only the generative model itself is untouched. The limitations section acknowledges only appearance and context failures, not the missing formalization. The conflicting project pages (ca-world vs sam3d-phys) reinforce that this is not a wording slip.\n\nThe body could become a respectable workshop-level pipeline paper if reframed honestly and evaluated more rigorously, but as submitted it is not the CA-World paper and should not consume referee time in this form. I would desk reject, but explicitly invite the authors to resubmit as SAM3D-Phys with an accurate abstract and stronger evidence.","headline":"The abstract sells a counterfactual alignment framework that doesn't exist in the body; what's actually there is a decent but under-validated pipeline paper under a mismatched title.","tokens_in":14800,"tokens_out":2513,"would_cite":false,"duration_ms":26167,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that interaction-ready 3D reconstruction can be driven by counterfactual consistency: removing objects and restoring them should recover the original scene, and enforcing that match yields appearance, spatial, and physical","keywords":["interaction-ready reconstruction","counterfactual alignment","3D reconstruction","physics simulation","generative 3D priors","object-scene alignment","multi-object interaction","material point method"],"falsifier":"Set up a scene with a known object, remove the object, inpaint the background, and reintegrate the exact original object at its exact pose. If the three consistency losses (appearance, spatial, physical) do not drop near zero—for example because the background inpainting changes lighting or shadows—then counterfactual consistency cannot serve as a faithful supervision signal for interaction readiness.","tokens_in":13821,"feed_emoji":"🕹️","tokens_out":9168,"duration_ms":85141,"temperature":0.7,"pith_summary":"The paper argues that a reconstructed 3D world is only truly ready for interaction if it anticipates change—objects can be moved, removed, and collided with. To build such worlds efficiently, it frames foreground-background decoupling as a visual intervention: removing objects and inpainting the background is a counterfactual edit, and reintegrating the completed objects is the inverse edit. Counterfactual consistency then says that reversing the intervention should return to the observed scene, which yields three supervision signals: appearance, spatial, and physical consistency. This lets each object be checked against the observed scene locally, avoiding joint optimization of all object states. The paper also presents a training-free implementation of this idea and tests it on real multi-object scenes.","feed_headline":"A scene edit that undoes itself yields interactive 3D worlds","feed_subtitle":"The paper frames object removal and reintegration as counterfactual pairs, using three consistency losses.","key_machinery":"The central mechanism is counterfactual consistency over a decoupling-reintegration cycle. Foreground-background separation is treated as a visual intervention; separate object generation and background inpainting are treated as counterfactual generation; and scene reintegration is treated as the inverse intervention. The paper claims that enforcing equality between reintegrated and original scenes in appearance, spatial relationship, and physical plausibility is a sufficient supervision signal, and that the locality of object-level interventions lets each object be aligned independently using the observed scene, without joint optimization. In the implementation, this appears as three post-p","core_discovery":"On the paper's own terms, the central claim is that interaction-ready reconstruction reduces to counterfactual alignment learning. The authors define object removal and background inpainting as a counterfactual intervention, object completion as counterfactual generation, and scene reintegration as the inverse intervention. The requirement that the reintegrated scene match the original in appearance, spatial layout, and physical plausibility provides direct supervision without ground-truth object states. Locality of object-level interventions means each counterfactual state is constrained by the observed scene rather than by jointly optimizing all objects, cutting cost and error accumulation","pith_inferences":["The abstract promises counterfactual alignment learning, but the full-text implementation is a training-free pipeline that uses a frozen generative prior and optimization-based post-processing; a direct test of the counterfactual claim would be to train the completion model itself with the three consistency losses and measure improvement.","The counterfactual framing naturally suggests an automatic evaluation protocol: perform an intervention (remove and reintegrate an object), then measure the three consistencies; scenes that score well should be the ones that simulate correctly. This could become a standard benchmark for interaction-ready reconstruction.","The locality argument implies that the method's benefit grows with scene clutter; in sparse scenes the counterfactual supervision may provide little signal, so the approach may be best suited for dense multi-object environments.","If the invertibility assumption is violated (e.g., shadows, specular highlights, or inpainting ambiguity), the consistency loss could misattribute generative ambiguity to causal error; extending the framework to model these effects is a natural next step."],"forward_implications":["If counterfactual consistency is a valid supervisory signal, interaction-ready 3D reconstruction can be trained from ordinary multi-view images, since the original scene serves as its own ground truth.","Locality of object-level interventions implies reconstruction cost scales with the number of objects rather than jointly with all object states, reducing computational load and error accumulation in cluttered scenes.","The three consistency objectives (appearance, spatial, physical) provide a concrete checklist for whether a reconstructed scene will behave correctly under interaction—not merely render well.","The training-free implementation can run on consumer hardware and support real-time multi-object physical simulation, making interactive reconstruction practical outside research labs.","Clean object-scene separation plus completion enables dynamic scene editing, such as removing, replacing, or re-simulating individual objects."],"fun_headline_variants":["Object edits as counterfactual pairs yield interactive 3D worlds","Counterfactual alignment trains interaction-ready 3D scenes","Undo-and-redo object edits to build interactive worlds","Scene reintegration as inverse intervention for 3D"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that removing and reintegrating an object is an invertible intervention—that any mismatch between the reintegrated scene and the original genuinely reflects incorrect completion or alignment, rather than generative ambiguity, lighting changes, or inpainting artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Object edits as counterfactual pairs yield interactive 3D worlds","Counterfactual alignment trains interaction-ready 3D scenes","Undo-and-redo object edits to build interactive worlds","Scene reintegration as inverse intervention for 3D"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00058,"raw_usage":{"total_tokens":2581,"prompt_tokens":770,"completion_tokens":1811,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":1742}},"tokens_in":514,"tokens_out":1811,"duration_ms":14790,"temperature":1.0,"reasoning_tokens":1742,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T04:54:47.544160+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Set up a scene with a known object, remove the object, inpaint the background, and reintegrate the exact original object at its exact pose. If the three consistency losses (appearance, spatial, physical) do not drop near zero—for example because the background inpainting changes lighting or shadows—then counterfactual consistency cannot serve as a faithful supervision signal for interaction readiness.","supporting_citations":[],"review_version":2}