{"id":"ed790fcb-eeeb-43df-9c1f-2e1617f37d5d","arxiv_id":"2607.19190","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Agentic Real2Sim automates real-to-sim conversion of robot interaction episodes using vision-language agents, achieving a 48% VLM-judged replay success rate on DROID-100 with a 31B open-weight model.","lead":"This paper presents a pipeline that uses vision-language agents to automatically turn real video recordings of robots interacting with objects into physics simulations that replay the interaction. The system, Agentic Real2Sim, is tested on 100 DROID manipulation episodes plus deformable-object and humanoid motion cases, showing that an open-weight model can match proprietary models on a VLM-based success metric at a fraction of the cost.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"VLM judge metric measures visual plausibility, not physical fidelity; start-pose drift is explicitly unscored, so the central preservation claim is unsubstantiated.","rationale":"The reader's weakest assumption is that the VLM-based replay-success metric is a valid measure of physical real-world alignment. This is indeed the single most load-bearing concern: every headline number (48/100 successes, backend comparability, 31.4× cost reduction) is filtered through this metric. The Section 4.1 rubric is explicit about what it checks—target object identity, final location, action, gripper location—and explicitly does not subtract start-pose drift. Given that the pipeline's grasp optimization stage intentionally modifies object placement to achieve grasping, VLM-judged success does not imply that the simulated episode preserves the real object's state, trajectory, or physical parameters. A physical audit on the successful episodes would directly test whether the concern lands. Since this confirms the reader's conditional verdict rather than escalating it, no change is needed.","tokens_in":9183,"tokens_out":4983,"duration_ms":56099,"concrete_test":"Take the 48 episodes scored as replay successes by the Gemma 4 31B backend. Using the real episode's FoundationPose object tracks and the robot's recorded end-effector poses, compute (a) initial-object-pose drift introduced by grasp optimization, (b) final object position error between real and simulated replay, and (c) end-effector trajectory overlap. If more than ~25% of 'success' episodes have final object position error >5 cm or initial pose drift >5 cm, the VLM metric is not validating the 'preserves object states' claim; re-run the cost comparison under a physically validated success label.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claim—that Agentic Real2Sim produces physically aligned episodic twins and that an open 31B VLM matches proprietary backends—rests entirely on the Section 4.1 replay-success metric. That metric defines success as at least one of three VLM judges assigning a best-candidate score ≥8, where judges compare real and simulated keyframes on target-object identity, final object location, action similarity, and end-gripper location, and where 'starting-pose drift is recorded as context rather than directly subtracted.' This is a visual-plausibility score, not a physical-fidelity score. It does not measure object trajectories, contact timing, forces, or inferred physical parameters, despite the abstract's claim that the framework 'preserves observations, geometries, robot interactions, and object states.' Moreover, the pipeline's simulator-in-the-loop grasp optimization (§3.1) deliberately shifts the object's placement to enable grasping; because the metric tolerates start-pose drift, a successful episode can have a different initial object state from the real one. The 48/100 Gemma success count and the 31.4× cost comparison therefore do not demonstrate real-world-aligned twins unless the VLM score is shown to track physical correspondence. No human validation, physical ground-truth check, or error bars are reported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Agentic Real2Sim, a VLM-orchestrated pipeline that converts real-world recordings of robot-object interactions into simulatable MuJoCo episode twins. The system decomposes conversion into four stages—visual processing, physical-prior inference, scene preparation, and simulator-in-the-loop grasp optimization—organized under a shared 'episode contract.' It is evaluated on DROID-100 rigid manipulation episodes using a VLM-judged replay-success metric, across four VLM backends, with reported success counts from 37/100 to 48/100 and model costs ranging from $2.62 to $82.30. The authors also demonstrate qualitative deformable (PhysTwin-style) and humanoid (BFM-Zero-style) adapters. The central claims are that the framework preserves observations, geometries, robot interactions, and object states; that an open 31B VLM achieves comparable results to proprietary backends at up to 31.4× lower cost; and that the same contract generalizes across domains.","tokens_in":9495,"tokens_out":4926,"duration_ms":50649,"significance":"If the quantitative claims are substantiated, Agentic Real2Sim would be a meaningful step toward automated real-to-sim conversion: it provides a modular, backend-agnostic pipeline with explicit artifact contracts and bounded agentic decision-making, and it suggests that open-weight models can serve as practical orchestrators. The 31.4× cost reduction claim is striking and relevant. However, the evidence is currently insufficient to support the 'physical alignment' and 'preserving object states' claims, because the success metric is a VLM visual-plausibility score with no ground-truth validation. The contribution is therefore conditional; the framework design is promising, but the evaluation must be strengthened before the claims can be accepted.","major_comments":[{"comment":"The central quantitative evaluation rests entirely on a VLM-judged metric that scores real-vs-sim keyframes on target-object identity, final location, action similarity, and end-gripper location. Starting-pose drift is explicitly 'recorded as context rather than directly subtracted,' and §3.1's grasp sweep deliberately shifts object placement. A successful episode can therefore have a different initial object state from the real one. This is a visual-plausibility score, not a measure of physical fidelity: it does not verify trajectories, contacts, forces, or inferred physical parameters. As written, the abstract's claim that the framework 'preserves ... object states' and the conclusion's claim that 'physical parameters' are preserved are not supported by the reported 48/100 success count. The authors should validate the metric against human judgments or physical ground-truth (e.g., traj","section":"§4.1 (replay-success metric)"},{"comment":"Each VLM backend is run exactly once on DROID-100; the success counts (48, 45, 43, 37) are not accompanied by variance, confidence intervals, or statistical tests. The differences among backends are within what could be run-to-run noise, and the headline 31.4× cost ratio is a point estimate from a single run. The claim that 'an open 31B VLM achieves comparable observed replay-success outcomes' is therefore not statistically established. Repeating the pipeline with multiple seeds or reporting bootstrap intervals would be necessary.","section":"§4.2 and §4.3"},{"comment":"The deformable and humanoid results are qualitative only, and the limitations paragraph in §5 states that future work will 'extend automated conversion to deformable episodes and systematically evaluate agentic components.' This contradicts the abstract's statement that the framework is 'evaluated' across these domains. The generalization claim is thus an over-claim; at minimum, the abstract and §1 should be revised to say these are preliminary stress tests, or quantitative deformable/humanoid metrics should be added.","section":"§4.4 and §5"}],"minor_comments":[{"comment":"The identities of the three judge VLMs are not specified. If a judge backend coincides with a conversion backend, the evaluation would be less convincing; please state the judge models explicitly.","section":"§4.3"},{"comment":"The left panel could include error bars or multiple-run markers; the right panel's log scale makes the 31.4× difference visually large, but the interpretation should mention the single-run nature.","section":"Fig. 3"},{"comment":"The episode-twin tuple includes symbols (S_{1:T}, Θ, B, M) that are partly explained only later; consider defining them in a table or in the text immediately after the equation.","section":"Eq. (1)"},{"comment":"No direct quantitative comparison with prior real2sim pipelines (e.g., Scalable Real2Sim [15], TwinAligner [5]) is provided. Adding a comparison on a common benchmark would help contextualize the 'first step toward scalable conversion' claim.","section":"Related work"},{"comment":"The caption refers to 'Fig. 4 (a)' and '(b)' but the figure panels are labeled A and B; unify the notation.","section":"Fig. 4 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper has a very large author list and spans several institutions; that is not an issue per se. The main concern for the editor is that the paper's headline claims exceed the evidence. The VLM-judged metric is a reasonable first-pass metric but cannot alone support 'physical alignment.' I recommend a major revision with additional validation, and I would be happy to review a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a careful referee, but the claims are ahead of the evidence. The genuinely new piece is the agentic orchestration layer for episode-level real2sim: a set of bounded, schema-constrained VLM queries wrapped around deterministic perception and simulation tools, all feeding a common episode contract that spans rigid manipulation, deformable objects, and humanoid motion. That separation between agent decisions and deterministic tools is a good architectural idea, and the paper describes it clearly enough to rebuild. The cost comparison is transparent about model bills, and the qualitative panel includes failures, which is honest.\n\nNow the soft spots. The replay-success metric used for all quantitative claims is defined by three VLM judges comparing real vs. simulated keyframes, scoring target identity, final location, action, and gripper pose, with start-pose drift recorded as context rather than subtracted. The grasp-optimization stage deliberately shifts object placement to enable grasping. So a 'success' can have an initial object state different from the real episode, and the judge is rewarding visual plausibility, not physical correspondence. The paper's claim that the twins 'preserve observations, geometries, robot interactions, and object states' is not supported by that metric. There are no physical ground-truth checks, no human validation, no error bars, and no statistical tests; each backend runs once. There is also no comparison to prior real2sim pipelines, so the 37–48 success counts out of 100 are hard to interpret. The deformable and humanoid sections are explicitly qualitative, and the paper's own limitation statement concedes the focus is on rigid DROID episodes.\n\nThat said, these weaknesses are addressable rather than fatal. Adding per-episode physical checks (e.g., comparing tracked object trajectories, contact timing, and final poses), a small human study, baselines, and independent runs would substantially strengthen the claims. The stress-test note about the metric is fair; I read the paper with that concern and it holds.\n\nWho this is for: people building automated real2sim pipelines or using DROID-style data for simulation; also anyone thinking about VLM-based evaluation metrics. It deserves serious peer review. I would not rely on the quantitative claims until they are backed by physical validation, but the architecture is worth citing. If I had it on my reading list, I'd bring it up.","headline":"Promising agentic real2sim architecture with a clear tool/agent split, but the VLM-judged replay metric leaves the physical-fidelity claims unproven.","tokens_in":10044,"tokens_out":3271,"would_cite":true,"duration_ms":35165,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Agentic Real2Sim converts real robot-demonstration videos into physics-simulated digital twins automatically.","keywords":["real-to-sim","vision-language agents","robot manipulation","digital twin","physics simulation","agentic pipeline","deformable objects","humanoid motion"],"falsifier":"Take a set of converted episode twins that pass the VLM replay-success threshold, then directly compare the simulated object trajectories and contact events against tracked real-world poses; if a substantial fraction of 'successful' twins show large quantitative divergence (for example, object position errors exceeding several centimeters) despite high VLM scores, the metric's claim to physical alignment is refuted.","tokens_in":9093,"feed_emoji":"🤖","tokens_out":5356,"duration_ms":57809,"temperature":0.7,"pith_summary":"The paper's central claim is that real-world recordings of robot-object interaction—not just static scenes—can be automatically converted into runnable physical simulations, called episode twins, that preserve the actors, objects, trajectories, and physical parameters of the original episode. The framework achieves this with a pipeline of vision-language agents that orchestrate deterministic perception tools (segmentation, depth reconstruction, pose tracking) and iteratively refine the simulated scene by replaying it in a physics engine. Evaluated on 100 real manipulation episodes, the framework succeeds in roughly half of cases, and an open 31-billion-parameter vision-language model matches proprietary backends in conversion success while cutting model cost by up to 31×. The same 'episode contract' generalizes to deformable objects and humanoid motion, suggesting a unified path toward scalable real-to-sim conversion without manual scene authoring.","feed_headline":"Automated pipeline converts real robot demos into physics-simulated twins","feed_subtitle":"Four vision-language agents orchestrate perception and simulation, with an open 31B model matching far pricier backends.","key_machinery":"The central mechanism is the 'episode twin' contract—a structured representation that standardizes what a conversion must preserve across domains, together with a VLM-driven agentic loop that makes bounded, schema-constrained decisions (object discovery, keyframe selection, mask acceptance, refinement choices) while delegating geometry, physics, and rendering to deterministic tools. This separation lets the pipeline transfer across vision-language backends without per-model tuning.","core_discovery":"This paper proposes that the bottleneck in real-to-sim conversion is not missing perception models but the glue between them. It introduces Agentic Real2Sim, a framework whose four agents—visual processing, physical-prior inference, scene preparation, and simulator-in-the-loop grasp optimization—convert a recorded interaction into a simulatable 'episode twin' represented as a structured tuple (observations, actors, geometry, states, parameters, simulator, metrics). The framework is evaluated on 100 real manipulation episodes: an open 31B vision-language model produces 48 successful, 8 partial, and 44 failed replay outcomes, with other backends scoring 37–45 successes, and the open model cost","pith_inferences":["A direct extension is to replace the VLM judge with quantitative physical metrics—such as object-pose error or contact forces—to validate whether 'replay success' truly corresponds to physical alignment.","The same agentic orchestration pattern could transfer to other domains that combine perception with simulation, such as biomechanics or autonomous driving scene reconstruction.","The cost-efficiency result suggests that when tool boundaries are cleanly separated, small open-weight models can substitute for frontier models, a lesson likely relevant to other agentic pipelines beyond real-to-sim.","If replay success is driven by visual plausibility rather than dynamics, the twin may be insufficient for contact-rich policy learning; a targeted study comparing policy performance in real versus simulated episodes would clarify this."],"forward_implications":["If conversion can be automated at a cost of about $2.62 per 100 episodes with an open model, large-batch simulation assets for policy learning become practical.","The episode contract provides a common interface across rigid, deformable, and humanoid domains, reducing the need for separate real-to-sim pipelines.","Observed replay success near 50% across backends indicates the remaining bottleneck is upstream perception and simulation fidelity, not vision-language reasoning power.","Simulator-in-the-loop refinement (grasp optimization) shows that a physics engine can correct imperfect vision estimates during conversion."],"fun_headline_variants":["Four AI agents turn robot demos into physics twins","Open 31B model matches pricey backends for real2sim","Agentic Real2Sim: self-driving pipeline for robot simulation","Physics twins from real demos: cheap VLM does it all","Robot demos to sim: agentic pipeline cuts cost, keeps quality"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The replay-success metric assumes that a vision-language judge's score of 8 or higher out of 10 for real-versus-simulated keyframe similarity is a valid measure of genuine physical alignment between the simulation and the real episode.","fun_headline_variants_meta":{"raw":{"variants":["Four AI agents turn robot demos into physics twins","Open 31B model matches pricey backends for real2sim","Agentic Real2Sim: self-driving pipeline for robot simulation","Physics twins from real demos: cheap VLM does it all","Robot demos to sim: agentic pipeline cuts cost, keeps quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000267,"raw_usage":{"total_tokens":1461,"prompt_tokens":765,"completion_tokens":696,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":621}},"tokens_in":509,"tokens_out":696,"duration_ms":6156,"temperature":1.0,"reasoning_tokens":621,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T13:08:32.548275+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of converted episode twins that pass the VLM replay-success threshold, then directly compare the simulated object trajectories and contact events against tracked real-world poses; if a substantial fraction of 'successful' twins show large quantitative divergence (for example, object position errors exceeding several centimeters) despite high VLM scores, the metric's claim to physical alignment is refuted.","supporting_citations":[],"review_version":1}