{"id":"599d1fdd-05c8-4626-810f-ed72b9b7220f","arxiv_id":"2412.14417","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Constraining removal volume per action makes a geometric grinding simulator accurate enough that diffusion-planned cutting sequences transfer to a real robot without real-world training data.","lead":"A robot grinding planner trained only in simulation can transfer directly to a real robot, as long as each grinding action removes only a tiny amount of material. This matters because it could cut the costly real-world data collection normally needed to automate industrial shaping processes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The sim-to-real transfer mechanism is confounded: closed-loop replanning and a pre-calibration step, not the small-removal constraint, may explain the real-robot accuracy, and no real-world ablation varies ε_vol.","rationale":"The reader's weakest_assumption already identifies the small-removal-volume premise and the opaque calibration step. My pass agrees with that core concern but sharpens it: the real-robot evaluation also uses closed-loop replanning every M=32 steps, so even a poor geometric transition model could yield low final Chamfer error through repeated correction. That means Table V does not by itself isolate the proposed mechanism. The proposed ablation is feasible, uses the paper's own metrics and constraints, and would settle whether the mechanism is causal. I do not recommend rejection: the real-robot results are genuinely encouraging and the method is clearly described enough to test. However, absent this ablation or a clear disclosure of what the calibration step contains, the central sim-to-real claim remains conditionally supported rather than established. A secondary concern, not pursued here, is that the training-data description in §VI.C mentions only Shape A as the reference shape; if the same diffuser is used for Shapes B and C, the paper should explain how target-shape generalization is obtained. That issue is worth checking but is secondary to the mechanism question.","tokens_in":12184,"tokens_out":6076,"duration_ms":55422,"concrete_test":"Run a real-robot ablation on Shape A/ASA with all other components fixed: train separate CSD diffusers on simulated data collected with ε_vol ∈ {0.5, 1.0, 2.0, 4.0} (matching δ_vol and c_vol), execute with the same M=32 closed-loop schedule and calibration, n≥5 trials each, and report final Chamfer error plus mean per-step absolute deviation between observed and GCM-predicted removal. The small-volume mechanism predicts monotonically increasing error and deviation with ε_vol; flat curves would indicate that closed-loop replanning or calibration carries the transfer, and would invalidate the central claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on Eq. 8 and §IV.B: keeping per-step removal volume below ε_vol makes the geometric cutting model accurate enough that a policy trained only on simulated cuts transfers. Table V is consistent with this, but it does not demonstrate it. Deployment is not open-loop: §IV.A replans every M=32 steps from the current observed point cloud, so each segment can correct accumulated GCM error. §VII adds that 'shape coordinate error between simulation and the real world was calibrated in advance'; if that calibration absorbs geometry or resistance mismatch, the 'no real data' claim is overstated. No real-robot experiment varies ε_vol, and no measurement of actual per-step removal volume vs grinding resistance is reported. The causal claim that small removal volumes reduce the reality gap is therefore not isolated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Cutting Sequence Diffuser (CSD), a diffusion-based trajectory planner for robotic belt grinding. Training data are collected in a geometric cutting simulator under a small per-step removal volume constraint (Eq. 8), point-cloud shapes are compressed by a VAE, and an H-step action sequence is generated by guided diffusion with a two-step guidance procedure, executed closed-loop in M-step blocks. The method is evaluated in simulation and on a real grinding cell with ASA and PC materials and three target shapes, comparing shape error and planning time against Const-RS and CSA-MBRL. The main claim is that constraining per-step removal volume reduces grinding resistance and thereby makes a purely geometric simulation model sufficient for sim-to-real transfer without real-robot data.","tokens_in":12329,"tokens_out":5704,"duration_ms":49360,"significance":"If the central claim were established, the paper would make a useful practical contribution: sim-only training for an irreversible process in which real data collection is expensive, combined with fast long-horizon planning and demonstration on a real robot. The real-robot experiments across two materials and three shapes, the inclusion of CSA-MBRL (a real-data baseline) as a comparison, and the honest discussion of limitations are strengths. However, the mechanism behind the alleged sim-to-real transfer is not isolated by the reported experiments: closed-loop replanning and a pre-calibration step are confounded with the small-removal-volume constraint, no real-robot ablation varies the threshold epsilon_vol, and several load-bearing experimental details are missing. The paper is therefore currently a promising feasibility study rather than a fully supported demonstration of its stated causal claim.","major_comments":[{"comment":"The central causal claim—that the small-removal-volume constraint in Eq. (8a) is what makes the geometric cutting model accurate enough for zero-real-data transfer—is not isolated by the reported experiments. Deployment in §IV.A replans every M=32 steps from the current observed point cloud, so each closed-loop segment can correct accumulated geometric-model error regardless of epsilon_vol. Section VII adds that \"shape coordinate error between the simulation and the real world was calibrated in advance,\" but the content, magnitude, and necessity of this calibration are not disclosed. No real-robot experiment varies epsilon_vol, and no experiment compares open-loop (M=T) execution with closed-loop execution. Table V is therefore consistent with the claim but does not demonstrate it. I request a real-robot ablation varying epsilon_vol (for example, 0.25, 0.5, 1.0, 2.0) under otherwise identical replanning, an open-loop variant, and a full disclosure of the calibration procedure and its parameters.","section":"§IV.A, §IV.B, §VII, Table V"},{"comment":"The virtual grinding resistance model used in the simulation experiment is not specified quantitatively. The text states that shape deformation due to grinding resistance \"was reproduced as the deviation of the cutting surface\" and that the model depends on robot action and removal volume, but no equation, parameter values, or validation against real grinding data are given. Since the simulation results are used to support the sim-to-real transfer claim, the fidelity of this resistance model is load-bearing. Please provide the complete model, report how its parameters were chosen, and state whether any calibration against real measurements was performed.","section":"§VI.A"},{"comment":"Key hyperparameters and algorithmic details needed for reproducibility are missing. The cost weights lambda and lambda_k in §IV.D are never reported, nor is the cut-volume threshold delta_vol in Eq. (11); §VI.C gives epsilon_vol=1.0 and d_vol=0.3 but not delta_vol or the cost weights. In addition, the two-step guidance in §IV.D requires generating a trajectory over the full task horizon (T=160) at t=0, while the diffusion model is trained with H=32; the text does not explain how a model trained for length-32 sequences is used to generate a 160-step trajectory. Please report all hyperparameters in the paper or a stable appendix and clarify the full-horizon generation procedure.","section":"§IV.D and §VI.C"},{"comment":"The real-robot comparison is statistically thin: each condition in Table V has only three trials, and no significance tests are reported. With n=3, the claimed equivalence between CSD and CSA-MBRL is fragile (for example, Shape C with ASA: CSD 1.16±0.02 versus CSA-MBRL 1.52±0.20). Please provide per-trial data or additional trials, and report an appropriate statistical comparison or explicitly temper the equivalence claim.","section":"Table V and §VII.A"}],"minor_comments":[{"comment":"The action smoothness cost contains a subscript typo: in the condition for d_len, \"a^j_{j+1}\" should be \"a^j_{l+1}\", and the index j in the denominator should be l. Please correct this.","section":"Eq. (12)"},{"comment":"In the robot experiments, the over-cutting value for \"CSD w/ Guide\" is identical (4.00±1.63) for Shape A with ASA and Shape A with PC, despite different action-smoothness and cut-volume-limit values. This looks like a copy-paste error or a reporting inconsistency; please verify and correct the table.","section":"Table VI"},{"comment":"The notation pθ(τ_{i-1}|τ_i,O=1) in Eq. (6) is inconsistent with the unconditional pθ defined in Eqs. (3)–(4); the conditional distribution should be written with a separate symbol (for example, p~_θ) or the conditioning should be introduced explicitly.","section":"Eq. (6)"},{"comment":"CSA-MBRL uses T=30, H=5, M=1 in the real experiments and T=40 in simulation, while CSD uses T=160, H=32, M=32; the planning-time comparison in Table V should be accompanied by a clarification of whether replanning frequency and observation cadence are comparable, since these directly affect wall-clock time.","section":"§VII.A"},{"comment":"The state constraint cost is written as a Dirac penalty with -∞, but the implementation replaces the state with the constraint value; please clarify how the -∞ is handled numerically and how the replacement interacts with the diffusion denoising updates.","section":"§IV.D and Eq. (9)"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a relevant and under-studied problem, and the real-robot results are encouraging. My main concern is that the headline claim—that the small-removal-volume constraint itself enables sim-to-real transfer—is not isolated from closed-loop replanning and an undisclosed calibration step. This is fixable with additional experiments and fuller reporting, so I recommend major revision rather than rejection. The missing hyperparameters and the unclear full-horizon generation in the two-step guidance also need to be resolved for reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a decent engineering paper with a genuinely useful idea, and the real-robot results are better than I expected. The claim that small per-step removal volumes shrink the sim-to-real gap for grinding is plausible and follows from grinding force theory, but the current experiments do not isolate that mechanism. Still, it deserves a serious refereeing.\n\nWhat's new: applying diffusion-based long-horizon planning to a removal process, with an action space constrained to small removal volumes. That combination is not present in prior diffuser or MBRL work, and the two-step guide for handling the goal-conditioned long horizon is a practical fix. The simulation ablations for the guide functions and the two-step guidance are clean and show clear effects.\n\nWhat's good: the real-robot evaluation covers two materials and three shapes, with CSD matching or beating CSA-MBRL (which trains on real data) on shape error and doing it much faster in planning and total execution. That is a meaningful result if it holds. The paper is also honest about limitations in the discussion.\n\nSoft spots, in order of importance. First, the mechanism. The \"no real data\" claim is partly carried by closed-loop replanning every M=32 steps and an advance calibration of shape-coordinate error between sim and real. That is not necessarily a flaw—closed-loop execution is standard—but it means the small-removal-volume constraint is not the only thing bridging the gap. There is no real-robot ablation varying epsilon_vol, so the central causal claim is underdetermined. Second, the grinding resistance model is a proxy; the proportionality between removal volume and resistance is cited from theory but not measured for this robot, belt, and materials. Third, reporting is thin: cost weights and key thresholds (epsilon_vol, delta_vol, lambda) are omitted from the main text; they are promised on a project page, but the paper should stand alone. Fourth, only three trials per real condition, and the CSA-MBRL comparison has different horizons and observation schedules, making the timing comparison largely unfair.\n\nNone of this kills the paper. The central idea is sound and the results are encouraging. But a referee should push for a direct test of the removal-volume hypothesis on the real robot—vary epsilon_vol with everything else fixed—and for the missing hyperparameters and calibration details.\n\nWho this is for: people working on sim-to-real transfer, contact-rich manipulation, or industrial grinding. I would bring it to a reading group if anyone in our group does manipulation. Recommendation: send it to peer review. It is not a desk reject.","headline":"Good engineering paper with a clean idea and real-robot results, but the current experiments do not isolate the small-removal-volume mechanism that is supposed to close the sim-to-real gap.","tokens_in":12856,"tokens_out":2280,"would_cite":true,"duration_ms":18097,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a grinding robot can learn to shape objects from pure geometric simulation, provided each cut removes only a small volume, and that a diffusion model can plan those long cutting sequences quickly enough for real use.","keywords":["grinding","sim-to-real transfer","diffusion model","object shaping","removal process","geometric cutting model","robot manipulation","motion planning"],"falsifier":"A reader could test the central claim by running the same CSD policies on a material or belt condition where measured single-pass grinding force is not linear in removal volume (for example, a worn belt or a much harder workpiece) and checking whether shape error degrades relative to CSA-MBRL. A second test is to repeat the real robot experiment with an uncalibrated sim-to-real coordinate frame; if shape error rises sharply, the calibration step, rather than the small-volume constraint, is carrying the transfer.","tokens_in":11986,"feed_emoji":"⚙️","tokens_out":5865,"duration_ms":46730,"temperature":0.7,"pith_summary":"This paper argues that the main obstacle to sim-to-real transfer in robotic grinding is grinding resistance, which grows with the volume removed in a single pass. If each grinding action removes only a small amount of material, the paper claims, the real shape change is well approximated by a purely geometric cutting model, so a diffusion-based planner trained entirely in that simplified simulation can be transferred to a real robot without real-world data. The proposed Cutting Sequence Diffuser learns from random action sequences collected under a small-removal-volume constraint and generates long-horizon cutting sequences with cost-guided denoising and a two-step planning guide. In both simulation and real robot experiments, it reaches shape errors comparable to a model-based reinforcement learning baseline that learns from real data, while planning much faster. The central idea is that careful action-space design, not a more realistic simulator, is what closes the reality gap.","feed_headline":"Small grinding bites let a simulation-trained robot shape real parts","feed_subtitle":"Diffusion planner trained only on geometric cutting matches a real-data method's accuracy in less planning time.","key_machinery":"The load-bearing object is the constrained action space defined by Eq. (8). Every sampled action must remove no more than $\\epsilon_{\\rm vol}$ volume per step and must not overcut the reference shape by more than $d_{\\rm vol}$, which keeps each transition within the regime where grinding resistance is small enough for the geometric cutting model to hold. On top of this, a diffusion model is trained over latent shape features obtained from a PointNet-based VAE, and during deployment it plans with two-step guidance: a full-task-horizon trajectory is generated first to supply intermediate shape targets, then shorter-horizon guided denoising replans against those targets, preventing large-volume compensating cuts.","core_discovery":"The central claim is that constraining each grinding action to a small removal volume, as in Eq. (8a) with $f_{\\rm col}(s_t,a_t)\\le\\epsilon_{\\rm vol}$, makes the real grinding process behave like the Geometric-Cutting Model used in simulation. Under this constraint, the shape transition is approximately a geometric splitting by a cutting surface, so the simulator's algebraic operations predict real shape changes well enough that a diffusion model trained only on simulated data can plan actions for a real robot. The paper reports that CSD achieves final shape error comparable to CSA-MBRL, which learns from real data, across two materials and three target shapes, while reducing action planning time to tens of seconds and lowering over-cutting and per-step volume violations when guided by the designed cost functions.","pith_inferences":["If the volume-proportional resistance assumption is what makes this work, the same small-step constraint could extend to other irreversible removal processes such as milling or filing, where a geometric cutting approximation is plausible.","The paper discloses that shape coordinate differences between simulation and reality were calibrated in advance; a natural follow-up is to measure how much of the transfer accuracy survives without that calibration.","A sharper test of the theory would be to estimate $\\epsilon_{\\rm vol}$ from a one-step real grinding force measurement and compare shape error across materials, rather than fixing the threshold once.","For worn belts or very hard workpieces, the linear resistance law likely degrades, so re-tuning the cut-volume-limit cost or adding force feedback would become necessary."],"forward_implications":["A grinding robot can be set up for a new material or target shape using only simulation data, avoiding irreversible and costly real-world data collection.","The diffusion planner's long-horizon sequences reduce the need for frequent shape observations, cutting total task execution time compared with the real-data baseline.","Cost-guided denoising measurably improves over-cutting, action smoothness, and per-step volume limit violations over unguided generation.","The two-step guide suppresses the large single-step cuts that would otherwise appear when the planning horizon is shorter than the task horizon."],"supporting_citations":[{"why":"Supplies the grinding-resistance theory that resistance is proportional to removal volume, the load-bearing premise for the small-volume action constraint.","marker":"[3]"},{"why":"Provides the CSA-MBRL baseline and the cutting-surface shape transition functions $\\Psi_s, \\Psi_w$ that CSD builds on and compares against.","marker":"[2]"},{"why":"Introduces the diffusion planning framework with cost-guided denoising that CSD uses to generate long-horizon trajectories.","marker":"[8]"},{"why":"Demonstrates long-horizon robot motion planning with diffusion models, informing the trajectory representation and training procedure.","marker":"[7]"},{"why":"Defines the denoising diffusion formulation and simplified noise-prediction loss used to train the diffuser.","marker":"[5]"},{"why":"Supplies the PointNet architecture used in the VAE to compress point-cloud shape states into latent features.","marker":"[24]"},{"why":"Provides the Chamfer discrepancy metric used to measure shape error in all simulation and robot experiments.","marker":"[25]"}],"fun_headline_variants":["Small bites let sim-trained diffusion planner grind real shapes","Diffusion planner trained on sim, constrained bites, grinds real parts","Constrained grinding steps make sim-to-real shaping work","Sim-only training with small removals beats real-data planning time","Tiny grinding steps shrink sim-to-real gap for robot shaping"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that grinding resistance is proportional to removal volume and that keeping each step's removal volume below a threshold makes the geometric cutting simulation accurate enough for the real robot, belt, and materials; the paper also notes that shape coordinate differences between simulation and reality were calibrated in advance, and that calibration is not specified.","fun_headline_variants_meta":{"raw":{"variants":["Small bites let sim-trained diffusion planner grind real shapes","Diffusion planner trained on sim, constrained bites, grinds real parts","Constrained grinding steps make sim-to-real shaping work","Sim-only training with small removals beats real-data planning time","Tiny grinding steps shrink sim-to-real gap for robot shaping"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1308,"prompt_tokens":918,"completion_tokens":390,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":305}},"tokens_in":534,"tokens_out":390,"duration_ms":3268,"temperature":1.0,"reasoning_tokens":305,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:15:14.252020+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could test the central claim by running the same CSD policies on a material or belt condition where measured single-pass grinding force is not linear in removal volume (for example, a worn belt or a much harder workpiece) and checking whether shape error degrades relative to CSA-MBRL. A second test is to repeat the real robot experiment with an uncalibrated sim-to-real coordinate frame; if shape error rises sharply, the calibration step, rather than the small-volume constraint, is carrying the transfer.","supporting_citations":[{"cited_title":"Modeling and experimental study of grinding forces in surface grinding,","cited_arxiv_id":null,"evidence_quote":"Supplies the grinding-resistance theory that resistance is proportional to removal volume, the load-bearing premise for the small-volume action constraint."},{"cited_title":"Learning to shape by grinding: Cutting-surface-aware model-based reinforcement learning,","cited_arxiv_id":null,"evidence_quote":"Provides the CSA-MBRL baseline and the cutting-surface shape transition functions $\\Psi_s, \\Psi_w$ that CSD builds on and compares against."},{"cited_title":"Planning with diffusion for flexible behavior synthesis,","cited_arxiv_id":null,"evidence_quote":"Introduces the diffusion planning framework with cost-guided denoising that CSD uses to generate long-horizon trajectories."},{"cited_title":"Motion plan- ning diffusion: Learning and planning of robot motions with diffusion models,","cited_arxiv_id":null,"evidence_quote":"Demonstrates long-horizon robot motion planning with diffusion models, informing the trajectory representation and training procedure."},{"cited_title":"Denoising diffusion probabilistic mod- els,","cited_arxiv_id":null,"evidence_quote":"Defines the denoising diffusion formulation and simplified noise-prediction loss used to train the diffuser."},{"cited_title":"Pointnet: Deep learning on point sets for 3d classification and segmentation,","cited_arxiv_id":null,"evidence_quote":"Supplies the PointNet architecture used in the VAE to compress point-cloud shape states into latent features."},{"cited_title":"Point-set distances for learning representations of 3d point clouds,","cited_arxiv_id":null,"evidence_quote":"Provides the Chamfer discrepancy metric used to measure shape error in all simulation and robot experiments."}],"review_version":1}