{"id":"b88af247-fa62-43df-92a5-60ac8b4c8e8d","arxiv_id":"2607.23473","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Explicit low-degree factorized polynomial proprioceptive features improve robot RL and imitation policies beyond matched-capacity MLPs and induce sensorless compliance-like contact behavior in simulation.","lead":"PRISM adds a compact factorized polynomial layer so robot policies can use products of joint, velocity, and command signals, not just linear mixes. In simulation it beats capacity-matched MLPs on humanoid walking and contact-rich manipulation and shows force-free compliance-like contact behavior.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The headline LIBERO gain (63.8%→91.0%) may largely be an artifact of an undertrained, early-checkpoint Diffusion Policy baseline; on the stronger SmolVLA backbone the identical change yields only +3.05 points on a single seed.","rationale":"In good faith: the paper's goal is to show that exposing factorized polynomial interactions among deployment-available observations helps motor policies beyond raw capacity, and the Humanoid-Gym evidence (param-matched control, 5 seeds, degree ablation in Table 18, linear probes in Table 4) genuinely supports that narrow claim. The reader correctly flagged low-degree sufficiency and the capacity-control stand-in, and I partially agree — but I judge the more load-bearing issue to be the evidentiary basis of the specific number the strongest claim quotes: 63.8%→91.0% on LIBERO. That number comes from the weakest experimental cell in the paper (20K-step task-specific DP, 10 rollouts/task), while the paper's own strongest-backbone experiment shows an order-of-magnitude smaller effect on one seed. This is not circularity or unsoundness; it is correctness risk on the magnitude of the headline effect. The reader's verdict of CONDITIONAL already anticipates this ('very large LIBERO delta', 'independent reproduction'), so I do not move the verdict — I sharpen the reproduction condition: reproduction must control for baseline convergence and seed variance, and the compliance narrative should be re-evaluated without success-conditioned diagnostics. If that control is run and the gap holds, the paper's central claim stands; if it collapses, the architectural-universality language ('standard architectural choice') needs to be scaled back to the locomotion result.","tokens_in":14674,"tokens_out":2797,"duration_ms":113878,"concrete_test":"Retrain the per-task LIBERO Diffusion Policy baseline to convergence under a standard published protocol (e.g., ≥100K steps or the original epoch budget, not 20K), then train PRISM with identical budget and evaluate both at 50 rollouts/task over ≥3 seeds. If the success gap collapses from +27 to low single digits (mirroring Table 14's +3.05), the headline LIBERO number is a baseline-strength artifact; if a double-digit gap persists at converged baselines, the concern is answered.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim leans on two empirical pillars: Humanoid-Gym (Table 1) and the LIBERO jump from 63.8% to 91.0% (Table 2/13). The Humanoid-Gym pillar is reasonably solid — 5 seeds, parameter-matched Larger MLP control, consistent direction across four metrics. The LIBERO pillar is much softer than the headline suggests. Per Appendix A.2 (Table 9), the task-specific Diffusion Policy experiment trains only 20K steps and evaluates 10 rollouts per task. The resulting baseline (Table 13: Object 36.0%, Avg 63.8%) sits far below published Diffusion Policy results on LIBERO under standard training budgets, which is the classic signature of an undertrained baseline. A representation change that acts partly as a better-conditioned/regularized proprioceptive pathway will show its largest apparent gains exactly when the baseline is far from convergence. The paper's own stronger-backbone check supports this reading: with SmolVLA (Table 14), the same PRISM swap yields only +3.05 points (63.50→66.55), on a single seed (1000), with per-suite movements of mixed sign relative to the larger-capacity control. If the converged-baseline gap is really ~3 points rather than ~27, then the 91.0% figure — and the 'sensorless compliance' narrative illustrated by one rollout (Fig. 4) and contact diagnostics that are explicitly success-conditioned (Appendix A.4, which biases force comparisons toward the higher-success method) — overstates the interaction-structure effect. The internal claim 'not capacity alone' survives via Table 1; the external 'standard architectural choice' claim, which the abstract hangs on the LIBERO delta, is the part at risk.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper proposes PRISM, a factorized elementwise polynomial module that exposes degree-K interactions among deployment-available proprioceptive variables while preserving the policy objective, action interface, and low-level controller. In PPO-based Humanoid-Gym locomotion, PRISM reportedly outperforms both a standard MLP and a parameter-matched larger MLP; in LIBERO imitation learning, replacing Diffusion Policy's linear proprioceptive conditioner with a degree-2 PRISM layer raises reported average success from 63.8% to 91.0%. Additional experiments test BFM-Zero and SmolVLA, ablate polynomial degree, probe linear recoverability of physical quantities, evaluate evaluation-time perturbations, and include a small SO-101 pilot. The authors also interpret lower contact-force behavior and factor ablations as evidence of sensorless, compliance-like interaction regulation.","tokens_in":15116,"tokens_out":2685,"duration_ms":28870,"significance":"If the results hold, the paper provides a useful and easily adopted architectural inductive bias: a compact, learnable interaction basis that requires no force, tactile, contact-label, or privileged deployment input. Particularly valuable are the parameter-count controls, degree ablation, evaluation on two architecturally stronger backbones, simulator-only post-hoc physical probes, and the preliminary real-robot pilot. The Humanoid-Gym comparison is promising evidence that interaction structure is not interchangeable with width. Establishing the LIBERO result at converged training budgets and the compliance result without success-conditioned selection would make the proposed standard architectural choice substantially more credible.","major_comments":[{"comment":"Tables 2, 9, and 13: the headline Diffusion Policy result (63.8% to 91.0%) is obtained after only 20K updates and 10 rollouts per task. No training-length comparison or convergence evidence is provided, and the companion experiment on the stronger SmolVLA backbone finds only +3.05 points at an aligned 80K checkpoint (Table 14), using seed 1000. Consequently, the present evidence cannot distinguish a genuine interaction-structure gain at convergence from a large short-training/regularization effect on an undertrained baseline. Because Table 2 anchors the manipulation and sensorless-compliance claims, the paper should report longer or converged task-specific Diffusion Policy training, ideally with learning curves and multiple seeds, and reconcile the result with the much smaller SmolVLA gain.","section":"§4.2 / Appendix A.2 / Tables 2, 9, 13–14"},{"comment":"Table 1 and Table 8 give inconsistent parameter counts for what appear to be the same Humanoid-Gym actors: Table 1 reports 0.926M for MLP and 1.321M for both Larger MLP and PRISM, while Table 8 reports 527K, 922K, and 922K, respectively, explicitly labeled as actor parameters. Since the capacity-control conclusion—that interaction structure rather than parameter count drives the gain—depends on exact matching, the authors must explain what additional modules are included in Table 1 or correct one table.","section":"§4.2 / Appendix A.1 / Tables 1 and 8"},{"comment":"The default degree and locomotion numbers are not internally consistent. §4.2 states that degree 2 is the default and degree 3 performs best, while Table 1's PRISM row reports episode length 2233.4 and survival 92.50. Under the stated same Humanoid-Gym setup, Table 18 reports PRISM-D2 as 2198.7/91.00 and PRISM-D3 as 2296.5/92.50. Table 1 therefore matches neither row exactly. Please identify which degree, checkpoint-selection rule, and seed protocol generated Table 1, and make the main result and degree ablation directly comparable.","section":"§4.2 / Appendix C.1 / Tables 1 and 18"},{"comment":"Appendix A.4 states that the contact diagnostics are success-conditioned, averaging only successful episodes. Since PRISM has a much higher success rate than Diffusion Policy and MCC-Sensorless, this selection rule can mechanically bias contact-force comparisons toward the method that succeeds more often; Figure 4 further relies on one representative rollout. The sensorless-compliance claim should instead be supported by contact-aligned force, impulse, power, and approach-speed distributions over all rollouts, or by comparisons stratified by success/failure and task, with confidence intervals.","section":"§4.2 / Figure 4 / Appendix A.4"},{"comment":"Several key comparative tables do not quantify uncertainty or replication. Table 1 is described as averaged over five seeds but reports no standard errors; Tables 2 and 13 use only 10 rollouts per task; Table 14 reports one seed for SmolVLA; and Table 15 uses 20 rollouts per perturbation condition. For the paper's central cross-domain claim, report seed/task-level confidence intervals or significance tests and state the number of independently trained policies for every entry.","section":"§4 / Tables 1–3 and 13–15"}],"minor_comments":[{"comment":"§1 and Fig. 2: the introduction says PRISM is applied 'after' the MLP backbone in RL, whereas §3.2–3.3 and Fig. 2 describe polynomial conditioning of the proprioceptive input before the downstream actor. Please use one consistent description and block diagram.","section":"§1 / §3.2–3.3 / Fig. 2"},{"comment":"Table 2 calls MCC-Oracle a non-deployable force-access ablation that provides an upper bound, but its measured success (64.5%) is below PRISM and only slightly above Diffusion Policy. 'Force-access ablation' is accurate; 'upper bound' is not established by the result.","section":"§4.2 / Table 2"},{"comment":"Table 4 supports improved linear recoverability, but the title's phrase 'Newtonian quantities' is stronger than the operational targets shown; slip, contact impulse, and contact work are simulator-computed proxies. Consider 'mechanics-inspired quantities,' as used in the caption.","section":"§4.3 / Table 4"},{"comment":"Table 5's post-hoc factor names are useful, but the mapping from two affine branches' largest weights to a named dominant input product is qualitative. Please give the exact ranking/threshold rule and, if possible, show the leading weight products for at least one factor.","section":"§4.3 / Table 5"},{"comment":"The SO-101 pilot is appropriately framed as preliminary, but Table 17 should include trial-level confidence intervals or otherwise avoid visual emphasis given only five demonstrations and ten trials per method.","section":"Appendix B.5 / Table 17"},{"comment":"Several typographical and formatting issues should be corrected, including 'Our approaches quickly before contact,' the missing space in the project URL sentence, and the inconsistent multiplication notation in Fig. 1 and Eq. (2).","section":"§4.2 and throughout"}],"recommendation":"major_revision","confidential_remarks":"The work is a good topical fit, but the manuscript's broad recommendation that polynomial representations become a standard architectural choice currently rests on a mix of strong locomotion evidence and substantially weaker short-budget, low-replication manipulation evidence. The editorial decision should therefore hinge less on novelty and more on whether the authors can reconcile the reported configurations and provide converged, replicated LIBERO/compliance results."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: put a compact factorized degree-2 (or 3) interaction layer on deployment-available proprio/command history, and you get consistent sim gains that a wider MLP does not buy you. That is the actual contribution—not a new theory of polynomials.\n\nWhat is new is the packaging and the controls. They keep the rest of PPO / Diffusion Policy / SmolVLA / BFM-Zero fixed, match parameter count with a larger MLP (Tables 1 and 3), ablate degree (Table 18), and show the same direction on two stronger backbones. Humanoid-Gym is the cleaner pillar: five seeds, episode length and survival roughly double, tracking error halves, larger MLP does almost nothing. Linear probes and factor ablations are honest analysis, not training labels. Citations to multiplicative/polynomial nets are in place; they do not overclaim invention of the math.\n\nSoft spots, in proportion. The 63.8%→91.0% LIBERO delta is almost certainly inflated by a weak baseline. Appendix training is 20K steps, 10 rollouts/task; Object at 36% is the classic undertrained signature. Their own SmolVLA check is only +3 points on one seed, mixed vs the larger conditioner. Fig. 4 and success-conditioned contact diagnostics then over-sell “sensorless compliance.” The real-robot pilot is five demos / ten trials—fine as a pilot, not a pillar. Limitations section already admits low-degree and missing geometry/material cases; the abstract’s “standard architectural choice” language still outruns the evidence.\n\nWho it is for: people who ship proprio-conditioned loco/manip policies and want a cheap inductive bias before adding force hardware. Math is elementary and fine; data protocol is reproducible in principle if they release the exact checkpoints.\n\nI would send it to referees. Ask them to demand converged Diffusion baselines, multi-seed SmolVLA, and toned universality claims. Worth engaging; do not treat 91% as the effect size.","headline":"Solid robotics packaging of factorized poly proprio conditioning with real capacity controls; the big LIBERO jump is the soft pillar, not the Humanoid-Gym one.","tokens_in":15943,"tokens_out":505,"would_cite":true,"duration_ms":9056,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Making products among robot sensors explicit beats simply making the policy network bigger.","keywords":["polynomial interaction","motor control","sensorless compliance","proprioceptive conditioning","humanoid locomotion","contact-rich manipulation","diffusion policy","embodied AI"],"falsifier":"Train the same locomotion and LIBERO setups with a width-matched MLP versus PRISM under identical data, rewards, and controllers: if the larger MLP matches or beats PRISM on episode length, tracking error, survival, and success, and contact-force traces show no sensorless yielding advantage, the claim that interaction structure cannot be replaced by capacity collapses.","tokens_in":15705,"feed_emoji":"🤖","tokens_out":972,"duration_ms":21858,"temperature":0.7,"pith_summary":"Robot policies usually feed joint angles, velocities, and commands into a plain multilayer network and hope it invents the right couplings. Many quantities that actually matter for control—power, slip, contact impulse, compliance—are products of those signals, not the signals alone. This paper argues that those interactions should be built into the policy as a compact, learnable polynomial layer rather than left for the network to discover. The method, PRISM, forms a few latent factors and multiplies them element-wise so higher-order structure is available without listing every monomial. On humanoid walking and contact-rich arm tasks it outperforms both standard and larger-capacity networks under the same sensors and controllers, and it produces yielding contact behavior without force or tactile hardware. The practical claim is that interaction structure is not interchangeable with parameter count, so polynomial proprioceptive conditioning should be a default architectural choice in embodied motor control.","feed_headline":"Polynomial sensor products beat bigger robot policy nets","feed_subtitle":"A compact interaction layer lifts walking and contact tasks without force sensors or extra capacity tricks","key_machinery":"PRISM: a factorized polynomial proprioceptive module. Two (or recursively K) learned affine maps of the deployable state form latent factors; element-wise products with a learned, near-zero-initialized gate yield ψ_K, which keeps a first-order path and adds compact higher-order interaction features before the unchanged downstream actor or diffusion conditioner.","core_discovery":"Exposing low-degree factorized polynomial interactions among deployment-available physical observations improves motor policies beyond what matched-capacity multilayer networks achieve, and can induce sensorless compliance-like contact regulation without force, wrench, tactile input, contact labels, or an admittance controller. On humanoid locomotion and contact-rich manipulation, that architectural change—not extra width alone—raises tracking, survival, and task success while making mechanics-inspired quantities more linearly readable from the policy representation.","pith_inferences":["If low-degree products of proprioception are this useful, many existing sim-to-real stacks may be leaving easy structure on the table simply by keeping linear state encoders.","A natural next stress test is tasks where failure is dominated by unobserved geometry or materials: PRISM should plateau while sensor-augmented baselines keep improving.","Adaptive or input-dependent polynomial degree could matter when some joints need quadratic coupling and others stay nearly linear.","The sensorless compliance result suggests re-examining ‘force-free’ industrial insertion and wiping stacks that currently bolt on admittance only because the policy never saw multiplicative state structure."],"forward_implications":["Default robot actors and visuomotor conditioners should expose factorized polynomial proprioceptive interactions rather than only linear or MLP state embeddings.","Matched-capacity ablations become necessary when claiming an architectural inductive bias in motor control; width alone is not a substitute for an interaction basis.","Contact-rich policies can seek compliance-like force regulation from proprioceptive products without adding force/tactile sensors or admittance loops at deployment.","Linear probes on frozen policy features become a diagnostic for whether mechanics-inspired quantities (slip, power, impulse, work) are more accessible after the representation change.","The same module can drop into both PPO-style locomotion actors and diffusion or vision-language-action proprioceptive pathways without changing actions or low-level controllers."],"fun_headline_variants":["PRISM factorized polynomials beat matched-capacity MLP robot policies","Explicit sensor products lift locomotion and contact success","Low-degree interaction layer yields sensorless compliant contact","Polynomial proprioception outperforms wider nets on humanoid tasks","Factorized poly features make mechanics cues linearly readable"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The paper assumes that the cues that matter for these tasks live mainly in low-degree products of signals the robot already has, so a compact polynomial layer is enough without force sensing, long history, or hidden contact geometry.","fun_headline_variants_meta":{"raw":{"variants":["PRISM factorized polynomials beat matched-capacity MLP robot policies","Explicit sensor products lift locomotion and contact success","Low-degree interaction layer yields sensorless compliant contact","Polynomial proprioception outperforms wider nets on humanoid tasks","Factorized poly features make mechanics cues linearly readable"]},"model":"grok-4.5","effort":"low","cost_usd":0.003287,"raw_usage":{"total_tokens":1137,"prompt_tokens":775,"num_sources_used":0,"completion_tokens":58,"cost_in_usd_ticks":32868000,"prompt_tokens_details":{"text_tokens":775,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":304,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":775,"tokens_out":58,"duration_ms":6776,"temperature":1.0,"reasoning_tokens":304,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T21:24:18.853037+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train the same locomotion and LIBERO setups with a width-matched MLP versus PRISM under identical data, rewards, and controllers: if the larger MLP matches or beats PRISM on episode length, tracking error, survival, and success, and contact-force traces show no sensorless yielding advantage, the claim that interaction structure cannot be replaced by capacity collapses.","supporting_citations":[],"review_version":1}