{"id":"aa47d92f-46f4-4ee5-9321-c9012109c2f0","arxiv_id":"2607.28405","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Aligning PTQ decisions to WAM structure, closed-loop rollouts, and the joint video–action objective yields W4A4 policies within 0.2–0.7 pp of FP16 on simulation benchmarks with ~29% block memory.","lead":"QuantWAMs is a post-training quantization recipe for World Action Models that keeps closed-loop robot success nearly at full precision under mostly 4-bit weights and activations. It matters because WAMs are expensive to run on robots, and ordinary LLM/DiT quantizers break when error feeds back through the control loop.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection beyond the reader’s already-flagged finite-sample surrogate limitation; central in-distribution claim holds.","rationale":"The reader’s strongest claim matches what Tables 1–2 and the measurement caveats actually support, and the weakest assumption (32-traj local surrogates vs. compounding closed-loop return) is the correct soft underbelly—explicitly flagged in §3.3 Eq. 21 and Limitations. Ablations (Tables 3–5) and pooling diagnostics (Fig. 3, Prop. 1) give coherent component-level support inside the QuantWAMs pipeline; experimental hygiene (disjoint roles, matched nominal budgets, paired seeds) is above average for PTQ systems work. Residual gaps—no code, block-level-only mem/speedup, underpowered 10-trial robot study, asterisk baselines that match counts not information—are grounds for CONDITIONAL, not for rejecting the reported means. Stress-testing does not surface a sharper internal break (e.g., an invalid pooling admissibility step or a leakage path from Dtest into fitting). Verdict stays CONDITIONAL at high confidence; no upgrade or downgrade.","tokens_in":15458,"tokens_out":636,"duration_ms":66801,"concrete_test":"Freeze masks, layer bit-allocations, and repaired step indices fitted on the standard 32+32 sets; evaluate closed-loop success on a task hold-out or a randomization tier materially stronger than Dcal (e.g., RoboTwin unseen task identities or held-out visual domain). If QuantWAMs–FP16 gap widens beyond ~5 pp while FP16 stays flat, the finite-sample local surrogates are not deployment-stable and the headline near-parity does not travel.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is an in-distribution empirical report (Tables 1–2): under the disclosed W4A4-dominant schedule, simulation means sit 0.2–0.7 pp from FP16. That report is directly backed by seed-paired, trajectory-disjoint cal/profile/val/test splits and large episode counts. The load-bearing soft spot is the same one the reader names—precision decisions are locked by local surrogates on only 32 cal trajectories plus 32 FP16 rollouts (shared-basis energy Top-K under the Prop. 1 working model, diagonal joint-Fisher layer scores, one-step unprotected replay in §3.3), while closed-loop error compounds through transition Jacobians the method never scores. Matched-budget Atom*/SVDQuant* control bit counts and protected-step cardinalities, not schedule indices or the co-training backward pass, so they understate how much of the gap is privileged calibration information rather than quantizer structure alone. Within the paper’s stated scope (benchmark-specific, in-distribution, block-level efficiency; Limitations explicit), this weakens transfer/deployment rhetoric more than the reported means themselves. No internal contradiction in the three strategies or the tables.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"QuantWAMs is a post-training quantization framework for World Action Models (WAMs), which jointly denoise future video and actions in closed loop. The authors argue that standard PTQ fails because it uses open-loop objectives, homogeneous-module assumptions, and calibration states that do not match deployment. They organize decisions around a calibration context (structural scope, rollout distribution, objective) and propose three components: (i) shared-basis outlier calibration that pools squared-energy channel statistics only across coordinate-compatible modules, with a random-effects crossover N★ guiding when pooling helps; (ii) co-training-objective saliency that builds a joint video–action empirical Fisher and upgrades the top 20% of candidate Linears by layer-level score; (iii) fixed-intervention rollout auditing that reassigns a fixed K protected denoising steps using one-step unprotected replay on FP16 closed-loop states. On Fast-WAM and LingBot-VA, under a W4A4-dominant schedule, simulation success means on RoboTwin 2.0 and LIBERO differ from FP16 by 0.2–0.7 pp, with ~29% of FP16 peak weight-and-activation memory and 1.4–1.6× block-level speedups on targeted blocks. Real-robot trials on an AgiBot G2 show feasibility on three tasks.","tokens_in":15882,"tokens_out":1480,"duration_ms":29052,"significance":"If the empirical claims hold, this is a useful systems contribution for deploying multi-stream diffusion robot policies under tight memory/latency budgets. The paper’s main conceptual value is making calibration context (coordinate admissibility, closed-loop state distribution, joint objective) explicit rather than treating WAMs as generic transformers. Strengths include trajectory-disjoint cal/profile/val/test roles, three protocol seeds with large episode counts, matched-budget Atom*/SVDQuant* controls, cumulative and factor ablations on LIBERO-Long, an explicit pooling crossover analysis (Prop. 1, Fig. 3), and appropriately cautious language on real-robot non-inferiority and block-level (not end-to-end) efficiency. The work is primarily empirical/engineering rather than a new theoretical quantization guarantee, but it is well scoped for the embodied-AI systems audience.","major_comments":[{"comment":"§4.2 and Tables 1–2: Matched-budget Atom* and SVDQuant* equalize candidate modules, W8 count budget, outlier fraction, and protected-step cardinalities, but not the co-training backward pass or Dprof/Dval schedule audit. The large gap to QuantWAMs therefore conflates (a) better quantizer structure with (b) privileged calibration signals unavailable to the baselines. A load-bearing control would give baselines the same joint-gradient saliency and/or fixed-intervention schedule indices (or ablate those signals off QuantWAMs while keeping its quantizer). Without that, the claim that the three strategies—not extra information—drive near-FP16 success is only partially supported.","section":"§4.2, Tables 1–2"},{"comment":"§3.3 Eqs. (18)–(21) and Limitations: Precision decisions are locked by local surrogates on 32 cal trajectories and 32 FP16 rollouts (energy Top-K, diagonal joint Fisher, one-step unprotected ℓt), while the paper correctly notes that closed-loop impact depends on products of transition Jacobians Aj, Bj that are never scored. Tables 1–2 establish strong in-distribution means, but the deployment rhetoric (abstract; §5) should be tightened to “benchmark-specific, in-distribution PTQ” unless the authors add a stress test under distribution shift (e.g., held-out task families, perturbed dynamics, or quantized-state replay for schedule selection). This is the central soft spot of the strongest claim, not an internal contradiction.","section":"§3.3, Eqs. (18)–(21), Limitations"},{"comment":"Table 6: Real-robot evidence is 10 trials per task (FP16 19/30, QuantWAMs 17/30, Atom* 12/30). The manuscript already disclaims equivalence, which is appropriate, but then “establish deployment feasibility” is doing a lot of work for a three-task, underpowered study on one platform. Either expand trials / report confidence intervals, or move real-robot results to a clearly labeled feasibility appendix and keep the primary claim on simulation Tables 1–2.","section":"§4.5, Table 6"}],"minor_comments":[{"comment":"Title and running header alternate “QuantWAMs” / “QuantW AMs” / “QuantW AMs”; normalize spelling throughout.","section":"Title, headers"},{"comment":"Figure 2 is dense; the three-column overview would benefit from a short caption walk-through mapping each panel to §§3.1–3.3.","section":"Figure 2"},{"comment":"Eq. (5)–(7): define how bσ²c and bτ²c are estimated in the main text (currently deferred entirely to Appendix A); a one-line estimator would help readers interpret N★ screens.","section":"§3.1, Prop. 1"},{"comment":"Tables 1–2 report “Speedup” and “Mem. (GB)” for targeted blocks only; add an explicit footnote on every table (not only §4.2) that these are not end-to-end control-loop metrics.","section":"Tables 1–2"},{"comment":"Clarify whether λv, λa and the 2%/20%/K hyperparameters were tuned on Dval or fixed a priori from training; free-parameter sensitivity belongs in the appendix if space is tight.","section":"§4.2 Quantization configuration"},{"comment":"Related Work on VLA quantization is brief; a sentence contrasting open-loop action-reconstruction metrics with closed-loop WAM success would sharpen the novelty claim.","section":"§2"}],"recommendation":"minor_revision","confidential_remarks":"Fit is appropriate for a systems/embodied-AI venue. Novelty is real but incremental relative to Atom/SmoothQuant/diffusion PTQ; the contribution is the calibration-context packaging for WAMs, not a new quantizer primitive. I would not block on the surrogate limitation if the authors narrow deployment language and add (or clearly refuse) the privileged-information control. No integrity concerns. Borderline major_revision only if the editor requires the matched-information baseline in the camera-ready; otherwise minor_revision is proportionate."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing: under a disclosed W4A4-dominant schedule, this gets Fast-WAM and LingBot-VA within 0.2–0.7 pp of FP16 on RoboTwin 2.0 and LIBERO, with real block-level memory/speed numbers and a small real-robot feasibility check. That is not vapor; the tables are seed-paired and the splits are trajectory-disjoint.\n\nWhat is actually new is not Atom, Hadamard, GPTQ, or timestep protection—those are borrowed. It is the calibration-context framing plus three operational fixes that match how WAMs are built: pool activation masks only when coordinates are literally shared (with a finite-sample crossover screen), score weights from the joint video–action gradient rather than fused marginals, and repair protected denoising steps by replaying FP16 rollout states under a fixed unprotected intervention. Prop. 1 and the recovered-energy plots make the pooling story falsifiable instead of hand-wavy. Matched-budget Atom*/SVDQuant* and the cumulative ladder on LIBERO-Long are the right controls for a methods paper of this type. Limitations section is unusually honest about local surrogates, block-level metrics, and underpowered robot trials.\n\nSoft spots, in proportion: the load-bearing assumption is exactly what they admit—32 cal trajectories and 32 FP16 rollouts lock masks, layer bits, and step indices via energy/Fisher/one-step replay, while closed-loop error compounds through Jacobians they never score. Matched bit budgets do not match the co-training backward pass or the privileged schedule audit, so some of the gap over Atom* is extra signal, not just better grouping. Efficiency is targeted blocks only; no end-to-end cycle or power story. Real robot is 10 trials/task—feasibility, not parity. No artifacts in the manuscript. Novelty is real but compositional.\n\nThis is for people shipping dual-stream or shared-backbone diffusion manipulators, and for PTQ folks who care about closed-loop multi-path models. Not a theory paper. I would send it to peer review; the central in-distribution claim holds and the experimental hygiene earns referee time. Engage if you work on embodied compression; skim the method figure and Tables 1–3 if you only need the takeaway.","headline":"Careful WAM-specific PTQ that actually closes most of the closed-loop gap under W4A4; novelty is compositional, hygiene is better than average, transfer claims stay soft.","tokens_in":16483,"tokens_out":585,"would_cite":true,"duration_ms":15685,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Post-training quantization for world action models works when calibration matches structure, rollout states, and the joint video–action objective.","keywords":["world action models","post-training quantization","closed-loop robotics","shared-basis calibration","joint video-action saliency","denoising-step protection","mixed precision","robot manipulation"],"falsifier":"On the same Fast-WAM or LingBot-VA checkpoints and RoboTwin/LIBERO protocols, a matched W4A4-dominant budget whose masks, layer upgrades, or protected denoising steps are chosen without shared-basis screening, joint co-training Fisher, or fixed-intervention replay would close most of the gap to QuantWAMs’ near-FP16 success; if it does not, or if success collapses under modest calibration-set or rollout-distribution shift, the central claim fails.","tokens_in":16335,"feed_emoji":"🤖","tokens_out":1090,"duration_ms":20072,"temperature":0.7,"pith_summary":"World Action Models predict future video and robot actions together, but iterative denoising and closed-loop control make full-precision deployment expensive. Standard post-training quantization fails here because it scores open-loop losses, treats the network as one homogeneous stream, and profiles states the robot never actually reaches. QuantWAMs argues that every precision choice is a finite-sample estimate fixed before deployment, and is only useful if three contexts stay right: which modules share a coordinate system, which closed-loop states are measured, and which joint training objective ranks importance. The method pools activation outliers only across coordinate-compatible modules, scores weight bits from the joint video–action gradient at layer granularity, and repairs which denoising steps stay high-precision by replaying real full-precision rollouts under a fixed unprotected intervention. On two WAM architectures and standard robot benchmarks, a mostly 4-bit schedule stays within a fraction of a percentage point of full precision in simulation, cuts targeted block memory to about 29% of FP16, and yields 1.4–1.6× block speedups, with real-robot trials showing the quantized policy can still run manipulation tasks.","feed_headline":"4-bit world action models stay within 0.7 points of FP16","feed_subtitle":"Calibrate pooling, joint gradients, and denoising steps to closed-loop context—not open-loop proxies.","key_machinery":"Calibration context: each quantization decision is treated as a finite-sample estimate that must match three axes at once—shared-basis pooling only where modules share quantizer coordinates, joint video–action empirical-Fisher saliency at layer granularity, and fixed-intervention replay that revises denoising-step protection on reachable FP16 rollout states without changing the bit budget.","core_discovery":"QuantWAMs claims that aligning post-training quantization with the structural, distributional, and objective calibration context of World Action Models recovers near–full-precision closed-loop success under a W4A4-dominant schedule: simulation means differ from FP16 by only 0.2–0.7 points while cutting targeted video/action block peak weight-and-activation memory to roughly 29% of FP16 and delivering 1.4–1.6× block-level speedups.","pith_inferences":["Any embodied policy whose early actions change later observations will inherit the same distributional mismatch unless calibration uses closed-loop states.","Coordinate-admissible pooling may generalize to other multi-expert or multi-modal transformers where literal channel indices do not mean the same thing across paths.","If one-step replay still misses long-horizon Jacobian products, future work may need multi-step counterfactual audits rather than larger bit budgets.","Benchmark-specific 32-trajectory calibration leaves open whether a single frozen schedule transfers to held-out tasks without re-profiling."],"forward_implications":["PTQ for closed-loop multi-stream robot policies should screen which modules may share activation statistics rather than pool by convenience.","Weight mixed-precision for jointly trained video–action models should score the combined gradient, not fuse single-stream scores after the fact.","Denoising-step protection schedules should be audited on reachable closed-loop states under a fixed intervention, not on synthetic open-loop inputs alone.","A mostly 4-bit WAM path can retain near-FP16 task success in simulation while cutting targeted block memory to about 29% of FP16.","The same calibration recipe can transfer across dual-stream and shared-backbone WAM architectures with architecture-specific grouping rules."],"fun_headline_variants":["QuantWAMs keep W4A4 WAMs within 0.7 pts of FP16","Align PTQ to WAM structure, rollouts, and joint objective","Shared-basis outliers and Fisher scores for closed-loop WAMs","W4A4 WAMs hit ~29% memory and 1.4–1.6× block speedups","Fixed-intervention audits protect WAM denoising under W4A4"],"cache_read_input_tokens":128,"weakest_assumption_plain":"Local proxies fitted on only 32 calibration trajectories and 32 full-precision rollouts—channel energy masks, layer Fisher scores, and one-step replay errors—remain the right precision choices at deployment even though the method never optimizes closed-loop task return and early errors compound through later states.","fun_headline_variants_meta":{"raw":{"variants":["QuantWAMs keep W4A4 WAMs within 0.7 pts of FP16","Align PTQ to WAM structure, rollouts, and joint objective","Shared-basis outliers and Fisher scores for closed-loop WAMs","W4A4 WAMs hit ~29% memory and 1.4–1.6× block speedups","Fixed-intervention audits protect WAM denoising under W4A4"]},"model":"grok-4.5","effort":"low","cost_usd":0.002575,"raw_usage":{"total_tokens":1081,"prompt_tokens":869,"num_sources_used":0,"completion_tokens":93,"cost_in_usd_ticks":25748000,"prompt_tokens_details":{"text_tokens":869,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":119,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":869,"tokens_out":93,"duration_ms":3347,"temperature":1.0,"reasoning_tokens":119,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T08:22:56.232511+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On the same Fast-WAM or LingBot-VA checkpoints and RoboTwin/LIBERO protocols, a matched W4A4-dominant budget whose masks, layer upgrades, or protected denoising steps are chosen without shared-basis screening, joint co-training Fisher, or fixed-intervention replay would close most of the gap to QuantWAMs’ near-FP16 success; if it does not, or if success collapses under modest calibration-set or rollout-distribution shift, the central claim fails.","supporting_citations":[],"review_version":1}