{"id":"7cb76f3c-e178-49bf-be51-2cf10f97653c","arxiv_id":"2501.05329","paper_version":1,"verdict":"REJECT","confidence":"LOW","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Reward distillation from a 317M TD-MPC2 teacher yields a 1M-parameter MT30 agent that scores 28.12 to 28.45, narrowly above a from-scratch baseline, with FP16 halving size.","lead":"This preprint distills a 317M-parameter TD-MPC2 reinforcement learning agent into a 1M-parameter student by adding a reward-prediction distillation loss, then applies FP16 quantization. On the 30-task MT30 benchmark it reports a normalized score up to 28.45, about 2.8% above a from-scratch 1M model.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 28.45 FP16 score appears in no experiment; all tabulated results and the quoted +48.5%/+2.77% gains correspond to 28.12, so the central SOTA claim is internally unsupported.","rationale":"The reader's weakest_assumption focuses on whether reward-prediction distillation is a better learning signal than environment rewards or latent representations. That is a plausible scientific concern about mechanism and generalization. However, the single most load-bearing issue for the paper's central claim is more basic and more decisive: the headline number 28.45 is not supported by any reported experiment. The introduction's quoted improvements, +48.5% and +2.77%, exactly match 28.12, the best tabulated distilled score, not 28.45. The FP16-quantized model, which the abstract credits with the 28.45 score, is never evaluated in the results section. This is an internal inconsistency that directly invalidates the abstract's state-of-the-art claim. A reader cannot assess whether the method works because the primary quantitative evidence is absent and the reported numbers contradict each other. The reader's overall REJECT verdict is appropriate, but I would reach it through this arithmetic and evidentiary gap rather than through the distillation-signal assumption. The suggested concrete test, obtaining and recomputing the FP16 evaluation, would settle whether the 28.45 is a simple typo or a fabricated/unsupported result; either way, the manuscript as written does not support its central claim. I therefore agree with the rejection but identify a different weakest point, hence 'disagree' on the specific weakest_assumption while affirming the verdict.","tokens_in":3635,"tokens_out":3890,"duration_ms":34789,"concrete_test":"Request the authors' training and evaluation logs for the FP16-quantized model, including the exact configuration (d_coef, batch size, training steps, seeds) and per-task scores. Independently recompute the normalized MT30 score and the two relative improvements from Table 2 and from the FP16 run. If the FP16 score equals 28.12 or is unavailable, then the abstract's 28.45 is either a typo or an unsupported claim, and the state-of-the-art result collapses. If the FP16 score is genuinely 28.45, the authors must supply the missing experimental row and correct the +48.5% and +2.77% percentages, which currently correspond to 28.12.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim is that FP16 quantization yields a 28.45 normalized MT30 score, surpassing the original 1M model by 48.5% and a from-scratch model by 2.77%. However, no experimental section reports an FP16 result. Table 2's best distilled score is 28.12 (batch 256, 1M steps, d_coef=0.45). Arithmetic shows 28.12/18.93 - 1 = 48.5% and 28.12/27.36 - 1 = 2.78%, exactly the percentages quoted in the introduction. Thus the abstract's 28.45 is not merely an omitted detail; it is inconsistent with the paper's own data. If the FP16 model actually scored 28.45, the improvement over 18.93 would be 50.3%, not 48.5%, and the improvement over 27.36 would be 4.0%, not 2.77%. Moreover, the claimed benefit of FP16 quantization, maintaining performance while reducing size, is never measured in any table, figure, or sentence beyond the abstract. Because the abstract and introduction assert state-of-the-art performance based on this number, the paper's central quantitative contribution has no supporting experimental evidence in the manuscript. This is a load-bearing internal inconsistency, not a matter of missing code or seeds: the numbers that are reported actively contradict the headline result.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a teacher-student knowledge distillation method for model-based RL, distilling a 317M-parameter TD-MPC2 teacher into a 1M-parameter student by adding a reward-distillation MSE loss to the original TD-MPC2 objective. The authors report normalized scores on the MT30 benchmark for different distillation coefficients, batch sizes, training durations, teacher sizes, and a latent-distillation variant, and they claim that an FP16-quantized version of the distilled model achieves a state-of-the-art normalized score of 28.45, surpassing the original 1M model by 48.5% and a from-scratch model by 2.77%. The paper also acknowledges limitations including the need for real-world deployment and the narrow focus on MT30.","tokens_in":3895,"tokens_out":4135,"duration_ms":36481,"significance":"If the headline claims were supported by the reported experiments, the paper would offer a useful recipe for compressing large world models into deployable 1M-parameter agents while preserving multi-task performance on a standard benchmark. The ablations over d_coef, batch size, training length, teacher capacity, and latent distillation are informative and provide a basis for follow-up work. However, the central quantitative claims are undermined by an internal inconsistency in the reported scores and by the complete absence of any experimental evaluation of FP16 quantization. The lack of multiple seeds or error bars further weakens the significance of the claimed 2.77% improvement over from-scratch training.","major_comments":[{"comment":"The abstract and introduction claim a normalized score of 28.45 and improvements of +48.5% and +2.77%, but no experimental section reports an FP16 score or a score of 28.45. The best distilled score in Table 2 is 28.12 (d_coef=0.45, batch size 256, 1M steps), and the arithmetic 28.12/18.93 - 1 = 48.5% and 28.12/27.36 - 1 = 2.78% matches the quoted percentages exactly. The headline number is therefore not merely an omitted detail; it is inconsistent with the paper's own data. If the true FP16 score were 28.45, the improvements would be 50.3% and 4.0%, not 48.5% and 2.77%. This inconsistency directly affects the paper's central claim of state-of-the-art performance.","section":"Abstract and Section 3 (Table 2)"},{"comment":"The abstract and introduction state that FP16 post-training quantization reduces model size by 50% while maintaining performance, and the introduction explicitly attributes the 28.45 state-of-the-art result to the FP16-quantized model. However, Section 3 contains no quantization experiment: no FP16 scores, no size measurements, no inference-speed or memory measurements. The only statement about FP16 in the methods is the sentence 'Lastly, FP16 quantization is applied.' Without any experimental evidence, the claimed benefit of FP16 quantization is unsupported.","section":"Section 2 and Section 3"},{"comment":"Table 2 reports each setup as a single normalized score with no standard deviation, confidence interval, or number of seeds. The central comparison between the distilled model (28.12) and the from-scratch model (27.36) is a difference of 0.76 points, which is likely within training noise for a single-seed run on a multi-task benchmark. Without repeated runs or variance estimates, the claimed +2.77% improvement is not statistically substantiated.","section":"Table 2 and Section 3"},{"comment":"The paper selects d_coef, batch size, and training length after inspecting the normalized scores on the MT30 benchmark: Section 3 states that values close to 0.5 yield the best results and that batch size 256 offers an optimal balance, but these choices are made from the same table used for the final comparison. This is a selected optimum, not a predictive evaluation. A held-out validation split or a correction for multiple comparisons would be needed to claim that distillation outperforms from-scratch training by the reported margin.","section":"Section 3 (Tables 1 and 2)"},{"comment":"The claim that distillation improves over from-scratch training is not supported by the full table: at 200K steps with batch size 1024, from-scratch achieves 18.70 while distillation achieves 18.11, and at 337K steps with batch size 1024, from-scratch achieves 26.94 while distillation achieves 25.44. The method therefore underperforms in several configurations, and the positive conclusion is based only on the batch-256, 1M-step setting. The paper should either explain why the method fails at larger batch sizes or temper the claim that distillation is generally beneficial.","section":"Table 2"}],"minor_comments":[{"comment":"There is a typo: 'the 317M-parameter TD-MPC2 model servers as the teacher' should be 'serves as the teacher.'","section":"Section 2"},{"comment":"The sentence 'Each task is scored on a scale of 1 to 1000, with the average sum divided by the number of tasks' is awkward; it should say the sum of task scores is divided by the number of tasks.","section":"Section 2"},{"comment":"Table 1 shows an optimal d_coef of 0.4, but Table 2's best result uses d_coef = 0.45, which does not appear in the d_coef sweep. The paper should clarify whether 0.45 was evaluated separately and how it relates to the sweep.","section":"Tables 1 and 2"},{"comment":"The figure caption refers to colors ('in red', 'in teal', 'in green', 'in blue'), but the figure is not visible in the manuscript text. Please ensure the figure is included and consider a color-blind-safe palette.","section":"Figure 1"},{"comment":"The sentence 'reward prediction directly aligns with task-specific performance' is supported only by two example tasks (pendulum-swingup and cup-catch). A task-level breakdown of the MT30 scores would strengthen this interpretation.","section":"Section 3"}],"recommendation":"reject","confidential_remarks":"The manuscript appears to be an extended abstract whose central quantitative claims are internally inconsistent: the abstract's 28.45 score and the associated percentages do not match any reported experiment, and the FP16 quantization that is central to the introduction's 'state-of-the-art' claim is never evaluated. The missing seeds/error bars are also a serious issue for a 0.76-point improvement claim. I see no way to fix these problems without substantially extending the experiments, so rejection is appropriate. If the authors later provide a full version with corrected numbers, repeated seeds, and actual quantization measurements, the distillation idea itself may be worth reconsidering."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper has a real potential nugget—reward-only distillation from a 317M TD-MPC2 teacher into a 1M student on MT30—but the abstract's headline number (28.45 FP16, +48.5%) is inconsistent with the paper's own tables. The best tabulated distilled score is 28.12, and the quoted +48.5% and +2.77% arithmetic works only for 28.12/18.93 and 28.12/27.36. No FP16 result appears anywhere in the results section. That is a load-bearing internal inconsistency, not a missing footnote.\n\nWhat's actually new: the specific application of teacher-student distillation to TD-MPC2, the choice of reward-prediction MSE as the distilled signal, and a clean teacher-size comparison (48M vs 317M teacher, 13.61 vs 17.85). The method section is concise, and the d_coef sweep in Table 1 is a reasonable empirical search. The negative result on latent distillation (7.69/8.78 vs 14.04) is honestly reported and useful.\n\nSoft spots: beyond the headline inconsistency, the evidence for state-of-the-art rests on a single configuration: batch 256, 1M steps, d_coef=0.45. Table 2 shows from-scratch training beating distillation at 200K with batch 1024 (18.7 vs 18.11) and at 337K (26.94 vs 25.44). That does not kill the method, but it means the advantage is not robust across the tested grid. The hyperparameters (d_coef, batch, steps) were selected after looking at benchmark scores, so the final comparison is partly a selected optimum. There is no code, no seeds, no error bars, and no ablation for FP16 quantization. The conclusion's limitations paragraph acknowledges the MT30-only focus, which is honest.\n\nWho it's for: researchers working on model compression for model-based RL, especially those targeting deployment on limited hardware. The idea is worth a serious look, but the analysis is not yet at a level that supports the abstract's claims.\n\nRecommendation: if this crossed my desk, I would send it to peer review, but with a clear request to fix the numerical inconsistency and provide the missing FP16 results and ablations before acceptance. As is, it needs heavy revision; the central claim is unsupported by the reported data.","headline":"The distillation recipe is sensible and the teacher-size comparison is useful, but the headline FP16 number appears in no table and the abstract's percentages match a different score, so the paper's central claim is unsupported.","tokens_in":4495,"tokens_out":2530,"would_cite":false,"duration_ms":21307,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a 1M-parameter model-based RL agent, trained by matching a 317M-parameter teacher's reward predictions and then applying FP16 quantization, achieves a 28.45 normalized score on MT30, beating the original 1M…","keywords":["model-based reinforcement learning","knowledge distillation","multi-task learning","TD-MPC2","MT30 benchmark","FP16 quantization","reward prediction","model compression"],"falsifier":"Run the same 1M-parameter student from scratch on MT30 with the identical training budget (1M steps, batch size 256) across multiple seeds; if the mean from-scratch score reaches or exceeds 28.12, the reported transfer gain is within run-to-run noise and the distillation loss is not the cause.","tokens_in":3379,"feed_emoji":"🤖","tokens_out":7200,"duration_ms":60345,"temperature":0.7,"pith_summary":"This paper tries to show that model-based reinforcement learning agents can be made dramatically smaller without losing competence, by having a small student imitate the reward predictions of a large teacher. The authors take a 317M-parameter world model and distill it into a 1M-parameter student using an extra loss that matches the teacher's predicted reward for each state-action pair. On the 30-task MT30 benchmark, the distilled student reaches a normalized score of 28.45 after FP16 quantization, up from 18.93 for the original 1M model and 27.36 for a from-scratch 1M model. The claim matters because it points to a cheap route from large, capable RL models to small models that can run on limited hardware, such as robots or edge devices. The load-bearing idea is that reward predictions, not latent representations, are the right thing to transfer.","feed_headline":"Distilled 1M RL agent scores 28.45 on MT30, beats from-scratch","feed_subtitle":"Compact student copies teacher's reward predictions and outperforms 1M models trained from scratch on 30 control tasks.","key_machinery":"The central mechanism is the reward distillation loss $L_{distill} = \\mathrm{MSE}(R_{teacher}(s,a), R_{student}(s,a))$, added to the original TD-MPC2 losses with a weighting coefficient $d_{coef}$. The teacher is a frozen 317M-parameter TD-MPC2 world model; the student is a trainable 1M-parameter version. This loss funnels the teacher's knowledge into a scalar reward channel, which the student can match even though its latent dimension is much smaller, and it is combined with FP16 post-training quantization to halve the model footprint.","core_discovery":"On its own terms, the paper's central discovery is that a compact agent can inherit multi-task behavior by supervising it to reproduce the teacher's reward estimates. Concretely, the student's total loss is the original TD-MPC2 loss (consistency, reward, and value) plus a term $d_{coef} \\cdot \\mathrm{MSE}(R_{teacher}(s,a), R_{student}(s,a))$, with $d_{coef}$ around 0.4-0.5 working best. With this reward-distillation loss, a 1M-parameter student trained for 1M steps with batch size 256 scores 28.12 on MT30, compared to 27.36 for an equivalent from-scratch model, and the final FP16-quantized model scores 28.45. The paper also finds that distilling next-state latent representations fails (scores of 7.69-8.78) because the teacher's 1376-dim latent space cannot be projected down to the student's 128-dim space without losing the information that matters. Reward prediction, by contrast, is a low-dimensional, task-aligned target, which is why the authors conclude it transfers better.","pith_inferences":["The reported margins over from-scratch are small (2.77% at 1M steps), and the paper reports no variance; a multi-seed replication would establish whether the benefit is robust or seed luck.","Because the teacher's reward predictions are learned and shaped by the teacher's world model, the distilled student may inherit the teacher's reward mis-specifications; a task with a deliberately miscalibrated teacher reward would test this.","A natural extension is to apply reward distillation in partially observable or sparse-reward settings, where reward predictions are less informative, and see whether the method's advantage shrinks.","The success of scalar reward targets over high-dimensional latent targets suggests a more general principle for distillation across capacities: match the output signal that is closest to the task metric, not the most information-rich representation."],"forward_implications":["Reward prediction is a viable knowledge channel for compressing model-based RL agents, so future compression pipelines can avoid expensive latent-space matching.","A 1M-parameter agent, after reward distillation and FP16 quantization, can be deployed on a single consumer GPU or edge device while retaining competitive multi-task performance.","Bigger teachers transfer more: the 317M teacher gives a 31.2% relative improvement over the 48M teacher under the same short distillation schedule.","Extended distillation (1M steps) with a moderate batch size is needed to surpass from-scratch training; shorter runs still beat the original 1M checkpoint but not from-scratch at equal steps.","Latent distillation is not a viable alternative when teacher and student latent dimensions differ, so reward-level targets are the pragmatic choice for heterogeneous architectures."],"supporting_citations":[{"why":"Provides the TD-MPC2 architecture, the 317M and 1M checkpoints, the MT30 benchmark, and the from-scratch baselines.","marker":"[2]"},{"why":"Introduces the soft-target distillation idea that the reward MSE adapts.","marker":"[3]"},{"why":"Supplies the policy-distillation background that motivates transferring behavior via a teacher-student objective.","marker":"[1]"},{"why":"Shows early multitask transfer through teacher-student actor mimicry.","marker":"[5]"},{"why":"Establishes policy distillation as a compression method for RL policies.","marker":"[6]"},{"why":"Provides a robust multitask distillation objective that regularizes student policies across tasks.","marker":"[8]"},{"why":"Underlies the FP16 mixed-precision training and quantization used to halve model size.","marker":"[4]"},{"why":"Defines the continuous-control task suite and the normalized-score metric used by MT30.","marker":"[7]"}],"fun_headline_variants":["Reward distillation: 1M RL agent beats from-scratch on MT30","Compact RL agent learns by copying teacher's reward estimates","317M to 1M: distilled RL agent scores 28.45 on MT30","Distilled RL: 1M params, FP16 halves size, retains performance","Reward prediction transfer beats latent distillation in RL"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole result depends on the idea that matching the teacher's reward predictions, instead of the environment's true rewards or the teacher's internal representations, gives a 1M-parameter student a consistently better learning signal across all 30 MT30 tasks.","fun_headline_variants_meta":{"raw":{"variants":["Reward distillation: 1M RL agent beats from-scratch on MT30","Compact RL agent learns by copying teacher's reward estimates","317M to 1M: distilled RL agent scores 28.45 on MT30","Distilled RL: 1M params, FP16 halves size, retains performance","Reward prediction transfer beats latent distillation in RL"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000843,"raw_usage":{"total_tokens":3670,"prompt_tokens":939,"completion_tokens":2731,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":2635}},"tokens_in":555,"tokens_out":2731,"duration_ms":19308,"temperature":1.0,"reasoning_tokens":2635,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:13:56.344206+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 1M-parameter student from scratch on MT30 with the identical training budget (1M steps, batch size 256) across multiple seeds; if the mean from-scratch score reaches or exceeds 28.12, the reported transfer gain is within run-to-run noise and the distillation loss is not the cause.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the policy-distillation background that motivates transferring behavior via a teacher-student objective."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides a robust multitask distillation objective that regularizes student policies across tasks."}],"review_version":1}