{"id":"dd0e275e-168e-429b-abb2-b337caa0ff34","arxiv_id":"2509.08826","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"RewardDance reframes visual reward modeling as a yes/no judgment task in a VLM and reports consistent gains in text-to-image, text-to-video, and image-to-video generation as the reward model scales from 1B to 26B.","lead":"RewardDance trains vision-language models to judge whether one generated image or video beats a reference, then uses those judgments to fine-tune generators. The paper reports that scaling these reward judges from 1B to 26B parameters improves generation quality and argues it suppresses reward hacking.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reward variance is an unvalidated proxy for reward-hacking resistance; the abstract's 'proving' overstates the correlational evidence in §4.4/Figs. 2,5,6, so the central anti-hacking claim is not yet established.","rationale":"The reader's weakest_assumption identifies exactly the load-bearing concern: high reward variance is treated as proof of anti-hacking, but no external validation connects variance to actual output diversity or to the absence of reward overoptimization. I agree this is the single most important weakness because it directly undermines the paper's strongest claim ('Crucially, we resolve the persistent challenge of reward hacking'). Other issues—proprietary benchmarks, missing artifacts, ID accuracy paradox—are real but secondary; they affect reproducibility and generalizability rather than the internal logic of the central claim. The scaling results themselves (Tables 2–4) are plausible and the GenEval improvements provide some external anchor, so I would not move the verdict to REJECT. The paper should either (a) validate the variance proxy with diversity and human-preference metrics, or (b) soften the 'proving' language to 'suggesting' and reframe the contribution as scaling evidence rather than a resolved anti-hacking claim. Since the reader already recommended CONDITIONAL, my assessment does not change the verdict; it reinforces the need for the stated conditions.","tokens_in":17951,"tokens_out":6678,"duration_ms":79996,"concrete_test":"Re-run the Seedream-3.0 RL fine-tuning with the 26B RM, but corrupt the reward signal by adding independent Gaussian noise to P(yes) with σ chosen to match the observed late-training variance (≈5.4e-2 from Figure 5). If the final alignment score and sample diversity (LPIPS/Vendi) do not significantly degrade compared to the clean 26B run, then matching reward variance is not sufficient to ensure anti-hacking. Alternatively, compute the per-iteration reward variance and the policy's generated-image diversity and human-preference gap on a fixed OOD prompt set for RM sizes 1B-26B; if variance does not positively correlate with diversity or negatively correlate with the reward-preference gap, the proxy fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.4 and Figures 2/5/6 treat the variance of the reward score over a sliding window of 1,000 RL iterations as a direct measure of exploration and resistance to reward hacking. The abstract goes further: 'Our large-scale RMs exhibit and maintain high reward variance during RL fine-tuning, proving their resistance to hacking.' This inference is not valid without external validation. Reward variance is a property of the RM's scoring function, not of the policy's output distribution. A noisy or miscalibrated RM—or a policy that oscillates between exploiting different loopholes—can also produce high temporal variance, while a diverse and high-quality policy might receive stable rewards from a well-calibrated RM. The paper reports no measure of output diversity (e.g., LPIPS, Vendi score) and no reward-overoptimization curve (reward vs. human preference on held-out OOD prompts). Indeed, the body text only says the relationship 'suggests' resistance (Section 4.4), which is inconsistent with the abstract's 'proving.' The empirical correlation between RM size and variance is interesting, but it does not establish the causal anti-hacking claim. Since the anti-hacking result is the headline 'Crucially' contribution, the central claim is load-bearing and currently unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces RewardDance, a generative reward modeling framework for visual generation. Instead of a regression head, the reward score is defined as a VLM's probability of predicting a 'yes' token for 'image 2 is better than image 1' under task-specific instructions, reference examples, and chain-of-thought reasoning. The framework is scaled from 1B to 26B parameters and evaluated on text-to-image, text-to-video, and image-to-video tasks under RL fine-tuning and test-time scaling. The main reported results are consistent scaling gains in alignment score, GenEval, Bench-240, and SeedVideoBench-1.0. The paper further claims that large reward models exhibit high reward variance during RL fine-tuning, which it interprets as proof of resistance to reward hacking and sustained output diversity.","tokens_in":18296,"tokens_out":4311,"duration_ms":49495,"significance":"If the scaling results hold, the paper makes a practically important contribution: it is the first systematic study of scaling a generative VLM-based reward model for visual generation, and it demonstrates gains across multiple base generators (FLUX.1-dev, Seedream-3.0, Seedance-1.0) and two optimization regimes (RL and test-time scaling). The OOD-accuracy discussion is a useful suggestion for reward-model evaluation. However, the headline 'Crucially, we resolve reward hacking' claim is not currently established: it rests entirely on an unvalidated proxy, reward variance, and the body text itself only says the evidence 'suggests' resistance. Because this is the central novelty emphasized in the abstract and introduction, the paper needs either direct validation of the anti-hacking mechanism or a substantial tempering of the claim.","major_comments":[{"comment":"The anti-hacking conclusion is based solely on reward-variance dynamics. The abstract says high reward variance 'proves' resistance to hacking, but §4.4 only says the relationship 'strongly suggests' it, and Fig. 1 calls variance an 'indicator.' No independent validation is provided: no output-diversity metric (e.g., LPIPS, Vendi), no reward-overoptimization curve on held-out human preference, and no comparison with a policy known to have been hacked. A noisy or miscalibrated reward model can produce high temporal variance even for a degenerate policy, while a well-calibrated model can assign stable rewards to diverse high-quality outputs. As the central 'Crucially' contribution, this needs direct evidence, not a proxy.","section":"Abstract; §4.4; Figs. 2, 5, 6; Eq. (2)"},{"comment":"The in-domain RM accuracy is non-monotonic with scale (64.70, 69.36, 65.37, 74.92, 78.44), and the paper introduces OOD accuracy post hoc as 'more critical' after observing that it correlates with the desired scaling trend. With only five RM sizes and no independent, pre-registered metric selection, this is an ad-hoc reinterpretation. OOD accuracy could indeed be the better predictor, but that claim needs support beyond a single correlation on the authors' own data, e.g., cross-validation across RM families or benchmarks, or an a priori argument for why OOD accuracy should be decisive.","section":"Table 2; §4.2"},{"comment":"Two of the three headline benchmarks, Bench-240 and SeedVideoBench-1.0, are developed by the same organization that produces the evaluated models (Seedream and Seedance). The paper reports no independent human evaluation, no confidence intervals, and no third-party replication for these benchmarks. The SOTA claims in Tables 5 and 6 therefore rest largely on internally constructed evaluation sets. Independent evaluation or release of the full prompt/rating protocol is needed before the SOTA claims can be accepted at face value.","section":"Tables 5 and 6; §4.1.2"}],"minor_comments":[{"comment":"Typo: 'It primarily due' should be 'It is primarily due.'","section":"Abstract"},{"comment":"The numbers next to the curves (e.g., '=7.2e-3') are not labeled. Please state explicitly that these are standard deviations of raw/smoothed reward scores, and report the sliding-window length (stated in §4.4 as 1,000) in the captions.","section":"Figs. 2, 5, 6"},{"comment":"The cell entries such as '+28%+32% +4%' are visually confusing. Use separate columns for GSB improvement and its uncertainty, or explain the notation in the caption.","section":"Table 3"},{"comment":"Typo: 'point-wisee' should be 'point-wise.'","section":"§3.3.2"},{"comment":"The weighted CE loss coefficient for the pointwise generative variant is described only as 'small.' Please give the exact value for reproducibility.","section":"§3.2.3"},{"comment":"The claim that larger DiT architectures benefit more from reward scaling rests on a single comparison (Seedream-Lite vs. Seedream). Add at least one more model pair or error bars before drawing a scaling-law conclusion.","section":"§4.4 and Fig. 7"}],"recommendation":"major_revision","confidential_remarks":"The paper is from ByteDance Seed, and several of the baseline models, evaluation benchmarks, and reference technical reports are the authors' own. Please consider whether the internal Bench-240 and SeedVideoBench-1.0 evaluations need independent corroboration before publication. The reward-variance-as-proof-of-no-hacking claim is the main novelty; it is not yet supported by the evidence in the manuscript. This is fixable in revision by adding direct diversity/overoptimization measurements or by rewriting the claim as a hypothesis, but as written it substantially overstates the result."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what is actually useful: this is the first systematic scaling study of a pairwise generative reward model for visual generation, from 1B to 26B, across T2I, T2V, and I2V. The P(yes) formulation is borrowed from DeepSeek-GRM, Pairwise RM, and UnifiedReward, but the scaling sweeps, the context-scaling recipe (instructions, reference examples, CoT), and the ablations (regressive vs generative, BoN reference quality, CoT finetuning) are genuinely new and informative. The GenEval numbers provide an external anchor, and the internal scaling trends are consistent across base models and tasks. Credit where it's due.\n\nThe big problem is the reward-hacking claim. Reward variance is a property of the RM's scoring function, not the policy's output diversity. A noisy RM, or a policy that oscillates between exploiting different loopholes, can also show high variance. The paper provides no diversity metric (LPIPS, Vendi, etc.) and no reward-overoptimization curve (reward vs. human preference on held-out prompts). So the abstract's \"proving their resistance to hacking\" is not supported; the body's \"strongly suggests\" is already a stretch. This is load-bearing because the abstract and contributions lean on it.\n\nRelated: the ID accuracy is non-monotonic (64.70, 69.36, 65.37, 74.92, 78.44) and the paper then declares OOD accuracy the \"more critical core metric.\" That is post hoc selection, and while it is an interesting observation, it is not a strong basis for the conclusion. Also, Bench-240 and SeedVideoBench-1.0 come from the same lab, so those comparisons are less independent.\n\nMinor: no uncertainty estimates on the scaling trends, no code/data released, and the BoN reference count and variance window length are free parameters. None of those are fatal to the scaling result.\n\nOverall, the scaling result is plausible and worth engaging with. But the anti-hacking conclusion should either be validated externally or removed/softened. This paper deserves a serious referee, and I would tell the authors to add diversity metrics, overoptimization curves, and release artifacts.","headline":"Useful scaling study; the 'proving' reward-hacking claim rests on an unvalidated variance proxy and should be softened.","tokens_in":18810,"tokens_out":2531,"would_cite":true,"duration_ms":25617,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"RewardDance claims that recasting reward as a yes-token probability in a VLM unlocks scaling to 26B parameters and yields resistance to reward hacking.","keywords":["reward models","visual generation","reinforcement learning from human feedback","reward hacking","scaling laws","vision-language models","text-to-image","text-to-video"],"falsifier":"Run an RL fine-tuning loop with a large generative reward model and measure actual output diversity (e.g., perceptual feature coverage or pairwise image distance) alongside the reward variance. If reward variance stays high while output diversity drops sharply, or if held-out human preference scores flatten or worsen, the anti-hacking claim would be refuted. Comparing a 26B generative RM against a matching 26B regression RM on identical reference images would isolate whether the benefit comes from the generative formulation or simply from parameter count.","tokens_in":17821,"feed_emoji":"🎨","tokens_out":5420,"duration_ms":47269,"temperature":0.7,"pith_summary":"RewardDance argues that the usual way to build visual reward models—a scalar regression head trained with Bradley-Terry loss—is what keeps them from scaling well. Instead, it reformulates reward as the probability that a vision-language model outputs the token 'yes' when asked whether one image beats another, which natively fits how VLMs generate text. This reformulation unlocks scaling along two axes: model parameters (from 1B to 26B) and input context (task instructions, reference images, and chain-of-thought reasoning). Across text-to-image, text-to-video, and image-to-video, the paper reports consistent quality gains from both kinds of scaling. It also reports that large reward models maintain high reward variance during RL fine-tuning, which it reads as evidence that the policy avoids reward hacking and mode collapse.","feed_headline":"Yes-token reward model scales to 26B and resists reward hacking","feed_subtitle":"Recasting reward as a VLM's 'yes' probability yields stable gains in image and video generation.","key_machinery":"The central object is the 'yes'-token probability: r(x1,x2,y,i) = P('yes' | x1,x2,y,i), where x1,x2 are images being compared, y is the prompt, and i is a task instruction. This turns reward prediction into the VLM's native autoregressive task and removes the regression head. Two scaling axes then become effective: model scaling (InternVL variants from 1B to 26B) and context scaling (task instructions, reference images, and chain-of-thought reasoning). The reward curves during RL fine-tuning, with their variance bands, serve as the diagnostic that larger models keep exploring rather than collapsing.","core_discovery":"RewardDance replaces the regression head of a VLM reward model with a generative question: 'Is image 2 better than image 1?' The reward is simply the predicted probability of the token 'yes'. This aligns the reward objective with the VLM's next-token prediction, and with that alignment the paper demonstrates that scaling model parameters from 1B to 26B and scaling context (instructions, reference images, chain-of-thought) yields consistent quality improvements in text-to-image, text-to-video, and image-to-video generation. The paper also observes that large reward models keep high reward variance during RL fine-tuning, interpreting this as evidence that the policy avoids reward hacking and m","pith_inferences":["The paper's anti-hacking evidence rests on reward variance as a proxy; a stronger test would measure output diversity directly (e.g., feature coverage or perceptual diversity) independently of the reward model used.","Scaling laws for reward models may interact with generator scale: the paper already notes that larger diffusion models benefit more from larger RMs, implying a joint-scaling recipe rather than an isolated RM-scaling law.","If high variance is indeed the key signal, then reward-model designers might deliberately tune calibration and output entropy, not just accuracy, when training RMs for RL.","The pairwise reference-image formulation introduces an N-way search cost at inference; a testable extension would be to amortize or distill the pairwise comparisons into a pointwise model that retains the scaling benefits."],"forward_implications":["If scaling RMs is the right principle, then visual generation systems should invest in larger VLM-based reward models rather than only larger generators.","Context scaling with reference examples and chain-of-thought provides a practical path to better reward signals without changing the base generator.","The reported variance signature gives a cheap, training-time early warning for reward hacking: a shrinking variance band during RL fine-tuning indicates the policy is collapsing.","The same generative reward formulation should carry over to other preference-based multimodal tasks, such as editing or audio-to-video, with minimal changes.","OOD accuracy of the reward model, not in-domain accuracy, is the metric that predicts downstream RL gains, pointing toward new benchmark design for RMs."],"fun_headline_variants":["Reward as 'yes' token: scaling to 26B resists hacking","Generative reward model scales to 26B, thwarts hacking","Yes-token reward scales to 26B, resists reward hacking","RewardDance: 'yes' token scaling to 26B for stable generation"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper treats high reward variance during RL fine-tuning as proof that the policy is not reward-hacking, but variance alone could also reflect a noisy or miscalibrated reward signal rather than genuinely broad exploration.","fun_headline_variants_meta":{"raw":{"variants":["Reward as 'yes' token: scaling to 26B resists hacking","Generative reward model scales to 26B, thwarts hacking","Yes-token reward scales to 26B, resists reward hacking","RewardDance: 'yes' token scaling to 26B for stable generation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000637,"raw_usage":{"total_tokens":2816,"prompt_tokens":831,"completion_tokens":1985,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":575,"completion_tokens_details":{"reasoning_tokens":1902}},"tokens_in":575,"tokens_out":1985,"duration_ms":15933,"temperature":1.0,"reasoning_tokens":1902,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T20:03:58.485216+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run an RL fine-tuning loop with a large generative reward model and measure actual output diversity (e.g., perceptual feature coverage or pairwise image distance) alongside the reward variance. If reward variance stays high while output diversity drops sharply, or if held-out human preference scores flatten or worsen, the anti-hacking claim would be refuted. Comparing a 26B generative RM against a matching 26B regression RM on identical reference images would isolate whether the benefit comes from the generative formulation or simply from parameter count.","supporting_citations":[],"review_version":1}