{"id":"07dd23cc-30cb-465b-9db7-db45eb4acc23","arxiv_id":"2508.11255","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":2,"one_line_summary":"A three-part system, Talking-Critic, Talking-NSQ, and TLPO, aligns diffusion portrait animation models to human preferences and improves lip-sync, motion naturalness, and visual quality.","lead":"The paper introduces a reward model, a 410K-pair preference dataset, and a multi-expert optimization method to make audio-driven animated faces more natural, lip-synced, and visually clean. A generalist might care because this is a step toward automated dubbing and realistic virtual avatars.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Talking-Critic both curates the 410K preference pairs and supplies the TLPO reward signal, so claimed human-preference gains may reflect reward hacking against the critic rather than genuine improvement.","rationale":"The paper's headline is an empirical claim about human preference. The internal logic makes Talking-Critic the sole source of preference information: it creates the dataset and scores the optimization. Unless an independent human-evaluation layer is present, the loop is closed and the claim cannot be distinguished from reward hacking. This is not a charge of bad faith; it is a standard failure mode of RLHF/DPO pipelines, and the manuscript provides no visible guard against it. The reader's weakest_assumption already points at critic faithfulness and label quality; I agree with that and would sharpen it: even a fairly accurate critic can be exploited by the optimization, and using the same critic to curate the training pairs removes the usual safeguard of a held-out human label set. If the intact paper shows a separate human study, the concern would be resolved; if not, the claim should be treated as conditional at best. Given that the full text is inaccessible, the appropriate disposition remains the reader's UNVERDICTED; hence the verdict is unchanged rather than escalated.","tokens_in":15343,"tokens_out":5506,"duration_ms":67573,"concrete_test":"Recover the intact manuscript and locate the evaluation section. If any main quantitative result for TLPO is computed using Talking-Critic scores or using preference pairs generated by Talking-Critic as the evaluation signal, the circularity concern lands. The decisive external check is a pre-registered forced-choice human study on a random sample of 200 held-out videos, comparing TLPO against the strongest baseline separately for lip-sync accuracy, motion naturalness, and visual quality. If TLPO does not achieve a statistically significant human win rate above 50% on the claimed dimensions, the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that TLPO achieves substantial improvements in lip-sync accuracy, motion naturalness, and visual quality relative to human preferences. The load-bearing condition is that Talking-Critic scores are trustworthy human-preference proxies at every point where they are used. In the abstract, Talking-Critic is used to curate Talking-NSQ ('leveraging this model, we curate Talking-NSQ') and then the same reward model drives TLPO. Any systematic bias in the critic—e.g., a preference for exaggerated mouth motion or over-smoothed texture—is injected twice: once into the preference-pair labels and once into the optimization objective. The DPO-style loss visible in the text, L = -E[log σ(-β/2(L(x_w,t)-L(x_l,t)))], is known to over-optimize reward proxies, so the model can improve on the critic while moving away from genuine human preference. The abstract does not state whether final numbers come from independent human raters or from Talking-Critic itself, and the supplied full text is too corrupted to verify the evaluation protocol. Thus the claim of simultaneous human-preference gains on all three dimensions is unestablished.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a three-component system for audio-driven portrait animation: (i) Talking-Critic, a multimodal reward model trained to score generated videos along motion naturalness, lip-sync, and visual quality; (ii) Talking-NSQ, a claimed 410K-pair multidimensional preference dataset curated using Talking-Critic; and (iii) TLPO, a timestep-layer adaptive multi-expert preference optimization method that uses DPO-style losses to align a diffusion-based talking-head model with these preferences. The abstract claims substantial improvements over baselines in all three dimensions, with both qualitative and quantitative support. The supplied full text is largely corrupted and unreadable, so only the abstract, a few equations, figure fragments, and table shells could be inspected. The central claim of human-preference-aligned gains is therefore not currently verifiable from the manuscript as submitted.","tokens_in":15606,"tokens_out":2622,"duration_ms":33666,"significance":"If the results hold, the paper would make a useful contribution: a 410K-pair multidimensional preference dataset for talking-head generation is a potentially valuable community asset, and the idea of decoupling preference dimensions into expert LoRAs with timestep/layer-dependent fusion is technically interesting and goes beyond simple scalar-reward fine-tuning. The paper also makes a plausible case that multi-dimensional preference conflicts are a real bottleneck in this application. However, the significance is conditional on evidence that the reward model and the optimization actually track human preferences rather than artifacts of the reward model itself. The current manuscript does not supply that evidence in a readable or verifiable form.","major_comments":[{"comment":"There is a load-bearing circularity risk. The abstract states 'leveraging this model, we curate Talking-NSQ' and describes TLPO as driven by reward signals from Talking-Critic. Thus Talking-Critic supplies both the preference labels for the 410K pairs and the optimization objective. If the final evaluation also uses Talking-Critic (or if no independent human evaluation is reported), the claimed improvements in 'lip-sync accuracy, motion naturalness, and visual quality' may reflect reward over-optimization rather than genuine human preference gains. The manuscript must clarify (i) whether the preference pairs are human-annotated or critic-generated, (ii) whether any held-out human evaluation was performed, and (iii) whether the final reported numbers come from humans or from Talking-Critic. This is not a minor omission; it is central to the paper's claim of human-preference alignment.","section":"Abstract; Dataset curation and reward model sections"},{"comment":"The quantitative results are not legible in the supplied manuscript. The table shells contain only arrows and unreadable entries; no metric values, baseline names, or error bars can be seen. The abstract's assertion of 'substantial improvements' and 'superior performance' is therefore unsupported by any extractable numerical evidence. The authors need to provide a fully readable experimental section with concrete numbers, standard deviations or confidence intervals, baseline descriptions, and statistical significance tests if applicable. As submitted, the central empirical claim cannot be checked.","section":"Experiments; Tables in the experimental section"},{"comment":"The displayed loss L = -E[log σ(-β/2(L(x_w_t,t)-L(x_l_t,t)))] is a DPO-style reward-maximizing objective. Such objectives are known to over-optimize the reward proxy when the reward model is also used to curate the preference data. The manuscript gives no safeguard (e.g., KL regularization, reward-model validation against held-out human judgments, or early-stopping based on an independent metric). The paper should either provide a theoretical or empirical justification that the multi-expert decomposition prevents proxy over-optimization, or add an independent human evaluation that confirms the optimized videos are preferred on all three dimensions. Without that, the claim that TLPO 'does not interfere' across dimensions is not established.","section":"Eq. (1) in the TLPO method"}],"minor_comments":[{"comment":"The table column headers use arrows (↑/↓) without defining which direction is better or what each acronym (e.g., 'LP', 'SS', 'FID') stands for. Please add a caption that explains all metrics and arrows.","section":"Table headers and notation in experiments"},{"comment":"Several figure captions and text blocks are corrupted or contain placeholders. For example, the qualitative comparison figure labels are unreadable. The authors should ensure the final PDF renders all text correctly, since this is necessary for review and for reproducibility.","section":"Figures and captions"},{"comment":"The preference-loss temperature β, the fusion gate weights W_fuse and bias b_fuse, and the low-rank dimension k are introduced without a sensitivity analysis. At minimum, report the chosen values and provide an ablation or reference to justify them.","section":"Free parameters and hyperparameters"}],"recommendation":"major_revision","confidential_remarks":"The supplied manuscript is severely corrupted: most of the full text is unreadable, tables are empty shells, and the only fully readable element is the abstract. I cannot verify the central empirical claim in the current form. The circularity concern is real and must be addressed: Talking-Critic is used to curate the preference dataset and to provide the reward signal for TLPO, so an independent human evaluation is essential. I recommend major revision rather than rejection only because the issues are fixable if the authors can supply a clean, complete manuscript with readable tables and explicit human evaluation. If the corruption is a pipeline artifact, please provide a clean copy before any decision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I’ll get straight to it. The paper proposes three concrete things: a multimodal reward model (Talking-Critic), a 410K-pair preference dataset (Talking-NSQ), and a timestep-layer adaptive multi-expert preference optimization (TLPO). On the abstract and the partial methods I can read, that’s a real combination, not a rename of standard DPO. The multi-expert design that routes different preference dimensions (motion naturalness, lip-sync, visual quality) to separate LoRA experts and fuses them per timestep and layer is the genuinely new piece, and the dataset alone could be useful to the portrait-animation community if released.\n\nWhat the paper does well: it identifies a real problem (conflicting preference objectives), proposes a concrete solution, and claims to measure outcomes on all three axes. The writing in the abstract is clear. The fragments of the method section show a DPO-style loss with a timestep-dependent weighting, which is at least internally consistent.\n\nThe soft spot, and it is the load-bearing one: the circularity. Talking-Critic is used both to curate the preference pairs in Talking-NSQ and to supply the reward signal for TLPO. If the final evaluation also relies on Talking-Critic, then the “substantial improvements” could be the model learning to game the critic, not a real gain in human preference. The abstract doesn’t say whether the reported numbers come from independent human raters or from the critic itself. That is the first thing a referee should ask. A second, smaller concern: the abstract claims simultaneous gains on all three dimensions with “no mutual interference,” but that is a strong claim from a method that is still, as far as I can tell, a weighted sum of expert outputs. The fusion gate might introduce trade-offs that the paper needs to ablate.\n\nThe evidence, as provided, is too corrupted for me to check the experiments or the derivation. But that’s an artifact of the upload—not evidence of sloppiness. The claims are plausible, the method is not obviously wrong, and the subfield is active. The dataset, if released, would be a contribution on its own.\n\nWho is this for: anyone working on preference alignment for generative video, or on reward models for subjective video quality. It deserves a serious referee, not a desk reject. The referee’s main job: determine whether the final numbers use human raters or the model itself, and whether the dataset curation protocol prevents the critic from reinforcing its own biases.","headline":"Plausible three-part contribution that deserves a referee to check whether the reward model is curating its own preferred outputs.","tokens_in":16077,"tokens_out":2377,"would_cite":false,"duration_ms":24979,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that audio-driven portrait animation can be aligned simultaneously with human preferences for motion naturalness, lip-sync accuracy, and visual quality by decomposing the preference objective into per-dimension experts and","keywords":["audio-driven portrait animation","preference optimization","reward model","diffusion model","lip-sync accuracy","motion naturalness","visual quality","multi-expert LoRA"],"falsifier":"A head-to-head human study on held-out videos: if human raters prefer the baseline over TLPO at chance frequency, or if Talking-Critic's win predictions disagree with majority human judgment on a fresh test set, the claimed preference gains fail. A simpler check: remove the timestep-layer gate and train with a plain sum of per-dimension losses; if performance does not drop, the adaptive fusion is not the active ingredient.","tokens_in":15231,"feed_emoji":"🎬","tokens_out":4889,"duration_ms":50075,"temperature":0.7,"pith_summary":"The paper sets out to show that audio-driven portrait animation can be aligned with three human preferences at the same time—lip-sync accuracy, motion naturalness, and visual quality—even though those preferences often fight each other. It builds this on a learned reward model, Talking-Critic, which scores generated videos along the three dimensions, and on Talking-NSQ, a 410K-pair preference dataset produced with that model. The proposed training scheme, TLPO, splits the reward into per-dimension expert modules and fuses them adaptively at each diffusion timestep and network layer. The authors report that the resulting model beats baselines on all three axes in both qualitative and quantitative evaluations. A sympathetic reader would take the paper's core bet to be that decomposing preferences and recombining them at many small decision points is a better route to human-aligned generation than a single aggregated objective.","feed_headline":"One training pass improves lip-sync, motion, and visual quality","feed_subtitle":"A 410K-pair preference dataset and a per-timestep expert-fusion loss drive the gains.","key_machinery":"TLPO's timestep-layer adaptive collaborative fusion. The method maintains separate LoRA expert modules for motion naturalness, lip-sync, and visual quality. At each diffusion timestep and network layer, a small gating network computes a weighted combination of the experts' output adjustments, with weights produced by a low-rank projection of the current timestep embedding and layer index, followed by softmax. A preference-optimization loss with a winner-loser contrast is then applied to the fused output, so the gate learns both which preference matters and where in the network it should act. Talking-Critic provides the reward signal that labels winners and losers in the data.","core_discovery":"The central claim is that TLPO simultaneously improves lip-sync accuracy, motion naturalness, and visual quality in diffusion-based portrait animation by decoupling the preference signal into three specialized LoRA experts and fusing them with a timestep- and layer-adaptive gate. The gate learns, for each denoising step and each network layer, how much weight to give each dimension's expert, so that the optimization pushes the generated frame to satisfy the right preference at the right moment. This is enabled by Talking-Critic, a multimodal reward model trained on binary question-answer judgments about whether a video's motion is natural, its lips are synchronized, and its visual quality is","pith_inferences":["Editorial inference: The gating weights learned by TLPO could be inspected to identify which network layers and denoising steps dominate each preference dimension; the paper does not analyze this, but the mechanism makes it directly measurable.","Editorial inference: The Talking-NSQ dataset, if released, could become a shared benchmark for reward-model evaluation in portrait animation, extending beyond the paper's own training use.","Editorial inference: A stress test on out-of-distribution content (different languages, speaking styles, or non-human faces) would reveal whether the preference alignment generalizes or overfits to the data distribution used to train the critic."],"forward_implications":["If TLPO works as reported, portrait-animation systems can be aligned on multiple preference dimensions without the usual trade-off, because each dimension gets its own expert and a learned gate decides where to apply it.","Learned reward models like Talking-Critic can replace expensive human annotators at scale, curating large preference datasets and enabling iterative self-improvement loops.","Timestep- and layer-adaptive fusion suggests that quality dimensions are best corrected at different phases of the diffusion process, which could inform future architecture design.","The same decouple-and-fuse pattern could be applied to any diffusion model facing conflicting objectives, not only talking-head generation."],"supporting_citations":[],"fun_headline_variants":["Adaptive expert fusion improves lip-sync, motion, and visual quality","Timestep-layer adaptive gating aligns portrait animation with preferences","Per-step expert weighting lifts lip-sync, motion, and quality","410K-pair preference dataset and adaptive experts improve portrait animation","Multidimensional preference tuning refines portrait animation in one pass"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that Talking-Critic's learned scores are faithful proxies for human preferences on unseen videos and that the 410K preference pairs in Talking-NSQ are correctly labeled and representative enough to train that critic.","fun_headline_variants_meta":{"raw":{"variants":["Adaptive expert fusion improves lip-sync, motion, and visual quality","Timestep-layer adaptive gating aligns portrait animation with preferences","Per-step expert weighting lifts lip-sync, motion, and quality","410K-pair preference dataset and adaptive experts improve portrait animation","Multidimensional preference tuning refines portrait animation in one pass"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00184,"raw_usage":{"total_tokens":7089,"prompt_tokens":782,"completion_tokens":6307,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":6219}},"tokens_in":526,"tokens_out":6307,"duration_ms":48003,"temperature":1.0,"reasoning_tokens":6219,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:01:57.734687+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A head-to-head human study on held-out videos: if human raters prefer the baseline over TLPO at chance frequency, or if Talking-Critic's win predictions disagree with majority human judgment on a fresh test set, the claimed preference gains fail. A simpler check: remove the timestep-layer gate and train with a plain sum of per-dimension losses; if performance does not drop, the adaptive fusion is not the active ingredient.","supporting_citations":[],"review_version":1}