{"id":"59bab092-ada1-490d-81ca-a403f73615c9","arxiv_id":"2605.28230","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Proprio uses flow residuals from latent perturbations in frozen video generators as a self-scoring signal for physical plausibility, yielding reported gains of 16.5% on Physics-IQ and 20.6% on VideoPhy2-hard.","lead":"Proprio lets a frozen video generator score its own outputs for physical plausibility by measuring how much motion flow changes under small latent perturbations. A smart generalist might read it to see whether existing AI video models can self-correct without retraining or external judges.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Flow-residual self-scoring may track generator consistency rather than real-world physics","rationale":"The reader's weakest_assumption directly identifies the load-bearing link between the internal signal and physical reality. The reported benchmark gains and human preference are consistent with the claim but do not isolate whether the gains arise from physics or from model-specific artifacts; the proposed test would falsify or support that link without relying on the same benchmarks.","tokens_in":1736,"tokens_out":307,"duration_ms":17475,"concrete_test":"On a set of 50 synthetic videos with controlled physics violations (e.g., gravity or collision errors) generated by an external simulator and never seen in training, compute Proprio scores and compare against ground-truth physics error; if the correlation between residual magnitude and true physics error is below 0.4, the self-scoring signal does not track physical plausibility.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that smaller, more stable flow residuals under latent perturbations reliably indicate better alignment with physical dynamics (not merely the generator's training distribution). The abstract states that 'samples that are better explained by the generator's learned dynamics induce smaller and more stable residuals' and uses this for scoring/refinement, but provides no direct evidence that the signal distinguishes real physical violations from in-distribution artifacts. If the generator was trained on data containing the same biases or shortcuts that the perturbations exploit, the residual metric could improve benchmark scores without improving physical fidelity.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes Proprio, a training-free framework for frozen video generators that uses the model's own flow residuals under controlled latent perturbations as a self-scoring signal for physical plausibility. Samples better aligned with the generator's learned dynamics produce smaller, more stable residuals; these are aggregated across timesteps and perturbations, masked to motion regions, and applied via best-of-N selection or gradient refinement. Experiments on text-to-video and image-to-video benchmarks report gains over VLM-based scoring and external world-model baselines, including Physics-IQ rising from 32.2 to 37.5 and VideoPhy2-hard from 45.6 to 55.0, plus human preference for physical plausibility in roughly two-thirds of pairwise comparisons.","tokens_in":1870,"tokens_out":614,"duration_ms":32328,"significance":"If the residual signal genuinely tracks physical dynamics rather than generator-internal consistency, the work offers a practical inference-time route to more plausible video synthesis without retraining or external supervisors. The training-free design and use of internal dynamics are strengths, as are the reported benchmark lifts and human-study results. However, the absence of direct validation that residuals correlate with real physical violations (as opposed to training-distribution artifacts) limits the strength of the central claim.","major_comments":[{"comment":"Abstract: The central assertion that smaller and more stable flow residuals indicate better physical plausibility rests on the premise that the signal distinguishes real-world physics violations from in-distribution artifacts, yet no controlled test (e.g., synthetic videos with explicit gravity or collision errors) is reported to establish this correlation independent of the generator's biases.","section":"Abstract"},{"comment":"Quantitative results (abstract): Improvements such as Physics-IQ +16.5% and VideoPhy2-hard +20.6% are stated without error bars, statistical significance, or ablation tables isolating the contribution of the dynamic mask, perturbation schedule, or aggregation method; this weakens the claim of consistent outperformance.","section":"Abstract"},{"comment":"Method description (abstract and §3): The scoring signal is derived entirely from the frozen generator's internal flow dynamics under self-induced perturbations, creating a self-referential loop; the manuscript provides no external physics oracle or cross-generator transfer experiment to confirm the residual measures physical fidelity rather than model-specific consistency.","section":"§3"}],"minor_comments":[{"comment":"The abstract mentions 'dynamic spatiotemporal mask' without specifying its exact formulation or sensitivity analysis; a brief equation or pseudocode would clarify how motion-relevant regions are isolated.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about flow residuals tracking generator consistency rather than physics lands directly: the self-referential construction makes external validation essential, and its absence is the primary load-bearing gap. The manuscript fits the journal scope but would benefit from explicit discussion of this circularity risk."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback and for recognizing the training-free design and reported benchmark gains. We address each major comment below with honest assessment of the manuscript's current evidence and planned revisions. The responses focus on substance and aim to strengthen the central claims where possible.","responses":[{"response":"We agree that a controlled experiment using synthetic videos with explicit, isolated physics violations (e.g., gravity or collision errors) would provide the strongest direct evidence that residuals track physical fidelity rather than generator-specific artifacts. Our validation currently rests on Physics-IQ and VideoPhy2 benchmarks, which are constructed around physical commonsense violations, together with human preference results (approximately two-thirds favoring Proprio). We will add an explicit limitations paragraph in the revision discussing this distinction and the self-referential nature of the signal. A full synthetic-video experiment is not feasible within the current experimental budget but could be noted as future work.","revision_made":"partial","referee_comment":"[Abstract] Abstract: The central assertion that smaller and more stable flow residuals indicate better physical plausibility rests on the premise that the signal distinguishes real-world physics violations from in-distribution artifacts, yet no controlled test (e.g., synthetic videos with explicit gravity or collision errors) is reported to establish this correlation independent of the generator's biases."},{"response":"The referee is correct that the abstract numbers are presented without accompanying error bars or significance tests. The full manuscript contains component ablations, but these are not summarized in the abstract and lack statistical reporting. In the revision we will (1) add error bars and significance statements to the abstract claims, (2) expand the ablation table to isolate the dynamic mask, perturbation schedule, and aggregation choices, and (3) ensure all quantitative statements in the abstract are directly supported by the main-text results.","revision_made":"yes","referee_comment":"[Abstract] Quantitative results (abstract): Improvements such as Physics-IQ +16.5% and VideoPhy2-hard +20.6% are stated without error bars, statistical significance, or ablation tables isolating the contribution of the dynamic mask, perturbation schedule, or aggregation method; this weakens the claim of consistent outperformance."},{"response":"The self-referential design is deliberate: the method is intended to operate without any external physics oracle or additional trained models, which is the core practical advantage. Validation occurs via external benchmarks (Physics-IQ, VideoPhy2) and human raters who judge physical plausibility independently of the generator. A cross-generator transfer study would be informative but lies outside the stated scope of demonstrating utility on a single frozen model. We will revise §3 to more explicitly state this design choice and its relation to the training-free constraint.","revision_made":"no","referee_comment":"[§3] Method description (abstract and §3): The scoring signal is derived entirely from the frozen generator's internal flow dynamics under self-induced perturbations, creating a self-referential loop; the manuscript provides no external physics oracle or cross-generator transfer experiment to confirm the residual measures physical fidelity rather than model-specific consistency."}],"tokens_in":1529,"tokens_out":663,"duration_ms":38212,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that this paper gives a way to let a frozen video generator score its own outputs for physical plausibility by looking at flow residuals after small latent perturbations, then use that for best-of-N selection or gradient refinement. They add a dynamic mask to focus on motion areas and aggregate across timesteps.\n\nWhat is new is the specific framing of those residuals as a proprioception-style self-signal, plus showing it works on both text-to-video and image-to-video tasks. The paper does well on the empirical side: with TurboWan2.2 it lifts Physics-IQ from 32.2 to 37.5 and VideoPhy2-hard from 45.6 to 55.0, beats VLM-based and external world-model baselines in several settings, and gets human raters to prefer the results in about two-thirds of cases.\n\nThe soft spots are around the core claim. The method assumes smaller and more stable residuals mean the output better matches real physical dynamics, but the evidence is indirect and rests on the generator's own distribution. Nothing in the abstract directly tests whether the signal catches genuine physics violations versus in-distribution artifacts or training shortcuts. The self-referential loop the stress-test flags is real here, and without error bars, detailed ablations, or external physics checks, the quantitative gains are harder to interpret. That concern lands.\n\nThis is for people building or using video generators who need inference-time fixes for physical issues in simulation or robotics work. A reader already working on self-consistency or latent-space diagnostics would get the most out of the implementation details.\n\nIt deserves peer review because the framework is fresh, the results are concrete, and the idea is worth testing further even with the validation gaps.","headline":"Proprio shows a training-free use of flow residuals under perturbations for scoring and refining video outputs, with reported benchmark gains, but the signal may track model consistency more than real physics.","tokens_in":2355,"tokens_out":434,"would_cite":false,"duration_ms":21463,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Frozen video generators contain internal flow-residual signals that can score and refine physical plausibility at inference time.","keywords":["video generation","physical plausibility","flow residuals","latent perturbations","self-scoring","inference-time refinement","best-of-N selection","gradient refinement"],"falsifier":"A test in which videos containing clear, measurable physical violations (such as object interpenetration or gravity reversal) consistently produce smaller residuals than physically correct videos.","tokens_in":2645,"feed_emoji":"📹","tokens_out":653,"duration_ms":27690,"temperature":0.7,"pith_summary":"The paper presents Proprio, a training-free method that lets a frozen video model judge its own outputs for physical realism. It measures the model's flow residual under small controlled changes to the input latents; videos that fit the model's learned dynamics produce smaller and steadier residuals. These residuals are aggregated over time steps and perturbations, masked to moving regions, and then used either to pick the best sample from several candidates or to guide a gradient-based refinement step. The approach improves standard physical-plausibility benchmarks on both text-to-video and image-to-video tasks and beats scoring methods that rely on external vision-language models or separate world models. Human raters also prefer the Proprio-chosen or refined videos for physical correctness in roughly two-thirds of direct comparisons.","feed_headline":"Frozen video models self-score physics via flow residuals","feed_subtitle":"A training-free method uses internal residuals to select and refine outputs, lifting Physics-IQ by 16 percent on benchmarks.","key_machinery":"Flow residual under controlled latent perturbations, aggregated and masked as an internal self-scoring signal for physical alignment.","core_discovery":"Proprio treats the model's flow residual under controlled latent perturbations as a self-scoring signal. Samples that are better explained by the generator's learned dynamics induce smaller and more stable residuals. Aggregating this signal across timesteps and perturbations, focusing it on motion-relevant regions with a dynamic spatiotemporal mask, and using it for best-of-N search, gradient-based self-refinement, or both yields consistent gains in physical plausibility on text-to-video and image-to-video benchmarks.","pith_inferences":["The same residual signal could be checked on generators trained on synthetic data lacking real physics to test whether the method still selects plausible outputs.","If the signal proves architecture-independent, it might serve as a lightweight internal consistency check for other generative tasks such as 3D scene synthesis.","Applying the perturbation schedule at different noise levels could reveal whether the physical signal is strongest at particular stages of the denoising process.","The method leaves open whether the residual signal can be used to steer sampling during generation rather than only after full samples are produced."],"forward_implications":["Raises Physics-IQ from 32.2 to 37.5 and VideoPhy2-hard physical commonsense from 45.6 to 55.0 on the tested generators.","Outperforms both VLM-based scoring and external world-model baselines in several benchmark settings.","Produces videos that human raters judge more physically plausible in about two-thirds of pairwise comparisons.","Works on both text-to-video and image-to-video tasks without any additional training."],"fun_headline_variants":["Proprio self-scores video physics with flow residuals","Latent perturbations score physical plausibility internally","Video models use flow residuals for self-assessment","Internal signals refine physically plausible video outputs"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That smaller and more stable flow residuals under latent perturbations indicate better alignment with real physical dynamics rather than the generator's own training biases or artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Proprio self-scores video physics with flow residuals","Latent perturbations score physical plausibility internally","Video models use flow residuals for self-assessment","Internal signals refine physically plausible video outputs"]},"model":"grok-4.3","cost_usd":0.005037,"raw_usage":{"total_tokens":2396,"prompt_tokens":709,"num_sources_used":0,"completion_tokens":54,"cost_in_usd_ticks":50365500,"prompt_tokens_details":{"text_tokens":709,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1633,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":709,"tokens_out":54,"duration_ms":18261,"temperature":1.0,"reasoning_tokens":1633,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T12:52:26.843900+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A test in which videos containing clear, measurable physical violations (such as object interpenetration or gravity reversal) consistently produce smaller residuals than physically correct videos.","supporting_citations":[],"review_version":1}