{"id":"29c138fc-0d2f-4126-9466-09a9dc2af363","arxiv_id":"2607.28362","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Cross-shadow prediction on appearance-resampled video pairs yields a unified latent dynamics interface that transfers demonstrated actions across environments better than prior latent-action and interactive world models.","lead":"ShadowDancer learns frame-level action control for video world models from paired clips that share motion but differ in appearance. That lets any demonstrated clip become a reusable action asset replayed in new scenes without labels or fine-tuning.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Table 1 gains vs Olaf-World confound cross-shadow pairing with a source-asset stream the baseline lacks; mechanism isolation is narrower than the central claim packages.","rationale":"The reader correctly flags A1 / pair constructibility as foundational for Theorem C.2 and the scoped any-action claim, and CONDITIONAL is the right posture given synthetic pairing dependence, self-pairs not identifying (Sec. 3.4, Supp. C.4), and VLM long-rollout judging. I agree that without true shadows the identification story fails, and the paper already scopes controllability to families where pairs can be built. The softer spot in the packaged strongest claim, though, is empirical attribution: Table 1 is the main quantitative pillar for “improved action transfer,” yet it compares a full assets-conditioned system to z-only Olaf, so it does not isolate the proposed representation. Table 3’s cleaner pairing ablation is real but small-N and partial-family. That does not justify REJECT—the idea, probes, mod transfer, and controlled +pairing effect still support an accept-shaped contribution if assets-matched full-split numbers and pairing scripts hold up—but it keeps confidence moderate and the verdict CONDITIONAL rather than a clean accept. A1 remains a practical brake for real-world any-clip use; the Table 1 confound is the sharper issue for whether current evidence fully carries the causal headline.","tokens_in":23185,"tokens_out":711,"duration_ms":84513,"concrete_test":"Evaluate a matched baseline: Olaf/self-reconstruction LAM with the same source-asset conditioning routes as ShadowDancer (assets+unpaired z), on the full held-out Table 1 sets and metrics for all five families. If residual gaps to ShadowDancer shrink to a small fraction of current Table 1 margins (e.g., human +4.2 dB, robot +8.6 dB largely vanish), headline transfer gains cannot be attributed primarily to cross-shadow pairing.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim credits cross-shadow prediction on shadow pairs for beating Olaf-World on every transfer family (Table 1) and enabling reusable action assets. Sec. 4.2 explicitly keeps Olaf on its original z-only design with no source-asset stream, while ShadowDancer conditions on a=(z,s): frozen 3D-VAE source latents via input channel-concat and cross-attention (Sec. 3.3, Supp. B.2). Table 3 shows assets alone already reach 15.07 PSNR vs Olaf-recipe z-only at 12.44; the only contrast that isolates pairing—assets+unpaired z (a′) vs full (d)—is +1.43 PSNR on a 12-pair subset of two families, not the full five-family Table 1 protocol. Probes (Table 5) and z-interventions (Table 7) support that paired z carries dynamics, but the primary showcase (Table 1 / Fig. 3) is a system comparison that bundles pairing, multi-head readout, and a strong motion-detail pathway Olaf never receives. If most visible transfer fidelity is asset copying plus architecture, the causal link from cross-shadow supervision to the reported transfer margins is weaker than claimed—without overturning the pairing idea itself.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"ShadowDancer targets any-action, frame-level control of interactive video world models by treating demonstration videos as dense action specifications. The core claim is representational: standard latent-action training on self-reconstruction entangles dynamics with appearance, so the paper introduces (i) shadow pairs—frame-synchronized renders of identical dynamics under independently resampled appearance, built via a multi-source Shadow Library—and (ii) cross-shadow prediction, in which a latent-action model encodes one shadow and predicts the other, so that only shared dynamics survive in z. This z, together with frozen 3D-VAE source assets s, forms reusable action assets a=(z,s) that condition a block-causal video diffusion world model. A formal identification result (Theorem C.2) states conditions under which the minimal cross-shadow-sufficient statistic is the shared dynamics. Experiments report large transfer gains vs a matched Olaf-World baseline across five dynamics families (Table 1), component ablations (Table 3), latent probes (Supp. D), and ~86% average blinded 2AFC win rates on long rollouts against Olaf-World, Yume-1.5, and LingBot-World 2.0.","tokens_in":23519,"tokens_out":1677,"duration_ms":35607,"significance":"If the mechanism holds, the work offers a practical and conceptually clean interface for interactive world models: control by demonstration without per-family estimators, action labels, or fine-tuning, with an operational definition of controllability (a family is controllable when shadow pairs can be built). Strengths that should be credited include an explicit identification theorem with stated observability/separation assumptions (Supp. C), a source-agnostic pairing protocol spanning human motion, games, robots, and camera trajectories, factor-selective cam/dyn/full readout, and ablations plus representation probes that go beyond pure generation metrics. The limitation that real video enters only as self-pairs is stated honestly. The contribution is significant for video world models and latent-action learning if the reported transfer margins can be cleanly attributed to cross-shadow supervision rather than bundled conditioning pathways.","major_comments":[{"comment":"Sec. 4.2 and Table 1 attribute large multi-family transfer gains primarily to shadow-trained latents, but the comparison confounds cross-shadow pairing with a source-asset stream the baseline lacks. The text states Olaf-World retains its original z-only design with no source-asset pathway, while ShadowDancer conditions on a=(z,s) via channel concatenation and cross-attention (Sec. 3.3, Supp. B.2). Table 3 shows assets alone already reach 15.07 PSNR vs Olaf-recipe z-only at 12.44; the contrast that isolates pairing—assets+unpaired z (a′) vs full (d)—is only +1.43 PSNR on a 12-pair subset of two families, not the five-family Table 1 protocol. For the central claim that cross-shadow prediction yields the transferable representation, Table 1 should include a matched baseline that receives the same s pathway (and, ideally, multi-head readout), and the pairing-isolation ablation should be repo","section":"Sec. 4.2, Table 1, Table 3, Sec. 3.3"},{"comment":"The abstract and introduction market “any-action” control, while the operational scope (a family is controllable exactly when constructible shadow pairs exist) and Supp. A/C.4 make clear that identifying supervision is synthetic: real video enters only as degenerate self-pairs that do not supply the Theorem C.2 guarantee. This scoping is scientifically honest in the body but under-signaled in the title/abstract claims and in the deployment narrative (“any demonstrated clip”). Please state the synthetic-pair precondition up front when claiming any-action generality, and either quantify degradation when assets come from unpaired real/modded footage without true shadows (beyond the qualitative Fig. 5) or temper the claim to families with re-renderable dynamics.","section":"Abstract, Sec. 1, Sec. 3.4, Supp. A, Supp. C.4"},{"comment":"Long-rollout evaluation (Sec. 4.3, Table 2) is system-level 2AFC judged by a VLM (Fable 5) over three action-centric axes, with only a 20% human audit mentioned in Supp. D.3. Because interfaces differ by design (latent assets vs text/keystrokes vs camera-pose+text), the comparison measures end-to-end command survival rather than matched information. That is a valid systems question, but the reported ~86% average win rate is load-bearing for the interactive-control claim: please report inter-annotator agreement (VLM vs human audit) per axis, raw win counts, and a sensitivity check with human-only judgments on the full set or a larger audited subset. Otherwise it is hard to know how much of Table 2 is judge noise or interface mismatch versus genuine control gains.","section":"Sec. 4.3, Table 2, Supp. D.3"}],"minor_comments":[{"comment":"Fig. 2 and Sec. 3.3: the dual role of s (high-frequency motion detail vs appearance) is easy to misread as appearance leakage. A short diagram or paragraph clarifying that shadow-pair training removes the reward for copying source appearance would help.","section":"Fig. 2, Sec. 3.3"},{"comment":"Notation: d vs D, and z vs a=(z,s), shift between “unified dynamics representation” and “action asset.” Keep one term for z alone throughout Sec. 3 and the experiments.","section":"Sec. 3"},{"comment":"Table 4 mixture weights and self-pair probability (0.5 for human body, ~1/3 overall) are important free choices; a one-sentence sensitivity pointer in the main text (even if details stay in Supp. D.7) would aid reproducibility.","section":"Sec. 3.4, Supp. D.1–D.2"},{"comment":"Typos/formatting: “Shadow Dancer-1.github.io” spacing in the abstract; “3D-V AE” broken across lines; Fréchet rendered as “Fr´echet” in places; “Fable 5” should be identified more clearly as the judge model.","section":"Abstract, Sec. 4"},{"comment":"Related work could briefly contrast with multi-view/invariance SSL citations already in Supp. C ([22],[54]) in the main Sec. 2, since the pairing-as-task-definition idea is central.","section":"Sec. 2"}],"recommendation":"major_revision","confidential_remarks":"The pairing idea and formalization are genuinely interesting and above incremental LAM work. My major_revision stance is driven by attribution (Table 1 vs assets confound) and evaluation reliability (VLM 2AFC), not by disbelief in the method. If the authors add a matched assets-enabled Olaf baseline on the full transfer suite and tighten the any-action scoping, this could clear a high bar. Fit for a top CV/ML venue is reasonable given interactive world-model interest; watch that the project page and “any action” framing do not outrun the synthetic-pair precondition in camera-ready text."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: the load-bearing idea is changing the data—renderer-side shadow pairs plus cross-shadow prediction—so the latent has to carry shared dynamics rather than co-occurring appearance. That is a clean operational fix for the non-identifiability problem latent-action models keep hitting, and they actually build the multi-source library and ship a block-causal world model on top of it.\n\nWhat is new is not “latent actions” or “interactive world models,” both of which they cite honestly. It is treating pairing as the definition of the controllable factor, implementing it across human, robot, game, and camera sources, and turning clips into variable-length action assets with no labels or estimators. Table 1 vs a matched Olaf-World backbone shows large, consistent transfer gains across five families; latent probes and FVD in the supplement push back on the “blurry win” worry; the appendix theorem states observability/separation assumptions and admits self-pairs do not identify. That is serious work.\n\nSoft spots, in proportion. The stress-test note is partly right: Sec. 4.2 keeps Olaf z-only while ShadowDancer conditions on a=(z,s), so the headline transfer margins are a system comparison, not a pure pairing ablation. Table 3 is the real isolation—assets alone already help a lot; pairing on top of assets is about +1.4 PSNR on a small two-family subset. That does not kill the idea (probes and z-interventions still say paired z carries dynamics), but the paper packages more credit onto cross-shadow supervision than the primary showcase strictly isolates. Long-rollout wins (~86%) rest on VLM 2AFC across heterogeneous interfaces; useful, not definitive. Real video only enters as self-pairs, so “any action” is honestly “any family you can re-render.” No code/data release is a practical brake.\n\nWho it is for: people building controllable video world models, LAMs, or sim-to-real action interfaces. Citation pattern looks fair; free parameters are ordinary training knobs, not hidden degrees of freedom. I would send this to referees. Engage with it—especially if they release pairing scripts—and push them to report full-protocol ablations with assets held fixed.","headline":"Solid systems+representation paper: shadow pairs make dynamics identifiable by construction, with real transfer gains—but Table 1 bundles pairing with a source-asset pathway the baseline never gets.","tokens_in":24159,"tokens_out":572,"would_cite":true,"duration_ms":14635,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Any demonstrated clip can become a reusable, frame-level action that transfers to new scenes once the same motion is seen under two different appearances.","keywords":["interactive video world models","latent actions","shadow pairs","cross-shadow prediction","dynamics representation","action transfer","block-causal video generation","demonstration-driven control"],"falsifier":"Hold out true shadow pairs for a family, extract the latent from one video, generate from the first frame of the other, and check whether transfer metrics (reconstruction or trajectory error) and blinded rollout preference collapse toward ordinary self-reconstruction latent-action baselines; if they do, the identifying claim of the pairing fails.","tokens_in":24017,"feed_emoji":"🎬","tokens_out":917,"duration_ms":19149,"temperature":0.7,"pith_summary":"Interactive video world models can explore virtual worlds but still lack a single way to specify how any kind of action should unfold frame by frame. Text and key presses are too loose; 3D tracks and pose streams are precise but family-specific and hard to get. Demonstration videos look like the answer, yet a video always shows motion through one particular look, so latents trained by reconstructing that clip entangle action with appearance and fail to transfer. ShadowDancer’s claim is that the fix is in the data, not another penalty: build shadow pairs that replay identical dynamics under independently resampled appearance, then train by predicting one video from the other. Whatever the pair resamples is discarded by construction; whatever it preserves becomes a unified dynamics latent. That latent turns any clip into a variable-length action asset that can be stored, composed, and replayed in new environments without labels, estimators, or fine-tuning, and the paper reports clear gains on transfer and long rollouts across human motion, games, robots, and camera control.","feed_headline":"Show a clip once; replay its motion in any new scene","feed_subtitle":"Paired renders of the same dynamics teach a latent that transfers actions without labels or fine-tuning","key_machinery":"Shadow pairs plus cross-shadow prediction: two frame-synchronized videos of the same dynamics under independently resampled appearance, with an encoder reading one and a decoder predicting the other, so invariance is a property of the supervision rather than a regularizer; the resulting latent conditions a block-causal video world model as reusable action assets.","core_discovery":"Cross-shadow prediction on shadow pairs yields a unified dynamics representation: by construction the latent keeps only what two independently appeared renders of the same motion share, so any demonstrated clip becomes a reusable, variable-length action asset that drives frame-level control in new environments without action labels, motion estimators, or fine-tuning.","pith_inferences":["If video editors or generators can manufacture faithful appearance-changed shadows of real footage, the same identifying signal could extend beyond engines into ordinary video.","Asset libraries may scale more with demonstration and generative coverage than with new end-to-end training runs, changing how interactive worlds are extended after deployment.","The same pairing protocol is a general recipe for making any chosen factor the controllable one, not only motion—suggesting analogous ‘shadow’ constructions for other entangled generative factors."],"forward_implications":["A dynamics family becomes controllable through one interface exactly when shadow pairs can be constructed for it—no new vocabulary, labels, or per-family estimators.","Any post-training demonstration can join the action library at the cost of one frozen encoder pass and be replayed or chained in new scenes.","The same reference clip can be read as pure camera, pure scene/body/arm motion, or both, by how pairs were built and which readout head is used.","Interactive world models can be driven by showing rather than telling, across entertainment and simulation-based training of embodied agents."],"fun_headline_variants":["Replay any clip’s motion in a new scene via shadow pairs","Cross-shadow prediction turns one demo into reusable action","Unified dynamics latent from paired shadows transfers actions","Same motion, new looks: learn actions without labels","Shadow pairs make demonstrated clips drive any-scene control"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"For each action family you care about, you must be able to build (or closely approximate) pairs that truly share the same frame-by-frame dynamics while fully resampling how they look—something easy in games and simulators but hard for real-world video.","fun_headline_variants_meta":{"raw":{"variants":["Replay any clip’s motion in a new scene via shadow pairs","Cross-shadow prediction turns one demo into reusable action","Unified dynamics latent from paired shadows transfers actions","Same motion, new looks: learn actions without labels","Shadow pairs make demonstrated clips drive any-scene control"]},"model":"grok-4.5","effort":"low","cost_usd":0.002549,"raw_usage":{"total_tokens":1048,"prompt_tokens":826,"num_sources_used":0,"completion_tokens":59,"cost_in_usd_ticks":25488000,"prompt_tokens_details":{"text_tokens":826,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":163,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":826,"tokens_out":59,"duration_ms":3646,"temperature":1.0,"reasoning_tokens":163,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T09:39:42.344298+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Hold out true shadow pairs for a family, extract the latent from one video, generate from the first frame of the other, and check whether transfer metrics (reconstruction or trajectory error) and blinded rollout preference collapse toward ordinary self-reconstruction latent-action baselines; if they do, the identifying claim of the pairing fails.","supporting_citations":[],"review_version":1}