REVIEW 4 major objections 42 references
Debiasing latent actions from unlabeled video makes robot world models follow commands with far less labeled data.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 04:49 UTC pith:VSSITAYM
load-bearing objection Solid empirical fix for confounded latent actions in LAM-based world models; the efficiency and intervention results are real, even if the causal packaging and SAM3-shared metrics overclaim a bit. the 4 major comments →
Causally Debiased Latent Action Model for Embodied Action Conditioned World Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Reconstruction-only latent actions entangle embodiment dynamics with action-irrelevant visual factors, confounding the downstream world model; three short fine-tuning objectives that force embodiment focus, action-aware neighborhoods, and calibrated non-collapse produce debiased latents that measurably improve action following, visual fidelity, and robot-action adaptation efficiency on both 2B and 14B action-conditioned world models.
What carries the argument
CD-LAM: a three-objective LAM fine-tuning loss (embodiment-centric weighted reconstruction, action-centric contrastive learning over coarse verb primitives, and latent-space calibration with free-bit KL plus zero-transition anchoring) applied in a three-stage pipeline that first debiases the latent action, then debiases the world model on those latents, then bridges executable robot actions into the same space.
Load-bearing premise
The method assumes that automatic embodiment masks and coarse caption-verb clusters are faithful enough proxies for true action factors that the three losses remove confounding rather than merely re-weighting toward another correlated visual cue.
What would settle it
Hold the world-model architecture fixed, swap only the latent-action encoder for a reconstruction-only baseline, and check whether mean foreground displacement error under identical robot-action sequences still drops by roughly thirty percent and whether zero-action and camera-shift diagnostics remain low; if the gains vanish or the diagnostics stay high, the causal-debiasing claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that reconstruction-only latent action models (LAMs) encode action-irrelevant confounders (background, non-interacted objects, camera-like factors) into the latent action z_t, which then confounds action-conditioned world models (ACWMs). It proposes CD-LAM, a three-stage fine-tuning pipeline whose Stage-1 LAM objectives—embodiment-centric weighted reconstruction (Eqs. 8–9), action-centric contrastive learning over 12-way caption-verb primitives (Eq. 10), and latent-space calibration via free-bit KL plus zero-transition anchoring (Eqs. 11–12)—produce debiased latents. On 2B and 14B DreamDojo-style ACWMs, the method reports lower FDCE after latent-action and robot-action conditioning, higher PSNR/SSIM, stronger zero-action and target-action interventions, and matching of the DreamDojo 50k-update reference with more than 12× fewer robot-action adaptation steps (final 3k/6k checkpoints). Supporting evidence includes a LAM confounding audit (Table I), multi-stage rollout tables (II–III), data-tier scaling (Table IV), objective ablations (Table V), and qualitative rollouts.
Significance. If the results hold under independent scrutiny, the work is a practically useful contribution to embodied world models: it isolates a concrete failure mode of reconstruction-trained LAMs, supplies diagnostic metrics (zero-transition response, camera-shift response, shortcut leakage, FDCE), and shows that a short, targeted LAM fine-tune can improve controllability and cut robot-action adaptation cost by more than an order of magnitude at both 2B and 14B. The multi-stage evaluation design (latent-only rollouts, robot-action adaptation, zero-action and target-transfer interventions), objective ablations that map each loss to a distinct failure mode, and the promised release of debiased LAMs/ACWMs, protocols, and code are genuine strengths. The efficiency finding—that debiasing the condition is cheaper than unlearning a confounded condition downstream—is of clear engineering value for robot world-model pipelines that rely on unlabeled video pretraining.
major comments (4)
- The primary action-following metric and a core training signal share the same tooling. Embodiment-centric reconstruction reweights pixels by SAM3 masks M_t (Eqs. 8–9); FDCE seeds tracks inside SAM3 foreground masks and scores only those tracks (Appendix A, Eq. A.4). Action-centric contrast and the shortcut-leakage diagnostic further share the 12-way caption-verb clusters (Appendix B). A model that concentrates capacity on SAM3-selected regions and verb-cluster neighborhoods can therefore improve headline FDCE and Table I diagnostics without necessarily purifying the causal factor A_t from C_t/V_t. Zero-action residual FDCE and full-frame PSNR provide partially independent evidence, but the claimed 35%/30% FDCE reductions and the “causally debiased” framing rest on a metric aligned with the training signal. Please add at least one mask-independent motion metric (e.g., full-frame or random
- Tables II and III (and the efficiency curves in Fig. 8) report point estimates only—no standard errors, bootstrap intervals, or multi-seed variance—despite multi-scale claims and percentage reductions that are central to the abstract. With 300 evaluation clips, seed or clip-level variability is estimable. Without it, it is hard to judge whether the 2B→14B baseline FDCE worsening, the 12× efficiency claim, or the per-action breakdowns in Fig. A.1 are stable. Please report uncertainty for the main FDCE/PSNR numbers and for the step at which CD-LAM crosses the DreamDojo reference.
- The causal analysis in §III (Eq. 5, Fig. 1d) and the title/abstract language (“causally debiased,” “confounding path”) go beyond what the experiments strictly establish. The interventions show improved sensitivity to the supplied action, and Table I shows reduced responses to static pairs and synthetic shifts, but there is no identification argument or interventional test that separates removal of C_t/V_t from reweighting toward correlated embodiment appearance. Table V’s footnote already notes that removing zero-transition calibration can lower FDCE while failing the camera-shift diagnostic—evidence that FDCE alone does not certify causal purity. Soften or operationalize the causal claims (e.g., “reduces measured action-irrelevant responses and improves action following under fixed context”) unless additional identification-style evidence is added.
- Comparisons are limited to DreamDojo with its original reconstruction-trained LAM (§V-A). The related-work section cites Genie, LAPO/LAPA, AdaWorld, Moto, IGOR, and ConLA, but none appear as empirical baselines for the LAM audit or for Stage-2/3 rollouts. At minimum, a reconstruction-only LAM re-finetuned for the same 1k steps without the three CD-LAM terms (or with only L_emb) should be reported as a compute-matched control beyond the partial ablations in Table V, so that gains are not confounded with extra Stage-1/2 fine-tuning budget alone.
Circularity Check
No derivation-level circularity; only mild training–diagnostic alignment on Stage-1 audits, while held-out FDCE/efficiency claims remain empirical.
specific steps
-
other
[Sec. IV-B3 Eq. (12); Table I zero-transition diagnostic; App. A Eq. (A.5)]
"L_zero = E_ot [ ( [ ∥z0_t∥2 / sg(s_Δ)+ε − m_zero ]_+ )^2 ]. ... Zero-transition response (static pair (o_t, o_t); rel. norm ↓) Median response 0.527 → 0.043"
L_zero is defined to push the norm of duplicated-frame latents below a margin times ordinary-transition RMS. Table I’s primary zero-transition audit is that same relative norm. Reporting a large drop after training with L_zero is measuring the optimized objective, not an independent causal prediction of purified A_t. Downstream FDCE claims do not inherit this reduction by construction.
-
other
[Sec. IV-B1 Eqs. (8)–(9); App. A FDCE definition; Sec. V-A Metrics]
"Let M_t ∈ [0,1]^{h×w} be the embodiment–object foreground mask obtained by SAM3... W_t = α_fg M_t + α_bg(1−M_t)... Foreground masks select embodiment and interacted-object regions using SAM3, and point tracks are computed only within valid foreground regions"
Embodiment-centric reconstruction reweights pixels by SAM3 masks; FDCE seeds and scores tracks inside SAM3 foreground masks. A model that concentrates capacity on SAM3-selected regions can improve FDCE partly by aligning to the same proxy used in training, so the causal reading ‘purified A_t vs C_t/V_t’ is not fully independent of the evaluation tooling. This is metric–training alignment, not a definitional identity of the headline efficiency claim.
-
other
[Sec. IV-B2 Eq. (10); Table I shortcut leakage; App. A Eq. (A.6)]
"L_ctr = (1/|P|) Σ softplus(−y_ij(τ v_i^⊤ v_j + b)) ... Shortcut leakage 0.151 → 0.014 ... L_shortcut = E[cos | same episode, diff. primitive] − E[cos | diff. episode, same primitive]"
Action-centric contrast pulls same-primitive pairs together and pushes different-primitive pairs apart using the 12-way caption-verb clusters. Shortcut leakage is exactly the cosine gap between those two pair types. Improving Table I’s shortcut-leakage number is therefore largely the direct effect of L_ctr, not an external test that the latent encodes A_t rather than verb-correlated context. Again, this does not force the held-out robot-action FDCE or adaptation-efficiency results.
full rationale
CD-LAM is an empirical methods paper: causal analysis motivates three fine-tuning losses, which are then evaluated on held-out EgoDex/AgiBot rollouts and action-replacement interventions. There is no uniqueness theorem, no self-citation load-bearing premise, no fitted constant renamed as a prediction of an independent quantity, and no first-principles derivation that reduces to its inputs by construction. The central claims (≈30–35% FDCE drop after robot-action adaptation, PSNR gains, >12× fewer adaptation updates) are measured outcomes on external robot data under fixed observation context, not algebraic identities. Mild circularity risk is confined to Stage-1 LAM audits that re-measure quantities the losses directly optimize (zero-transition norm under L_zero; primitive-neighborhood structure under L_ctr) and to shared SAM3 tooling between embodiment-centric training weights and FDCE foreground selection. Those are validity/interpretation caveats, not forced reductions of the main system-level results. Score 2 reflects that minor diagnostic–objective alignment only.
Axiom & Free-Parameter Ledger
free parameters (7)
- λ_ctr(k) contrastive weight schedule
- λ_cal calibration weight
- α_fg / α_bg foreground–background reconstruction weights
- m_zero zero-transition margin
- contrastive temperature τ and bias b
- free-bits KL floor (per-dimension)
- Stage-1/2/3 step budgets (1k / 2k / 3k–6k) and data-tier hours
axioms (5)
- domain assumption Next-frame reconstruction sufficiency alone admits action-irrelevant factors into z_t (Eq. 5 factorization into A_t, C_t, V_t).
- domain assumption SAM3 masks identify embodiment and interacted-object regions that should dominate the action latent.
- domain assumption 12-way caption-verb clusters are a valid coarse action-primitive space for contrastive structure without executable actions.
- domain assumption A lightweight MLP bridge g_η can map executable robot actions into the debiased latent space without reintroducing the original confounders.
- standard math Standard conditional video / diffusion training losses and latent-action encoder–decoder form (Eqs. 1–4).
invented entities (3)
-
CD-LAM three-objective package (L_emb + L_ctr + L_cal)
no independent evidence
-
FDCE (Foreground Displacement Chamfer Error)
no independent evidence
-
LAM confounding audit suite (zero-transition response, camera-shift response, shortcut leakage)
no independent evidence
read the original abstract
Action-conditioned world models (ACWMs) aim to simulate future observations conditioned on embodied actions, offering a promising foundation for robot planning, policy evaluation, and data augmentation. However, learning controllable ACWMs requires large-scale action-labeled data, which remains costly to collect in the real world. Latent action models (LAMs) mitigate this bottleneck by inferring latent actions from unlabeled videos, but existing LAMs are typically trained with reconstruction-only objectives and therefore entangle action-relevant dynamics with action-irrelevant visual factors such as backgrounds and untouched objects. In this work, we identify this action-irrelevant bias as a key obstacle to controllable ACWMs and introduce evaluation metrics to measure latent-action bias, action following, and robustness. We propose CD-LAM, a causally debiased framework for LAM-based ACWMs. CD-LAM introduces three efficient fine-tuning objectives: embodiment-centric reconstruction, action-centric contrastive learning, and latent space calibration, which together encourage embodiment-focused, action-aware, and calibrated non-collapsed latent action representations. Experiments on 2B and 14B ACWM backbones show that CD-LAM substantially improves latent-action controllability, downstream robot-action following, visual fidelity, and adaptation efficiency, requiring only 6k fine-tuning steps and more than 12$\times$ fewer robot-action adaptation updates than the baseline.
Figures
Reference graph
Works this paper leans on
-
[1]
D. Ha and J. Schmidhuber, “World models,”arXiv preprint arXiv:1803.10122, 2018
Pith/arXiv arXiv 2018
-
[2]
Mastering diverse do- mains through world models,
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap, “Mastering diverse do- mains through world models,”arXiv preprint arXiv:2301.04104, 2023
Pith/arXiv arXiv 2023
-
[3]
ACWM-Phys: Investigating generalized physical interaction in action- conditioned video world models,
H. Xue, Y . Chen, L. Ma, Z. Zhao, L. Moukheiber, Y . Zhu, and Y . Chen, “ACWM-Phys: Investigating generalized physical interaction in action- conditioned video world models,”arXiv preprint arXiv:2605.08567, 2026
Pith/arXiv arXiv 2026
-
[4]
iVideoGPT: Interactive VideoGPTs are scalable world models,
J. Wu, S. Yin, N. Feng, X. He, D. Li, J. Hao, and M. Long, “iVideoGPT: Interactive VideoGPTs are scalable world models,” inAdvances in Neu- ral Information Processing Systems (NeurIPS), 2024, arXiv:2405.15223
Pith/arXiv arXiv 2024
-
[5]
Open X-Embodiment: Robotic learning datasets and RT-X models,
Open X-Embodiment Collaboration, “Open X-Embodiment: Robotic learning datasets and RT-X models,”arXiv preprint arXiv:2310.08864, 2023
Pith/arXiv arXiv 2023
-
[6]
DROID: A large-scale in-the-wild robot manipulation dataset,
A. Khazatsky, K. Pertschet al., “DROID: A large-scale in-the-wild robot manipulation dataset,” inRobotics: Science and Systems (RSS), 2024
2024
-
[7]
Ego4D: Around the world in 3,000 hours of egocentric video,
K. Grauman, A. Westbury, E. Byrneet al., “Ego4D: Around the world in 3,000 hours of egocentric video,” inIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022
2022
-
[8]
Ego-Exo4D: Understand- ing skilled human activity from first- and third-person perspectives,
K. Grauman, A. Westbury, L. Torresaniet al., “Ego-Exo4D: Understand- ing skilled human activity from first- and third-person perspectives,” inIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[9]
Scaling egocentric vision: The EPIC-KITCHENS dataset,
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kaza- kos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray, “Scaling egocentric vision: The EPIC-KITCHENS dataset,”European Conference on Computer Vision (ECCV), 2018
2018
-
[10]
Genie: Generative interactive environ- ments,
J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y . Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, Y . Aytar, S. Bech- tle, F. Behbahani, S. Chan, N. Heess, L. Gonzalez, S. Osindero, S. Ozair, S. Reed, J. Zhang, K. Zolna, J. Clune, N. de Freitas, S. Singh, and T. Rocktäschel, “Genie: Generative interactive environ- ments,” inInternatio...
Pith/arXiv arXiv 2024
-
[11]
Learning to act without actions,
D. Schmidt and M. Jiang, “Learning to act without actions,” inInter- national Conference on Learning Representations (ICLR), 2024, lAPO; arXiv:2312.10812
Pith/arXiv arXiv 2024
-
[12]
Latent action pretraining from videos,
S. Ye, J. Jang, B. Jeon, S. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y .-W. Chao, B. Y . Lin, L. Liden, K. Lee, J. Gao, L. Zettlemoyer, D. Fox, and M. Seo, “Latent action pretraining from videos,” inInternational Conference on Learning Representations (ICLR), 2025
2025
-
[13]
AdaWorld: Learning adaptable world models with latent actions,
S. Gao, S. Zhou, Y . Du, J. Zhang, and C. Gan, “AdaWorld: Learning adaptable world models with latent actions,” inInternational Conference on Machine Learning (ICML), 2025
2025
-
[14]
DreamDojo: A generalist robot world model from large-scale human videos,
S. Gao, W. Liang, K. Zheng, A. Malik, S. Ye, S. Yu, W.-C. Tseng, Y . Dong, K. Mo, C.-H. Lin, Q. Ma, S. Nah, L. Magne, J. Xiang, Y . Xie, R. Zheng, D. Niu, Y . L. Tan, K. R. Zentner, G. Kurian, S. Indupuru, P. Jannaty, J. Gu, J. Zhang, J. Malik, P. Abbeel, M.-Y . Liu, Y . Zhu, J. Jang, and L. Fan, “DreamDojo: A generalist robot world model from large-scale...
Pith/arXiv arXiv 2026
-
[15]
What do latent action models actually learn?
C. Zhang, T. Pearce, P. Zhang, K. Wang, X. Chen, W. Shen, L. Zhao, and J. Bian, “What do latent action models actually learn?” inAdvances in Neural Information Processing Systems (NeurIPS), 2025
2025
-
[16]
Latent action learning requires supervision in the presence of distractors,
A. Nikulin, I. Zisman, D. Tarasov, N. Lyubaykin, A. Polubarov, I. Kise- lev, and V . Kurenkov, “Latent action learning requires supervision in the presence of distractors,” inInternational Conference on Machine Learning (ICML), 2025
2025
-
[17]
ConLA: Contrastive latent action learning from human videos for robotic manipulation,
W. Dai, K. Lan, J. Zhou, B. Zhao, X. Su, J. Tong, W. Guan, and S. Yang, “ConLA: Contrastive latent action learning from human videos for robotic manipulation,”arXiv preprint arXiv:2602.00557, 2026
arXiv 2026
-
[18]
AgiBot-World-Contributors, Q. Bu, J. Cai, L. Chen, X. Cui, Y . Ding, S. Feng, S. Gao, X. He, X. Hu, X. Huang, S. Jiang, Y . Jiang, C. Jing, H. Li, J. Li, C. Liu, Y . Liu, Y . Lu, J. Luo, P. Luo, Y . Mu, Y . Niu, Y . Pan, J. Pang, Y . Qiao, G. Ren, C. Ruan, J. Shan, Y . Shen, C. Shi, M. Shi, M. Shi, C. Sima, J. Song, H. Wang, W. Wang, D. Wei, C. Xie, G. Xu...
Pith/arXiv arXiv 2025
-
[19]
SAM 3: Segment anything with concepts,
N. Carion, L. Gustafson, Y .-T. Hu, S. Debnath, R. Hu, D. Suris, C. Ryali, K. V . Alwala, H. Khedr, A. Huanget al., “SAM 3: Segment anything with concepts,”arXiv preprint arXiv:2511.16719, 2025
Pith/arXiv arXiv 2025
-
[20]
Sigmoid loss for language image pre-training,
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” inIEEE/CVF International Conference on Computer Vision (ICCV), 2023
2023
-
[21]
EgoDex: Learning dexterous manipulation from large-scale egocentric video,
R. Hoque, P. Huang, D. J. Yoon, M. Sivapurapu, and J. Zhang, “EgoDex: Learning dexterous manipulation from large-scale egocentric video,” 2025. [Online]. Available: https://arxiv.org/abs/2505.11709
Pith/arXiv arXiv 2025
-
[22]
CoWTracker: Tracking by warping instead of correlation,
Z. Lai, E. Insafutdinov, E. Sucar, and A. Vedaldi, “CoWTracker: Tracking by warping instead of correlation,” 2026. [Online]. Available: https://arxiv.org/abs/2602.04877
arXiv 2026
-
[23]
Imitating la- tent policies from observation,
A. D. Edwards, H. Sahni, Y . Schroecker, and C. L. Isbell, “Imitating la- tent policies from observation,” inInternational Conference on Machine Learning (ICML), 2019, iLPO; arXiv:1805.07914
Pith/arXiv arXiv 2019
-
[24]
Video PreTraining (VPT): Learning to act by watching unlabeled online videos,
B. Baker, I. Akkaya, P. Zhokhov, J. Huizinga, J. Tang, A. Ecof- fet, B. Houghton, R. Sampedro, and J. Clune, “Video PreTraining (VPT): Learning to act by watching unlabeled online videos,” in Advances in Neural Information Processing Systems (NeurIPS), 2022, arXiv:2206.11795
Pith/arXiv arXiv 2022
-
[25]
Moto: Latent motion token as the bridging language for learning robot manipulation from videos,
Y . Chen, Y . Ge, W. Tang, Y . Li, Y . Ge, M. Ding, Y . Shan, and X. Liu, “Moto: Latent motion token as the bridging language for learning robot manipulation from videos,” inIEEE/CVF International Conference on Computer Vision (ICCV), 2025
2025
-
[26]
X. Chen, J. Guo, T. He, C. Zhang, P. Zhang, D. C. Yang, L. Zhao, and J. Bian, “IGOR: Image-GOal representations are the atomic con- trol units for foundation models in embodied AI,”arXiv preprint arXiv:2411.00785, 2024
Pith/arXiv arXiv 2024
-
[27]
Learning additively compositional latent actions for embodied AI,
H. Wei, X. Chen, C. Zhang, T. Pearce, J. Chen, A. Lamb, L. Zhao, and J. Bian, “Learning additively compositional latent actions for embodied AI,”arXiv preprint arXiv:2604.03340, 2026
Pith/arXiv arXiv 2026
-
[28]
Co- evolving latent action world models,
Y . Wang, F. Zhang, D.-C. Zhan, L. Zhao, K. Wang, and J. Bian, “Co- evolving latent action world models,”arXiv preprint arXiv:2510.26433, 2025
Pith/arXiv arXiv 2025
-
[29]
MVP-LAM: Learning action-centric latent action via cross-viewpoint reconstruction,
J. M. Lee, D. Lee, S. Ju, T. Cho, J. W. Koo, L. Zhao, S. Hong, and J. Lee, “MVP-LAM: Learning action-centric latent action via cross-viewpoint reconstruction,”arXiv preprint arXiv:2602.03668, 2026
Pith/arXiv arXiv 2026
-
[30]
Masked autoencoders are scalable vision learners,
K. He, X. Chen, S. Xie, Y . Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” inIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022
2022
-
[31]
Learning interactive real-world simula- tors,
M. Yang, Y . Du, S. K. S. Ghasemipour, J. Tompson, L. P. Kaelbling, D. Schuurmans, and P. Abbeel, “Learning interactive real-world simula- tors,” inInternational Conference on Learning Representations (ICLR), 2024, uniSim; arXiv:2310.06114
Pith/arXiv arXiv 2024
-
[32]
Cosmos world foundation model platform for physical AI,
NVIDIA, “Cosmos world foundation model platform for physical AI,” arXiv preprint arXiv:2501.03575, 2025
Pith/arXiv arXiv 2025
-
[33]
Towards accurate generative models of video: A new metric and challenges,
T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly, “Towards accurate generative models of video: A new metric and challenges,”arXiv preprint arXiv:1812.01717, 2018, fVD
Pith/arXiv arXiv 2018
-
[34]
MotionPro: A precise motion controller for image-to-video generation,
Z. Zhang, F. Long, Z. Qiu, Y . Pan, W. Liu, T. Yao, and T. Mei, “MotionPro: A precise motion controller for image-to-video generation,” inIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
2025
-
[35]
TAP-Vid: A benchmark for tracking any point in a video,
C. Doersch, A. Gupta, L. Markeeva, A. Recasens, L. Smaira, Y . Aytar, J. a. Carreira, A. Zisserman, and Y . Yang, “TAP-Vid: A benchmark for tracking any point in a video,” inAdvances in Neural Information Processing Systems (NeurIPS), 2022, arXiv:2211.03726
Pith/arXiv arXiv 2022
-
[36]
TAPIR: Tracking any point with per-frame initialization and temporal refinement,
C. Doersch, Y . Yang, M. Vecerik, D. Gokay, A. Gupta, Y . Aytar, J. a. Carreira, and A. Zisserman, “TAPIR: Tracking any point with per-frame initialization and temporal refinement,” inIEEE/CVF International Con- ference on Computer Vision (ICCV), 2023, arXiv:2306.08637
Pith/arXiv arXiv 2023
-
[37]
CoTracker: It is better to track together,
N. Karaev, I. Rocco, B. Graham, N. Neverova, A. Vedaldi, and C. Rupprecht, “CoTracker: It is better to track together,” inEuropean Conference on Computer Vision (ECCV), 2024
2024
-
[38]
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, P. Dollár, and R. Gir- shick, “Segment anything,” inIEEE/CVF International Conference on Computer Vision (ICCV), 2023, sAM; arXiv:2304.02643
Pith/arXiv arXiv 2023
-
[39]
SAM 2: Segment anything in images and videos,
N. Ravi, V . Gabeur, Y .-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafsonet al., “SAM 2: Segment anything in images and videos,”arXiv preprint arXiv:2408.00714, 2024
Pith/arXiv arXiv 2024
-
[40]
M. Arjovsky, L. Bottou, I. Gulrajani, and D. Lopez-Paz, “Invariant risk minimization,”arXiv preprint arXiv:1907.02893, 2019
Pith/arXiv arXiv 1907
-
[41]
Shortcut learning in deep neural networks,
R. Geirhos, J.-H. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann, “Shortcut learning in deep neural networks,”Nature Machine Intelligence, vol. 2, pp. 665–673, 2020
2020
-
[42]
Causal confusion in imita- tion learning,
P. de Haan, D. Jayaraman, and S. Levine, “Causal confusion in imita- tion learning,” inAdvances in Neural Information Processing Systems (NeurIPS), 2019. APPENDIXA METRICDETAILS PSNR Reporting.We report PSNR as the visual-fidelity metric, computed on full frames in dB. For images normalized to[0,1], PSNR(x,ˆx) = 10 log10 1 MSE(x,ˆx).(A.1) Fig. 1(b) report...
2019
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.