Pith. sign in

REVIEW 3 major objections 5 minor 115 references

Visual General Intelligence: A White Paper

T0 review · 3 major / 5 minor · reviewed 2026-08-27 · deepseek-v4-flash

Pith's one-line read Visual experience, not language, may be the path to general intelligence, and scaled generative video models are this white paper's leading candidate route.

desk verdict A well-written, honest multi-perspective white paper on vision-first AI; no new results, but a useful map of current bets and a clear statement of the field's central unresolved tension. read the letter →

arxiv 2608.25924 v1 pith:LZIZQCGG submitted 2026-08-26 cs.CV

classification cs.CV
keywords visualgeneralintelligencevisionfoundationmodelsgenerativevideoworldembodiedspatialAIcontinuallearninganalysisbysynthesis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Visual general intelligence (VGI), as this white paper frames it, is the thesis that intelligence can emerge from visual experience alone, not just from language modeling. The contributors decline to settle on a single definition, but they share one premise: vision should not be reduced to an input channel for a language-centered system. The paper's most concrete bet is that training large models on video with generative objectives will reproduce the language-model story—broad zero-shot transfer, in-context learning, and ultimately general reasoning about objects, geometry, and physics. Its synthetic proposal is that sequential prediction, open-ended generation, and reconstruction are three complementary objectives whose convergence could carry vision foundation models toward general intelligence. If the agenda is right, the field should redirect resources toward scaling generative video, persistent spatial world models, embodied interaction, and benchmarks that measure transfer, creativity, and physical consistency rather than static task accuracy.

What carries the argument

The mechanism that carries the argument is the generative visual objective at scale, with video as the general case. The paper's argument is that generating visual data is hard in the right way: a classifier can pass by exploiting texture or background, but a generative model is punished unless it accounts for the object, the background, the lighting, and the motion together. On top of this, the paper's named mechanism for full VGI is the convergence of three learning objectives: sequential prediction (anticipating the next visual state), open-ended generation (imagining what could exist), and reconstruction (recovering hidden structure from partial observations). The surrounding mechanism is analysis by synthesis—treating perception as inference over latent scene structure—which grounds the reconstruction objective in a long visual-computing tradition.

What would settle it

Probe a pixel-trained video model for recoverable physical parameters. Concretely: freeze the model, then attempt to estimate mass, friction, and contact states from its internal features or from its generated rollouts; use those estimates to predict the outcome of a novel physical intervention, such as tilting a table or adding a push, and compare against real video. If the model produces visually flawless videos yet fails intervention-based predictions that appearance-only baselines also fail, the central mechanism—generation forces physical structure—is refuted.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is that visual experience is a sufficient substrate for general intelligence. It argues that the same recipe that produced language intelligence—large-scale data, a simple generative objective, and aggressive scaling—can be transferred to video, the most general form of visual data. Several contributors add that a genuinely visually intelligent system must keep learning over time, maintain a persistent and revisable spatial world model, act to gather informative experience, and generate coherent alternatives, not just one correct answer. The paper's integrative hypothesis is that three objectives—autoregressive prediction, open-ended generation, and reconstruction—jointly define a route from visual foundation models to visual intelligence. It explicitly stops short of declaring that this route is already established, and it states that whether pixel prediction alone recovers physical structure remains an open question.

Load-bearing premise

The load-bearing premise is that predicting or generating pixels forces a model to represent the physical structure of the world—mass, friction, contact, geometry—rather than only how things look, a premise one section of the paper itself challenges with the example of a convincingly rendered falling cup that may contain no representation of mass, contact, or friction.

Editorial extensions

If this is right

  • Task-specific vision models would become components or evaluations of a single generative video foundation model rather than ends in themselves.
  • Vision would stop being a peripheral encoder for language models and become a source of grounded world structure that language can query, steer, and teach.
  • Benchmarks would have to measure transfer, persistent world knowledge, continual adaptation, active information seeking, creativity, and physical consistency instead of static task accuracy.
  • Scaling visual models could yield emergent abilities analogous to those of language models, including visual in-context learning in which a few visual examples define a new task with no parameter updates.
  • The proposed route from vision to general intelligence is the convergence of sequential prediction, open-ended generation, and reconstruction; each objective alone may be useful, but their joint optimization is the paper's candidate path.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A consequence the paper leaves implicit: if the three-objective convergence is the real story, the natural benchmark is a single scene task requiring all three—predict the next frame, generate an alternative intervention, and reconstruct occluded geometry—rather than measuring each objective separately.
  • The paper's internal disagreement points to a synthesis it does not explicitly endorse: pixel-level generation may supply broad appearance priors, while explicit structure such as programs, code, or simulation and embodied interaction supply the physical quantities pixel loss does not guarantee.
  • A testable extension suggested by the proposal: measure whether physical parameters such as mass or friction become more recoverable from a video model's internals as video scale increases; if they do not, the scaling-generation route likely needs a structural auxiliary objective.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The manuscript is a white paper advancing 'visual general intelligence' (VGI) as a research agenda: intelligence that emerges from visual experience and that is capable of prediction, reconstruction, generation, and adaptation across tasks, with vision treated as a core pathway toward AGI rather than as a sensor feeding language-centric systems. It collects ten individually authored perspectives (Sections 2.1-2.10) covering generative video models, algorithmic creativity, continual lifelong learning, multimodal generative efficiency, scientific discovery, Spatial AI, embodied intelligence, compositional physical structure via code, vision-native interaction, and the convergence of prediction, imagination, and reconstruction. Section 3 summarizes the perspectives in Table 1, acknowledges that they do not converge, and discusses scaling, generation, physical structure, continual learning, and evaluation. The paper explicitly declines to offer a single definition or roadmap, presenting this plurality as a central conclusion.

Significance. The paper makes explicit bets that are useful for shaping the computer vision research agenda: generative video scaling as a 'version 1.0' of visual foundation models (Section 2.1), creativity as a test of VGI (Section 2.2), continuous one-datum learning (Section 2.3), and the recovery of compositional physical structure (Section 2.8). It also cites concrete demonstrations, notably the Veo 3 zero-shot task-transfer results in reference [93], and it is honest about the lack of convergence among the perspectives. If the agenda is correct, it would redirect substantial resources toward generative video modeling, persistent spatial world models, embodied learning, and new evaluation dimensions. The paper does not claim machine-checked proofs or fitted parameters, which is appropriate for a position piece; the main scientific risk is the unresolved empirical bet about whether pixel-level generation is sufficient for recovering world structure, a point the authors themselves raise but do not resolve.

major comments (3)
  1. [Section 2.1 vs. Section 2.8 and Section 3.2] The load-bearing claim that a generative objective 'forces the model to get everything right' (Section 2.1) is directly contradicted by Section 2.8, which states that a video model rendering a convincing falling cup 'has not shown that mass, contact, or friction appear anywhere inside it.' Section 3.1 concedes that the perspectives 'do not converge,' and Section 3.2 only says that representations 'may need' to be inspectable, editable, simulated, and verified. Because the paper's strongest concrete route to VGI depends on the assumption that pixel-level generation is sufficient for recovering latent physical and compositional structure, the manuscript should either propose a discriminating experiment (for example, intervention or counterfactual prediction tests of the type cited in references [25] and [87]) or explicitly re-frame the generative route as one open hypothesis among several rather than 'the key to solving visual intelligence.' As written, the paper contains both an assertion and its negation without explaining why the VGI agenda should nevertheless be prioritized.
  2. [Section 2.1, reference [93]] The only direct empirical support for the flagship claim that generative video models are visual foundation models is the Veo 3 zero-shot demonstration described in Section 2.1 and reference [93]. As Section 2.1 itself notes, machine learning models 'love to learn shortcuts,' and the appearance-versus-structure distinction drawn in Section 2.8 applies directly to these transfer results: performing edge detection, segmentation, or maze solving through image-to-video generation could in principle rely on memorized or generic image priors rather than on a general visual world model. The paper does not discuss what controls or probes were used in [93] to rule out shortcut solutions, such as generalization to novel objects or consistency under intervention. This evidence is too weak, as presented, to support the inference that video models constitute a 'visual foundation model' in the VGI sense; the limits of the demonstration should be stated explicitly.
  3. [Section 3.2 and Section 2.2] Section 3.2 lists desirable evaluation dimensions for VGI, and Section 2.2 gives concrete metrics for creativity (coherence, structural diversity, originality, utility), but no concrete metric or benchmark is proposed that could falsify the central claim that pixel-level video generation is sufficient for visual general intelligence. The manuscript would be substantially strengthened by specifying at least one operational probe of physical-structural understanding—for example, action-conditioned prediction accuracy, edit/simulation consistency, or out-of-distribution counterfactual correctness—and stating in advance which outcome would count against the scaling-generative hypothesis. Without such a test, the disagreement between Section 2.1 and Section 2.8 remains a matter of intellectual taste rather than a resolvable empirical question, and the VGI agenda risks being unfalsifiable as stated.
minor comments (5)
  1. [Section 2.5] The text refers to 'Maxell's equations'; this should be corrected to 'Maxwell's equations.'
  2. [Section 2.3 and elsewhere] Several informal asides, such as 'Happy to be wrong, I like simplicity' in the footnotes, are out of keeping with the register of a journal white paper; consider removing them or moving the sentiment into the main text more formally.
  3. [Table 1] Table 1 would be easier to use if the 'Position' column included the corresponding section number for each contributor (e.g., 'Section 2.4') to allow quick cross-referencing.
  4. [References] Many key references are 2025-2026 arXiv preprints; for the camera-ready version, the authors should replace them with published versions where available and verify the description of the Veo 3 demonstration in [93] independently of the contributing author's summary.
  5. [Section 2.9] The estimate that visual learning may require '1,000x or even 10,000x' more compute and data than language pre-training is presented without any supporting argument or citation; since the section is explicitly speculative, this should be labeled more clearly as an order-of-magnitude guess rather than a projection with evidence.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a position piece whose concrete claims are hypotheses or empirical citations, and the absence of a derivation chain means there is nothing that reduces to its own inputs.

full rationale

This is a white paper synthesizing invited perspectives; it contains no fitted parameters, no equations, and no prediction derived from a model. The central assertions are framed as bets or open questions, not consequences: Section 2.1 says Geirhos is 'convinced' that scaling generative video models 'holds the key,' and Section 2.8 explicitly states that 'Whether these quantities eventually emerge from pixel prediction alone... remains an open question.' The many self-citations (e.g., [93] for Veo 3 zero-shot behavior, [89] for compositional generalization, [16,17] for Spatial AI) are empirical or prior-research references presented transparently ('In our recent work'), and the strongest one is corroborated by subsequent external works cited in the same paragraph ([37,45,57,69,80,86,110,115]). None of these citations is used as a uniqueness theorem or as a substitute for an argument; the paper's own discussion sections acknowledge non-convergence (Section 3.1) and prematurity (Section 4). The definition of VGI as intelligence emerging from visual experience is a framing choice, not a derivation that assumes the conclusion. Therefore no circular step can be exhibited.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper introduces the term VGI, but treats it as a label for an existing cluster of research directions rather than a new theoretical entity. No particles, forces, or formal objects are invented. The ledger instead consists of five domain-level assumptions about learning, scaling, and the primacy of vision that the white paper's agenda rests on. These are load-bearing for the paper's recommendations, and at least one (the generative objective implying world structure) is explicitly contested inside the paper itself.

assumptions (5)
  • domain assumption Large-scale generative pretraining on visual data yields general-purpose visual intelligence.
    Invoked most strongly in Section 2.1 as the central bet that video generation is the key to visual intelligence, and in Section 2.10 as part of the prediction/imagination/reconstruction triad. It is asserted as a hypothesis, not derived or tested within this paper.
  • domain assumption Vision is a privileged modality for grounding world structure, with a phylogenetic priority over language.
    Section 1 argues from Cambrian eyes to the recent evolution of language that vision is foundational; Sections 2.3, 2.8, and 2.9 repeat versions of this premise. It motivates the whole agenda without being independently established.
  • domain assumption A generative objective forces a model to learn underlying world structure, not only appearance statistics.
    Used in Section 2.1 to argue that generative video models will learn more than classifiers do. Directly contested by Section 2.8, which states that renderable appearance does not imply internalized physical quantities. The paper leaves the contradiction standing.
  • domain assumption Scaling laws and emergent abilities observed in language models transfer to visual modalities.
    Section 1 and Section 2.1 extrapolate from the GPT trajectory to video models. The paper acknowledges the transfer is not guaranteed ('the success of language models cannot simply be transferred to visual models') but proceeds as if the analogy is a useful guide.
  • domain assumption Alternative definitions of intelligence as symbolic or language-centric are insufficient for AGI.
    The paper's framing assumes that a vision-independent path is worth pursuing as an AGI route. This is an interpretive stance, argued from evolution and the existence of unverbalized physical knowledge (Section 1, Section 2.8), not a proven fact.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Visual General Intelligence: A White Paper." pith.science (2026). https://pith.science/paper/LZIZQCGG

@misc{pith2026260825924,
  author       = {Pith},
  title        = {Pith review of: Visual General Intelligence: A White Paper},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LZIZQCGG}},
  note         = {Machine review of arXiv:2608.25924}
}
read the original abstract

This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward AGI. In the language domain, beginning with the introduction of the Transformer architecture, the GPT series has demonstrated transfer to unseen tasks through autoregressive language modeling on web-scale text combined with aggressive scaling. This raises a natural question, namely, what capabilities and forms of intelligence can emerge from visual modalities such as images, videos, and geometry? In this paper, we discuss whether visual intelligence can serve as a pathway toward AGI, referred to in this paper as visual general intelligence (VGI), by bringing together contributors from diverse standpoints and affiliations. Our aim is not to offer a single definition of visual intelligence, but to clarify the principles that computer vision should pursue in the AGI era, the visual input modalities, the benchmarks, the learning paradigms, and the relationship between vision, when taken as the core, and other modalities such as language.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

115 extracted references · 59 canonical work pages

  1. [93]

    Video mod- els are zero-shot learners and reasoners.arXiv preprint arXiv:2509.20328, 2025

    Thadd ¨aus Wiedemer, Yuxuan Li, Paul Vicol, Shixi- ang Shane Gu, Nick Matarese, Kevin Swersky, Been Kim, Priyank Jaini, and Robert Geirhos. Video mod- els are zero-shot learners and reasoners.arXiv preprint arXiv:2509.20328, 2025. 3

  2. [25]

    WorldScore: A unified evaluation benchmark for world generation

    Haoyi Duan, Hong-Xing Yu, Sirui Chen, Li Fei-Fei, and Jiajun Wu. WorldScore: A unified evaluation benchmark for world generation. InICCV, 2025. 11

  3. [87]

    ENACT: Evaluating em- bodied cognition with world modeling of egocentric inter- action

    Qineng Wang, Wenlong Huang, Yu Zhou, Hang Yin, Tian- wei Bao, Jianwen Lyu, Weiyu Liu, Ruohan Zhang, Jiajun Wu, Li Fei-Fei, and Manling Li. ENACT: Evaluating em- bodied cognition with world modeling of egocentric inter- action. InICLR, 2026. 11

  4. [1]

    A review of learning-based dynam- ics models for robotic manipulation.Science Robotics, 10 (105), 2025

    Bo Ai, Stephen Tian, Haochen Shi, Yixuan Wang, Tobias Pfaff, Cheston Tan, Henrik I Christensen, Hao Su, Jiajun Wu, and Yunzhu Li. A review of learning-based dynam- ics models for robotic manipulation.Science Robotics, 10 (105), 2025. 11

  5. [2]

    Flamingo: A visual language model for few-shot learn- ing

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, An- toine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Se- bastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikołaj Bi ´...

  6. [3]

    Neural machine translation by jointly learning to align and translate

    Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. InInternational Conference on Learning Repre- sentations, 2015. 1

  7. [4]

    Dana H. Ballard. Animate vision.Artificial Intelligence, 48 (1):57–86, 1991. 12

  8. [5]

    Baugh and Thomas Cable.A History of the En- glish Language

    Albert C. Baugh and Thomas Cable.A History of the En- glish Language. Routledge, 6 edition, 2013. 1

Show all 115 references
  1. [6]

    Recogni- tion in terra incognita

    Sara Beery, Grant Van Horn, and Pietro Perona. Recogni- tion in terra incognita. InProceedings of the European con- ference on computer vision (ECCV), pages 456–473, 2018. 3

  2. [7]

    A neural probabilistic language model

    Yoshua Bengio, R ´ejean Ducharme, Pascal Vincent, and Christian Janvin. A neural probabilistic language model. Journal of Machine Learning Research, 3:1137–1155,

  3. [8]

    Rates of passerine body plan evolution in time and space.Nature Ecology & Evolution, pages 1–15, 2026

    Jacob S Berv, Charlotte M Probst, Santiago Claramunt, J Ryan Shipley, Matt Friedman, Stephen A Smith, David F Fouhey, and Brian C Weeks. Rates of passerine body plan evolution in time and space.Nature Ecology & Evolution, pages 1–15, 2026. 7

  4. [9]

    Combining labeled and unlabeled data with co-training

    Avrim Blum and Tom Mitchell. Combining labeled and unlabeled data with co-training. InProceedings of the eleventh annual conference on Computational learning the- ory, pages 92–100, 1998. 6

  5. [10]

    Observation of free oscillations of the sun.Nature, 259(5539):92–95,

    JR Brookes, GR Isaak, and HB Van der Raay. Observation of free oscillations of the sun.Nature, 259(5539):92–95,

  6. [11]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Nee- lakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Je...

  7. [12]

    Large video planner enables generalizable robot control.arXiv preprint arXiv:2512.15840, 2025

    Boyuan Chen, Tianyuan Zhang, Haoran Geng, Caiyi Zhang, Peihao Li, Kiwhan Song, William T Freeman, Jiten- dra Malik, Pieter Abbeel, Russ Tedrake, et al. Large video planner enables generalizable robot control.arXiv preprint arXiv:2512.15840, 2025. 10

  8. [13]

    Mark CM Cheung, P Boerner, CJ Schrijver, P Testa, F Chen, H Peter, and A Malanushenko. Thermal diagnostics with the atmospheric imaging assembly on board the solar dynamics observatory: a validated method for differential emission measure inversions.The Astrophysical Journal, ...

  9. [14]

    Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots.arXiv preprint arXiv:2402.10329, 2024

    Cheng Chi, Zhenjia Xu, Chuer Pan, Eric Cousineau, Ben- jamin Burchfiel, Siyuan Feng, Russ Tedrake, and Shu- ran Song. Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots.arXiv preprint arXiv:2402.10329, 2024. 5

  10. [15]

    Daniels and William Bright, editors.The World’s Writing Systems

    Peter T. Daniels and William Bright, editors.The World’s Writing Systems. Oxford University Press, 1996. 1

  11. [16]

    A. J. Davison. FutureMapping: The computational structure of Spatial AI systems.arXiv preprint arXiv:1803.11288, 2018. 8

  12. [17]

    A. J. Davison and J. Ortiz. FutureMapping 2: Gaus- sian Belief Propagation for Spatial AI.arXiv preprint arXiv:1910.14139, 2019. 8

  13. [18]

    Anymate: A dataset and baselines for learning 3D object rigging

    Yufan Deng, Yuhao Zhang, Chen Geng, Shangzhe Wu, and Jiajun Wu. Anymate: A dataset and baselines for learning 3D object rigging. InSIGGRAPH, 2025. 11

  14. [19]

    BERT: Pre-training of deep bidirectional trans- formers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional trans- formers for language understanding. InProceedings of the 2019 Conference of the North American Chapter of the As- sociation for Computational Linguistics: Human La...

  15. [20]

    Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duck- worth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussai...

  16. [21]

    Compositional genera- tive modeling: A single model is not all you need.arXiv preprint arXiv:2402.01103, 2024

    Yilun Du and Leslie Kaelbling. Compositional genera- tive modeling: A single model is not all you need.arXiv preprint arXiv:2402.01103, 2024. 9

  17. [22]

    Curious represen- tation learning for embodied intelligence

    Yilun Du, Chuang Gan, and Phillip Isola. Curious represen- tation learning for embodied intelligence. InProceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 10408–10417, 2021. 10

  18. [23]

    Learning object-based state estimators for household robots

    Yilun Du, Tomas Lozano-Perez, and Leslie Pack Kael- bling. Learning object-based state estimators for household robots. In2022 IEEE/RSJ International Conference on In- telligent Robots and Systems (IROS), pages 12558–12565. IEEE, 2022. 9

  19. [24]

    Learning universal policies via text-guided video genera- tion.Advances in neural information processing systems, 36:9156–9172, 2023

    Yilun Du, Sherry Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Josh Tenenbaum, Dale Schuurmans, and Pieter Abbeel. Learning universal policies via text-guided video genera- tion.Advances in neural information processing systems, 36:9156–9172, 2023. 10

  20. [26]

    Modality forcing for scalable spatial generation.arXiv preprint arXiv:2606.13676, 2026

    Bardienus Pieter Duisterhof, Deva Ramanan, Jeffrey Ich- nowski, Justin Johnson, and Keunhong Park. Modality forcing for scalable spatial generation.arXiv preprint arXiv:2606.13676, 2026. 6

  21. [27]

    Tecumseh Fitch.The Evolution of Language

    W. Tecumseh Fitch.The Evolution of Language. Cam- bridge University Press, 2010. 1

  22. [28]

    Neuse: Neural se (3)-equivariant em- bedding for consistent spatial understanding with objects

    Jiahui Fu, Yilun Du, Kurran Singh, Joshua B Tenenbaum, and John J Leonard. Neuse: Neural se (3)-equivariant em- bedding for consistent spatial understanding with objects. arXiv preprint arXiv:2303.07308, 2023. 9

  23. [29]

    CaP-X: A frame- work for benchmarking and improving coding agents for robot manipulation

    Letian Fu, Justin Yu, Karim El-Refai, Ethan Kou, Haoru Xue, Huang Huang, Wenli Xiao, Guanzhi Wang, Dantong Niu, Li Fei-Fei, Guanya Shi, Jiajun Wu, Shankar Sastry, Yuke Zhu, Ken Goldberg, and Linxi Fan. CaP-X: A frame- work for benchmarking and improving coding agents for robot...

  24. [30]

    ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness

    Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A Wichmann, and Wieland Bren- del. ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In International conference on learning representations, 2019. 3

  25. [31]

    Shortcut learning in deep neural net- works.Nature Machine Intelligence, 2(11):665–673, 2020

    Robert Geirhos, J ¨orn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Fe- lix A Wichmann. Shortcut learning in deep neural net- works.Nature Machine Intelligence, 2(11):665–673, 2020. 3

  26. [32]

    Birth and death of a rose

    Chen Geng, Yunzhi Zhang, Shangzhe Wu, and Jiajun Wu. Birth and death of a rose. InCVPR, 2025. 11

  27. [33]

    NeuROK: Generative 4d neural object kinematics

    Chen Geng, Guangzhao He, Yue Gao, Yunzhi Zhang, Shangzhe Wu, and Jiajun Wu. NeuROK: Generative 4d neural object kinematics. InCVPR, 2026. 11

  28. [34]

    Gibson.The Ecological Approach to Visual Per- ception

    James J. Gibson.The Ecological Approach to Visual Per- ception. Houghton Mifflin, Boston, 1979. 12

  29. [35]

    Artificial general intelligence: Concept, state of the art, and future prospects.Journal of Artificial Gen- eral Intelligence, 5(1):1–48, 2014

    Ben Goertzel. Artificial general intelligence: Concept, state of the art, and future prospects.Journal of Artificial Gen- eral Intelligence, 5(1):1–48, 2014. 1

  30. [36]

    JHU Press,

    Ulf Grenander.Elements of pattern theory. JHU Press,

  31. [37]

    Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark

    Ziyu Guo, Xinyan Chen, Renrui Zhang, Ruichuan An, Yu Qi, Dongzhi Jiang, Xiangtai Li, Manyuan Zhang, Hong- sheng Li, and Pheng-Ann Heng. Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark. InProceedings of the IEEE/CVF Conference on Com...

  32. [38]

    The solar cycle.Living reviews in solar physics, 12(1):4, 2015

    David H Hathaway. The solar cycle.Living reviews in solar physics, 12(1):4, 2015. 7

  33. [39]

    Movement-produced stimula- tion in the development of visually guided behavior.Jour- nal of Comparative and Physiological Psychology, 56(5): 872–876, 1963

    Richard Held and Alan Hein. Movement-produced stimula- tion in the development of visually guided behavior.Jour- nal of Comparative and Physiological Psychology, 56(5): 872–876, 1963. 10

  34. [40]

    The era5 global reanalysis.Quarterly journal of the royal mete- orological society, 146(730):1999–2049, 2020

    Hans Hersbach, Bill Bell, Paul Berrisford, Shoji Hirahara, Andr´as Hor ´anyi, Joaqu ´ın Mu ˜noz-Sabater, Julien Nicolas, Carole Peubey, Raluca Radu, Dinand Schepers, et al. The era5 global reanalysis.Quarterly journal of the royal mete- orological society, 146(730):1999–2049, 2020. 7

  35. [41]

    Fast and accurate emulation of the sdo/hmi stokes inversion with uncertainty quantification.The Astrophysical Journal, 911(2):130, 2021

    Richard EL Higgins, David F Fouhey, Dichang Zhang, Spiro K Antiochos, Graham Barnes, J Todd Hoeksema, KD Leka, Yang Liu, Peter W Schuck, and Tamas I Gombosi. Fast and accurate emulation of the sdo/hmi stokes inversion with uncertainty quantification.The Astrophysical Journal, ...

  36. [42]

    To recognize shapes, first learn to gen- erate images.Progress in brain research, 165:535–547,

    Geoffrey E Hinton. To recognize shapes, first learn to gen- erate images.Progress in brain research, 165:535–547,

  37. [43]

    The helioseismic and magnetic imager (hmi) vector magnetic field pipeline: Overview and performance.Solar Physics, 289(9):3483– 3530, 2014

    J Todd Hoeksema, Yang Liu, Keiji Hayashi, Xudong Sun, Jesper Schou, Sebastien Couvidat, Aimee Norton, Monica Bobra, Rebecca Centeno, KD Leka, et al. The helioseismic and magnetic imager (hmi) vector magnetic field pipeline: Overview and performance.Solar Physics, 289(9):3483– ...

  38. [44]

    What’s left? concept grounding with logic-enhanced foundation models

    Joy Hsu, Jiayuan Mao, Joshua B Tenenbaum, and Jiajun Wu. What’s left? concept grounding with logic-enhanced foundation models. InNeurIPS, 2023. 11

  39. [45]

    Collabvr: Collaborative video reasoning with vision- language and video generation models.arXiv preprint arXiv:2605.08735, 2026

    Joowon Kim, Seungho Shin, Joonhyung Park, and Eunho Yang. Collabvr: Collaborative video reasoning with vision- language and video generation models.arXiv preprint arXiv:2605.08735, 2026. 3

  40. [46]

    Segment anything

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. InProceedings of the IEEE/CVF international conference on computer vision, pages 4015–4026, 2023. 3

  41. [47]

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural net- works. InAdvances in Neural Information Processing Sys- tems, 2012. 2

  42. [48]

    Picture: A probabilistic program- ming language for scene perception

    Tejas D Kulkarni, Pushmeet Kohli, Joshua B Tenenbaum, and Vikash Mansinghka. Picture: A probabilistic program- ming language for scene perception. InProceedings of the ieee conference on computer vision and pattern recogni- tion, pages 4390–4399, 2015. 9

  43. [49]

    A collection of definitions of intelligence, 2007

    Shane Legg and Marcus Hutter. A collection of definitions of intelligence, 2007. 1

  44. [50]

    BLIP: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

    Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. BLIP: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation. InPro- ceedings of the 39th International Conference on Machine Learning, pages 12888–12900, 2022. 2

  45. [51]

    BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. InPro- ceedings of the 40th International Conference on Machine Learning, pages 19730–19742, 2023. 2

  46. [52]

    Particulate: Feed-forward 3D object articulation

    Ruining Li, Yuxin Yao, Chuanxia Zheng, Christian Rup- precht, Joan Lasenby, Shangzhe Wu, and Andrea Vedaldi. Particulate: Feed-forward 3D object articulation. InCVPR,

  47. [53]

    Instruct-Particulate: Scaling feed-forward 3d ob- ject articulation with kinematic control.arXiv preprint arXiv:2606.14699, 2026

    Ruining Li, Yuxin Yao, Matt Zhou, Chuanxia Zheng, Chris- tian Rupprecht, Joan Lasenby, Shangzhe Wu, and Andrea Vedaldi. Instruct-Particulate: Scaling feed-forward 3d ob- ject articulation with kinematic control.arXiv preprint arXiv:2606.14699, 2026

  48. [54]

    Learning the 3D fauna of the web

    Zizhang Li, Dor Litvak, Ruining Li, Yunzhi Zhang, Tomas Jakab, Christian Rupprecht, Shangzhe Wu, Andrea Vedaldi, and Jiajun Wu. Learning the 3D fauna of the web. InCVPR,

  49. [55]

    Structured 4d latent predictive model for robot plan- ning.arXiv preprint arXiv:2607.01166, 2026

    Zhiyi Li, Peilin Wu, Xiaoshen Han, Ruojin Cai, and Yilun Du. Structured 4d latent predictive model for robot plan- ning.arXiv preprint arXiv:2607.01166, 2026. 9

  50. [56]

    Visual instruction tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. InAdvances in Neural In- formation Processing Systems, 2023. 2

  51. [57]

    Can world simulators reason? gen-vire: A generative visual reasoning benchmark.arXiv preprint arXiv:2511.13853, 2025

    Xinxin Liu, Zhaopan Xu, Ming Li, Kai Wang, Yong Jae Lee, and Yuzhang Shang. Can world simulators reason? gen-vire: A generative visual reasoning benchmark.arXiv preprint arXiv:2511.13853, 2025. 3

  52. [58]

    World action verifier: Self-improving world models via forward-inverse asymmetry.arXiv preprint arXiv:2604.01985, 2026

    Yuejiang Liu, Fan Feng, Lingjing Kong, Weifeng Lu, Jinzhou Tang, Kun Zhang, Kevin Murphy, Chelsea Finn, and Yilun Du. World action verifier: Self-improving world models via forward-inverse asymmetry.arXiv preprint arXiv:2604.01985, 2026. 10

  53. [59]

    A decade’s battle on dataset bias: Are we there yet?arXiv preprint arXiv:2403.08632,

    Zhuang Liu and Kaiming He. A decade’s battle on dataset bias: Are we there yet?arXiv preprint arXiv:2403.08632,

  54. [60]

    Self-improving loops for visual robotic planning

    Calvin Luo, Zilai Zeng, Mingxi Jia, Yilun Du, and Chen Sun. Self-improving loops for visual robotic planning. InInternational Conference on Learning Representations,

  55. [61]

    Chore- ographing a world of dynamic objects

    Yanzhe Lyu, Chen Geng, Karthik Dharmarajan, Yunzhi Zhang, Hadi Alzayer, Shangzhe Wu, and Jiajun Wu. Chore- ographing a world of dynamic objects. InCVPR, 2026. 11

  56. [62]

    The neuro-symbolic concept learner: Interpreting scenes, words, and sentences from nat- ural supervision

    Jiayuan Mao, Chuang Gan, Pushmeet Kohli, Joshua B Tenenbaum, and Jiajun Wu. The neuro-symbolic concept learner: Interpreting scenes, words, and sentences from nat- ural supervision. InICLR, 2019. 11

  57. [63]

    Build- ing intelligent agents with neuro-symbolic concepts.Com- munications of the ACM, 69(2), 2026

    Jiayuan Mao, Joshua B Tenenbaum, and Jiajun Wu. Build- ing intelligent agents with neuro-symbolic concepts.Com- munications of the ACM, 69(2), 2026. 11

  58. [64]

    David Marr.Vision: A Computational Investigation into the Human Representation and Processing of Visual Informa- tion. W. H. Freeman, San Francisco, 1982. 10, 12

  59. [65]

    Michael McCloskey and Neal J. Cohen. Catastrophic inter- ference in connectionist networks: The sequential learning problem. InPsychology of Learning and Motivation, pages 109–165. Academic Press, 1989. 10

  60. [66]

    Efficient estimation of word representations in vector space, 2013

    Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space, 2013. 1

  61. [67]

    On the computational architecture of the neocortex: I

    David Mumford. On the computational architecture of the neocortex: I. the role of the thalamo-cortical loop.Biologi- cal cybernetics, 65(2):135–145, 1991. 6, 10

  62. [68]

    Murai, E

    R. Murai, E. Dexheimer, and A. J. Davison. MASt3R- SLAM: Real-time dense SLAM with 3D reconstruction pri- ors. InCVPR, 2025. 8

  63. [69]

    Video models reason early: Exploiting plan commitment for maze solving.arXiv preprint arXiv:2603.30043, 2026

    Kaleb Newman, Tyler Zhu, and Olga Russakovsky. Video models reason early: Exploiting plan commitment for maze solving.arXiv preprint arXiv:2603.30043, 2026. 3

  64. [70]

    Free Press, 2003

    Andrew Parker.In the Blink of an Eye: How Vision Sparked the Big Bang of Evolution. Free Press, 2003. 1

  65. [71]

    J. Pearl. Theoretical impediments to machine learning, with seven sparks from the causal revolution. Technical report, University of California, Los Angeles, 2017. Technical Re- port R-275. 8

  66. [72]

    Improving language understanding by genera- tive pre-training

    Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by genera- tive pre-training. OpenAI Technical Report, 2018. 1

  67. [73]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. InProceedings of the...

  68. [74]

    Berg, and Li Fei-Fei

    Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. Imagenet large scale visual recogni- tion challenge.International Journal of Computer Vision, 11...

  69. [75]

    Jona Ruthardt, Manu Gaur, Deva Ramanan, Makarand Tapaswi, and Yuki M. Asano. Steerable visual represen- tations.ECCV, 2026. 5

  70. [76]

    3D-Generalist: Vision-language-action mod- els for crafting 3D worlds

    Fan-Yun Sun, Shengguang Wu, Christian Jacobsen, Thomas Yim, Haoming Zou, Alex Zook, Shangru Li, Yu- Hsin Chou, Ethem Can, Xunlei Wu, Clemens Eppner, Valts Blukis, Jonathan Tremblay, Jiajun Wu, Stan Birchfield, and Nick Haber. 3D-Generalist: Vision-language-action mod- els for ...

  71. [77]

    Ilya Sutskever, Oriol Vinyals, and Quoc V . Le. Sequence to sequence learning with neural networks. InAdvances in Neural Information Processing Systems, 2014. 1

  72. [78]

    Richard S. Sutton. The bitter lesson.http : / / www . incompleteideas . net / IncIdeas / BitterLesson.html, 2019. 1, 3, 11

  73. [79]

    Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural in- formation processing systems, 37:84839–84865, 2024

    Keyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Li- wei Wang. Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural in- formation processing systems, 37:84839–84865, 2024. 6

  74. [80]

    Thinking with video: Video generation as a promising multimodal reason- ing paradigm

    Jingqi Tong, Yurong Mou, Hangcheng Li, Mingzhe Li, Yongzhuo Yang, Ming Zhang, Qiguang Chen, Tianyi Liang, Xiaomeng Hu, Yining Zheng, et al. Thinking with video: Video generation as a promising multimodal reason- ing paradigm. InProceedings of the IEEE/CVF Conference on Compute...

  75. [81]

    Beyond language modeling: An exploration of multimodal pretraining.arXiv preprint arXiv:2603.03276, 2026

    Shengbang Tong, David Fan, John Nguyen, Ellis Brown, Gaoyue Zhou, Shengyi Qian, Boyang Zheng, Th ´eophane Vallaeys, Junlin Han, Rob Fergus, et al. Beyond language modeling: An exploration of multimodal pretraining.arXiv preprint arXiv:2603.03276, 2026. 5

  76. [82]

    Alan M. Turing. Computing machinery and intelligence. Mind, 59(236):433–460, 1950. 1

  77. [83]

    Visual routines.Cognition, 18(1–3):97– 159, 1984

    Shimon Ullman. Visual routines.Cognition, 18(1–3):97– 159, 1984. 12

  78. [84]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdvances in Neural Information Processing Systems, 2017. 1

  79. [85]

    CVPR 2026 Workshop on Visual General Intelligence: Vision Research Toward the AGI Era.https://cvpr2026- vgi- workshop

    VGI Workshop Organizers. CVPR 2026 Workshop on Visual General Intelligence: Vision Research Toward the AGI Era.https://cvpr2026- vgi- workshop. limitlab.xyz/, 2026. Accessed: 1 August, 2026. 2

  80. [86]

    A very big video rea- soning suite.arXiv preprint arXiv:2602.20159, 2026

    Maijunxian Wang, Ruisi Wang, Juyi Lin, Ran Ji, Thadd ¨aus Wiedemer, Qingying Gao, Dezhi Luo, Yaoyao Qian, Lianyu Huang, Zelong Hong, et al. A very big video rea- soning suite.arXiv preprint arXiv:2602.20159, 2026. 3

  81. [88]

    Super- synthia: Physics-ready full-disk vector magnetograms from hmi, hinode, and machine learning.The Astrophysical Jour- nal, 970(2):168, 2024

    Ruoyu Wang, David F Fouhey, Richard EL Higgins, Spiro K Antiochos, Graham Barnes, J Todd Hoeksema, Yang Liu, Peter W Schuck, and Tamas I Gombosi. Super- synthia: Physics-ready full-disk vector magnetograms from hmi, hinode, and machine learning.The Astrophysical Jour- nal, 970...

  82. [89]

    Composi- tional scene understanding through inverse generative mod- eling.arXiv preprint arXiv:2505.21780, 2025

    Yanbo Wang, Justin Dauwels, and Yilun Du. Composi- tional scene understanding through inverse generative mod- eling.arXiv preprint arXiv:2505.21780, 2025. 9

  83. [90]

    Shared morphological consequences of global warming in north american migratory birds.Ecol- ogy Letters, 23(2):316–325, 2020

    Brian C Weeks, David E Willard, Marketa Zimova, As- pen A Ellis, Max L Witynski, Mary Hennen, and Ben- jamin M Winger. Shared morphological consequences of global warming in north american migratory birds.Ecol- ogy Letters, 23(2):316–325, 2020. 7

  84. [91]

    Skeletal trait measurements for thousands of bird species.Scientific Data, 12(1):884, 2025

    Brian C Weeks, Zhizhuo Zhou, Charlotte M Probst, Jacob S Berv, Bruce O’Brien, Brett W Benz, Heather R Skeen, Mark Ziebell, Louise Bodt, and David F Fouhey. Skeletal trait measurements for thousands of bird species.Scientific Data, 12(1):884, 2025. 7

  85. [92]

    Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

    Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Bar- ret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. Emergent abilities of large language mod- el...

  86. [94]

    MIT Press, 1948

    Norbert Wiener.Cybernetics: Or Control and Communi- cation in the Animal and the Machine. MIT Press, 1948. 1

  87. [95]

    PhD thesis, Massachusetts Institute of Technology, 2019

    Jiajun Wu.Learning to See the Physical World. PhD thesis, Massachusetts Institute of Technology, 2019. 10

  88. [96]

    Physical scene understanding.AI Magazine, 45 (1):121–129, 2024

    Jiajun Wu. Physical scene understanding.AI Magazine, 45 (1):121–129, 2024. 10

  89. [97]

    Galileo: Perceiving phys- ical object properties by integrating a physics engine with deep learning

    Jiajun Wu, Ilker Yildirim, Joseph J Lim, William T Free- man, and Joshua B Tenenbaum. Galileo: Perceiving phys- ical object properties by integrating a physics engine with deep learning. InNeurIPS, 2015. 11

  90. [98]

    Learning to see physics via vi- sual de-animation

    Jiajun Wu, Erika Lu, Pushmeet Kohli, William T Freeman, and Joshua B Tenenbaum. Learning to see physics via vi- sual de-animation. InNeurIPS, 2017. 11

  91. [99]

    Discovering hybrid world representations with co-evolving foundation models

    Jiajun Wu, Yunzhi Zhang, Hong-Xing Yu, Joy Hsu, and Ji- ayuan Mao. Discovering hybrid world representations with co-evolving foundation models. InAAAI Conference on Ar- tificial Intelligence, Emerging Trends in AI, 2026. 11

  92. [100]

    MagicPony: Learning articu- lated 3D animals in the wild

    Shangzhe Wu, Ruining Li, Tomas Jakab, Christian Rup- precht, and Andrea Vedaldi. MagicPony: Learning articu- lated 3D animals in the wild. InCVPR, 2023. 11

  93. [101]

    Depth anything: Un- leashing the power of large-scale unlabeled data

    Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Ji- ashi Feng, and Hengshuang Zhao. Depth anything: Un- leashing the power of large-scale unlabeled data. InPro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10371–10381, 2024. 3

  94. [102]

    Video as the new language for real-world decision making.arXiv preprint arXiv:2402.17139, 2024

    Sherry Yang, Jacob Walker, Jack Parker-Holder, Yilun Du, Jake Bruce, Andre Barreto, Pieter Abbeel, and Dale Schu- urmans. Video as the new language for real-world decision making.arXiv preprint arXiv:2402.17139, 2024. 3

  95. [103]

    3d-mem: 3d scene memory for embodied exploration and reasoning

    Yuncong Yang, Han Yang, Jiachen Zhou, Peihao Chen, Hongxin Zhang, Yilun Du, and Chuang Gan. 3d-mem: 3d scene memory for embodied exploration and reasoning. In Proceedings of the Computer Vision and Pattern Recogni- tion Conference, pages 17294–17303, 2025. 9

  96. [104]

    Yarbus.Eye Movements and Vision

    Alfred L. Yarbus.Eye Movements and Vision. Plenum Press, New York, 1967. 12

  97. [105]

    Gel- sight: High-resolution robot tactile sensors for estimating geometry and force.Sensors, 17(12):2762, 2017

    Wenzhen Yuan, Siyuan Dong, and Edward H Adelson. Gel- sight: High-resolution robot tactile sensors for estimating geometry and force.Sensors, 17(12):2762, 2017. 5

  98. [106]

    Vision as bayesian infer- ence: analysis by synthesis?Trends in cognitive sciences, 10(7):301–308, 2006

    Alan Yuille and Daniel Kersten. Vision as bayesian infer- ence: analysis by synthesis?Trends in cognitive sciences, 10(7):301–308, 2006. 9, 10

  99. [107]

    Understanding bias in large-scale visual datasets

    Boya Zeng, Yida Yin, and Zhuang Liu. Understanding bias in large-scale visual datasets. InAdvances in Neural Infor- mation Processing Systems, 2024. 13

  100. [108]

    Seeing a rose in five thousand ways

    Yunzhi Zhang, Shangzhe Wu, Noah Snavely, and Jiajun Wu. Seeing a rose in five thousand ways. InCVPR, 2023. 11

  101. [109]

    The scene language: Representing scenes with programs, words, and embeddings

    Yunzhi Zhang, Zizhang Li, Matt Zhou, Shangzhe Wu, and Jiajun Wu. The scene language: Representing scenes with programs, words, and embeddings. InCVPR, 2025. 11

  102. [110]

    Are image-to-video models good zero-shot image editors? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2090– 2103, 2026

    Zechuan Zhang, Zhenyuan Chen, Zongxin Yang, and Yi Yang. Are image-to-video models good zero-shot image editors? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2090– 2103, 2026. 3

  103. [111]

    Bottjer, Shixue Hu, Zongjun Yin, and Maoyan Zhu

    Fangchen Zhao, David J. Bottjer, Shixue Hu, Zongjun Yin, and Maoyan Zhu. Complexity and diversity of eyes in early cambrian ecosystems.Scientific Reports, 3:2751, 2013. 1

  104. [112]

    Tesseract: learning 4d embodied world models.arXiv preprint arXiv:2504.20995,

    Haoyu Zhen, Qiao Sun, Hongxin Zhang, Junyan Li, Siyuan Zhou, Yilun Du, and Chuang Gan. Tesseract: learning 4d embodied world models.arXiv preprint arXiv:2504.20995,

  105. [113]

    Articraft: An agentic system for scalable articulated 3D asset generation.arXiv preprint arXiv:2605.15187, 2026

    Matt Zhou, Ruining Li, Xiaoyang Lyu, Zhaomou Song, Zhening Huang, Chuanxia Zheng, Christian Rupprecht, Andrea Vedaldi, and Shangzhe Wu. Articraft: An agentic system for scalable articulated 3D asset generation.arXiv preprint arXiv:2605.15187, 2026. 12

  106. [114]

    Robodreamer: Learning compo- sitional world models for robot imagination.arXiv preprint arXiv:2404.12377, 2024

    Siyuan Zhou, Yilun Du, Jiaben Chen, Yandong Li, Dit-Yan Yeung, and Chuang Gan. Robodreamer: Learning compo- sitional world models for robot imagination.arXiv preprint arXiv:2404.12377, 2024. 9

  107. [115]

    Video models can reason with verifiable rewards.arXiv preprint arXiv:2605.15458, 2026

    Tinghui Zhu, Sheng Zhang, James Y Huang, Selena Song, Xiaofei Wen, Yuankai Li, Hoifung Poon, and Muhao Chen. Video models can reason with verifiable rewards.arXiv preprint arXiv:2605.15458, 2026. 3

Pith tools

Reviewed August 27, 2026 · model on record in the stance chip above.