Pith. sign in

Paper Citation Record · LEDGER

Video models are zero-shot learners and reasoners

As of 4 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 100 inbound Pith citation observations for arXiv:2509.20328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.20328 v2

Coverage vector

measured 98 of 98 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-14T02:16:45.554252Z

measured 198 of 198 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 100 of 100 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T08:20:10.044650Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

98 of 98 outbound references displayed

  • verified exact34
  • verified fuzzy55
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bef73ba4-bf2d-41db-8f33-085b65e67f4f · outbound

This paper cites A Survey on Large Language Models for Code Generation.

Video models are zero-shot learners and reasoners A Survey on Large Language Models for Code Generation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-14T02:16:45.828955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:6ebfc9451696743b5a0a335b0f6e69f06d2f666fd0f499bbd73e588425135a9b

Observation 4f41cd84-640e-4566-b877-8b0cad30b2a1 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Video models are zero-shot learners and reasoners Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-14T02:16:45.630561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:fc6d788b1f8901a83eaab098c5dd5038a58d8a94f77dc0a85356f90196b12f0b

Observation 70f7216d-28ca-4688-ad18-9e2d03da39b4 · outbound

This paper cites Weaver: Foundation Models for Creative Writing.

Video models are zero-shot learners and reasoners Weaver: Foundation Models for Creative Writing

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.686151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:207fa40de68005591048586eaaaa8c4b5e6c1c61fae4c225caf38168db9cf272

Observation 33061bc7-b30a-4fc4-bdee-440561bd778b · outbound

This paper cites Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis.

Video models are zero-shot learners and reasoners Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.707178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:8d5956001eb79cb326da6c9c1165309938d9dd9a975bb8972797e9210d953fc3

Observation 1f8e1a78-c68f-4e95-adc5-ac257f4defc3 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

Video models are zero-shot learners and reasoners The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-14T02:16:45.758015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:bfe37aa2e3079c776e7760db6e55584c03a2e9ad2ac8ceacb1ea46539ba03f85

Observation 1e4b0727-f9e6-4e01-ace3-8b1d8e75b720 · outbound

This paper cites Agent Laboratory: Using LLM Agents as Research Assistants.

Video models are zero-shot learners and reasoners Agent Laboratory: Using LLM Agents as Research Assistants

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:05:19.002260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:363d14c0f663fdb6d9c7bd14414771cf3963c95bc84df1e077bc0c119a498628

Observation f01a4d9f-b346-4837-a3ad-f958976b4d7e · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901.

Video models are zero-shot learners and reasoners Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.945383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:9942c5c4430626195696ca1824be7007fc81c8c3d86aeceb0f771512a2ecb0a2

Observation 348889bd-e45a-4be7-be25-ee6416d71be5 · outbound

This paper cites Emergent Abilities of Large Language Models.

Video models are zero-shot learners and reasoners Emergent Abilities of Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-14T02:16:45.636483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:3fc38e83e279e33c051436f00362ed90c8b8520d97561d64387ef8f30a28fa06

Observation ef3203b8-916f-402c-b6dd-615aa45ae9b8 · outbound

This paper cites A Survey on In-context Learning.

Video models are zero-shot learners and reasoners A Survey on In-context Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-14T02:16:45.660933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:844ead87b5c33940db0904cb79957b5c89d348653080c7a30e6ce775e32033e6

Observation a12b7684-345e-4894-930a-8cb96a6a7e17 · outbound

This paper cites Large language models are zero-shot reasoners.Advances in neural information processing systems, 35: 22199–22213.

Video models are zero-shot learners and reasoners Large language models are zero-shot reasoners.Advances in neural information processing systems, 35: 22199–22213

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.949791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:9cf7bfdf4bc9afa442e8003d5b5936a04e7dccb6ea09beafa9b6508884454717

Observation 07e10f6a-0a33-436b-8da5-831d03b9d893 · outbound

This paper cites Segment anything.

Video models are zero-shot learners and reasoners Segment anything

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.953874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:0bea7c57d89f4817e9548bb7700e6d75a9e0c469098197b9ddc4695a4212c76d

Observation 31b2fcc2-bc71-4d27-a17d-8578eacd7d05 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Video models are zero-shot learners and reasoners SAM 2: Segment Anything in Images and Videos

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-14T02:16:45.736064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:0ac58d2490f8808684c5603c67e6e817882cf238c8a2bd8068e5044673d6fb92

Observation 6e7e27cb-928d-4b5a-a0ac-b0f675908781 · outbound

This paper cites You only look once: Unified, real-time object detection.

Video models are zero-shot learners and reasoners You only look once: Unified, real-time object detection

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.957495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:36d00f2e17d99658bcef99fcf539a58143e19e1503f10d0ded9a1e17876f3345

Observation 599fb6f1-94b3-435c-aaf5-cf4dcb48c2cc · outbound

This paper cites YOLOv11: An Overview of the Key Architectural Enhancements.

Video models are zero-shot learners and reasoners YOLOv11: An Overview of the Key Architectural Enhancements

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-14T02:16:45.782899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:7791067a37d72c4f1d1cccfea21cd851db0b728439c57c080255d5b7dce7615b

Observation 88ce5459-4164-4a44-985c-3275a995d336 · outbound

This paper cites From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models.

Video models are zero-shot learners and reasoners From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.793856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:6c581f04e626f4916fbd30099c80d22374f50f8fad1d3b0b9154c8954b6cc1a2

Observation 6db6f449-ee4c-4925-96a1-26781f0063bc · outbound

This paper cites Taskonomy: Disentangling task transfer learning.

Video models are zero-shot learners and reasoners Taskonomy: Disentangling task transfer learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.961504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:97b8a78f68d1f418da5596cc211a20b0e72bf9a70daa6d1505ffeb1424fe131d

Observation cb05e227-021d-4edc-b7b1-623258346c5a · outbound

This paper cites RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models.

Video models are zero-shot learners and reasoners RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.816914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:f008a3775e208f8b6fbbc33689a913086df173a323f223fe34fe06aeaad0f275

Observation 296b9fa4-276c-4e55-9db6-add78bb5826a · outbound

This paper cites Visualcloze: A universal image generation framework via visual in-context learning.

Video models are zero-shot learners and reasoners Visualcloze: A universal image generation framework via visual in-context learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.824763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:e1ac02aa2d75854cc8f1b0ff6f0da8b2bc330c282991447f59a678f85f680fa2

Observation ae17744b-1166-4a7c-b573-310a46abdc7b · outbound

This paper cites Images speak in images: A generalist painter for in-context visual learning.

Video models are zero-shot learners and reasoners Images speak in images: A generalist painter for in-context visual learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.965080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:2810d5fb07afe4db8363dc8d3728e3d0fcdcbf6bf7de53c12971acaf26085259

Observation 02ae06f0-e66e-4bea-9b6e-cb5f3c7669ab · outbound

This paper cites Test- time visual in-context tuning.

Video models are zero-shot learners and reasoners Test- time visual in-context tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.968994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:943e6a402e312c175bdbe4001a53d7f77e6cd95f9d64e3e3d261e2ffd000b45e

Observation e9fd3736-1373-466c-bc23-fbf9e0aba84c · outbound

This paper cites PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions.

Video models are zero-shot learners and reasoners PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.643098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:b014728fbbeb6afcdf29fc3b3f12176677575ef39d25b7783376af036833e7d6

Observation b3e254ba-791e-47a7-9ea8-d5cb42603a52 · outbound

This paper cites Omnigen: Unified image generation.

Video models are zero-shot learners and reasoners Omnigen: Unified image generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.972880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:6ecfdbb7d9f1179c7577b2b944d14abe0cc5c83efbb55fdeb5aa7beecff7e6f0

Observation c66a504e-7a54-40b2-9460-8134ed831263 · outbound

This paper cites One diffusion to generate them all.

Video models are zero-shot learners and reasoners One diffusion to generate them all

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.976845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:d5e30ecdd2adcb1d5ad7313376e6c4eda7c7334f793f72e859f597adea2a019c

Observation cc555498-692d-46d6-8300-b1e47755234d · outbound

This paper cites Dreamix: Video Diffusion Models are General Video Editors.

Video models are zero-shot learners and reasoners Dreamix: Video Diffusion Models are General Video Editors

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.692914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:92667a2308e00f47181380b1295642f054fd3c7ac2bf6e64b0802cc901756a59

Observation c4cebdcd-5d64-4fac-b9bd-0e9abc855778 · outbound

This paper cites Scalingproperties of diffusion models for perceptual tasks.

Video models are zero-shot learners and reasoners Scalingproperties of diffusion models for perceptual tasks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.980114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:c262e7b03818358f0d2aa804427c060679bda8dcd8731df9729fbbbf01fd9272

Observation 4a4b10f8-090a-45da-8050-70aa98552be8 · outbound

This paper cites Video as the New Language for Real-World Decision Making.

Video models are zero-shot learners and reasoners Video as the New Language for Real-World Decision Making

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.722401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:038e90f972e582ee70f58718b0bb7ca8ae5b5e9bf0dfcf2b2d3dd6b67b305c7e

Observation b033b0c7-d2a0-4392-960a-97a413181f01 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837.

Video models are zero-shot learners and reasoners Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.984159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:498543070f91db0d5a2f55b4098bb4c5a2f9c53658bfdc4f826e650a198bd654

Observation 4edf34c7-13a5-44d1-8391-472090970482 · outbound

This paper cites Large language models are human-level prompt engineers.

Video models are zero-shot learners and reasoners Large language models are human-level prompt engineers

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.996125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:64c464e2612c7e8427b6ce0663c29d7553681581a8ff472e4019e826e6ddbeca

Observation 56fcc7f1-f44a-42d9-93c1-8b4f9da2b148 · outbound

This paper cites Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing.ACM computing surveys, 55(9):1–35.

Video models are zero-shot learners and reasoners Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing.ACM computing surveys, 55(9):1–35

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.001183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:0e512733d27893f901440686988d170a1d664c55fbb9521722f82354fb7b0ae8

Observation abbdbd02-e4b8-4427-b500-ab4bafa21be2 · outbound

This paper cites Vertex AI Veo Prompt Rewriter.https://cloud.google.com/vertex-ai/ generative-ai/docs/video/turn-the-prompt-rewriter-off#prompt-rewriter.

Video models are zero-shot learners and reasoners Vertex AI Veo Prompt Rewriter.https://cloud.google.com/vertex-ai/ generative-ai/docs/video/turn-the-prompt-rewriter-off#prompt-rewriter

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.005034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:abb3718ced23f7c61d245cf059eb6a78e2fe8b2903b4923ddb69986e6c40bf45

Observation d24e5fc7-f1b8-4028-9193-a70fe8262488 · outbound

This paper cites an unresolved cited work.

Video models are zero-shot learners and reasoners Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-14T02:16:46.008932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:955a3f896ba9da18cbb3c3f631674448ce83813f33455ba2032535a5a7b10acf

Observation 28ab6dec-88e1-475e-8054-73ef29444676 · outbound

This paper cites Lmsys org text-to-video leaderboard.https://lmarena.ai/leaderboard/t ext-to-video, September 2025.

Video models are zero-shot learners and reasoners Lmsys org text-to-video leaderboard.https://lmarena.ai/leaderboard/t ext-to-video, September 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.013734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:b572a0d1633702eb870365250d05220820c8d09a5a321b982bb0f3687181d653

Observation 8376f9d4-3c43-4aba-a2c3-94e1c960dc6d · outbound

This paper cites Veo 2 announcement.https://blog.google/technology/google-labs/vide o-image-generation-update-december-2024/.

Video models are zero-shot learners and reasoners Veo 2 announcement.https://blog.google/technology/google-labs/vide o-image-generation-update-december-2024/

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.016928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:6c02b2128a63f108e815a29197dd7a52aee7d61a014a7e1a5c1813a73d4e884a

Observation 909cac90-fa50-496e-b8b9-e400b1cc3e07 · outbound

This paper cites Veo 2 launch.https://developers.googleblog.com/en/veo-2-video-gen eration-now-generally-available/.

Video models are zero-shot learners and reasoners Veo 2 launch.https://developers.googleblog.com/en/veo-2-video-gen eration-now-generally-available/

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.020541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:56e6bda4068bfdaa999ecfbe00fcae0fb491511588445376a4a9c3c61095841e

Observation 6954cf1d-45d1-4eeb-ae65-5546e6d0dbe1 · outbound

This paper cites Veo 3 announcement.https://blog.google/technology/ai/generative-m edia-models-io-2025/.

Video models are zero-shot learners and reasoners Veo 3 announcement.https://blog.google/technology/ai/generative-m edia-models-io-2025/

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.024333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:7325569f6f5a0d7944bd8adf3d49d333f9246d0e3759f428118408b7247339da

Observation ae0d721a-5a0d-48e1-b0cd-12dad6264034 · outbound

This paper cites Veo 3 launch.

Video models are zero-shot learners and reasoners Veo 3 launch

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.028191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:147d33de4d79c8a6199f6e46f97e57e9e4f8360beb21cc3fee70a0abe7a64322

Observation 83e868e0-cbe2-4cd9-9898-a5030c80e7c3 · outbound

This paper cites Holistically-nested edge detection.

Video models are zero-shot learners and reasoners Holistically-nested edge detection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.031839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:e5534383f32fcd2e6b880f9340c69baa5bb6dfd0fa344e4bfa92addf289f889e

Observation 662da419-d5dc-44b8-90e1-60a2fc0f1461 · outbound

This paper cites IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning.

Video models are zero-shot learners and reasoners IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T02:16:45.673040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:b981778597359b13668f28dfbfec1abfadb2593c4868c3f0c0c62cc73bac95c7

Observation 71f277fe-c5d9-4842-a170-1f6c587c8070 · outbound

This paper cites Bear, Elias Wang, Damian Mrowca, Felix J.

Video models are zero-shot learners and reasoners Bear, Elias Wang, Damian Mrowca, Felix J

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.035489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:f167502103fa25ff59f14a45c7336bd38abdf8e577351d35afc0a933d3146663

Observation 6f0f6882-7402-49be-a0c9-578581e2ce98 · outbound

This paper cites Benchmarking progress to infant-level physical reasoning in ai.

Video models are zero-shot learners and reasoners Benchmarking progress to infant-level physical reasoning in ai

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.039510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:5ac03738e90d8ad2eb683b196ff369dd5b396d393bf67f88b72d10e6594e72a4

Observation e67d384f-5157-445e-8efd-cab7b1cd49a2 · outbound

This paper cites GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models.

Video models are zero-shot learners and reasoners GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T02:16:45.697614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:2316aa43cbfc85995400fedb79ecc16e14f6118bcdcbf2c762e6eafcfbd91449

Observation 88cbd2c3-72a9-4dc1-b803-bc9136d59245 · outbound

This paper cites Physion++: Evaluating physical scene understanding that requires online inference of different physical properties.Advances in Neural Information Processing Systems, 36.

Video models are zero-shot learners and reasoners Physion++: Evaluating physical scene understanding that requires online inference of different physical properties.Advances in Neural Information Processing Systems, 36

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.043452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:4a76cfee0852670c7d7d0c2c74e705c03817e8769c1b51c73092db8202ac26ff

Observation d4da3c3e-8496-43bd-8189-9107dfec051c · outbound

This paper cites Videophy: Evaluating physical commonsense for video generation.

Video models are zero-shot learners and reasoners Videophy: Evaluating physical commonsense for video generation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.047617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:139c4b713a9308673db7b05a11fb788c81ae3b2d52e5a572ec01ca8906bec23f

Observation 0bb14a58-a96f-4d6b-b8f7-35fef17abfa4 · outbound

This paper cites LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines.

Video models are zero-shot learners and reasoners LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-14T02:16:45.732601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:86dbc27228cefe949ac08fbbccd47fb0da3b232ea0a106fa82038367ced98d46

Observation 67e73ec4-8a3f-4025-a1c0-ba330521b86e · outbound

This paper cites Towards world simulator: Crafting physical commonsense- based benchmark for video generation.

Video models are zero-shot learners and reasoners Towards world simulator: Crafting physical commonsense- based benchmark for video generation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.051461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:3f63d1139ec874c9870dcbd6c417ab238d586542987a228903841aeb338a0d17

Observation b46f4546-d677-490e-88a4-123891b1c737 · outbound

This paper cites How Far is Video Generation from World Model: A Physical Law Perspective.

Video models are zero-shot learners and reasoners How Far is Video Generation from World Model: A Physical Law Perspective

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:13:41.491170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:d9028def47562a7727ef4f0b0da3fff1170ed1f6b01cc59a77e515520026ffca

Observation 11174655-76a8-4901-8b29-a845e33e090a · outbound

This paper cites Do generative video models understand physical principles?.

Video models are zero-shot learners and reasoners Do generative video models understand physical principles?

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:47:06.031259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:43667c5666c29103c7b99de5d420d19eee7d18bcaae1feff017f050d0460d76d

Observation 658d0d9f-56e1-443c-aea4-d5a136be12ee · outbound

This paper cites Generative Physical AI in Vision: A Survey.

Video models are zero-shot learners and reasoners Generative Physical AI in Vision: A Survey

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.752438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:6de0b07705283dd19edf5c452293adcf7c10f2980e7514a9f09af90f25820e09

Observation cd3b184b-2f2c-4dc2-b44e-bb5f7f648a13 · outbound

This paper cites Visual cognition in multimodal large language models.Nature Machine Intelligence, pages 1–11.

Video models are zero-shot learners and reasoners Visual cognition in multimodal large language models.Nature Machine Intelligence, pages 1–11

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.055526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:8892c6b7cad58a2f5cc910c23d8c7672c92b654c85a2c574c71b80823f586b0e

Observation f958c9a8-4a82-449e-bf1f-b2406c8a9c42 · outbound

This paper cites Evaluating Newtonian Mechanics in Video Generative Models with Real Physical Systems.

Video models are zero-shot learners and reasoners Evaluating Newtonian Mechanics in Video Generative Models with Real Physical Systems

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T03:17:04.940243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:b2ac99528ad050540e63cb97dd42a57bdade71d0fd59e2062245370549004dfe

Observation 69a5a721-c0c7-4f28-9e95-6fb80f7ebe5c · outbound

This paper cites Intuitive physics understanding emerges from self-supervised pretraining on natural videos.

Video models are zero-shot learners and reasoners Intuitive physics understanding emerges from self-supervised pretraining on natural videos

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.772679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:de2d58006cc8f882c0c443620f956972b8c1528311f82544c5c22ae819c6c105

Observation 92c33d57-4426-4dd7-aaa5-6de33846f577 · outbound

This paper cites Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting.

Video models are zero-shot learners and reasoners Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.778635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:819880f1ef1d43d8cbb3df24068a7fa8a37870706658ef5c2d84d606cf007a0c

Observation c478925f-bc14-4f6d-8c21-88a2922abdb7 · outbound

This paper cites Human-level concept learning through probabilistic program induction.Science, 350(6266):1332–1338.

Video models are zero-shot learners and reasoners Human-level concept learning through probabilistic program induction.Science, 350(6266):1332–1338

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.059600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:929e98a6dce5f958f58838891c407a5ac3e4613dfc2ce564e7b5c36eb43fd387

Observation 96208e04-2720-4f38-9ca4-5006ef5ab42c · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Video models are zero-shot learners and reasoners V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-14T02:16:45.787940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:efc84f159dc5d4a60fe01fb0410dbcf98dd60313bc0342a09b41327249ad81d9

Observation 04c05d67-43d5-4323-94cd-b5b60de22f9b · outbound

This paper cites Nano Banana: Gemini Image Generation Overview.https://gemini.google/ov erview/image-generation/.

Video models are zero-shot learners and reasoners Nano Banana: Gemini Image Generation Overview.https://gemini.google/ov erview/image-generation/

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.065188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:9cd0c47a99b8cca90d171c8bbce8e0ceff52ac051955f0cb50fa546e70d3d2bc

Observation 819c2c9b-35fe-4b93-ba1e-7ad3611949f1 · outbound

This paper cites Text-to-image diffusion models are zero shot classifiers.Advances in Neural Information Processing Systems, 36:58921–58937.

Video models are zero-shot learners and reasoners Text-to-image diffusion models are zero shot classifiers.Advances in Neural Information Processing Systems, 36:58921–58937

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.068567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:b5ad12d9c23edd47f01bc3eafa11b9c754556f96a00ac0d3aead321eb1ce4283

Observation cc1b3590-b6be-4fa7-9507-ff61347bde4f · outbound

This paper cites Intriguing properties of generative classifiers.

Video models are zero-shot learners and reasoners Intriguing properties of generative classifiers

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.072441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:47cfb9e71028d0f1d4839db9c4bfce36275ec7e4c2c0facecfbce285bb1efd61

Observation d0b8f05f-29c6-4fde-828f-30cdf38962e9 · outbound

This paper cites Peekaboo: Text to Image Diffusion Models are Zero-Shot Segmentors.

Video models are zero-shot learners and reasoners Peekaboo: Text to Image Diffusion Models are Zero-Shot Segmentors

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T02:16:45.821415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:fe141942a9983fb6a3e4c61cf451db7fcb969480a4c8f85f64986078c62633c9

Observation c52699cf-9069-4241-a370-53baac5fbf99 · outbound

This paper cites Text2video-zero: Text-to-image diffusion models are zero-shot video generators.

Video models are zero-shot learners and reasoners Text2video-zero: Text-to-image diffusion models are zero-shot video generators

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.075862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:b0d766ed9c07db2dd3170861d493705cbd71ca4919989a3e9629637ef747008d

Observation 5d6a7d5a-7fb4-475c-a741-d2b810d252b4 · outbound

This paper cites Dense extreme inception network: Towards a robust CNN model for edge detection.

Video models are zero-shot learners and reasoners Dense extreme inception network: Towards a robust CNN model for edge detection

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.080483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:c3f00dac59016d22c0fedb594a4be8c37053f4a900188acfacf31285f58fcbb6

Observation 3048c93f-6aae-4f61-b8df-6a717a2cf6b6 · outbound

This paper cites Dense extreme inception network for edge detection.Pattern Recognition, 139:109461.

Video models are zero-shot learners and reasoners Dense extreme inception network for edge detection.Pattern Recognition, 139:109461

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.624252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:97545283f9fab964ace14c8902765713cbc108eeff97a2405141eaa01cab0179

Observation 66e36985-accc-4d25-b3eb-a604f46f84a8 · outbound

This paper cites LVIS: A dataset for large vocabulary instance segmentation.

Video models are zero-shot learners and reasoners LVIS: A dataset for large vocabulary instance segmentation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.084493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:ab00661dbcaf45aad3531df0ba5942249ef5d50d0b2a05f78d7d64905936d094

Observation 90b5de49-3b14-4f96-be59-40327be201fc · outbound

This paper cites Emu edit: Precise image editing via recognition and generation tasks.

Video models are zero-shot learners and reasoners Emu edit: Precise image editing via recognition and generation tasks

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.088704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:1c1e82f121c777e86a48fce7c98bf39b4a9e15b3881acdb8f8379f94aa0fe2b4

Observation 85e9d970-c35b-4f80-8a49-45453bb18c8a · outbound

This paper cites Diffusion Model-Based Video Editing: A Survey.

Video models are zero-shot learners and reasoners Diffusion Model-Based Video Editing: A Survey

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.649401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:96bcb8b8f931068e87d676b0b172b654f75d1a2c48cf41d6eff7bdec933e879a

Observation e1cbb1d1-877c-44da-8271-f7955efe7489 · outbound

This paper cites VEGGIE: instructional editing and reasoning video concepts with grounded generation.

Video models are zero-shot learners and reasoners VEGGIE: instructional editing and reasoning video concepts with grounded generation

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.655323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:6061f3ec9714cb718a481e4c1d6a7e634430131a8a13f19483ba965f8ab33750

Observation 7d99fe5e-b18c-43bb-a476-59b7987e8e6a · outbound

This paper cites Pathways on the image manifold: Image editing via video generation.

Video models are zero-shot learners and reasoners Pathways on the image manifold: Image editing via video generation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.092664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:5f3643fb3899901782224c826d33adcc8267c95d1d9465d021f10357843b1b63

Observation 562bad6e-290e-434a-bb7f-6e19ffefdb60 · outbound

This paper cites Kiva: Kid-inspired visual analogies for testing large multimodal models.

Video models are zero-shot learners and reasoners Kiva: Kid-inspired visual analogies for testing large multimodal models

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.667062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:cf8e1a68cc06b4ffde02661b49f228411cf2c0ec6cf4028d3c1af560bc5717cd

Observation 82df084a-7573-43b6-8f1a-caa524c8f283 · outbound

This paper cites ImageNet classification with deep convolutional neural networks.Advances in neural information processing systems, 25.

Video models are zero-shot learners and reasoners ImageNet classification with deep convolutional neural networks.Advances in neural information processing systems, 25

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.096756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:aac4902a10eaecd9bb954c9191e916b0c2f73ba531a5900418b1bdfbd8d64d8a

Observation e5ac8b7b-b2a7-462c-b506-d6c6c3b849da · outbound

This paper cites The broader spectrum of in-context learning.

Video models are zero-shot learners and reasoners The broader spectrum of in-context learning

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T02:16:45.679766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:d1b552d52afa90d1bc4553339e0ae0b049d3f9fa5232aac1ca4aec239cc519ba

Observation 19f0b701-2afd-4426-896c-e7054f8baa9e · outbound

This paper cites Performance vs.

Video models are zero-shot learners and reasoners Performance vs

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.100499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:b1627ecfb1843a5f4dece841f27ed41347c0f3a3cfce1acc771cbefc556f0ad4

Observation 32f7ea35-5384-41ba-9d67-137c080d06f2 · outbound

This paper cites Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673.

Video models are zero-shot learners and reasoners Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:46.104113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:bb41bc5643ebd324f039adb3e6de9ba903a46a83943814848011afaa42d5c8fb

Observation 7b0824cf-c67a-4100-95d4-2fb9af23cecd · outbound

This paper cites LLM inference prices have fallen rapidly but unequally across tasks, march 2025.

Video models are zero-shot learners and reasoners LLM inference prices have fallen rapidly but unequally across tasks, march 2025

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.832087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:ba849023130e3c53477d9daea3d81a26a491cfd8faeee17801c4081c2fd8c655

Observation df72a32f-a177-4c83-8b72-91c9b6d470e8 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Video models are zero-shot learners and reasoners Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-14T02:16:45.702329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:bba3669465934254b239cc95f952fdccdb5b1cf5eb2696325fda0a350944fd5f

Observation e8623ba1-15bd-4c91-b129-e55e413cfbd1 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594.

Video models are zero-shot learners and reasoners Self-refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534–46594

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.835856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:33f5da118c75b59cdf74751a460a3df4dc950fcc570b96501c42c6eaeb938e7f

Observation 5bd3cd9d-904f-446a-a968-c1c1e5f1c99e · outbound

This paper cites OpenAI o1 System Card.

Video models are zero-shot learners and reasoners OpenAI o1 System Card

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-14T02:16:45.711478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:de604c15ac51ac885d1e5ad1bbe18cee13a4e6e0007727d5bcd6fdd34fc58781

Observation c3da3f50-69ff-463b-9fb1-17618de2aa3e · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Video models are zero-shot learners and reasoners Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-14T02:16:45.716616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:8aad4320368ccfec5dfa92a3c9efd412ab2b182347a5d96413d598bcb4848b54

Observation 8c1e15e0-3515-4b12-9dc9-f0a7d598cb99 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35: 27730–27744.

Video models are zero-shot learners and reasoners Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35: 27730–27744

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.838946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:ca1e18cd42490dcaa7c4cd9b3cce7da32bc6bb0d2c817a2dc70954ab7ee5e653

Observation f1b08678-78a0-49f1-b3c5-0af7bc882d9c · outbound

This paper cites Instruction Tuning with GPT-4.

Video models are zero-shot learners and reasoners Instruction Tuning with GPT-4

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:04:18.193148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:4b07c1018f37591ab9fee65fb78bb203ee65395d8a114aca633a6969428a067c

Observation 2490a2bf-1199-4680-b469-ec23e4abd2c4 · outbound

This paper cites Sparse gradient regularized deep retinex network for robust low-light image enhancement.IEEE Transactions on Image Processing, 30:2072–2086.

Video models are zero-shot learners and reasoners Sparse gradient regularized deep retinex network for robust low-light image enhancement.IEEE Transactions on Image Processing, 30:2072–2086

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.842507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:7bd9177071ae43462450c52b7d4f377c41c8349bcd2bd890a64b5d400f6a7563

Observation 129d23ba-5da9-4787-a816-bca8f930feac · outbound

This paper cites Under- standing the limits of vision language models through the lens of the binding problem.Advances in Neural Information Processing Systems, 37:113436–113460.

Video models are zero-shot learners and reasoners Under- standing the limits of vision language models through the lens of the binding problem.Advances in Neural Information Processing Systems, 37:113436–113460

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.846108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:4d12d89ab0a50c957860ac16083534e039893c8afe65a8e9cdfc386413c740af

Observation 57826310-2a63-4706-a644-d42444dbfa92 · outbound

This paper cites an unresolved cited work.

Video models are zero-shot learners and reasoners Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-05-14T02:16:45.849438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:4071398375dedefc3ed39d24c474e3d38fbedfa48550723c3a03f4abebb4b9ea

Observation b6c999c7-920d-4a66-8efc-081d7d7e94b2 · outbound

This paper cites The intelligent eye.

Video models are zero-shot learners and reasoners The intelligent eye

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.852676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:eae341cd585ee052c521b692d508eab761422886d8eae3f428ab9cab54e50ebd

Observation 9c7b4fd0-bcc6-4d1c-b8be-2081d169ced8 · outbound

This paper cites ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness.

Video models are zero-shot learners and reasoners ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.856289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:dc4ad54e5d4f52fb4bc272b5389745af6742b3e81772adda3c8eb0f9b572fc40

Observation e2ea640d-bf28-41bb-a0b7-198814fce6af · outbound

This paper cites Objaverse-xl: A universe of 10m+ 3d objects.Advances in Neural Information Processing Systems, 36:35799–35813.

Video models are zero-shot learners and reasoners Objaverse-xl: A universe of 10m+ 3d objects.Advances in Neural Information Processing Systems, 36:35799–35813

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.860867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:d4d755778d4e545a2576dd36b331e40c25f350fc7f20605005bcb9900fa5f1d4

Observation 3540fdfa-a7d1-4c1c-8c83-db344f8fa203 · outbound

This paper cites On the Measure of Intelligence.

Video models are zero-shot learners and reasoners On the Measure of Intelligence

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-14T02:16:45.762658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:206efc6d4c21602a481366912e8cedabd27a481f2cb4c70a3154759f30b6c08c

Observation 0117641c-90ad-48ed-9529-5ca00c388108 · outbound

This paper cites Lawrence Zitnick.

Video models are zero-shot learners and reasoners Lawrence Zitnick

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.866068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:53a0aba138db2550a888de40587af1d7cf9fa30e0f7c514b563b357ae789ed90

Observation 9dee6f10-07ab-47ed-90d3-c01d8328272c · outbound

This paper cites Lawrence Zitnick.

Video models are zero-shot learners and reasoners Lawrence Zitnick

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.870013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:86c9baeb38b7d59d066fa2f99ccc3b3b80003c49afe6e609080493491c5cd594

Observation b2f09b07-9408-4916-9701-b83fc85fe857 · outbound

This paper cites Lawrence Zitnick and Piotr Dollár.

Video models are zero-shot learners and reasoners Lawrence Zitnick and Piotr Dollár

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.873460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:6dcc407f199c57e26b92b308b5d32853bb20da9fee51755f0c2bfe492de56b80

Observation 22148205-b5f1-43a3-98d7-96f2b2b93fc6 · outbound

This paper cites Superedge: Towards a generalization model for self-supervised edge detection.CoRR.

Video models are zero-shot learners and reasoners Superedge: Towards a generalization model for self-supervised edge detection.CoRR

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.877117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:4fb528a3cd288baa8d715480eb47c83e45d165b6b2af205b84f15e0ee63ae7c7

Observation a2962c2f-0e77-42b3-aef7-fcbac00a5771 · outbound

This paper cites Maze dataset.https://pypi.org/project/maze-dataset/0.3.4/.

Video models are zero-shot learners and reasoners Maze dataset.https://pypi.org/project/maze-dataset/0.3.4/

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.880750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:0626965c1f81a05fec1b7e98aa44d04501ddb478de784ae9f21ec380aaaa09e3

Observation ab715fa7-ffc0-4b98-b2c7-296648a4e7e3 · outbound

This paper cites 15 Video models are zero-shot learners and reasoners.

Video models are zero-shot learners and reasoners 15 Video models are zero-shot learners and reasoners

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.884397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:9e6f9e00e29476bc42a4cfa82a68e861f94c9c02b9d1f45cf96eb2e37f236281

Observation a2c54ea0-4e4c-4a4d-b816-275a92d5c5d2 · outbound

This paper cites Diffusion classifiers understand compositionality, but conditions apply.arXiv preprint arXiv:2505.17955, 2.

Video models are zero-shot learners and reasoners Diffusion classifiers understand compositionality, but conditions apply.arXiv preprint arXiv:2505.17955, 2

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.799062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:f0d2f3a2b5680f5f20d57610712bf10bf68f8bf0da3c86b9d52e3a4b3867d34e

Observation cddfd6bc-0e33-40b3-91b4-c2c9ba743516 · outbound

This paper cites Force prompting: Video generation models can learn and gen- eralize physics-based control signals.

Video models are zero-shot learners and reasoners Force prompting: Video generation models can learn and gen- eralize physics-based control signals

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:45.804128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:2358659df82560db3faf415a900d94e7d9a1577aa56ffc7f7616b0e49729e3ed

Observation f114cad5-3a23-437f-b7fc-9982b8b31f6e · outbound

This paper cites Motion Prompting: Controlling Video Generation with Motion Trajectories.

Video models are zero-shot learners and reasoners Motion Prompting: Controlling Video Generation with Motion Trajectories

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T02:16:45.808266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:39fff061ae1a9fd45c91150211788019907c472ed37dce35dc23b6eb8967beda

Observation 7856f664-b4a5-4927-b09e-3cfdc3488dd3 · outbound

This paper cites target image.

Video models are zero-shot learners and reasoners target image

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.888735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:54719c32069f1e215819f0b487ff40d428b474e59bf884c5e46b6e3d17635f1a

Observation 7ba182f8-0dc4-48e4-af9f-9607b4e24084 · outbound

This paper cites Focus on the object {color}.

Video models are zero-shot learners and reasoners Focus on the object {color}

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.892324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:1b82b9b57b24c0c65eac552841c14d040a90f7c20c3fc3b1b092111dca50ef90

Observation b8e22924-356d-4135-82a2-70431328e59a · outbound

This paper cites For example, if the target image shows a dog, and the choices show a cat, the object types are considered different.

Video models are zero-shot learners and reasoners For example, if the target image shows a dog, and the choices show a cat, the object types are considered different

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T02:16:45.896481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:6a65751dcaa85593b3645752940380702dfb3c2841fa3e01c1484ad3ede7d29a

Observation 70fc672a-27aa-4c97-86d8-861cd9af2f00 · outbound

This paper cites Final Answer: [answer].

Video models are zero-shot learners and reasoners Final Answer: [answer]

Reference 98

Resolution
malformed identifier
raw_fallback, observed 2026-05-14T02:16:45.911481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T02:16:45.554252Z digest=sha256:0158d3e522a72b0f3b9e6065aa1beb0fdce1fd4cd7cd99f4f5096207902d1398

Pith citing papers

Observation 00d0a81b-6dc7-4915-9968-81d4395ce2ad · inbound

Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging? cites this paper.

Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging? Video models are zero-shot learners and reasoners

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-18T07:26:02.994509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T07:25:04.166352Z digest=sha256:cd18b15741ee4e5689847746afd9c3137b00dc7ade37e22865998739dde6d6dc

Observation 5deacf6f-7f5c-4b23-8a73-47a005454c83 · inbound

Epipolar Geometry Improves Video Generation Models cites this paper.

Epipolar Geometry Improves Video Generation Models Video models are zero-shot learners and reasoners

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T08:20:10.044650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:20:10.044650Z digest=sha256:74b336136cf2f8c76ad8c31d7ff500d561484549f3eea7023f0843b1eed19921

Observation 9980435e-eb6a-4904-aba3-28082a0df574 · inbound

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm cites this paper.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Video models are zero-shot learners and reasoners

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.077440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:661e134ebaa26abac2e5334538cf478f514ed7a2bac04a913e3795c57b3d9b0c

Observation 6e03a2a8-bdb5-47e1-b374-98b9b0fffd5d · inbound

Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets? cites this paper.

Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets? Video models are zero-shot learners and reasoners

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:02:04.211945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:00:20.895672Z digest=sha256:2d5a5a9ac51f6bf88cd2d13c69820c35aa0a8d163807d660df081a4246ac0ea0

Observation bb196b9e-e861-4f40-abcd-3af5ff1f6414 · inbound

PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding cites this paper.

PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding Video models are zero-shot learners and reasoners

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T20:16:13.958140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:16:13.958140Z digest=sha256:1e6b6b1e992d00b8b42c79521b441c1d182241bae99f12c49b5f08a079444bc1

Observation a76490f6-51a5-4fbc-a34f-6206a431792d · inbound

Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos cites this paper.

Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos Video models are zero-shot learners and reasoners

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T19:11:52.916096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:11:52.916096Z digest=sha256:a5b60ac2cd5b48296fca28dfdca5ba402939a7179b2684f7507d83a957eb13cd

Observation 0aeee703-2ac3-4b4b-beac-aedc66bd6505 · inbound

Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation cites this paper.

Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation Video models are zero-shot learners and reasoners

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-16T18:17:55.050614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T18:17:54.943863Z digest=sha256:0ba216140ef46ae81a00b1519c78d96d732bfe754bdf3a0c3683549610b32a8e

Observation 297246b8-0278-49a7-862c-6a00eee4c807 · inbound

VideoCoF: Unified Video Editing with Temporal Reasoner cites this paper.

VideoCoF: Unified Video Editing with Temporal Reasoner Video models are zero-shot learners and reasoners

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:08:43.190117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T00:08:09.479706Z digest=sha256:2f0650522e6128866e45aafb155cae8d9f4246238f69a7973229f2c671f2ca49

Observation ebe77d37-8928-48fa-8bc9-0858cb0472b3 · inbound

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling cites this paper.

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling Video models are zero-shot learners and reasoners

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:29:56.460133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T14:29:56.348733Z digest=sha256:a76d27d84e979d0138e5e90f7077dfb7b27787d9691d1ce43f8bac54873e7dca

Observation c9dc0b19-4e8f-4c0e-a3cb-92c671ce8b7e · inbound

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling cites this paper.

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling Video models are zero-shot learners and reasoners

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T16:11:17.633640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:11:17.633640Z digest=sha256:88a26f331dac0b74bed120e5bbdccef4562578e0638be513499a8430dd5077b6

Observation e49a4b71-0c03-4438-a892-7a4551b8a785 · inbound

mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs cites this paper.

mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs Video models are zero-shot learners and reasoners

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:41:00.243691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T10:41:00.142543Z digest=sha256:1a1e4b0f1f03a0ac87181610290899ea58441791103db76b273f57f524192eb7

Observation f80e95d7-ac90-4e57-ba48-0adcc831f0e5 · inbound

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling cites this paper.

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling Video models are zero-shot learners and reasoners

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T15:46:12.576413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:46:12.576413Z digest=sha256:eb117dd37bb17d3f9a546c83925c10856bd7d408f23d12ea130fd06a907001f8

Observation 7c56be35-e66c-4d41-bbbf-538729048392 · inbound

Kling-Omni Technical Report cites this paper.

Kling-Omni Technical Report Video models are zero-shot learners and reasoners

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:00:58.607229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T21:00:58.473043Z digest=sha256:4edd46e6ec081ffff1abd925aee8718a9975bd5db1c6689c78a3ecfd53e9a348

Observation eb08817f-69c5-4bf1-a53c-249ee7977af4 · inbound

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World cites this paper.

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World Video models are zero-shot learners and reasoners

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:38:21.076677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T19:34:39.518649Z digest=sha256:84cab5758a0ec07d716a07827bcfaef01771446153001c17787711dcbc46ffb4

Observation c412de83-6508-4305-90d1-588dfac72819 · inbound

Rewriting Video: Text-Driven Reauthoring of Video Footage cites this paper.

Rewriting Video: Text-Driven Reauthoring of Video Footage Video models are zero-shot learners and reasoners

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:08:02.125676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T15:04:59.050484Z digest=sha256:faf9e1046e89c32eecb3c805e742fa63f153c3ff01173d40a0bc81bca6792255

Observation 8cff95ed-22f9-420a-a778-817638137559 · inbound

Walk through Paintings: Egocentric World Models from Internet Priors cites this paper.

Walk through Paintings: Egocentric World Models from Internet Priors Video models are zero-shot learners and reasoners

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T08:56:10.974820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:56:10.974820Z digest=sha256:4dba8be429bbf865e4a5036708c1e51197173d2c5c69974898a1f20b8cc8e3ae

Observation 7e1b603e-3425-4ea3-a639-ce4b70dcfe05 · inbound

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery cites this paper.

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery Video models are zero-shot learners and reasoners

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T05:23:36.362402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:23:36.362402Z digest=sha256:f15c420ba7036a8b1456f0c7a7fd060a9040d8f389f610816244771070c1f2d5

Observation 42e9d9c3-35f6-4d26-9826-aeb083b53ec1 · inbound

PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation cites this paper.

PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation Video models are zero-shot learners and reasoners

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:30:12.715082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:26:57.909496Z digest=sha256:f04d79b2023452e1bc1d636493c26eff5e373b8dfcc77e77439a1a0f9a4856ef

Observation f02c4f4e-950a-42bb-bc05-791687929a38 · inbound

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos cites this paper.

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos Video models are zero-shot learners and reasoners

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-16T17:02:34.202125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T17:02:33.997887Z digest=sha256:5e899ddb7b832d7a5cd1e9c97c12f3f21ab10514ea3b6c05a79d4ba820f0e2b9

Observation d4271ad1-e621-4140-b318-795f780d4f68 · inbound

Olaf-World: Orienting Latent Actions for Video World Modeling cites this paper.

Olaf-World: Orienting Latent Actions for Video World Modeling Video models are zero-shot learners and reasoners

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T01:20:05.597563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:20:05.597563Z digest=sha256:fff3efd66742ff5a321f0c41fc74a62bafc1b629ea6ca31777a3ef6b7bf12a2f

Observation 9b18cd75-4f0e-42f7-9e8b-5b7d318d266e · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning Video models are zero-shot learners and reasoners

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:accfbed8a96468996b7febd23ebf04cca43a3a6dc2399392a9a1d7cb46585b14

Observation 235c9488-a005-442c-aae0-b67a4bbaab26 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning Video models are zero-shot learners and reasoners

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:59.549335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:59.549335Z digest=sha256:517642a8958083375f19a835bc752576adb579e8e6eedf05e3ce65fb06162234

Observation 1b78b9f2-52b7-4f9b-81f4-541eb2d1fe70 · inbound

Generative Control as Optimization: Time Unconditional Flow Matching for Adaptive and Robust Robotic Control cites this paper.

Generative Control as Optimization: Time Unconditional Flow Matching for Adaptive and Robust Robotic Control Video models are zero-shot learners and reasoners

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T08:39:52.189510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T08:39:03.661087Z digest=sha256:b38749709b42f6fc96840da2b0a12e346b38c5b14ec2075b0cca1ddf79fb5606

Observation ba64a69c-d493-454d-91a5-a60a73667133 · inbound

Pretrained Video Models as Differentiable Physics Simulators for Urban Wind Flows cites this paper.

Pretrained Video Models as Differentiable Physics Simulators for Urban Wind Flows Video models are zero-shot learners and reasoners

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:05:11.529257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T07:01:08.953346Z digest=sha256:8a7a826830dd40020277b7c6afce933d8051ea90f6a37e52e62fa9569c1fbdf7

Observation a5cc8712-4230-45de-b03c-9803796d2771 · inbound

Stepper: Stepwise Immersive Scene Generation with Multiview Panoramas cites this paper.

Stepper: Stepwise Immersive Scene Generation with Multiview Panoramas Video models are zero-shot learners and reasoners

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:17:59.786293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:13:40.764990Z digest=sha256:7b2f6a0bf3b19751b6b95d8dc8ff319aab71d1d84e530e9fe4cbda2d3a1686fd

Observation 46d4b097-e5b1-45cb-9639-013698813b3f · inbound

LivingWorld: Interactive 4D World Generation with Environmental Dynamics cites this paper.

LivingWorld: Interactive 4D World Generation with Environmental Dynamics Video models are zero-shot learners and reasoners

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T14:17:59.568752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:17:59.568752Z digest=sha256:b33d6b23a5491e59698a7e1a65565bb12365ba4d761502a386c39e1d1cb4c381

Observation 925b46ba-2822-46e1-8030-7a14036a260e · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Video models are zero-shot learners and reasoners

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T19:36:42.100191Z digest=sha256:685d903723a1ea0e0411abfa82b211dc2b06ad420a44f3977e13666c31b0bde0

Observation 243eeb01-d7d0-4e85-8212-34c02326e6c8 · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Video models are zero-shot learners and reasoners

Reference 132

Resolution
unresolved
no resolver link, observed 2026-07-13T09:42:23.808691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:42:23.808691Z digest=sha256:f9f8a4033fe287170402c3038bd8eae95b745051d6f7bdd67a8b1be0839424fa

Observation e17e2b26-c9dd-43f8-9081-405113acdb3f · inbound

Neural Computers cites this paper.

Neural Computers Video models are zero-shot learners and reasoners

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T20:01:09.594072Z digest=sha256:5c62011af57800073e0f63695ced940fbaaa7c2dadc7a53520590b83dfc21b8c

Observation 1cf11679-e11a-4b4b-a1b7-60848a6e8e6b · inbound

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis cites this paper.

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis Video models are zero-shot learners and reasoners

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:16:31.378588Z digest=sha256:a432de5a912e250b447d97cedc407845e562c1974ca04f190f553de3bf57b55a

Observation 984ba074-04f9-4fe4-b1f8-ff33c0703bb3 · inbound

Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories cites this paper.

Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories Video models are zero-shot learners and reasoners

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:19:50.116736Z digest=sha256:4d53da26598aef9ebcb0959e53c900d295f8643a14ea813300b1b6451a6242bf

Observation 87becae0-5a4f-45a6-85cd-7f1ed152188d · inbound

Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories cites this paper.

Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories Video models are zero-shot learners and reasoners

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T23:14:56.161802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:14:56.161802Z digest=sha256:f92f7880e15a8123dfd205b3e0965ae245b1667154aa9c80aff06a84df3d22c5

Observation c4aeb11e-a523-4b42-9660-0de421449b27 · inbound

GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models cites this paper.

GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models Video models are zero-shot learners and reasoners

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:38:10.956726Z digest=sha256:3a3dee454fdf7aaf8fad6450f923055eaa5070d1a33eadd01efa95fdfb4d2709

Observation 05dc2edd-95c2-4580-b303-04baf8f8cb4e · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Video models are zero-shot learners and reasoners

Reference 185

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:276ebb862ada40413597a64d8562cd7a99246b69ec4bd8adf6e3162eef30e134

Observation ac169a5b-0bba-4a6a-b831-076d4fa5b746 · inbound

VibeFlow: Versatile Video Chroma-Lux Editing through Self-Supervised Learning cites this paper.

VibeFlow: Versatile Video Chroma-Lux Editing through Self-Supervised Learning Video models are zero-shot learners and reasoners

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T13:15:35.258754Z digest=sha256:c318d9ab6351280fbf9658f8dc44c7834ac9104486afdcebc848e1a2c2f91418

Observation 1b480e70-761c-44a2-9e30-df8eb095ca8e · inbound

Motif-Video 2B: Technical Report cites this paper.

Motif-Video 2B: Technical Report Video models are zero-shot learners and reasoners

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:50:04.001693Z digest=sha256:79e345b9dccb48867c0853a66f4ef3485b7992a552b72f3f48bac43fbb647b98

Observation 735057b6-9c5a-4d3b-b333-b96c82f07fab · inbound

Motif-Video 2B: Technical Report cites this paper.

Motif-Video 2B: Technical Report Video models are zero-shot learners and reasoners

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-21T00:13:53.063109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T00:12:23.096146Z digest=sha256:bd7e2dcc9358aa70878065d5ec0f2708aef2b96986d013bfcb982a8a7e3ff263

Observation 23898561-f7fc-4679-b269-768414dbec70 · inbound

ViPS: Video-informed Pose Spaces for Auto-Rigged Meshes cites this paper.

ViPS: Video-informed Pose Spaces for Auto-Rigged Meshes Video models are zero-shot learners and reasoners

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T05:39:49.102468Z digest=sha256:ea8f74579c75629035bb98761e49c6eb2be78a10e830037fc3766917847b5c96

Observation 0c5c945e-32b6-4df0-b193-7bbd5026eb61 · inbound

ViPS: Video-informed Pose Spaces for Auto-Rigged Meshes cites this paper.

ViPS: Video-informed Pose Spaces for Auto-Rigged Meshes Video models are zero-shot learners and reasoners

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T11:31:29.234533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T11:31:26.240607Z digest=sha256:ac676fd397fe923d627c0f9d4dd68f11ec6ad8bc8e0fc9b587ea5a1e0587c860

Observation 32b338c2-0eae-413b-b304-250bb2a6c282 · inbound

Grokking of Diffusion Models: Case Study on Modular Addition cites this paper.

Grokking of Diffusion Models: Case Study on Modular Addition Video models are zero-shot learners and reasoners

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T05:28:08.221886Z digest=sha256:0489e0619abd881b1df6bf1762eee48f14f369268d3b27af1ff9142757826a5a

Observation 47e19029-4357-4292-be91-2dfa8b99743b · inbound

Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation cites this paper.

Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation Video models are zero-shot learners and reasoners

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:49:44.525757Z digest=sha256:2cf1248116173fc0c4298dd8e8d5fe2fe29c44c4e9d0211492c7fd39b18f5c0a

Observation 4daf5438-bf57-48ba-a6d8-7872d39fc58c · inbound

How Far Are Video Models from True Multimodal Reasoning? cites this paper.

How Far Are Video Models from True Multimodal Reasoning? Video models are zero-shot learners and reasoners

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:320bf24cd44fd239d08b79ab14eb443809509a416088caf288e30b379f0ee09c

Observation 3b411e37-1b73-450a-b891-3cdfb1d70d9e · inbound

Image Generators are Generalist Vision Learners cites this paper.

Image Generators are Generalist Vision Learners Video models are zero-shot learners and reasoners

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:14:05.034951Z digest=sha256:c856f6072e467f25bbf1a97702afb3d270903bbbe0767e7f4f8fc364060d9fe9

Observation 45cefcfd-2ef4-4145-93d7-257f374f74b8 · inbound

Image Generators are Generalist Vision Learners cites this paper.

Image Generators are Generalist Vision Learners Video models are zero-shot learners and reasoners

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:45:14.717890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T07:40:46.090808Z digest=sha256:a2e73809fb122d331a2bb9b5007102ad2e35c51e68b4eb14bfbd51128a6dfbf4

Observation 18e25512-3d4f-42c9-ad82-d92c2eefc363 · inbound

Open-Source Image Editing Models Are Zero-Shot Vision Learners cites this paper.

Open-Source Image Editing Models Are Zero-Shot Vision Learners Video models are zero-shot learners and reasoners

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T16:44:53.286118Z digest=sha256:bff3bf8a5cc591f755a3298d66ed8a9ba7c1ebeae804630938a8bc6807c6c89c

Observation 0b894adf-48d8-4de1-b4e2-34205f2d0048 · inbound

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency cites this paper.

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency Video models are zero-shot learners and reasoners

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T13:30:47.547011Z digest=sha256:d3573e06977973f76bdf281603bfde86f60ca86989f875bc0c1b04e39e9da034

Observation cf7d03c2-7f76-4988-a158-239e9d3dcd04 · inbound

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency cites this paper.

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency Video models are zero-shot learners and reasoners

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:10:23.616363Z digest=sha256:18ffc6f9f84b7b95452f7cf38049775f546459e096534624988c9a72091b71c6

Observation 389e8009-68a0-4f98-92c5-77ccaefd2b8c · inbound

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency cites this paper.

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency Video models are zero-shot learners and reasoners

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T06:58:57.520289Z digest=sha256:f1b40576d6b2f0f537a3c50be468935f8aa3f6c0d091f024edf035b218cbb9ac

Observation 04aa48ef-e24c-40e1-88ea-b60f033b90f2 · inbound

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency cites this paper.

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency Video models are zero-shot learners and reasoners

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:25:44.802759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T23:25:13.965611Z digest=sha256:4f6d4bf39b46ddc880ec5c359cb77b08e4d24557695c0b5bfb2e7e74de96d4e9

Observation 697e63d6-4fd3-42be-9818-d8fab8f63292 · inbound

CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models cites this paper.

CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models Video models are zero-shot learners and reasoners

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:17:30.227755Z digest=sha256:f23f19e9360b5706e8210d90f658c5c4ea7e4bb8c7df8c623b7359087d748c18

Observation 5be5ac0f-af33-4fd4-8469-02e0762a63c6 · inbound

Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs cites this paper.

Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs Video models are zero-shot learners and reasoners

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:57:57.791351Z digest=sha256:7b7eee2b82fe485b22071c6ecc292dc1a11de2b3114ce4d7d1c6eb59cf7ee368

Observation 32b52f5c-3574-4a18-aa09-c33c549bf975 · inbound

Do multimodal models imagine electric sheep? cites this paper.

Do multimodal models imagine electric sheep? Video models are zero-shot learners and reasoners

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:24:01.933339Z digest=sha256:7f44beb1848ee0d57cbc5ee6bb84e111d9ffa65b81e1aba0c52c17c357f1511b

Observation 145f546c-a99d-4267-aeb7-7ef1b2de59ae · inbound

Progressive Photorealistic Simplification cites this paper.

Progressive Photorealistic Simplification Video models are zero-shot learners and reasoners

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T03:57:50.443100Z digest=sha256:6caeeca0886e9c7b7e040b7661ed714fd911d62137c559c4a3db349a42635f01

Observation 425ab599-d4ba-4dc6-8df2-a395cc433f46 · inbound

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors cites this paper.

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors Video models are zero-shot learners and reasoners

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:16:46.105477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:39:01.254916Z digest=sha256:cab814636bd425b87745c604ae54db806a96cb935e03e838f3d3ae2bd5f3561f

Observation 99de3e62-1ccb-4dcf-81dd-0fa80f3c5d32 · inbound

Video Models Can Reason with Verifiable Rewards cites this paper.

Video Models Can Reason with Verifiable Rewards Video models are zero-shot learners and reasoners

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-19T15:07:37.122050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T15:03:14.894952Z digest=sha256:cb26926cd56a58777f6c5820de2eca614859c7006eb87190cd8e4c9ad00ec3de

Observation 5d5396a0-1d46-41f5-8166-401dbc58a4e7 · inbound

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration cites this paper.

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration Video models are zero-shot learners and reasoners

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:18:18.293926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:15:54.413960Z digest=sha256:2ee950d1dad71c1702d461d18993688c86a78cc2a0216bcb09b97345351c6d0e

Observation 030c2cfc-0d4b-4ecf-acc3-44152f3c6e53 · inbound

RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic Manipulation cites this paper.

RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic Manipulation Video models are zero-shot learners and reasoners

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T12:38:17.021889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T12:34:11.744980Z digest=sha256:ed28ace6a19dfa20671cd0a0504f4bdaa3b2c71844b12ca9d2b5c878a778589d

Observation bacafae6-d485-41d0-bd49-fa5299c96311 · inbound

GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation cites this paper.

GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation Video models are zero-shot learners and reasoners

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T11:08:13.567132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T11:06:09.367559Z digest=sha256:d3a035c93d557e7a9d2288de933419589393d315e2493442b3b2d24f7ad82046

Observation 69863998-34f2-4cb3-8c98-2f504caefe43 · inbound

PhyWorld: Physics-Faithful World Model for Video Generation cites this paper.

PhyWorld: Physics-Faithful World Model for Video Generation Video models are zero-shot learners and reasoners

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:33:07.612602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T07:28:20.248452Z digest=sha256:dd25a1c02a751eb8fc849d0e07beae56f7d87eca60961bda137032f791a99ff2

Observation 1bffb555-5f83-4d1b-98dd-642440800660 · inbound

Scalable, Energy-Efficient Optical-Neural Architecture for Multiplexed Deepfake Video Detection cites this paper.

Scalable, Energy-Efficient Optical-Neural Architecture for Multiplexed Deepfake Video Detection Video models are zero-shot learners and reasoners

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:33:05.467216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T06:30:12.892474Z digest=sha256:92cb24900ff8a5f42c0150d6c25cc5a313023b5cd478b78e55125b5e4626a29f

Observation 29bc56db-8e81-4077-b4cb-7a4792f2284f · inbound

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis cites this paper.

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis Video models are zero-shot learners and reasoners

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:14:42.762103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T07:12:02.612292Z digest=sha256:67925df7292dfa8581f548efdd7ce23cf59d6239a4c0dcd3bd8a1f451a912046

Observation 6c87aaae-e097-4264-b3d9-d151a9ac0836 · inbound

MotiMotion: Motion-Controlled Video Generation with Visual Reasoning cites this paper.

MotiMotion: Motion-Controlled Video Generation with Visual Reasoning Video models are zero-shot learners and reasoners

Reference 93

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T05:54:38.705592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-22T05:51:17.890788Z digest=sha256:22ae98f42be940979137650b2d152d4935c2d1395a2c9025bae01a4e28e479ec

Observation e8e8d0ad-deb8-487e-8684-fefbe74f49a3 · inbound

Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation cites this paper.

Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation Video models are zero-shot learners and reasoners

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.894400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-29T18:17:52.284353Z digest=sha256:26a75253d01f7fd47bed495513d554b64022b67d794ecfbecb05384445eee438

Observation 74fa3fa4-a446-41b2-b1b6-4b3f4c61aac9 · inbound

Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players cites this paper.

Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players Video models are zero-shot learners and reasoners

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:53:29.102899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T13:34:03.020051Z digest=sha256:2b4438eb3bfc2b2a962576acc805e5d0de4d38f3e271dac47fbf1aff4c8ae39a

Observation 76e35556-fa49-42d9-8c92-057d85e24d5d · inbound

StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement cites this paper.

StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement Video models are zero-shot learners and reasoners

Reference 110

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:25:59.817219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:45:58.629263Z digest=sha256:73ae8edad3a6e8c3c43d41f7d884d93121a0d93a45ca54ddc70dfa080124da80

Observation 1dc5ba28-b9fc-4ef2-bf5e-601f76404a6c · inbound

OptiWorld: Optimal Control for Video World Generation under Physical Constraints cites this paper.

OptiWorld: Optimal Control for Video World Generation under Physical Constraints Video models are zero-shot learners and reasoners

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:32:35.523797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:02:51.848742Z digest=sha256:4beb1c97b8906f854bc3e51441a10ec1e2ed2e242e43dc100a0dd08d765c37bf

Observation 1b146f1f-0cb9-4312-8b2c-a0c5c6d7bdc2 · inbound

AlbedoEdit: Unified Instance-Level Video Editing with Albedo Guidance cites this paper.

AlbedoEdit: Unified Instance-Level Video Editing with Albedo Guidance Video models are zero-shot learners and reasoners

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:56:16.370934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T15:55:23.583463Z digest=sha256:96c32c56f29e5d0916811ed5c1a826970198cc63ea33f038ee29d5d16219bbb6

Observation 871081c3-33ae-4a0e-ba68-9dc2fedbbb32 · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization Video models are zero-shot learners and reasoners

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.218172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T15:26:21.284810Z digest=sha256:f543df87d8b1c07c7706c66faa5b65b1ae9cd2dca5bfca3f3d1e94f794225405

Observation 997a36fb-569a-42a8-b392-62f02be18087 · inbound

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization cites this paper.

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization Video models are zero-shot learners and reasoners

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-30T10:44:36.802140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T10:38:22.619277Z digest=sha256:4b0adb14148b398622698d06d3ae40882b1f31ff415abbda6a7945f0a5d7fb1d

Observation 0ed2161f-86dc-48bf-80ec-fcf968de0120 · inbound

Cosmos 3: Omnimodal World Models for Physical AI cites this paper.

Cosmos 3: Omnimodal World Models for Physical AI Video models are zero-shot learners and reasoners

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:46:18.400971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:a579353002a6d0fd9d586fefec0c847942baf44b7b2771ed6e9b0431cb11e4bd

Observation 52cef885-4b75-4b3c-8992-f7eeac9d89a1 · inbound

PointAction: 3D Points as Universal Action Representations for Robot Control cites this paper.

PointAction: 3D Points as Universal Action Representations for Robot Control Video models are zero-shot learners and reasoners

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:16:34.972421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T10:09:48.280446Z digest=sha256:f84764cad0e13b2e5b875757a3b845e2de76bf9e941bf890e5dc809140ca5f15

Observation 59b46e92-a43b-454f-8eab-708b51a08e82 · inbound

OmniTryOn: Video Try-On Anything at Once! cites this paper.

OmniTryOn: Video Try-On Anything at Once! Video models are zero-shot learners and reasoners

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:17:26.380398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T19:00:44.243335Z digest=sha256:f51caa36dc728a3dcbd3f42526142c4f81c5c30f9656990d6b448f1792d74111

Observation 851a4ec4-c9d4-4a56-922d-7e0f3139c5a7 · inbound

Data-Driven Automation cites this paper.

Data-Driven Automation Video models are zero-shot learners and reasoners

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:27:36.195390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T13:58:40.370152Z digest=sha256:f209c0921d597c4e225d1dfb13ca28a08535f6589b142f7889bc1d64049c7b18

Observation 6ce2e47a-a03b-4ff9-835b-5cad02df67d7 · inbound

Quo Vadis, Visual In-Context Learning? A Unified Benchmark Across Domains and Tasks cites this paper.

Quo Vadis, Visual In-Context Learning? A Unified Benchmark Across Domains and Tasks Video models are zero-shot learners and reasoners

Reference 99

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T05:07:39.113513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:24:48.138643Z digest=sha256:89281e93ff35b8189b21a107d75720773c6f5feda601e5eec891860c67608a13

Observation 2b1b0719-1c48-459f-b10e-4a5201d8595e · inbound

WorldOlympiad: Can Your World Model Survive a Triathlon? cites this paper.

WorldOlympiad: Can Your World Model Survive a Triathlon? Video models are zero-shot learners and reasoners

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:47:41.341692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:05:26.397711Z digest=sha256:387c1d2ae48a0e6c272a010d0d8d0a121a0f13cb76e559d8945e80b6b41bf586

Observation 6aad9cc6-d226-435e-90f7-3bc87b10614d · inbound

VICX: Generalizable Robot Manipulation via Video Generation and In-Context Operator Network cites this paper.

VICX: Generalizable Robot Manipulation via Video Generation and In-Context Operator Network Video models are zero-shot learners and reasoners

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:28:04.540519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T09:32:03.146878Z digest=sha256:6320236903f280cdd6a9bd813a1e2bcff2618630b17bf4865270b197a8b316a6

Observation 3cf84998-9a19-450d-a02b-c440560e3faf · inbound

World Model Self-Distillation: Training World Models to Solve General Tasks cites this paper.

World Model Self-Distillation: Training World Models to Solve General Tasks Video models are zero-shot learners and reasoners

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:07:55.828142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T10:16:35.511426Z digest=sha256:a467ad54bd11b1dca096c1832282ed5d658afb3eeadc417d6710dc921273e86a

Observation 2765f2ef-669a-46b6-8033-9fd7646e9c92 · inbound

WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation cites this paper.

WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation Video models are zero-shot learners and reasoners

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:38:34.121144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T06:22:47.622007Z digest=sha256:155086125506b667dd58cfad1455c4d52b0fda24d9381e362b7a131c701bcc98

Observation 900e083e-6dc4-4831-9237-075a15147f80 · inbound

NEXUS: Neural Energy Fields for Physically Consistent Contact-Rich 3D Object Dynamics cites this paper.

NEXUS: Neural Energy Fields for Physically Consistent Contact-Rich 3D Object Dynamics Video models are zero-shot learners and reasoners

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:18:43.360556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T04:25:27.067362Z digest=sha256:6558d6629e2823854ae22b65092c573bce50069cf24612b3fc98526ad9ba0338

Observation 24a22b57-39c7-4bfd-bb47-1c06a3136b9f · inbound

Physics-IQ Verified cites this paper.

Physics-IQ Verified Video models are zero-shot learners and reasoners

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:29:15.593809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:13:36.568700Z digest=sha256:369f13eba91df76615efbb0512392fcc6d78c187af0b8459cca873e9773915f4

Observation dfa7dfc3-070c-45e7-bb53-d6ad772fdd4d · inbound

FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation cites this paper.

FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation Video models are zero-shot learners and reasoners

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:49:58.078372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T00:11:11.534272Z digest=sha256:753c463d8576d0406c97420e984130e8f835ddc78bf9a67b484215bb36212497

Observation 0b66c0aa-f0f9-44b7-a1ef-ad9f4fd0a122 · inbound

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation cites this paper.

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation Video models are zero-shot learners and reasoners

Reference 75

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:19:51.050647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T05:16:53.011837Z digest=sha256:b80cd646f135959cb465d34848d01649ae359a4b7030a6adbd7239df4190ea1c

Observation 5759d396-7558-45f8-b65a-d45dedb64175 · inbound

A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks cites this paper.

A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks Video models are zero-shot learners and reasoners

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-07-01T15:45:48.900946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-30T00:59:50.478584Z digest=sha256:83cc7510ae19bf70a8407a595ab2f6a03358f2b4fc13fa7ec97966b02ca6944e

Observation 981ed76b-0c7d-4df0-a74e-595e213c2758 · inbound

Bridging Video Understanding and Generation in a Unified Framework cites this paper.

Bridging Video Understanding and Generation in a Unified Framework Video models are zero-shot learners and reasoners

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:05:40.347769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T05:57:54.653504Z digest=sha256:3d909625d544144cbbc1fe1603ee4e4b5e3cf5e2fd8c3feb1d1f4061fce6ec6e

Observation e8ddcdf1-dfc9-41a5-b97d-2a8bbc1f8a1d · inbound

Global Pose Control for Generative View Synthesis in Normalized Object Coordinate Space cites this paper.

Global Pose Control for Generative View Synthesis in Normalized Object Coordinate Space Video models are zero-shot learners and reasoners

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-12T07:34:02.473153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:34:02.473153Z digest=sha256:25d368bf6e9fb3a75691f4e9095a4aab1985f9458f5586a1d91b7b9b4cf2c152

Observation 0db48952-a698-4bfa-939b-d25a5512eb4c · inbound

Self-Improving Diffusion Classifiers with Minority Preference Optimization cites this paper.

Self-Improving Diffusion Classifiers with Minority Preference Optimization Video models are zero-shot learners and reasoners

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-12T00:04:25.954963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:04:25.954963Z digest=sha256:40a8dc1a537d5a9326889dd96e873c875fc8c430da4b9dfe6b26cb1eafe2c4e5

Observation b8277d48-383d-42a7-9771-029834d3907c · inbound

Video Generation Models Are Inherent Lighting Estimators cites this paper.

Video Generation Models Are Inherent Lighting Estimators Video models are zero-shot learners and reasoners

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-11T15:23:22.681709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T15:23:22.681709Z digest=sha256:11f561356aa307be7d59c8533eb5d8f283f87247b7277454285db73f5252ac7c

Observation 0c660030-abd4-41db-ab64-750e25cad946 · inbound

Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models cites this paper.

Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models Video models are zero-shot learners and reasoners

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-11T07:04:37.865482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:04:37.865482Z digest=sha256:c3b3945930247adad30aa95af8c9f938d5840ac5d5772303d09df7a694369337

Observation aa4ea3b1-c690-484e-a86f-6a14c1f70746 · inbound

Gen4U: Unifying Video Generation and Understanding via Diffusion cites this paper.

Gen4U: Unifying Video Generation and Understanding via Diffusion Video models are zero-shot learners and reasoners

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:26:39.639035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T00:16:43.190961Z digest=sha256:9c2e146afdf5cdb3f8439202005ea3ed2fbca35c8f17757bee9714d9617e600f

Observation 8657b9ef-67e8-4500-a5f8-38582f8953d2 · inbound

OpenCoF: Learning to Reason Through Video Generation cites this paper.

OpenCoF: Learning to Reason Through Video Generation Video models are zero-shot learners and reasoners

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.972750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:7ad052fed21569d116d8006a76ed49b3e3e03c72554306f5e2a942ada8df2414

Observation 8e37c9d8-47b1-4b2e-aff0-3e4f726cd7f5 · inbound

Video Generation Models are General-Purpose Vision Learners cites this paper.

Video Generation Models are General-Purpose Vision Learners Video models are zero-shot learners and reasoners

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-13T00:56:18.867382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:56:18.867382Z digest=sha256:6d141994066dbae0af4e4f3f5958639d3701f51d36a3884f8c7140d4a26c5593

Observation cb1be0d7-326c-4df3-92ca-1544e8a006fb · inbound

From Pixels to States: Rethinking Interactive World Models as Game Engines cites this paper.

From Pixels to States: Rethinking Interactive World Models as Game Engines Video models are zero-shot learners and reasoners

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T02:52:18.530088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:52:18.530088Z digest=sha256:558345f0b4801fd1029645967db255f0ac3fa54cdf7bf8e7314edba3e2094fca

Observation 8c482106-cb9a-469d-9c97-c392dce55833 · inbound

Privacy-Aware Synthetic Video Benchmarking and Relational Evaluation for Worker-Under-Suspended-Load Detection cites this paper.

Privacy-Aware Synthetic Video Benchmarking and Relational Evaluation for Worker-Under-Suspended-Load Detection Video models are zero-shot learners and reasoners

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T22:40:12.822532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:40:12.822532Z digest=sha256:8b0a6f387141b52eb23978819bf2752351dcd8fe05f771fadb41a85332285491

Observation 6792c75c-d84d-4134-b62d-4259ad71952f · inbound

Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence cites this paper.

Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Video models are zero-shot learners and reasoners

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:27.838962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:27.838962Z digest=sha256:1f23b8c5c25e4ad7b0ab4109214980f65ffd64912032d5ee9da05f87ddb76f44

Observation 3250ba63-0662-4c3d-a75d-c3d83c42d811 · inbound

Between Safe Boundaries: Exploiting Temporal Consistency for Jailbreaking Text-To-Video Generation Models cites this paper.

Between Safe Boundaries: Exploiting Temporal Consistency for Jailbreaking Text-To-Video Generation Models Video models are zero-shot learners and reasoners

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T18:33:06.207028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:33:06.207028Z digest=sha256:69c2ad94611acab5114d262b982e7f72f1b6a92ab42ce79e24b50ae47f780e45

Observation 127f1619-a1a4-436b-bd5c-7b1487b9bdc1 · inbound

Diffusion ReRoll: Revisable Denoising for Robotic Sequential Prediction cites this paper.

Diffusion ReRoll: Revisable Denoising for Robotic Sequential Prediction Video models are zero-shot learners and reasoners

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T11:24:29.782087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:24:29.782087Z digest=sha256:94421f6823764ee5494b28735bef29776dc44616be86869abbc41ddd684342d5

Observation 9846cd65-116c-4b57-9aa4-0b11093d535f · inbound

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation cites this paper.

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation Video models are zero-shot learners and reasoners

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T20:13:33.415740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T20:13:33.415740Z digest=sha256:f9bd67c5003e9594dffc1ec4b24dba94d59941918520435a3ad76b784715deee

Observation 0654ae9d-0209-4ca9-b583-f60db8edc336 · inbound

Visual prompt engineering for video models cites this paper.

Visual prompt engineering for video models Video models are zero-shot learners and reasoners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T02:13:11.573564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:13:11.573564Z digest=sha256:2f4db488b27db6169814f9ab0030a7de4590ae2538c572473ed4fbf42bdef3da

Observation 6bf596d3-03c3-4ea5-b621-d9e868882fec · inbound

Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory cites this paper.

Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory Video models are zero-shot learners and reasoners

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-30T20:12:49.623361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T20:12:49.623361Z digest=sha256:522ea9da423783798e253b4cff9a33a487da3782d2dec1e253dcc7a6ecfea03f

Observation 06ddc16a-919d-4051-9d31-01dab965ebab · inbound

World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models cites this paper.

World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models Video models are zero-shot learners and reasoners

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T04:59:38.639705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:59:38.639705Z digest=sha256:1104827e3723328075bd87ae72d63d720b83a63cf1c42ca2b09791f74de160c1