Pith. sign in

Paper Citation Record · LEDGER

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

As of 21 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 6 inbound Pith citation observations for arXiv:2501.09781.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09781 v2

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:49:20.143466Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:38:30.073837Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T09:12:14.462558Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved35
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e41be6c8-24c6-4f54-a4ee-8221fa64521e · outbound

This paper cites GPT-4 Technical Report.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.817710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.817710Z digest=sha256:8ac8109828f6f83200ecefb9352dc6cd4ed4cb9999968a35e2f114400762f38f

Observation a8b8759c-ab4d-496b-b07d-4baae7c2c10a · outbound

This paper cites PaLM 2 Technical Report.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos PaLM 2 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.822689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.822689Z digest=sha256:033ee6d34966d956ccb270726ea5ed09b34c90262a3daa225ff98cea9140d51f

Observation f58e48ca-8e25-4261-a96b-1cb45f9f70d9 · outbound

This paper cites Qwen Technical Report.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.827287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.827287Z digest=sha256:810944418790056e934432dea2c33e8f0654f63348766e922c8c635ef2644af7

Observation 515a1532-fd31-431b-a2cb-601d882ae4cc · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.832831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.832831Z digest=sha256:e365ca57e4722b4d81b9a731b03d5c8db27caa956cb6c48ca8453e0fffa552ac

Observation 881021ae-2ff1-4ff5-b603-b94af3f0ca44 · outbound

This paper cites Knowledge Representation and Rea- soning.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Knowledge Representation and Rea- soning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:21.144987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.837748Z digest=sha256:8d4794df45d147a6cc81e86ffc85cda4c7ae08679021d1a220c731c2ddbd7109

Observation 9caf91c9-d8cd-45cc-aa48-69ea44b2f68b · outbound

This paper cites Video generation models as world simulators.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Video generation models as world simulators

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.842204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.842204Z digest=sha256:71842c9a1a31dda2ce0ca5244ca393251c8671bf095935044e93082169c47be9

Observation 5cce4b29-cded-4205-a80f-5008c7a0d4d2 · outbound

This paper cites Lan- guage models are few-shot learners.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Lan- guage models are few-shot learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.846448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.846448Z digest=sha256:bc0e4297d1abe07ee1186f8f7be94f0ca22a9eb4a835046ef90f46ae6cf8349f

Observation 28597d7f-e4c5-43a4-b6bb-fc51de01bf32 · outbound

This paper cites Ge- nie: Generative interactive environments.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Ge- nie: Generative interactive environments

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.850424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.850424Z digest=sha256:a6d466f8965a1e7ab806f51827e0d1c7da91f57250de9b60845472c80fca842f

Observation 1fe0800f-a28e-4f72-b322-e1adcd63bbad · outbound

This paper cites Learning by imi- tation: A hierarchical approach.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Learning by imi- tation: A hierarchical approach

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:21.105981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.854657Z digest=sha256:15a00e03995ad1242bf896cb76893a7391673c5f5199a82ad80b53c697cbfb94

Observation a43855a2-f1f9-49aa-84bb-6782bbf0529a · outbound

This paper cites Efficient se- quential decision making with large language models.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Efficient se- quential decision making with large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:21.091274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.858592Z digest=sha256:5be62cb9e6d5370fab647dc0312778416c8b07cbf1bc0f4221db60467b05df9c

Observation a0454179-a608-4425-a70e-548a0e3195f5 · outbound

This paper cites Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T19:49:20.400601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.862040Z digest=sha256:2d23c3a02e8afbeab4005a643cef211ad18cfd72966362368f7ab37b905ed76c

Observation 096e9a8a-fed9-4d04-aba3-bfb6c37b2b81 · outbound

This paper cites Generative pre- training from pixels.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Generative pre- training from pixels

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:21.076784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.866950Z digest=sha256:2c7fd32f2a65a083c132e3c69472be482a36d1c4d155fb5e077e9d6603e42163

Observation 3fc5a54f-7a8e-4615-af5a-294ca7fe3acb · outbound

This paper cites Palm: Scaling language modeling with pathways.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Palm: Scaling language modeling with pathways

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.871345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.871345Z digest=sha256:f10b812f5182ff9d23d9ac60d719ff7907bb91a96371e14f3f12cd55bd3334b0

Observation e6e2454e-38e7-486b-9a09-265d228d529b · outbound

This paper cites Whole-history rating: A bayesian rating sys- tem for players of time-varying strength.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Whole-history rating: A bayesian rating sys- tem for players of time-varying strength

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:21.055090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.875731Z digest=sha256:80314df981e188529864e24ab94b3d78b67f7b28779a57ad33fe26a91eb2d3ea

Observation 1ac38fb9-afd7-480a-82e9-3512f634f22d · outbound

This paper cites Tenenbaum, Dale Schuurmans, and Pieter Abbeel.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Tenenbaum, Dale Schuurmans, and Pieter Abbeel

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:21.040792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.879960Z digest=sha256:f151cdf018dfde23400b17ebb034e0067a82afaf46a74ab4578d098247537a59

Observation f8424e22-545e-4ca0-a20e-2012085075bb · outbound

This paper cites Tenenbaum, Dale Schuurmans, and Pieter Abbeel.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Tenenbaum, Dale Schuurmans, and Pieter Abbeel

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:21.026783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.884620Z digest=sha256:6a8d873271003f948713054b9eaa4f6290cb9f2d6e057d3595791a6f2000238e

Observation 1321d493-c8c3-46ca-a30c-f058ff72495d · outbound

This paper cites Determinants of LLM-assisted Decision-Making.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Determinants of LLM-assisted Decision-Making

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.888581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.888581Z digest=sha256:27587af336616bb252b234481ab12b0c1f4e2112ff030494d840bebfd6290805

Observation 31f74718-c0c3-4cd6-8fba-fc7c8e289530 · outbound

This paper cites Chessgpt: Bridging policy learning and language modeling.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Chessgpt: Bridging policy learning and language modeling

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:21.014027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.893518Z digest=sha256:8944f4d8f9c6b4a905ea308a67659100a91874c53732cb80d00d7ce2a8e95f1e

Observation ec04f2ad-f6b9-4cdd-a0c0-eeb80af971b3 · outbound

This paper cites Text with knowledge graph aug- mented transformer for video captioning.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Text with knowledge graph aug- mented transformer for video captioning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.897550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.897550Z digest=sha256:f87cdf978db358a36d13f817aa700c28a9fd55dcc0fb908b90b9f64cef72c5ef

Observation c305ee51-bb62-40cb-bf58-76653770bdd1 · outbound

This paper cites Context-guided spatio-temporal video grounding.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Context-guided spatio-temporal video grounding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.991321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.902586Z digest=sha256:416eeb37c067464973ed9fe8195b9f5c2f6aa4d84f3fcb7c0832fdc32587fc0d

Observation 24715193-772f-4b18-98aa-9a21339da084 · outbound

This paper cites I2v-adapter: A general image-to-video adapter for diffusion models.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos I2v-adapter: A general image-to-video adapter for diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.977244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.906780Z digest=sha256:81aaa7c0e87be408eb1a4a517de9f2b9fa724f8690a4cd0955cad1a24b79233f

Observation d38394b7-6777-4c9c-b638-0f37d3df3296 · outbound

This paper cites Rose: Revolutionizing open-set dense segmentation with patch-wise perceptual large multimodal model, 2024.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Rose: Revolutionizing open-set dense segmentation with patch-wise perceptual large multimodal model, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.962801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.911613Z digest=sha256:3238568affc57b3b43e5099189f329c65d8b06c807e4b764b2c51462cc7dd8bf

Observation d351d8bb-4310-4a0b-879d-69e4d1b00f30 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Teaching Large Language Models to Reason with Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.915886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.915886Z digest=sha256:038ef7bb1b4146c192af46ba095b72018ba1ac8da3f1ae5a8a49112b10f0b8d8

Observation 4ebfae70-2795-4484-9dbc-d0a891c0ffd1 · outbound

This paper cites Denoising dif- fusion probabilistic models.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Denoising dif- fusion probabilistic models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.920884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.920884Z digest=sha256:292b5bc544ae29d5eab671b7e98a053555ee9e5e7ed79bea2848a998da16e41f

Observation 33090167-de52-4f8d-a08f-61d76549a7a7 · outbound

This paper cites Social learning: an in- troduction to mechanisms, methods, and models.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Social learning: an in- troduction to mechanisms, methods, and models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.937878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.924855Z digest=sha256:78cf44dda4bffdb570bf2ed90e90c8a34c52670b0b901fe684527e69b4335655

Observation 1333289f-afcb-4695-9766-ffaa06073270 · outbound

This paper cites Language models as zero-shot planners: Extract- ing actionable knowledge for embodied agents.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Language models as zero-shot planners: Extract- ing actionable knowledge for embodied agents

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.923223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.929601Z digest=sha256:c272165704dfcee9e0cb6a019ef9716413a743d017aa369940c280aba290aef5

Observation e508a831-83e2-4719-88af-a249efd2990a · outbound

This paper cites Language models as zero-shot planners: Extract- ing actionable knowledge for embodied agents.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Language models as zero-shot planners: Extract- ing actionable knowledge for embodied agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.933739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.933739Z digest=sha256:030928ddd037f61170f53900ca6e42d8ee00ed5333e0081afd817cc0188fcdad

Observation edd99fc8-d003-49f8-b546-f61a2d9817b0 · outbound

This paper cites an unresolved cited work.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:49:20.898972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.938302Z digest=sha256:ca06fc891ecc98a64661546a1a6445ffab9ea7d686f3d8325a04d68e1585474a

Observation 71fae37c-8ad4-47f0-a7c6-5ae9db03a36e · outbound

This paper cites Mistral 7B.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Mistral 7B

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.942574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.942574Z digest=sha256:a09f430f38c5abd53eaba3e5fa2b7d89e5d1cb4847b20974684fdccc1610d8db

Observation 8cef78a2-819c-4f32-af42-ad7486661ce3 · outbound

This paper cites Collaborative vision-text rep- resentation optimizing for open-vocabulary segmentation,.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Collaborative vision-text rep- resentation optimizing for open-vocabulary segmentation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.884179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.947417Z digest=sha256:c5d5f1e82e5d5647ccacef4c9596252a5271559c50a0e96dc01f8d6bd77a08b7

Observation ccdbb9a3-d273-41fe-8c0b-a29d4b9ebb12 · outbound

This paper cites Auto-Encoding Variational Bayes.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Auto-Encoding Variational Bayes

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.951451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.951451Z digest=sha256:5e8728c342b1b12acbc2839d78410c3c2496c95d57f30bf74157a893439df1b4

Observation fcfddf05-edf2-41d2-9682-4e3037d7dd79 · outbound

This paper cites Large language models are zero-shot reasoners.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Large language models are zero-shot reasoners

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.871227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.955055Z digest=sha256:c9837686088c77b9c027adb7fe069069e25727240010d020b40a59e5672ca015

Observation e3c7bd50-bf1f-49c9-aac0-31f471257989 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.958690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.958690Z digest=sha256:8136cb88d3d23e52f9d1f871e1112f74e1181dec1f224e931f03034dacbd8d70

Observation 3d2140df-16be-4240-995a-cf04ad467630 · outbound

This paper cites Umap: Uniform manifold approximation and projection.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Umap: Uniform manifold approximation and projection

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.858822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.962270Z digest=sha256:2e6f2b12720ccd883c9b2961ef3df8a1f7552d97a2ae383f1195116acd1ac849

Observation cba34e5f-e41a-474f-b7da-f177a8836939 · outbound

This paper cites Emergent world representations: Exploring a sequence model trained on a synthetic task.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Emergent world representations: Exploring a sequence model trained on a synthetic task

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.844852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.965832Z digest=sha256:d138840deb1aa702b3f87036cfc716cbd0fe6fd6d0249fab235e9189a30ff3fc

Observation f03606a6-5714-439a-ae34-27e017aded36 · outbound

This paper cites DeLLMa: Decision Making Under Uncertainty with Large Language Models.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos DeLLMa: Decision Making Under Uncertainty with Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.970782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.970782Z digest=sha256:a11d8c3f9d098eff2986f60aa0e88140953af342d7d2f7e7a29ec8b616bf5bd4

Observation a4e8c22a-3f01-4f82-a606-4ccd2edb4b96 · outbound

This paper cites Decoupled weight decay regularization, 2019.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Decoupled weight decay regularization, 2019

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:19.975620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:19.975620Z digest=sha256:088d4107aad72e7efdc72805e27a88cb7ee42d82dadc9a989ce9834bcfae7b72

Observation e3b57fee-2a96-44d7-b818-6d71e15f135f · outbound

This paper cites Language conditioned imitation learning over unstructured data, 2021.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Language conditioned imitation learning over unstructured data, 2021

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.821068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.980225Z digest=sha256:cd756c77044e7e8780a8ce968e99b00af46a99ae1bf30b14cfaae282469fbcb7

Observation 4dbb2d9f-4b19-44aa-838d-48af7f3bebf1 · outbound

This paper cites What matters in language conditioned robotic imitation learning over unstructured data, 2022.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos What matters in language conditioned robotic imitation learning over unstructured data, 2022

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.806761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.984276Z digest=sha256:ebf948a621a99490ee1db9fb9bde08e67d46db5e112ff5ba6aa90d0147ed04a0

Observation cf4f358c-6b44-4410-ad28-6700c2917ab1 · outbound

This paper cites Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manipu- lation tasks.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manipu- lation tasks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.792358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.989119Z digest=sha256:eb9b2e700540a482032261042efd9e480e5d3fe46c7a6cbddb05d54ae9e2563c

Observation 81e4bcfd-0104-455d-8781-bf333e4ffdf9 · outbound

This paper cites Finite scalar quantization: Vq-vae made simple.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Finite scalar quantization: Vq-vae made simple

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.777926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.993247Z digest=sha256:de3fc6b4d070411cb22da7a34475af9234c7da36a5d710cabc782f357b7310ea

Observation c2f94ea1-14dd-49a7-9d48-a86c6dffbe1a · outbound

This paper cites Online-go.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Online-go

Reference 42

Resolution
parse uncertain
raw_fallback, observed 2026-08-10T19:49:20.765409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:19.998055Z digest=sha256:1f94a8660a2bcaab33e3c9ea9a73066123d1851452b38f78f923d94b61f37a1c

Observation eff7438e-e7d9-483c-8914-06750e306e25 · outbound

This paper cites Training language models to follow instructions with human feedback.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Training language models to follow instructions with human feedback

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.002150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.002150Z digest=sha256:0afe46667b5ced3745acec2d049004eaad3ce00f70f0ba1ad2d727ccbbb740e3

Observation cf8ce8f9-e524-4618-a986-974f52b51f63 · outbound

This paper cites Scalable diffusion models with transformers.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Scalable diffusion models with transformers

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.007090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.007090Z digest=sha256:c933b5cc26be74ad29da86d09902e1943d0681fa6499eb0e42d2e824dbf15fdb

Observation 5dffd414-fc1b-493a-ae33-9d02f8ad3b2c · outbound

This paper cites Rio: A benchmark for reasoning intention-oriented objects in open environments,.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Rio: A benchmark for reasoning intention-oriented objects in open environments,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.733523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.011605Z digest=sha256:fb09d6b90fa58fbfeaf966538afe46603b3dde34b21fc53ba346c537623ae069

Observation ce8854b8-4787-4137-8343-d8d1ef6fbacb · outbound

This paper cites Chatvtg: Video temporal grounding via chat with video dialogue large language models, 2024.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Chatvtg: Video temporal grounding via chat with video dialogue large language models, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.719309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.016250Z digest=sha256:ce1b8feeb658757560590d30afbf8144687f291c908eec7d6ed420cbaeaf3d10

Observation 7a0a349a-f395-4abb-9536-98d26809cac6 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model, 2024.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Pixellm: Pixel reasoning with large multimodal model, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.703983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.020421Z digest=sha256:bf81480fceac7a6d5022485b07df84c518c84afd4f8a40c9dbaa6c2154ba0f0f

Observation 525b558e-f441-4958-ae23-922fcef7a00c · outbound

This paper cites Lewis, Joel Veness, and Tim Genewein.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Lewis, Joel Veness, and Tim Genewein

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.689848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.025251Z digest=sha256:8d63c75e3eb9d981b42f8eae6a4396ba0b656347ea5ad7772ddfd5d5fe11553f

Observation 033ba0c8-65d4-43af-a29e-ddcfecccd477 · outbound

This paper cites Artificial intelligence: a modern approach.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Artificial intelligence: a modern approach

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.675287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.029532Z digest=sha256:05a19acfff8fea7ea9622ef57bc7db7a7b62159b73f8d9f1f4b9f613523773de

Observation ef2c1478-d408-4587-a798-fee96424e5c7 · outbound

This paper cites Learning to Act without Actions.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Learning to Act without Actions

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.034312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.034312Z digest=sha256:072027b7fd40594bcc869c1797d8676cf62025a0e5d881bf7cf631a5f70fc6b9

Observation 226e67cf-9024-4159-8235-dbc96947471a · outbound

This paper cites Large Language Models are Learnable Planners for Long-Term Recommendation.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Large Language Models are Learnable Planners for Long-Term Recommendation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.038665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.038665Z digest=sha256:6a544d88e32734bbb0dd8832c49fbce3028c64c34a32954df74c1d42899fe3d4

Observation 21417d79-4337-4cc7-8d95-dd1ca345e558 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Mastering the game of go with deep neural networks and tree search

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.661412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.043583Z digest=sha256:a42a34a27a4d41c1bcd92e46f2a61ef5bb20659d273d546bb83c545d8109e0f8

Observation 5a3e945f-d0d1-43bf-a353-310b7b133431 · outbound

This paper cites Llm-planner: Few-shot grounded planning for embodied agents with large language models.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Llm-planner: Few-shot grounded planning for embodied agents with large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.647864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.047967Z digest=sha256:b91ad3ba60909d82aa6ff4c6e26449aedae1a4afb94754f17dc9de68d2a66f9a

Observation 7f310ef1-80a7-4221-8eab-0260494737d6 · outbound

This paper cites Score-based generative modeling through stochastic differential equa- tions.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Score-based generative modeling through stochastic differential equa- tions

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.052647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.052647Z digest=sha256:8e8ab3363f623600c0b972ba0297307beffb4205031c5efd51418bda1b93af5d

Observation efb117ab-dbab-4d50-afb5-acc1564fd84e · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Gemini: A Family of Highly Capable Multimodal Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.056747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.056747Z digest=sha256:710722e3bb3eddb9f3e775333bd8a06e092faee8085cf81a44a5be50b8addedc

Observation 00be1d5a-28c7-4415-8110-90bd7f4435cd · outbound

This paper cites Llama: Open and efficient foundation lan- guage models, 2023.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Llama: Open and efficient foundation lan- guage models, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.625095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.061008Z digest=sha256:eb2f92236b855f43bbffdb9e85a756d930de46204e521a087803181ddad12c1c

Observation 476846c5-dd49-487f-a234-724be54a52b0 · outbound

This paper cites Neural discrete representation learning.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Neural discrete representation learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.612817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.064641Z digest=sha256:c2f43190730e1771267cace7f751460d79d3d6270b640efc06de9e5e77e1968c

Observation b40cdd45-0750-45bb-9839-0a85376824a2 · outbound

This paper cites Neural discrete representation learning.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Neural discrete representation learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.068261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.068261Z digest=sha256:baaf55af596becbe9bcb70b138aa9b5799381049354cb898e8b3467301df6b20

Observation 4ed976ce-5d4b-444e-bf0f-ccde6729576a · outbound

This paper cites Attention is all you need.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Attention is all you need

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.071768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.071768Z digest=sha256:82d788570f16927ee916c6eef1c7a55fa359f303f2741e88365398d0f8c5e71c

Observation 8e09dab6-d5f2-489e-bb03-ecbe6182dc2e · outbound

This paper cites World to code: Multi- modal data generation via self-instructed compositional cap- tioning and filtering.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos World to code: Multi- modal data generation via self-instructed compositional cap- tioning and filtering

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.580572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.075247Z digest=sha256:fb1b997a313e72acb6d024906e14ba73e81b66cd51f33e631f3c2d5f05fc1401

Observation f5b589f5-b792-44b6-b1f1-18517a0997eb · outbound

This paper cites Images speak in images: A generalist painter for in-context visual learning.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Images speak in images: A generalist painter for in-context visual learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.566395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.078916Z digest=sha256:987ed9b81818f335b40e80d442e6e9132d3191baf4a647c01c7244a7352e3429

Observation b98495dc-fcde-4f1c-a264-76bc8d266395 · outbound

This paper cites Describe, explain, plan and select: interactive planning with large lan- guage models enables open-world multi-task agents.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Describe, explain, plan and select: interactive planning with large lan- guage models enables open-world multi-task agents

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.552038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.083234Z digest=sha256:4f20fd34eca64e83d56d2c4007f40c85dcd86617e6b44588c96ee24a57cfb7b9

Observation 3e9192bc-65fc-4252-83a8-72a138d6f435 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.087272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.087272Z digest=sha256:379f8971b6760548ed4ae4a02439c9626b260cdfe02cbe3950763b8d59fd73ef

Observation 220511c3-bbcb-4195-899c-5636fdd2fcb4 · outbound

This paper cites Conformity to cultural norms of tool use in chimpanzees.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Conformity to cultural norms of tool use in chimpanzees

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.527349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.092288Z digest=sha256:aa8116d685d84d76afda199635606c032def23b58656d0999f6e71ee5d8275c5

Observation c0096517-0ca5-41fd-bb3b-8138693be782 · outbound

This paper cites Accelerating Self-Play Learning in Go.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Accelerating Self-Play Learning in Go

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.096548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.096548Z digest=sha256:899871d642be8ff69117c47a0914e5725ce789d5eecb160214c3a522783782dc

Observation 6a2ff479-028d-4e5f-a8eb-5484f520be7f · outbound

This paper cites Unleashing large-scale video generative pre-training for visual robot manipulation, 2023.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Unleashing large-scale video generative pre-training for visual robot manipulation, 2023

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.512895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.101474Z digest=sha256:afe7c7797a844310dc0cddfef540cc8fcfbddb1ff4f00aeff131719e1943d6b4

Observation 9e133f33-bbdd-46b0-a2a9-42195704c809 · outbound

This paper cites Qwen2 Technical Report.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Qwen2 Technical Report

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.105651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.105651Z digest=sha256:de9be40c3d806ad2bd943de04931923d1c8daf7d7ee0b5dd2ded2895dec2cfc6

Observation cc8839f7-62bd-437a-92f1-1ae556bb0999 · outbound

This paper cites Video as the New Language for Real-World Decision Making.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Video as the New Language for Real-World Decision Making

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.110723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.110723Z digest=sha256:7820accf0b6c294369e3858ffb26552bcb51ab9e1ba9774ea29dda67e4cfed98

Observation c7bd887a-47e3-46f4-806b-273586744e63 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.116281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.116281Z digest=sha256:b9772e1f85075cd811fa52464c8cb6091070ded2dc7a12bf33b00dd9f8bb6ac5

Observation 8d85f7be-f30d-452d-afaf-7424a3f19a12 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Tree of thoughts: Deliberate problem solving with large language models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.498012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.120632Z digest=sha256:3918085cbcfc89e24c7ea6afb93b7c0b069ec4bdcf563cbe9cffdd693cc0123a

Observation db227515-27e7-4c8e-a034-1bb4e386e33c · outbound

This paper cites Latent Action Pretraining from Videos.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Latent Action Pretraining from Videos

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.125221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.125221Z digest=sha256:dd49ddfd061278a418fac8cc58655381121b8dad8c7303fbd6b8e8513c4ec954

Observation a74b4d63-c5cd-404e-adf6-89012786fe48 · outbound

This paper cites Magvit: Masked generative video transformer.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Magvit: Masked generative video transformer

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.484103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.129648Z digest=sha256:3912d6f542944b8215fd02bdae4c2e86eebaf46233bec4ec470a06a6b234d34e

Observation 2e456eaf-853f-455b-bb17-28176b15eeff · outbound

This paper cites Gundavarapu, Luca Ver- sari, Kihyuk Sohn, David Minnen, Yong Cheng, Vigh- nesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Gundavarapu, Luca Ver- sari, Kihyuk Sohn, David Minnen, Yong Cheng, Vigh- nesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:49:20.470995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T19:49:20.134457Z digest=sha256:d9b1368dc474b9a91c4c956dd75ff3a0ba89f181bf89ebeff63af97bf92ccf8d

Observation 4a54c6a8-65d1-4f45-b9b5-e850d0fd7f98 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Advancing LLM Reasoning Generalists with Preference Trees

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.138864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.138864Z digest=sha256:e57b1021ee81d593fb165d68fef25ca9d56f2514d1f646bb53c9618102682b45

Observation 27c27c54-7a09-4279-a0c9-c8000260a2f9 · outbound

This paper cites ControlVideo: Training-free Controllable Text-to-Video Generation.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos ControlVideo: Training-free Controllable Text-to-Video Generation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.143466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.143466Z digest=sha256:6300772a67707de5be6096ab0df3c21b1041a044bee84b2408e1c93319e05d4a

Pith citing papers

Observation e0e03382-e30f-4da4-93b7-ef5ff2975789 · inbound

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots cites this paper.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:09:10.250194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:a992284eb929a2d46d6ed3561816a90dd6aac5a0ea25f16f249c843529f70fc6

Observation 19bddcdc-3e9f-4d20-b56c-f482502d069f · inbound

Video-GPT via Next Clip Diffusion cites this paper.

Video-GPT via Next Clip Diffusion VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:30.073837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:38:30.073837Z digest=sha256:3e1765ba5219437bd170193a76129eee07faacb6c0c862fa3012421295d245f2

Observation e3700806-74e9-4a5f-9ac7-bcd038bd9f6a · inbound

DreamGen: Unlocking Generalization in Robot Learning through Video World Models cites this paper.

DreamGen: Unlocking Generalization in Robot Learning through Video World Models VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:50:45.623987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T23:50:45.332466Z digest=sha256:5c55cafcfaa104851b1bc50c52ac4c7b4e98ea56e2cac29364df662449ee57ce

Observation 782ca756-58ec-49ea-b0e6-d7f12cec7bf2 · inbound

Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency cites this paper.

Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:06.038673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:41:06.038673Z digest=sha256:f92d3649743968fd9fae14448c0d9eb475ab20a505caee089a8c720e4161984f

Observation d9c13907-1a94-44bd-9014-925d640868c8 · inbound

VGR: Visual Grounded Reasoning cites this paper.

VGR: Visual Grounded Reasoning VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:12:14.466283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T09:11:00.295700Z digest=sha256:fe15f26ddd8bba36e4631bfafde03074cca435e7b810cf3a696e517d923b23e6

Observation 5f6028fb-1b20-4a91-8432-462ded9e3bd2 · inbound

LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation cites this paper.

LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T22:04:35.640439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:04:35.640439Z digest=sha256:e0e1990219ce361c724e54aa9f73b05d2ff19345d750218a5b00ea437755ddcd