Pith. sign in

Paper Citation Record · LEDGER

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics

As of 7 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2506.06006.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06006 v3

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:07:24.277177Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8863294b-c256-4205-9463-1d4538c03e41 · outbound

This paper cites Recurrent world models facilitate policy evolution.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Recurrent world models facilitate policy evolution

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.201072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:23.988932Z digest=sha256:79037928f46f42ea99df89cad1a4779bdaf702b761d8e93189f346a9746d8eb3

Observation 4286e6d7-37ae-4216-b2cc-2fbec26dba99 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Cosmos World Foundation Model Platform for Physical AI

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:23.995257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:23.995257Z digest=sha256:3653ecd7e03a760708cb9cacc80806f8b8e142d92d792bb99c4561de4ed080f4

Observation 0ca95a54-c6cc-4406-a062-1c01f424057b · outbound

This paper cites Genie: Generative interactive environments.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Genie: Generative interactive environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.001012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.001012Z digest=sha256:b0d86ade7a3cf86820bcdab362bd38a4be78210ad797bbe936558459399479ac

Observation 7d0c545d-536c-4143-ba7b-cbf786b12461 · outbound

This paper cites Video generation models as world simulators.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Video generation models as world simulators

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.173532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.006263Z digest=sha256:16676aa5b2f990e0f00160dd2078937e3ec7c3babc7eead871bc454da56e7c47

Observation ea51fd94-ebf7-46bc-bed2-42afdc6a2250 · outbound

This paper cites WorldSimBench: Towards Video Generation Models as World Simulators.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics WorldSimBench: Towards Video Generation Models as World Simulators

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.010922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.010922Z digest=sha256:e23d843c98df70a0165f9e8b100ad7c3450f6264c4f4a0f5a85195e6250fd1d0

Observation 7c89477f-64fd-4e9c-acd9-ef939b47b253 · outbound

This paper cites Do as I can, not as I say: Grounding language in robotic affordances.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Do as I can, not as I say: Grounding language in robotic affordances

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.155966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.015924Z digest=sha256:410c3fee4564d1e329ac4b1c0b0f91b762f633e57aeeaf0ee1d9d29387771784

Observation 27e59029-a26d-4f79-bcae-5794c0d4232c · outbound

This paper cites Inner Monologue : Embodied reasoning through planning with language models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Inner Monologue : Embodied reasoning through planning with language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.137732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.020456Z digest=sha256:4cd3cec3e7b809c43d0d9632905bd96dc24723e962f37e4651c3e1f17dbaf2bc

Observation 3092e8bf-2fe7-43ca-b59b-c0bbd45e614e · outbound

This paper cites Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.025967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.025967Z digest=sha256:f62f7b4162e346774232c72dc6e597c96bbb8bba6a6e16e4f507cc4122a1222a

Observation 2a29bc56-079a-476f-adc6-e35ffee01026 · outbound

This paper cites A Generalist Agent.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics A Generalist Agent

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.031045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.031045Z digest=sha256:893f49dca2b1e77d05548815d07fe27d4d87a9e5eb389ccc332c2cc664579fb6

Observation 32537f05-b388-4691-98de-c7c20ef1cf32 · outbound

This paper cites Learning Interactive Real-World Simulators.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Learning Interactive Real-World Simulators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.035996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.035996Z digest=sha256:809ad2027f575cd7e15588a861075cf98e656bf993ede3e38e0d7c8694483315

Observation 78cea173-cdcd-4395-8461-ecbc2d080cdc · outbound

This paper cites Mastering diverse control tasks through world models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Mastering diverse control tasks through world models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.122718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.042286Z digest=sha256:dc8d10fc2fef1aef654220f90c688f9527878170234c90eb36064cc1a410e902

Observation d0ab5390-6057-4e67-ae4b-ba144b5846da · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.046755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.046755Z digest=sha256:6e12333d0f9b2aba36012798307c4adaf1d4f99cfaddecbdbd38fd88ba9eb305

Observation d3c9dccb-f84f-44ff-a2fa-27e1dc9191db · outbound

This paper cites Do generative video models understand physical principles?.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Do generative video models understand physical principles?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.051408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.051408Z digest=sha256:01906cd0007053fde685a33a6061f31eea237e028d343fe8051c70f357f7c05e

Observation 6cf3a3ba-b5db-450e-806f-021a7abf32c7 · outbound

This paper cites Physically grounded vision-language models for robotic manipulation.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Physically grounded vision-language models for robotic manipulation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.107730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.056924Z digest=sha256:90761eab5a9a5afa0e343a5d0fb272f2749d582e95dbf087ca4d2afe49b9fae9

Observation 69c4f1b3-c04f-4d8c-9e03-c6280b7f88ff · outbound

This paper cites an unresolved cited work.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:07:25.091676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.061561Z digest=sha256:003a043b744f01026f6b35f2a5cdf1345273b58ef9a58e1e48ee33a0503f5098

Observation 85becdf0-b936-4438-b842-bbc53d393128 · outbound

This paper cites Can language models encode perceptual structure without grounding? A case study in color.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Can language models encode perceptual structure without grounding? A case study in color

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.075935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.066153Z digest=sha256:eb66fb589e5eeb4411cf09c3ffb1535004e56e00442127f36b0234829ff075b8

Observation 3a0aab83-bf9d-49ef-9bf7-5f8aecfbd6df · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.072246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.072246Z digest=sha256:c0e117707372c0995085b1dc7e78739fcd0468391e5bc45cd1acf9b970782c0e

Observation 0f58401b-5e6a-42a0-aa18-07a5686137cb · outbound

This paper cites Learning Action and Reasoning-Centric Image Editing from Videos and Simulations.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Learning Action and Reasoning-Centric Image Editing from Videos and Simulations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.060218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.078717Z digest=sha256:81f7e7bf0a3bd7b5ed2e241476cd5deb208009a6d6ae19eb79e1358b3122cbb1

Observation 01bca877-3c16-4f23-92ba-7269b995fbf4 · outbound

This paper cites Video PreTraining (VPT) : Learning to act by watching unlabeled online videos.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Video PreTraining (VPT) : Learning to act by watching unlabeled online videos

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.043321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.084572Z digest=sha256:af917be3739be8a3be54c1627b018b337c8543db8b3a248f5bdf7a445b9aec9c

Observation 59a6acba-88a1-4f93-8150-cd732c3975b0 · outbound

This paper cites Moments in time dataset: one million videos for event understanding.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Moments in time dataset: one million videos for event understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.027199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.094006Z digest=sha256:4bf6d0a7066f9d851231b008e62a046541ad43f0d47d9eb194b7882780c01da4

Observation 6f02e7ac-5174-4f45-ab1a-a94afb616404 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics The Kinetics Human Action Video Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.106289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.106289Z digest=sha256:b65075fe2db7557eee7caa1143fea0f31ea5d0327e8687ae509ac92ad4c8d911

Observation e580f276-799e-4ede-9b90-7b96d4523b84 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics A Short Note on the Kinetics-700 Human Action Dataset

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.112945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.112945Z digest=sha256:70331dde8879e12152f24a6620cf207a2f9a64449cfa447bf1356d3e65a01474

Observation addd9681-e6dd-4c3d-9596-6b0341cff92a · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.117780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.117780Z digest=sha256:862ab91d10a8c1a08abbe0bdd7e5f2accb4c38275bf391ac9b5f18594b946305

Observation 845beb56-b69f-4cf9-b273-296b9a23170e · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.123574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.123574Z digest=sha256:c12a544768540832bbeb07f15e32e8503fe7b3c18414e20a79603ee92a20f502

Observation b7a16e98-9bca-4d02-9e3f-2ebaeb863cd8 · outbound

This paper cites Scaling egocentric vision: The EPIC-KITCHENS dataset.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Scaling egocentric vision: The EPIC-KITCHENS dataset

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:25.009280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.131233Z digest=sha256:cd53678f391a3729fd9947eeb94700c204ec89b2834b3a33599ad6c64723ee83

Observation 8b21755f-24ef-46e9-81fb-a454489583d8 · outbound

This paper cites s1: Simple test-time scaling.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics s1: Simple test-time scaling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.139015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.139015Z digest=sha256:14b34c78f52a0ad6185f61a17dbbce5f30f23ebd522a23e520c706a53d4f8a14

Observation cee2bfe3-7a20-43d3-9550-c4581fa7b974 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.146649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.146649Z digest=sha256:64d86b99ba456b8f539c36864959d348d173f75c74d26bfd359806e2c1aa2202

Observation abc8ad40-06ad-4a4f-9ead-752ccaf4dcea · outbound

This paper cites Kubric: A scalable dataset generator.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Kubric: A scalable dataset generator

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.993415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.154506Z digest=sha256:f84d69a7ef5c37fb2de5072b5cb66a3d3e5b0ce13c7f8bdc7426271e1f374f24

Observation d90e9194-b20e-49c3-aa98-da96adc23f23 · outbound

This paper cites InstructPix2Pix : Learning to follow image editing instructions.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics InstructPix2Pix : Learning to follow image editing instructions

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.978148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.160136Z digest=sha256:e9a4ebda53f61f79ef2b14b8a158a5d350874395bdadf61ee8d1925ba55abf75

Observation e85fb9d6-0103-4439-b9cc-e33646806b97 · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.165875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.165875Z digest=sha256:69891c19e242aa5030bcd759e62548b9eaab47e8811df6c974d0c3f510c9988a

Observation e0c0d67e-b944-490e-99fc-7cb50afc0217 · outbound

This paper cites SmartEdit : Exploring complex instruction-based image editing with multimodal large language models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics SmartEdit : Exploring complex instruction-based image editing with multimodal large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.961540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.171421Z digest=sha256:04fc8760db7e82209ef5ed966e52f923027532a69b1d3b55720bf24a3f93da49

Observation 9f064c14-cceb-4b80-a95a-21b137f4c6f1 · outbound

This paper cites BERTScore : Evaluating text generation with bert.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics BERTScore : Evaluating text generation with bert

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.943531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.177171Z digest=sha256:7d15ba1a717d7c23dd29b56f406e3554f7b442afe90a57b43e9f9742fe537178

Observation 3a054d53-0926-4a4b-8846-47f5df5cce80 · outbound

This paper cites ROUGE : A package for automatic evaluation of summaries.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics ROUGE : A package for automatic evaluation of summaries

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.926323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.181794Z digest=sha256:0d6b8b3f671b0591cb095ede515d04eb9f510a11f36d122ee57842aecadf8207

Observation 40bd7703-77e1-4a64-8af0-24e30a93a50c · outbound

This paper cites BLEU : a method for automatic evaluation of machine translation.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics BLEU : a method for automatic evaluation of machine translation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.910375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.186374Z digest=sha256:afc695ec86436d12235f603f5a57025343bcc699a3cd13413010f997db84fed1

Observation 22b337b0-2348-4812-b120-9203aac9fc09 · outbound

This paper cites Variational best-of-n alignment.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Variational best-of-n alignment

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.894651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.192078Z digest=sha256:049b104e243d58db0725b49c2dd50b540673c63713923a855a62c57b564554bc

Observation 04175720-5e5d-4a3d-bbb1-8affe0c9027f · outbound

This paper cites Learning to predict by the methods of temporal differences.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Learning to predict by the methods of temporal differences

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.879924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.196626Z digest=sha256:9c4686a8a9c2978b121c1ed4334bf8002d869358d1a1ad73a9574a5abf491256

Observation d3207217-fac6-436e-a918-954b2067100b · outbound

This paper cites Learning latent dynamics for planning from pixels.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Learning latent dynamics for planning from pixels

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.863771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.201635Z digest=sha256:9d5fbd00ed3c8dd9ee7fad28204901218d82b073b8c8c8597af8f1e9651c46f8

Observation e934c709-0463-4043-b79d-342cc8da7aa4 · outbound

This paper cites Transformers are sample-efficient world models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Transformers are sample-efficient world models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.848709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.206106Z digest=sha256:5ab22fbc80f3f28a847f3896aca8931c6603c7482a463012339f221e9b15706f

Observation f5d5c6b4-4972-4c15-90a1-8e701d569918 · outbound

This paper cites Transformer-based World Models Are Happy With 100k Interactions.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Transformer-based World Models Are Happy With 100k Interactions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.210816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.210816Z digest=sha256:62f1efd3a85bc1331edf61eb042126a2360bc47bb239087eda46009a8a527304

Observation 92f9b374-e22a-4e4f-8497-f805085aa920 · outbound

This paper cites Diffusion for world modeling: Visual details matter in atari.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Diffusion for world modeling: Visual details matter in atari

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.832798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.215773Z digest=sha256:05e4ab610727515a301d58cf4d18cfd975ca9cc4678899fee54cc14335e6fe66

Observation 1ec03a46-1e03-43c3-9ea4-a5a150fe8928 · outbound

This paper cites Video-LLaVA : Learning united visual representation by alignment before projection.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Video-LLaVA : Learning united visual representation by alignment before projection

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.816123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.220051Z digest=sha256:0c203c0c1ff1cc847c578ff50f2094b86edb59ede161f6f044f474c8f4c78686

Observation 0ee50795-2a17-4ce8-b0d3-d58b82240769 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.224463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.224463Z digest=sha256:747f4687b518182cb56b70c168b960d5bca4dbd1632d5884144a22957f927908

Observation 705237c3-ef8a-4da8-ab19-11f7323d799e · outbound

This paper cites iVideoGPT : Interactive VideoGPTs are scalable world models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics iVideoGPT : Interactive VideoGPTs are scalable world models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.799182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.229666Z digest=sha256:e2f1e8f689ffa69925c8d26a9bd350e7ef47520257f8b1b4b219f0cb6430a0b8

Observation 3d3f5554-3e1e-4183-b63f-9c8d3a053a5e · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Decision transformer: Reinforcement learning via sequence modeling

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.780990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.234370Z digest=sha256:dd7146fae18816ec8f74e1572d2c57135996d55d27a737c5a83fd345023c173b

Observation 0d753bfb-16d2-407e-bd88-cf44f27b7454 · outbound

This paper cites Vision-language models provide promptable representations for reinforcement learning.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Vision-language models provide promptable representations for reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.764728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.238717Z digest=sha256:48f98024c8af3ce96be56015c7a885039a34a657c5d952c1353781bce3e0f092

Observation d8e610c3-5eba-4198-89e6-b50ebdb6ae98 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.748169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.244183Z digest=sha256:6be7f77928fb242489d219232d8c8a42ccf8d3498406b515dd8aad7c2f0492e8

Observation bd6a97b2-880f-4f4a-81a5-1f4dba080336 · outbound

This paper cites Video as the New Language for Real-World Decision Making.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Video as the New Language for Real-World Decision Making

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.250329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.250329Z digest=sha256:1d320ed4a7529dcfcfdf3fea30cf83c95050702677b902985e0f62bba3dbfbd1

Observation 13d75b34-ae20-4441-806d-4682344e22b4 · outbound

This paper cites VideoAgent: Self-Improving Video Generation.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics VideoAgent: Self-Improving Video Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.256971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.256971Z digest=sha256:1d3f9484539d00ae714a9693a7214b271ec74ec470fec817f7b02a95d952d7f9

Observation e3caf459-b8b1-403e-a755-12a927bd2ff1 · outbound

This paper cites ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.262152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.262152Z digest=sha256:e2fd97ab84a8573f072e64e640d822011241f5a916cbbdb1868c3265537c23ee

Observation 32172cfe-5218-4da7-9b8b-0b6637d226e4 · outbound

This paper cites LoRA : Low-rank adaptation of large language models.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics LoRA : Low-rank adaptation of large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.731994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.267271Z digest=sha256:233a49b2133bc3dd4792a2b9b85a8852d149ec6267c4c3f7a7fdf74fe9d4f3b0

Observation 28e71d01-3eb4-46c8-9e6f-0f23089554fc · outbound

This paper cites Transformers: State-of-the-art natural language processing.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics Transformers: State-of-the-art natural language processing

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:07:24.715787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T06:07:24.272518Z digest=sha256:0fdef18592fee03528719f40a7c5c1f75a6578e847ed47b9db091587209b4a21

Observation 24fe666c-1b17-4bb7-8922-5bf0e22de8bf · outbound

This paper cites write newline.

Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics write newline

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:24.277177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:24.277177Z digest=sha256:67a2edaa4d3031ac1c5b29b5c9260cf5fe1123a0a9eceb4b18a4df01850f203f

Pith citing papers

No inbound Pith citation observations are available.