Pith. sign in

Paper Citation Record · LEDGER

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

As of 17 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 4 inbound Pith citation observations for arXiv:2506.21876.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21876 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:22:39.392445Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:04:43.145985Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 96d09218-983c-4fd2-8ab9-aa20774187c0 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:44.639025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.684245Z digest=sha256:549be5552e2328cde4a0a4f739246f7a96647c77b4dac1563580b4f337f94737

Observation 165c8bd4-30ce-460b-b4e2-37a3124c2fa5 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:44.365198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.707507Z digest=sha256:340297007d64f77f60726020c6e91ade6f862accfdfcb04688374dffa2fb77c6

Observation 3b6f71db-95d1-49e5-9e86-1c708b75a125 · outbound

This paper cites The topological order of the graph is determined by time and world dynamics.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation The topological order of the graph is determined by time and world dynamics

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:44.089308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.734779Z digest=sha256:329ff08e563d92daf59a06e8d283b2e010efd5faba87d7687d893ba8f74265c5

Observation afe4b3b5-5f30-46f6-b85c-2d8ea1bffb13 · outbound

This paper cites These multi-view tasks emphasize the model’s ability to synthesize distinct viewpoints into a coherent three-dimensional representation of object arrangements.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation These multi-view tasks emphasize the model’s ability to synthesize distinct viewpoints into a coherent three-dimensional representation of object arrangements

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:42.690342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.917251Z digest=sha256:33bfbce2b9082083f0d0b8e9212ade49880455244f8a2cca76a5ed16ca375198

Observation 93f01c19-7bb6-40be-94fd-880cabfa4775 · outbound

This paper cites SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.200429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.200429Z digest=sha256:aa1d4b5731678dd015f582e1b4079a87086498169cbb4ba92f0723e3d8b1a12d

Observation 515bd20e-6045-42d8-9c10-e33ef8999192 · outbound

This paper cites Why think step by step? Reasoning emerges from the locality of experience.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Why think step by step? Reasoning emerges from the locality of experience

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.379328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.379328Z digest=sha256:b4cc847dde12f56b4051fcc4f2e1b0bba6ad561a636fbfd06ed64f54984bb83f

Observation e34fa55c-165e-4d18-b51e-fd7d9700d18d · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Vision language models are blind: Failing to translate detailed visual features into words

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.423858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.423858Z digest=sha256:d6e3cc35d73f3f0563a49dc2a89026aeffd2ddfcbffc78ddd4edac9e763dc55a

Observation aeed72a1-4d00-48c2-8431-55b489ce3dc5 · outbound

This paper cites IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.457730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.457730Z digest=sha256:fd96f0aac61df56d4c0e2c05b170568e222a20d72347f342dfad379491ba7c38

Observation 8098a857-5480-4c86-831a-40b6fb1c9afb · outbound

This paper cites Diffusion Dynamics Models with Generative State Estimation for Cloth Manipulation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Diffusion Dynamics Models with Generative State Estimation for Cloth Manipulation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.511154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.511154Z digest=sha256:005494c367ad1e63bbb23e34795152d74bc6d32bcc73ac3a872726329bf954b7

Observation b42919ff-a75d-4e1c-bf9c-1db86f018aab · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.580685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.580685Z digest=sha256:2cd5aaf64b601de2e9783954a5f2f8aabc98be73f8456cb4dca7b0abb9447c3b

Observation 3ccfcbb2-0c3b-4b86-ac09-df84f94e9a8b · outbound

This paper cites VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.649660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.649660Z digest=sha256:528544fd9a809165a77462c3046c10d6cce5e6bef22e8c3cd9985957e9031f28

Observation d2eb0d2c-2e7e-46ed-b259-451264b35f48 · outbound

This paper cites This task evaluates whether the model can accurately discern spatial relationships based on visual cues.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This task evaluates whether the model can accurately discern spatial relationships based on visual cues

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.479831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.794107Z digest=sha256:417b84fddf3dd19725a73f5d1bfde8b3fcf94ad0ffe2aa543f14b1aeb23dbdc3

Observation f3fab12a-a7f6-4d52-bcfc-09507104ee3b · outbound

This paper cites Thus testing the model’s understanding of spatial constraints.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Thus testing the model’s understanding of spatial constraints

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.162217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.824097Z digest=sha256:06bcb5c580017688650bd706365a308d7bee26ffb4e37e8c50a05a15ecd6c8ee

Observation e601943c-4656-44f8-99b2-02a51961ab65 · outbound

This paper cites Thus testing the model’s understanding of object size.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Thus testing the model’s understanding of object size

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:42.943046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.864031Z digest=sha256:8abae249e2ac5d86cf024451f94e9f58a093d1182efe1f3ec787a88af18e2ca8

Observation 155d7142-72dd-44bc-8a7c-1ce312aabd2d · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:42.551464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.953529Z digest=sha256:9abf1082e5660002afaba1fcd65f7521149ee37914999e3d2f3aee66b3870e3d

Observation 47f848a7-016a-4c1b-9e7f-817bd4eb82cd · outbound

This paper cites Collectively, these tasks investigate the aptitude of a model to maintain consistent temporal representations, estimate durations, and infer the correct order of events.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Collectively, these tasks investigate the aptitude of a model to maintain consistent temporal representations, estimate durations, and infer the correct order of events

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:42.369495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.991794Z digest=sha256:191850ef1a723f9467a531d6185fbbcf8b03acab29a7b8b99e3c38f095800e3b

Observation 2a7c8717-0104-4d8b-b9ce-b329c2b846e8 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:42.257805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:38.024311Z digest=sha256:f6c939651c22a829df1fb9b49ddf5afae66dbb726835477e34c96996fd450648

Observation e521baee-d6a4-4daa-b710-1e437ba9f704 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:42.135747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:38.094325Z digest=sha256:e5c8374433c36023ac10656d83377828687d53fcf1178acf20a4598c28ce68b7

Observation 2d1df94b-2d08-4c4c-bc53-87f17ad2d684 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:42.036339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:38.154224Z digest=sha256:2553e58f67fdf34853b0a5ef97adcd9aca165c94ac95e1d3178a1b0a6c5eaf3a

Observation 42c72287-8263-46a2-b8fb-3256ec1db670 · outbound

This paper cites This setup tests the model’s capacity to identify the moving action of objects.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup tests the model’s capacity to identify the moving action of objects

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.857927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:38.224945Z digest=sha256:3eb38d61c4b7394ac0aa0324a0b37b7f4c1cf84f1b0c4f1fe2955c9808c91267

Observation e8c37cfd-8fe3-4aa4-a7dc-816b6751e3c5 · outbound

This paper cites This setup tests the model’s capacity to track position changes over time and estimate relative velocity.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup tests the model’s capacity to track position changes over time and estimate relative velocity

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.741165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:38.287916Z digest=sha256:7666d1f170590869bc540fa2d0d0255f4ac302b7ad140cc07e7f47aeca66f794

Observation 958617c8-47ad-4c21-ae53-9e26396b54d7 · outbound

This paper cites This setup tests the model’s 20 capacity to track position changes over time and estimate relative moving direction.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup tests the model’s 20 capacity to track position changes over time and estimate relative moving direction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.615188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:38.367788Z digest=sha256:ebd185678e71cd1450b851a414a825ad20f68d4c5e5c2ca23fab16b9aeae259a

Observation 6c57eddc-4f41-46e4-b015-a3257e87c33d · outbound

This paper cites Quantitative Perception.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Quantitative Perception

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.495654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:38.459906Z digest=sha256:d6ea1638df264f7efa02bbbbdda5240dac930ffcaebe57bd07c93740625d396b

Observation 8b0665ea-32c4-4eb6-b3a6-141e4b19dccb · outbound

This paper cites This setup evaluates the model capacity for discrete numerical estimation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity for discrete numerical estimation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.375521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:38.534512Z digest=sha256:5605c39ccbaceba8d9810ae9a708f747851dc832779cc33a44a69112326ed589

Observation 74281348-1b85-4e0e-aee5-7c50ec805ee0 · outbound

This paper cites an unresolved cited work.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:22:41.268379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:38.610245Z digest=sha256:5549c98f85345afda978e8785963e21736874f7ca708946c5647e7acf4b0cdf9

Observation e73263ab-5223-4843-a247-4e8fe27ccb9e · outbound

This paper cites This setup probes counting skills, numerical reasoning, and perceptual comparisons in a visual context.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup probes counting skills, numerical reasoning, and perceptual comparisons in a visual context

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.156533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:38.684429Z digest=sha256:4e52e86a54bbc47ee2c01e23db97d05933e25832f41b6b7e91f785650d1cbe8a

Observation 5f530ea3-642d-45e2-a267-7e1b50930ba5 · outbound

This paper cites This setup evaluates the model capacity in physical reasoning.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in physical reasoning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:41.056090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:38.774412Z digest=sha256:0be0c1c72a2e27245362e2f29c172b112e8448681d58d72594a75c307f0546ca

Observation 4ffb2d8c-db14-4262-8d1a-663616202112 · outbound

This paper cites This setup evaluates the model capacity in predictive reasoning.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in predictive reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:40.842243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:38.869516Z digest=sha256:794d7054d8d1a4c430d2a52deeb217cc307242a2e1b49d64ba0b67bdebad00c8

Observation f77b19c7-0baa-4c21-93b7-78fb3eb07adc · outbound

This paper cites This setup evaluates the model capacity in predictive reasoning for robot manipulation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in predictive reasoning for robot manipulation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:40.672138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:38.977727Z digest=sha256:b84aad2410894d1beabead395572a142fd83de948d34fc46129f935a60440e10

Observation 241f8ca8-00e2-4180-a1c0-d8ec06f2aceb · outbound

This paper cites This setup evaluates the model capacity in predictive reasoning for autonomous navigation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in predictive reasoning for autonomous navigation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:40.525701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:39.048655Z digest=sha256:0dde89047479bea2d4a005dfc0069a996eca18c4ca32c3973c81657af8d8026f

Observation 0ee026b0-4639-4502-97dd-395333f84139 · outbound

This paper cites This setup evaluates the model capacity in multi-step predictive reasoning for robotic manipulation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in multi-step predictive reasoning for robotic manipulation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:40.350422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:39.122549Z digest=sha256:ba2ed54622d9e906d7ad4124354c56f0cb414de4790a8bc63fd9b0eb775de244

Observation e656d49b-4ad5-447f-b6f2-a0f0e3964994 · outbound

This paper cites This task evaluates the model’s ability to perform compositional inferences about physical causality and object behavior.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This task evaluates the model’s ability to perform compositional inferences about physical causality and object behavior

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:40.162338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:39.211208Z digest=sha256:d942e5d1b24ff84f8c5723597aaaaefbbaae64823709c148b658e3634c4d16dc

Observation 6de38852-8989-4b01-8495-8a41ab86f89d · outbound

This paper cites This setup evaluates the model capacity in concurrent action predictive reasoning for robotic manipulation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation This setup evaluates the model capacity in concurrent action predictive reasoning for robotic manipulation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:39.987746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:39.298648Z digest=sha256:f4464699b7f5b89d80d9c5af166e964ed16a7a3c84059c884ad0dd1354ac5490

Observation dc74911f-a30f-40e4-8d70-b6e3e09c7389 · outbound

This paper cites yes” or “no.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation yes” or “no

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:39.809002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:39.392445Z digest=sha256:affa1765444185eee439960864ee4c9e4213494abb5ed1cf17534d27fa3812cb

Observation af3719b3-b7c1-466c-8f02-660250ebc0c0 · outbound

This paper cites John Wiley & Sons Hoboken, NJ.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation John Wiley & Sons Hoboken, NJ

Reference 2004

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:03.324343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:36.849611Z digest=sha256:642b5d2b843d96db628e375e3f4ec3d3e9af2b8f69c9e4f1f85b0d1d70b4c3b1

Observation cab77872-0d73-4b5b-b584-3375f1655071 · outbound

This paper cites S(n) t to denote the whole relationship between n component states and complex states.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation S(n) t to denote the whole relationship between n component states and complex states

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:43.816986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.758890Z digest=sha256:6fe7194f37fbf6d03a0f552bf09a68b61bb245b147a7be33f70a7c28fb0f5b33

Observation 399c084c-8c05-4a5e-b08e-811189840f70 · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation GAIA-1: A Generative World Model for Autonomous Driving

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.281149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.281149Z digest=sha256:2633e0439d64186bdecae8ba83e89ad9d497b490e9aebc51d54041066c21a1c2

Observation ee900967-a049-4644-9de2-b0cda7747bd3 · outbound

This paper cites Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.339688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.339688Z digest=sha256:20333fca353170727c0dc34115dc2328628ab70fad74475305ab985a5923efac

Observation 2f2ff74f-0605-4c25-919e-08b5cca1b525 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:36.559873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:36.559873Z digest=sha256:5734d1d74651c23a1b16a35c908b9833bfa3e0affe601359bcb21da14f87ec2a

Observation 7f3f7446-c88c-416c-a422-478b4e23c352 · outbound

This paper cites SpatialSense: An Adversarially Crowdsourced Benchmark for Spatial Relation Recognition.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation SpatialSense: An Adversarially Crowdsourced Benchmark for Spatial Relation Recognition

Reference 2019

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:22:39.582059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.555217Z digest=sha256:a64f978f625155bd05cb8d3febe3a62495c9424eaf35d5172318ff9adace8046

Observation 890028db-ee63-46c2-ab8e-277326e274e3 · outbound

This paper cites In Advances in Neural Information Processing Systems, volume 33, pages 10514–10525.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation In Advances in Neural Information Processing Systems, volume 33, pages 10514–10525

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:02.907075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.101582Z digest=sha256:dce5b1d74f9077430ee481203293521e06788efa8d05c4cbca89d569b624cf75

Observation 6c279471-d033-4ad8-8d46-16a9da93dcc7 · outbound

This paper cites CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:36.996473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:36.996473Z digest=sha256:826755c22c3cd1df41f223bdb19811eb04809a6f1c33c4a4f65e624f4b0bc5b7

Observation bc149e21-7257-48a7-ad3a-38db00cca78b · outbound

This paper cites Faith and Fate: Limits of Transformers on Compositionality.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation Faith and Fate: Limits of Transformers on Compositionality

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:36.904449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:36.904449Z digest=sha256:0f43f32882cfe20d8d80965074b950efba6c927e9833b88c4d1405012a77c5d0

Observation 802dd24f-d6f8-43b7-b55e-d9e7c290a502 · outbound

This paper cites How Far is Video Generation from World Model: A Physical Law Perspective.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation How Far is Video Generation from World Model: A Physical Law Perspective

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.307022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.307022Z digest=sha256:d4739773cd3dce02e845e19592ec078b1350fae230309adfe716326e4cff0628

Observation 3392a115-99fe-46c2-9cc2-5a7dc56cf3e8 · outbound

This paper cites In The Thirteenth International Conference on Learning Representations.

Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation In The Thirteenth International Conference on Learning Representations

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:22:44.942959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:22:37.626475Z digest=sha256:56ddd7007b561904c2249ad4c66c0b40e56c6c5f6a62e0f7b71080625101fba7

Pith citing papers

Observation 770e4f03-85ca-428b-9829-82fbbaaf69ce · inbound

Egocentric Bias in Vision-Language Models cites this paper.

Egocentric Bias in Vision-Language Models Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.145985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.145985Z digest=sha256:d46b1cc9db35a060ea6c5946c75beca7ec1cd5c3b1ea6ccee30c4e019307bb53

Observation 17d4b390-cd08-4e15-a2f8-2390ef800198 · inbound

DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics cites this paper.

DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T21:36:42.378149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T21:36:42.378149Z digest=sha256:1685a213070dee003a055a85646df5068e3931536d60b1c24d1333b0b93f7586

Observation f6bf9e4e-5780-48d2-9252-45f48f43f995 · inbound

RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning cites this paper.

RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T07:27:08.281893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:27:08.281893Z digest=sha256:17146972ea296c9aa3b3bf79327982664362ae1ae8862a71db80a6734072a4cf

Observation d023fc94-0bd3-4f98-b89b-fe4b2cc8c1be · inbound

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment cites this paper.

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T05:16:36.966727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:16:36.966727Z digest=sha256:e2fd6c577cbd4959562e93e43ff24ddcfc4d6b416038eb449fdaa9cb9d48586e