Pith. sign in

Paper Citation Record · LEDGER

Vision language models have difficulty recognizing virtual objects

As of 19 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2505.10453.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10453 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:12:12.530827Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 61e1e4e4-c887-4805-92bb-6492584a9f6c · outbound

This paper cites The imaginative mind.Human brain map- ping, 37(11):4197–4211, 2016.

Vision language models have difficulty recognizing virtual objects The imaginative mind.Human brain map- ping, 37(11):4197–4211, 2016

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:13.125204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.254145Z digest=sha256:e74cb06cc49ae6c243dd5d56c6afaf690ecb5f54c1d606a59718adfd171695ef

Observation f91c4ed3-7143-4b53-91ee-ca32f1fbabd8 · outbound

This paper cites Mapping the imaginative mind: Charting new paths forward.Current Directions in Psychological Science, 30(1):82–89, 2021.

Vision language models have difficulty recognizing virtual objects Mapping the imaginative mind: Charting new paths forward.Current Directions in Psychological Science, 30(1):82–89, 2021

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:13.112778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.258689Z digest=sha256:02c7ec7a5b4fb718eea46939d0dbe5afa7871fd4097ca05bec6743d88d363632

Observation 9a92cd4f-23b8-4d0e-9bb5-bf5ca575c1d8 · outbound

This paper cites Humans predict liquid dynamics using probabilistic simulation.

Vision language models have difficulty recognizing virtual objects Humans predict liquid dynamics using probabilistic simulation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:13.101947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.402073Z digest=sha256:42568df59075a1a602526f9bf9175a358e81c4a2fb76f50f37a250aaffa4d5d4

Observation 6f4f7625-2b7b-4975-b1bb-fcbc65845f1b · outbound

This paper cites Non- commitment in mental imagery.Cognition, 238:105498,.

Vision language models have difficulty recognizing virtual objects Non- commitment in mental imagery.Cognition, 238:105498,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:13.090673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.406714Z digest=sha256:3919029dbbc1ea5292ebb3fd9fe02e3a4fc8856afce8713e575efdd0afaddc5f

Observation 30cd9e77-9830-4a90-b16e-64e50895bf85 · outbound

This paper cites Infusing perception with imagination.Per- ceptual imagination and perceptual memory, pages 133– 160, 2018.

Vision language models have difficulty recognizing virtual objects Infusing perception with imagination.Per- ceptual imagination and perceptual memory, pages 133– 160, 2018

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:13.078240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.410727Z digest=sha256:6522136e3878d7c55ec8b8fb6580a46c6f9f905d76a78fd8a6b6008c1c36f6fa

Observation 9bd19b81-4e66-40ab-8a77-898be6cdd1ac · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

Vision language models have difficulty recognizing virtual objects SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:12.414626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:12.414626Z digest=sha256:d57de03e579157e152017104ba6e18c3c12d3cd8dc3b3082af6c0cdf9fa3d7ad

Observation d5cd8748-daf4-48bc-bf0c-0e753bab194b · outbound

This paper cites The artist as neuroscientist.Nature, 434 (7031):301–307, 2005.

Vision language models have difficulty recognizing virtual objects The artist as neuroscientist.Nature, 434 (7031):301–307, 2005

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:13.062649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.418121Z digest=sha256:399ab1c97faf597072d5160f32a133b79128b29d469d9aa88360acd029ac6b7e

Observation 7f57a4cb-e8d9-48c4-9c6a-c05b41646c12 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

Vision language models have difficulty recognizing virtual objects Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:12.421825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:12.421825Z digest=sha256:eb0795c741d12eea648eadb067068b404c98228dc77428eff2454cf2953e8e27

Observation b457c641-7738-4ea9-bcd7-2e409852a23c · outbound

This paper cites Large language models are visual reasoning coordinators.Ad- vances in Neural Information Processing Systems, 36, 2024.

Vision language models have difficulty recognizing virtual objects Large language models are visual reasoning coordinators.Ad- vances in Neural Information Processing Systems, 36, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:13.043677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.425291Z digest=sha256:a8dcfa810d501e9273f748c099011b12a6be68bb71f75727c71f674ee0abd422

Observation 2fb3764b-8f06-4347-b9f1-de1549c1d379 · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

Vision language models have difficulty recognizing virtual objects SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:12.428655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:12.428655Z digest=sha256:16258172c6c6e8e53dc32952593bedcbae82f1827d60c64debce094b3e659ff9

Observation 9ad80f3b-ef34-4abd-a3f0-b9d4cf5ec100 · outbound

This paper cites What makes mental modeling difficult? normative data for the multidimensional relational reasoning task.Frontiers in psychology, 12:668256, 2021.

Vision language models have difficulty recognizing virtual objects What makes mental modeling difficult? normative data for the multidimensional relational reasoning task.Frontiers in psychology, 12:668256, 2021

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:13.030909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.432415Z digest=sha256:9374137175c6f94b328cf45ced66b81c59dd98a0bba4a889364e135f5cca511f

Observation 9bfc39e8-cd9b-4a0e-83a1-9f0711354e6b · outbound

This paper cites A survey on multimodal large lan- guage models for autonomous driving.

Vision language models have difficulty recognizing virtual objects A survey on multimodal large lan- guage models for autonomous driving

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:12.436019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:12.436019Z digest=sha256:0bcd79b5156932b221cf6acdb62df6ebf92b7c9d286648b7adfd5162f83783a3

Observation 0a817b15-cc02-4583-9785-ba8b4851fe32 · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

Vision language models have difficulty recognizing virtual objects Objaverse: A universe of annotated 3d objects

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:12.439269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:12.439269Z digest=sha256:ecd6848f68cf6721b43157f6860d97eef6d847f18dd5da2b9f19727f343fc5ea

Observation e734e45a-98b1-4815-8ca0-3c556d5ec889 · outbound

This paper cites Psychology press, 2014.

Vision language models have difficulty recognizing virtual objects Psychology press, 2014

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.999904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.442642Z digest=sha256:1037c1c513545f79b36d1ac00174f2b9ebd41d4e73ab522db06ab7464b7e5041

Observation 164277a5-3b5f-4f02-a53e-2c9829f9401c · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Vision language models have difficulty recognizing virtual objects Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:12.445767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:12.445767Z digest=sha256:d7e3435be7aecc8feea9eb7c295bacb0732a3c57f5ee7ae7167f3ccac74b0cce

Observation 8711b5e9-2115-4fc6-8791-60ae69e25e4a · outbound

This paper cites Exploring the frontier of vision- language models: A survey of current methodologies and future directions.arXiv preprint arXiv:2404.07214, 2024.

Vision language models have difficulty recognizing virtual objects Exploring the frontier of vision- language models: A survey of current methodologies and future directions.arXiv preprint arXiv:2404.07214, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:12.449129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:12.449129Z digest=sha256:ecaa109fff37b5b88cbf030be39e1f2d513093d8b4cbfa22f8f8f66b960b2abb

Observation 4eb97da6-8ce5-497d-a1d8-6740840ccc4d · outbound

This paper cites Mental animation: Inferring motion from static displays of mechanical systems.Journal of experi- mental psychology: learning, memory, and cognition, 18(5): 1084, 1992.

Vision language models have difficulty recognizing virtual objects Mental animation: Inferring motion from static displays of mechanical systems.Journal of experi- mental psychology: learning, memory, and cognition, 18(5): 1084, 1992

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.986759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.452908Z digest=sha256:6725e460f88d799ec992231cc8ac5ed9db5ced8219ad9be11c14c4924e1ff6ad

Observation a5664f86-3b87-45e0-aafa-3b3a0e1a0ad8 · outbound

This paper cites Components of spatial intelligence.

Vision language models have difficulty recognizing virtual objects Components of spatial intelligence

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.975912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.456118Z digest=sha256:3a97330d4b6f4a5a08ea1c12f6ad76bda3e4bb56d070f4ecc5965ad5c51cb5c1

Observation d0e1af8d-6469-49a4-8168-ac44de2f4fdf · outbound

This paper cites Correctness comparison of chatgpt-4, gemini, claude-3, and copilot for spatial tasks.Transactions in GIS, 2024.

Vision language models have difficulty recognizing virtual objects Correctness comparison of chatgpt-4, gemini, claude-3, and copilot for spatial tasks.Transactions in GIS, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.965247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.459075Z digest=sha256:6ba8c8ed59b3ca2a7570cb173609eb40776f901208ade5b566186bc958711e3f

Observation aec52deb-391d-4a65-9802-a6b75c2ac3fa · outbound

This paper cites Vcoder: Ver- satile vision encoders for multimodal large language models.

Vision language models have difficulty recognizing virtual objects Vcoder: Ver- satile vision encoders for multimodal large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.953207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.462237Z digest=sha256:b8943d4b52f2a2d5df5fafae386792bd1769462abc06a030780c9b3936bd30d9

Observation 10bb8773-193f-4eb9-86dd-819740fed691 · outbound

This paper cites Imagery, visualization, and think- ing.Perception and cognition at century’s end, pages 441– 467, 1998.

Vision language models have difficulty recognizing virtual objects Imagery, visualization, and think- ing.Perception and cognition at century’s end, pages 441– 467, 1998

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.942096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.465394Z digest=sha256:99a947b24dc49475595ceea0f049bfb6a538e2c53cc86b8c261ee5db0b3cec94

Observation fde784fd-c61c-48d3-9335-170a4de56184 · outbound

This paper cites Harrison, Wallace E.

Vision language models have difficulty recognizing virtual objects Harrison, Wallace E

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.931118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.468859Z digest=sha256:8eadf43c5ff98f437cebebe07cc268ad375dae5508b45d376f6f3549e98d51ba

Observation ac046d2e-a6a4-4f84-a67b-ab23c3895bf1 · outbound

This paper cites Kinematic mental sim- ulations in abduction and deduction.proceedings of the na- tional academy of sciences, 110(42):16766–16771, 2013.

Vision language models have difficulty recognizing virtual objects Kinematic mental sim- ulations in abduction and deduction.proceedings of the na- tional academy of sciences, 110(42):16766–16771, 2013

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.920554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.472083Z digest=sha256:b4cabfac09b83a997a74935848cb3f0fd00a1d5f7411a6136f02141c08497f30

Observation 46fb6041-f54a-47e4-810c-97446058807d · outbound

This paper cites Mit Press, 2013.

Vision language models have difficulty recognizing virtual objects Mit Press, 2013

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.908492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.475261Z digest=sha256:33b072c5a8fd355935051299a7d2e3aab61fa00d52de272293bcf77db9b6f8ba

Observation 76e15394-aa56-4769-8216-912b6c7bd466 · outbound

This paper cites Visual imagery can impede reasoning.Memory & cognition, 30:363–371,.

Vision language models have difficulty recognizing virtual objects Visual imagery can impede reasoning.Memory & cognition, 30:363–371,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.897222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.478376Z digest=sha256:5cb75eed1ee42ecca3bb4e53652aaf9a8f1238ca0feef7dd49d7c407294e3852

Observation 21114088-c913-4f85-ad9b-7ae1baa0ece0 · outbound

This paper cites tracking.

Vision language models have difficulty recognizing virtual objects tracking

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.885986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.481449Z digest=sha256:8f36a80aab769a49d079036e39fc78e2a8c3a1a456c2b54c7985134e130383b3

Observation 141662ea-b9a3-4ea4-9dd4-2f1f2f6aff36 · outbound

This paper cites What matters when building vision-language models?.

Vision language models have difficulty recognizing virtual objects What matters when building vision-language models?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:12.484775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:12.484775Z digest=sha256:1300a4e7f2a6a1e1b3fb4cfcd5b7867fa89b4058faafca6d1a5d576333adbe4d

Observation da5c2a36-8920-42df-8271-ae264f5165ff · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Vision language models have difficulty recognizing virtual objects BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.875302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.488512Z digest=sha256:892954487377df4711b42a4ca9db6333a10eac32b8ebe491ef65d432df9d94fd

Observation b94335cc-8ad5-47a5-a88f-f0f2c730c53e · outbound

This paper cites A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends.

Vision language models have difficulty recognizing virtual objects A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:12.491634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:12.491634Z digest=sha256:2bebc67d603215d7ced39dd8aebc3ddf2b9a8a3a31183bee2e78fdb3b8ef73cd

Observation 807d42e8-5609-4e4e-8e72-8cb10b988d6e · outbound

This paper cites Visual spa- tial reasoning.Transactions of the Association for Computa- tional Linguistics, 11:635–651, 2023.

Vision language models have difficulty recognizing virtual objects Visual spa- tial reasoning.Transactions of the Association for Computa- tional Linguistics, 11:635–651, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.864609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.494765Z digest=sha256:948bb49767cbd73d7e3c63e5f1e3abd063f3b024e0911c79f32efb8f834a9565

Observation 82f5e51e-8551-45a3-874b-97ddd95a782e · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Vision language models have difficulty recognizing virtual objects A Survey on Hallucination in Large Vision-Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:12.497566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:12.497566Z digest=sha256:ca9920c26c288cd8f7ba29ed547ebe223d58b099afe6df4528c2ba08e90eeabe

Observation 94fa8e14-06a5-4714-bcf7-f397c6c1dd3c · outbound

This paper cites 5 Enhancing visual reasoning with autonomous imagination in multimodal large language models.arXiv preprint arXiv:2411.18142, 2024.

Vision language models have difficulty recognizing virtual objects 5 Enhancing visual reasoning with autonomous imagination in multimodal large language models.arXiv preprint arXiv:2411.18142, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:12.500995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:12.500995Z digest=sha256:1e310396d29c05f1e82e2c3f253d140677052f685199c598e695305e051fda5b

Observation 5c04e663-9ad8-49bf-8f0d-1d495d31001a · outbound

This paper cites The human imagination: the cognitive neu- roscience of visual mental imagery.Nature reviews neuro- science, 20(10):624–634, 2019.

Vision language models have difficulty recognizing virtual objects The human imagination: the cognitive neu- roscience of visual mental imagery.Nature reviews neuro- science, 20(10):624–634, 2019

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.853676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.503869Z digest=sha256:f39866012bff659dd5e125bae434be23969b055cb2e1531ef10ce85b28f84dff

Observation 310301e0-6938-4449-949b-61240fddea3d · outbound

This paper cites Mental rotation of three-dimensional objects.Science, 171(3972):701–703,.

Vision language models have difficulty recognizing virtual objects Mental rotation of three-dimensional objects.Science, 171(3972):701–703,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:12.506859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:12.506859Z digest=sha256:2dbb1de8d9063933f6e0343c98052d7ec659813a7c68e8ffe7d0c1c06a7d4106

Observation f95d7ee2-e416-4371-8d95-f61529d9a29c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Vision language models have difficulty recognizing virtual objects Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:12.509993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:12.509993Z digest=sha256:03c63518c560714d9a183e8c90a65c39dcbc3b284e5c63f74e8490884d7ab888

Observation 45b7ea2c-5398-4e21-a6d6-9dc468a9a7c5 · outbound

This paper cites Visuospatial reasoning.

Vision language models have difficulty recognizing virtual objects Visuospatial reasoning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.835650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.513731Z digest=sha256:9191c945ad64702deb1e66a259aa7957f93cdbf1cb17c0d015c0c75ac4704337

Observation e4a4ded8-de26-4fe7-9487-b3961f580a70 · outbound

This paper cites Learning physical parameters from dynamic scenes.Cognitive psychology, 104:57–82,.

Vision language models have difficulty recognizing virtual objects Learning physical parameters from dynamic scenes.Cognitive psychology, 104:57–82,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.825131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.517167Z digest=sha256:1af5504c733b0ae6bf02d9201948c30c125771af5982c5879c88e14943a861a4

Observation 75f919fd-5d57-4978-bd47-fd163b16edfa · outbound

This paper cites Multimodal large language models: A sur- vey.

Vision language models have difficulty recognizing virtual objects Multimodal large language models: A sur- vey

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.814212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.520724Z digest=sha256:bb72e42423c65675b4f55a0ae3f808b5e6d66c84f6e560235866068f3789ea3c

Observation 53f8aa8d-70b3-4747-b4f4-65376980ecc1 · outbound

This paper cites an unresolved cited work.

Vision language models have difficulty recognizing virtual objects Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:12:12.802280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.524013Z digest=sha256:ad8d0bddd6258e14c914a166e3047af2a24f369a849a13ef022ff3e384e95936

Observation 79ff480f-2082-47aa-86dc-3b8e118755d7 · outbound

This paper cites Vision-language models for vision tasks: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence,.

Vision language models have difficulty recognizing virtual objects Vision-language models for vision tasks: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:12.791134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:12:12.527114Z digest=sha256:70750f076759089a1b404121e695b94e1a3f08830d9787d17fe8fac4bf1df7d6

Observation 4531f783-7981-4220-9a1f-54e597539126 · outbound

This paper cites ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination.

Vision language models have difficulty recognizing virtual objects ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:12.530827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:12.530827Z digest=sha256:1c9eb0be6b217c81498f197d9ef6c70458629960f5c11a3f60a459e030cf4907

Pith citing papers

No inbound Pith citation observations are available.