Pith. sign in

Paper Citation Record · LEDGER

Do large language vision models understand 3D shapes?

As of 14 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2412.10908.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10908 v5

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:32:59.803624Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3034be33-4458-4393-be03-3d806614462b · outbound

This paper cites Creating such a system is one of the main goals of computer vision and can enable autonomous robots, cars, and numerous other applications[6].

Do large language vision models understand 3D shapes? Creating such a system is one of the main goals of computer vision and can enable autonomous robots, cars, and numerous other applications[6]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.672489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.617247Z digest=sha256:577da0bcf13a59220b5f76eab229d99f25be9afe0a2653011cf1b0847124b093

Observation 0a98dca0-2506-4489-b256-8723388d0a50 · outbound

This paper cites The use of CGI allows for massive amounts of object materials and environments as well as easy replacement of object materials.

Do large language vision models understand 3D shapes? The use of CGI allows for massive amounts of object materials and environments as well as easy replacement of object materials

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.655761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.623477Z digest=sha256:4ad9767f11c8838dfd4b1c6541c11369be06c9c616ca9abf5b204e4b4f914f83

Observation dc14ceb1-d4df-4d9f-8fd7-075eb122aa50 · outbound

This paper cites an unresolved cited work.

Do large language vision models understand 3D shapes? Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:33:00.637457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.629201Z digest=sha256:861d7bd9a37723fc4154ba434ce038b4732744fc4dbd79e071836a8c9173c02d

Observation a36ba2db-b468-4525-b700-c15386f9024f · outbound

This paper cites an unresolved cited work.

Do large language vision models understand 3D shapes? Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:33:00.617632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.635538Z digest=sha256:f0118cf2285595590b26840130061f0af5402714acb0e556b98acd7ea9a54580

Observation 501796c8-e85a-4d6b-9aea-6ed22d964a8b · outbound

This paper cites an unresolved cited work.

Do large language vision models understand 3D shapes? Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:33:00.598836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.640895Z digest=sha256:484e2846c9f4d7a2dd9d62da013c9dfe2676e3719abf8501dff1f9e30b153f92

Observation e686b0bb-aee8-43e1-9337-462d38ce1b52 · outbound

This paper cites Note that tests 2,4,6 force the model to rely only on the 3D shape for matching, while tests 1,3,5,7 allow the model to use the color/texture or the 2D projection for recognition.

Do large language vision models understand 3D shapes? Note that tests 2,4,6 force the model to rely only on the 3D shape for matching, while tests 1,3,5,7 allow the model to use the color/texture or the 2D projection for recognition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.579705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.646636Z digest=sha256:853a0c776e378e5ed270dd736a71fc56b76b53ac0f6d795daf4e3dff5bf27ddb

Observation abc6e0d9-52d2-4f76-9914-571fe80d2593 · outbound

This paper cites Which of the panels contains an object with an identical 3D shape to the object in panel A. Your answer must come as a single letter.

Do large language vision models understand 3D shapes? Which of the panels contains an object with an identical 3D shape to the object in panel A. Your answer must come as a single letter

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.562114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.653051Z digest=sha256:19a293d001bdfa71183287c97396654b03e379d79aaea60f015d0e3460e3938a

Observation 6368588a-5d7d-42a8-9a86-55ea9f8c417d · outbound

This paper cites 2) Has a different orientation compared to the object in panel A.

Do large language vision models understand 3D shapes? 2) Has a different orientation compared to the object in panel A

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.544037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.658106Z digest=sha256:01ae7e7d9a54ece6363817f0115012e7e655c663d8a5af7b918977dce23b496c

Observation 680de0eb-c39b-4424-bf54-467d886c0059 · outbound

This paper cites Gemini and GPT 4o clearly excel in this task but all models show some level of understanding (Table 1).

Do large language vision models understand 3D shapes? Gemini and GPT 4o clearly excel in this task but all models show some level of understanding (Table 1)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.525021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.663124Z digest=sha256:e40c9cfdf06af288ad72d67de4efc1fae7141b848209392079b6cecd5b893346

Observation 8a5c8630-6636-4aac-a126-99e5bc22f2eb · outbound

This paper cites These results are consistent with previous works which show that despite their impressive performance Vision Language Models (VLM) often miss basic aspects of reality[8-12].

Do large language vision models understand 3D shapes? These results are consistent with previous works which show that despite their impressive performance Vision Language Models (VLM) often miss basic aspects of reality[8-12]

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.503699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.668844Z digest=sha256:d32768363e71c28b604e6c8a34a2582cbe73e45c47a9ec5eccdbda65f2c8e7e6

Observation 2e6bfac0-3978-4ddf-a952-d83db8758918 · outbound

This paper cites Code used to evaluate the models on the benchmark available at this URL.

Do large language vision models understand 3D shapes? Code used to evaluate the models on the benchmark available at this URL

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.477553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.674052Z digest=sha256:e8fb07df24587377d22875ad218b9fd44558d474e70f29b69820b10a7153a035

Observation 7c8821e9-3f3f-4f04-8fd5-2cb6bb35516d · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Do large language vision models understand 3D shapes? Flamingo: a visual language model for few-shot learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.458914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.679122Z digest=sha256:5ddfb0953048bc9624e365a6c83fb7795c9fe164bdea18edccaa3d019410424b

Observation 50ea74b2-1b65-4e03-8042-7766b9cbe2cd · outbound

This paper cites GPT-4 Technical Report.

Do large language vision models understand 3D shapes? GPT-4 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.684975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.684975Z digest=sha256:20712b5dc53c12748afc34759c9b011a05fda2aca1e7ff76742dc44f265af7e9

Observation a9899b30-3309-445f-8474-9cff7f7ab3f7 · outbound

This paper cites Real-world robot applications of foundation models: A review.

Do large language vision models understand 3D shapes? Real-world robot applications of foundation models: A review

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.441207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.692018Z digest=sha256:9277e276b8a73520069e1dedaa75b38d4bd3da3cc17eddbd6f5f4ce0e4f5c072

Observation 0ab174d4-2dda-4e41-8d9f-03f3439000f4 · outbound

This paper cites Approaching human 3D shape perception with neurally mappable models.

Do large language vision models understand 3D shapes? Approaching human 3D shape perception with neurally mappable models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.700235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.700235Z digest=sha256:cafcb8b6e0844ea2024c827cc62b2a43c95bc9a7b0881d6d4f3600fcffff11bc

Observation c95c0070-cf76-4ea2-a227-819545c5a26f · outbound

This paper cites Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models.

Do large language vision models understand 3D shapes? Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.421219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.706021Z digest=sha256:e477e843188671a8614695b259ced29b479d259cd584c2f647422386862dc522

Observation b51dd177-3386-4663-a34b-04f75132c30f · outbound

This paper cites VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information.

Do large language vision models understand 3D shapes? VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.711784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.711784Z digest=sha256:f7a752e2dd2ecf98a0eb1f9ee2e405ba4f0e19f9366fbc538ded9308445fb6a8

Observation 19fba42e-ae53-433a-832d-7bb50df268e8 · outbound

This paper cites Can We Talk Models Into Seeing the World Differently?.

Do large language vision models understand 3D shapes? Can We Talk Models Into Seeing the World Differently?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.717893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.717893Z digest=sha256:2c136ffee268f57f124fce75b31695982c4627d407d945f31172646c1b4a9ecd

Observation 9b2d850c-06bd-4181-be6d-f1398fa68724 · outbound

This paper cites Exploring the frontier of vision-language models: A survey of current methodologies and future directions.

Do large language vision models understand 3D shapes? Exploring the frontier of vision-language models: A survey of current methodologies and future directions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.723920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.723920Z digest=sha256:c897debd197179b2e979284ccf77c167945dee930ec1b04ef8b1b448e51ac379

Observation b98382ba-30bc-42b8-915d-d86fbd198ffc · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Do large language vision models understand 3D shapes? BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.729389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.729389Z digest=sha256:2ea61690b4c3639408fcf877d35f3c9dbd0ff6af0326819b8d2600f93b4d05ef

Observation 89992980-f643-4bcd-9190-ff61916e22ca · outbound

This paper cites Can 3D Vision-Language Models Truly Understand Natural Language?.

Do large language vision models understand 3D shapes? Can 3D Vision-Language Models Truly Understand Natural Language?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.735368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.735368Z digest=sha256:0b8368bf870cfb19e6fcdf03439a16a849c61638a195308420a859dbabfa1659

Observation 7767142a-5c19-4f6a-a0aa-7a0d04da4b3c · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Do large language vision models understand 3D shapes? Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.741436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.741436Z digest=sha256:ea41660a339ca6a73dc5de280ba9491e9c7a6f4dc697585ed59af06bfd62981a

Observation 7ed68ce4-b4cd-4062-aa72-90accc550ae2 · outbound

This paper cites GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?.

Do large language vision models understand 3D shapes? GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.747061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.747061Z digest=sha256:7663b89b5c0d8161a7218a2d6a41cbb57201942bd65caeaf0783f71962df1ef1

Observation ec3a7477-e733-413f-8262-f7e521146d38 · outbound

This paper cites A general protocol to probe large vision models for 3d physical understanding.

Do large language vision models understand 3D shapes? A general protocol to probe large vision models for 3d physical understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.384059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.752641Z digest=sha256:037bff4b8a88e85d656f4888782600e4050f8448d588e26917af7cdecb74a1ab

Observation f96fe009-55fb-4756-a18d-1f54a1fb6376 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Do large language vision models understand 3D shapes? A Survey on Hallucination in Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.757593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.757593Z digest=sha256:2c8d54bc99dd1aa51b706ce6ed0bf58308ce6f43d70294b84df3d275f2944927

Observation 4765c8e0-0de6-46d9-9039-00709859032f · outbound

This paper cites When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models.

Do large language vision models understand 3D shapes? When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.764351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.764351Z digest=sha256:850bd55b7953f259a70f7d3f2531f6d5754eedcffe8b09455624928cd9fa936d

Observation add2d72e-3b04-461f-bfcc-1d9e03910ff2 · outbound

This paper cites Language-Image Models with 3D Understanding.

Do large language vision models understand 3D shapes? Language-Image Models with 3D Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.771499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.771499Z digest=sha256:8b94fcd5a1d9565e9827758fad3380e0f1be8c1bd509c7f37e85f5775614ff7a

Observation 56fee059-30fe-4cdb-86e0-03dacf78109a · outbound

This paper cites Shapellm: Universal 3d object understanding for embodied interaction.

Do large language vision models understand 3D shapes? Shapellm: Universal 3d object understanding for embodied interaction

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.362487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.778276Z digest=sha256:28f51f1ecf00a5421ced7f2431a313789da548386742390591fef0cf6b13ae29

Observation e5721b71-513f-4a10-bdab-8aa799b477dc · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

Do large language vision models understand 3D shapes? Objaverse: A universe of annotated 3d objects

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.345067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.786119Z digest=sha256:5314afd4fcfeffd2a7c5e91a6fa550024ca8898943e8b24b58969395555d2d00

Observation 2af7571c-6ed6-4cee-82fc-5a8b44b5bc94 · outbound

This paper cites Vastextures: Vast repository of textures and PBR materials extracted from real-world images using unsupervised methods.

Do large language vision models understand 3D shapes? Vastextures: Vast repository of textures and PBR materials extracted from real-world images using unsupervised methods

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.792476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.792476Z digest=sha256:1f32076f29c6be788aa76ee69877c6b142e0ab7fc4161f8cf56d13e7f1347569

Observation bf7155e3-a1fc-48f5-ab46-ed90f98c84ec · outbound

This paper cites Infusing Synthetic Data with Real-World Patterns for Zero-Shot Material State Segmentation.

Do large language vision models understand 3D shapes? Infusing Synthetic Data with Real-World Patterns for Zero-Shot Material State Segmentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.325661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.798061Z digest=sha256:30fe9091e4b21fa8d4deb5f10f13d18dad803fb6d2ada478ffa9b19cc6250834

Observation cc5021e5-b612-4a76-af9f-6b71d0dfdd6a · outbound

This paper cites Which panel contains an object that has identical 3D shape to the object in panel A, but different in orientation and texture. Explain.

Do large language vision models understand 3D shapes? Which panel contains an object that has identical 3D shape to the object in panel A, but different in orientation and texture. Explain

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.307386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T15:32:59.803624Z digest=sha256:fc12d62a65ba49bb7b80d843ab95ca894377792e5750a222a02fc5bf6d6e1115

Pith citing papers

No inbound Pith citation observations are available.