Pith. sign in

Paper Citation Record · LEDGER

Do large language vision models understand 3D shapes?

As of 14 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2412.10908.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10908 v5

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:32:59.803624Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3034be33-4458-4393-be03-3d806614462b · outbound

This paper cites Creating such a system is one of the main goals of computer vision and can enable autonomous robots, cars, and numerous other applications[6].

Do large language vision models understand 3D shapes? Creating such a system is one of the main goals of computer vision and can enable autonomous robots, cars, and numerous other applications[6]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.672489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.617247Z digest=sha256:2b569d47cb6dae99213020581284c3af8df8e954e8d76b93744438bcc2395176

Observation 0a98dca0-2506-4489-b256-8723388d0a50 · outbound

This paper cites The use of CGI allows for massive amounts of object materials and environments as well as easy replacement of object materials.

Do large language vision models understand 3D shapes? The use of CGI allows for massive amounts of object materials and environments as well as easy replacement of object materials

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.655761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.623477Z digest=sha256:c0e90e0f495946fd9ea58453d129aa689cda172bcde753afc4d55afc143216f0

Observation dc14ceb1-d4df-4d9f-8fd7-075eb122aa50 · outbound

This paper cites an unresolved cited work.

Do large language vision models understand 3D shapes? Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:33:00.637457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.629201Z digest=sha256:3fc35b920097c228a6667d9af96dad1435985425c84fb74e205e73186d2954ba

Observation a36ba2db-b468-4525-b700-c15386f9024f · outbound

This paper cites an unresolved cited work.

Do large language vision models understand 3D shapes? Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:33:00.617632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.635538Z digest=sha256:b891a14b93d51ca32d2fe83ea719a1fd71f3a9ad5da6b1cba7a8f223b0caf620

Observation 501796c8-e85a-4d6b-9aea-6ed22d964a8b · outbound

This paper cites an unresolved cited work.

Do large language vision models understand 3D shapes? Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:33:00.598836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.640895Z digest=sha256:6721f106493329a7a426e4b462bc6fb176bcd054a1d5bc98fd5e02fd6e9c91bc

Observation e686b0bb-aee8-43e1-9337-462d38ce1b52 · outbound

This paper cites Note that tests 2,4,6 force the model to rely only on the 3D shape for matching, while tests 1,3,5,7 allow the model to use the color/texture or the 2D projection for recognition.

Do large language vision models understand 3D shapes? Note that tests 2,4,6 force the model to rely only on the 3D shape for matching, while tests 1,3,5,7 allow the model to use the color/texture or the 2D projection for recognition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.579705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.646636Z digest=sha256:e8e6bd8129142e0e48658d4b49696683ad3eb60270410c561d5b8b1c9e2d3180

Observation abc6e0d9-52d2-4f76-9914-571fe80d2593 · outbound

This paper cites Which of the panels contains an object with an identical 3D shape to the object in panel A. Your answer must come as a single letter.

Do large language vision models understand 3D shapes? Which of the panels contains an object with an identical 3D shape to the object in panel A. Your answer must come as a single letter

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.562114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.653051Z digest=sha256:e5c511910a796ac878b1e4366141d7b68332b895fb33a03346729b0f895e1754

Observation 6368588a-5d7d-42a8-9a86-55ea9f8c417d · outbound

This paper cites 2) Has a different orientation compared to the object in panel A.

Do large language vision models understand 3D shapes? 2) Has a different orientation compared to the object in panel A

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.544037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.658106Z digest=sha256:cb33326523290b79ba37a24f0b1dc49c0d65d306fcf7abd6a69605f1d88d314e

Observation 680de0eb-c39b-4424-bf54-467d886c0059 · outbound

This paper cites Gemini and GPT 4o clearly excel in this task but all models show some level of understanding (Table 1).

Do large language vision models understand 3D shapes? Gemini and GPT 4o clearly excel in this task but all models show some level of understanding (Table 1)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.525021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.663124Z digest=sha256:7b573bdaab4b8744f2492c7be64150bbe894b9bf653fb7aa79b9a556c94190ee

Observation 8a5c8630-6636-4aac-a126-99e5bc22f2eb · outbound

This paper cites These results are consistent with previous works which show that despite their impressive performance Vision Language Models (VLM) often miss basic aspects of reality[8-12].

Do large language vision models understand 3D shapes? These results are consistent with previous works which show that despite their impressive performance Vision Language Models (VLM) often miss basic aspects of reality[8-12]

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.503699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.668844Z digest=sha256:53d86290ab801f367c6b6845d05e743ff3801659318dbbcdd7546f903bb03b0d

Observation 2e6bfac0-3978-4ddf-a952-d83db8758918 · outbound

This paper cites Code used to evaluate the models on the benchmark available at this URL.

Do large language vision models understand 3D shapes? Code used to evaluate the models on the benchmark available at this URL

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.477553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.674052Z digest=sha256:05308696bec851b582887b60c6728ae45fccbe72a6421cfcf598f90685820ae1

Observation 7c8821e9-3f3f-4f04-8fd5-2cb6bb35516d · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Do large language vision models understand 3D shapes? Flamingo: a visual language model for few-shot learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.458914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.679122Z digest=sha256:9436d4dc71e0504147e71c7f7fc0542c1d4702bc72ce96b70f12e48955afd357

Observation 50ea74b2-1b65-4e03-8042-7766b9cbe2cd · outbound

This paper cites GPT-4 Technical Report.

Do large language vision models understand 3D shapes? GPT-4 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.684975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.684975Z digest=sha256:20712b5dc53c12748afc34759c9b011a05fda2aca1e7ff76742dc44f265af7e9

Observation a9899b30-3309-445f-8474-9cff7f7ab3f7 · outbound

This paper cites Real-world robot applications of foundation models: A review.

Do large language vision models understand 3D shapes? Real-world robot applications of foundation models: A review

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.441207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.692018Z digest=sha256:36e3a7166e4d874b8eb1c56ade0c6c5f536a54796f87c99324d5c0cff79d758b

Observation 0ab174d4-2dda-4e41-8d9f-03f3439000f4 · outbound

This paper cites Approaching human 3D shape perception with neurally mappable models.

Do large language vision models understand 3D shapes? Approaching human 3D shape perception with neurally mappable models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.700235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.700235Z digest=sha256:cafcb8b6e0844ea2024c827cc62b2a43c95bc9a7b0881d6d4f3600fcffff11bc

Observation c95c0070-cf76-4ea2-a227-819545c5a26f · outbound

This paper cites Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models.

Do large language vision models understand 3D shapes? Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.421219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.706021Z digest=sha256:708dc607e70cbc99e649582ca9537dab2e2ab94fa57058a47cfdca070343f964

Observation b51dd177-3386-4663-a34b-04f75132c30f · outbound

This paper cites VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information.

Do large language vision models understand 3D shapes? VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.711784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.711784Z digest=sha256:f7a752e2dd2ecf98a0eb1f9ee2e405ba4f0e19f9366fbc538ded9308445fb6a8

Observation 19fba42e-ae53-433a-832d-7bb50df268e8 · outbound

This paper cites Can We Talk Models Into Seeing the World Differently?.

Do large language vision models understand 3D shapes? Can We Talk Models Into Seeing the World Differently?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.717893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.717893Z digest=sha256:2c136ffee268f57f124fce75b31695982c4627d407d945f31172646c1b4a9ecd

Observation 9b2d850c-06bd-4181-be6d-f1398fa68724 · outbound

This paper cites Exploring the frontier of vision-language models: A survey of current methodologies and future directions.

Do large language vision models understand 3D shapes? Exploring the frontier of vision-language models: A survey of current methodologies and future directions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.723920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.723920Z digest=sha256:c897debd197179b2e979284ccf77c167945dee930ec1b04ef8b1b448e51ac379

Observation b98382ba-30bc-42b8-915d-d86fbd198ffc · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Do large language vision models understand 3D shapes? BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.729389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.729389Z digest=sha256:2ea61690b4c3639408fcf877d35f3c9dbd0ff6af0326819b8d2600f93b4d05ef

Observation 89992980-f643-4bcd-9190-ff61916e22ca · outbound

This paper cites Can 3D Vision-Language Models Truly Understand Natural Language?.

Do large language vision models understand 3D shapes? Can 3D Vision-Language Models Truly Understand Natural Language?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.735368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.735368Z digest=sha256:34790685e20b9c443743b95cb2033ed070aef0fa9c010c181bf6917f3bea58ef

Observation 7767142a-5c19-4f6a-a0aa-7a0d04da4b3c · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Do large language vision models understand 3D shapes? Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.741436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.741436Z digest=sha256:ea41660a339ca6a73dc5de280ba9491e9c7a6f4dc697585ed59af06bfd62981a

Observation 7ed68ce4-b4cd-4062-aa72-90accc550ae2 · outbound

This paper cites GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?.

Do large language vision models understand 3D shapes? GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.747061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.747061Z digest=sha256:7663b89b5c0d8161a7218a2d6a41cbb57201942bd65caeaf0783f71962df1ef1

Observation ec3a7477-e733-413f-8262-f7e521146d38 · outbound

This paper cites A general protocol to probe large vision models for 3d physical understanding.

Do large language vision models understand 3D shapes? A general protocol to probe large vision models for 3d physical understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.384059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.752641Z digest=sha256:743ded7edb34b5809ada3284c69769131de21e8bcd008de9514a1d3b41929a18

Observation f96fe009-55fb-4756-a18d-1f54a1fb6376 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Do large language vision models understand 3D shapes? A Survey on Hallucination in Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.757593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.757593Z digest=sha256:2c8d54bc99dd1aa51b706ce6ed0bf58308ce6f43d70294b84df3d275f2944927

Observation 4765c8e0-0de6-46d9-9039-00709859032f · outbound

This paper cites When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models.

Do large language vision models understand 3D shapes? When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.764351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.764351Z digest=sha256:850bd55b7953f259a70f7d3f2531f6d5754eedcffe8b09455624928cd9fa936d

Observation add2d72e-3b04-461f-bfcc-1d9e03910ff2 · outbound

This paper cites Language-Image Models with 3D Understanding.

Do large language vision models understand 3D shapes? Language-Image Models with 3D Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.771499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.771499Z digest=sha256:8b94fcd5a1d9565e9827758fad3380e0f1be8c1bd509c7f37e85f5775614ff7a

Observation 56fee059-30fe-4cdb-86e0-03dacf78109a · outbound

This paper cites Shapellm: Universal 3d object understanding for embodied interaction.

Do large language vision models understand 3D shapes? Shapellm: Universal 3d object understanding for embodied interaction

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.362487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.778276Z digest=sha256:6bb386c549bde16eed6fca75d4d889218756c13610476f8a823fdd3ca4bfd098

Observation e5721b71-513f-4a10-bdab-8aa799b477dc · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

Do large language vision models understand 3D shapes? Objaverse: A universe of annotated 3d objects

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.345067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.786119Z digest=sha256:5a89089a4f2f3cb011aa935904d9928a544e3fc18cd1c2d66778814b663b5353

Observation 2af7571c-6ed6-4cee-82fc-5a8b44b5bc94 · outbound

This paper cites Vastextures: Vast repository of textures and PBR materials extracted from real-world images using unsupervised methods.

Do large language vision models understand 3D shapes? Vastextures: Vast repository of textures and PBR materials extracted from real-world images using unsupervised methods

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:32:59.792476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:32:59.792476Z digest=sha256:1f32076f29c6be788aa76ee69877c6b142e0ab7fc4161f8cf56d13e7f1347569

Observation bf7155e3-a1fc-48f5-ab46-ed90f98c84ec · outbound

This paper cites Infusing Synthetic Data with Real-World Patterns for Zero-Shot Material State Segmentation.

Do large language vision models understand 3D shapes? Infusing Synthetic Data with Real-World Patterns for Zero-Shot Material State Segmentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.325661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.798061Z digest=sha256:d5f3eb06f7f9ae21d5fcec455179930e13a5f6e360a44ba23c0d3cd1e189303c

Observation cc5021e5-b612-4a76-af9f-6b71d0dfdd6a · outbound

This paper cites Which panel contains an object that has identical 3D shape to the object in panel A, but different in orientation and texture. Explain.

Do large language vision models understand 3D shapes? Which panel contains an object that has identical 3D shape to the object in panel A, but different in orientation and texture. Explain

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:33:00.307386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:32:59.803624Z digest=sha256:3a6495736f414c2565eddfcd897347892eb7fdb4997d3f8af73cdd3b55b3bcf3

Pith citing papers

No inbound Pith citation observations are available.