Pith. sign in

Paper Citation Record · LEDGER

3D Primitives are a Spatial Language for VLMs

As of 13 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2605.12586.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12586 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-14T21:32:52.151978Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact11
  • verified fuzzy5
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29923a26-1915-402c-9687-486c095ff4dd · outbound

This paper cites The spatial blindspot of vision-language models.arXiv preprint arXiv:2601.09954.

3D Primitives are a Spatial Language for VLMs The spatial blindspot of vision-language models.arXiv preprint arXiv:2601.09954

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:58.971089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:32:52.151978Z digest=sha256:2c732924fb069fbe46c6e490763e5a355f1b3e6478467470787adfb500aa94e1

Observation 88bb82c5-b168-4f3c-851e-13051a645db3 · outbound

This paper cites Program Synthesis with Large Language Models.

3D Primitives are a Spatial Language for VLMs Program Synthesis with Large Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:59.011813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:32:52.151978Z digest=sha256:0f9cf544a7276c1226b26fa5c1867ebb27200cdc67cf2228ef6eeb7f84df5b8a

Observation ddaba844-5d0e-4118-92c3-703ae95b5d4e · outbound

This paper cites Evaluating Large Language Models Trained on Code.

3D Primitives are a Spatial Language for VLMs Evaluating Large Language Models Trained on Code

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:58.984612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:32:52.151978Z digest=sha256:9cf1a9379f5e75ca07b68c35b4e8e7c86d2f6e03cb413e3f76a91e8f0cd71ccb

Observation 6ce0409e-659a-400d-a9b0-c2e9eeda9b7e · outbound

This paper cites Why is spatial reasoning hard for vlms? an attention mechanism perspective on focus areas.

3D Primitives are a Spatial Language for VLMs Why is spatial reasoning hard for vlms? an attention mechanism perspective on focus areas

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:58.976091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:32:52.151978Z digest=sha256:39d482e5a06922e2d03a8ee9d933e7ddf33b4451705ee05f2d1f8131a8fe4066

Observation 6e819d70-b53c-4a08-a7e8-fd370f09b088 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

3D Primitives are a Spatial Language for VLMs SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:58.966580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:32:52.151978Z digest=sha256:7bc245bde38f76f6abe60e326ee4c414d759fb5539155bc266f490d36ae62192

Observation 3919e01c-597e-449d-8440-06452b2dcc1b · outbound

This paper cites Linear mechanisms for spa- tiotemporal reasoning in vision language models.arXiv preprint arXiv:2601.12626.

3D Primitives are a Spatial Language for VLMs Linear mechanisms for spa- tiotemporal reasoning in vision language models.arXiv preprint arXiv:2601.12626

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:58.989163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:32:52.151978Z digest=sha256:1e36d7cbeb82cf3c1895e441ebf494514723590ff84487044c5f826a1f88e135

Observation 8251b0b6-aeef-4f99-8a9e-9ef63a733bc7 · outbound

This paper cites Llava-st: A multimodal large language model for fine-grained spatial-temporal understanding.

3D Primitives are a Spatial Language for VLMs Llava-st: A multimodal large language model for fine-grained spatial-temporal understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:58.980495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:32:52.151978Z digest=sha256:00147f00d4b6c820e3bc81bbfe51dfbe5874b0e30dca0825951ae0a5115aee01

Observation d3bef577-88ac-470e-aec5-a1f658d10457 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

3D Primitives are a Spatial Language for VLMs MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:58.993340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:32:52.151978Z digest=sha256:f48b976450780772454fc174726e338d4a2753aa59db0530edf25a60232bab45

Observation 5718c137-5411-46f9-aa94-662b9553a17e · outbound

This paper cites Beyond semantics: Rediscovering spatial awareness in vision-language models.

3D Primitives are a Spatial Language for VLMs Beyond semantics: Rediscovering spatial awareness in vision-language models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.007937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:32:52.151978Z digest=sha256:7c2ca6e74a0bb9c789f920164fb189faa693d01141c5fe3e9ab5374a39a9fd4d

Observation 85ec4c39-6a78-462e-ba08-a55bebc4d37c · outbound

This paper cites Design2code: Benchmarking multimodal code generation for automated front-end engineering.

3D Primitives are a Spatial Language for VLMs Design2code: Benchmarking multimodal code generation for automated front-end engineering

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T21:33:00.156018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:32:52.151978Z digest=sha256:95f1c15615fcfb8b5caa5ad8ee952785bca9a1c3d97cd8c5ab532a7c1de76217

Observation b26cd689-e332-4ab7-b3b7-eb13164ff188 · outbound

This paper cites SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition.

3D Primitives are a Spatial Language for VLMs SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:58.996935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:32:52.151978Z digest=sha256:31886bef21f4df9ee4f0f1a4c8a82ad455fcb028fcfe42e5acfb0cea75a26cd1

Observation 22ac36c4-05e4-483a-a7f6-901f8403ec98 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

3D Primitives are a Spatial Language for VLMs Instruction-Following Evaluation for Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:32:58.961463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:32:52.151978Z digest=sha256:1a2f192d5a4144492c512cdd94c6a76a73f009744eea9ae26d2060c73404fda5

Observation 420c7853-ffb1-46d8-bd5a-dca623eecb9d · outbound

This paper cites scene_id.

3D Primitives are a Spatial Language for VLMs scene_id

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T21:33:00.159014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:32:52.151978Z digest=sha256:bd5ec1b62d50804cedc9cadf43d2ffaecc4caf03fb4931017bc309cdbeadc773

Observation 53b51ad9-1b5b-46ad-a946-946cf8570ae0 · outbound

This paper cites best_cc /worst_cc use each model’s primitive-best / primitive-worst scene-code language from Table 1.∆best-direct = best_cc−direct(pp).

3D Primitives are a Spatial Language for VLMs best_cc /worst_cc use each model’s primitive-best / primitive-worst scene-code language from Table 1.∆best-direct = best_cc−direct(pp)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T21:33:00.145737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:32:52.151978Z digest=sha256:d74b5a1c50518bb0731642830a9c9ccb8a8452f84929fabba798315909531f37

Observation 05425b58-f0fc-4d3f-b2ce-6a9a31f3ae01 · outbound

This paper cites single-view S3-FT lifts SpatialBabel-QA by ∼+7% averaged across modes on Qwen3-VL-8B.

3D Primitives are a Spatial Language for VLMs single-view S3-FT lifts SpatialBabel-QA by ∼+7% averaged across modes on Qwen3-VL-8B

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T21:33:00.149167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:32:52.151978Z digest=sha256:6f729642cb7015e77d2e1e69d5afbe8e4d5d23d6cbb094169533db6df8274e15

Observation 3e247290-6e19-4148-894d-e046f4108f84 · outbound

This paper cites Pass rates: Qwen3-VL-8B 82%, Qwen2.5-VL-7B 84%.

3D Primitives are a Spatial Language for VLMs Pass rates: Qwen3-VL-8B 82%, Qwen2.5-VL-7B 84%

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T21:33:00.152345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:32:52.151978Z digest=sha256:1c7b095cb8b9ef68cb4a2ebffb3f6077650a9361b6dfd64a148065ea8ca62a17

Pith citing papers

No inbound Pith citation observations are available.