Pith. sign in

Paper Citation Record · LEDGER

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

As of 23 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2607.21072.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21072 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:39:41.730369Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:54:58.326785Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T00:54:58.682244Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 751917e2-5586-4b9f-ab71-c3993b054f13 · outbound

This paper cites GPT-4 Technical Report.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:36.172840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:36.172840Z digest=sha256:268b7b024c7dec6f59da0b9c4a7a46446c0ded67c0b289e98a9230bc3b0b3590

Observation aa48be1d-8c24-4979-8ea7-c6e4f7a090c1 · outbound

This paper cites Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:36.595969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:36.595969Z digest=sha256:1dc7500f34cc88d3c1060860375aac0f0cd2931ee5da55262d0f7aaf6246cd56

Observation cfe40d53-b601-4d15-800a-1e0873088288 · outbound

This paper cites Emu3.5: Native Multimodal Models are World Learners.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Emu3.5: Native Multimodal Models are World Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:36.784561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:36.784561Z digest=sha256:0403c9f4a5fc18cc9b7bf441142dfc3b76b1ba8946f534f3c130f2989d2df3ce

Observation 21c8bbc7-3954-478e-bafd-bb833e4922bd · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:37.227319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:37.227319Z digest=sha256:e8c8e580e91ac6b557d8a81fbdcb237143f24e5ef84ce8995b739c510534bb2e

Observation 0450b52d-99f1-4deb-9547-b9dc204e2f84 · outbound

This paper cites Image Generators are Generalist Vision Learners.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Image Generators are Generalist Vision Learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:37.393050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:37.393050Z digest=sha256:fd4e78f1ae842caeb6db0da3dd5d65c1bce4433899b78f37a52f6609be2c3894

Observation 6856bfed-fbff-426f-83f5-a7da9ed8eca5 · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:37.870420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:37.870420Z digest=sha256:ec90557e1888a01735edf7040a47d6e15ea1f94240805df7b41e26ad93db3d4f

Observation df4d622e-7669-45e5-9a32-dddcb688a954 · outbound

This paper cites Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models.arXiv preprint arXiv:2506.03135 ,.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models.arXiv preprint arXiv:2506.03135 ,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:38.021033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:38.021033Z digest=sha256:572b8e6506dec407ef6a5a8c84691699487a1afbe2e43262c56b52d8e587dc50

Observation 3097d538-4df7-47cd-8105-22a90588c10a · outbound

This paper cites Joyai-image: Awakening spatial intelligence in unified multimodal understanding and generation.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Joyai-image: Awakening spatial intelligence in unified multimodal understanding and generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:38.201400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:38.201400Z digest=sha256:9121d57cef340b77ae9698b90223aa50503152aecf3280266b3bc459be01e3d4

Observation cc3edf69-dad9-4353-8630-c22152300c63 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:38.403605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:38.403605Z digest=sha256:7b155c69ed5b00bd04aa584cbf61aa91fd54ce7c7654a0ebf6197114e87ef3ef

Observation 61260bf7-9534-4614-98bc-bd70f4d3562e · outbound

This paper cites GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:38.560857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:38.560857Z digest=sha256:3ea3c85a156db5154e12773512ddba0bd6e79f070bad03567f30b45c3e951b83

Observation b15d9875-5b1b-46dc-bf3f-38f2e4209192 · outbound

This paper cites Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:38.723115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:38.723115Z digest=sha256:fee057813ed128ec8e000f11eab6f4cbd26af6021dbe9974a3fb20f86b626f79

Observation bacd8e71-3b47-47d5-a705-9188e002bd1c · outbound

This paper cites Ssr: Enhancing depth perception in vision-language models via rationale-guided spatial reasoning.arXiv preprint arXiv:2505.12448,.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Ssr: Enhancing depth perception in vision-language models via rationale-guided spatial reasoning.arXiv preprint arXiv:2505.12448,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:38.919903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:38.919903Z digest=sha256:266f03bc1ef648275aa60b795e8103dea738026ed210e6e7b1da3f213f128914

Observation 0ecfc0cc-3e16-461d-aa21-0829e0c2e238 · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.240342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.240342Z digest=sha256:cf0289dc8ef0aadfb6279c64759e9693b1d819105d2cbe197e3b752efc141cc2

Observation 32aad53a-00ba-4532-a80a-cc3a6becb46d · outbound

This paper cites Sat: Dynamic spatial aptitude training for multimodal language models.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Sat: Dynamic spatial aptitude training for multimodal language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.342345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.342345Z digest=sha256:1bee360329948629cf87e8917b88b4cd3ce6d8dbb1abd006e2cb7ddd01372190

Observation 7f54c391-f3e8-4db5-8355-8b6c97d1afbc · outbound

This paper cites Seedream 4.0: Toward Next-generation Multimodal Image Generation.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Seedream 4.0: Toward Next-generation Multimodal Image Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.437251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.437251Z digest=sha256:5539ccf13355125df3ae6b3bc8f6140efac0df98e3f285ea3a55929a77d390cd

Observation ce3e6dab-3e1f-4d8f-a056-01001d16d180 · outbound

This paper cites Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.509832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.509832Z digest=sha256:9299b5c2430aa3e23583cf23a1d9adb336134ffbbe7956da9daeaf2eb9d05bbb

Observation abf7b80b-6f2e-4bb7-b642-502c0a25c831 · outbound

This paper cites Yingbo Tang, Lingfeng Zhang, Shuyi Zhang, Yinuo Zhao, and Xiaoshuai Hao.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Yingbo Tang, Lingfeng Zhang, Shuyi Zhang, Yinuo Zhao, and Xiaoshuai Hao

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.626649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.626649Z digest=sha256:c189eb0561f747f59191a156941133fcc83ad6cb575527360452cf393bec7ee7

Observation 8d4eb2e2-0fdc-4ac0-ab67-deb7ac899a1d · outbound

This paper cites org/10.1145/3746027.3758209.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text org/10.1145/3746027.3758209

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.740636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.740636Z digest=sha256:6e69c1f1d9d8521eda9653c01d345424bf548beac1f422e0196f395e591acf34

Observation c1a2063d-680f-4928-86e8-7520052c9167 · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.852110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.852110Z digest=sha256:49260e62fd154f2e81fa9d26ed99ec6128d85cd8d9d80f24bc26bb2f63692187

Observation 162d9b19-9942-4e80-a14c-7f2a8ed002b9 · outbound

This paper cites Mindcube: Spatial mental modeling from limited views, 2026a.https://arxiv.org/abs/2506.21458.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Mindcube: Spatial mental modeling from limited views, 2026a.https://arxiv.org/abs/2506.21458

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.988118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.988118Z digest=sha256:03309b6f7cb3885d3889f6683b8f1d4cc645e138e8cc70ce828d7e1e1831f055

Observation 91b4fefa-9314-4d1d-8710-b28411d75609 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:40.120708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:40.120708Z digest=sha256:1199fa175d2e2d69203ea2ab1c780e7465cda576c8be6a936a6ca4092964e18b

Observation c9d49f14-f828-49f3-b6b7-335078e0de3d · outbound

This paper cites Qiucheng Wu, Handong Zhao, Michael Saxon, Trung Bui, William Yang Wang, Yang Zhang, and Shiyu Chang.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Qiucheng Wu, Handong Zhao, Michael Saxon, Trung Bui, William Yang Wang, Yang Zhang, and Shiyu Chang

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:40.241042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:40.241042Z digest=sha256:dbd5139ca409b66ab95bcad589fda17f2497507a64043d591161d2059401a552

Observation c2253254-bb0b-4b8b-b93a-96a24195c4dd · outbound

This paper cites SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:40.396402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:40.396402Z digest=sha256:84fafcf88f05c37314599eba4d626c87e31716b0c1393bc710125e9839b151ff

Observation d395aebd-a326-4e01-9ce4-ee26ee7bef62 · outbound

This paper cites Sphere: Unveiling spatial blind spots in vision-language models through hierarchical evaluation.ACL, 2025a.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Sphere: Unveiling spatial blind spots in vision-language models through hierarchical evaluation.ACL, 2025a

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:40.520878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:40.520878Z digest=sha256:8c2c805e795cf5eee889aac9e4d70beb60ac7fe8cc5fbe11bca12c7a4edcdac0

Observation 9a094702-8da5-4b8d-b8e7-8868f3660ee5 · outbound

This paper cites GEM: Generative Supervision Helps Embodied Intelligence.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text GEM: Generative Supervision Helps Embodied Intelligence

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:40.704498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:40.704498Z digest=sha256:1c138185f1b7dc8b7c0a76f445416dd0aaabe90bc333800561070edc9ffb90a5

Observation 275e632a-55d9-4624-9b16-655157e579a5 · outbound

This paper cites Roborefer: Towards spatial referring with reasoning in vision-language models for robotics.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Roborefer: Towards spatial referring with reasoning in vision-language models for robotics

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:40.853401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:40.853401Z digest=sha256:321ba0a1a0831e07ff54ad262d74f125233a871c015ce627bec93003201ff001

Observation 719611ca-6e89-463f-8885-d5f9cfc487d4 · outbound

This paper cites • Appendix B documents data sources, schema, and quality control.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text • Appendix B documents data sources, schema, and quality control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:41.054054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:41.054054Z digest=sha256:71227c19a8edccd13183158af69ffd31bbe7616dbc2740d57f7955dee014d5ae

Observation e0f07522-73ef-4bab-822e-ab3c1a7b9294 · outbound

This paper cites GPT-5.4 OpenAI-compatible chat API; request model gpt-5.4; direct structured answer; deterministic request settings where supported.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text GPT-5.4 OpenAI-compatible chat API; request model gpt-5.4; direct structured answer; deterministic request settings where supported

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:41.232662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:41.232662Z digest=sha256:4af574b7376b2e820f6481d5e244a43b832496d57dbc6e5ebc77b5e874a6c7f8

Observation 80b07f73-2746-4696-a9e8-5840ce7f5e30 · outbound

This paper cites VLM-3R-7B; SpatialRGPT-8B; Spatial-MLLM; SpatialBot-3B (Fan et al., 2025; Cheng et al., 2024; Wu et al., 2025b; Cai et al.,.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text VLM-3R-7B; SpatialRGPT-8B; Spatial-MLLM; SpatialBot-3B (Fan et al., 2025; Cheng et al., 2024; Wu et al., 2025b; Cai et al.,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:41.411567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:41.411567Z digest=sha256:2118a3b64d593e309f3964a46ffd71e3668ba0cd3c06985505bad32fd2484cd2

Observation d6a9c072-c79c-434e-83f4-5eeea6c1547f · outbound

This paper cites SenseNova-Vision-7B-MoT; Janus-Pro-7B; Janus-1.3B (Han et al., 2026; Chen et al., 2025; Wu et al.,.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text SenseNova-Vision-7B-MoT; Janus-Pro-7B; Janus-1.3B (Han et al., 2026; Chen et al., 2025; Wu et al.,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:41.612334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:41.612334Z digest=sha256:1beb69c27ea3fcc7dbba20a1da80ad3e92dad24e3bf1319422aba9ac6030ec82

Observation 362e470c-d8fc-4a98-83f6-52048146213b · outbound

This paper cites 32 Table 17 continued.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text 32 Table 17 continued

Reference 1024

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:41.730369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:41.730369Z digest=sha256:34a570d35147603d4b247ac1a7418476b172d3967795e2e6cd561a3315ea76e8

Observation 23f1e4ba-6321-4b1e-a023-24f3724d5ea7 · outbound

This paper cites Vision as Unified Multimodal Generation.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Vision as Unified Multimodal Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:37.738395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:37.738395Z digest=sha256:67866a3e8556d326a113112307864b359b4d8ac8f20d0a29a2f2f1c65f0a2e77

Observation d397bc5e-c954-44b6-a872-c5f290657a1a · outbound

This paper cites Benchmarking Spatial Relationships in Text-to-Image Generation.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:37.518805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:37.518805Z digest=sha256:5e8e5541bcecc48f7737edfd2bbbc0ac3952b905189b725fae49ae637b369b15

Observation 6892ec34-3a4b-4b7b-9bc9-800038bb5662 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:36.464785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:36.464785Z digest=sha256:32edb91235608070de8220606634f4b86333238fe9cf26a195b19986cbf01dcd

Observation 7262283c-de5f-4f5b-b544-fee1b27583e0 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:36.317366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:36.317366Z digest=sha256:fbabc2be70401c879b2e657b9de44bd210c707e45dc532c73d8b9827ccfddb2c

Observation 32c9d74b-f98e-463d-964f-e35f4f97a71d · outbound

This paper cites 14 Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, et al.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text 14 Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, et al

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:37.090693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:37.090693Z digest=sha256:cacf74bba45d777d77191b4ec016a8aaaffd6f76d97117fc6c336929c5e53f3f

Pith citing papers

Observation 7dc54b9e-9f3e-4f75-a76b-d0ca4a5dc8e9 · inbound

Image-Space Rule Discovery cites this paper.

Image-Space Rule Discovery Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T00:54:58.688729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T00:54:58.326785Z digest=sha256:f3997ca27d9038a46cf171aeac845f14689ad26d53bfb99fb063dc44b4a0b334