Pith. sign in

Paper Citation Record · LEDGER

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

As of 9 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2607.21072.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21072 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:39:41.730369Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:54:58.326785Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T00:54:58.682244Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 751917e2-5586-4b9f-ab71-c3993b054f13 · outbound

This paper cites GPT-4 Technical Report.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:36.172840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:36.172840Z digest=sha256:ce5460bad16b433cca960325f121f2a1f06fec91d51a8d9f57784539fcaa69a5

Observation aa48be1d-8c24-4979-8ea7-c6e4f7a090c1 · outbound

This paper cites Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:36.595969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:36.595969Z digest=sha256:473f283ea9557a51311f86c20a597ab10b7dab35c553447de7cd7d09938cebec

Observation cfe40d53-b601-4d15-800a-1e0873088288 · outbound

This paper cites Emu3.5: Native Multimodal Models are World Learners.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Emu3.5: Native Multimodal Models are World Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:36.784561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:36.784561Z digest=sha256:e3c8ba85680415334b3d1b4647f7f20d745eb755baee4b728be6f2d9367ea26e

Observation 21c8bbc7-3954-478e-bafd-bb833e4922bd · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:37.227319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:37.227319Z digest=sha256:3486d42295f9c28919914558ad60c817c92a6fdb0389e6fcb3857e986f3c2352

Observation 0450b52d-99f1-4deb-9547-b9dc204e2f84 · outbound

This paper cites Image Generators are Generalist Vision Learners.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Image Generators are Generalist Vision Learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:37.393050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:37.393050Z digest=sha256:1ee2596d19bbc2fa0a016fce8a8605d7894b8ed84aac24397832edb10e7b0144

Observation 6856bfed-fbff-426f-83f5-a7da9ed8eca5 · outbound

This paper cites RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:37.870420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:37.870420Z digest=sha256:62e1f6018ed7d4743a8616ac2708694ca4bc3335f25e59eb5b3d1c7a5718b537

Observation df4d622e-7669-45e5-9a32-dddcb688a954 · outbound

This paper cites Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models.arXiv preprint arXiv:2506.03135 ,.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Omnispatial: Towards comprehensive spatial reasoning benchmark for vision language models.arXiv preprint arXiv:2506.03135 ,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:38.021033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:38.021033Z digest=sha256:9982a1654ce760f6a719f2dbf3d0ab37d146dd3d2a51e7e0bc5d8664c6b144ff

Observation 3097d538-4df7-47cd-8105-22a90588c10a · outbound

This paper cites Joyai-image: Awakening spatial intelligence in unified multimodal understanding and generation.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Joyai-image: Awakening spatial intelligence in unified multimodal understanding and generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:38.201400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:38.201400Z digest=sha256:fcd4059e328401ed95db927a5e8ad15a3add4b2c4f1fda8412154f0fe4ca9778

Observation cc3edf69-dad9-4353-8630-c22152300c63 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:38.403605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:38.403605Z digest=sha256:6d3ca2ea6573b176b3f2e2e89cd449e949a01809de291c61cfe8e39191992fa6

Observation 61260bf7-9534-4614-98bc-bd70f4d3562e · outbound

This paper cites GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:38.560857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:38.560857Z digest=sha256:c0d890c492eec301ef0f105a52d3ebd2dc0bd14dae61c4e40878c44c762613ec

Observation b15d9875-5b1b-46dc-bf3f-38f2e4209192 · outbound

This paper cites Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:38.723115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:38.723115Z digest=sha256:34a9a78ae32d8202f5f6fed0be77cf0c77959f10fa5b9646b45cdd0882066b48

Observation bacd8e71-3b47-47d5-a705-9188e002bd1c · outbound

This paper cites Ssr: Enhancing depth perception in vision-language models via rationale-guided spatial reasoning.arXiv preprint arXiv:2505.12448,.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Ssr: Enhancing depth perception in vision-language models via rationale-guided spatial reasoning.arXiv preprint arXiv:2505.12448,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:38.919903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:38.919903Z digest=sha256:55b8a4b16b98755cddc65df78ccc1094f9d101db2509f51e54864d1baa7a4e87

Observation 0ecfc0cc-3e16-461d-aa21-0829e0c2e238 · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.240342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.240342Z digest=sha256:9d64b5720f185bd7ce73055b43d988d3d086c48d62eb3e9d132a8e9697926a08

Observation 32aad53a-00ba-4532-a80a-cc3a6becb46d · outbound

This paper cites Sat: Dynamic spatial aptitude training for multimodal language models.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Sat: Dynamic spatial aptitude training for multimodal language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.342345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.342345Z digest=sha256:0fe8abdd46f38a0570ef5fb9a6bce8fbfec6ee5f7684310f385af922cf0e577e

Observation 7f54c391-f3e8-4db5-8355-8b6c97d1afbc · outbound

This paper cites Seedream 4.0: Toward Next-generation Multimodal Image Generation.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Seedream 4.0: Toward Next-generation Multimodal Image Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.437251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.437251Z digest=sha256:9b0abc64c4042479b64e2b6ab2b0ac100eec67c1ffab0532348343514d2a499c

Observation ce3e6dab-3e1f-4d8f-a056-01001d16d180 · outbound

This paper cites Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.509832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.509832Z digest=sha256:0d18199cda300c8ac5e7f3aedb3dc59a16a9e5da24a2db560d0e28dfb5225716

Observation abf7b80b-6f2e-4bb7-b642-502c0a25c831 · outbound

This paper cites Yingbo Tang, Lingfeng Zhang, Shuyi Zhang, Yinuo Zhao, and Xiaoshuai Hao.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Yingbo Tang, Lingfeng Zhang, Shuyi Zhang, Yinuo Zhao, and Xiaoshuai Hao

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.626649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.626649Z digest=sha256:8bb6e436d77b8474a94a6f790b9fe8243faf69c465f6168daed3bd1a79ea3ad5

Observation 8d4eb2e2-0fdc-4ac0-ab67-deb7ac899a1d · outbound

This paper cites org/10.1145/3746027.3758209.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text org/10.1145/3746027.3758209

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.740636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.740636Z digest=sha256:0fb0ab7829f5b695a114401322f203581760efe27de1b0230f0e5cc2646e09e2

Observation c1a2063d-680f-4928-86e8-7520052c9167 · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.852110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.852110Z digest=sha256:32de594f438fd2991e84c1dfee2bc74eebde0e0125345841444af115ed571156

Observation 162d9b19-9942-4e80-a14c-7f2a8ed002b9 · outbound

This paper cites Mindcube: Spatial mental modeling from limited views, 2026a.https://arxiv.org/abs/2506.21458.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Mindcube: Spatial mental modeling from limited views, 2026a.https://arxiv.org/abs/2506.21458

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:39.988118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:39.988118Z digest=sha256:87fae98334e8898423660247190d5ec75af2b14c5fd2107eab0c4bd64ce57a32

Observation 91b4fefa-9314-4d1d-8710-b28411d75609 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:40.120708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:40.120708Z digest=sha256:7b0af94b3e7da84b12905647eb94da9dcd78130994b4452d7edaaefba6e487ec

Observation c9d49f14-f828-49f3-b6b7-335078e0de3d · outbound

This paper cites Qiucheng Wu, Handong Zhao, Michael Saxon, Trung Bui, William Yang Wang, Yang Zhang, and Shiyu Chang.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Qiucheng Wu, Handong Zhao, Michael Saxon, Trung Bui, William Yang Wang, Yang Zhang, and Shiyu Chang

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:40.241042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:40.241042Z digest=sha256:8e81f7507d6d854727029a9240a89e5fe7ae71cdcd127cab7e2e1806b3e5ca31

Observation c2253254-bb0b-4b8b-b93a-96a24195c4dd · outbound

This paper cites SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:40.396402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:40.396402Z digest=sha256:af3e44e06ee6d6b826ef567e6d8c61405934ed61552fec8e18639c5c458db3eb

Observation d395aebd-a326-4e01-9ce4-ee26ee7bef62 · outbound

This paper cites Sphere: Unveiling spatial blind spots in vision-language models through hierarchical evaluation.ACL, 2025a.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Sphere: Unveiling spatial blind spots in vision-language models through hierarchical evaluation.ACL, 2025a

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:40.520878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:40.520878Z digest=sha256:63fe7003ae81904b964bb3acf9b99f69c283111bd8e0d39dc04e86752f5fe988

Observation 9a094702-8da5-4b8d-b8e7-8868f3660ee5 · outbound

This paper cites GEM: Generative Supervision Helps Embodied Intelligence.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text GEM: Generative Supervision Helps Embodied Intelligence

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:40.704498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:40.704498Z digest=sha256:ebd2a55a4de58b9b5f1f4de16b1f7f3a3dcc27e98d6e74a7759d744a36f9c7ab

Observation 275e632a-55d9-4624-9b16-655157e579a5 · outbound

This paper cites Roborefer: Towards spatial referring with reasoning in vision-language models for robotics.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Roborefer: Towards spatial referring with reasoning in vision-language models for robotics

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:40.853401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:40.853401Z digest=sha256:c25f23d47ddca25303ac75cfc78937119bcba658c1c3226d07fc4ec47d5c7eae

Observation 719611ca-6e89-463f-8885-d5f9cfc487d4 · outbound

This paper cites • Appendix B documents data sources, schema, and quality control.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text • Appendix B documents data sources, schema, and quality control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:41.054054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:41.054054Z digest=sha256:3adbc77bc1288329d4feda4cad23762877b068e97e522dfabcb71bae038a5aa0

Observation e0f07522-73ef-4bab-822e-ab3c1a7b9294 · outbound

This paper cites GPT-5.4 OpenAI-compatible chat API; request model gpt-5.4; direct structured answer; deterministic request settings where supported.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text GPT-5.4 OpenAI-compatible chat API; request model gpt-5.4; direct structured answer; deterministic request settings where supported

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:41.232662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:41.232662Z digest=sha256:ef6b151118514cd16526243286c5c519b6533e6058987854a8a36653578bd0a8

Observation 80b07f73-2746-4696-a9e8-5840ce7f5e30 · outbound

This paper cites VLM-3R-7B; SpatialRGPT-8B; Spatial-MLLM; SpatialBot-3B (Fan et al., 2025; Cheng et al., 2024; Wu et al., 2025b; Cai et al.,.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text VLM-3R-7B; SpatialRGPT-8B; Spatial-MLLM; SpatialBot-3B (Fan et al., 2025; Cheng et al., 2024; Wu et al., 2025b; Cai et al.,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:41.411567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:41.411567Z digest=sha256:3dd8376ea5b5c8bb49f77ad3f0b8c23864f32c3f153309db59b3dbb75614b45b

Observation d6a9c072-c79c-434e-83f4-5eeea6c1547f · outbound

This paper cites SenseNova-Vision-7B-MoT; Janus-Pro-7B; Janus-1.3B (Han et al., 2026; Chen et al., 2025; Wu et al.,.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text SenseNova-Vision-7B-MoT; Janus-Pro-7B; Janus-1.3B (Han et al., 2026; Chen et al., 2025; Wu et al.,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:41.612334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:41.612334Z digest=sha256:2f81ab8ba88972ffec2cc43a8542d133b5d821e8ec768fca00e41a0f3f8afab2

Observation 362e470c-d8fc-4a98-83f6-52048146213b · outbound

This paper cites 32 Table 17 continued.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text 32 Table 17 continued

Reference 1024

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:41.730369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:41.730369Z digest=sha256:ebff8a327903eb44684dc4741e8252abc203f253fa62447080132d121df5caff

Observation 23f1e4ba-6321-4b1e-a023-24f3724d5ea7 · outbound

This paper cites Vision as Unified Multimodal Generation.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Vision as Unified Multimodal Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:37.738395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:37.738395Z digest=sha256:8a6fa9e334d079e67b551ca9c0e808b63edb42a078ec2f039d1b78423c82b360

Observation d397bc5e-c954-44b6-a872-c5f290657a1a · outbound

This paper cites Benchmarking Spatial Relationships in Text-to-Image Generation.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:37.518805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:37.518805Z digest=sha256:430a7c7032b73f4189823d5f62dbb8b2569af20a8ce7bca00ce924bd44cff342

Observation 6892ec34-3a4b-4b7b-9bc9-800038bb5662 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:36.464785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:36.464785Z digest=sha256:31072a301be68d07b1095bebbd6650fadff6b38bc896deaff15da46770bb4823

Observation 7262283c-de5f-4f5b-b544-fee1b27583e0 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:36.317366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:36.317366Z digest=sha256:942c8775c6d4ea12d505b23a7e975c388d68900bbb7010e20dacda9d950036e4

Observation 32c9d74b-f98e-463d-964f-e35f4f97a71d · outbound

This paper cites 14 Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, et al.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text 14 Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, et al

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:37.090693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:37.090693Z digest=sha256:cbde5eedbea363147212dd7b4acc9899c96ac3167e373860ad310fca5f5e8c12

Pith citing papers

Observation 7dc54b9e-9f3e-4f75-a76b-d0ca4a5dc8e9 · inbound

Image-Space Rule Discovery cites this paper.

Image-Space Rule Discovery Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T00:54:58.688729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T00:54:58.326785Z digest=sha256:8978d8a9988903a6be1e6e4a79d20f314b722d7b83d2fe04be50b656bb98576f