Pith. sign in

Paper Citation Record · LEDGER

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

As of 11 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 16 inbound Pith citation observations for arXiv:2601.11109.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.11109 v3

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T13:50:19.988813Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:13:38.497502Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T14:08:21.569823Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact10
  • verified fuzzy39
  • unresolved8
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch13

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 371c5de1-0fc7-4f98-8964-caa73b782aa0 · outbound

This paper cites Open-Universe Indoor Scene Generation using LLM Program Synthesis and Uncurated Object Databases.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Open-Universe Indoor Scene Generation using LLM Program Synthesis and Uncurated Object Databases

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:51:00.618680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:4c96ef088fd42da55a708d62d56fff38ca6e4ac751a5feaf8a81f906d2bf7a52

Observation b527588f-3ccc-42be-973b-84b89c6335e3 · outbound

This paper cites Retrieved from https://deepmind.google/models/gemini/pro/ 22.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Retrieved from https://deepmind.google/models/gemini/pro/ 22

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.798004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:254675cf73455c48eb3233d0956989b5071a80963554f65444c58f2ad322fbf7

Observation 164f08a6-a54c-4c04-9b02-566c5954b76b · outbound

This paper cites Retrieved from https://docs.cloud.google.com/vertex- ai/generative-ai/docs/partner-models/claude/sonnet-4 22.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Retrieved from https://docs.cloud.google.com/vertex- ai/generative-ai/docs/partner-models/claude/sonnet-4 22

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.802405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:667869e58c89ec10af1e49c1b451b2e2e75ba9209503c1361fcaa98689dbc965

Observation 8ecbb215-b6d2-4639-ba5c-a0f0e681b0e1 · outbound

This paper cites an unresolved cited work.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:51:00.793704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:eb1737ed6336f1f5f9263b5ff649592ebf8e89970a786f9d0d68a581e844ad17

Observation 7cf3085e-b683-4840-b211-be58240804eb · outbound

This paper cites Reflection-Based Memory For Web navigation Agents.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Reflection-Based Memory For Web navigation Agents

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:51:00.615430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:2db6a4832fb34999e0b18f4b919dc4660c1acbaff4d7696c0fd3ea039e3dc8d7

Observation 0b501b7a-0aa3-49d4-abcc-4513611c61b1 · outbound

This paper cites In: Mahamood, S., Minh, N.L., Ippolito, D.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Mahamood, S., Minh, N.L., Ippolito, D

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.808281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:e4f901ac8f6eb6b4ff0bb8bee103d8c9833a8e0f29716cc639e3a1e3f9b133d2

Observation 8dc5bddb-07ea-4b2c-9b08-e6c8fda876aa · outbound

This paper cites https://doi.org/10.18653/v1/2024.inlg-main.184.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning https://doi.org/10.18653/v1/2024.inlg-main.184

Reference 7

Resolution
verified exact
doi, observed 2026-05-16T13:51:00.548634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:18fbb0c011e2a7a6ed9656d871c79ec1c060fa72b90125bbb41d73385859ea22

Observation d609d1c3-e535-4552-92b3-414fcdb573ee · outbound

This paper cites Seminal Graphics Papers: Pushing the Boundaries, Volume 2 (1999),https: //api.semanticscholar.org/CorpusID:2037052112.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Seminal Graphics Papers: Pushing the Boundaries, Volume 2 (1999),https: //api.semanticscholar.org/CorpusID:2037052112

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.800364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:2e61da0d8f04b2463cd139009158e31860f3cd8cf16d887c9d6e3ee717d23039

Observation d5af8d3f-9186-4b77-9d55-809644fa4f07 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.791578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:ec390041b5866e33c9796b4ce0b0f3a544e6b2d28180ad2f987ac2f41126af32

Observation 407a0dc7-594c-45b0-81ec-cf1d4d6d2aa7 · outbound

This paper cites In: NeurIPS (2024) 4.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: NeurIPS (2024) 4

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.795979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:28440468732ae895c2e8c038b25ba5cebb28da6c3bdf2377b0a13b01f879b956

Observation 5f81eebe-1ad7-4115-ae69-72cad6d5a42d · outbound

This paper cites In: NeurIPS (2022), outstanding Paper Award 4.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: NeurIPS (2022), outstanding Paper Award 4

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.814147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:2a70e9143b420c80d1fddaa7f59d92c8b0ba5c4fc732d8e39e2aa49ba7c97b8f

Observation 444e5b12-2622-44a3-9302-58fee1755a50 · outbound

This paper cites In: Proceedings of the British Machine Vision Conference (BMVC).

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Proceedings of the British Machine Vision Conference (BMVC)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.804251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:3d284be8f6faa5f0ff2fb33db3a0af2379218d79274e85bbfd33b69f79c6a811

Observation 4343d635-5364-4de9-bb4d-7a95fe1c92b0 · outbound

This paper cites In: Proceedings of the 37th International Conference on Neural Information Processing Systems.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Proceedings of the 37th International Conference on Neural Information Processing Systems

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.806393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:841ef34e96d227edf4343ae316114a3e35d05d5833edf17eb994e53825cde867

Observation 484d6337-c0f9-43b1-8518-502f5028c185 · outbound

This paper cites ACM Transactions on Graphics (ToG), Proc.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning ACM Transactions on Graphics (ToG), Proc

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.810281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:515586ffbdd6e49cedf4d7de1b8f51d8007b41c3e71d5a0a022687f427484b8b

Observation 541e9ff6-9545-4429-8ce9-bfd7d4e6639a · outbound

This paper cites In: European Conference on Computer Vision.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: European Conference on Computer Vision

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.860950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:64ee59322da8e9480fdcfc2b1349c3979728f1a52d4f771bf28473d6f4f52e42

Observation 60d50e7d-a8b3-4a7f-ac73-c75935132fc0 · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.850281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:7cf06f23d1cb0b8e9111614fb4e13c198c1aa037cf08300a21c38b4c6488b31f

Observation 9409b5e1-e2b5-4269-bda1-2c4e5a47e506 · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.867204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:1573406b6708967ff215139b2e513f29b012c9f0c39632e3f5102490430ca0ab

Observation f8130f4e-dbe3-494f-8107-6995eefdc330 · outbound

This paper cites Founda- tions and Trends® in Programming Languages 4(1-2), 1–119 (2017).

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Founda- tions and Trends® in Programming Languages 4(1-2), 1–119 (2017)

Reference 18

Resolution
verified exact
doi, observed 2026-05-16T13:51:00.539776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:ad6199581c4a804c02f5b57094a2bf453649d9c4b028f5e21f8aba2102c1bccd

Observation cd414a45-69a3-43e1-b9e2-6445f6a81cc7 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.886490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:6fcabbd509e3a5961025bedc86ab8bedd9dd09f36e374c866f0591031f8be61f

Observation c5ff00fa-d536-4160-8ace-9c389acc50dd · outbound

This paper cites Llm-based multi-agent systems for software engineering: Literature review, vision, and the road ahead.ACM Trans.Softw.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Llm-based multi-agent systems for software engineering: Literature review, vision, and the road ahead.ACM Trans.Softw

Reference 20

Resolution
verified exact
doi, observed 2026-05-16T13:51:00.530952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:293dd4278816011ebef4d389345546a61102a318c684460c4497e63cac9203df

Observation 21fa6a76-cd7c-4afa-839e-26bdf5d9b1ca · outbound

This paper cites an unresolved cited work.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:51:00.854817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:136d50292d3c093f2957e4a71d50423cd429e42789a8a2fe494a40a4ec466d35

Observation f340bcf5-1739-47a8-936d-06356552470b · outbound

This paper cites In: Advances in Neural Information Processing Systems (NeurIPS).

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Advances in Neural Information Processing Systems (NeurIPS)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.865280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:d5ef74b9fc89fea24d40c1e15631c04774502a772764f0c81d3237d2d272bbdb

Observation 878fd20e-0dbd-47ff-8fdf-d7f80e3ca6ef · outbound

This paper cites an unresolved cited work.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:51:00.855680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:90b1cadee53c677749063bbeb47303ebd04c0e10cb0681a6cd608b72efda7fe6

Observation 8a98f5e5-2177-4678-80a4-f37001827aec · outbound

This paper cites In: Forty-first International Conference on Machine Learning (2024) 4.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Forty-first International Conference on Machine Learning (2024) 4

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.852689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:79d156919bb55ee57ecc7359d4da644fc34b8efce5d6cfbf8e0765b3cd598b0a

Observation 5c79f49c-924f-4bcc-8aed-a1a2bf98b6b0 · outbound

This paper cites In: European Conference on Computer Vision.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: European Conference on Computer Vision

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.848218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:267e367cef7be234c068c57a22bb4a2073ec5f19a6604430913b21ee623d3caa

Observation e8976720-1c81-4e36-8a37-811d8449837d · outbound

This paper cites In: ACM SIGGRAPH Asia 2023 Conference Papers.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: ACM SIGGRAPH Asia 2023 Conference Papers

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:51:00.543037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:29acab6791530a9fe92d3788d62cc4332bbdfe9c6a17fbc042dea1c68fa46846

Observation 6fb05802-7256-427e-8fda-30bf651d87e6 · outbound

This paper cites In: ICLR 2024 Workshop on Large Language Model (LLM) Agents (2024),https://openreview.net/ forum?id=RPKxrKTJbj4.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: ICLR 2024 Workshop on Large Language Model (LLM) Agents (2024),https://openreview.net/ forum?id=RPKxrKTJbj4

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.863043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:730f8d22e4c5802f605715f0b7b1753ad17426d7ff6ec1bb41d6f9fbe117b00e

Observation 207aa4ea-dbdb-4e23-a81b-c0d571f9d4e3 · outbound

This paper cites Transactions on Machine Learning Research (2024) 5.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Transactions on Machine Learning Research (2024) 5

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.888571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:74f99c29ee233cc70e96aaccfb4dfa573d10e2bf1ce8f05577621d44c67abc74

Observation f399d897-2f0e-42e8-8823-72affdf89085 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.869170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:8f86d979b329c42a2c8b11107c46bf44d7bc0e1171d148521a2c304dcae95ae6

Observation 26543a02-6378-46d7-9a85-33f1673708c1 · outbound

This paper cites Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:51:00.560832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:95097fa2438992846102df3572d63c844df95d54e94caad890f3f2185cdbb2c2

Observation 5c744b66-c031-4f5a-8bb9-f76855bc56a4 · outbound

This paper cites Transactions of the Association for Computational Linguistics12, 157–173 (2024).https://doi.org/10.1162/tacl_a_006388.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Transactions of the Association for Computational Linguistics12, 157–173 (2024).https://doi.org/10.1162/tacl_a_006388

Reference 31

Resolution
verified exact
doi, observed 2026-05-16T13:51:00.528281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:88f39309b92eaaa5917ed6db2d0aa7586e056496de6afe003f43c14576a707c6

Observation 13af9cb6-c9c3-4893-919f-bc8dc4bd4168 · outbound

This paper cites 2019 IEEE/CVF International Con- ference on Computer Vision (ICCV) pp.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning 2019 IEEE/CVF International Con- ference on Computer Vision (ICCV) pp

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.840621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:6110eabb1adb28eaa1efdc821b49f6dc0983fb6eea9011a581d732ff46d6c120

Observation c81bc145-271a-486c-9224-9aec7c6d6165 · outbound

This paper cites In: European Conference on Computer Vision (2014), https : / / api.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: European Conference on Computer Vision (2014), https : / / api

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.856973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:e3b1cc3aac9cca5e13aff4ccc462f454e86c5670cc029c1d35be6bcaa2bd818d

Observation 9ca72ab7-c5ef-402f-be68-5254b961c7b6 · outbound

This paper cites an unresolved cited work.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:51:00.873364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:a90d697ceda26101183c6aa292d4095fa519b63aca38d60481712e1bfe78030c

Observation 8a58ab86-189b-461a-b252-ff97ea8c8174 · outbound

This paper cites In: Proc.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Proc

Reference 35

Resolution
malformed identifier
doi_truncated, observed 2026-05-16T13:51:00.537582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:b53eb86a4a9a43c915181274157c260c1275a59d0ddde387fcf217604c2bc370

Observation 8fe158ad-1e68-4968-ace2-4c000babd0fc · outbound

This paper cites Retrieved from https://cdn.openai.com/gpt-4o-system-card.pdf 22.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Retrieved from https://cdn.openai.com/gpt-4o-system-card.pdf 22

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.842666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:5e732bf322964664d1e47abf7d6f13b5fb59228aaf285887cbec031b9bd8af46

Observation e53af463-6432-487f-8608-e63c38be9800 · outbound

This paper cites an unresolved cited work.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:51:00.835682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:e004e40e134e9a27452f52fc363ed75387946c3fff760db3f6ec9a8de0215e4e

Observation aec74e1c-1346-44ef-8790-e3bd404c777c · outbound

This paper cites 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) pp.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) pp

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.837965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:7495537041dd859c4e4e05e1e3b12062e7e6c99aafa34d6587d861ca19ca9fac

Observation ec5bcde4-39a0-4ec3-8bb4-aa7f4460780f · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T13:51:00.606417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:ff42e91e5cd405f5bfad935499e0936e199ba93cc924928666e0ec1c82ac7a2f

Observation 1c76c414-2e31-4329-b358-5c0f8eb4cb8d · outbound

This paper cites J.: Infinite photorealistic worlds using procedural generation.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning J.: Infinite photorealistic worlds using procedural generation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.845206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:2aff5945f764a894c4e723c9a110b04bbe68f881029ae3f6bace2f7c41071c01

Observation cbeda6c2-20d7-42c4-aa65-6af3af57ef49 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.818648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:6d88ce425b44b538c248ca8054fbb5078c251a696eed50ec7f2925f32ca23067

Observation dcb0fb20-8b0e-4a97-b101-5e514784ca79 · outbound

This paper cites Computer Graphics Forum42(2), 545–568 (2023),https://onlinelibrary.wiley.com/doi/ abs/10.1111/cgf.147754.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Computer Graphics Forum42(2), 545–568 (2023),https://onlinelibrary.wiley.com/doi/ abs/10.1111/cgf.147754

Reference 42

Resolution
verified exact
doi, observed 2026-05-16T13:51:00.546030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:ece7579f65ca229a8f193ace16a9b0d68b934af57bc5e8ece7822141e04800f4

Observation fed75475-b447-4ee4-9a42-05b0c2057468 · outbound

This paper cites an unresolved cited work.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:51:00.884547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:c1ab3ca7b152edd7b7a8a39873a5c06871717940aa2d814de13e569a941460e9

Observation 9f6a2258-fcd0-4f7b-9d1e-b470ce550375 · outbound

This paper cites Modular Visual Question Answering via Code Generation.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Modular Visual Question Answering via Code Generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:51:00.599385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:5b52fa99e097873d960f5dcc1560d3e678fe36ab56adcc7838631df870ffc04d

Observation 3c184239-d8fc-4f6c-96a5-5f73dcc8795d · outbound

This paper cites In: 2025 International Conference on 3D Vision (3DV).

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: 2025 International Conference on 3D Vision (3DV)

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.820589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:8f42efb6b073d9913c2388a88fd1a41fefb4b06d93f0e94606dc40847d769e08

Observation 4ba3e97b-74b5-4a40-a4b8-793b50e84fa1 · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.875369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:61c0e12c9e6e019a672565ea55a8ca3281f130ade35015384f9efc8b8f893f6f

Observation f4882744-93ec-4c58-b0f6-95ce98fe92b5 · outbound

This paper cites 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:51:00.591108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:d9b8beceb0aa91b9f3ff13eeae901007716be5e18d5b5f7faab88be013f34b93

Observation a5d730dc-0dda-4224-9629-1aac5f84fbeb · outbound

This paper cites In: Proceedings of the IEEE/CVF international conference on computer vision.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Proceedings of the IEEE/CVF international conference on computer vision

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.828073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:292fc3458fd60728073351bfa7171edbf7d8ef93677aa2f88e61ddd93afa3ec5

Observation e940b29a-a7b7-49af-9b20-ca6569117ba6 · outbound

This paper cites Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:51:00.612599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:406becbee32d4a82c0ab91710acf1bb82d49f546f43aa04aa9cccd88c2aaf9c6

Observation 2bd1b1b6-6899-4aa2-b738-1a9fd0ffbb9d · outbound

This paper cites SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from Design

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:51:00.573620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:c51d9fca313ca69655f0a3fdf2d30fc1a67ae262cd996645a68aadf5dfba15db

Observation b39c00d0-844e-43db-ba82-e0a3289bfaf6 · outbound

This paper cites arXiv preprint arXiv:2510.03463 , year=.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning arXiv preprint arXiv:2510.03463 , year=

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:51:00.602888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:41a6a3ee41583830e3ddead9a8f64eff2cfc6649e1e9cfa1c4a360fb4fb13f5b

Observation 6e5053e8-0703-41f2-8bac-514772dd65a9 · outbound

This paper cites SAM 3D: 3Dfy Anything in Images.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning SAM 3D: 3Dfy Anything in Images

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T13:51:00.584496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:1073940c123eb4208e9b525a6c18b584a9a383b056531ff18c4f40b554f66b87

Observation e64402fb-efde-4794-b56c-3899fddbd8a4 · outbound

This paper cites an unresolved cited work.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:51:00.829895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:124869b0eec0ab4af97bc9c5b0f1775068283f6e830aa294f50baa4eb0e2f67b

Observation acc1807f-a41c-45d3-ad2b-5c4f825ee005 · outbound

This paper cites In: Proceedings of the 33rd ACM International Conference on Multimedia.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Proceedings of the 33rd ACM International Conference on Multimedia

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:51:00.534303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:3a33e8b447959f781342722bb8ad7e7b536c027f93fc2e9f01c613954d93be50

Observation a67acf95-4f45-4594-b75c-bc98d67f28bf · outbound

This paper cites In: Proceedings of the 38th International Conference on Neural Information Processing Systems.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Proceedings of the 38th International Conference on Neural Information Processing Systems

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.877766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:d06bb4f7acf73ac1a5f4ec826a5f92bde33a0a3a0aea2b2dd833197ad535a12b

Observation a8e9b2eb-b72c-440d-8f6c-3e59bb729acb · outbound

This paper cites In: Forty-second International Conference on Machine Learning (2025),https: //openreview.net/forum?id=NTAhi2JEEE4.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Forty-second International Conference on Machine Learning (2025),https: //openreview.net/forum?id=NTAhi2JEEE4

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.879889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:33121ef7ccad1464e1715e765096abc0b9d8b3695b695f0bfdf5dfce33060c5d

Observation e7e59146-2991-4fad-8d4c-aab0a9d2c9ae · outbound

This paper cites In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2017) 2, 5.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2017) 2, 5

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.825895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:3089f3af3c37c7588f867fdf7adc20cff8015c109517332b60e58f4b5863519c

Observation 878097ad-f5d3-440c-8bd3-27ab10c34382 · outbound

This paper cites 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) pp.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) pp

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.882064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:bd0a1ef69d16e799cf521ef3d6aae168ec8ccf683a2cb8561d1651b23d62c987

Observation 7a0011f9-d010-4d99-9c6a-1fb655c36175 · outbound

This paper cites In: The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track (2024), https://openreview.net/forum?id=tN61DTr4Ed4.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track (2024), https://openreview.net/forum?id=tN61DTr4Ed4

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.890671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:8a0e311be711b33736d4056c7cd75de7739a671bb252e4d1f3aac412e85bda0a

Observation 1d861f24-5261-44cb-b1b8-8d61f35c727c · outbound

This paper cites Empowering LLMs to Understand and Generate Complex Vector Graphics.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Empowering LLMs to Understand and Generate Complex Vector Graphics

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:51:00.595047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:432f9a1552643e31dbb7792916fdc5d80dfad2282cde6991cb810cf214822326

Observation 9aaaaaba-843d-4444-8f2f-a042b049bfc3 · outbound

This paper cites In: Advances in Neural Information Processing Systems (2025) 4.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Advances in Neural Information Processing Systems (2025) 4

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.833127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:b7fd6dcbfd34aad076f2cd5d34ce2d3b02a38f31db1fead4ef38f819b848a200

Observation b8621659-e1ca-4a32-986e-7a374ee44afa · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.858950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:fc7bf316b3355b8e1a8c173a8dad8399fe9889a43a13dc00ac2420b860bd288a

Observation 90fd00fc-410b-4101-b98d-442049d7b322 · outbound

This paper cites Advances in Neural Information Processing Systems37, 50528– 50652 (2024) 4 20 Shaofeng Yin et al.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Advances in Neural Information Processing Systems37, 50528– 50652 (2024) 4 20 Shaofeng Yin et al

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.894993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:3b7bd6b7c0b3b9f1cf8fa2aca27472d967e85a5ad8ae23dedfca6c354978d41a

Observation 1d700583-cc2f-4797-933a-23f6f382254e · outbound

This paper cites Omnisvg: A unified scalable vector graphics generation model.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Omnisvg: A unified scalable vector graphics generation model

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:51:00.565341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:473218dde4057d26fcdf50a5c2151a1f63ad330df13179f4b0523ebdf5c1fca3

Observation d4262a74-aeb6-45c6-bebf-bce1881ad9db · outbound

This paper cites Advances in Neural Information Processing Systems35, 20744–20757 (2022) 4.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Advances in Neural Information Processing Systems35, 20744–20757 (2022) 4

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.871326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:9ada6ef9d713f082af0dae754b799003560762081f9dbd31812e514f5b535f0c

Observation bbfacc65-4d64-4c1d-bc03-cb59c9df8d4f · outbound

This paper cites Task Memory Engine: Spatial Memory for Robust Multi-Step LLM Agents.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Task Memory Engine: Spatial Memory for Robust Multi-Step LLM Agents

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:51:00.587813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:78cce2400d3c3bf91b9e1fd00d8eec64deb1dd128ba2a5eab00a4eb6fecbdd85

Observation 57730518-1e78-4bd4-b881-0a2144cf83f0 · outbound

This paper cites In: Advances in Neural Information Processing Systems (NeurIPS) 31.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning In: Advances in Neural Information Processing Systems (NeurIPS) 31

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:51:00.816324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:e9f4add09ffae40fb1708e90b7ddf1e3a6e7f68624e84f8baf5e1beecea458af

Observation 1ee5d6c3-6c23-4704-b481-cccf5bd1ac9c · outbound

This paper cites Building Cooperative Embodied Agents Modularly with Large Language Models.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Building Cooperative Embodied Agents Modularly with Large Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:51:00.568852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:89c0a30ffa912e8b9931a6daf9c9551fbeec08157be1f2811962429b2189dd76

Observation f232206d-afa5-4bd5-bc4e-c3d84b801b4c · outbound

This paper cites Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T13:51:00.581200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:6f1f9399158fd9a9575d108c846528f3fb53947fff68b6b94cc4a58992210b69

Observation 7038641c-9f60-4ded-b1f1-ca46e5bdd22c · outbound

This paper cites PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:51:00.609420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:fe10e3852f819b42a5815958bd89790d013364afea41c25024cd6529d2c9e90e

Observation 5886d0d9-e7ed-498d-aaee-4502c81eba68 · outbound

This paper cites an unresolved cited work.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:51:00.892950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:3d081d9496c2aab2230211e9d035e25ae308d536422a62aefdca24c6fb31806b

Observation 2d4cad1f-a241-43ed-a3a2-e9299db4027c · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 72

Resolution
malformed identifier
local_arxiv, observed 2026-05-16T13:51:00.577861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:2c572bcef251f64fdf59aa59fe8fba86064535de4aa1697efead502d55729231

Pith citing papers

Observation 12434dbf-3219-4d9f-b3e0-4d18c75df05f · inbound

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents cites this paper.

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T07:51:01.973346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:55:48.506049Z digest=sha256:46fe84cb79cfee73e72df65737649b2eb90dd024dd1ab13c019cafcb10367285

Observation ad32293e-0ed6-4ddb-b58e-7c4ef210f2e1 · inbound

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents cites this paper.

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T23:07:40.699400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:07:40.699400Z digest=sha256:94e728bddfe472cbecf36439a41fc3b3b1fe80ef0d36b572c4c9e8d329eaaaa0

Observation 73ccf823-7d85-42f2-8bec-3f0b0dffde40 · inbound

SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation cites this paper.

SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T12:31:06.664620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T03:22:30.012610Z digest=sha256:86dd63d9f8837d65bcc11b634665b97bb6b214b5e3ab849220b91a5c10a19f14

Observation 05bf3e08-cbd7-44fa-8a2c-7a5eeeb91c75 · inbound

LychSim: A Controllable and Interactive Simulation Framework for Vision Research cites this paper.

LychSim: A Controllable and Interactive Simulation Framework for Vision Research Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:27:24.568631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T06:26:14.825783Z digest=sha256:2ff13864c9cc27ebd317d91afa7bdb4db85dad61dc9edd34c5ba716bf12f4758

Observation 233093c2-e75f-402c-bc90-12aede6a698b · inbound

Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis cites this paper.

Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:53:15.009674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T11:49:53.113415Z digest=sha256:b478aecee3536272005db0ef33c648276b25f26d965ade214887e2bcae6af277

Observation ec967844-92aa-481f-86ff-3215e6faa23b · inbound

REST3D: Reconstructing Physically Stable 3D Scenes from a Single Image cites this paper.

REST3D: Reconstructing Physically Stable 3D Scenes from a Single Image Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Reference 44

Resolution
malformed identifier
local_arxiv, observed 2026-06-29T07:53:14.077572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T07:44:41.469789Z digest=sha256:a066335a0fe678478eeaeaeda93fb5bdbf546558c2f73b98d5ae06b8d165f111

Observation 7d1a75a9-4996-482e-bee2-7cdeaae4ebeb · inbound

Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models cites this paper.

Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:46:19.340650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T15:04:01.773923Z digest=sha256:c690d999c3b527a116ec4fb50f321a5dc3760c28491b29350fd725baf76a04b0

Observation 3e449962-bc7c-47f3-9cf4-681470997143 · inbound

SceneConductor: 3D Scene Generation from a Single Image with Multi-Agent Orchestration cites this paper.

SceneConductor: 3D Scene Generation from a Single Image with Multi-Agent Orchestration Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:07:26.681136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T19:08:08.329817Z digest=sha256:93853c579cd199a257b0b91ca614ca69d93a712edc148bb2c745447dcf13a031

Observation bf1e5520-b144-4cee-a9a0-b5aad6b724ac · inbound

P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning cites this paper.

P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:37:37.018009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T13:48:07.845039Z digest=sha256:f8210ab0951b9df86b814ab1f32e7cc99df94b14dc34f0698396bb24c0e32a11

Observation b992f767-7967-42cf-ad78-a615661714b5 · inbound

One Video, One World: Turning Monocular Video into Physical 4D Scenes cites this paper.

One Video, One World: Turning Monocular Video into Physical 4D Scenes Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Reference 110

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T10:15:44.975015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-01T05:38:19.541391Z digest=sha256:7c3838ab2f52bbdde6680677c528086a049b0dedd684c618b005a354960dde9a

Observation e891bd6f-f983-46fb-9bdb-574c2688185d · inbound

SimWorlds: A Multi-Agent System for Dynamic 3D Scene Creation cites this paper.

SimWorlds: A Multi-Agent System for Dynamic 3D Scene Creation Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:08:21.570989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-03T13:58:48.681254Z digest=sha256:3cef06018a2390121a186241317879fe1554a59f829d7aec6f6b17f9edb48bbe

Observation d4df9963-96db-47a8-9f5c-bbf1c540e153 · inbound

RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning cites this paper.

RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T07:27:08.281893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:27:08.281893Z digest=sha256:637ec13b24c7a7cfede368e76a263f6edaad9f598669977a8530b0184dcd0b2f

Observation e9f83500-712b-42d9-94c2-2f4dd2aaf774 · inbound

PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation cites this paper.

PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T22:27:52.771740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:27:52.771740Z digest=sha256:2c37c18040b4d5653740cb31ebba97bf380fc903ba7a875b86f08c16cd28c3dc

Observation 933db3c9-f77e-4290-99c6-26cf3c1ba743 · inbound

Engine-Native Editable 3D World Reconstruction with Objects and Lighting cites this paper.

Engine-Native Editable 3D World Reconstruction with Objects and Lighting Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T09:10:35.430779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:10:35.430779Z digest=sha256:e74725c08e15d4f1d435b67efcb46fbe9e21165388e8ce6a2b2a61d408b97ad2

Observation 9d8488da-0138-4911-b589-d0f22ddd9d42 · inbound

GS-Agent: Creating 4D Physical Worlds With Generative Simulation cites this paper.

GS-Agent: Creating 4D Physical Worlds With Generative Simulation Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T07:14:42.538868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:14:42.538868Z digest=sha256:45bc8791f63a6d6aa8b8918210709464d6e235fd4e0d6ebf7e85ac39e73eaee8

Observation 3b8bc932-0f4f-4ce2-b3f4-bd13d13cddaf · inbound

WorldClaw: Agentic 3D Open-World Generation at Scale cites this paper.

WorldClaw: Agentic 3D Open-World Generation at Scale Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-08T17:13:38.497502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:13:38.497502Z digest=sha256:6cd769f587a7e904f3f16449ee37d13ab710faaac8749ac21970da7a576c3fe6