Pith. sign in

Paper Citation Record · LEDGER

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

As of 7 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 4 inbound Pith citation observations for arXiv:2506.23329.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23329 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:52:19.061479Z

measured 93 of 93 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:42.147646Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:46:19.349816Z

Reference resolution

89 of 89 outbound references displayed

  • verified exact0
  • verified fuzzy40
  • unresolved48
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 70399b4d-b6e2-4b73-94bf-bad20488e404 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering LLaVA-OneVision: Easy Visual Task Transfer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.055004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.055004Z digest=sha256:74961cd7022ec748ad5ebb5abecbf16c5c404571dce55ca886d009b4ac1feaa4

Observation e399d4ea-d573-4ec4-8742-5321185e750e · outbound

This paper cites GPT-4o System Card.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering GPT-4o System Card

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.097608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.097608Z digest=sha256:e7561383c791c23d015064018337d3622738e3899a400da089a8c1899cd4be7a

Observation c1bf531f-f280-4fa3-83ef-1b4ba3d07b95 · outbound

This paper cites The Claude 3 Model Family: Opus, Sonnet, Haiku.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering The Claude 3 Model Family: Opus, Sonnet, Haiku

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.169741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.169741Z digest=sha256:7dfc25c6a1f168d67cdcde51b159e9b8878faa2b6c39a8c508a2e0f08498c63c

Observation c5f399e7-11ab-498b-a352-6b5424b05bb8 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Gemini: A Family of Highly Capable Multimodal Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.224504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.224504Z digest=sha256:9e994e650b22591a13e8f24d85c4c7e0d71357739988342e4fc310298208aebf

Observation b4095e18-2038-4477-a16d-c90e47d0c61d · outbound

This paper cites Qwen2.5-VL Technical Report.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.285409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.285409Z digest=sha256:9adc03b9340cb56de683e8e6c12f4f31576bcb0afc593f60667c35ff634938d8

Observation d393dec2-7886-404f-bd47-c76fdd7b1f14 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.354626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.354626Z digest=sha256:c93b0b7be7c1a70fe05110d65e4155e6527885116776f6298bdc4beeb186a309

Observation 623e3470-0483-491f-8878-f676cb8fedf9 · outbound

This paper cites Perceptions as hypotheses.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Perceptions as hypotheses

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.396632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.396632Z digest=sha256:8955c90d4c47210518fa47991f1817bc51cf7a8eb0907926432070ecb9188502

Observation b67c78c0-8f4d-430b-9a8a-c6a05d6a7535 · outbound

This paper cites Yuille and Daniel Kersten.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Yuille and Daniel Kersten

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.441337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.441337Z digest=sha256:d8bced7ef0ad5f6f5bb6fd39d521230e97534053d91f0b4c6a205e2f95e24cb1

Observation 4ba89f9d-b7fd-48cc-94e5-6035ea2f9f20 · outbound

This paper cites Efficient and robust analysis-by-synthesis in vision: A computational framework, behavioral tests, and modeling neuronal representations.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Efficient and robust analysis-by-synthesis in vision: A computational framework, behavioral tests, and modeling neuronal representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.504617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.504617Z digest=sha256:d7e08dd362dc16aa1373e76d42c4b1e2f8131552ee1183a807bf3c39449c7a2e

Observation 04895bc9-16e4-42ab-b6cb-fc311ebde2ea · outbound

This paper cites Bever and David Poeppel.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Bever and David Poeppel

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.569097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.569097Z digest=sha256:eb3740a547f1bc5d3ae4526f7ec691e2010b5b46df4cb89dc281157e2ec28717

Observation 2bd2439d-548f-4935-b8eb-ebc1323b56a8 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Toolformer: Language models can teach themselves to use tools

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.626028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.626028Z digest=sha256:84af416fc4e5740d8457a2dedce6b33ea6c066a951ae40ab1148b3569686a046

Observation 41d94738-46fb-42ca-aa07-daad572b27a0 · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Visual programming: Compositional visual reasoning without training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.687738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.687738Z digest=sha256:5779a57e44073ff91fc78809837dd426a7501acfd8bd01fb59524a000d353d89

Observation 2371c8ed-4db0-4785-b740-808dc0eddae7 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Vipergpt: Visual inference via python execution for reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:11.750586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:11.750586Z digest=sha256:6f1677d73ba6c940a97bd329347e8f436ac8f460bedc68cf94ceac8ae36f5994

Observation 759e214b-494b-47d0-a7ff-29d41b221fb4 · outbound

This paper cites Gorilla: Large language model connected with massive apis.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Gorilla: Large language model connected with massive apis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:27.564130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:11.803713Z digest=sha256:835e81fb2a09a9703bd9c2d140e806c7878685b7ce797bd93c83956da4ac1089

Observation 58910aac-d9e8-4948-9f22-a5cf9a3c51dd · outbound

This paper cites Blender - a 3D modelling and rendering package, 2016.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Blender - a 3D modelling and rendering package, 2016

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:27.281964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:11.874949Z digest=sha256:f8c9e1ce4421a785e11a01c20fe60f0de7426190a3fcdd2c8b3622e73c145a8f

Observation 4f94bb30-7534-47ce-b4a6-8f2c24fe64e2 · outbound

This paper cites https://github.com/modelcontextprotocol, 2024.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering https://github.com/modelcontextprotocol, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:27.013539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:11.913246Z digest=sha256:eeeb9f11405d2423df38f3acec07b7322c4c0ead5ddb803f9b80456416669eec

Observation ecc096df-bc12-4ee8-b717-625b2d4f4528 · outbound

This paper cites Blendermcp.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Blendermcp

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:26.613500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:11.977613Z digest=sha256:9a68a7b0356e91964d4c331abac1166311b3fe5c72664ca0f3e526b1dd6bff0e

Observation c343af7a-bbde-476a-9168-91e73e1b1192 · outbound

This paper cites Embodied question answering.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Embodied question answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:26.204239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:12.022798Z digest=sha256:0b284bd7afaf468190c1c1bc72c8b6f195834a6de75c843702e249434bd2134f

Observation 24c50811-3954-4c4c-8c9a-3443545a9ca4 · outbound

This paper cites Multi-target embodied question answering.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Multi-target embodied question answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:12.083938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:12.083938Z digest=sha256:3576029d28361c85a96460a88a9361c703f7fd691274d55143da48a15293e9f9

Observation 5b4e46d7-ef6f-4292-8bc7-a0adc8407f6b · outbound

This paper cites 3d concept learning and reasoning from multi-view images.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3d concept learning and reasoning from multi-view images

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:25.777892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:12.138262Z digest=sha256:70db58c5b2b2db0a3fe1b4d514b4aac2cead1dc33ef5646b4ccaf9d60a9e1b82

Observation 9a512965-3142-40a5-a3af-38bb88cef8f6 · outbound

This paper cites One step at a time: Long-horizon vision-and-language navigation with milestones.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering One step at a time: Long-horizon vision-and-language navigation with milestones

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:25.671215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:12.194541Z digest=sha256:f50536d2a685ffc306b5004dc64bc229409d975b53f319d5b6204bfc5b5e5582

Observation c6d80a6f-3acd-410b-b3bb-eb48b0cd720d · outbound

This paper cites Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:12.241831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:12.241831Z digest=sha256:195fb176e97a0d9ff1cf4190d3637673529ed4b22df32fc54b9513f1868e5eeb

Observation de9ccc8a-466a-4a37-afa9-fe5ffb61f905 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:12.279337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:12.279337Z digest=sha256:6bd3d0a898153c904be840290317c52d3b412f015f0ad3c7846397d3db073614

Observation e9a6c378-4606-4592-94b9-abeafdd0c1ae · outbound

This paper cites Barrow and Jay M.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Barrow and Jay M

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:25.549695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:12.293443Z digest=sha256:104ec883d826072cf4944f7a411f2d1fbf213559b53cc35991d944c55d09e3fb

Observation 05dad22f-a4c9-4d5e-b07d-273a93a6fc3a · outbound

This paper cites Soft rasterizer: A differentiable renderer for image-based 3d reasoning.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Soft rasterizer: A differentiable renderer for image-based 3d reasoning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:25.418707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:12.327295Z digest=sha256:bc889dcda69ad9b94e9fe58679ad177ab21d6ab4c13cd93ce9b7eebaf45dcf29

Observation 47e3950a-d8e2-4c22-ac40-8243dea9bbdb · outbound

This paper cites Carr, Jonathan Ragan-Kelley, and Frédo Durand.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Carr, Jonathan Ragan-Kelley, and Frédo Durand

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:25.281806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:12.364961Z digest=sha256:df2b3c91d5b0b6534f6bc6f81548f9354556789f339f5543afd50a0f91c27d47

Observation 03cb4dd2-5833-44e6-a443-1ded4a1cbcc2 · outbound

This paper cites Differentiable vector graphics rasterization for editing and learning.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Differentiable vector graphics rasterization for editing and learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:25.156345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:12.397916Z digest=sha256:8612ebdf46576bb23f95074c8c94b09e1031cef103ec5dc5f072e03621756ac0

Observation fdf89410-523c-4f03-82e4-6e26fc4b84b2 · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view synthesis.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Nerf: Representing scenes as neural radiance fields for view synthesis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:25.044121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:12.474675Z digest=sha256:e48cfc59ef9ef343f187d581c23d86caa8855d7e5b25b3d4f82c00cd7ab45c36

Observation 185ef9c7-49ac-46b2-984d-870a0eff56e2 · outbound

This paper cites Unisurf: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Unisurf: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:24.925554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:12.567066Z digest=sha256:da246c563f34e067e95419469b24a59427d3ce426d6b54e4cef87e8b90d8c8d9

Observation 170cb4bb-be8d-4702-b529-96c80dc09769 · outbound

This paper cites V olume rendering of neural implicit surfaces.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering V olume rendering of neural implicit surfaces

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:24.826114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:12.634446Z digest=sha256:2275d20c9e433d4cff650f6dc3386180e831e28692a41ec101a9c27575a1ffdb

Observation fe13a610-3f42-47f0-be62-c42423e06960 · outbound

This paper cites Synsin: End-to-end view synthesis from a single image.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Synsin: End-to-end view synthesis from a single image

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:24.683222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:12.704011Z digest=sha256:18067fd3b18f85796f57e6303f632d23d672a0397b103af96ff7fe10f6df600d

Observation a40ed8df-92b3-4d8e-9c4f-eb74c9b4898b · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3d gaussian splatting for real-time radiance field rendering

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:24.597407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:12.762820Z digest=sha256:bc5bc45ddbf71e5f2a69e42efbcda3bc4b655144ce0d88c61224a6fa2cbc6957

Observation 120bdca9-0787-411c-af13-41bf6860ac8f · outbound

This paper cites Path-space differentiable rendering.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Path-space differentiable rendering

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:24.443779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:12.822049Z digest=sha256:e5a95847598a286a7123f25c4de36dec472ecd79fed3df2911cf3e628b178ce9

Observation 4a82b5d2-f0fa-4dcd-90ad-0e7dafa35cb0 · outbound

This paper cites Inst3d-lmm: Instance-aware 3d scene understanding with multi-modal instruction tuning, 2025.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Inst3d-lmm: Instance-aware 3d scene understanding with multi-modal instruction tuning, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:24.322002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:12.878226Z digest=sha256:580d95867c6584c95bcb1689ff04e93014e0754b4c119129b23b54c7660a6b11

Observation baa76daf-d0cc-4046-867a-c8c4e18cb39f · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:12.924934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:12.924934Z digest=sha256:6c5cee7197aad72c985c32d03ae61f89cb33ebc04d124cb2d0cac6a9b2107745

Observation af2108fa-3480-4223-889d-1ddbf54662d7 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Scanqa: 3d question answering for spatial scene understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:24.123602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:12.978758Z digest=sha256:ccb4b404b181abc0f420cceb3c7d9c9f0960cfc95b40c36e6fa791fdd8f7fb6a

Observation 0fdde439-780f-4367-abce-abcb0b2ba4ac · outbound

This paper cites 3d-vista: Pre- trained transformer for 3d vision and text alignment.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3d-vista: Pre- trained transformer for 3d vision and text alignment

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:23.895150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:13.066768Z digest=sha256:f1dc7c071503a773cb872468836b6c79a8f692e6bf0b8781113652c15b7fa29c

Observation 804d0774-75fd-4430-b308-0a344c7c0b98 · outbound

This paper cites ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:13.125119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:13.125119Z digest=sha256:6402ac996438b6df7423e5e8d885c04051bb5afe0421bfc03739797ea820e2b5

Observation 069e343b-baa2-47c1-ab21-fa116d0be223 · outbound

This paper cites an unresolved cited work.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:23.716990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:13.176031Z digest=sha256:a2c1924d45968e9569d6d030a404dba6fca2e04aea6d1ae080f37ff501be4f83

Observation 8a82e53a-4cde-4c52-8b39-2781b3ce7187 · outbound

This paper cites an unresolved cited work.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:23.514382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:13.257778Z digest=sha256:57fb5931e6f42368cd26ca21dfd67aea7d8f5be399229f49e14132261c5842f1

Observation 91a71c99-fc19-4bfb-9743-85f859313aca · outbound

This paper cites Scalable 3d captioning with pretrained models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Scalable 3d captioning with pretrained models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:13.307929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:13.307929Z digest=sha256:7326ffbcb0d3e2f85b043d06d760f3454a994c43693513b10fbc657f1f84e87d

Observation f8e232af-a3db-4d52-a83a-5d822a3dd8a3 · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Openeqa: Embodied question answering in the era of foundation models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:23.340070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:13.367178Z digest=sha256:a526f3e3da2069a0e081b4362c507b5db1430ab537d21a6838820034a7ebd4e6

Observation 5364b1bc-adb6-4a12-a480-5b8ade7cb1cb · outbound

This paper cites Alfred: A benchmark for interpreting grounded instructions for everyday tasks.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Alfred: A benchmark for interpreting grounded instructions for everyday tasks

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:23.158643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:13.427092Z digest=sha256:f35affb640250348af75309c5cc094d574e5b2646757f9e883172008d3ac2d9f

Observation 79c150fb-dcb2-4dda-9b34-946805354b5f · outbound

This paper cites Sceneverse: Scaling 3d vision-language learning for grounded scene understanding.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Sceneverse: Scaling 3d vision-language learning for grounded scene understanding

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:22.964985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:13.473093Z digest=sha256:1ed23df744d99c76d9fb893e84b9f20f51bcbfd7b2ae4a36c9179997e1a140ec

Observation 2243b079-f024-4250-87c2-d590e88832e9 · outbound

This paper cites Spatialrgpt: Grounded spatial reasoning in vision-language models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Spatialrgpt: Grounded spatial reasoning in vision-language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:22.738013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:13.520082Z digest=sha256:cdefd138b04dc5f93048ac169f7aca43abdaf7bcc460b4202303fb6c5204d50d

Observation 13c423ae-e694-4b4d-9483-84ce12cea5bc · outbound

This paper cites ConceptFusion: Open-set Multimodal 3D Mapping.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering ConceptFusion: Open-set Multimodal 3D Mapping

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:13.629446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:13.629446Z digest=sha256:dfa12640f5e7a70ad980c7c6f6f998291044d606b8e8067e3092549d855d5bdf

Observation 0854f686-091f-4bf1-9322-c38faaea3951 · outbound

This paper cites Context-aware entity grounding with open-vocabulary 3d scene graphs.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Context-aware entity grounding with open-vocabulary 3d scene graphs

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:22.553902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:13.689722Z digest=sha256:8fccc307f341e6e2611c1567c5d382227153706116bd4cf9f722ecf00ee239e0

Observation 09505462-0c34-481b-9759-994420a4609d · outbound

This paper cites Tenenbaum, Antonio Torralba, Florian Shkurti, and Liam Paull.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Tenenbaum, Antonio Torralba, Florian Shkurti, and Liam Paull

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:22.310278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:13.772311Z digest=sha256:0467dc25a642ae2b9826b5155712a3e5eb8441e3503864e9d795ae70595d4193

Observation 5d0d8d06-ed63-48ce-ab2b-6ad2bebe8242 · outbound

This paper cites 3D-LLM: Injecting the 3D World into Large Language Models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3D-LLM: Injecting the 3D World into Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:13.851642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:13.851642Z digest=sha256:fe0f02cc42b1ea3bd6634da187506a8143f7d0c0b517c9167b681b96681a167b

Observation af0980fc-8aae-4b20-8ca6-af1ebf75083a · outbound

This paper cites Agent3D-Zero: An Agent for Zero-shot 3D Understanding.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Agent3D-Zero: An Agent for Zero-shot 3D Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:13.961478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:13.961478Z digest=sha256:c73d37f03c778b8a2d15b154932db8f24bf78bbf713d86d1dd8e8cc0d5c0858b

Observation 477b6121-dda0-4230-a3af-7ebb78398485 · outbound

This paper cites 3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:14.100701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:14.100701Z digest=sha256:57c55d89e82b581ecd2250c619faad7b5de3bd6db6d97c8da01dc081d0b731cf

Observation aa631a06-2c28-496b-934d-0e48b671585c · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Openeqa: Embodied question answering in the era of foundation models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:14.285145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:14.285145Z digest=sha256:9f25bba0d7546b0f2493025ab17e27638ef3f00903c671c97645cb511d2752cf

Observation bf2e2b69-3430-422d-9178-84203d7d112f · outbound

This paper cites Kulkarni, Pushmeet Kohli, Joshua B.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Kulkarni, Pushmeet Kohli, Joshua B

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:22.056195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:15.001949Z digest=sha256:3e30d0a3010133146793a1973ef493f27f6b74b184dd0c434f94446803bbc4ff

Observation 92033c46-13f7-4403-9222-3acc43d758a9 · outbound

This paper cites Learning to infer graphics programs from hand-drawn images.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Learning to infer graphics programs from hand-drawn images

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:21.901078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:15.895509Z digest=sha256:b8d045fe6fc0d0975771e5511cbd5b36035198659d3c45d82d41397c5db86a56

Observation ab25332c-a816-4f5b-8b68-f98bd8ce40d5 · outbound

This paper cites Learning to infer and execute 3d shape programs.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Learning to infer and execute 3d shape programs

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:21.706626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:16.005434Z digest=sha256:be779d72d1977f9a1d677345247afe648ec87ff21b245325e66eea449ef0666b

Observation 370dc2bc-3632-4845-839c-20e7bf03bfeb · outbound

This paper cites Shapeassembly: Learning to generate programs for 3d shape structure synthesis.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Shapeassembly: Learning to generate programs for 3d shape structure synthesis

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:21.534646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:16.170005Z digest=sha256:87d693ce6fa1044eef22b1cab69a5b95b83a1434a3bb27ea0b7c76e0e457887d

Observation 616311f7-9e01-49cd-bbf1-2bee4be39133 · outbound

This paper cites Scenecraft: An llm agent for synthesizing 3d scenes as blender code.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Scenecraft: An llm agent for synthesizing 3d scenes as blender code

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:21.372486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:16.338669Z digest=sha256:c1eb769e3737e9620a8ed826290dd9e79b6f76d759b55a530f6593375a786a5e

Observation a1d01f54-7b0d-4d4d-ac16-f51c5f1babf8 · outbound

This paper cites The Scene Language: Representing Scenes with Programs, Words, and Embeddings.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering The Scene Language: Representing Scenes with Programs, Words, and Embeddings

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:16.453656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:16.453656Z digest=sha256:93b26de386d611babf0847663cc3f83a6db43027d6ffbc4dc1b81eab88fb79f3

Observation a6f75e8e-096c-4c4f-82dd-00bfbd5e88a2 · outbound

This paper cites 3D-GPT: Procedural 3D Modeling with Large Language Models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3D-GPT: Procedural 3D Modeling with Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:16.602426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:16.602426Z digest=sha256:98d1c05f827e4734287a3b76068feb8e83b5aa471af53ead9eec1fa0d44b7da1

Observation c0e1084c-6cc2-4184-af33-3f98af61bc67 · outbound

This paper cites SceneMotifCoder: Example-driven Visual Program Learning for Generating 3D Object Arrangements.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering SceneMotifCoder: Example-driven Visual Program Learning for Generating 3D Object Arrangements

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:16.700146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:16.700146Z digest=sha256:b6f7a86e1e59d17c19079e2a837fd89eac1bb7223a2b40d11b6f52c2b74e0b2f

Observation 2abac4fd-99b5-4555-9d70-d6c9bc572221 · outbound

This paper cites L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:16.816617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:16.816617Z digest=sha256:c0b78d12d54181b4ec8e3674ed2112b567f311d960131bb82d4c5707ab590d96

Observation d6757f77-95e2-4fef-8b62-63e1a9ee112f · outbound

This paper cites Creative agents: Empowering agents with imagination for creative tasks.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Creative agents: Empowering agents with imagination for creative tasks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:16.883085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:16.883085Z digest=sha256:57cd081364dabc5e2184acd851cc3a2798421dc927b06b6de7a23d1748fc79fc

Observation 79bd3400-c7c6-46cb-b80a-a8beebee358c · outbound

This paper cites Scenex: Procedural controllable large-scale scene generation via large-language models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Scenex: Procedural controllable large-scale scene generation via large-language models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:21.193430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:16.958079Z digest=sha256:e1fcca1911a4b467ebde648a0c643f1d23f61e9fa0a1e8a079106298e0ea1368

Observation c70a0e45-e2d4-4063-9d60-a6cd36e02871 · outbound

This paper cites Program-guided image manipulators.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Program-guided image manipulators

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:20.997606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:17.034681Z digest=sha256:0fd9ee37b161f2a1b7321e393f6279e1a663b18d60e538649aeae1295ce87f9a

Observation 56d960d6-db08-4b58-8524-19e5526477d4 · outbound

This paper cites LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:17.103306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:17.103306Z digest=sha256:f54fe1499c87ef3b07f5d3b0cbce82b7048490bd9799c4a11ec9531c9cf79a44

Observation ef41b7f0-8b70-486f-bd19-089ce59ad090 · outbound

This paper cites Virtualhome: Simulating household activities via programs.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Virtualhome: Simulating household activities via programs

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:20.787871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:17.197318Z digest=sha256:18fe50a36fa65a1d44bbb1f4a41edc7cfe54b5eabf36f4eb9afaf36b9716b8c4

Observation bfd141a2-30a4-4447-96ec-8d97ac1962a1 · outbound

This paper cites Xia, Peng Xu, Karol Hausman, Brian Ichter, Peter R.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Xia, Peng Xu, Karol Hausman, Brian Ichter, Peter R

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:20.637410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:17.287964Z digest=sha256:c2c23afe30a12a025e6b39405881f625504a603ef88637ae9090bc1230ae9fab

Observation a65d344b-a34a-4ad8-8d9e-6228213936f2 · outbound

This paper cites MONet: Unsupervised Scene Decomposition and Representation.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering MONet: Unsupervised Scene Decomposition and Representation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:17.363407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:17.363407Z digest=sha256:7d2ad96a904eb261c90ea3a2ab01ad7fb50532cae877b9c5642a6942e2cf261c

Observation 6d8a2049-58fe-4748-bb8a-c73c40b2b546 · outbound

This paper cites GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent Representations.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent Representations

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:17.455262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:17.455262Z digest=sha256:8312c922fec79af1d728ed74c3135c92519320cdd01bf10d17dc461949f4fd0f

Observation a4f98124-b414-415b-a5d4-c4825010d2c2 · outbound

This paper cites Giraffe: Representing scenes as compositional genera- tive neural feature fields.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Giraffe: Representing scenes as compositional genera- tive neural feature fields

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:20.430818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:17.543148Z digest=sha256:1e50f42b4bfa61d0e700df56a25db94484e58e10202b280eb5eb8ce65631ed52

Observation e66a6f52-80df-424c-844f-adb73f893c23 · outbound

This paper cites A simple neural network module for relational reasoning.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering A simple neural network module for relational reasoning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:20.210637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:17.637002Z digest=sha256:27df25f6fc334685b925bf5c725cd7ad2c307b4919b3bee773b80a60e85bfd91

Observation 0c3d53fe-8b34-406e-a0f3-bcdcb4df24c8 · outbound

This paper cites Compositional Attention Networks for Machine Reasoning.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Compositional Attention Networks for Machine Reasoning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:17.725712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:17.725712Z digest=sha256:86305d190038bc86d6b3f8c71829b1967535755b18ff0058b0d834bd706bf0e2

Observation b22eba6e-6dc0-4ade-97a4-8535b2058c59 · outbound

This paper cites Learning transferable visual models from natural language supervision.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Learning transferable visual models from natural language supervision

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:17.799376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:17.799376Z digest=sha256:275c388e042407844698327c9f70839194edaa5398fb67307d8e13d29bf2ce54

Observation e8e258f8-9e15-4947-b421-1b5902e24e86 · outbound

This paper cites Segment anything.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Segment anything

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:17.870461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:17.870461Z digest=sha256:5f6b03e027074617d66b00b7b9f1e9547e30ab6e40ce28bec929ce6c333fa15d

Observation 3eda3f9c-55df-4792-a0c4-65ce2a8f4966 · outbound

This paper cites GPT-4 Technical Report.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering GPT-4 Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:17.945267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:17.945267Z digest=sha256:a0cb023d3b748db37dc3aebae3fa9034d89ed8ba1e8173b8aca62d0140d8da83

Observation 0d9e3326-48fc-4d65-ad73-61666236654d · outbound

This paper cites The Gemini 2 Model Family: Google Deepmind.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering The Gemini 2 Model Family: Google Deepmind

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:19.994411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:18.015298Z digest=sha256:fac65e6fd16a235ffce7e17398b548067b2f0877feb7b9381ab4a6e5b8fdd5a6

Observation 3c804746-c846-4535-9ac2-932819b67dd7 · outbound

This paper cites The Grok Model Family: xAI.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering The Grok Model Family: xAI

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:19.789753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:18.083136Z digest=sha256:0e55315c60354596ca64d5c3799a8ad4728e8de7caf917c13c02bce44a7e0810

Observation fd58f6e1-3696-4dba-9f93-904ae09d03a6 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.167303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.167303Z digest=sha256:d5b2ff622dbe13d3281e266047b47448f978e9abece8874eaea35f59fd774650

Observation 70584297-2efd-4ea7-a61b-8ecc834cb15a · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, 2024.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Llava-next: A strong zero-shot video understanding model, 2024

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:19.606170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:52:18.247348Z digest=sha256:275c8008cb8ae202a959ea2d8d32ec3494795f69e761430ab0c1bd70b3958ac5

Observation 4ba9f124-075a-4a11-962e-08cfa35bf935 · outbound

This paper cites The Llama 3 Herd of Models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering The Llama 3 Herd of Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.335284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.335284Z digest=sha256:147c345b495cc6e8a9ef0506aed9bb0231327be69b4f39bff6215e1a279f9ada

Observation ddad7a53-b8f2-4202-8256-79c24da2b0b9 · outbound

This paper cites H2OVL-Mississippi Vision Language Models Technical Report.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering H2OVL-Mississippi Vision Language Models Technical Report

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.427568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.427568Z digest=sha256:a675792f3cba34d4f237596c89cb992b19bdabff351585700c4c58109a1552f4

Observation 33af74a0-f683-449d-ad39-85af9eb13b2c · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.492837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.492837Z digest=sha256:28cc155919620f4c2ffdea343480ed51d282a83a87cd1afa3cfbfd4d81df0d71

Observation f76e75a6-afe5-4cc8-b853-38164d5b91a9 · outbound

This paper cites Pixtral 12B.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Pixtral 12B

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.563315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.563315Z digest=sha256:362eee8f2fb4a3b604399d89ee79202efcadf227a7e29885dd574f2866f7d09b

Observation 84e3669f-09fa-4e89-bf38-aecdd5b59f7a · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.651138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.651138Z digest=sha256:51f527d2b94ed9ba66ebb6c3b8d09dc6346f0e2222182f23ece598460104e608

Observation 66574c34-677a-4271-8694-6f96025fe335 · outbound

This paper cites OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.730268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.730268Z digest=sha256:b2cb47dc4df405b75d64db80fa5eb945e019ee44b7699c6c87f020ceaef609e8

Observation b05f4901-3c8d-45b0-beb2-d76b513b1823 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.821199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.821199Z digest=sha256:4f00341559a6ea90f1bc24ff0878a2c931957c1f29b83d7ff9ff1285ec47cc57

Observation 862d55c2-a190-4aa5-983e-591ccce7ca14 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.924233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.924233Z digest=sha256:bcf468c15ba0deb7183cb0f1626b0f8c2771c822bfc48b9c03c266468236797d

Observation b1dbb342-9b92-4151-94cc-b8c1d253e240 · outbound

This paper cites Qwen2.5 Technical Report.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Qwen2.5 Technical Report

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:18.990164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:18.990164Z digest=sha256:fad2256e8e9f584c731111a873683df7b02526bc9408290a7f78db4e4c25874d

Observation bbb341f6-1094-4204-bbad-c1ad333a3a6b · outbound

This paper cites Mistral 7B.

IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Mistral 7B

Reference 90

Resolution
malformed identifier
no resolver link, observed 2026-08-06T21:52:19.061479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:19.061479Z digest=sha256:a172d9d58c2be15d2bd4ed61b0087cf5167185a9dea64fd255d7b0e15bc99013

Pith citing papers

Observation 58f1947f-c76c-4444-8630-cb9fb76e6631 · inbound

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs cites this paper.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.147646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.147646Z digest=sha256:19cbd3e8487c37c3d55a39c9a759f58a7c805cd291be3584a43590b1640e945a

Observation 2a8e50f9-e90a-41bd-9177-58400071a09e · inbound

SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis cites this paper.

SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:03.448702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:36:33.149921Z digest=sha256:ac672eb446a010a9d2a7ef86b85624dfb544446e02ec7149be52989440d358b2

Observation d5b28e77-12f4-45cc-8b14-f7dcdfcd832d · inbound

Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models cites this paper.

Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:46:19.352179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T15:04:01.773923Z digest=sha256:ef52a5d17f5c7eb156d201498c42cda8b72abe7e829d60f79f5da4ef5a34d1a5

Observation 239479b0-2d39-4cf4-b714-2e34cd9b6a29 · inbound

IDEAL-Bench: Indoor Dataset and Evaluation suite for Analyzing 3D Layout reasoning cites this paper.

IDEAL-Bench: Indoor Dataset and Evaluation suite for Analyzing 3D Layout reasoning IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T01:11:59.721089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:11:59.721089Z digest=sha256:c1bda079964cf9afc03d3bf49df23afebc50158c7162e6cdbf35b6b30db1d320