Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:52:19.061479Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 4 inbound Pith citation observations for arXiv:2506.23329.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:52:19.061479Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:42.147646Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T22:46:19.349816Z
89 of 89 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 70399b4d-b6e2-4b73-94bf-bad20488e404 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering LLaVA-OneVision: Easy Visual Task Transfer
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e399d4ea-d573-4ec4-8742-5321185e750e · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering GPT-4o System Card
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1bf531f-f280-4fa3-83ef-1b4ba3d07b95 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering The Claude 3 Model Family: Opus, Sonnet, Haiku
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5f399e7-11ab-498b-a352-6b5424b05bb8 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Gemini: A Family of Highly Capable Multimodal Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4095e18-2038-4477-a16d-c90e47d0c61d · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Qwen2.5-VL Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d393dec2-7886-404f-bd47-c76fdd7b1f14 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 623e3470-0483-491f-8878-f676cb8fedf9 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Perceptions as hypotheses
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b67c78c0-8f4d-430b-9a8a-c6a05d6a7535 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Yuille and Daniel Kersten
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ba89f9d-b7fd-48cc-94e5-6035ea2f9f20 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Efficient and robust analysis-by-synthesis in vision: A computational framework, behavioral tests, and modeling neuronal representations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04895bc9-16e4-42ab-b6cb-fc311ebde2ea · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Bever and David Poeppel
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bd2439d-548f-4935-b8eb-ebc1323b56a8 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Toolformer: Language models can teach themselves to use tools
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41d94738-46fb-42ca-aa07-daad572b27a0 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Visual programming: Compositional visual reasoning without training
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2371c8ed-4db0-4785-b740-808dc0eddae7 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Vipergpt: Visual inference via python execution for reasoning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 759e214b-494b-47d0-a7ff-29d41b221fb4 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Gorilla: Large language model connected with massive apis
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 58910aac-d9e8-4948-9f22-a5cf9a3c51dd · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Blender - a 3D modelling and rendering package, 2016
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4f94bb30-7534-47ce-b4a6-8f2c24fe64e2 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering https://github.com/modelcontextprotocol, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ecc096df-bc12-4ee8-b717-625b2d4f4528 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Blendermcp
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c343af7a-bbde-476a-9168-91e73e1b1192 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Embodied question answering
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24c50811-3954-4c4c-8c9a-3443545a9ca4 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Multi-target embodied question answering
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b4e46d7-ef6f-4292-8bc7-a0adc8407f6b · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3d concept learning and reasoning from multi-view images
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a512965-3142-40a5-a3af-38bb88cef8f6 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering One step at a time: Long-horizon vision-and-language navigation with milestones
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c6d80a6f-3acd-410b-b3bb-eb48b0cd720d · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de9ccc8a-466a-4a37-afa9-fe5ffb61f905 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9a6c378-4606-4592-94b9-abeafdd0c1ae · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Barrow and Jay M
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05dad22f-a4c9-4d5e-b07d-273a93a6fc3a · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Soft rasterizer: A differentiable renderer for image-based 3d reasoning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47e3950a-d8e2-4c22-ac40-8243dea9bbdb · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Carr, Jonathan Ragan-Kelley, and Frédo Durand
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03cb4dd2-5833-44e6-a443-1ded4a1cbcc2 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Differentiable vector graphics rasterization for editing and learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fdf89410-523c-4f03-82e4-6e26fc4b84b2 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Nerf: Representing scenes as neural radiance fields for view synthesis
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 185ef9c7-49ac-46b2-984d-870a0eff56e2 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Unisurf: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 170cb4bb-be8d-4702-b529-96c80dc09769 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering V olume rendering of neural implicit surfaces
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe13a610-3f42-47f0-be62-c42423e06960 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Synsin: End-to-end view synthesis from a single image
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a40ed8df-92b3-4d8e-9c4f-eb74c9b4898b · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3d gaussian splatting for real-time radiance field rendering
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 120bdca9-0787-411c-af13-41bf6860ac8f · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Path-space differentiable rendering
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a82b5d2-f0fa-4dcd-90ad-0e7dafa35cb0 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Inst3d-lmm: Instance-aware 3d scene understanding with multi-modal instruction tuning, 2025
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation baa76daf-d0cc-4046-867a-c8c4e18cb39f · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af2108fa-3480-4223-889d-1ddbf54662d7 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Scanqa: 3d question answering for spatial scene understanding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0fdde439-780f-4367-abce-abcb0b2ba4ac · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3d-vista: Pre- trained transformer for 3d vision and text alignment
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 804d0774-75fd-4430-b308-0a344c7c0b98 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 069e343b-baa2-47c1-ab21-fa116d0be223 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a82e53a-4cde-4c52-8b39-2781b3ce7187 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91a71c99-fc19-4bfb-9743-85f859313aca · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Scalable 3d captioning with pretrained models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8e232af-a3db-4d52-a83a-5d822a3dd8a3 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Openeqa: Embodied question answering in the era of foundation models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5364b1bc-adb6-4a12-a480-5b8ade7cb1cb · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79c150fb-dcb2-4dda-9b34-946805354b5f · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Sceneverse: Scaling 3d vision-language learning for grounded scene understanding
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2243b079-f024-4250-87c2-d590e88832e9 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Spatialrgpt: Grounded spatial reasoning in vision-language models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13c423ae-e694-4b4d-9483-84ce12cea5bc · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering ConceptFusion: Open-set Multimodal 3D Mapping
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0854f686-091f-4bf1-9322-c38faaea3951 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Context-aware entity grounding with open-vocabulary 3d scene graphs
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 09505462-0c34-481b-9759-994420a4609d · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Tenenbaum, Antonio Torralba, Florian Shkurti, and Liam Paull
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d0d8d06-ed63-48ce-ab2b-6ad2bebe8242 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3D-LLM: Injecting the 3D World into Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af0980fc-8aae-4b20-8ca6-af1ebf75083a · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Agent3D-Zero: An Agent for Zero-shot 3D Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 477b6121-dda0-4230-a3af-7ebb78398485 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa631a06-2c28-496b-934d-0e48b671585c · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Openeqa: Embodied question answering in the era of foundation models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf2e2b69-3430-422d-9178-84203d7d112f · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Kulkarni, Pushmeet Kohli, Joshua B
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92033c46-13f7-4403-9222-3acc43d758a9 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Learning to infer graphics programs from hand-drawn images
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab25332c-a816-4f5b-8b68-f98bd8ce40d5 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Learning to infer and execute 3d shape programs
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 370dc2bc-3632-4845-839c-20e7bf03bfeb · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Shapeassembly: Learning to generate programs for 3d shape structure synthesis
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 616311f7-9e01-49cd-bbf1-2bee4be39133 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Scenecraft: An llm agent for synthesizing 3d scenes as blender code
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1d01f54-7b0d-4d4d-ac16-f51c5f1babf8 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering The Scene Language: Representing Scenes with Programs, Words, and Embeddings
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6f75e8e-096c-4c4f-82dd-00bfbd5e88a2 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering 3D-GPT: Procedural 3D Modeling with Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0e1084c-6cc2-4184-af33-3f98af61bc67 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering SceneMotifCoder: Example-driven Visual Program Learning for Generating 3D Object Arrangements
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2abac4fd-99b5-4555-9d70-d6c9bc572221 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6757f77-95e2-4fef-8b62-63e1a9ee112f · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Creative agents: Empowering agents with imagination for creative tasks
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79bd3400-c7c6-46cb-b80a-a8beebee358c · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Scenex: Procedural controllable large-scale scene generation via large-language models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c70a0e45-e2d4-4063-9d60-a6cd36e02871 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Program-guided image manipulators
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 56d960d6-db08-4b58-8524-19e5526477d4 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef41b7f0-8b70-486f-bd19-089ce59ad090 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Virtualhome: Simulating household activities via programs
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bfd141a2-30a4-4447-96ec-8d97ac1962a1 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Xia, Peng Xu, Karol Hausman, Brian Ichter, Peter R
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a65d344b-a34a-4ad8-8d9e-6228213936f2 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering MONet: Unsupervised Scene Decomposition and Representation
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d8a2049-58fe-4748-bb8a-c73c40b2b546 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent Representations
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4f98124-b414-415b-a5d4-c4825010d2c2 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Giraffe: Representing scenes as compositional genera- tive neural feature fields
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e66a6f52-80df-424c-844f-adb73f893c23 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering A simple neural network module for relational reasoning
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c3d53fe-8b34-406e-a0f3-bcdcb4df24c8 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Compositional Attention Networks for Machine Reasoning
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b22eba6e-6dc0-4ade-97a4-8535b2058c59 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Learning transferable visual models from natural language supervision
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8e258f8-9e15-4947-b421-1b5902e24e86 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Segment anything
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eda3f9c-55df-4792-a0c4-65ce2a8f4966 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering GPT-4 Technical Report
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d9e3326-48fc-4d65-ad73-61666236654d · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering The Gemini 2 Model Family: Google Deepmind
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3c804746-c846-4535-9ac2-932819b67dd7 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering The Grok Model Family: xAI
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd58f6e1-3696-4dba-9f93-904ae09d03a6 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70584297-2efd-4ea7-a61b-8ecc834cb15a · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Llava-next: A strong zero-shot video understanding model, 2024
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ba9f124-075a-4a11-962e-08cfa35bf935 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering The Llama 3 Herd of Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddad7a53-b8f2-4202-8256-79c24da2b0b9 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering H2OVL-Mississippi Vision Language Models Technical Report
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33af74a0-f683-449d-ad39-85af9eb13b2c · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f76e75a6-afe5-4cc8-b853-38164d5b91a9 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Pixtral 12B
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84e3669f-09fa-4e89-bf38-aecdd5b59f7a · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66574c34-677a-4271-8694-6f96025fe335 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b05f4901-3c8d-45b0-beb2-d76b513b1823 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 862d55c2-a190-4aa5-983e-591ccce7ca14 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1dbb342-9b92-4151-94cc-b8c1d253e240 · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Qwen2.5 Technical Report
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbb341f6-1094-4204-bbad-c1ad333a3a6b · outbound
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering Mistral 7B
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58f1947f-c76c-4444-8630-cb9fb76e6631 · inbound
STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a8e50f9-e90a-41bd-9177-58400071a09e · inbound
SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5b28e77-12f4-45cc-8b14-f7dcdfcd832d · inbound
Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 239479b0-2d39-4cf4-b714-2e34cd9b6a29 · inbound
IDEAL-Bench: Indoor Dataset and Evaluation suite for Analyzing 3D Layout reasoning IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.