Pith. sign in

Paper Citation Record · LEDGER

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

As of 23 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 5 inbound Pith citation observations for arXiv:2507.06484.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06484 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:09:30.946930Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T23:11:00.840340Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:56:34.752902Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 757fdf72-0fd8-42f4-8226-9a5f310f6c49 · outbound

This paper cites GPT-4 Technical Report.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:25.060519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:25.060519Z digest=sha256:c3ebcb471206ea13c2ec8ac2b243f246e942852f6a5e836c41ff091647fb2d75

Observation 56ab8fd7-a601-4246-b55f-bec25da8fa23 · outbound

This paper cites Open-Universe Indoor Scene Generation using LLM Program Synthesis and Uncurated Object Databases.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Open-Universe Indoor Scene Generation using LLM Program Synthesis and Uncurated Object Databases

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:25.147440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:25.147440Z digest=sha256:472ef102b6c16db28ca667e63b0e04fd2b4c585d5105eebff711378f14c836f9

Observation 3e3f0ed2-7c48-42d1-a7a2-3999f5c443ca · outbound

This paper cites I-design: Personal- ized llm interior designer.arXiv preprint arXiv:2404.02838,.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds I-design: Personal- ized llm interior designer.arXiv preprint arXiv:2404.02838,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:25.251909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:25.251909Z digest=sha256:f73765d9a134f27fc4a39a7f287d965e57096bca4c4a762f4d689eb468620de4

Observation 22b496ac-4bef-4dfc-a5d2-e0b4a47a2a15 · outbound

This paper cites Learning spatial knowledge for text to 3d scene generation.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Learning spatial knowledge for text to 3d scene generation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:36.332637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:25.319361Z digest=sha256:57abda475a6fabbe71df618351e37cc1a8139ce8ef078529477031203711e887

Observation 40790c0b-4ea2-4831-8d58-c1c6cbee8dec · outbound

This paper cites SceneSeer: 3D Scene Design with Natural Language.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds SceneSeer: 3D Scene Design with Natural Language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:25.395511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:25.395511Z digest=sha256:583febdfc43d8959a1446c7227fe1f332698c8c92281e17f168ac1ef9ec73731

Observation 82379160-45a2-46f9-b488-2652080775ae · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:36.194302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:25.490051Z digest=sha256:5d4400da1c955be7d99853dc8578d62cbf5a0931efd872fc1f4acddaa52eaa12

Observation b60dd022-5bcc-4f0e-bfd5-3d99bb6bc0da · outbound

This paper cites URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:25.571972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:25.571972Z digest=sha256:6d951dd8cdfc2a4299630cdca05a96ab17406f5f0ce2ae57d80bb5c693881e45

Observation aa4cd893-691c-4236-96ff-78c5eb4d968e · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:25.638480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:25.638480Z digest=sha256:8067ae51635eee18d75f79340ba411f9ffbcd6c61f31567f3fb8b65d08a264cf

Observation 5e35302d-176e-4e4c-af07-ec4ccdfc46a8 · outbound

This paper cites Wordseye: An automatic text-to-scene conversion system.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Wordseye: An automatic text-to-scene conversion system

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:36.160901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:25.731961Z digest=sha256:14ba279b40ed3cfe72d681cfdc2802766bc123dab9f82406cfb6cf4ace2fa65c

Observation 0bb1a632-f669-463d-8509-c1b34df47882 · outbound

This paper cites Procthor: Large-scale embodied ai using procedural generation.Ad- vances in Neural Information Processing Systems, 35:5982– 5994, 2022.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Procthor: Large-scale embodied ai using procedural generation.Ad- vances in Neural Information Processing Systems, 35:5982– 5994, 2022

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:36.048234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:25.878293Z digest=sha256:cf64d833ab33f754cd2544684b192b9ef555370366d3873b712b794da0405097

Observation 06aed207-0e9d-41d1-96a6-bb47042fd37c · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Objaverse: A universe of annotated 3d objects

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:35.866121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:25.968577Z digest=sha256:21c8aa2455964817e67d992913af65f10c44c3f70d2db629cbee908142a59fd5

Observation 4f681eb1-6109-42ce-b1ff-eddb7344277b · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:26.052398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:26.052398Z digest=sha256:4375fd5b2f643f7c85704f1da46d6087399cfa91ececcd0f34ae1ce920ab1e24

Observation b692d7ec-d7ff-4038-bd19-782c9f7bbc46 · outbound

This paper cites ImageNet: A large-scale hierarchical im- age database.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds ImageNet: A large-scale hierarchical im- age database

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:35.702757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:26.151933Z digest=sha256:f97298c7815c4db0ac6a9f4352903a89a68daa9d7eaebc0ac0544d4739b84084

Observation d4d23b0e-ff5c-4ecd-8eba-2997edfc0767 · outbound

This paper cites Disentangled 3D Scene Generation with Layout Learning.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Disentangled 3D Scene Generation with Layout Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:26.240368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:26.240368Z digest=sha256:5c90beab5685f28028f7c1f696d0100bb146cf93e6c7891759f04625a77895b6

Observation f04a234c-3512-43dd-a05c-32384fc3740b · outbound

This paper cites Diffusion360: Seamless 360 Degree Panoramic Image Generation based on Diffusion Models.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Diffusion360: Seamless 360 Degree Panoramic Image Generation based on Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:26.363654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:26.363654Z digest=sha256:77c4bf58be3768ffcbfabd45ff69e5327f73772d5d6931ced8106b35f91561e0

Observation 94d2422b-0527-4e9a-acfc-f877a188431a · outbound

This paper cites Layoutgpt: Compositional visual plan- ning and generation with large language models.Advances in Neural Information Processing Systems, 36, 2024.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Layoutgpt: Compositional visual plan- ning and generation with large language models.Advances in Neural Information Processing Systems, 36, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:35.644267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:26.459243Z digest=sha256:088fe7e03da8fdced42f60a4d3f5f462cd4523d040e210a4f5c10c300c2c4599

Observation 9c076fbb-9d0b-460a-aea4-3cf6dbdf70b7 · outbound

This paper cites Example-based synthesis of 3d object arrangements.ACM Transactions on Graphics (TOG), 31(6):1–11, 2012.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Example-based synthesis of 3d object arrangements.ACM Transactions on Graphics (TOG), 31(6):1–11, 2012

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:35.565521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:26.586131Z digest=sha256:a2af157587aaa5f04b0b032735f2a213b3cd72d49fe6fd64a228f884ca2d10f9

Observation dda20418-a8e1-4c05-b2de-3dbee326ae00 · outbound

This paper cites Any- home: Open-vocabulary generation of structured and tex- tured 3d homes.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Any- home: Open-vocabulary generation of structured and tex- tured 3d homes

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:35.391203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:26.694531Z digest=sha256:d1c1364b5ad5cb3ddb8b7bca852d52e706fcb08c3c9fc05fcb2dcb0704666e8c

Observation 12864e82-d745-4f54-8ff6-76f09672ce55 · outbound

This paper cites Rel3D: A minimally contrastive benchmark for grounding spatial relations in 3d.Advances in Neural Information Pro- cessing Systems, 33:10514–10525, 2020.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Rel3D: A minimally contrastive benchmark for grounding spatial relations in 3d.Advances in Neural Information Pro- cessing Systems, 33:10514–10525, 2020

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:35.207373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:26.771115Z digest=sha256:9e6be01698438a9bf06dab175279a9a8b7deefc5890b4a949a7934fea1580a5d

Observation 8fc49ff0-f424-47f4-8370-fbfdad7a956c · outbound

This paper cites Text2Room: Extracting textured 3D meshes from 2D text-to-image models.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Text2Room: Extracting textured 3D meshes from 2D text-to-image models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:34.922274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:26.844439Z digest=sha256:87acd6f10d0d9b0dc86839b06257b54b1f849f5a421ec718c6b427cb34e06bdb

Observation 73641011-fd9f-4384-8daa-bb76568dc2b9 · outbound

This paper cites 3D-LLM: In- jecting the 3D world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds 3D-LLM: In- jecting the 3D world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:34.695667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:26.935072Z digest=sha256:ca6630e10ceedd2cf1cdc10209e55cdc8987672fd66eecc5792a65a1460e1b98

Observation 628e9190-c971-4d58-a671-8761e57772cd · outbound

This paper cites Scenecraft: An llm agent for synthesizing 3d scenes as blender code.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Scenecraft: An llm agent for synthesizing 3d scenes as blender code

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:34.559575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:27.045961Z digest=sha256:5fbdb7cff83669b62a5d97814f51b4e395c948c225f4daf0047a67c3d68af608

Observation 6e0d37c2-cbfe-42a5-8923-7c1e5662c247 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Large Language Models Cannot Self-Correct Reasoning Yet

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:27.114885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:27.114885Z digest=sha256:f8f54fc6c043873a0cb3289b19f47397442a0fc4928fcfca2e9f81a7402b64bb

Observation 400aedfb-5b36-4ce2-be5b-4d624cbe5ae6 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Training Language Models to Self-Correct via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:27.212385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:27.212385Z digest=sha256:400a22be2758607b08cdc3ec9b34d5088faa3b5ac5b51ecd86b9fd0c699c6295

Observation 97f7ff7f-42c6-4291-b3e2-842bfb087700 · outbound

This paper cites InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:27.266519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:27.266519Z digest=sha256:6f799031c14d5b817ee8f437893228a0bb59da3a9be4df34ae8e03e6a72f9f5b

Observation 8c6b66b2-27c5-46ea-9dd3-d96b532ec92b · outbound

This paper cites Microsoft coco: Common objects in context.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Microsoft coco: Common objects in context

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:27.354684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:27.354684Z digest=sha256:f8f29691bb32539157ad0cc6491f0c1c09ea58ac0f30ecaa031e332794e09a93

Observation 649621b1-eddf-4348-9ed5-86f25b3237cf · outbound

This paper cites Material palette: Extraction of materials from a single image.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Material palette: Extraction of materials from a single image

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:34.404330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:27.417306Z digest=sha256:90d404005aa0d78b9da5203956472c33ff2bdf7aff83e08017b62701ab1f9a2b

Observation c6654d33-4ca3-4af9-8118-56fa931bdbab · outbound

This paper cites Language-driven synthe- sis of 3d scenes from scene databases.ACM Transactions on Graphics (TOG), 37(6):1–16, 2018.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Language-driven synthe- sis of 3d scenes from scene databases.ACM Transactions on Graphics (TOG), 37(6):1–16, 2018

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:34.331089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:27.482068Z digest=sha256:045d7228a765c4c77608e3566599d16f5d51847109cebc420c144873a3c363c6

Observation ecdc951c-3b1f-4264-8d72-89b6ac9ace65 · outbound

This paper cites RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:27.557334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:27.557334Z digest=sha256:696fe199b178d580dfe02b0f1753b608c6be7cc424eb2797faab234e0fe6f80e

Observation f0a61184-4834-46a2-80c6-f0dc2d834564 · outbound

This paper cites Edify 3D: Scalable High-Quality 3D Asset Generation.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Edify 3D: Scalable High-Quality 3D Asset Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:27.659339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:27.659339Z digest=sha256:b05519eafbbe4bbaf00b59603058d309d89973d68c5e9ccdb964537bba0adacf

Observation d5d1dbc3-e9e4-4b1a-a653-ef2e27248c83 · outbound

This paper cites Atiss: Autoregres- sive transformers for indoor scene synthesis.Advances in Neural Information Processing Systems, 34:12013–12026,.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Atiss: Autoregres- sive transformers for indoor scene synthesis.Advances in Neural Information Processing Systems, 34:12013–12026,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:27.762897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:27.762897Z digest=sha256:89385eb186c4110df4d544f86e893996512c2aa0fe088437463c40193bd84d72

Observation 4166ae79-89be-490f-a2ca-63425b49735e · outbound

This paper cites Advances in data- driven analysis and synthesis of 3d indoor scenes.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Advances in data- driven analysis and synthesis of 3d indoor scenes

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:34.204046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:27.847004Z digest=sha256:721da78d1e5aabcdb8d0f9d5034b4a4a9cb008f3bb65e7e6c6bc54c5c3d0e781

Observation ac336ab7-8668-4344-ac3f-b09e09c91a16 · outbound

This paper cites Compositional 3d scene generation using locally conditioned diffusion.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Compositional 3d scene generation using locally conditioned diffusion

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:34.091543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:27.896624Z digest=sha256:c30010d40bc6486a91c99e173a9ca1301f5bac7958f4008fe1f9752f9424af4b

Observation 6ed5852a-8429-4350-bd96-029b7ddbe82e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Learning transferable visual models from natural language supervi- sion

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:27.959566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:27.959566Z digest=sha256:a89eb6e528ed233a92bbc4442b204e801afb1f2f301aa52901fcc0dfc2d345f8

Observation 45cd910c-6ffe-4d68-8f44-932d3a284a16 · outbound

This paper cites Lay-A-Scene: Personalized 3D Object Arrangement Using Text-to-Image Priors.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Lay-A-Scene: Personalized 3D Object Arrangement Using Text-to-Image Priors

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:09:31.267876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:28.043773Z digest=sha256:8c0a3fcc25f7fd27a003570c29331dc5bf5fea733409b08ceb5007af58368466

Observation 9924f5f2-4431-475e-af85-12b9231f8995 · outbound

This paper cites Grounded sam: Assembling open-world models for diverse visual tasks,.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Grounded sam: Assembling open-world models for diverse visual tasks,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:28.095933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:28.095933Z digest=sha256:8d296887c39964ebf56c0ec0d3176f1057afb5c0ac9ba3bdc9a77247b83dd9f9

Observation 22cfc7e8-f557-498f-ac5c-98c37baee8f6 · outbound

This paper cites Susskind.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Susskind

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:28.195189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:28.195189Z digest=sha256:ba39b00439600a60b533575de98c5159521fdc4029760f52dedaed520f4143c3

Observation 68c8327f-2d4c-433f-85f0-231279d08c57 · outbound

This paper cites Object Hallucination in Image Captioning.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Object Hallucination in Image Captioning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:28.253128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:28.253128Z digest=sha256:a615452519bed5b4e8b02e9863991aa755ee2ec9d15b3ebfaedb61e57cea569d

Observation df90daee-1c8d-4296-a659-20304ca21b12 · outbound

This paper cites Controlroom3d: Room gen- eration using semantic proxy rooms.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Controlroom3d: Room gen- eration using semantic proxy rooms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:28.335751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:28.335751Z digest=sha256:35923da22c053d4dc831a006ad5700c187fadb603414969dd245bdcbf533f3ce

Observation 378a64ea-5841-49a3-ad39-8a4a146a5199 · outbound

This paper cites Real-time automatic 3d scene generation from natural language voice and text de- scriptions.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Real-time automatic 3d scene generation from natural language voice and text de- scriptions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:33.928249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:28.456742Z digest=sha256:b0e07ad7bc24f7660712810a9dfd995eb243b1a518d98c3dcc97aa42a9ef3eed

Observation 5941c165-779d-4786-9f71-77114fbce186 · outbound

This paper cites Horizonnet: Learning room layout with 1d represen- tation and pano stretch data augmentation.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Horizonnet: Learning room layout with 1d represen- tation and pano stretch data augmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:33.777795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:28.580659Z digest=sha256:bd25cf4663dca0a002bb1ee392bbcf9c3555f3f4c1303e604a44374ca83dc409

Observation 189c25ce-7354-44c2-8df8-900b39caf607 · outbound

This paper cites LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:28.680768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:28.680768Z digest=sha256:e412f445bc58bd34cf9ec3ee40515a6070ef5f17bd5cfe7a1a3d0633e9361e56

Observation 31197d64-302c-4235-afef-3475e3b4b745 · outbound

This paper cites Factorsim: Generative simulation via factorized rep- resentation.Advances in Neural Information Processing Sys- tems, 37:87438–87472, 2024.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Factorsim: Generative simulation via factorized rep- resentation.Advances in Neural Information Processing Sys- tems, 37:87438–87472, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:33.665930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:28.768156Z digest=sha256:64445f03e3d7d37c64e4b2194aea9d5b273c3969a641eb67a2aa2d6a7e5ad23a

Observation f236d975-a3cd-4f93-b853-7c59bce4e246 · outbound

This paper cites Partial-view object view synthesis via filtering inversion.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Partial-view object view synthesis via filtering inversion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:33.571398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:28.869353Z digest=sha256:07ed357341c8b18afc5464057b16597a2c00e63085b13d915a5ab0c20a419b44

Observation 61e176e0-4271-4ed5-ac51-1d454af5b741 · outbound

This paper cites Lgm: Large multi-view gaussian model for high-resolution 3d content creation.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Lgm: Large multi-view gaussian model for high-resolution 3d content creation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:33.460132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:28.965667Z digest=sha256:0b5123cd348d597eba133dee5ead61c781abf85b674085f1e6f065851f84aa2c

Observation 0a787edc-b8d6-4976-adc7-9cb6b0b6db5f · outbound

This paper cites DiffuScene: Denoising diffu- sion models for generative indoor scene synthesis.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds DiffuScene: Denoising diffu- sion models for generative indoor scene synthesis

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:33.329825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:29.068290Z digest=sha256:96dd41c1cb120f97cfa57b89655a131c9d6a098aa214a69c3ef5db904c7fa668

Observation 5fb754f0-75b4-4bf3-bd21-ee0657327702 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:29.197304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:29.197304Z digest=sha256:b17dedaf629ee3ad7b5d0e2d8263301b4d0e4cdba9507323488509e0f4faf602

Observation 3c692faa-d191-487c-b32e-a6db1f62d42f · outbound

This paper cites Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:29.291799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:29.291799Z digest=sha256:71863d60f864836473677494c06550f29d34c684fcc8eea6ed03686dbccfb633

Observation 36be4184-03cd-4347-a035-b2a495046f64 · outbound

This paper cites RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative Simulation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:29.396038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:29.396038Z digest=sha256:148bcff9740fc77ccacb572776578b5e3c4810163a3f6ecd0670bd84831927e6

Observation faed3a08-26c4-42f0-bddc-03853aa784b0 · outbound

This paper cites Architect: Generating vivid and interactive 3d scenes with hierarchical 2d inpainting, 2024.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Architect: Generating vivid and interactive 3d scenes with hierarchical 2d inpainting, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:33.206158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:29.490104Z digest=sha256:2a68d8854e86add4d6b8cd666f4df900d2241c14942f128080f3f9b106ebee3a

Observation d4598024-9fcb-44ce-b65b-52ff61dc3d11 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:33.093312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:29.607181Z digest=sha256:972c1b4ac2825396151275fabc8c95d7138b0685cd5507ec89348893ad1b19ae

Observation 685240fa-fe05-4b7f-be0e-61bef0ae514d · outbound

This paper cites InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:29.717119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:29.717119Z digest=sha256:fc19e0b4eccfffb253b0034a6f5fd95c279a280a212e34d9fbfda1f08e52bdf3

Observation c76dbc7a-d1d5-4e2d-ba4e-b6f1566e51e7 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:29.817253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:29.817253Z digest=sha256:31497efe004736bc1d829227a60e55c3403a3db9d9b0e5ef1559a120041d049e

Observation ff5bf40f-6a21-4fa3-8ea0-bf0eaa3b2871 · outbound

This paper cites Physcene: Physically interactable 3d scene synthe- sis for embodied ai.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Physcene: Physically interactable 3d scene synthe- sis for embodied ai

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:32.914444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:29.929380Z digest=sha256:7162815b9909aacb3cd0233fa233a09b6bfdaaca108c91196ddea4b00a308991

Observation 32381829-480e-4e4d-b2bf-515812bb4628 · outbound

This paper cites Holodeck: Language guided gen- eration of 3d embodied ai environments.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Holodeck: Language guided gen- eration of 3d embodied ai environments

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:32.707862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:30.026747Z digest=sha256:87328206635224dfe259474bac8004aebccf87c190dfb9ae993db40320998ff1

Observation d65d69dd-45a4-4cc0-889a-bf0928bf6c8f · outbound

This paper cites The clutterpalette: An interactive tool for detailing indoor scenes.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds The clutterpalette: An interactive tool for detailing indoor scenes

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:32.521254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:30.106167Z digest=sha256:7ccc50fea890c5936a8d6cf84a1626d453f51049d60da2f69097f2267004fa8b

Observation e7f87d50-a65e-4b5d-aeff-8e8484634865 · outbound

This paper cites RLF-V: Towards trustworthy MLLMs via behavior alignment from fine-grained correctional hu- man feedback.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds RLF-V: Towards trustworthy MLLMs via behavior alignment from fine-grained correctional hu- man feedback

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:32.329272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:30.186377Z digest=sha256:ae74fc7c8d011e0b70b0566c6a54e2ff530d3ed85e46d8310d899e60acc6f466

Observation e4dc9c68-d7aa-4626-8768-134154e5c463 · outbound

This paper cites Clay: A controllable large-scale generative model for creat- ing high-quality 3d assets.ACM Transactions on Graphics (TOG), 43(4):1–20, 2024.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Clay: A controllable large-scale generative model for creat- ing high-quality 3d assets.ACM Transactions on Graphics (TOG), 43(4):1–20, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:30.324696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:30.324696Z digest=sha256:ec072bd36aff1af5c9df03e0b9d95d113904aec35744e8441b405f1e3fdc9f02

Observation 7f84a514-1eeb-472b-84de-235ba23992e0 · outbound

This paper cites SceneWiz3D: Towards Text-guided 3D Scene Composition.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds SceneWiz3D: Towards Text-guided 3D Scene Composition

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:30.414347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:30.414347Z digest=sha256:ab87e10d9b90b4a23f5884be2cdcbf1a5bb6a5d18e68cd1a941f39c2154442ac

Observation a2e42970-38ae-4611-a5ed-21dbd4680c14 · outbound

This paper cites DreamScene360: Uncon- strained text-to-3D scene generation with panoramic gaus- sian splatting.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds DreamScene360: Uncon- strained text-to-3D scene generation with panoramic gaus- sian splatting

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:32.122004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:30.558770Z digest=sha256:21c9dc608ec687f2222952a8e2d6cb380092f7143cc52b3e967950b4738793e4

Observation c99a6018-6cf9-4342-94d3-ed9242d464bf · outbound

This paper cites GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:30.649221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:30.649221Z digest=sha256:b8f72e6b4a999c682d5166721d11d993366dd38dbf082a9e6ccb52c5f97fea20

Observation 4c5e14bb-7747-43ea-9a01-5b16d95a5721 · outbound

This paper cites placements.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds placements

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:31.918596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:30.753677Z digest=sha256:f49961e25244a67a8a3fff08ea10ba1f793b2e6d33e97768936ff4dffcb807c0

Observation 68d92fdb-d18e-4240-bd51-c81cf5b9e13a · outbound

This paper cites an unresolved cited work.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:09:31.791694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:30.834018Z digest=sha256:05949d580c8a88bb08860763a23a8101dc58d11c848aa2cd68aa33025aef1062

Observation e54cadc9-f2ba-4df3-a054-1a519a52ea5c · outbound

This paper cites receptacle objects.

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds receptacle objects

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:09:31.585806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T19:09:30.946930Z digest=sha256:8f096c5629d5fc15544ae3a34409e4b7f9fbd4909c3410909a5fb389f11aa68b

Pith citing papers

Observation f4882744-93ec-4c58-b0f6-95ce98fe92b5 · inbound

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning cites this paper.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:51:00.591108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:7adb33fe868b87b026a3421ba3b1e3de3bd2215c74a4a883edb1564900347334

Observation 7b8f4dcd-56b7-4aaa-859a-2ba50a53d977 · inbound

StoryBlender: Inter-Shot Consistent and Editable 3D Storyboard with Spatial-temporal Dynamics cites this paper.

StoryBlender: Inter-Shot Consistent and Editable 3D Storyboard with Spatial-temporal Dynamics 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:28:26.411745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T23:24:26.557732Z digest=sha256:d8e3ca5c7b927e51ae1b02f5a962b2858ad279725fb10c9a5d122902b48dde9c

Observation ce742356-f2b1-44fe-a795-630313bf90dc · inbound

Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis cites this paper.

Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:53:14.996405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T11:49:53.113415Z digest=sha256:9657f7332670a3edb36630e9a3f8786bb838d3f5727cbb0aca262d318204bc00

Observation adb0c979-3899-46e7-aaa7-e10b964af779 · inbound

Function2Scene: 3D Indoor Scene Layout from Functional Specifications cites this paper.

Function2Scene: 3D Indoor Scene Layout from Functional Specifications 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:12:46.730083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:11:00.840340Z digest=sha256:a7d620414840f2f07d2a05597ce24d01aff92a76084093e59973e74bd549e193

Observation 667ff8d3-a95f-4923-85bc-f27ea0d09903 · inbound

PerceptTwin: Semantic Scene Reconstruction for Iterative LLM Planning and Verification cites this paper.

PerceptTwin: Semantic Scene Reconstruction for Iterative LLM Planning and Verification 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:56:34.754685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T09:34:34.220944Z digest=sha256:5474ecc6ccc24c875c702d9898a71c73cdd1774adb60e44f40d25cd9d3e39f6b