Pith. sign in

Paper Citation Record · LEDGER

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2505.02836.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02836 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:21:43.104278Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:09:37.256195Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5b32ae89-6483-435b-9982-647d8edc2a7b · inbound

LLM-BT-Terms: Back-Translation as a Framework for Terminology Standardization and Dynamic Semantic Embedding cites this paper.

LLM-BT-Terms: Back-Translation as a Framework for Terminology Standardization and Dynamic Semantic Embedding Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:43.104278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:43.104278Z digest=sha256:3893b67a3f5e5beaa0915a63569d33dbea7c305bb493bbb4fbcaa84bec18f93c

Observation a31b2567-2df7-46dd-ab93-3d480503aff5 · inbound

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills cites this paper.

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:32.850883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:32.850883Z digest=sha256:1373f82770d3f483b35ffec79d6e470629153bf6ce4725eb7b466da0df59c6ab

Observation f97302bc-978c-48e6-a955-c770b23c387b · inbound

Video Perception Models for 3D Scene Synthesis cites this paper.

Video Perception Models for 3D Scene Synthesis Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:59.729438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:59.729438Z digest=sha256:39e8c417754d51be17c3963123413fb0687ed3ada6d1532f2228ccf72e6dc971

Observation a4081951-88c8-47fd-88d2-e67f97dafcec · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:42.305835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:42.305835Z digest=sha256:f2ac0a15001a19a5cab33daf43454562820a4c3bc002c14d7f66d23b6ad8bb1a

Observation c0daa79a-4e47-481e-a554-80e5d332ed29 · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:36.368585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:36.368585Z digest=sha256:11b1ce216db8be311d84c794f2819758d388f77087ff82c55b82538c244d6c55

Observation 85383ed5-aad4-4b7c-b011-5beca377713b · inbound

TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation cites this paper.

TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T19:19:20.992326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:19:20.992326Z digest=sha256:6c9ec3c7186526664da467cc6d1e8c1eb6f8062429a1204f212b7a2ec859b700

Observation 36e4d326-5360-4fd9-9be8-22eb11c93a07 · inbound

Real2Sim via Active Perception with Behavior Trees Automatically Generated by VLMs cites this paper.

Real2Sim via Active Perception with Behavior Trees Automatically Generated by VLMs Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.648939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T15:07:54.930273Z digest=sha256:0789e84f20065f732b0a99444d3616126b4b763cfcc489dbbf5f88e8f4d5ab0d

Observation 26543a02-6378-46d7-9a85-33f1673708c1 · inbound

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning cites this paper.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:51:00.560832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:a9c5605fad04aaddb53b1265e904c6b55803117c93752f0ade9adbfc176b5bc0

Observation e1a06504-1b43-4f3b-9a11-30a8982adfca · inbound

Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models cites this paper.

Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:16:05.932363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:34:15.916685Z digest=sha256:972fdfd7e58b2144a0044a1b6c8f3b6ad2f0fc56540dc0d6f3c1d2d43d325674

Observation e693439b-16c8-4d60-83b0-4fc46678f9d3 · inbound

SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis cites this paper.

SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:11:03.657154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:36:33.149921Z digest=sha256:ca8e84a26184a0fe57faae8095b849e99a605b79d1985d8fca5404d9bb45fbcf

Observation 3d6782b4-b96a-4ba2-bdf7-ad05b5150363 · inbound

EmbodiedClaw: Conversational Workflow Execution for Embodied AI Development cites this paper.

EmbodiedClaw: Conversational Workflow Execution for Embodied AI Development Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.232494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:34:58.288245Z digest=sha256:6f9e975763776ada51be2cafb775d3a95083916a9d158409d231f0e148181841

Observation 057d08e1-e449-482e-bb80-fe41c9e25c24 · inbound

Repurposing 3D Generative Model for Autoregressive Layout Generation cites this paper.

Repurposing 3D Generative Model for Autoregressive Layout Generation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:12:26.732757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:09:00.779456Z digest=sha256:31aa178a0abc0454214fcef807a7557b6627417a7add7b8adef72c55e6401e9e

Observation f5417681-3361-451a-9452-ae11b24206a1 · inbound

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion cites this paper.

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:36.978937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T09:22:10.651922Z digest=sha256:933e6b274f479431a10fa33e7a1a306d2dce093774ce3b6ea7f1a71035d353b3

Observation 05f27cc3-095e-4f35-800e-ad8400f5a2e8 · inbound

SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation cites this paper.

SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:31:06.580065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T03:22:30.012610Z digest=sha256:5c0a2fdde8b842849ebe0fb89864582fa6d48bc18d40e2ff156080419ec3d2cc

Observation 4112b671-dff6-4520-b981-7c75a03eb637 · inbound

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models cites this paper.

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:59:36.508177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:55:40.370796Z digest=sha256:92fe0f8fd987637a50dd23871933549c85183c2b9f4b54058a5e89c0852ef5cb

Observation d46764da-2d11-4041-85f1-8298ebb0421b · inbound

Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation cites this paper.

Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:01.323641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T23:11:59.282627Z digest=sha256:ea28c6aef3a5ca0105a437c26672753046f7aafa9b261442e05f739b9c26a731

Observation 975ec25c-94f1-4a7a-a71b-d301f9dc1282 · inbound

SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image cites this paper.

SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:26.366072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:59:17.553098Z digest=sha256:8f400c1fee6110a13560fba851cefa33eb07de127e59877a4553ee9ddb1fda70

Observation 8ffe7006-363f-4c1e-ba7a-92183b4b77d6 · inbound

SceneConductor: 3D Scene Generation from a Single Image with Multi-Agent Orchestration cites this paper.

SceneConductor: 3D Scene Generation from a Single Image with Multi-Agent Orchestration Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:07:26.723110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T19:08:08.329817Z digest=sha256:821e2b02028c594d241cb122159ab438244d531e89c2ddd1700ccea6ba852d00

Observation ee327f72-b816-4837-b399-d85f8f7a3d3a · inbound

$\phi$-Scene: Physically Grounded Image-to-3D Scene Reconstruction cites this paper.

$\phi$-Scene: Physically Grounded Image-to-3D Scene Reconstruction Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:09:37.258248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T14:50:16.317301Z digest=sha256:6b1c51e1cfc8c3c70738dd0021130f7e704132828879370f478fad4b903f9935

Observation d0407fcf-3f38-47fc-a0ff-d6f490c6f095 · inbound

NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation cites this paper.

NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:14:25.460428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T08:12:30.693936Z digest=sha256:560b84fd026b1db5ec46e7f9a938f3c6e6d6780c86441ee4389a4c56ce2e02fc

Observation 117dbc72-dd2b-4e11-9aad-123f302ebc8f · inbound

One Video, One World: Turning Monocular Video into Physical 4D Scenes cites this paper.

One Video, One World: Turning Monocular Video into Physical 4D Scenes Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:45.028897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T05:38:19.541391Z digest=sha256:c39d915876a70c1ad3e6d5da7d0118c9acd79203add7b10a6693315a720604ce

Observation 9b69de31-9219-4245-be32-20415ad32f18 · inbound

InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360{\deg} Image cites this paper.

InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360{\deg} Image Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T22:29:10.666143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:29:10.666143Z digest=sha256:de2b50d6724345b45255d57b113ca9c9daac8cda081042f26931f6ed6e34a397

Observation 0973aaf3-5216-40b5-b911-1f59fe8edfde · inbound

Text2Villa: Hierarchical Generation of 3D Indoor Environments with Physics-Aware Analysis-by-Synthesis cites this paper.

Text2Villa: Hierarchical Generation of 3D Indoor Environments with Physics-Aware Analysis-by-Synthesis Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T18:57:31.940710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:57:31.940710Z digest=sha256:de73d1047b131e09b3951843aad24e37a7f7bfc76096950d2dbbf65d60884cc6