Pith. sign in

Paper Citation Record · LEDGER

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2505.02836.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02836 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:21:43.104278Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:09:37.256195Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5b32ae89-6483-435b-9982-647d8edc2a7b · inbound

LLM-BT-Terms: Back-Translation as a Framework for Terminology Standardization and Dynamic Semantic Embedding cites this paper.

LLM-BT-Terms: Back-Translation as a Framework for Terminology Standardization and Dynamic Semantic Embedding Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:43.104278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:43.104278Z digest=sha256:3893b67a3f5e5beaa0915a63569d33dbea7c305bb493bbb4fbcaa84bec18f93c

Observation a31b2567-2df7-46dd-ab93-3d480503aff5 · inbound

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills cites this paper.

RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:32.850883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:32.850883Z digest=sha256:1373f82770d3f483b35ffec79d6e470629153bf6ce4725eb7b466da0df59c6ab

Observation f97302bc-978c-48e6-a955-c770b23c387b · inbound

Video Perception Models for 3D Scene Synthesis cites this paper.

Video Perception Models for 3D Scene Synthesis Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:59.729438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:59.729438Z digest=sha256:39e8c417754d51be17c3963123413fb0687ed3ada6d1532f2228ccf72e6dc971

Observation a4081951-88c8-47fd-88d2-e67f97dafcec · inbound

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth cites this paper.

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T05:15:42.305835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:15:42.305835Z digest=sha256:f2ac0a15001a19a5cab33daf43454562820a4c3bc002c14d7f66d23b6ad8bb1a

Observation c0daa79a-4e47-481e-a554-80e5d332ed29 · inbound

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models cites this paper.

VLM4D: Towards Spatiotemporal Awareness in Vision Language Models Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:36.368585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:12:36.368585Z digest=sha256:11b1ce216db8be311d84c794f2819758d388f77087ff82c55b82538c244d6c55

Observation 85383ed5-aad4-4b7c-b011-5beca377713b · inbound

TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation cites this paper.

TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T19:19:20.992326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:19:20.992326Z digest=sha256:6c9ec3c7186526664da467cc6d1e8c1eb6f8062429a1204f212b7a2ec859b700

Observation 36e4d326-5360-4fd9-9be8-22eb11c93a07 · inbound

Real2Sim via Active Perception with Behavior Trees Automatically Generated by VLMs cites this paper.

Real2Sim via Active Perception with Behavior Trees Automatically Generated by VLMs Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.648939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T15:07:54.930273Z digest=sha256:614ddfd35863f4535e50661850ecacaaa678cad09a91a6de9151e921a884acbe

Observation 26543a02-6378-46d7-9a85-33f1673708c1 · inbound

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning cites this paper.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:51:00.560832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:50:19.988813Z digest=sha256:8ce2b358d24e6cd5ba43f6c40d0fbb207dbecd361614d8802b525ad2673415a0

Observation e1a06504-1b43-4f3b-9a11-30a8982adfca · inbound

Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models cites this paper.

Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:16:05.932363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:34:15.916685Z digest=sha256:5012c61ad43b7ff9597d30038c9406c30055df188faef752fd2df2af39568533

Observation e693439b-16c8-4d60-83b0-4fc46678f9d3 · inbound

SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis cites this paper.

SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:11:03.657154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:36:33.149921Z digest=sha256:704f83120e7030e4aef762a6dd91b745217a6b2d759397f2b43d5a52a55fab60

Observation 3d6782b4-b96a-4ba2-bdf7-ad05b5150363 · inbound

EmbodiedClaw: Conversational Workflow Execution for Embodied AI Development cites this paper.

EmbodiedClaw: Conversational Workflow Execution for Embodied AI Development Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.232494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:34:58.288245Z digest=sha256:834575b34adab6544937d54f3ef32e142c6017f44b9a2a2739dd4b451098e3f7

Observation 057d08e1-e449-482e-bb80-fe41c9e25c24 · inbound

Repurposing 3D Generative Model for Autoregressive Layout Generation cites this paper.

Repurposing 3D Generative Model for Autoregressive Layout Generation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:12:26.732757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:09:00.779456Z digest=sha256:47bab2db03ae1c7113fe07088e74089ad6dfda5b779d253630a5cfc9f8e26852

Observation f5417681-3361-451a-9452-ae11b24206a1 · inbound

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion cites this paper.

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:36.978937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T09:22:10.651922Z digest=sha256:2285e125707875c28df2758b039d568b383542d84d9e2050aad6d65074b39841

Observation 05f27cc3-095e-4f35-800e-ad8400f5a2e8 · inbound

SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation cites this paper.

SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:31:06.580065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T03:22:30.012610Z digest=sha256:f0d23cca904e4d0216df390804a1c76dff4529b911ef36d0c2b344b5b3c88a9e

Observation 4112b671-dff6-4520-b981-7c75a03eb637 · inbound

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models cites this paper.

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:59:36.508177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T04:55:40.370796Z digest=sha256:a4defd9ff0db10fed88f7dd66be78a291a90b398bc56daf07b9e7629da2c17f2

Observation d46764da-2d11-4041-85f1-8298ebb0421b · inbound

Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation cites this paper.

Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:01.323641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T23:11:59.282627Z digest=sha256:849a3b7feec18739e9d1c07a57561bf7b57a27632788e930b1d236c6f052212e

Observation 975ec25c-94f1-4a7a-a71b-d301f9dc1282 · inbound

SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image cites this paper.

SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:26.366072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:59:17.553098Z digest=sha256:e19ea6937dd822e0221d33634fa2037be9baac45a97ab835e16c745ced0ad074

Observation 8ffe7006-363f-4c1e-ba7a-92183b4b77d6 · inbound

SceneConductor: 3D Scene Generation from a Single Image with Multi-Agent Orchestration cites this paper.

SceneConductor: 3D Scene Generation from a Single Image with Multi-Agent Orchestration Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:07:26.723110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T19:08:08.329817Z digest=sha256:e951c755fb721542af127cb26bb39b9d4c02895f0e6b5e29003f76d2e2be03c5

Observation ee327f72-b816-4837-b399-d85f8f7a3d3a · inbound

$\phi$-Scene: Physically Grounded Image-to-3D Scene Reconstruction cites this paper.

$\phi$-Scene: Physically Grounded Image-to-3D Scene Reconstruction Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:09:37.258248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T14:50:16.317301Z digest=sha256:2b0a92717bce25d36685836604d1ba4f4674dbe6e8a57959146577ad92fc1c13

Observation d0407fcf-3f38-47fc-a0ff-d6f490c6f095 · inbound

NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation cites this paper.

NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:14:25.460428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T08:12:30.693936Z digest=sha256:755b38800a8f41281384ec9a8824e87a78f679265ee0c7d21ff7b08813ecc937

Observation 117dbc72-dd2b-4e11-9aad-123f302ebc8f · inbound

One Video, One World: Turning Monocular Video into Physical 4D Scenes cites this paper.

One Video, One World: Turning Monocular Video into Physical 4D Scenes Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:45.028897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:38:19.541391Z digest=sha256:667eba97046cdf1fbf6be5d6d74ecc124ef4efbeeba0012097d3f577dcac66e9

Observation 9b69de31-9219-4245-be32-20415ad32f18 · inbound

InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360{\deg} Image cites this paper.

InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360{\deg} Image Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T22:29:10.666143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:29:10.666143Z digest=sha256:de2b50d6724345b45255d57b113ca9c9daac8cda081042f26931f6ed6e34a397

Observation 0973aaf3-5216-40b5-b911-1f59fe8edfde · inbound

Text2Villa: Hierarchical Generation of 3D Indoor Environments with Physics-Aware Analysis-by-Synthesis cites this paper.

Text2Villa: Hierarchical Generation of 3D Indoor Environments with Physics-Aware Analysis-by-Synthesis Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T18:57:31.940710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:57:31.940710Z digest=sha256:de73d1047b131e09b3951843aad24e37a7f7bfc76096950d2dbbf65d60884cc6