Pith. sign in

Paper Citation Record · LEDGER

Video Perception Models for 3D Scene Synthesis

As of 10 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 0 inbound Pith citation observations for arXiv:2506.20601.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20601 v1

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:49:06.433344Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

84 of 84 outbound references displayed

  • verified exact2
  • verified fuzzy48
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bdc588d8-eba1-4acb-90c8-5ead96a99bf7 · outbound

This paper cites Augmented reality meets computer vision: Efficient data generation for urban driving scenes.International Journal on Computer Vision (IJCV), 2018.

Video Perception Models for 3D Scene Synthesis Augmented reality meets computer vision: Efficient data generation for urban driving scenes.International Journal on Computer Vision (IJCV), 2018

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:56.034466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:56.034466Z digest=sha256:7a72c429242505145a90b794e832ab39bb7691315582a0a9786b448f2627cf13

Observation 6c0ce686-1fc4-4440-9273-079e55b61f08 · outbound

This paper cites GPT-4 Technical Report.

Video Perception Models for 3D Scene Synthesis GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:56.109342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:56.109342Z digest=sha256:c9ba49ae4e3cf66481b9ccecad4c3703070758fdbf289862a604f4d1b3e853a1

Observation a23ec9aa-a0d6-4663-a18f-89a11bf97dd5 · outbound

This paper cites Open-Universe Indoor Scene Generation using LLM Program Synthesis and Uncurated Object Databases.

Video Perception Models for 3D Scene Synthesis Open-Universe Indoor Scene Generation using LLM Program Synthesis and Uncurated Object Databases

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:56.325321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:56.325321Z digest=sha256:2297947354a68307e3d36a107a98b8058cfde8d1cce09b8dc9a030364c2e65e9

Observation 1bb3a597-23de-44b0-a891-8dbf4f86a499 · outbound

This paper cites I-design: Personalized llm interior designer.

Video Perception Models for 3D Scene Synthesis I-design: Personalized llm interior designer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:56.433828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:56.433828Z digest=sha256:d0f661deb636f80876bda8732f0a26538b9f6de0c4a027b7bf66554bc5a34f3d

Observation 5d4dd171-2eb3-477c-8894-f23e63e2c3fe · outbound

This paper cites Meshgen: Generating pbr textured mesh with render-enhanced auto-encoder and generative data augmentation.

Video Perception Models for 3D Scene Synthesis Meshgen: Generating pbr textured mesh with render-enhanced auto-encoder and generative data augmentation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:56.569470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:56.569470Z digest=sha256:c4f3c53513346f6d1f50c8c03a4f3306fec85f7a9c7b1a5121c546f192b5c9c7

Observation ef4f1bb2-8ffd-4076-80df-4e906229a473 · outbound

This paper cites Procthor: Large-scale embodied ai using procedural generation.International Conference on Neural Information Processing Systems (NeurIPS), 2022.

Video Perception Models for 3D Scene Synthesis Procthor: Large-scale embodied ai using procedural generation.International Conference on Neural Information Processing Systems (NeurIPS), 2022

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:56.682731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:56.682731Z digest=sha256:d23b1ab1fdacb56b1de9203a81cb87f6772f308bcd767049bea3d4efe790840a

Observation 9c05f9c5-5028-4ebd-b67d-a6e6611bee77 · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

Video Perception Models for 3D Scene Synthesis Objaverse: A universe of annotated 3d objects

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:22.679269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:56.851170Z digest=sha256:5d9f64ce6f92603a3188e2f55bc2558a34acc6d61a821099b103f44a9d77f24b

Observation b4a49ec1-16c5-49d5-84bb-4ea6060af26a · outbound

This paper cites Global-Local Tree Search in VLMs for 3D Indoor Scene Generation.

Video Perception Models for 3D Scene Synthesis Global-Local Tree Search in VLMs for 3D Indoor Scene Generation

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:49:07.416388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:56.973998Z digest=sha256:9d281c852c3469d615aff83acc64997a1cf354871e393b1fe0a100fe77467324

Observation feef66ba-7c26-419a-bac0-599015288700 · outbound

This paper cites Layoutgpt: Compositional visual planning and generation with large language models.International Conference on Neural Information Processing Systems (NeurIPS), 2023.

Video Perception Models for 3D Scene Synthesis Layoutgpt: Compositional visual planning and generation with large language models.International Conference on Neural Information Processing Systems (NeurIPS), 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:22.403723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:57.096462Z digest=sha256:504cd2e4282e97af3a568d1ad00ae4273f781b6d633fc2b2bac0976df9979d9e

Observation ecdce60d-aa25-493c-8423-8ce2e8341c8a · outbound

This paper cites Anyhome: Open-vocabulary generation of structured and textured 3d homes.

Video Perception Models for 3D Scene Synthesis Anyhome: Open-vocabulary generation of structured and textured 3d homes

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:22.109966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:57.253288Z digest=sha256:4fb4a8015f9492087d84baa390ea0721c4f5e5921762d19311498018855f3de4

Observation 7cb0f7e7-7b8c-4138-bfc3-f472b9537ac9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Video Perception Models for 3D Scene Synthesis DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:57.384307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:57.384307Z digest=sha256:4783210cd51725b4c14258e0671c0394c80efbfa74a1c6507074dccefdf2997b

Observation 4b4cf6d6-f117-4f1d-ba02-2685b950ea19 · outbound

This paper cites MaterialMVP: Illumination-Invariant Material Generation via Multi-view PBR Diffusion.

Video Perception Models for 3D Scene Synthesis MaterialMVP: Illumination-Invariant Material Generation via Multi-view PBR Diffusion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:57.574455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:57.574455Z digest=sha256:25b4daa8f46df133bc078f0897b74bc5ed1e34bacc86cc748de9969a9a00ddea

Observation 78f124c1-e2d9-43e5-9c03-bcadbc90086f · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Video Perception Models for 3D Scene Synthesis CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:57.707381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:57.707381Z digest=sha256:a6cc02d6dc41dddd0009f217e60b24b2bc409028b3e12eb14de2e719a22f7740

Observation 3cff6095-1fba-404a-a0ac-36b51fb79992 · outbound

This paper cites Video diffusion models.International Conference on Neural Information Processing Systems (NeurIPS), 2022.

Video Perception Models for 3D Scene Synthesis Video diffusion models.International Conference on Neural Information Processing Systems (NeurIPS), 2022

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:21.819721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:57.877525Z digest=sha256:a3808750929fbbbfdeec3f77ac0308a24198fb7d9823f520c37506b6387d9c90

Observation 00a7e75a-6f1c-4fd5-9c50-c03864799e7c · outbound

This paper cites Text2room: Extracting textured 3d meshes from 2d text-to-image models.

Video Perception Models for 3D Scene Synthesis Text2room: Extracting textured 3d meshes from 2d text-to-image models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:21.568856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:57.986754Z digest=sha256:428418e3a8c8c684dddd06bb6c0a61f829d81257b4c83352f9fcc0292a29b3e4

Observation a9680f05-bb78-44e7-a071-7ef28d9cf7f8 · outbound

This paper cites Material anything: Generating materials for any 3d object via diffusion.

Video Perception Models for 3D Scene Synthesis Material anything: Generating materials for any 3d object via diffusion

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:21.334865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:58.118324Z digest=sha256:e75f9ba1d2c4ab1b51895a4eabf8bebc86eea2042c6579a6fa9e2a7a856c45f9

Observation ee439aa8-fae9-441a-8fc9-ccdef25e4653 · outbound

This paper cites Midi: Multi-instance diffusion for single image to 3d scene generation.

Video Perception Models for 3D Scene Synthesis Midi: Multi-instance diffusion for single image to 3d scene generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:21.045100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:58.214384Z digest=sha256:9394fa2939de540962532afb4ce66bc06ab95142865369edd7d5875249805a99

Observation d16667b1-3237-4af7-8c36-64a367edbc0e · outbound

This paper cites GPT-4o System Card.

Video Perception Models for 3D Scene Synthesis GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:58.317839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:58.317839Z digest=sha256:1f95d8e2d2aa697736343b23b4974bde2f6f261733957bf0b8adf9c0c50706c5

Observation 078e3dd9-fa81-4640-97a6-aaab2b6a09c3 · outbound

This paper cites How Far is Video Generation from World Model: A Physical Law Perspective.

Video Perception Models for 3D Scene Synthesis How Far is Video Generation from World Model: A Physical Law Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:58.420289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:58.420289Z digest=sha256:3f2d51d2671e2a95acea91104dbe5a611efc47b4a7861f810475aeaf40e0ecf6

Observation 67ec5481-40f2-4565-8a04-b6c8369ae152 · outbound

This paper cites Repurposing diffusion-based image generators for monocular depth estimation.

Video Perception Models for 3D Scene Synthesis Repurposing diffusion-based image generators for monocular depth estimation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:20.852164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:58.524678Z digest=sha256:4f0a1bf674b820a123f84b6f80af49ec7e5d0e65f20ace8a28efe0ba5345de5b

Observation 2c8933bd-b7ee-41cc-9dd6-e47e2ff07113 · outbound

This paper cites A new measure of rank correlation.Biometrika, 1938.

Video Perception Models for 3D Scene Synthesis A new measure of rank correlation.Biometrika, 1938

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:20.531539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:58.635530Z digest=sha256:bec7ee17a173432e6b4dbe83ef671c622bfa1fdd89918a253d9baa4e3f70ea11

Observation 4d293d99-2b30-4bc8-b891-6b735ca9077c · outbound

This paper cites Kling, 2025.https://klingai.com/global/.

Video Perception Models for 3D Scene Synthesis Kling, 2025.https://klingai.com/global/

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:20.240539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:58.758859Z digest=sha256:fe8e4ff805c7d177dfe9fd08cd199ac995c7d15ba64345b060ed1a6a3924b6f5

Observation 6d14f9a5-8818-4ef0-87fe-b72cc2882aa8 · outbound

This paper cites Scenecraft: automating interactive narrative scene generation in digital games with large language models.

Video Perception Models for 3D Scene Synthesis Scenecraft: automating interactive narrative scene generation in digital games with large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:19.971505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:58.912878Z digest=sha256:2b4bff5577c8c70b1e25ab67f3ea023b48a35f481debea97f5c26d8d3dd51c98

Observation 7141c404-6bc3-48c1-b05a-402fdbe6a5c0 · outbound

This paper cites Grounding image matching in 3d with mast3r.

Video Perception Models for 3D Scene Synthesis Grounding image matching in 3d with mast3r

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:19.675974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:59.036279Z digest=sha256:eeaa4aaf05647c568a8e0f2161840bffba8072793e4e1197643abb1848dd30bd

Observation 4ba632b1-4fbf-47e6-9a0f-c3944096c9cd · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Video Perception Models for 3D Scene Synthesis Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:19.356793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:59.156100Z digest=sha256:8551ad33d125d84e46bcfefe29a2e299401e5e53997a76489e00aa98892587ea

Observation d1072fe8-b9aa-4365-839f-d2aebb0c6b62 · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Video Perception Models for 3D Scene Synthesis Open-Sora Plan: Open-Source Large Video Generation Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:59.227544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:59.227544Z digest=sha256:c1bf341d014b930e7bb0c7f900708337bcd94ba9cb0247716d6fc3e105564ced

Observation 59efb5a2-0ad3-44cb-ab18-770ff980b27c · outbound

This paper cites Instructscene: Instruction-driven 3d indoor scene synthesis with semantic graph prior.

Video Perception Models for 3D Scene Synthesis Instructscene: Instruction-driven 3d indoor scene synthesis with semantic graph prior

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:19.115134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:59.327193Z digest=sha256:214ef66f43a9d4df1ca9e6d0fe027de99c7e7abc1283271bd2fc2b8782130e36

Observation 4d58fdb4-733c-43ba-ad29-8219903d0c94 · outbound

This paper cites Towards Language-guided Interactive 3D Generation: LLMs as Layout Interpreter with Generative Feedback.

Video Perception Models for 3D Scene Synthesis Towards Language-guided Interactive 3D Generation: LLMs as Layout Interpreter with Generative Feedback

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:49:07.082148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:59.441221Z digest=sha256:fa1d658077873ca79108e2a6ecc293f674ab91bbb759a35bcbec4310cea2f640

Observation 9cd6a9d2-61cd-4bec-b331-04d945f596ba · outbound

This paper cites Evaluating text-to-visual generation with image-to-text generation.

Video Perception Models for 3D Scene Synthesis Evaluating text-to-visual generation with image-to-text generation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:18.780712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:59.588234Z digest=sha256:e8db08ae152b221e7659cba35d32d8c533b3fa9fccdc75f717534e2495927c8d

Observation f97302bc-978c-48e6-a955-c770b23c387b · outbound

This paper cites Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation.

Video Perception Models for 3D Scene Synthesis Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:48:59.729438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:48:59.729438Z digest=sha256:e41f4cfe3353a7c4e8508604a974a2838a2be21a41e9828fd550fe8897d0b045

Observation 096379e4-aa1a-445b-b605-635ee1f33d39 · outbound

This paper cites Visual instruction tuning.International Conference on Neural Information Processing Systems (NeurIPS), 2023.

Video Perception Models for 3D Scene Synthesis Visual instruction tuning.International Conference on Neural Information Processing Systems (NeurIPS), 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:18.438665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:59.883976Z digest=sha256:0afecc88f43034adf8b567509c16ebc75a5c0816acc239c7d08c807921113670

Observation 6774bcdf-5cbf-4eac-9975-b2bfe16c0319 · outbound

This paper cites Dream machine, 2024.https://lumalabs.ai/dream-machine.

Video Perception Models for 3D Scene Synthesis Dream machine, 2024.https://lumalabs.ai/dream-machine

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:18.149801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:48:59.996398Z digest=sha256:914a8a29eebf52bc10401327bd58affb8e9d8eb236e5003a2ad6ddf8534456f5

Observation e592531a-c502-4c4c-9d8b-7e7fee694cc7 · outbound

This paper cites Mimicgen: A data generation system for scalable robot learning using human demonstrations.

Video Perception Models for 3D Scene Synthesis Mimicgen: A data generation system for scalable robot learning using human demonstrations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:00.135090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:00.135090Z digest=sha256:fd253730739125063c5d8ccff1b84aa4c50ee18a26a369649ea18cb39dbd394e

Observation 295ddaec-02f9-4422-97f0-a26d71cc39a1 · outbound

This paper cites Cosmos, 2024.https://www.nvidia.com/en-us/ai/cosmos/.

Video Perception Models for 3D Scene Synthesis Cosmos, 2024.https://www.nvidia.com/en-us/ai/cosmos/

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:17.856800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:00.295561Z digest=sha256:1b871586eaf877d2c53b347eb686828b8a95ebd30672156616eec171019cbefa

Observation 3e1c6d8c-b28b-42b5-9554-3b4058f97021 · outbound

This paper cites Global structure-from-motion revisited.

Video Perception Models for 3D Scene Synthesis Global structure-from-motion revisited

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:17.552287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:00.419618Z digest=sha256:153b2de6c3cc4367306c101c83946e5f95567266e17335366e9726b499266235

Observation b00247bf-c869-472b-ab9c-02a3701d902d · outbound

This paper cites Scalable diffusion models with transformers.

Video Perception Models for 3D Scene Synthesis Scalable diffusion models with transformers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:00.559290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:00.559290Z digest=sha256:589f74823d8a658ee007026ef90715d74392ab5ba9da943259ce43ce27b52960

Observation ca4999cb-614a-40ac-b9f8-3c69bffd5235 · outbound

This paper cites Unidepth: Universal monocular metric depth estimation.

Video Perception Models for 3D Scene Synthesis Unidepth: Universal monocular metric depth estimation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:17.262660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:00.674905Z digest=sha256:2366727e78af705cd99a25b59bba441a08fda9685186001b9dc71caa9300dd7c

Observation 69d131c1-9a2f-4024-931c-00a31e58f5e8 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Video Perception Models for 3D Scene Synthesis SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:00.817729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:00.817729Z digest=sha256:e0b82472a0ccf392dc93cf11d5dffec09127b4aa85f671e695bf67d186bffa11

Observation 42dae668-ca86-4684-8a33-165a554bc651 · outbound

This paper cites Hsm: Hierarchical scene motifs for multi-scale indoor scene generation.arXiv preprint arXiv:2503.16848, 2025.

Video Perception Models for 3D Scene Synthesis Hsm: Hierarchical scene motifs for multi-scale indoor scene generation.arXiv preprint arXiv:2503.16848, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:00.962603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:00.962603Z digest=sha256:34ec8e551e7be2806abe5b15e5188303adbf8dce8775510c7e0f03858710ec25

Observation 2b40e8a6-715c-4530-8ba2-fb2be1579047 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Video Perception Models for 3D Scene Synthesis Learning transferable visual models from natural language supervision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:01.094481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:01.094481Z digest=sha256:72659e84537d4f2afa19809873c4a9d190ae305e885db1056d6fcea8da5d1490

Observation c5265f9a-dae3-45be-8446-8f307fc5da0b · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

Video Perception Models for 3D Scene Synthesis Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:01.255095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:01.255095Z digest=sha256:7e0406ded783dc466b305373ffd578a906e8d2f1e802cac13c76900741fc8bf4

Observation 7e69a5c9-7c99-463c-a1ea-ead24b717ccc · outbound

This paper cites Gen 3, 2024.https://runwayml.com/research/introducing-gen-3-alpha.

Video Perception Models for 3D Scene Synthesis Gen 3, 2024.https://runwayml.com/research/introducing-gen-3-alpha

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:16.976271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:01.452576Z digest=sha256:dac60ca0639233b2e650e45f01749596adefdd547e5400e0764dcefeb8dfedc0

Observation 81d745df-128b-4d59-a902-398fba294688 · outbound

This paper cites Temporal generative adversarial nets with singular value clipping.

Video Perception Models for 3D Scene Synthesis Temporal generative adversarial nets with singular value clipping

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:16.774408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:01.582860Z digest=sha256:2d64985e1ff218b5b92c7c5853c272064f06651f0c1f5c029e58b6a11eb6f497

Observation 89e92793-3925-4390-b103-ad26a0591667 · outbound

This paper cites Structure-from-motion revisited.

Video Perception Models for 3D Scene Synthesis Structure-from-motion revisited

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:16.410380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:01.723523Z digest=sha256:0cdecbb6bf4d2a798192dc3a20a014422468a83c862e221c2180799782f92692

Observation 8d5c0edc-ea20-4a5a-a728-4abda4406df2 · outbound

This paper cites Pixelwise view selection for unstructured multi-view stereo.

Video Perception Models for 3D Scene Synthesis Pixelwise view selection for unstructured multi-view stereo

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:01.844001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:01.844001Z digest=sha256:c475f5a3b493f0fbc4d43d97f18b18caedad351f0eccb8b67cdce0e785d560b0

Observation 789763b4-cbdb-47ad-a6b8-4e1174f40558 · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation.

Video Perception Models for 3D Scene Synthesis Mask3d: Mask transformer for 3d semantic instance segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:16.118056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:01.963818Z digest=sha256:c2eee80ab680657f8474f50b1f5a21622d2a0b323bebe75337eafdb0039bd888

Observation 7fc4bf38-d6d7-4613-ab77-9aa48e64cb9c · outbound

This paper cites Controlroom3d: Room generation using semantic proxy rooms.

Video Perception Models for 3D Scene Synthesis Controlroom3d: Room generation using semantic proxy rooms

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:15.738656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:02.073504Z digest=sha256:2f22bd114586fd4399fd95d943beb1028b1f847de1862b13b53de30350b01882

Observation ffaa7004-b558-46a1-8dd9-ec650a2b8439 · outbound

This paper cites LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models.

Video Perception Models for 3D Scene Synthesis LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:02.194039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:02.194039Z digest=sha256:65c022c3cb2f33573288f0114c1f8b6b5d08d57d646fe06b7d32bbe1ea76d5e4

Observation b84f749e-fb70-46cb-8c54-c7f92a46caf8 · outbound

This paper cites Neuralrecon: Real-time coherent 3d reconstruction from monocular video.

Video Perception Models for 3D Scene Synthesis Neuralrecon: Real-time coherent 3d reconstruction from monocular video

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:15.360158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:02.368698Z digest=sha256:ebf4a4c526b0d6fc3e9f3b7df54ab6ae9b14396a2dce2e295a65fe4b919f19e7

Observation 086cb0cb-ae5a-41e2-a28a-a526b7fd10ed · outbound

This paper cites Retargetable ar: Context-aware aug- mented reality in indoor scenes based on 3d scene graph.

Video Perception Models for 3D Scene Synthesis Retargetable ar: Context-aware aug- mented reality in indoor scenes based on 3d scene graph

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:15.023686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:02.516076Z digest=sha256:4451750328fe3b12d6d226af03888e4968bdc369f618766df4763d06d4702863

Observation 15581ad0-f09c-490a-8f83-0bf3bc29b6dc · outbound

This paper cites Diffuscene: Denoising diffusion models for generative indoor scene synthesis.

Video Perception Models for 3D Scene Synthesis Diffuscene: Denoising diffusion models for generative indoor scene synthesis

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:14.671189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:02.639098Z digest=sha256:40ab1f139f6175a74ddf471287a845cbbe6e75ad9c1e188cc8f29bc765d233b7

Observation 412170fa-a4a6-4c48-a97f-e1af8e75ce5e · outbound

This paper cites Mv-dust3r+: Single-stage scene reconstruction from sparse views in 2 seconds.

Video Perception Models for 3D Scene Synthesis Mv-dust3r+: Single-stage scene reconstruction from sparse views in 2 seconds

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:14.292896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:02.747210Z digest=sha256:b79384eb9ed4c28b32e4737c9c1990832217edaf48c24c0e58d15cc2ea5d4a94

Observation b2ac455d-8a75-4cde-92bf-4e366d6ead53 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Video Perception Models for 3D Scene Synthesis Gemini: A Family of Highly Capable Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:02.887877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:02.887877Z digest=sha256:66859a3dda91f93741fd3130e0f74f151634ceeecd0f974d914f57dbb924e970

Observation 5d41ce81-9ff0-4809-94c6-83bb2827a7e0 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Video Perception Models for 3D Scene Synthesis Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:03.027631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:03.027631Z digest=sha256:43b59e682279f6c50e6128abc20d36fa1fc2c2b6a3cf44e099f883315668ed42

Observation 79c1c8c8-a465-43f7-9129-6a2ecadb1a76 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Video Perception Models for 3D Scene Synthesis LLaMA: Open and Efficient Foundation Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:03.173489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:03.173489Z digest=sha256:31fee948ac255a59cecb6f5bdee420c303d2c50deebb60abbfebbb7c7009e412

Observation c2ab6f0c-d7d5-41e9-92c9-f6993f74b811 · outbound

This paper cites Mocogan: Decomposing motion and content for video generation.

Video Perception Models for 3D Scene Synthesis Mocogan: Decomposing motion and content for video generation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:13.997067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:03.287796Z digest=sha256:eb0adc080bd79e4f325a22ccc578e944a88f161dd7f3deafe431bf7cd7482471

Observation 25e8da4d-53fb-41ee-8ed8-14c1cb5d1f42 · outbound

This paper cites 3D Reconstruction with Spatial Memory.

Video Perception Models for 3D Scene Synthesis 3D Reconstruction with Spatial Memory

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:03.416261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:03.416261Z digest=sha256:a9c61b04e88131c1d1510d539b6293371bd83002fd57c8da03e9c82058e1ae7b

Observation 0c142265-b44a-45e5-b3ae-9a325d739254 · outbound

This paper cites Vggt: Visual geometry grounded transformer.

Video Perception Models for 3D Scene Synthesis Vggt: Visual geometry grounded transformer

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:13.689544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:03.528516Z digest=sha256:cb60113d745be9983bc0c7278c5ede0874215ae0971297fdf52539e956771406

Observation 1f456c54-5401-43d5-8b1f-f57ab4488a20 · outbound

This paper cites Dust3r: Geometric 3d vision made easy.

Video Perception Models for 3D Scene Synthesis Dust3r: Geometric 3d vision made easy

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:13.250785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:03.673884Z digest=sha256:a6b2d6d790b52c8468e66ccb2abe394ed374b389f357a422eef765983dbe7062

Observation 96a772b7-3716-4c54-84a7-64f6d6cb3b80 · outbound

This paper cites Architect: Generating vivid and interactive 3d scenes with hierarchical 2d inpainting.International Conference on Neural Information Processing Systems (NeurIPS), 2025.

Video Perception Models for 3D Scene Synthesis Architect: Generating vivid and interactive 3d scenes with hierarchical 2d inpainting.International Conference on Neural Information Processing Systems (NeurIPS), 2025

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:12.902837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:03.792935Z digest=sha256:4e7cfa577b6b8d6f8045943f6730aa5a734e6addc0ca0abe564d09591546775f

Observation 3cc91a60-96f1-4f1b-a8a4-62860ef561d0 · outbound

This paper cites Gpt-4v (ision) is a human-aligned evaluator for text-to-3d generation.

Video Perception Models for 3D Scene Synthesis Gpt-4v (ision) is a human-aligned evaluator for text-to-3d generation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:12.607199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:03.921117Z digest=sha256:bda33dcb573e2ee9c3ab26edc616209570806d91b5cc5906dc5666407e98d14e

Observation 267e7927-cf2e-4a8b-8803-652185a80bdb · outbound

This paper cites Amodal3R: Amodal 3D Reconstruction from Occluded 2D Images.

Video Perception Models for 3D Scene Synthesis Amodal3R: Amodal 3D Reconstruction from Occluded 2D Images

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:04.080421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:04.080421Z digest=sha256:461b17bfa613565b4154b849170fdf4b0ba35b6eb885c1d58f81a8fc863b9665

Observation 953f9aa6-e44b-4d39-ad75-7e537b631fd2 · outbound

This paper cites Structured 3D Latents for Scalable and Versatile 3D Generation.

Video Perception Models for 3D Scene Synthesis Structured 3D Latents for Scalable and Versatile 3D Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:04.242789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:04.242789Z digest=sha256:ba42e5ad7c79abb533e1ac9a20ead6784e09ebdf4c261ec838e7e4c54afc1e82

Observation b88f15fa-60cc-4598-b9b8-4c66dd999aa3 · outbound

This paper cites Fast3r: Towards 3d reconstruction of 1000+ images in one forward pass.

Video Perception Models for 3D Scene Synthesis Fast3r: Towards 3d reconstruction of 1000+ images in one forward pass

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:12.305027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:04.360714Z digest=sha256:f71992c0afb80dd426de5551a340ff038151bb444c6234b9d4ba727735e040b0

Observation daeac370-867f-483e-a321-6a35fedc94de · outbound

This paper cites Diffusion probabilistic modeling for video genera- tion.Entropy, 2023.

Video Perception Models for 3D Scene Synthesis Diffusion probabilistic modeling for video genera- tion.Entropy, 2023

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:12.034297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:04.474716Z digest=sha256:f36acf992bd443db4ed0500a10b3d49416c0be1cf2e30941f00337e840500eb7

Observation ac31c551-4d32-44eb-90ae-00ab253b091e · outbound

This paper cites Hunyuan3D 1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation.

Video Perception Models for 3D Scene Synthesis Hunyuan3D 1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:04.592991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:04.592991Z digest=sha256:7c87a049996762a84054381ba8dcf238f563974afcf90289b4d3e88abe0c70ff

Observation 3db59cb8-b271-4f4c-ba32-e6c7e1eba7d2 · outbound

This paper cites Physcene: Physically interactable 3d scene synthesis for embodied ai.

Video Perception Models for 3D Scene Synthesis Physcene: Physically interactable 3d scene synthesis for embodied ai

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:11.719868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:04.724919Z digest=sha256:2c5d6d53eeb782767378f77b5fc11d90822dc0997eb47d2c60ccb2a799c3da3f

Observation dc169188-0774-49f1-8a28-75310eaaa968 · outbound

This paper cites LLplace: The 3D Indoor Scene Layout Generation and Editing via Large Language Model.

Video Perception Models for 3D Scene Synthesis LLplace: The 3D Indoor Scene Layout Generation and Editing via Large Language Model

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:04.838427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:04.838427Z digest=sha256:0431ddaae0460a01b63da8cd0ece14d0d5f1186ff90bf58d034e9dcf0fd7eb17

Observation 6c00a28b-971f-4fc7-8231-84e460b7805b · outbound

This paper cites Holodeck: Language guided generation of 3d embodied ai environments.

Video Perception Models for 3D Scene Synthesis Holodeck: Language guided generation of 3d embodied ai environments

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:11.385857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:04.950955Z digest=sha256:294d015607590f618e0f90d14992be9758452f3745699b06fe2e9d15954bf912

Observation 268c505b-4a19-4435-86ff-a130797b150f · outbound

This paper cites Mmgdreamer: Mixed-modality graph for geometry-controllable 3d indoor scene generation.

Video Perception Models for 3D Scene Synthesis Mmgdreamer: Mixed-modality graph for geometry-controllable 3d indoor scene generation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:11.143907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:05.048179Z digest=sha256:2a7435e9730a65feea151d70242cb1e00a9d04e6cf21304b5c1e3cfb5b2777e5

Observation 160a7dba-0b69-4ef1-b15d-f6f88ed94274 · outbound

This paper cites Commonscenes: Generating commonsense 3d indoor scenes with scene graphs.

Video Perception Models for 3D Scene Synthesis Commonscenes: Generating commonsense 3d indoor scenes with scene graphs

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:10.828072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:05.114934Z digest=sha256:38e54e26f00982bf34f57b7f93eca601a96b0711bd133e7cccd6926a9d0d1432

Observation 16f1fc80-347d-499c-b0c5-30549f0ee919 · outbound

This paper cites Sg-bot: Object rearrangement via coarse-to-fine robotic imagination on scene graphs.

Video Perception Models for 3D Scene Synthesis Sg-bot: Object rearrangement via coarse-to-fine robotic imagination on scene graphs

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:10.523588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:05.197964Z digest=sha256:1611cb484d051114ecc5b290cc759e061c2bbc124d39316ae52b4bff24fa0dfc

Observation 62ce0091-5e36-4a11-b0cb-8ecf6ee4d1dc · outbound

This paper cites Echoscene: Indoor scene generation via information echo over scene graph diffusion.

Video Perception Models for 3D Scene Synthesis Echoscene: Indoor scene generation via information echo over scene graph diffusion

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:10.260467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:05.283459Z digest=sha256:c436a73eab62d7eb1253ce2ff5bee87ca98abc8638151fcbc2f71d1f13ec97a0

Observation 0922fa1d-eaa2-4702-a525-287dbaba9062 · outbound

This paper cites Fast and robust iterative closest point.Transactions on Pattern Analysis and Machine Intelligence (PAMI), 2021.

Video Perception Models for 3D Scene Synthesis Fast and robust iterative closest point.Transactions on Pattern Analysis and Machine Intelligence (PAMI), 2021

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:09.952923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:05.360049Z digest=sha256:8f4a988023404227c9cc807d31a6b6477c3cf28bd69bd6d9506640e7df7a5a0f

Observation 6f03eb44-d4e8-4f61-adb6-16029f73e40c · outbound

This paper cites Dreammat: High-quality pbr material generation with geometry-and light-aware diffusion models.ACM Transactions on Graphics (TOG), 2024.

Video Perception Models for 3D Scene Synthesis Dreammat: High-quality pbr material generation with geometry-and light-aware diffusion models.ACM Transactions on Graphics (TOG), 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:09.580841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:05.427830Z digest=sha256:14fa896706e579d38cd3d0abd1bf0741052a8cd6edfb707b6dca5bbd8f432590

Observation f7278319-137d-4168-a834-36f490f26e3b · outbound

This paper cites Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation.

Video Perception Models for 3D Scene Synthesis Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:05.537270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:05.537270Z digest=sha256:51372d4af9e4fd2b8230ec73869acc62199b517b749d544830e6d65a3d3029bf

Observation 664c9b8d-1ed5-4aa3-827f-5f4adb60627b · outbound

This paper cites A bedroom with a large bed, two nightstands, a floor lamp, a wardrobe, and a big window.

Video Perception Models for 3D Scene Synthesis A bedroom with a large bed, two nightstands, a floor lamp, a wardrobe, and a big window

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:09.281404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:05.600034Z digest=sha256:b47f9cf4049c25af5f03d6f2720eca0a4b2c1666379118c4f92c6adf42985651

Observation 08229b1f-c9b6-4f00-b24b-2313b132b9bf · outbound

This paper cites realistic.

Video Perception Models for 3D Scene Synthesis realistic

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:08.607372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:05.855801Z digest=sha256:4d7a9c1f15ff859691697e69a90c09f4ea6075ac1cfe4adae250542e5e90d246

Observation 6334d86f-3f18-4d4f-b935-3219c1f75f4d · outbound

This paper cites an unresolved cited work.

Video Perception Models for 3D Scene Synthesis Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:49:08.417324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:05.931372Z digest=sha256:82303c8c7b575cf4cedd3d1ec8747d49ee85794c5d766585c254679207d4b132

Observation 5b21a4d5-b83e-47ec-be48-517a76730fd3 · outbound

This paper cites an unresolved cited work.

Video Perception Models for 3D Scene Synthesis Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:49:08.143616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:05.991925Z digest=sha256:696d5c6a42325b00245d5b3340994c19de8f128f4146bdc57368e2d700a0ddc4

Observation ff73c05a-d407-4b98-8a60-9cb47410ace4 · outbound

This paper cites Final answer: The first one: x x x The second one: x x x The third one: x x x (where x denotes ranks 1–3) (Please strictly follow the format above.

Video Perception Models for 3D Scene Synthesis Final answer: The first one: x x x The second one: x x x The third one: x x x (where x denotes ranks 1–3) (Please strictly follow the format above

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:07.871415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:06.094151Z digest=sha256:ee2c028c4b9c57ef88b29e604328fac08eaaf88bb2c7c3044508efc8d6056fda

Observation 2e09f92e-d940-4554-8091-04fb0df7f51c · outbound

This paper cites an unresolved cited work.

Video Perception Models for 3D Scene Synthesis Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:49:09.032891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:06.231711Z digest=sha256:9e9742e8ca9b0596081639d816115aabe93cbd5e88cb5addf2433dfcb86af555

Observation ab8ab2b6-81c9-421b-9e28-9815bbcd9c88 · outbound

This paper cites Consider object positions, orientations, and user convenience.

Video Perception Models for 3D Scene Synthesis Consider object positions, orientations, and user convenience

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:49:08.822295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:06.317452Z digest=sha256:4f389bf41019e26d01ccbf391c22ed0c75c0a8e3a0f5eef21bfb36e1d19730b5

Observation 84953c7e-408a-4611-9eff-c406841802cc · outbound

This paper cites an unresolved cited work.

Video Perception Models for 3D Scene Synthesis Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:49:07.649512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:49:06.433344Z digest=sha256:f98bb4bae97cbc9f3e88d52a9ed4ef1892a3b5dad02aff178553210b6a6ad7de

Pith citing papers

No inbound Pith citation observations are available.