Pith. sign in

Paper Citation Record · LEDGER

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

As of 20 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 4 inbound Pith citation observations for arXiv:2412.03558.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03558 v3

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:19:24.924471Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:57.655659Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T20:03:57.298307Z

Reference resolution

86 of 86 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e235fb43-7527-4703-8188-ee4147644d52 · outbound

This paper cites Neural rgb-d surface reconstruction.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Neural rgb-d surface reconstruction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.470319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.470319Z digest=sha256:08a6f5168d146ba280a086812a07488daffd2a269e37e4c21d154a2919ac4154

Observation 6e2bcb3f-4e3d-4f15-9319-b75e40814483 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.477120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.477120Z digest=sha256:74a2341e9e2a8f4f56e376dba442376c3cd3317b2d11f5fa3d8e0ef154bddb46

Observation a26b3620-e4d2-481e-ae07-4d385608d747 · outbound

This paper cites Matterport3D: Learning from RGB-D Data in Indoor Environments.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Matterport3D: Learning from RGB-D Data in Indoor Environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.483193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.483193Z digest=sha256:fdbc51fddc388ee913e0d9ece09a6c1c2a024c3ac11eec1d57c555ebb0ad8364

Observation 1e37fa38-5f5e-4003-808c-a92f919788e9 · outbound

This paper cites Single-view 3d scene reconstruc- tion with high-fidelity shape and texture.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Single-view 3d scene reconstruc- tion with high-fidelity shape and texture

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.489199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.489199Z digest=sha256:7fdeaef9a4cfdb40e35f40979524ed42e0e5283267db1b66bcbbad763b4106d0

Observation 8f6e5658-3c90-42d0-b9d5-457b21d44953 · outbound

This paper cites ComboVerse: Compositional 3D Assets Creation Using Spatially-Aware Diffusion Guidance.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation ComboVerse: Compositional 3D Assets Creation Using Spatially-Aware Diffusion Guidance

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.494401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.494401Z digest=sha256:38dd20e99148acfca1fd2f27a0659b059a0fe57abbfbc11bbf2e4138f4f8e030

Observation c02f3fd2-ad73-4121-8653-e66339e13dda · outbound

This paper cites Buol: A bottom-up framework with occupancy-aware lifting for panoptic 3d scene reconstruction from a single image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Buol: A bottom-up framework with occupancy-aware lifting for panoptic 3d scene reconstruction from a single image

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.500289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.500289Z digest=sha256:838845eca07e4eccfac16b828a87e9aaf6b85b312e24269fc50a79a3ac89f413

Observation 239bf3cb-c715-4bd7-9433-650664bdb1d0 · outbound

This paper cites Panoptic 3d scene reconstruction from a single rgb image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Panoptic 3d scene reconstruction from a single rgb image

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.506633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.506633Z digest=sha256:975ed109a6a71352bdc4f9a1f261992320442a218b4fe1c08be324e9808920c4

Observation 09b903b6-84d8-4024-975f-5276c71ff7d4 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.512113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.512113Z digest=sha256:ea532ae57be40f3a8ed9e610011a8b11a353c0b3b5157883eb30437b7a3b740f

Observation 94ab5b39-77d7-4abd-aa7f-914bbc3bc830 · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Objaverse: A universe of annotated 3d objects

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.517348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.517348Z digest=sha256:e838cf70953927782131add08fdeb79146109cb20bd802a720895b78b27fa8ff

Observation 1905d084-a603-4c44-bcba-fe7c466462dc · outbound

This paper cites Objaverse-xl: A universe of 10m+ 3d objects.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Objaverse-xl: A universe of 10m+ 3d objects

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.257528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.523049Z digest=sha256:27f37f3da18ea6bee7c2b1ff16fc32177065a1c0edbfb39a2fddc60886cf7be8

Observation 1228f12e-7aaa-4928-a7dd-409e7e3bbfcb · outbound

This paper cites Gen3DSR: Generalizable 3D Scene Reconstruction via Divide and Conquer from a Single View.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Gen3DSR: Generalizable 3D Scene Reconstruction via Divide and Conquer from a Single View

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.527729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.527729Z digest=sha256:8cd2700f2c445d719b920c1effe74215d6e18801300c221e5d0ec49fe44a174e

Observation c68931e9-63d4-46ab-ba3e-2eec7022ddd0 · outbound

This paper cites Tela: Text to layer-wise 3d clothed human generation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Tela: Text to layer-wise 3d clothed human generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.532509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.532509Z digest=sha256:c16ebacfbeea6263362faa8285b54f2c0121aa5c0583d2cc04ac9818a2fd306c

Observation ba352958-b5c6-444b-a471-1507a2a77061 · outbound

This paper cites Omnidata: A scalable pipeline for making multi- task mid-level vision datasets from 3d scans.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Omnidata: A scalable pipeline for making multi- task mid-level vision datasets from 3d scans

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.538986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.538986Z digest=sha256:7f209a5192db826e4d34c597983b371ad1041c25232573a885819c177f9edff3

Observation 595beabd-0d2e-40fb-8395-fd5b260a82c8 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.544028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.544028Z digest=sha256:fe12a19fd9c91e9ef98a6a9fcb86993706880f4a6004712569c5b7b99c6da1d5

Observation 034999a5-8e4a-43a0-8fb2-22483d787c19 · outbound

This paper cites 3d-front: 3d furnished rooms with layouts and semantics.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation 3d-front: 3d furnished rooms with layouts and semantics

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.548492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.548492Z digest=sha256:c97354d0e2909076e416ffdefcad03c43f7329dea821e70ce6b8f3c6c6cfb5a8

Observation 5c6d46ed-3de8-4569-956d-101af54ab5ac · outbound

This paper cites 3d-future: 3d fur- niture shape with texture.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation 3d-future: 3d fur- niture shape with texture

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.192855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.553425Z digest=sha256:d0ffff0e5e43786816b82536369466776c0680554baa4a22cace79bf4b26c5ee

Observation aba27113-874c-4f8d-9746-34ba865c5f5b · outbound

This paper cites Diffcad: Weakly-supervised probabilistic cad model retrieval and alignment from an rgb image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Diffcad: Weakly-supervised probabilistic cad model retrieval and alignment from an rgb image

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.174826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.559147Z digest=sha256:b924408352794b7939bd69dbe8c909e928e004898c263ae0cfffaaf00b2aa39d

Observation a063ab1c-879f-4b75-9296-6f0737c091c9 · outbound

This paper cites Learn- ing 3d object shape and layout without 3d supervision.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Learn- ing 3d object shape and layout without 3d supervision

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.156177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.565160Z digest=sha256:c3c1d036d94cf6b0c1c1786ba7ae8003f58f565d486dc69b63506a24add6d10e

Observation 3ad51016-5d78-4210-bcce-c97abcc04d6b · outbound

This paper cites Roca: Ro- bust cad model retrieval and alignment from a single image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Roca: Ro- bust cad model retrieval and alignment from a single image

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.136435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.570414Z digest=sha256:009e22cb0034fa1dcba6a2f6abba4430507af1ad2ae2726c9095fc33f69a772a

Observation 7d3f575f-4163-4bcf-adc6-c83283dab3c2 · outbound

This paper cites threestudio: A unified framework for 3d content generation, 2023.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation threestudio: A unified framework for 3d content generation, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.118486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.575723Z digest=sha256:af9165a3f7eb9ef9c04c4b3dcfaed6b7ae3bdb51b85dbc47df6db48cd4f3a9a3

Observation 3d281103-59ca-4932-ac68-e7c9aa24e9bd · outbound

This paper cites REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout Alignment.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout Alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.580426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.580426Z digest=sha256:0984580c86c7a839ad84daedeaa9d7ec59574be6e7fb8094fab837a38ca6082f

Observation 5fa7230e-a9e5-43f4-99e9-d81214e9651e · outbound

This paper cites Classifier-Free Diffusion Guidance.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Classifier-Free Diffusion Guidance

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.585916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.585916Z digest=sha256:39ace87b127f5c05641149fdbc4266f6c0ee45fbd98748dc47bf5ebc282ebbc8

Observation 69933ff9-1972-4ba0-befc-fd8bbbb83f8f · outbound

This paper cites Denoising dif- fusion probabilistic models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Denoising dif- fusion probabilistic models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.592499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.592499Z digest=sha256:219f21a4c8902d2b1bb6ff1db54b112f8a0b8656cb9d32847deec8947ca82261

Observation 5d2bf12b-69e0-4362-b26e-f0b1ed265da8 · outbound

This paper cites LRM: Large Reconstruction Model for Single Image to 3D.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation LRM: Large Reconstruction Model for Single Image to 3D

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.598267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.598267Z digest=sha256:f7fea7cb0b26614b22c5ee6b2e21c1aeea871194aaf809e3b0da754b31ba1a6c

Observation 2018dd2e-4c69-4d98-9878-fcb0e982f5fe · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation LoRA: Low-Rank Adaptation of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.603228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.603228Z digest=sha256:3eb3d7307d3fff1fb8c8a8fdd8d6cbf488c946e20802657554dd407f78fe51dd

Observation 96c46aaf-8813-4829-8cc1-e09b9da2b740 · outbound

This paper cites MV-Adapter: Multi-view Consistent Image Generation Made Easy.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation MV-Adapter: Multi-view Consistent Image Generation Made Easy

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.608784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.608784Z digest=sha256:db73c453bd01bb2e518645c051e389e324137f50f387b397c978261be131fbab

Observation 3097edca-607d-4802-b4d4-0549bfba768e · outbound

This paper cites Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.613872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.613872Z digest=sha256:48882def10c5766ea2d8948d2524189173b4da655dff5cd52d285eb021e7d89d

Observation 36675878-c973-4ca6-8952-84f4c963f89c · outbound

This paper cites an unresolved cited work.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:19:26.078879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.618869Z digest=sha256:017bef331fd29e4496728a5d3964c76361d92fea4a6bfa7b3b210425417358cd

Observation 7109ec4f-ecb9-494d-8c84-0b40e7bc6c64 · outbound

This paper cites Shap-E: Generating Conditional 3D Implicit Functions.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Shap-E: Generating Conditional 3D Implicit Functions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.624279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.624279Z digest=sha256:1369f2416c96cefa5d2773ccdf92fefac7189e243af157cb845899c56b91149f

Observation bb96c0ff-5170-430f-9afd-99b179c31daf · outbound

This paper cites Auto-Encoding Variational Bayes.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Auto-Encoding Variational Bayes

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.629396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.629396Z digest=sha256:1b5c27cf7718d00d219e5849c5c1031e54a52a2117775aa0cf875e4897c440a5

Observation 6c3d10de-360d-47b4-8c78-afe325186f7a · outbound

This paper cites Segment Anything.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Segment Anything

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.634421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.634421Z digest=sha256:cadaef0dd35622ee91ecda451d698f1ddcb83b29e18a286ab2492455101004fd

Observation 54f7fe89-3a06-41ea-b913-8f226b7e47b6 · outbound

This paper cites Mask2cad: 3d shape prediction by learning to segment and retrieve.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Mask2cad: 3d shape prediction by learning to segment and retrieve

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.062304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.639549Z digest=sha256:8b7d7634cc2ac0045a9d15f033c409dfc3a4d2f7c49e92d3f6a70a9bf405cfb5

Observation 6a93a930-8502-4b47-854d-d93ef3fb6723 · outbound

This paper cites Patch2cad: Patchwise embedding learning for in-the- wild shape retrieval from a single image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Patch2cad: Patchwise embedding learning for in-the- wild shape retrieval from a single image

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.040938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.645115Z digest=sha256:1ce861aa2813018666a4af3af5abe102d97ab23262401cfa2b1409e0cc9687a3

Observation 4f859191-a6a3-40d4-b8d9-8bccf58e4bd8 · outbound

This paper cites SPARC: Sparse Render-and-Compare for CAD model alignment in a single RGB image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation SPARC: Sparse Render-and-Compare for CAD model alignment in a single RGB image

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.651197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.651197Z digest=sha256:dc725d6e6f5406987e7916d6280c9b760ba356d964efb005e91d1269b5cde2d0

Observation 75e3988e-fddb-49f1-b8ac-d10803d98fcc · outbound

This paper cites CraftsMan3D: High-fidelity Mesh Generation with 3D Native Generation and Interactive Geometry Refiner.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation CraftsMan3D: High-fidelity Mesh Generation with 3D Native Generation and Interactive Geometry Refiner

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.657182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.657182Z digest=sha256:ea5f484185b783e80f652cf90118ed1dd67126ad8d5554591a54b793cc06836a

Observation 175e174b-e84c-42c9-9015-f4f6c702b2f2 · outbound

This paper cites TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.662780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.662780Z digest=sha256:644a4e690b1c3ec0ea918376a85bbf4f97a4773e9a90a4aecf477e536a83ddd4

Observation 4e14715a-3bd3-4850-a60d-10d45d5eac70 · outbound

This paper cites Part123: part-aware 3d reconstruction from a single-view image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Part123: part-aware 3d reconstruction from a single-view image

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.024483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.668803Z digest=sha256:5eb4f0505aa5974f48de82444c15b5466c08cdbfd611c694d69d75cced5c2a42

Observation 98b75164-459c-48c0-9671-a6fdb90a1060 · outbound

This paper cites Towards high-fidelity single-view holistic reconstruction of indoor scenes.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Towards high-fidelity single-view holistic reconstruction of indoor scenes

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:26.006488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.673792Z digest=sha256:c0b1b95f7b2611d808666ce570be720d56b658938beb5ac78d553e679758461d

Observation d0fb7e2a-d33e-4a7b-8772-ac2af0b768ca · outbound

This paper cites One-2-3-45++: Fast single im- age to 3d objects with consistent multi-view generation and 3d diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation One-2-3-45++: Fast single im- age to 3d objects with consistent multi-view generation and 3d diffusion

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.990012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.678786Z digest=sha256:9d76c4b5fe7da87dcb2cff8fc6b299d89e506bb286192adf91c5159dfc939790

Observation cda20c72-bd78-4aab-a7fe-d0a97c94dbd3 · outbound

This paper cites One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimiza- tion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimiza- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.973929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.683680Z digest=sha256:f08601630286093f981d180b1dab69db21e7382fccd80cd79e42b731d9e2055e

Observation 47648afb-9b6b-4fc5-8461-efd9a8bac810 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.688612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.688612Z digest=sha256:18308fbfec0b3d38b14635ff055d727f38b3e341aae97d1499d6faffac637520

Observation 13d820e1-6bc3-44ec-9861-0c422768086e · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.694510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.694510Z digest=sha256:a47f8430471221beed30a12754950af8950a0918dfaeb8f3efddffc6f1bd8aa2

Observation e927f0aa-afb3-4d16-a9d1-5d3f42350950 · outbound

This paper cites SyncDreamer: Generating Multiview-consistent Images from a Single-view Image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation SyncDreamer: Generating Multiview-consistent Images from a Single-view Image

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.700568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.700568Z digest=sha256:10e32e4dec44e133de34f33f60ae2f49c72ff8746665ffe22c7a9c25de33ad7a

Observation ddea7528-a219-40ec-ba4c-d59833c031cc · outbound

This paper cites Wonder3d: Sin- gle image to 3d using cross-domain diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Wonder3d: Sin- gle image to 3d using cross-domain diffusion

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.706386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.706386Z digest=sha256:83213885f8ecb537199bcecd0676ae4a0615a9a4f8699c44cb6f6946a8a22bf1

Observation 019a47af-dc24-4573-902e-3ea9154f3954 · outbound

This paper cites Marching cubes: A high resolution 3d surface construction algorithm.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Marching cubes: A high resolution 3d surface construction algorithm

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.946424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.711368Z digest=sha256:c122b10bdb8fb9d9db03eba9da2178c1009e8c9984de41c7e0908fec111e15a7

Observation 666a8032-ff54-4e60-b1a5-b0abaec7e79d · outbound

This paper cites LT3SD: Latent Trees for 3D Scene Diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation LT3SD: Latent Trees for 3D Scene Diffusion

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.716309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.716309Z digest=sha256:2925571fa0847bd6858a63bd6e3ecdded71e649033c36af7236541394b1a856a

Observation e9f1308a-1369-420f-a75d-8593b327dbac · outbound

This paper cites GLIDE: towards photorealis- tic image generation and editing with text-guided diffusion models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation GLIDE: towards photorealis- tic image generation and editing with text-guided diffusion models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.928013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.721777Z digest=sha256:0606375f0cdfc48d71bed99e8acefb2013a6000d0d612f6c43e940330f591a1d

Observation 54ec66ed-b173-448b-b89e-1c23a5b7e666 · outbound

This paper cites Total3dunderstanding: Joint lay- out, object pose and mesh reconstruction for indoor scenes from a single image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Total3dunderstanding: Joint lay- out, object pose and mesh reconstruction for indoor scenes from a single image

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.910886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.727319Z digest=sha256:5f3b31acf7994f338a4303090015ed175803fcf9bbbe88f0e549ae80a57f7582

Observation 0c0282a3-f776-4280-8189-4e258e7b3018 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.731797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.731797Z digest=sha256:4ad22d377cab9057dc088974204cbdf677a79ad6fd0010cfd2e3932348e8246d

Observation e24e5a4d-0044-4f4d-a413-796cca81d393 · outbound

This paper cites Atiss: Autoregres- sive transformers for indoor scene synthesis.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Atiss: Autoregres- sive transformers for indoor scene synthesis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.737191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.737191Z digest=sha256:c77005591a4a06b298fedc10a43eb72126c27a0ac7800ba2f1f11d953fc86e4a

Observation 66b52f76-11e1-42e9-89e2-5415084b1480 · outbound

This paper cites Scalable diffusion models with transformers.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Scalable diffusion models with transformers

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.742085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.742085Z digest=sha256:9f63baae09f61c11a9868d328a3b8bba217bea715a74161cc117eff329a9a973

Observation b8817661-99de-421b-b153-8fd0e05d1cf8 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.746930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.746930Z digest=sha256:f3b40269d4b73e02072138699f65e5c3940d5e575bb6a8a94f3a278052e2c322

Observation ea90cd2f-9f21-4682-bce3-cd66911dcd04 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Learning transferable visual models from natural language supervi- sion

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.752142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.752142Z digest=sha256:e50cc7f9e0b6555b2c6c40f276e0742a03da57a57b65247f085f5569cbbbc089

Observation 0aee56db-70e7-46d5-986f-f0326180cebe · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.756723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.756723Z digest=sha256:649c881d2f8765a1c26e9902212d174c76b44d03c8b174cf997e11d9861c2408

Observation 5910e255-592b-4ac9-9f78-605d0f2f09c8 · outbound

This paper cites Grounded sam: Assembling open-world models for diverse visual tasks,.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Grounded sam: Assembling open-world models for diverse visual tasks,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.761588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.761588Z digest=sha256:46472671d5af421b71ce86a6ea527755980f8c18105fb8924656748c2849a3a1

Observation b1c76509-8895-41e3-b0d9-73ab3a7b02e6 · outbound

This paper cites L3DG: Latent 3D Gaussian Diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation L3DG: Latent 3D Gaussian Diffusion

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.766543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.766543Z digest=sha256:13001c4c68c4127397fb1a9761ab32a605c9db3ff08e67a8fa8d26e90e14a5fe

Observation 1b3cd6f3-fc8a-4551-9bd2-193a6ce60af8 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation High-resolution image synthesis with latent diffusion models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.771589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.771589Z digest=sha256:e9405c0f8bbf15e0e12d7543fd2a265fdd2383a192f6be4e85341e3a6933da1a

Observation 2efe727b-e8ed-4d56-85b9-9871961d6cf0 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.776858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.776858Z digest=sha256:4845e309b50e1ddd10fc41862a67ed9cfc23a00d711ea9249c221dc5b2ec6197

Observation c28f9010-5b4c-4dc1-8fbf-5bfff51f1762 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Deep unsupervised learning using nonequilibrium thermodynamics

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.781731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.781731Z digest=sha256:9ad3128a23b73b12890afc039035bfb791cec4c7b06bbdbb1035608bf13c30ae

Observation b00e9ecb-db77-40ea-ba34-1c2ce12e8eb3 · outbound

This paper cites Denoising Diffusion Implicit Models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Denoising Diffusion Implicit Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.787127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.787127Z digest=sha256:7a80203cbb87266d377c0bd2f63d42b7f6da111349a69a6afd1035c19d410df8

Observation 2f01a310-0bf5-4fbd-8e7f-fb13c0a66753 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Score-Based Generative Modeling through Stochastic Differential Equations

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.793007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.793007Z digest=sha256:11eb99a1982764e714f6945186371ba86b708f89fedd542fe12b1b1e2b1e33bd

Observation 85b091e5-86cd-459a-81a0-68a15fcd8638 · outbound

This paper cites The Replica Dataset: A Digital Replica of Indoor Spaces.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation The Replica Dataset: A Digital Replica of Indoor Spaces

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.797905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.797905Z digest=sha256:8492178c855524bbd22541e53a31b1f79fadf2b3c1fe0f6515df96dd4eb0beaa

Observation 021566ec-dc81-4c05-aca8-dcac1652e558 · outbound

This paper cites Diffuscene: Denoising diffu- sion models for generative indoor scene synthesis.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Diffuscene: Denoising diffu- sion models for generative indoor scene synthesis

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.803354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.803354Z digest=sha256:36de7b4dc0f6d3df94d4d0a2ca13b6eb4fe8b8c16caefdab3bb5392b23a4e3bf

Observation 9305c0fb-9476-48f6-a03b-e66c1acd340a · outbound

This paper cites Lgm: Large multi-view gaussian model for high-resolution 3d content creation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Lgm: Large multi-view gaussian model for high-resolution 3d content creation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.811998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.809546Z digest=sha256:827925565c73cf2c3255be36139b35f8e152a44979ab73f021db26461ff09596

Observation 2387520d-6b36-4f6e-9be4-bfe118c6ff55 · outbound

This paper cites TripoSR: Fast 3D Object Reconstruction from a Single Image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation TripoSR: Fast 3D Object Reconstruction from a Single Image

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.814273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.814273Z digest=sha256:581facb3562fbd9c6ae70481fae705a26eff32ba04315eceb283084eea14ffc0

Observation 93544d0d-ed6c-4fc2-b764-d07cf498ddd6 · outbound

This paper cites Sv3d: Novel multi-view syn- thesis and 3d generation from a single image using latent video diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Sv3d: Novel multi-view syn- thesis and 3d generation from a single image using latent video diffusion

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.793990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.819595Z digest=sha256:2ec108feb1368579ecf00ecee02f929416d8178bd31e47fea30cb800340097d4

Observation 62907540-adfa-4392-b6dc-bd2287856f7c · outbound

This paper cites Neus: Learning neural im- plicit surfaces by volume rendering for multi-view recon- struction.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Neus: Learning neural im- plicit surfaces by volume rendering for multi-view recon- struction

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.776076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.824339Z digest=sha256:4e7ac29920f57e6f97062f9b51186e1db791795de4102d47de325c09edd5fa21

Observation 57786f34-0475-4391-a4cc-3fc6e4136d45 · outbound

This paper cites CRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction Model.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation CRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction Model

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.829216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.829216Z digest=sha256:455edcd431a484d679a43b9cd207b5ec360c580b6eb48150df0b916288921ad1

Observation 7bd221cb-40b3-47df-bb91-eaf98485e198 · outbound

This paper cites Ouroboros3D: Image-to-3D Generation via 3D-aware Recursive Diffusion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Ouroboros3D: Image-to-3D Generation via 3D-aware Recursive Diffusion

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.833706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.833706Z digest=sha256:25917a0c2faf5605fda7a2eb7461233d05e11d8f187a23474bc77ba5e6538031

Observation 0f8074ca-6a1c-4761-9079-777815bf6a79 · outbound

This paper cites Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.839041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.839041Z digest=sha256:a03d1ea1950ce1a231e3ef2c7e17a981d17bc5ca62adccc367b7b8ddaef97d43

Observation fa8909cb-667a-4e9d-babe-276b712589e8 · outbound

This paper cites Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.845335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.845335Z digest=sha256:d25f5b889a7d75499060d5ef4f4c4b13716734223305ad9d6792cf944549a268

Observation fba5c971-0aef-460f-a9ab-59ce9745411b · outbound

This paper cites Blockfusion: Expandable 3d scene gen- eration using latent tri-plane extrapolation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Blockfusion: Expandable 3d scene gen- eration using latent tri-plane extrapolation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.758795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.850436Z digest=sha256:8cb7f6488eb155ac2a742e98ca24a97e3ba7400e7eb0cac391baf90275e37652

Observation 9776c8a8-70a4-4eb2-bb2f-4ef7cc0b03b8 · outbound

This paper cites InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.855123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.855123Z digest=sha256:1b533674496102cdcc7f8f5e847d6219f8584cfb4386eff97319a2a97ba36e21

Observation ba131941-921c-4897-b769-bb2b2e7efed5 · outbound

This paper cites GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.861066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.861066Z digest=sha256:dce020a0492641299ed6f3ac2869875d6cfa1fa65cb18d2dc86b34969915d8f7

Observation 919b44c2-db8c-4e0a-828a-906f139bf3bf · outbound

This paper cites Hifi-123: Towards high-fidelity one image to 3d content gen- eration.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Hifi-123: Towards high-fidelity one image to 3d content gen- eration

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.866504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.866504Z digest=sha256:b463a1334a73c1e36ddcde269abb5d9fbc95e2d4360a801f9a8180fb65926094

Observation 7a460b27-9180-4f30-b0c1-b4ee2e977304 · outbound

This paper cites 3dshape2vecset: A 3d shape representation for neu- ral fields and generative diffusion models.ACM Transactions On Graphics (TOG), 42(4):1–16, 2023.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation 3dshape2vecset: A 3d shape representation for neu- ral fields and generative diffusion models.ACM Transactions On Graphics (TOG), 42(4):1–16, 2023

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T22:19:24.873513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:19:24.873513Z digest=sha256:fc41bbd4910f19e5c5710b8401bd49267cff9b67bd7bd0bc9900daedbb31f462

Observation 803b17c4-b102-4c33-89ff-fa26962839e1 · outbound

This paper cites Holistic 3d scene un- derstanding from a single image with implicit representation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Holistic 3d scene un- derstanding from a single image with implicit representation

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.720197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.878453Z digest=sha256:c341db90233ef26e0227d7a6914e8d89ab0bc296dc9d238649511370e1b6e9ac

Observation 5e6ee03e-db88-4d40-9444-2807829d1d27 · outbound

This paper cites Clay: A controllable large-scale generative model for creat- ing high-quality 3d assets.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Clay: A controllable large-scale generative model for creat- ing high-quality 3d assets

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.704010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.883741Z digest=sha256:00e4edb7ae1dcd7b41c472a70933badb69ac9680b4efff7d8f80b1467958da2a

Observation b54915e5-e69d-4a60-9f04-67cacd5c000f · outbound

This paper cites Uni-3d: A universal model for panoptic 3d scene reconstruc- tion.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Uni-3d: A universal model for panoptic 3d scene reconstruc- tion

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.685715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.888271Z digest=sha256:3f88a99b773166a0cdbc0f2cc2ae348b822ed91f7adfc1427c4db7f143578fff

Observation e329b9b4-f34f-4268-9b47-398e1a188dc8 · outbound

This paper cites Michelangelo: Conditional 3d shape generation based on shape-image-text aligned latent representation.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Michelangelo: Conditional 3d shape generation based on shape-image-text aligned latent representation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.667464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.893202Z digest=sha256:14a5d6adec3821446b4c171360758d7c0b530585a1f86f3cb8fa5616062d158c

Observation 23183069-f4b1-4e97-bea8-398ee8cea31f · outbound

This paper cites Zero-shot scene reconstruction from single images with deep prior as- sembly.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Zero-shot scene reconstruction from single images with deep prior as- sembly

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.649304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.897649Z digest=sha256:34c3bcb21d47fac995a8a23d05b1b146fef2bead72d87250cc91e01441fa634b

Observation e4009483-3653-444a-b7f4-ee62d19a5571 · outbound

This paper cites Triplane meets gaussian splatting: Fast and generalizable single-view 3d reconstruction with transformers.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Triplane meets gaussian splatting: Fast and generalizable single-view 3d reconstruction with transformers

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.630479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.903549Z digest=sha256:6c782fac198b01ec0c851988512eae7235ea9dac7c539096f49fb71985362288

Observation d0bc4880-5eca-48fc-adce-dc3268818d0a · outbound

This paper cites Following scalable 3D object generation methods [35, 71, 78, 80], we firstly trains a V AE to com- press 3D geometric representations into a low-dimensional latent space.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Following scalable 3D object generation methods [35, 71, 78, 80], we firstly trains a V AE to com- press 3D geometric representations into a low-dimensional latent space

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.611579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.908649Z digest=sha256:ce6d1eaa2e91c40eac555964cb51b2a339d5357f7ed3bba2350cd1de66bc1a17

Observation 41c32b14-175a-423f-977a-abc73eef6e11 · outbound

This paper cites we trained MIDI to simultaneously generate up to N = 7instances.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation we trained MIDI to simultaneously generate up to N = 7instances

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.594995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.914229Z digest=sha256:6b7340624cff2ca86f3367fa554f59b4e44c6ff627bba07740526903c50bdde6

Observation c9d72829-63f6-48ce-ac3e-402ab691a560 · outbound

This paper cites compositional generation methods.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation compositional generation methods

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:19:25.577736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.919288Z digest=sha256:ec0d435bef4e654418c227ecf161ed3df684826112b07fc8167cae4245d42290

Observation f827c2f4-932a-458f-9e23-5d056656da2f · outbound

This paper cites an unresolved cited work.

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:19:25.560639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T22:19:24.924471Z digest=sha256:d67bbf6a55351baf3f69420ba5c55dbf635217527ad666775a82d494daf864c4

Pith citing papers

Observation ea8209b2-d077-4e41-8e51-922df5ca57d8 · inbound

PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers cites this paper.

PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:57.655659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:57.655659Z digest=sha256:433e0a8f0fdf3b3e29c4e299f5bf65a79243e072dddde066afd68e706aced789

Observation 5582339e-cde9-40f7-9916-29b6ffea2854 · inbound

InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts cites this paper.

InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:11:40.108634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T17:11:26.524195Z digest=sha256:bd7eb569db326f97fc5567e9d8feee0ca0bb967c839c3b0e6f042e84b2007f5a

Observation 22391441-1ba5-497e-b491-b8c7d723cbec · inbound

STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics-Physics Dual System cites this paper.

STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics-Physics Dual System MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:18:54.649538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:14:48.069637Z digest=sha256:46ef6ee4a26048dd551a7c8fb49a49480032c7f5b9bf665c32f42bf52429c1ec

Observation 4cdd353a-d9dd-482d-86a7-fe77014bcc66 · inbound

ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures cites this paper.

ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:03:57.299791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T04:30:56.107507Z digest=sha256:9c8670a48835ace70dfa95a07c9d2b6fc162338aafe75955ffbbf03dc8fc7612