Pith. sign in

Paper Citation Record · LEDGER

Unified Semantic Transformer for 3D Scene Understanding

As of 11 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 3 inbound Pith citation observations for arXiv:2512.14364.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.14364 v3

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T16:13:56.722386Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T18:06:09.895072Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-01T08:15:32.510741Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved59
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef85ed60-9e7b-4dc6-8349-a9adae6951d7 · outbound

This paper cites SAB3R: Semantic-Augmented Backbone in 3D Reconstruction.

Unified Semantic Transformer for 3D Scene Understanding SAB3R: Semantic-Augmented Backbone in 3D Reconstruction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:51.829966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:51.829966Z digest=sha256:48fb7b7de384322f8b5f195de41c20dd9b813017a5992ec437f9ccf776af5e19

Observation 6eb83102-cedd-477d-9a4f-a762c57bde73 · outbound

This paper cites Schwing, Alexan- der Kirillov, and Rohit Girdhar.

Unified Semantic Transformer for 3D Scene Understanding Schwing, Alexan- der Kirillov, and Rohit Girdhar

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:51.907699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:51.907699Z digest=sha256:eac25f2eb4ab43bd21929fa555de632b57e44f45bc36369b4eec54898e3ae8c2

Observation 424fab8b-906a-46c0-83ac-61f53a770ec8 · outbound

This paper cites Scannet: Richly- annotated 3d reconstructions of indoor scenes.

Unified Semantic Transformer for 3D Scene Understanding Scannet: Richly- annotated 3d reconstructions of indoor scenes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:51.961952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:51.961952Z digest=sha256:87a206af184cab7d634ee602b20c20abbac8133428b2af2f2ebf91308c8979e6

Observation 9b4672a3-0cae-48a9-b65e-bd0960154908 · outbound

This paper cites Scenefun3d: Fine-grained functionality and affordance un- derstanding in 3d scenes.

Unified Semantic Transformer for 3D Scene Understanding Scenefun3d: Fine-grained functionality and affordance un- derstanding in 3d scenes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:52.021178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:52.021178Z digest=sha256:14299ad2300c9039176035edd818f93b356b762d518c9cebeae791c3847128b3

Observation 85ae6983-9706-4b92-8329-5ad9c02c53fd · outbound

This paper cites The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,.

Unified Semantic Transformer for 3D Scene Understanding The llama 3 herd of models.arXiv e-prints, pages arXiv–2407,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:52.110622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:52.110622Z digest=sha256:576a26bf7a5d61a2aeadae0fdfaa55ae0e43608a43fe2542eb31fe9e24a75835

Observation d5d2fec5-9bb1-41e4-9c7c-469fdf61d7a8 · outbound

This paper cites OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views.

Unified Semantic Transformer for 3D Scene Understanding OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:52.205158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:52.205158Z digest=sha256:f5356f3308caaf47084aa6fbe4f0a402864bc1265fe777136f8518979d460fc3

Observation 96b3029e-afdb-4ccc-b64a-9b4cc0b90b8a · outbound

This paper cites A density-based algorithm for discovering clusters in large spatial databases with noise.

Unified Semantic Transformer for 3D Scene Understanding A density-based algorithm for discovering clusters in large spatial databases with noise

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:52.265144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:52.265144Z digest=sha256:06c700af2463731d3bab69cab7d4422c7b3fb1e1a37c4e0e51c7fa8b64f4363d

Observation edc85827-0bbf-43fb-b4ee-037302f166c3 · outbound

This paper cites Large spatial model: End-to-end un- posed images to semantic 3d.Advances in neural information processing systems, 37:40212–40229, 2024.

Unified Semantic Transformer for 3D Scene Understanding Large spatial model: End-to-end un- posed images to semantic 3d.Advances in neural information processing systems, 37:40212–40229, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:52.322807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:52.322807Z digest=sha256:44255d304f5d4812f77dbeedac0e419d4eb9ea559c3c60f823f81260c5addfd5

Observation 2d142051-adaf-4af1-97e8-7920c3b247ab · outbound

This paper cites Efficient graph-based image segmentation.International journal of computer vision, 59(2):167–181, 2004.

Unified Semantic Transformer for 3D Scene Understanding Efficient graph-based image segmentation.International journal of computer vision, 59(2):167–181, 2004

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:52.498620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:52.498620Z digest=sha256:5d9e55d1b0361b4256c3e28a16de0986e5e49ca1a0f47226b66f2c0feb1daee5

Observation 58e58c7a-6cbe-4f62-8dd5-ed673b760b9b · outbound

This paper cites Scal- ing open-vocabulary image segmentation with image-level labels.

Unified Semantic Transformer for 3D Scene Understanding Scal- ing open-vocabulary image segmentation with image-level labels

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:52.657195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:52.657195Z digest=sha256:133e90f319028a99ac726b1e44b9acbf4810d1d5d54aa5188b736f099290a8b4

Observation 6ea98a13-d9d7-4c94-9d98-24889373891d · outbound

This paper cites Ov3r: Open- vocabulary semantic 3d reconstruction from rgb videos.

Unified Semantic Transformer for 3D Scene Understanding Ov3r: Open- vocabulary semantic 3d reconstruction from rgb videos

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:52.810249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:52.810249Z digest=sha256:242f3541393dfa341e5c8499d765419e09eebdcd17266681a589afe9c0049291

Observation 60af9cad-2ed4-48d1-912a-c2c7cda28f5d · outbound

This paper cites Con- ceptgraphs: Open-vocabulary 3d scene graphs for perception and planning.

Unified Semantic Transformer for 3D Scene Understanding Con- ceptgraphs: Open-vocabulary 3d scene graphs for perception and planning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:52.993582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:52.993582Z digest=sha256:dc64bc393e344ac1d5c78ef181a1ea8cbfe0226c03a4edc60b973a40ef6febb6

Observation 249f3950-ee5d-48a7-9ae5-850bfd836bbe · outbound

This paper cites Occuseg: Occupancy-aware 3d instance segmentation.

Unified Semantic Transformer for 3D Scene Understanding Occuseg: Occupancy-aware 3d instance segmentation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:53.107545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:53.107545Z digest=sha256:0c2edc73821a7e0f7dcadd5e053c116e23628308d30a9d5e8fdca44f5ebbb7b9

Observation 288868a2-74dd-46dc-bf22-97e7a7ac4f5c · outbound

This paper cites Algorithm as 136: A k-means clustering algorithm.Journal of the royal statistical society.

Unified Semantic Transformer for 3D Scene Understanding Algorithm as 136: A k-means clustering algorithm.Journal of the royal statistical society

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:53.228912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:53.228912Z digest=sha256:358eaaad55247dd81d866253a89a9122071d3743e38b74289fda68665e12c911

Observation a356fbf0-f95b-413d-b262-b427303e81c2 · outbound

This paper cites Mask r-cnn.

Unified Semantic Transformer for 3D Scene Understanding Mask r-cnn

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:53.339763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:53.339763Z digest=sha256:7f81c8421b9ec552e83f92bd272bdc9299952dacedcc4220cc25e160c42fde73

Observation f0ec8152-4ea5-410d-bdda-6cf1eb576747 · outbound

This paper cites Pe3r: Perception- efficient 3d reconstruction.arXiv preprint arXiv:2503.07507,.

Unified Semantic Transformer for 3D Scene Understanding Pe3r: Perception- efficient 3d reconstruction.arXiv preprint arXiv:2503.07507,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:53.457616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:53.457616Z digest=sha256:9576ca2e534cfc33c0813138b89fb4eff00442f8fbaa7de351ee7b25d5edf65c

Observation 2ee57efb-aabd-4240-81fd-bb0644917e7e · outbound

This paper cites Tenenbaum, Celso Miguel de Melo, Madhava Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba.

Unified Semantic Transformer for 3D Scene Understanding Tenenbaum, Celso Miguel de Melo, Madhava Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:53.567985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:53.567985Z digest=sha256:68f0528bf6df81645a96f2494d978287d78c5189f3dd04b934f20636364f37a0

Observation 0782c0c3-e47f-4c34-9eae-a22c7b8cb127 · outbound

This paper cites Opd: Single-view 3d openable part detection.

Unified Semantic Transformer for 3D Scene Understanding Opd: Single-view 3d openable part detection

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:53.662385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:53.662385Z digest=sha256:20ed5ba121cf0e2737b4bd95163423ad27c22ce4f4d54d7974aeabd9d6f99e3a

Observation 62c8885e-ee33-494e-94ac-d21d8360e5ae · outbound

This paper cites Open-vocabulary 3d semantic segmentation with foundation models.

Unified Semantic Transformer for 3D Scene Understanding Open-vocabulary 3d semantic segmentation with foundation models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:53.754487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:53.754487Z digest=sha256:7eff0b071ced4eb8d176ec13ac209251298161ca05bdf549efd1b64c6d4db7c7

Observation aa0d2542-7483-4769-824a-868473523534 · outbound

This paper cites MapAnything: Universal Feed-Forward Metric 3D Reconstruction.

Unified Semantic Transformer for 3D Scene Understanding MapAnything: Universal Feed-Forward Metric 3D Reconstruction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:53.815387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:53.815387Z digest=sha256:6cef0d44d8bf4d2409ecda4200edc60b4b75ad6910519ce52346097029e701e6

Observation c521f90f-e88a-4e3a-9826-2184c8f62dc7 · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.ACM Transactions on Graphics, 42(4), 2023.

Unified Semantic Transformer for 3D Scene Understanding 3d gaussian splatting for real-time radiance field rendering.ACM Transactions on Graphics, 42(4), 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:53.892781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:53.892781Z digest=sha256:8844a125fe7eaae61af644efe5aba9b0a31c2299d78c73baf5a02414ce703d44

Observation b4361fb3-6dce-4bfc-a07b-8b845fea12da · outbound

This paper cites Lerf: Language embedded radiance fields.

Unified Semantic Transformer for 3D Scene Understanding Lerf: Language embedded radiance fields

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:53.975692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:53.975692Z digest=sha256:7e12755e537ee503c24276d0bee5a5bc56bc6768f14fe1cacb94919e1e6be732

Observation 93f20ea8-530d-42d4-ba0d-07cc762a7c3b · outbound

This paper cites Segment any- thing.

Unified Semantic Transformer for 3D Scene Understanding Segment any- thing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:54.064122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:54.064122Z digest=sha256:3411086f85f1323d00bccf4f9e38e2674502bf1c82b952d9a511a156d87eb195

Observation 78a6cce2-eb0e-49ab-b0bb-4da3e40d7447 · outbound

This paper cites Open3dsg: Open-vocabulary 3d scene graphs from point clouds with queryable objects and open-set relationships.

Unified Semantic Transformer for 3D Scene Understanding Open3dsg: Open-vocabulary 3d scene graphs from point clouds with queryable objects and open-set relationships

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:54.144779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:54.144779Z digest=sha256:e1802cada0da47b459014a8d0e4918c84d9386879e2153e7a9b5c08fc079392f

Observation 4c5e345e-82c2-4402-9f60-8a1f3a829437 · outbound

This paper cites Relationfield: Relate anything in radiance fields.

Unified Semantic Transformer for 3D Scene Understanding Relationfield: Relate anything in radiance fields

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:54.208533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:54.208533Z digest=sha256:4fd6636294bca2f410942acc7fa5b4d0b1b8aa183735934d6b6273464ca93a52

Observation 0a3f2329-a3a1-4b86-8a8c-11336779d843 · outbound

This paper cites Iggt: Instance- grounded geometry transformer for semantic 3d reconstruc- tion.arXiv preprint arXiv:2510.22706, 2024.

Unified Semantic Transformer for 3D Scene Understanding Iggt: Instance- grounded geometry transformer for semantic 3d reconstruc- tion.arXiv preprint arXiv:2510.22706, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:54.285222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:54.285222Z digest=sha256:bad60ee18b6fef192bb481d27ed75d77937fba5868279b63153472cab405e19a

Observation 7a6e91a4-8f2e-4cfd-a5e0-c7cdb9cdd006 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Unified Semantic Transformer for 3D Scene Understanding Improved baselines with visual instruction tuning, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:54.359880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:54.359880Z digest=sha256:490a954d0e50e3c009500a829a2121909e2b1714b2450787cad5de044dea90f5

Observation 0a1e28ef-cf82-423a-bd77-957a254ad257 · outbound

This paper cites Visual instruction tuning.

Unified Semantic Transformer for 3D Scene Understanding Visual instruction tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:54.466426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:54.466426Z digest=sha256:cceccc70725bbf592f74819dbbc5cf02656e359464c7a95ed663e52f71657aa2

Observation 7a7ffcf9-ad16-4c70-9dfe-b2981a46c459 · outbound

This paper cites World- mirror: Universal 3d world reconstruction with any-prior prompting.arXiv preprint arXiv:2510.10726, 2025.

Unified Semantic Transformer for 3D Scene Understanding World- mirror: Universal 3d world reconstruction with any-prior prompting.arXiv preprint arXiv:2510.10726, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:54.577239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:54.577239Z digest=sha256:bd0412bfc162cf69ceb8f380b21e4e46b2f44a7b7ca3bf483478c61dda65555f

Observation 55877cbf-2255-4365-8313-7e625c65e563 · outbound

This paper cites Multiscan: Scalable rgbd scanning for 3d environments with articulated objects.Advances in neural information processing systems, 35:9058–9071, 2022.

Unified Semantic Transformer for 3D Scene Understanding Multiscan: Scalable rgbd scanning for 3d environments with articulated objects.Advances in neural information processing systems, 35:9058–9071, 2022

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:54.638086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:54.638086Z digest=sha256:59fb9e96cd556799f4a3309757e195212aea360e03b3e7351bccc459dfd394de

Observation a09c34d0-1163-4440-a4bd-d9078f4cde84 · outbound

This paper cites Accelerated hierarchical density based clustering.

Unified Semantic Transformer for 3D Scene Understanding Accelerated hierarchical density based clustering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:54.681470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:54.681470Z digest=sha256:e170f9cf5fe6ca7408f506d1d4597f06698b9dde0dc2161791d6a17c212a744b

Observation 56c84c05-3f24-4e44-8750-b2eb225e0bbe · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view syn- thesis.Communications of the ACM, 65(1):99–106, 2021.

Unified Semantic Transformer for 3D Scene Understanding Nerf: Representing scenes as neural radiance fields for view syn- thesis.Communications of the ACM, 65(1):99–106, 2021

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:54.761229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:54.761229Z digest=sha256:2b798b89fa1e02cf7b38b3769b185f5065617496bdc6eaec57baba82cba0b588

Observation 97a1356d-de04-44b6-8ff0-3f374ad4223b · outbound

This paper cites An End-to- End Transformer Model for 3D Object Detection.

Unified Semantic Transformer for 3D Scene Understanding An End-to- End Transformer Model for 3D Object Detection

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:54.839360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:54.839360Z digest=sha256:ebef7b843d6af1870be4aa6273e496b8287b53a6d2f55607ded93809b17725e7

Observation d693942f-2975-4438-a8f4-4abe9a05a7fb · outbound

This paper cites Mix3d: Out-of-context data augmenta- tion for 3d scenes.

Unified Semantic Transformer for 3D Scene Understanding Mix3d: Out-of-context data augmenta- tion for 3d scenes

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:54.920220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:54.920220Z digest=sha256:939a69c26d9755064098e13bb94eaff30463fdb6596170155ad7ebcc6c5eec0c

Observation 1bed2d4d-5498-4940-b614-252e007e293a · outbound

This paper cites an unresolved cited work.

Unified Semantic Transformer for 3D Scene Understanding Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:54.999557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:54.999557Z digest=sha256:6505888761cfd5c07652e1ae2f20ad7b153cb588aed08b50ba8935373161b452

Observation 4f64203a-d069-43a6-8378-54e31d43e383 · outbound

This paper cites an unresolved cited work.

Unified Semantic Transformer for 3D Scene Understanding Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:55.079614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:55.079614Z digest=sha256:cb6478f91544c30f20bfbd9fcdbf9af78f1aae4bce2921a6db4346ab821df2dd

Observation b67eb9f8-c877-4118-ade2-d2b01bc38843 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Unified Semantic Transformer for 3D Scene Understanding Openscene: 3d scene understanding with open vocabularies

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:55.130971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:55.130971Z digest=sha256:4b49b2638af448453fa4e3178150a699a99a7d79f97a5d55cd592247582c3be8

Observation 34b4d3e1-eb9f-4bf8-afdd-75651c99e4bd · outbound

This paper cites Deep hough voting for 3d object detection in point clouds.

Unified Semantic Transformer for 3D Scene Understanding Deep hough voting for 3d object detection in point clouds

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:55.190541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:55.190541Z digest=sha256:b7b28f8e8fad69a304f39325dd68e956401f10b080ffb8103a62e44c2d0a9bc3

Observation 1cc2bd24-f2bb-479e-911c-ba97cd9a7bb6 · outbound

This paper cites Langsplat: 3d language gaussian splatting.

Unified Semantic Transformer for 3D Scene Understanding Langsplat: 3d language gaussian splatting

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:55.253104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:55.253104Z digest=sha256:86e2789d6370bac010297b015706f710663df1ebd0ea2cd44fb87842bf2d8026

Observation 587c3746-fd74-4de7-bee2-5ba5ef7c1ab3 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Unified Semantic Transformer for 3D Scene Understanding Learning transferable visual models from natural language supervi- sion

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:55.321312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:55.321312Z digest=sha256:5a91b642d76aa822357ae71e11ec62c920510f9a6d71f40c5ad07ea806ff07d3

Observation 6275c937-c084-4334-b9d7-6a09b5862f07 · outbound

This paper cites Vi- sion transformers for dense prediction.

Unified Semantic Transformer for 3D Scene Understanding Vi- sion transformers for dense prediction

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:55.394637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:55.394637Z digest=sha256:a762adc68ca05922576fd4ea18b7ddd8c2b0527d3d3f705095443b18bec012a7

Observation fb895f79-8e47-44a7-ade9-5bdbcd283d84 · outbound

This paper cites Language- grounded indoor 3d semantic segmentation in the wild.

Unified Semantic Transformer for 3D Scene Understanding Language- grounded indoor 3d semantic segmentation in the wild

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:55.463373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:55.463373Z digest=sha256:a23cbd659e5c04b305a6b763dc67bc5df9b64a0b06e06b442ac82f41a9173b7d

Observation 6c767002-9f34-4d7a-8596-7e552b8ea99d · outbound

This paper cites Mask3D: Mask Trans- former for 3D Semantic Instance Segmentation.

Unified Semantic Transformer for 3D Scene Understanding Mask3D: Mask Trans- former for 3D Semantic Instance Segmentation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:55.546176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:55.546176Z digest=sha256:0c443e4b4db7b6777c7106ec13b1810f73cd0a17352504712269af12432b4cde

Observation d57d49a6-0b0b-411f-aa6e-ee776e4793b4 · outbound

This paper cites The Replica Dataset: A Digital Replica of Indoor Spaces.

Unified Semantic Transformer for 3D Scene Understanding The Replica Dataset: A Digital Replica of Indoor Spaces

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:55.626779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:55.626779Z digest=sha256:8ef2b6eab410c06305992b4c540fa4a5c45a38ceba17f6b86c5f08ab69978da9

Observation 1095a2ca-80b8-404d-9af7-d456b4adb6a7 · outbound

This paper cites Uni3r: Unified 3d reconstruction and semantic understanding via generalizable gaussian splatting from unposed multi-view images, 2025.

Unified Semantic Transformer for 3D Scene Understanding Uni3r: Unified 3d reconstruction and semantic understanding via generalizable gaussian splatting from unposed multi-view images, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:55.700796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:55.700796Z digest=sha256:2755a2483921d419643baabf940456ac02cd0cad05ced339c86de41ee3233281

Observation dfdce46d-b22f-46bc-955d-4a8d542bdf15 · outbound

This paper cites Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann.

Unified Semantic Transformer for 3D Scene Understanding Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:55.782117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:55.782117Z digest=sha256:8b76c6ce1ff6452a35207a7889f5ca1127803a25aa8edac88fdb9ffea8371ee6

Observation 31d3ce92-6e99-4236-87df-f871200d8923 · outbound

This paper cites Search3d: Hierarchical open-vocabulary 3d segmentation.

Unified Semantic Transformer for 3D Scene Understanding Search3d: Hierarchical open-vocabulary 3d segmentation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:55.847135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:55.847135Z digest=sha256:20fe62350393e283f8551dff5f7b34e6e0ec52194d28a0475b79f4c9cf487fa4

Observation 14fce24c-d453-41b8-94e8-9b7307ab980f · outbound

This paper cites Softgroup for 3d instance segmentation on point clouds.

Unified Semantic Transformer for 3D Scene Understanding Softgroup for 3d instance segmentation on point clouds

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:55.937835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:55.937835Z digest=sha256:c1ed57ad24bc8d450fe8c2b7a057f86216f04680527c1269efbd1109f5eb236d

Observation ffe0354a-a5c4-4de8-a4df-bfd1c8b455ce · outbound

This paper cites Rio: 3d object instance re- localization in changing indoor environments.

Unified Semantic Transformer for 3D Scene Understanding Rio: 3d object instance re- localization in changing indoor environments

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:56.010022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:56.010022Z digest=sha256:7426226e5d4f0499120aff189f598c25d24c3d5b193a5570427d07bcb59c0873

Observation c5e2d57a-5077-4b3c-94f3-961e83d9f30d · outbound

This paper cites Vggt: Visual geometry grounded transformer.

Unified Semantic Transformer for 3D Scene Understanding Vggt: Visual geometry grounded transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:56.092128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:56.092128Z digest=sha256:a136804b67cbd08efc6754919cedf1c387a56b3a46fd551dfe93c766c0f9ab63

Observation d9d3d5ec-61eb-4c2c-9ffb-37ee10a01787 · outbound

This paper cites Dust3r: Geometric 3d vision made easy.

Unified Semantic Transformer for 3D Scene Understanding Dust3r: Geometric 3d vision made easy

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:56.135198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:56.135198Z digest=sha256:92f4b1c826007624e265a86e5a8fa9fe71aa4f1e5707cbfa68bf2bd913be85c6

Observation cb635117-6fef-4a67-9386-2b41fface530 · outbound

This paper cites Shape2motion: Joint analysis of motion parts and attributes from 3d shapes.

Unified Semantic Transformer for 3D Scene Understanding Shape2motion: Joint analysis of motion parts and attributes from 3d shapes

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:56.207555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:56.207555Z digest=sha256:79e3920d4b63d53a0c57782e718d4a5ec9a80e9e0afb39c1f7ca1ef93c8cc498

Observation f5c3fdb6-b9ca-4cfd-8ce3-0cb30b2eb07d · outbound

This paper cites π3: Scalable permutation-equivariant visual geometry learning, 2025.

Unified Semantic Transformer for 3D Scene Understanding π3: Scalable permutation-equivariant visual geometry learning, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:56.284094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:56.284094Z digest=sha256:f73a1abddd46eae49cf57189ecbaa3b4fa59df00e59b64536ec7b210fe68e53c

Observation e7369080-302c-469a-88cf-4b98bcfc31f6 · outbound

This paper cites Hierarchical open- vocabulary 3d scene graphs for language-grounded robot nav- igation.Robotics: Science and Systems, 2024.

Unified Semantic Transformer for 3D Scene Understanding Hierarchical open- vocabulary 3d scene graphs for language-grounded robot nav- igation.Robotics: Science and Systems, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:56.342595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:56.342595Z digest=sha256:d01a9f99ec5cf72a68ceac83644f27a165ceb620234c18c848a84388f0805acc

Observation 2e5e07dd-7513-409c-b292-1195f0d89a6a · outbound

This paper cites Point transformer v3: Simpler, faster, stronger.

Unified Semantic Transformer for 3D Scene Understanding Point transformer v3: Simpler, faster, stronger

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:56.394164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:56.394164Z digest=sha256:d9ffbf66455839fb3dc67ec0c160d352c97810147833d9b36379fca253c627a3

Observation ca80aa73-bd3c-4702-b421-38b6fd74a210 · outbound

This paper cites SAM3D: Segment Anything in 3D Scenes.

Unified Semantic Transformer for 3D Scene Understanding SAM3D: Segment Anything in 3D Scenes

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:56.470579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:56.470579Z digest=sha256:462be4b710f666c3bf91e7063b698dc8ebe8e822aae6db80362cf41bb72fc456

Observation 11fde6eb-4f65-4dc1-aed3-5c7ea40d2e3e · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d indoor scenes.

Unified Semantic Transformer for 3D Scene Understanding Scannet++: A high-fidelity dataset of 3d indoor scenes

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:56.525303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:56.525303Z digest=sha256:7d998c1dec50f722bbdab088aeea09c22b74f4bf486353acd4bfd4405cd1d934

Observation 013acef2-1648-466c-89bc-e02dc179d507 · outbound

This paper cites Sigmoid loss for language image pre-training.

Unified Semantic Transformer for 3D Scene Understanding Sigmoid loss for language image pre-training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:56.587667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:56.587667Z digest=sha256:94fd1a0cb0d2c01c34889659a85f8697749f5163bb5454285fc68a2f10044ae5

Observation 40cdb8e9-6ca5-4634-98c0-a27f110eb2f9 · outbound

This paper cites Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip.

Unified Semantic Transformer for 3D Scene Understanding Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T16:13:56.647218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:56.647218Z digest=sha256:1d40e1d6b3fc87b06839a94b0b3783ae93a981c244a589c242fcb3ccdba49b95

Observation 630d8038-1d29-42fa-95ed-5c17b4e7a421 · outbound

This paper cites Panst3r: Multi-view consistent panoptic segmentation.

Unified Semantic Transformer for 3D Scene Understanding Panst3r: Multi-view consistent panoptic segmentation

Reference 60

Resolution
malformed identifier
no resolver link, observed 2026-08-03T16:13:56.722386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:13:56.722386Z digest=sha256:a91ab5548e1dc5aeb972af9ec1afead573daa75a3d2e886bab9c23681960cd96

Pith citing papers

Observation 7c711c68-0cd5-4ef3-98d1-cbe8c254770f · inbound

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models cites this paper.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models Unified Semantic Transformer for 3D Scene Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.895072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.895072Z digest=sha256:a559915f6c5deecb41a4d9abcc34636afb0cf7d9ab5d9ada65688dd3aabd32bd

Observation b666dc28-a108-471d-841b-1baa1d22533f · inbound

FOUND-IT: Foundation-model-first Task-driven 3D Scene Graphs with Granularity on Demand cites this paper.

FOUND-IT: Foundation-model-first Task-driven 3D Scene Graphs with Granularity on Demand Unified Semantic Transformer for 3D Scene Understanding

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:14:00.415456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T22:04:21.666471Z digest=sha256:825d05b09bc40ecfe1d5209623e3b60cbf730e7cc344e19dcbafb4bb88e267a3

Observation 32359263-b489-48f0-97a4-b9925aff81e4 · inbound

Pano3D: Unified 3D Reconstruction and Panoptic Segmentation cites this paper.

Pano3D: Unified 3D Reconstruction and Panoptic Segmentation Unified Semantic Transformer for 3D Scene Understanding

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.512870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-01T07:23:24.807344Z digest=sha256:a72fbbbcd2d8168fbedb0e9563e13a693e743251d221def4cb439e8154672376