Pith. sign in

Paper Citation Record · LEDGER

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition

As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2507.11261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11261 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:18:14.922481Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99a5221d-5176-43af-abae-2f73a0eb4af0 · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:23.475431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:10.932060Z digest=sha256:fd3a59bdcd1a451c61ea55a05356dd2282e0ab6f824a841519c8cf9f44f31653

Observation f975c64a-485f-4d78-9f8d-dc42fbe8e48a · outbound

This paper cites Cot3dref: Chain-of-thoughts data-efficient 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Cot3dref: Chain-of-thoughts data-efficient 3d visual grounding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:23.222477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:11.018378Z digest=sha256:560a6ad3457df145af0effb2d61ed68caec2d23993914740dfd140da8fea7ebc

Observation f4e61f26-db46-490f-aba8-0eb1df005f13 · outbound

This paper cites Visual question answering from another perspective: Clevr mental rotation tests.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Visual question answering from another perspective: Clevr mental rotation tests

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.964487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:11.096372Z digest=sha256:85f63bb4dee1d8e046f8deb495c3f4b6960d12970de228cf8c45cf43f77a00d7

Observation d4dc29a6-0669-4bdd-976d-eb467977906c · outbound

This paper cites Assertiveness-based agent communica- tion for a personalized medicine on medical imaging diag- nosis.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Assertiveness-based agent communica- tion for a personalized medicine on medical imaging diag- nosis

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.729024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:11.162571Z digest=sha256:dc0868ca1d800d395f9b1a96a98036e3115ddc5b2cb5cd7f493a2e04478a682d

Observation a17f72e6-33c7-48fc-a038-94af86b7efcf · outbound

This paper cites Mikasa: Multi-key-anchor & scene-aware transformer for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Mikasa: Multi-key-anchor & scene-aware transformer for 3d visual grounding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.578516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:11.232514Z digest=sha256:a2e6acd56cc387540d60242e99d7ec1fde751dd38a97e21f373e600532a702e3

Observation 526e06e0-2a1c-4bec-b757-a07af60d434d · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.419009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:11.272693Z digest=sha256:d11cf6ad637291426511d094eba9127b98f77111da604f47b626c8aaab98bb62

Observation 59de1d01-e895-4fa6-baea-6d5f6edca71a · outbound

This paper cites SCJD: Sparse Correlation and Joint Distillation for Efficient 3D Human Pose Estimation.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition SCJD: Sparse Correlation and Joint Distillation for Efficient 3D Human Pose Estimation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:11.315194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:11.315194Z digest=sha256:501097c4be844802cbc66f14fb568487bc292f1f7b428eaa51b6ed2387956ec7

Observation eccb3eba-6be2-470e-bbeb-bf59790be25d · outbound

This paper cites Talk2bev: Language-enhanced bird’s- eye view maps for autonomous driving.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Talk2bev: Language-enhanced bird’s- eye view maps for autonomous driving

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.228420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:11.372304Z digest=sha256:0fb19110f4dc0840349ea0e239df3433eee54f3e9de39aea219487d969c43f71

Observation 1dc3fd23-afd9-4983-abd1-963c1e72f9a5 · outbound

This paper cites Drive as you speak: Enabling human-like interac- tion with large language models in autonomous vehicles.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Drive as you speak: Enabling human-like interac- tion with large language models in autonomous vehicles

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:22.060228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:11.459556Z digest=sha256:106fb7ab7571b4f7c088ed5d3a8bbf5fda18923ac21f211f2392f6d416d24bd6

Observation 1c2babb7-5748-4946-a129-a39f62ca7b64 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:11.526902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:11.526902Z digest=sha256:18c32d318c3a5b03b5e1b73a30959ca8d63e5be67b5a42f594733058a47338e4

Observation 0ffb4a29-cc63-41ed-8328-6490f340850e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:11.558820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:11.558820Z digest=sha256:7d2a8eeb05525c4f67a6be37dadad702a401beb1dda088766226857b8f621d64

Observation ade96116-986d-4715-878b-0a4040aa684d · outbound

This paper cites Scenegenie: Scene graph guided diffusion models for image synthesis.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Scenegenie: Scene graph guided diffusion models for image synthesis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.859773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:11.609396Z digest=sha256:ab973cfa7668a5d4cff4c1fb5ae8994068ebbc45039ec3e41cd51e3fb671685a

Observation 94dd748e-1e5b-4db3-abdf-d3114273a179 · outbound

This paper cites Dense reinforce- ment learning for safety validation of autonomous vehicles.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Dense reinforce- ment learning for safety validation of autonomous vehicles

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.686087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:11.725233Z digest=sha256:fe09acf5a337cded6ca999bae9d9a2d642739c237f75a698b9d74e6067b1050c

Observation af1fb14e-7516-443a-8274-73fcc2fbc43c · outbound

This paper cites Viewinfer3d: 3d visual ground- ing based on embodied viewpoint inference.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Viewinfer3d: 3d visual ground- ing based on embodied viewpoint inference

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.525736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:11.814980Z digest=sha256:89840572ea6795cc06d66758d05f6591b69eb3c83e009e8ef02ae1888654c2cb

Observation e05dd4c4-9cf5-4bd0-8f03-d88fee1038db · outbound

This paper cites Viewrefer: Grasp the multi-view knowledge for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Viewrefer: Grasp the multi-view knowledge for 3d visual grounding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.382853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:11.879963Z digest=sha256:9b546c6b38b31bf10358a58fbaf044cc75c0a49ba542a25e07afcfa06298616a

Observation 10802696-30ec-4432-802d-8d1ad2a2a2ba · outbound

This paper cites Transrefer3d: Entity-and- relation aware transformer for fine-grained 3d visual ground- ing.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Transrefer3d: Entity-and- relation aware transformer for fine-grained 3d visual ground- ing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.222781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:11.960630Z digest=sha256:bdcc2b0480b2a6ecd4ee4971dae2d59f971a39400cb4a128db3bfaa3418cbac5

Observation 871a281a-822f-48d2-b87e-f017dd6f71da · outbound

This paper cites 3d-llm: In- jecting the 3d world into large language models.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition 3d-llm: In- jecting the 3d world into large language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:12.036499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:12.036499Z digest=sha256:9f0f6ead71fdeb1a7816d0b31f122569d86c88ee84b446080dfe865915bb13d8

Observation be0bd385-f4ef-4432-8a79-37595d71b9e9 · outbound

This paper cites Multi- view transformer for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi- view transformer for 3d visual grounding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:21.108849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:12.169841Z digest=sha256:04cf48a7e50f8ae41583fa991bd6c068d83eda5b2f795e3bc954d31f0e7e2a9c

Observation fe61d437-8966-45a3-b4d8-9ffb6c95ed08 · outbound

This paper cites Structure-clip: Towards scene graph knowledge to enhance multi-modal structured repre- sentations.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Structure-clip: Towards scene graph knowledge to enhance multi-modal structured repre- sentations

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.956537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:12.226913Z digest=sha256:38203eb111ab2658729462f81d35bc22f6b2ab4fba6833f21c6f7eef724e99ab

Observation 1447d34a-0364-45f7-9836-3487acfcbaec · outbound

This paper cites Nan-detr: noising multi-anchor makes detr better for object detection.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Nan-detr: noising multi-anchor makes detr better for object detection

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.759454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:12.302592Z digest=sha256:46d38f5bdc80ec4dd6b673528c5ec7b0ffb49bf1baf9fb8ebdf7326a8671a95e

Observation c7234e13-59d2-4e77-a0e8-722b6bfeb5ee · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Bottom up top down detection transform- ers for language grounding in images and point clouds

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.616504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:12.337407Z digest=sha256:af966866f452ca04fce946d1c2cd1f47a3d505f22b1b0a2f1f1d8b40f347e3ac

Observation 20737413-972d-41be-8b38-9cf6e05427b4 · outbound

This paper cites Sita: Struc- turally imperceptible and transferable adversarial attacks for stylized image generation.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Sita: Struc- turally imperceptible and transferable adversarial attacks for stylized image generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.474489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:12.420336Z digest=sha256:7035dd5c9fa7ec6cfebab93aa3483a4b8d6033b0ed9121eabc41889f85fe03c8

Observation 0a65a579-0f29-4f63-8d2d-aea7735b381e · outbound

This paper cites Is-ggt: Itera- tive scene graph generation with generative transformers.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Is-ggt: Itera- tive scene graph generation with generative transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.356956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:12.464267Z digest=sha256:e7ab6283917f1fd39d9059c3a6c074e3a93ff878e94b3164732fae2dbe17b0b8

Observation 379d5fb8-6792-49bd-994f-0a4a458fb316 · outbound

This paper cites Panogen: Text-conditioned panoramic environment generation for vision-and-language navigation.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Panogen: Text-conditioned panoramic environment generation for vision-and-language navigation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.153322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:12.567526Z digest=sha256:98ec528a61fd86f2f09d3c8ee4f0df35a7e449be2b123ef8b7ebf336369bbe7a

Observation 10e9e04c-bdb5-43e1-9a72-f80606081a36 · outbound

This paper cites Delving into in- visible semantics for generalized one-shot neural human ren- dering.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Delving into in- visible semantics for generalized one-shot neural human ren- dering

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:20.017204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:12.610425Z digest=sha256:bac78fce769452236c29cef73c480caeea74101807e349f8b001966d1e4cc480

Observation eac386b2-18cb-4782-833e-b7e25bd217ad · outbound

This paper cites Multi- modal situated reasoning in 3d scenes.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi- modal situated reasoning in 3d scenes

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.881172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:12.672393Z digest=sha256:f1f2697684adb6330e9cb16244b1c355fafb5f57a013534a4320fa2944c27beb

Observation 92964dcf-8d01-48f2-8a93-fcf256401bc5 · outbound

This paper cites DeepSeek-V3 Technical Report.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition DeepSeek-V3 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:12.730320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:12.730320Z digest=sha256:d188c69c1e7539a93a7fbb2f3c3ee6d007f3831c096a6e4c9bc2cd8e99e8f547

Observation 69934a9d-b77c-4eec-bdd4-0c747b923a4b · outbound

This paper cites Rotation-adaptive point cloud domain generalization via intricate orientation learning.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Rotation-adaptive point cloud domain generalization via intricate orientation learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.793645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:12.792481Z digest=sha256:b63bb1cdb6a47860929e0150651e56552469096fa8511140e2a61197db984b40

Observation 8d336ad7-7f9f-40a3-919c-4fe67cba887e · outbound

This paper cites Decoupled Weight Decay Regularization.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Decoupled Weight Decay Regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:12.856045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:12.856045Z digest=sha256:16d4f84f2907bcf826feb5cbbfe3f56fc13266f34895a6c8f9ae4c1edae16aab

Observation f2e1b9c6-b79e-4826-b41e-d6d679641ab6 · outbound

This paper cites Mmscan: A multi-modal 3d scene dataset with hierarchical grounded language annota- tions.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Mmscan: A multi-modal 3d scene dataset with hierarchical grounded language annota- tions

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.641153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:12.907511Z digest=sha256:6d43a1f7425a04e186c48380cccd3265e765cb8817f997fdecc80a83778156b3

Observation 6cb05e79-39f8-48c6-83b4-0778956ae5e7 · outbound

This paper cites Situa- tional awareness matters in 3d vision language reasoning.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Situa- tional awareness matters in 3d vision language reasoning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.530994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:13.002554Z digest=sha256:6ea7087c7791d39194b090fa3542c220d3afdc7a39942539a58b16748b94a695

Observation b1d8da55-3518-489d-841a-e30ffcb33ad9 · outbound

This paper cites Textrank: Bringing order into text.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Textrank: Bringing order into text

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.415622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:13.049968Z digest=sha256:1c950afff069f09e54f9389dd026fc914016851b66bf988bc198ed948dbb6b56

Observation 7baa5367-9431-4ebf-9dab-b0e614a4469c · outbound

This paper cites Gaussian prompter: Link- ing 2d prompts for 3d gaussian segmentation.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Gaussian prompter: Link- ing 2d prompts for 3d gaussian segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.242942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:13.117883Z digest=sha256:e37a4f415a8ec20823eaefce64196f759e8d8926e516ce3f1af07af76657a4f7

Observation c0cb8c5d-fc43-440c-8122-b81e486fd0af · outbound

This paper cites An approach to gener- ate a caption for an image collection using scene graph gen- eration.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition An approach to gener- ate a caption for an image collection using scene graph gen- eration

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:19.100794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:13.165411Z digest=sha256:158ad6d82bcd8fb643d63a887c361888efeb55156875157323ce1a6dcd72b1c2

Observation 9d956151-940c-4880-8a83-724d8a8d0321 · outbound

This paper cites Pointnet++: Deep hierarchical feature learning on point sets in a metric space.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Pointnet++: Deep hierarchical feature learning on point sets in a metric space

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:13.209464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:13.209464Z digest=sha256:cdceb5a6b7d8f87ce418efc79698e2d2fc8ddd32a76147cdceaa55a057991a6b

Observation 7aca918f-4f06-4bbe-8c3a-8ccfbe288b3a · outbound

This paper cites Languagerefer: Spatial-language model for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Languagerefer: Spatial-language model for 3d visual grounding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:18.899318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:13.264346Z digest=sha256:1723341aec7454695811faf4130c6b8184100b05e3f878817d6ae90d1db9d02d

Observation 0dc6fb11-c63d-4a44-8000-0772f5480f73 · outbound

This paper cites Aware visual grounding in 3d scenes.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Aware visual grounding in 3d scenes

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:18.591055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:13.302172Z digest=sha256:61b21946f7c65c84eaa2d73428986291ebe259b2dc099f6af35c75265ddb1ae3

Observation c8ca42ed-25ea-4e09-bf48-939444c1df02 · outbound

This paper cites Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:13.353401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:13.353401Z digest=sha256:48ed40c130ab4b2e5b8e878e14d9f2994e5491aa764a48c4578842fdbc066c85

Observation 1478cb02-15b0-46d6-92fe-72defd0eb2b3 · outbound

This paper cites Attention is all you need.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Attention is all you need

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:13.395000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:13.395000Z digest=sha256:729c022189212ecb2458bff00d551743e75998a5a0aa747a729f638df422bf41

Observation 3fb44b95-ce78-4a3d-a504-3ebc3c94a86f · outbound

This paper cites OpenChat: Advancing Open-source Language Models with Mixed-Quality Data.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition OpenChat: Advancing Open-source Language Models with Mixed-Quality Data

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:13.463357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:13.463357Z digest=sha256:125e388d9a0650f441a838652e4bb0f676f18422abec9f854dde698aa2fa6560

Observation 4abf7b5a-2bff-4eba-b2a7-666fe6aa4823 · outbound

This paper cites Gˆ 3-lq: Marry- ing hyperbolic alignment with explicit semantic-geometric modeling for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Gˆ 3-lq: Marry- ing hyperbolic alignment with explicit semantic-geometric modeling for 3d visual grounding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:18.360406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:13.530305Z digest=sha256:fab553615f6f565d63b934eca22fd3f2cbd409ed3292849966b36df9e293b163

Observation c3063547-1f1d-4b65-bd7c-d838fb358d06 · outbound

This paper cites Eda: Explicit text-decoupling and dense alignment for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Eda: Explicit text-decoupling and dense alignment for 3d visual grounding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:18.114493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:13.577268Z digest=sha256:febba984250f2fc461404ab3fa7ef6f642393583a261291a90e1db710046c718

Observation 4f8a1019-7534-49a9-91ef-647eee7f552d · outbound

This paper cites Multi- scale flow-based occluding effect and content separation for cartoon animations.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi- scale flow-based occluding effect and content separation for cartoon animations

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:17.875702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:13.656275Z digest=sha256:43080cd334ec7b2d13ecf55ebf70f28598f6fc80a201806d5bbfc5239fd78faa

Observation dc991a12-da9e-4040-aea2-80f2cd5f3ce3 · outbound

This paper cites Multi-attribute interactions matter for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi-attribute interactions matter for 3d visual grounding

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:17.663023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:13.741285Z digest=sha256:172cc3ed9bfa87034d4ee7c974aa45e2aacce61c9b29e0ba89ce4257998ab20a

Observation 4f1a0181-73b4-4da4-9ebb-d7a7579bb0f1 · outbound

This paper cites Learning with unreliability: Fast few-shot voxel radiance fields with relative geometric consistency.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Learning with unreliability: Fast few-shot voxel radiance fields with relative geometric consistency

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:17.436049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:13.829101Z digest=sha256:02a44a76fb8845c633dfa47071eeaba4d2a8308c1af7998621038d68ea2ab13c

Observation 8d5e569b-edb4-48de-9a4f-f24c278d71fa · outbound

This paper cites Qwen2.5 Technical Report.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Qwen2.5 Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:13.861692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:13.861692Z digest=sha256:fc3b4881f3b814f5086b6c66c1e475322a20c4ec2c29f2ae7f967b29fee9129d

Observation dee07287-1d11-4c9a-8419-d1521f62ed39 · outbound

This paper cites G2face: High-fidelity reversible face anonymization via gen- erative and geometric priors.IEEE Transactions on Informa- tion Forensics and Security, 2024.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition G2face: High-fidelity reversible face anonymization via gen- erative and geometric priors.IEEE Transactions on Informa- tion Forensics and Security, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:17.218042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:13.971665Z digest=sha256:b41b95503899c6f895d393771991c6e40d4fcae1ad4fee9a7297d0230a5ff011

Observation 410dc413-a061-453c-afe6-3e74887c004c · outbound

This paper cites Sat: 2d semantics assisted training for 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Sat: 2d semantics assisted training for 3d visual grounding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:17.017863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:14.062628Z digest=sha256:c017cc4d16f7c92430fee25a0a3edd63b1732da22969d3290fa33e5e6c2ce091

Observation c51e1a3b-11e9-4557-aa96-26c9f73191b1 · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition AppAgent: Multimodal Agents as Smartphone Users

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:14.165736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:14.165736Z digest=sha256:b43c602272ffe7ce711e9a128c10565930657b59e4e532b43cae9acbf824e74b

Observation 7285a952-e166-495d-96f7-3e13a0c27a45 · outbound

This paper cites Visually-prompted language model for fine-grained scene graph generation in an open world.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Visually-prompted language model for fine-grained scene graph generation in an open world

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:16.729026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:14.205944Z digest=sha256:65a210ca141012f355125313da73ecd75201f843d297070f2c55aeb5955fd85b

Observation 50924109-6dab-4f3e-b74f-a1aaabe18993 · outbound

This paper cites Visual programming for zero-shot open-vocabulary 3d visual grounding.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Visual programming for zero-shot open-vocabulary 3d visual grounding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:16.560839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:14.295905Z digest=sha256:7095ae44640f960991bdfb0f79189060dbeeb408eda67bc52d783d41920efff9

Observation 0d2a8417-3822-4740-b6da-d5ad400d68f9 · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d ob- jects.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Multi3drefer: Grounding text description to multiple 3d ob- jects

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:16.325888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:14.401402Z digest=sha256:39dad95f63ed17961b90b2c76a8db48671e0fe1eee9707ebca75a40364974498

Observation d7249bfb-ba41-47a9-91ba-e54441ddf6c2 · outbound

This paper cites Towards clip-driven language-free 3d visual grounding via 2d-3d relational en- hancement and consistency.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Towards clip-driven language-free 3d visual grounding via 2d-3d relational en- hancement and consistency

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:16.133141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:14.470853Z digest=sha256:fe6213473e5d006aeac37b18e2988fc87f78f94b451eebe1c839417cf3891a9a

Observation 4fe4caae-2393-4b23-ac9e-4e82e6bb8ea6 · outbound

This paper cites 3dvg- transformer: Relation modeling for visual grounding on point clouds.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition 3dvg- transformer: Relation modeling for visual grounding on point clouds

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:15.958457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:14.561773Z digest=sha256:a53749c777b8444230c563d808374307b9b3cd171b3c9c477823a3ec58d88382

Observation b2d39ab7-26b5-4deb-b54a-4aa8bcd4585e · outbound

This paper cites Recdreamer: Consistent text-to-3d generation via uniform score distillation.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Recdreamer: Consistent text-to-3d generation via uniform score distillation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:15.769685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:14.688627Z digest=sha256:80277418692059fae7e561c0683fff52b9bc8eb3f9deeb2e1cb06462212e0973

Observation 0f8264bd-dc26-47cd-bb8e-d9ba862e5af3 · outbound

This paper cites Learning an interpretable stylized subspace for 3d-aware animatable artforms.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Learning an interpretable stylized subspace for 3d-aware animatable artforms

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:15.560384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:14.768209Z digest=sha256:4a862d31b1b98c26e815c26ff2b447c06656080040849c3a95c5905b9cef8b92

Observation 24fadd34-555f-417a-b47b-443df573e8d0 · outbound

This paper cites Navgpt: Explicit reasoning in vision-and-language navigation with large lan- guage models.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Navgpt: Explicit reasoning in vision-and-language navigation with large lan- guage models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:15.365406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:14.861449Z digest=sha256:301fe0657b3222600926e666ac5ea1065532b8292c19a4dd73e7caa78bae5522

Observation 39697c49-7872-4fd2-bc57-0ee034b57ec5 · outbound

This paper cites Unifying 3d vision-language understanding via prompt- able queries.

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition Unifying 3d vision-language understanding via prompt- able queries

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:18:15.154024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:18:14.922481Z digest=sha256:92a5e58dbd65da95f881b14ddd58791aff90a54eef2e3b5de56e972d60ae6af4

Pith citing papers

No inbound Pith citation observations are available.