Pith. sign in

Paper Citation Record · LEDGER

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations

As of 23 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2412.06322.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06322 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:51:43.669537Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T08:47:25.712575Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T08:51:18.523931Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy40
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 39497bb6-902b-404c-8fc9-8b2b36052a7d · outbound

This paper cites GPT-4 Technical Report.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.221709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.221709Z digest=sha256:7b97a1b1701f5df7569b6de28ca3dd6b42b853a26fd81b820f2ef3cd99972914

Observation 922e3bcb-3b7b-4ba1-9a1e-5a871276ec37 · outbound

This paper cites Llama 3 model card.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Llama 3 model card

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.704174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.226887Z digest=sha256:aeb610cf9393897b5a37c2a19b3951ab507e4c3d3ea3592895f632df8fab1a60

Observation 3b93f7d4-901e-4ece-b0a1-2ded9a7b569a · outbound

This paper cites Vqa: Visual question answering.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Vqa: Visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.231390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.231390Z digest=sha256:8773227c74112ddc6e473864452ec50d21c3a6ad2e5bf615edd8c41579943a9c

Observation 69b3aa84-8263-4ba0-8bde-0a8db499ad73 · outbound

This paper cites Adabins: Depth estimation using adaptive bins.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Adabins: Depth estimation using adaptive bins

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.681644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.359521Z digest=sha256:ee9adfa48da74a1fd2e074a041936598fd954c41a691715bf8a5c95a7cf2b366

Observation 46e608b6-3100-4d2f-8f7e-f0cf91a01ea8 · outbound

This paper cites Transformerfusion: Monocular rgb scene reconstruction using transformers.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Transformerfusion: Monocular rgb scene reconstruction using transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.667241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.364262Z digest=sha256:58705b7ce65674bde60539e019aff34d685e170f72345c30564d3686a7f0de1c

Observation c99b669d-4cd7-4f08-ad7b-172a98bb1040 · outbound

This paper cites Language Models are Few-Shot Learners.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Language Models are Few-Shot Learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.370053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.370053Z digest=sha256:7735d31b987e333f324715687911ce3c59411af0e4ee84b7c03c1b2577013389

Observation ab2b56cd-7348-40e8-88b9-1ea909cb40ba · outbound

This paper cites Sift flow: Dense correspondence across different scenes.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Sift flow: Dense correspondence across different scenes

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.653775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.376028Z digest=sha256:c227c8b432c9bf53c1c314df72b717157e4260d5a1ff9ff02778c3f534ed5082

Observation 48438ee1-5f7e-4dd8-b156-f1dc65cc05a8 · outbound

This paper cites SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.385412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.385412Z digest=sha256:89dd8ca34f3969befb7b928c8323285d05608c4a6c7978a7fcfa97db440d38e0

Observation 0fc116a0-cd53-4837-8dca-9d578651ef64 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.391170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.391170Z digest=sha256:92b1320f7d7aa7a2c3a00227b67558801d514d6aac350dd58100b0534c85c3e6

Observation 670d3a78-c783-48c4-8554-aedf8d70167b · outbound

This paper cites Knowledge-embedded routing network for scene graph gen- eration.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Knowledge-embedded routing network for scene graph gen- eration

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.641632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.396263Z digest=sha256:746158c25edf632465fe6b4b1640d49d31e857d4ed1a1ea0bb55ef26d38b783d

Observation 650499dc-1685-49f1-845d-2a0110608a6c · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.400805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.400805Z digest=sha256:51b5294759ef337df29004646830eecb8ef573c5d90890e5e12d2b1601818cc6

Observation c8d9fd59-c5d3-4f34-9177-bf9d8526f891 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.406478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.406478Z digest=sha256:077db3c24e1d20ffad6471dae11d13060a94b6283376d8303fe076e7bd5c1673

Observation 8110f618-5fc2-478b-9e0a-cbfeec625a61 · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.417197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.417197Z digest=sha256:5f8e582a18b92c0d73260acf258d46cbbf7303bdc8b61c58be68229f9415a4e3

Observation a36950d6-2f13-41f6-9284-96ccef36a17d · outbound

This paper cites Deep- videomvs: Multi-view stereo on video with recurrent spatio- temporal fusion.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Deep- videomvs: Multi-view stereo on video with recurrent spatio- temporal fusion

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.622180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.422437Z digest=sha256:940213056f4b549800b5b8141f4e63db09ae915452bdaa4c5523f1e1f9663980

Observation f2c82a5a-d708-4297-bdfe-78c76707b1fa · outbound

This paper cites Depth map prediction from a single image using a multi-scale deep net- work.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Depth map prediction from a single image using a multi-scale deep net- work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.427327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.427327Z digest=sha256:4f7db431b944e0612e397fb80f7cdc9f7c1f37179e93e8b50d42f0e6818cdf02

Observation e79127c0-45e7-494e-9e5e-90f4e4ca4778 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.431469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.431469Z digest=sha256:ed4e8786c19b9ce946bc3eef2cd7cb9c4d89a235f8acda697291a872171b46e5

Observation 74396314-a112-445e-939d-6c2433987710 · outbound

This paper cites Recov- ering surface layout from an image.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Recov- ering surface layout from an image

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.594946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.435650Z digest=sha256:44fcd7dbcf69b1bf17e4e848e0b292d2a640426ac4bb859ccf9f662778167c0c

Observation fe669a19-ecd7-4604-b2ac-83923e1cdf11 · outbound

This paper cites Language is not all you need: Aligning perception with language mod- els.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Language is not all you need: Aligning perception with language mod- els

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.578612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.440020Z digest=sha256:3f05f918782f9e0fa50de8b2f5e813cf5fac4ea784992a963bf040859548892e

Observation 7032a545-41f7-45ea-85f4-b89444bfd56f · outbound

This paper cites Vi- sual prompt tuning.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Vi- sual prompt tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.444911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.444911Z digest=sha256:5fe7487426288e9b2c601bc6078bd2e4fad94bd64e7a68fa7083ff46e1f0743f

Observation 6f2678cb-59a8-4ca0-86f5-de1f36e3533a · outbound

This paper cites Image retrieval using scene graphs.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Image retrieval using scene graphs

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.551368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.449317Z digest=sha256:a0b99876326317d8c917f1b8866c10151870479e651ce66a265745a13a435b7d

Observation 2143d770-f82e-4d3f-8b4b-55ef52194054 · outbound

This paper cites Poisson surface reconstruction.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Poisson surface reconstruction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.537140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.454016Z digest=sha256:97e9252918293bed543a770e4e570fa97fe63bc71bbe6788dbd8d7f838f3e29d

Observation 9f8d8a15-300d-4ab4-b22b-2917dd973665 · outbound

This paper cites Two algorithms for constructing a delaunay triangulation.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Two algorithms for constructing a delaunay triangulation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.523178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.459550Z digest=sha256:b5c692e03d34a0c2268da582015442df20d25944bceb40c945979b81ab3973df

Observation 85e8f3ae-b47d-4feb-b8ae-76a2fc8df492 · outbound

This paper cites Depth and surface normal estimation from monocular images using regression on deep features and hi- erarchical crfs.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Depth and surface normal estimation from monocular images using regression on deep features and hi- erarchical crfs

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.507374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.465056Z digest=sha256:4d5c52ef801c1d56b77fa420d6e93d2600fe1a8876607fa87ed0335fd3d4b01c

Observation 2a7ec07c-d549-41d0-a158-58b9bfeb0ac4 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.470880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.470880Z digest=sha256:caf6e4a4b29816a9fe0195f2e356273d2811422afd5bac158b45b3ec913fb10d

Observation 8d30914f-4e13-40c4-8ecc-5b38f0ae760d · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Align before fuse: Vision and language representation learn- ing with momentum distillation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.475422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.475422Z digest=sha256:0c2a83b9854dc643e4f5819f4fabd6c842b1cc426e2c7274b29f8276bf705318

Observation 762cab21-7500-4bd9-b824-b44f7b0e6eb6 · outbound

This paper cites Factorizable net: an efficient subgraph-based framework for scene graph generation.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Factorizable net: an efficient subgraph-based framework for scene graph generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.474513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.481428Z digest=sha256:d4fc523ec6df0ce8e4d06f2601c442b6366f6f37c5f3e3df5b5e8bbac106ccd2

Observation 27cce3b6-fb70-4c0d-ba24-7b8a2dc85877 · outbound

This paper cites Scene graph generation from objects, phrases and region captions.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Scene graph generation from objects, phrases and region captions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.460319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.485917Z digest=sha256:68b9c728d083ec2972773c1ab24a1b41603ff76cfea3c6eec330e8fd85803f5d

Observation 195e6fd4-f7b0-4573-967a-ec81e199c2de · outbound

This paper cites StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.490848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.490848Z digest=sha256:be7b210d8b74e2794a6071ce5f99bd51f5ed3cdb37b2af28bb8ad7c9e94e0aea

Observation 1ac2b93e-21a5-4eef-9700-f62c2c0c8090 · outbound

This paper cites Binsformer: Revisiting adaptive bins for monocular depth estimation.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Binsformer: Revisiting adaptive bins for monocular depth estimation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.447512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.497784Z digest=sha256:04ae7bd6102882db40daab085ee942c685f09e3c11c07429089ae23b8869c323

Observation 184d2689-0012-46ea-bbdf-ba2f9adeaa3c · outbound

This paper cites Microsoft coco: Common objects in context.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Microsoft coco: Common objects in context

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.433886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.504058Z digest=sha256:a1a0f5589e4a89b839104ea9f20a04c56009023d7053f8366d9e4a29c547ba49

Observation ef52ed51-0af0-4a9d-b0e9-6271e2fa447e · outbound

This paper cites Gps-net: Graph property sensing network for scene graph generation.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Gps-net: Graph property sensing network for scene graph generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.419902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.508520Z digest=sha256:502b2f81f44bceaede201ebc0b865c435fc39c9748238978081d6efbea69e233

Observation 2c555bb6-1d02-41c6-9066-5bbf5502c84f · outbound

This paper cites Visual spa- tial reasoning.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Visual spa- tial reasoning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.404912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.515343Z digest=sha256:f9d08040415e46b92ece61091ca5bf4489ef797371ce450822a3113f3876035e

Observation d0be3b8e-dca5-4a8b-8c68-24d4a02c009c · outbound

This paper cites Improved baselines with visual instruction tuning.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Improved baselines with visual instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.387245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.519382Z digest=sha256:5d084eb65f7dfb9694d2a7c29ea14d35d428f4af5c6a3b00f33c0c302c48c8c2

Observation 5caa1d29-3e0b-48b1-af09-dd80b0c8628f · outbound

This paper cites Visual instruction tuning.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Visual instruction tuning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.371907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.523062Z digest=sha256:a7b5118c19aa750b9911d7e3d471aba5c74f740ad152229ecb547640437c99d8

Observation 76d7b4f5-95ba-4e8b-9667-7db0c0ec5e21 · outbound

This paper cites Visual relationship detection with language priors.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Visual relationship detection with language priors

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.356274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.527875Z digest=sha256:4810384605be5c79d7a941abd7a2b765d2d67ff310523bc9df8151e57ee39c52

Observation 770f9547-6fa9-4961-98e3-16b7d98339fb · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Ocr-vqa: Visual question answering by reading text in images

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.532032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.532032Z digest=sha256:b7d003aed6b1afa81532ccd0382691ffd849f3bb239f544f4b383f32fb2bd482

Observation 388776f7-ad93-4878-9ab0-d4563d735b80 · outbound

This paper cites Atlas: End- to-end 3d scene reconstruction from posed images.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Atlas: End- to-end 3d scene reconstruction from posed images

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.330590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.536493Z digest=sha256:fe9a19d88db10c3050b73e4e413017b4b854edb08e1680d0cdb9aca22aaf0069

Observation b8ea3c9d-09e5-405a-803a-fddc0297c8d6 · outbound

This paper cites Gpt-4o system card, August 2024.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Gpt-4o system card, August 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.316179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.540496Z digest=sha256:d4d01ee8856a4889989b8e4e99b371d99082da4a54ce8f76fba00ff5828e56db

Observation fcbfdd98-0f94-4fdf-bfbb-869e7826715a · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.545125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.545125Z digest=sha256:6c19ca37b66f4c6efbd6df961dcc2eb1f9eca58fa94b59b280c3a151a9736e53

Observation 288bab9d-ed7f-4ea4-9257-110d2d00eda7 · outbound

This paper cites Spatial-temporal knowledge-embedded transformer for video scene graph generation.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Spatial-temporal knowledge-embedded transformer for video scene graph generation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.300536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.549234Z digest=sha256:1b430cf17e011761b4a58c54fb393f7dc7f0888f1e033f82978fcf9777dab3f1

Observation a643bf50-030d-4567-9baa-e9d6b10d7062 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Qwen2.5: A party of foundation models, September 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.285145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.553123Z digest=sha256:cc4b90d4e30c67c91717b00173105cc8e6eb31eb6995caa3e3364bf14cf70d0e

Observation 169f72b4-033c-4f9e-8414-54abfebeb68a · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Learning transferable visual models from natural language supervi- sion

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.557283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.557283Z digest=sha256:8c7d419286615985df2ba9c786702cb9dc5c13cbf0552985f3ec59a6ca6d9988

Observation 0efdc85c-801b-4652-b9e9-409e725b6a86 · outbound

This paper cites Pixelwise view selection for unstructured multi-view stereo.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Pixelwise view selection for unstructured multi-view stereo

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.258940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.561715Z digest=sha256:265b7b41cef211fa5834cb9b12c5da90bd95f592850dcf54c795cb183b8a9d5d

Observation b106f35c-b582-4689-ab6a-726bde78c5ca · outbound

This paper cites Structured query- based image retrieval using scene graphs.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Structured query- based image retrieval using scene graphs

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.242378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.566847Z digest=sha256:0e7e1b6d74313e72bc769d630a3ea6830a36f27cdd7e9d00fc45d4a88b6dac74

Observation 15eed070-521a-42e8-a319-925a3bdd0ea9 · outbound

This paper cites Nddepth: Normal-distance as- sisted monocular depth estimation.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Nddepth: Normal-distance as- sisted monocular depth estimation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.227858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.572047Z digest=sha256:8d692b424cf4edb10801b1d0e62f8e4382efcc9dcdbb6c4f9bf3b7d2c26b1015

Observation 79c5c60c-2864-4287-8fd6-a76207b26152 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.214065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.578410Z digest=sha256:a536072c1009aded30749be9e62c8959d651e6fb9ca95f05b505da60857c2606

Observation 3e4fcaee-4e80-4299-a842-5c8d743c588c · outbound

This paper cites Towards vqa models that can read.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Towards vqa models that can read

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.198839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.582799Z digest=sha256:5d1ef0e5c2a5b804119f5ddab6c4f206009e383f055037ef0540acdea99a88b1

Observation 7cf09185-a045-46bb-9aba-c76e1462d5e0 · outbound

This paper cites Neuralrecon: Real-time coherent 3d re- construction from monocular video.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Neuralrecon: Real-time coherent 3d re- construction from monocular video

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.185172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.587503Z digest=sha256:c9ca1bb4b33293ce60738daae4cd700bb148bc6814f4cf80b68dfa825a9abbac

Observation 86281eb0-a7f4-4e34-aa7a-313b5666903b · outbound

This paper cites Learning to compose dynamic tree structures for visual contexts.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Learning to compose dynamic tree structures for visual contexts

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.167320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.592567Z digest=sha256:9b7020705a7de42086da93525491da7417ad70f7fd8760ab60e7dd81f1d0105b

Observation 33f13e65-90fc-43ab-9aad-e1ae3025a9a5 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.598156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.598156Z digest=sha256:dc747bfeb7524791d26493f840f56458af6c6c41fe149ae445206482a58688c1

Observation a669df6c-77bd-4c71-b69b-ea830c6f785f · outbound

This paper cites The All-Seeing Project V2: Towards General Relation Comprehension of the Open World.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations The All-Seeing Project V2: Towards General Relation Comprehension of the Open World

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.606576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.606576Z digest=sha256:48a1571809d5092ed0a08c1bc2b900c80cdaac216f663213571c66b4f2d5a378

Observation 66a83fdb-39d3-41e2-a7e4-6ea556099b3d · outbound

This paper cites The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.611325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.611325Z digest=sha256:47615fd7421bcc4986d9e37c5a0f861e9b21452c01a76535bbf7ba54736dc5f7

Observation b975ad39-7d7f-4869-9957-1c16109e1168 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.616939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.616939Z digest=sha256:e65b90dadd5e4ccc2048e78e4c8c2b157b73b876e254d98dd5952fc2cb9814d3

Observation 1393d05e-ed00-4506-b704-174f2aeba5de · outbound

This paper cites Scene graph generation by iterative message passing.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Scene graph generation by iterative message passing

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.139873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.620809Z digest=sha256:ab714f62d7aecb8716b3dd88a4c0df4ad763d86bd2f86ab23717acc280557bd9

Observation 40815618-0aff-4db9-880d-3037723aaa58 · outbound

This paper cites Panoptic scene graph gen- eration.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Panoptic scene graph gen- eration

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.121968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.624729Z digest=sha256:cd871492876523814fb036fd8826db3b9a21d5eec22a1d6c2a90607ced5292aa

Observation 432d6546-9082-4975-9dc3-7d4a7950b967 · outbound

This paper cites Graph r-cnn for scene graph generation.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Graph r-cnn for scene graph generation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.100017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.628324Z digest=sha256:708971b3613d17eb05fceed9ac06d3d8a4508a741d66e27fdb01f1d6d19acc5c

Observation e289b01d-f49f-4889-8f13-97769b19c86d · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Depth anything: Unleashing the power of large-scale unlabeled data

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.081519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.632227Z digest=sha256:c1088b23acfb5a5196120107e84c0a8dbb96d737fffe2418988feacb6bc6bfa0

Observation 577ea4c0-40ea-49a8-958c-b66a7f600c46 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.636055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.636055Z digest=sha256:b9e807ec7beaa72b57a13b4c79bfc9101c789fdd246bb7421871eccf7f1cad39

Observation a34d4f19-8543-4087-8c2b-35f8ac3172b9 · outbound

This paper cites Neural motifs: Scene graph parsing with global con- text.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Neural motifs: Scene graph parsing with global con- text

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.063272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.640447Z digest=sha256:45582e7ada0d38e902947de097bbb5aa18b982d9272f040c70adbd77d4eea5a0

Observation 579ed9c9-d3d0-4fd5-86d3-107fe9b2fdaa · outbound

This paper cites Conceptual and syntactical cross-modal alignment with cross-level consistency for image-text match- ing.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Conceptual and syntactical cross-modal alignment with cross-level consistency for image-text match- ing

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.046260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.645236Z digest=sha256:6a2e351c1341743f7b053098b80918d22024c628a08401b84c19063c9a21547b

Observation ee6f49af-a391-40af-8c86-dab709790b46 · outbound

This paper cites Textpsg: Panoptic scene graph generation from textual descriptions.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Textpsg: Panoptic scene graph generation from textual descriptions

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:44.028125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.650150Z digest=sha256:e06f5fea58906ab72340bd870daa2916fdef3e2663deea31890ba72e2e40bd34

Observation 36274a89-045f-44c4-be4d-64fe38596e20 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T19:51:43.654766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:51:43.654766Z digest=sha256:05a52257bde0c3361e368bfcc565ccb78d26a70d61aca72ed583da100205f9e3

Observation 5bc93447-2041-4245-99cd-62e25ffd18a2 · outbound

This paper cites an unresolved cited work.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:51:44.012852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.659817Z digest=sha256:14dcdd93a536d4c334973f2373789b11af08ba8cc2509e9fc7e05e3ba4e1725a

Observation 012915af-4d29-4a0a-a1e5-4fab99d9bf9e · outbound

This paper cites an unresolved cited work.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:51:43.996331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.664561Z digest=sha256:a494d14f1d2e1d95d426a36a5f4fc96ea3225d370723ed66cfc80a4f0df0d604

Observation 576fa07c-ae81-4b70-9232-468f177291c2 · outbound

This paper cites house” as an example): object labels(“house.

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations house” as an example): object labels(“house

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:51:43.976253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T19:51:43.669537Z digest=sha256:491cd48bd5428fe88aa0a8891b4860e99f0203edc5c0b808856b5b4f53012c9e

Pith citing papers

Observation 12af512c-a7af-4ad2-b8f3-0f9f707a467e · inbound

SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching cites this paper.

SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:51:18.526818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T08:47:25.712575Z digest=sha256:11830e8d5c826834d5b1dd2d2ab17e6c8ba14aecd1e8706bccf5004327017f63