Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:42:01.350617Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2505.12194.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:42:01.350617Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8b4be10b-1382-44bd-b4af-51394726bd5b · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation da35ada0-b447-4d57-b7e9-3dacc66b8a36 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed9d6ede-b858-48e3-b7f7-dbce13430eea · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Paligemma: A versatile 3b vlm for transfer, 2024
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6acb85eb-f62b-4b90-b140-5ed25d4dc631 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Meyer, Yuning Chai, and Yong Jae Lee
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6d3a2c34-fd32-4b0c-a84d-2f92adad4e50 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding End-to-End Object Detection with Transformers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87d26fe7-bffa-416e-ab89-53b1cee56357 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Chang, and Matthias Nießner
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f83ae1f4-d6ec-4ce9-a251-543cbc481efd · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Reproducible scaling laws for contrastive language-image learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 39df7291-0c35-4e9b-8dad-793358ac81fd · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Gonzalez, Ion Stoica, and Eric P
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 762ea977-ff9d-4b23-8948-1f2264653156 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding PaLM: Scaling Language Modeling with Pathways
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3c3f399-48b3-460c-b636-d75fea4626d4 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Scaling Instruction-Finetuned Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8130732-5a40-464b-a7ca-5f1466da59d3 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c45436c2-9cb1-4049-a484-6e16fa69a7d2 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Qlora: Efficient finetuning of quantized llms, 2023
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 37ccbc78-2b94-4772-9e69-0970e51f125a · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Bert: Pre-training of deep bidirectional transformers for language understanding, 10 2018
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c9f7267a-0eec-42da-bda3-570e8bd66d9e · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Eva: Exploring the limits of masked visual representation learning at scale, 12 2022
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cf39f242-ab4c-42d8-a1da-c4bfbcc6a177 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Sebastian Borgeaud
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b58ed854-ed16-400c-9f2d-071448d86b51 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Arzen-llm: Code-switched egyptian arabic-english translation and speech recognition using llms
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a5bc6b3b-bfbb-46e5-be73-e7cf9a4bf894 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Lawrence Zitnick, and Ross Girshick
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1e6b80dc-c27f-46a2-9f2c-b95a4af42833 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding What's "up" with vision-language models? Investigating their struggle with spatial reasoning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdb5dbec-3f04-4245-9e98-ac561733b61f · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Referitgame: Referring to objects in photographs of natural scenes
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3927bc2f-8dbe-4b33-a61c-38541ce2b2d6 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Computational genera- tion of referring expressions: A survey
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 50d03fe8-9096-468a-9ad8-5caeda702d2e · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding When to retrieve: Teaching llms to utilize information retrieval effectively, 05 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5c85146a-578c-46bf-bcea-0d30960c5668 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models, 2023
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 27718a0e-9a3d-4103-a71c-5aeb284cd0f7 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, 2022
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f3c152d-a8c4-4ad4-94e0-c62a001d1e4c · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Selvaraju, Akhilesh Deepak Gotmare, Shafiq Joty, Caiming Xiong, and Steven Hoi
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d69a8043-cef0-40f3-93c5-14e17c9fe60a · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Lawrence Zitnick, and Piotr Doll´ar
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ff3f4944-11d7-4a4e-9e26-782abbd2c978 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Visual spatial reasoning, 2023
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e54f4f4-558c-4328-af62-5476b10f019b · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Refer-it-in-rgbd: A bottom-up approach for 3d visual grounding in rgbd images
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 378b87a3-6f1a-429d-8baf-4ab029211aa2 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Improved baselines with visual instruction tuning, 2023
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e2f3754-f5cc-4591-bcb2-dcc28e77cb6b · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Visual instruction tuning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ad51824-6f5f-4c72-9c8a-aec9b6d663c0 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Clevr- ref+: Diagnosing visual reasoning with referring expressions
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation abaaab58-51d3-4651-a77c-354ec57b547f · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Roberta: A robustly optimized bert pretraining approach, 2019
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7891a45e-b672-4cea-9da9-f46cbd335944 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding MMBench: Is Your Multi-modal Model an All-around Player?
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2dd3b37-32f9-4cbf-adc4-c732769c27d4 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding An embarrassingly simple approach for llm with strong asr capacity
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 14052c44-5cfb-4b4e-bfb2-aa3f316cb49a · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Generation and comprehension of unambiguous object descriptions, 2016
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8f8cf1a7-e58c-4095-8fa7-3a664ea5b488 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Sun-spot: An rgb-d dataset with spatial referring expressions
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6212bf1c-060d-4f35-8ddf-b9d89230f993 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Domain terminology integration into machine translation: Leveraging large language models, 2023
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 172f7d6a-b936-4f74-98e3-f8da32adf00f · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a287a7dd-1a97-4e49-97df-e94acd94d29e · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding GPT-4 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ef06a3e-801d-43e5-bbec-4cab472d661c · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Reverie: Remote embodied visual referring expression in real indoor environments, 2020
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation aa93edda-840a-48b5-a640-a9d3aa962d3b · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Learning Transferable Visual Models From Natural Language Supervision
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dac8927-2b91-471b-8687-8912fdb530b0 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Sun rgb-d: A rgb-d scene understanding benchmark suite
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e05419f-41f1-41e1-b87b-2e32960c3ae2 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Eva- clip: Improved training techniques for clip at scale, 03 2023
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ba500ec0-b31a-4592-9f9c-68fd8781acca · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Self-retrieval: Building an information retrieval system with one large language model
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7cc07cf8-ec25-4113-a282-d956bbc8a967 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding LLaMA: Open and Efficient Foundation Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5505bb0-e851-4ef3-a301-8ecbf0ea6c15 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3f3ad556-9c77-41fb-b97d-86ec82214b33 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding The generation of natural descrip- tions: corpus-based investigations of referring expressions in visual domains
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 90b8c13a-315c-4806-982e-dacf13fa1c13 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v, 2023
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d377aad-26fb-44de-89cb-11dafc29afbe · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Xlnet: Generalized autoregressive pretraining for language understanding, 2019
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 200d4cde-550a-45b7-a4d3-9273d800c70f · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Modeling context in referring expressions
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0361d342-2114-4f40-aa67-82e7f06e30b6 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding A joint speaker-listener-reinforcer model for referring expressions
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6ceb79ce-3570-4b12-aa65-c996ba87e7e2 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Prompting large language model for machine translation: A case study, 2023
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d2ff4c5a-c55a-4eb0-bf61-eb5f79fce442 · outbound
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding Toolqa: A dataset for llm question answering with external tools
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
No inbound Pith citation observations are available.