Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:02:23.305959Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2412.11396.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:02:23.305959Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c44423f8-885c-44e4-9be8-222d7bb39dc7 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Gold- berg, Y., Kozareva, Z., Zhang, Y
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8df79f80-f676-46c1-8f88-4f1c50adc7f9 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 79725cf8-9740-479a-9ea1-60e64f91eb80 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Krause, A., Brunskill, E., Cho, K., Enge lhardt, B., Sabato, S., Scarlett, J
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c8a35adb-04b1-4a50-a56b-dacada09c633 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Proceedings of the 17th Conference of the Eu ropean Chapter of the Association for Computational Linguistics
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1750bb6b-8846-4c6b-af9b-a4d987990dcc · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ae961f3-1ee4-435c-b66a-1bb38611e1e2 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Findings of the Association for Comput ational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11- 16, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 525537a2-f56c-40d0-a742-1c6a16bfa4f5 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Findings of the Association for Computational Ling uistics: EACL 2023
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e0c1e7c9-a6ee-4500-843e-89878219bd55 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Findings of the Associ ation for Computational Linguistics: ACL 2023
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 191b4070-d104-4d94-8f3c-489bb9b5f391 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In : Proceedings of the AAAI Conference on Artificial Intelligence
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59d8c182-bd55-4414-8dea-36d1cc1d3839 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: International confer- ence on machine learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9b64a958-ae5d-4437-b526-841b1381d5cc · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4996202a-9291-4e88-9038-0737510d766d · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Oh, A., Naumann, T., Glo berson, A., Saenko, K., Hardt, M., Levine, S
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fc6484dd-47e6-41ce-ad46-f7132074b2b4 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccd8abb0-c467-4232-8c81-21890308095f · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Pro- ceedings of the 2021 Conference of the North American Chapte r of the Associa- tion for Computational Linguistics: Human Language Techno logies
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eaf28fd-90cc-4403-befc-67767904cdd4 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes : Modeling event- pair relations in external knowledge graphs for script reas oning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f4c610b-c044-42b6-a754-0d0ffa837559 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee4e5f92-be7c-4ccd-8e9f-f9154fdc5a6e · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Cai, J., Kankanhalli, M.S., Prabhakaran, B., Boll, S., Subramanian, R., Zheng, L., Singh, V.K., César , P., Xie, L., Xu, D
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7a0b849a-5ef0-475b-85ba-f424723da061 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14ef6d89-a516-4e00-8e59-0c5756570613 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a44c310a-daa2-46a1-a8bd-10804d827395 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes Meta Knowledge for Retrieval Augmented Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7083dd3-a474-4945-a193-6a637e572be6 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Wooldridge, M.J., Dy, J.G., Natarajan, S
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 357f6e0a-625a-4afa-b30c-f72a3bf62267 · outbound
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes 8845–8854
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.