Pith. sign in

Paper Citation Record · LEDGER

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes

As of 16 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2412.11396.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11396 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:02:23.305959Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c44423f8-885c-44e4-9be8-222d7bb39dc7 · outbound

This paper cites In: Gold- berg, Y., Kozareva, Z., Zhang, Y.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Gold- berg, Y., Kozareva, Z., Zhang, Y

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.191882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.191882Z digest=sha256:fecf1c57902ff274517013fa528680dd12acbd3ad0932fdb0b25877209cd95e9

Observation 8df79f80-f676-46c1-8f88-4f1c50adc7f9 · outbound

This paper cites Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T15:02:23.425177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:02:23.197396Z digest=sha256:46b1923d248ee870341a40b9a5b89a1ba46ca48625e432f852b9a0d7f74eddff

Observation 79725cf8-9740-479a-9ea1-60e64f91eb80 · outbound

This paper cites In: Krause, A., Brunskill, E., Cho, K., Enge lhardt, B., Sabato, S., Scarlett, J.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Krause, A., Brunskill, E., Cho, K., Enge lhardt, B., Sabato, S., Scarlett, J

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:02:23.753606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:02:23.202468Z digest=sha256:98d9eb6d4384b6d73efab84183400c78bb5898b157d5ebed50f10f82d9232989

Observation c8a35adb-04b1-4a50-a56b-dacada09c633 · outbound

This paper cites In: Proceedings of the 17th Conference of the Eu ropean Chapter of the Association for Computational Linguistics.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Proceedings of the 17th Conference of the Eu ropean Chapter of the Association for Computational Linguistics

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:02:23.733656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:02:23.207057Z digest=sha256:b1cc4debde024088ab2736223e200e52e499e69b87f33e4dc5afcaccd7b473c1

Observation 1750bb6b-8846-4c6b-af9b-a4d987990dcc · outbound

This paper cites Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.211881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.211881Z digest=sha256:f2d0a73ac28341188661eca01ae8063820c65d870f1bc78be4d22146ccbc8c12

Observation 1ae961f3-1ee4-435c-b66a-1bb38611e1e2 · outbound

This paper cites In: Findings of the Association for Comput ational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11- 16, 2024.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Findings of the Association for Comput ational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11- 16, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:02:23.716292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:02:23.217477Z digest=sha256:85cf30bb94db1138108ff77e91f1728078db9b0e798b3d69d218ce64b46ed8d9

Observation 525537a2-f56c-40d0-a742-1c6a16bfa4f5 · outbound

This paper cites In: Findings of the Association for Computational Ling uistics: EACL 2023.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Findings of the Association for Computational Ling uistics: EACL 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:02:23.699918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:02:23.223857Z digest=sha256:16111b3ee8dcfb39c4e0a193e3f42ff8fe9be78b4318455c3f5af15b17645ba2

Observation e0c1e7c9-a6ee-4500-843e-89878219bd55 · outbound

This paper cites In: Findings of the Associ ation for Computational Linguistics: ACL 2023.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Findings of the Associ ation for Computational Linguistics: ACL 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.229370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.229370Z digest=sha256:0f730b765b9db24d46a61da69910b458460c7551938c2849d7c4cf9522889ac5

Observation 191b4070-d104-4d94-8f3c-489bb9b5f391 · outbound

This paper cites In : Proceedings of the AAAI Conference on Artificial Intelligence.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In : Proceedings of the AAAI Conference on Artificial Intelligence

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.234520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.234520Z digest=sha256:23fe0de7171c3f7795ce329573c56e817d3b3526e447cd85b3f1f8705877f65f

Observation 59d8c182-bd55-4414-8dea-36d1cc1d3839 · outbound

This paper cites In: International confer- ence on machine learning.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: International confer- ence on machine learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:02:23.665352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:02:23.239294Z digest=sha256:679eaf1d4017e267c54811b5766ee6ea834c041259c661177548caf2865bf912

Observation 9b64a958-ae5d-4437-b526-841b1381d5cc · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.245449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.245449Z digest=sha256:79f9c2ca1fa982768779673b29f3ca5b57937b1f527619111c524a99cbeedb6f

Observation 4996202a-9291-4e88-9038-0737510d766d · outbound

This paper cites In: Oh, A., Naumann, T., Glo berson, A., Saenko, K., Hardt, M., Levine, S.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Oh, A., Naumann, T., Glo berson, A., Saenko, K., Hardt, M., Levine, S

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:02:23.650212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:02:23.252442Z digest=sha256:8d8241dcb12ccb685cda5636148c1ea1b1ea3bb134b3a7946a57582d02aee90c

Observation fc6484dd-47e6-41ce-ad46-f7132074b2b4 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.256879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.256879Z digest=sha256:1c016952dbefc4360a2d543884dd530c1b6f02a2afccb99de4f36ae39233b24f

Observation ccd8abb0-c467-4232-8c81-21890308095f · outbound

This paper cites In: Pro- ceedings of the 2021 Conference of the North American Chapte r of the Associa- tion for Computational Linguistics: Human Language Techno logies.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Pro- ceedings of the 2021 Conference of the North American Chapte r of the Associa- tion for Computational Linguistics: Human Language Techno logies

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.262416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.262416Z digest=sha256:77b432248889e13a841ab4ec08dabc77d4229855741abfa38cc5f5d4a32db7f6

Observation 1eaf28fd-90cc-4403-befc-67767904cdd4 · outbound

This paper cites : Modeling event- pair relations in external knowledge graphs for script reas oning.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes : Modeling event- pair relations in external knowledge graphs for script reas oning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.267112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.267112Z digest=sha256:f5de2ee5a18c10291be7ec31f11ce90cf8281add6f678d99048dfffaaffa129b

Observation 9f4c610b-c044-42b6-a754-0d0ffa837559 · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.270947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.270947Z digest=sha256:6720fc70fb85d7e591181b6a2f84f2436079fc62d6698fe69b357adcdcf5c3be

Observation ee4e5f92-be7c-4ccd-8e9f-f9154fdc5a6e · outbound

This paper cites In: Cai, J., Kankanhalli, M.S., Prabhakaran, B., Boll, S., Subramanian, R., Zheng, L., Singh, V.K., César , P., Xie, L., Xu, D.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Cai, J., Kankanhalli, M.S., Prabhakaran, B., Boll, S., Subramanian, R., Zheng, L., Singh, V.K., César , P., Xie, L., Xu, D

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:02:23.615837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:02:23.275985Z digest=sha256:bf4a8232ed0f4c7a177a656b728d3ceb7a70439f3c0b5a109970e4e711eceb53

Observation 7a0b849a-5ef0-475b-85ba-f424723da061 · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.285580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.285580Z digest=sha256:e62be8f3feb34935c9fb0ee3232e663c41d067774ef529a244a10fa0a9907fc8

Observation 14ef6d89-a516-4e00-8e59-0c5756570613 · outbound

This paper cites Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.290415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.290415Z digest=sha256:42fb204a15687b6473e27f888d582002b9d9f57344261e133ca08f7b0906bc27

Observation a44c310a-daa2-46a1-a8bd-10804d827395 · outbound

This paper cites Meta Knowledge for Retrieval Augmented Large Language Models.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes Meta Knowledge for Retrieval Augmented Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.295039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.295039Z digest=sha256:ddff71ff6e037ce5ac8336dff25bee96be02abb7a2313163652e18dc85b315f8

Observation e7083dd3-a474-4945-a193-6a637e572be6 · outbound

This paper cites In: Wooldridge, M.J., Dy, J.G., Natarajan, S.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes In: Wooldridge, M.J., Dy, J.G., Natarajan, S

Reference 21

Resolution
verified exact
doi, observed 2026-08-11T15:02:23.595819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:02:23.305959Z digest=sha256:65910b2bf501c3cb219917200d2fe91dc68da6f9dbd1de9a0b6c167178f05970

Observation 357f6e0a-625a-4afa-b30c-f72a3bf62267 · outbound

This paper cites 8845–8854.

Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes 8845–8854

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:23.281390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:02:23.281390Z digest=sha256:d75886a11047082245253ebcfb139a0589178149f752f9858ada8fff098dfd5e

Pith citing papers

No inbound Pith citation observations are available.