Pith. sign in

Paper Citation Record · LEDGER

Retrieval-Augmented Multimodal Language Modeling

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2211.12561.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.12561 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:34:14.457752Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

29
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 00b934be-6f7a-4749-956f-f7cd406540f9 · inbound

REPLUG: Retrieval-Augmented Black-Box Language Models cites this paper.

REPLUG: Retrieval-Augmented Black-Box Language Models Retrieval-Augmented Multimodal Language Modeling

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T12:41:54.147344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-17T12:41:53.833754Z digest=sha256:1fb4b800aa87fc79c384a2a15e395ebe4816fe7cfecba46c26a93b83125b24b8

Observation 9add77e6-99b8-4d24-8f4c-05972f515c61 · inbound

Language Is Not All You Need: Aligning Perception with Language Models cites this paper.

Language Is Not All You Need: Aligning Perception with Language Models Retrieval-Augmented Multimodal Language Modeling

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:32:22.906269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T18:32:22.813668Z digest=sha256:09d064462fe96ee5798b9a16fa501511b00cc55e361331d9b77e381b8c174dce

Observation 437f27d8-a5f0-49ff-8856-32d0ac7c6fee · inbound

Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection cites this paper.

Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection Retrieval-Augmented Multimodal Language Modeling

Reference 106

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T14:15:11.295731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T14:15:10.907921Z digest=sha256:5a9385823d4ae93e72e4943327bbe75f1153e97accdca8422fb4ea9fc5d848d5

Observation 00e2512f-96ed-4a1b-84a1-db95c39edd01 · inbound

Retrieval-Augmented Generation for Large Language Models: A Survey cites this paper.

Retrieval-Augmented Generation for Large Language Models: A Survey Retrieval-Augmented Multimodal Language Modeling

Reference 176

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:13:56.483980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T05:10:25.171044Z digest=sha256:99b1d250019da9b347caf4047e935dde98abcc5202920fbdbf1cfb0aecdfb5fb

Observation a442456b-ee74-40b2-b22b-aa57346c2b3b · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Retrieval-Augmented Multimodal Language Modeling

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:20:36.381184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:9ea756ca2c2e9b43e8feb781cd8c2fee60446d1be9af73aefb7fdbe65011769a

Observation 93cc96aa-2d47-47f9-b54f-f2071ded5392 · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems Retrieval-Augmented Multimodal Language Modeling

Reference 212

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T11:52:16.195672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:c2f2dc4602afdc1af45dec719abb03fc31e055b0cf1cae6a49f9df5ea42afc37

Observation 9aac474c-5315-4279-94fa-81bbd9350349 · inbound

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs cites this paper.

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs Retrieval-Augmented Multimodal Language Modeling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:14.457752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:14.457752Z digest=sha256:6fb9706689bcf8884a129b252a42e2bbe7b71dcc7dcd95a38f76c6e9a9590b7e

Observation 5a0c5e72-11f4-45ec-996a-10d0e88c35af · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Retrieval-Augmented Multimodal Language Modeling

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:53.705938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:53.705938Z digest=sha256:b4e2af59b2f1bf78fede7aa699e8d7c5ca4e5cf2011fbf64cf641d46500af148

Observation e5830b19-0053-4593-a880-68bf53e36a20 · inbound

Attention Grounded Enhancement for Visual Document Retrieval cites this paper.

Attention Grounded Enhancement for Visual Document Retrieval Retrieval-Augmented Multimodal Language Modeling

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:55:15.322419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:53:15.920563Z digest=sha256:10b07a9d3cced483d02bad44a35bda003a7aebdbba068a280060f54b22b1e7cd

Observation 2811251e-3007-413d-949f-7a9332fa0d32 · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs Retrieval-Augmented Multimodal Language Modeling

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:22.630225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T18:53:06.494640Z digest=sha256:e23ad10319362bcd202ad4773e67e93eadf1308601300d848f5f60ec065d94d6

Observation 0685a0b1-6834-4025-91b9-7be507767028 · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs Retrieval-Augmented Multimodal Language Modeling

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:50:51.598864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:49:15.136031Z digest=sha256:b819f96b0623fa269d52af2f8f0553646ac8e134f30a9f5e0e078e743f193058

Observation a1f8f8e3-7f02-4325-8f2a-2175f202321a · inbound

Differentiable Optimization Layers for Guaranteed Fairness in Deep Learning cites this paper.

Differentiable Optimization Layers for Guaranteed Fairness in Deep Learning Retrieval-Augmented Multimodal Language Modeling

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T15:08:24.781453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T15:06:53.377360Z digest=sha256:1c0769c7604ff623b1cef163eab2b21812d172a713afe1d099d48b9ac5e28b57

Observation 432ce052-a6d2-4896-b08c-5b9c4fc652ee · inbound

Qiskit Code Migration with LLMs cites this paper.

Qiskit Code Migration with LLMs Retrieval-Augmented Multimodal Language Modeling

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-06-26T16:29:35.554353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T16:24:25.357338Z digest=sha256:806caec7ed276542c6ed8b20466791949e48912ef8fc1161df9ec0b7a31a89b9

Observation 448d343d-28bb-44cb-a1e1-046f9f495a9b · inbound

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity cites this paper.

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity Retrieval-Augmented Multimodal Language Modeling

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:30:07.797148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-25T21:15:12.242955Z digest=sha256:9ce35d4c5106023b3e42b64cb2cd3d1bad9b6403ea39e6f6119e91696bceaee3

Observation e3bcc432-69a3-4971-ba94-5b99f908cc33 · inbound

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity cites this paper.

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity Retrieval-Augmented Multimodal Language Modeling

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:51.238082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T05:07:28.537679Z digest=sha256:56e359f40b40cc2fd1b8a09f0f6756d15a932e08d6bfbb1387c39f83d49cf405

Observation 792a6c93-babc-49f8-9a85-20b537027144 · inbound

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation cites this paper.

PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation Retrieval-Augmented Multimodal Language Modeling

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:19:51.071227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T05:16:53.011837Z digest=sha256:afc06ef2f3efa4ce9b73c00a559fd8480b13b4738e89dcbbca4ed60cf9264638

Observation 8e01de3e-ee56-47bc-96a1-57137851a3d4 · inbound

Failing to See or Failing to Know? Attributing Errors in Vision-Language Models cites this paper.

Failing to See or Failing to Know? Attributing Errors in Vision-Language Models Retrieval-Augmented Multimodal Language Modeling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T15:14:25.526254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T15:14:25.526254Z digest=sha256:18fdf78ae6f928e15abd700b040693de8938c8a3aa3576719386163714c78d8b

Observation 870d4085-1188-4380-8f83-5111b41c2d59 · inbound

Failing to See or Failing to Know? Attributing Errors in Vision-Language Models cites this paper.

Failing to See or Failing to Know? Attributing Errors in Vision-Language Models Retrieval-Augmented Multimodal Language Modeling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T06:57:34.446788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:57:34.446788Z digest=sha256:5a0f1bf2afdeb4fe42ca96341cae3e9f3782a4d267947e07d09d1113f17f5fe4