Pith. sign in

Paper Citation Record · LEDGER

CoMemo: LVLMs Need Image Context with Image Memory

As of 11 August 2026, this Paper Citation Record lists 100 of 114 outbound references and 2 inbound Pith citation observations for arXiv:2506.06279.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06279 v1

Coverage vector

measured 100 of 114 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:02:30.468386Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T01:49:15.136031Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T16:01:23.018232Z

Reference resolution

100 of 114 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 79d773ee-f286-40d3-bd8a-4f084176938d · outbound

This paper cites write newline.

CoMemo: LVLMs Need Image Context with Image Memory write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.127567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.127567Z digest=sha256:47cb9ac89bc4cf892b5e302cbd4dae83dd7f71e9e2d7a9f956852d0ed83222f9

Observation 1b37105c-5bfe-48d5-a435-308641a2b8b8 · outbound

This paper cites Nocaps: Novel object captioning at scale.

CoMemo: LVLMs Need Image Context with Image Memory Nocaps: Novel object captioning at scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.132542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.132542Z digest=sha256:31b0cdfab425697f0d07db9480d11da563402118e8919b1d2cc26ed7e1fbd32a

Observation 7646e2aa-2a92-4878-a2de-6d3de73342b7 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

CoMemo: LVLMs Need Image Context with Image Memory Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.136019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.136019Z digest=sha256:6e0d424950d639bf05832bd17be2fe613305596a81e6ee6f03fe4deca981fa21

Observation f750bb2f-0dcb-4ec5-bf3c-732c7d48c15b · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

CoMemo: LVLMs Need Image Context with Image Memory MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.139466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.139466Z digest=sha256:b52b146e5ea0d5c6204f568453bacab493532e1f85ae16380b04291107e2923a

Observation def89704-aae5-45d8-a5ac-d1f50160c601 · outbound

This paper cites A., Datla, V.

CoMemo: LVLMs Need Image Context with Image Memory A., Datla, V

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.143129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.143129Z digest=sha256:f87c16e8e87cdddbfa47040bfbdfcc0950d6b522bd925eb201342a5ae2f0df47

Observation 2e30b908-a5a0-41d9-9b70-8ab804a9e781 · outbound

This paper cites F., Tito, R., Mafla, A., Gomez, L., Rusinol, M., Valveny, E., Jawahar, C., and Karatzas, D.

CoMemo: LVLMs Need Image Context with Image Memory F., Tito, R., Mafla, A., Gomez, L., Rusinol, M., Valveny, E., Jawahar, C., and Karatzas, D

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.146609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.146609Z digest=sha256:75da8e689525878521c647acdfcf625a863e8611f1fad3c6b94813bed2633584

Observation b8d49dac-16a7-4aa6-825c-c39d28fffe1e · outbound

This paper cites Coyo-700m: Image-text pair dataset.

CoMemo: LVLMs Need Image Context with Image Memory Coyo-700m: Image-text pair dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.150059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.150059Z digest=sha256:ee72e0d078689d3affa29a707c2efbe6e6e37597d0c866564145e4060c4d60ec

Observation 8a9ee46c-ef58-4c1f-ae7b-45ef3d6f2469 · outbound

This paper cites and Xiao, J.

CoMemo: LVLMs Need Image Context with Image Memory and Xiao, J

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.154200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.154200Z digest=sha256:0713d3036e4e0f63b9816f8fb71016ec81c59962669019f3dd2bb8089075a3f9

Observation 05435e20-be77-4581-9970-3db59625d8a3 · outbound

This paper cites Textocr-gpt4v.

CoMemo: LVLMs Need Image Context with Image Memory Textocr-gpt4v

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.157676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.157676Z digest=sha256:f40a4bbce39b6d0cacac1d4083f6001fbfdfabb8a1f145b9e9f4fc2e313e017f

Observation ea0cc5cf-0707-45eb-84f8-9d044095792f · outbound

This paper cites MapQA: A Dataset for Question Answering on Choropleth Maps.

CoMemo: LVLMs Need Image Context with Image Memory MapQA: A Dataset for Question Answering on Choropleth Maps

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.160890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.160890Z digest=sha256:e5b4b0d0f64bd313700899b687e7303866c2e565e8a8021da59b94618d66f6c3

Observation e07abd60-3b54-47fa-9573-77bf4aa61664 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

CoMemo: LVLMs Need Image Context with Image Memory ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.164986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.164986Z digest=sha256:c6f82d6875a042083fb06bf88b0487fbf13d67456c2571c496c7bc123df1634a

Observation 59f67add-d782-4305-b878-8925ec32c11e · outbound

This paper cites UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression.

CoMemo: LVLMs Need Image Context with Image Memory UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.168722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.168722Z digest=sha256:c64ecdb25e81dabd5a283e182dd66f716fa4a256631bc27437d7927c1cf657aa

Observation ae2a3243-366e-4ca9-bc2f-638aae7bf3b1 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

CoMemo: LVLMs Need Image Context with Image Memory Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.172306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.172306Z digest=sha256:d8ad3bc6b2cafda4f3c174f054b022957bb2c8f9b6112403ad7e824ee27cd5a1

Observation aa14a1d1-29be-4aeb-89cf-76468b19416b · outbound

This paper cites EVLM: An Efficient Vision-Language Model for Visual Understanding.

CoMemo: LVLMs Need Image Context with Image Memory EVLM: An Efficient Vision-Language Model for Visual Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.175851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.175851Z digest=sha256:21fd82b769896b8e6493056d7126c8d84b93549147d4b97e357eaa54f4f91f03

Observation bbc65fc0-80bd-44e0-a9d6-e376db1066f4 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

CoMemo: LVLMs Need Image Context with Image Memory Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.179318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.179318Z digest=sha256:86ead4daac642fb7e9e10a1dc610f387c6807602dc05946a1bbeef6489fb05fc

Observation de2765e1-0281-4a57-8c39-37fe9e9ebe9e · outbound

This paper cites Complicated Table Structure Recognition.

CoMemo: LVLMs Need Image Context with Image Memory Complicated Table Structure Recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.182794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.182794Z digest=sha256:8ae399c34541c1cdd0ab0282f3cd6db0dd788b85b6501fb460a15bc1f169f07c

Observation 8b42769e-724d-4fc0-9d0c-1b09306a9c0f · outbound

This paper cites K., Liu, Y., Sun, Y., Ng, C.

CoMemo: LVLMs Need Image Context with Image Memory K., Liu, Y., Sun, Y., Ng, C

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.187512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.187512Z digest=sha256:1bbb133e4b4b9abea839c86c852aefc24dd5831d26aa5810105a2251ded41433

Observation cd94c8a4-0aba-436e-a7a4-f13cea61dff6 · outbound

This paper cites Simple and Effective Multi-Paragraph Reading Comprehension.

CoMemo: LVLMs Need Image Context with Image Memory Simple and Effective Multi-Paragraph Reading Comprehension

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.190767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.190767Z digest=sha256:44cd73a61a75e8504cc3169ebfd8bed57de9dab47bc9b867216bc4ed4422d663

Observation 33df9bd5-09bf-4ca6-8a80-9c1818466996 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

CoMemo: LVLMs Need Image Context with Image Memory NVLM: Open Frontier-Class Multimodal LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.194240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.194240Z digest=sha256:cfb0cb5dcd34d150fef8b6171a52c02a66314a24f9e57d4654783542ff4a7b02

Observation ff4a96ac-f9d2-4b61-bde0-f07a2c0002e3 · outbound

This paper cites Deep visual template-free form parsing.

CoMemo: LVLMs Need Image Context with Image Memory Deep visual template-free form parsing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.197809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.197809Z digest=sha256:5891bc6e12160835eaa60cb17f9ddb7a42173a47d657813bee1b0dd30d08e23e

Observation 1af4d204-7229-42ac-9f8c-b5f8dca7c0f4 · outbound

This paper cites The Llama 3 Herd of Models.

CoMemo: LVLMs Need Image Context with Image Memory The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.201168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.201168Z digest=sha256:1bba56d9431506437edfdc7b1fa7f99e4f808a25c98677b8b5cf3d0340535340

Observation 6ae6ab93-5751-4fb2-98a7-9ddc968e5796 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

CoMemo: LVLMs Need Image Context with Image Memory MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.204597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.204597Z digest=sha256:efe3b62ded5a4d4a256d407655786898940c0b1c94bf5ebdc17f607906bf4c53

Observation 763727de-789a-4026-a5f0-6c930777ce2c · outbound

This paper cites A., Ma, W.-C., and Krishna, R.

CoMemo: LVLMs Need Image Context with Image Memory A., Ma, W.-C., and Krishna, R

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.208209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.208209Z digest=sha256:a7329a27ead0be92e25cac9fe9900420c544c2b1aeb88cf045e059c50ee042ca

Observation 4544efaf-0813-4a67-9ba4-cbd5ad03d612 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

CoMemo: LVLMs Need Image Context with Image Memory Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.211435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.211435Z digest=sha256:6ac900b369170b4a6d390f35c06c4e1ad16af946ad44361e80da05dcd7565934

Observation 12fc4599-f778-4b46-a339-87cc2d6b3369 · outbound

This paper cites Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark.

CoMemo: LVLMs Need Image Context with Image Memory Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.214909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.214909Z digest=sha256:e718050c0a45c693d6383842b56f7e54df54b9c297a65da0502c7f5c766aeb81

Observation 52e019b4-ae1e-4d76-a27d-ed70e62f19f6 · outbound

This paper cites Eaten: Entity-aware attention for single shot visual text extraction.

CoMemo: LVLMs Need Image Context with Image Memory Eaten: Entity-aware attention for single shot visual text extraction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.218140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.218140Z digest=sha256:1ccb1bbfb9c1061e1649b5a0a3cf7d3b324ef1b0c01b39d639d5a158d91b3f43

Observation 2430febf-ee02-49cd-b57f-6ac84cc8b47f · outbound

This paper cites Icpr2018 contest on robust reading for multi-type web images.

CoMemo: LVLMs Need Image Context with Image Memory Icpr2018 contest on robust reading for multi-type web images

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.221368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.221368Z digest=sha256:9906a7b011f78babef709e18846cf3e1cce84936cd43f1d56718fe8d2be7f429

Observation 209f2471-82e4-450e-8fe7-45bb600d254e · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

CoMemo: LVLMs Need Image Context with Image Memory PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.224674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.224674Z digest=sha256:5b8d40c10ac228b6db69ebf764499504c6836d1e86c36433d0d7856085dcb9ff

Observation 669db1a4-b54e-49e4-8bfe-4d8b997731fb · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.228260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.228260Z digest=sha256:6fba5f0a9d11f964b57730bbe03a0398635957f41b1250b59763ce318e895106

Observation 6ca98ae8-e705-436f-9692-368ec67430e2 · outbound

This paper cites Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment.

CoMemo: LVLMs Need Image Context with Image Memory Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.231512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.231512Z digest=sha256:0c79eb19a8a5e12e090a199431455ca8bf2d2fdcabce03b82b36634a4f4484b6

Observation 13ce348c-6f4e-4ff0-a75c-36eed2ec4bee · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

CoMemo: LVLMs Need Image Context with Image Memory mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.234729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.234729Z digest=sha256:c0c38167789ddfe53f652155acaffbcce55bb6d72a98dbb3cd4e4e7c93b2f08d

Observation 8c75405c-2f20-4413-a079-1700b6a00b72 · outbound

This paper cites Medical-diff-vqa: a large-scale medical dataset for difference visual question answering on chest x-ray images, 2023.

CoMemo: LVLMs Need Image Context with Image Memory Medical-diff-vqa: a large-scale medical dataset for difference visual question answering on chest x-ray images, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.238109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.238109Z digest=sha256:419bd08022ea150d23440fc00ce27d6eaba0bf7a0f766fd818da97148cc76f30

Observation beed86e7-8ef0-4c38-b874-721f76f683c1 · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

CoMemo: LVLMs Need Image Context with Image Memory Movienet: A holistic dataset for movie understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.241444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.241444Z digest=sha256:2bc8d5f6490633fc4cf2821036bc78d600b381279c5f3626a1e41822de20f7ca

Observation a3283491-6cde-4272-a8e9-6a8b47581690 · outbound

This paper cites Icdar2019 competition on scanned receipt ocr and information extraction.

CoMemo: LVLMs Need Image Context with Image Memory Icdar2019 competition on scanned receipt ocr and information extraction

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.244739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.244739Z digest=sha256:0516821eb7279d56f5e1f53479690fe85478bc052a3c8fcddc1a34939fc03ff6

Observation 4dfa8682-7aa0-4224-8ea3-1847b10f9a60 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.247942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.247942Z digest=sha256:3877b0e3d6b1bae2d1ea7876f100e2bd8689dd9cb4255963e06c06a6e8acc0bd

Observation 174f77cd-8052-47e7-9770-9eba5b8372c6 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

CoMemo: LVLMs Need Image Context with Image Memory MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.251165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.251165Z digest=sha256:c3dafaecf2220d26cd167660278ac779ce4971128174875fae2471b9a5b0f9d1

Observation 80d25d45-dc03-4fc4-a455-9dc0d7bf3c4c · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

CoMemo: LVLMs Need Image Context with Image Memory Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.254523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.254523Z digest=sha256:81ee9c8ad70dcc844ae2b8436b78a10209a10d861f17b2a904e6ad7812b544b3

Observation b5cdb447-51fa-47f8-a580-c592747f2c7e · outbound

This paper cites Dvqa: Understanding data visualizations via question answering.

CoMemo: LVLMs Need Image Context with Image Memory Dvqa: Understanding data visualizations via question answering

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.258111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.258111Z digest=sha256:a6bdf9e05dfe98b28f325e794d4a62ba6f8c4f91566619ba9ec554524990b130

Observation d8be9ef5-2188-4f56-a41a-7957667c4179 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

CoMemo: LVLMs Need Image Context with Image Memory FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.261341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.261341Z digest=sha256:32fafb04fa81c2982d3be652db4ad1e942c7f14eafae01e380f868dbba8f7bf8

Observation c1474023-82b6-4b32-be90-265dbf797445 · outbound

This paper cites Chart-to-Text: A Large-Scale Benchmark for Chart Summarization.

CoMemo: LVLMs Need Image Context with Image Memory Chart-to-Text: A Large-Scale Benchmark for Chart Summarization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.264586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.264586Z digest=sha256:a0107faf37ef27f895234f495759bedda42c239f10f335a0ed658cca421a269b

Observation 8277c4b0-01fe-4f01-b02a-6ea0cf332505 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

CoMemo: LVLMs Need Image Context with Image Memory Referitgame: Referring to objects in photographs of natural scenes

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.268275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.268275Z digest=sha256:0b211ac6491293b8938044d12926648de47b60c83188338ac806e06913fb1e12

Observation 8f63e000-09da-4d75-827b-ff1de325a7e9 · outbound

This paper cites A diagram is worth a dozen images.

CoMemo: LVLMs Need Image Context with Image Memory A diagram is worth a dozen images

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.271379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.271379Z digest=sha256:d70d8d2aa6eb9772dae9ad651e7546f2028d044c7bdbc67015a5341a9cf1d43e

Observation f1cc699c-fcae-4550-87ec-cf8c16069995 · outbound

This paper cites Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension.

CoMemo: LVLMs Need Image Context with Image Memory Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.274553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.274553Z digest=sha256:a24ce523087088d7148c7b0e260865b70cc2a9d959d424df8950e5d9a33228ef

Observation 57ad395e-beeb-4c37-a87b-ed1205ca2522 · outbound

This paper cites Visual information extraction in the wild: practical dataset and end-to-end solution.

CoMemo: LVLMs Need Image Context with Image Memory Visual information extraction in the wild: practical dataset and end-to-end solution

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.278069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.278069Z digest=sha256:35af12e2148e68f70d21fe109691f816693aa3c1d0a69aa3a369fb7e8984ab74

Observation 810c6589-e670-4112-a806-0b314790c112 · outbound

This paper cites Laion-gpt4v dataset.

CoMemo: LVLMs Need Image Context with Image Memory Laion-gpt4v dataset

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.281255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.281255Z digest=sha256:d71c35e0a825cf1c3a2881e1c35d26fbeb6f5e1842a8f7c1ad64e458b421035e

Observation dfc59059-e7aa-4d95-a722-6d38b63e3d24 · outbound

This paper cites J., Gayen, S., Ben Abacha, A., and Demner-Fushman, D.

CoMemo: LVLMs Need Image Context with Image Memory J., Gayen, S., Ben Abacha, A., and Demner-Fushman, D

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.284375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.284375Z digest=sha256:36f801c80bb494ff4f9221bc622a6177913a1a6668fa23a0c772debffc5f5ee1

Observation 9a24351c-1ad4-46be-9297-e4b072587cfb · outbound

This paper cites What matters when building vision-language models?.

CoMemo: LVLMs Need Image Context with Image Memory What matters when building vision-language models?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.287592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.287592Z digest=sha256:8b3527081763f68d4ad125a1b512682fc914a0d952a87c464343bff0c64d026d

Observation cc7246f4-6e8f-4389-a348-e5da0dbb0458 · outbound

This paper cites G., and Lov \'o n Melgarejo, J.

CoMemo: LVLMs Need Image Context with Image Memory G., and Lov \'o n Melgarejo, J

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.290994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.290994Z digest=sha256:efe5143f069f0524add9c138cf57dcf2803f168c9d411e45366ec69bd5a8b640

Observation 36f591ec-7866-4929-8473-e3c6ff0d02ff · outbound

This paper cites Chemvlm: Exploring the power of multimodal large language models in chemistry area.

CoMemo: LVLMs Need Image Context with Image Memory Chemvlm: Exploring the power of multimodal large language models in chemistry area

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.294273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.294273Z digest=sha256:5648f4883174d6ae15ec8799795d46f41aad0a7c4e7c0ebffd7e9b833236c306

Observation 9322a4e8-564e-40c4-a543-064dc566e5e2 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.297412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.297412Z digest=sha256:2bb7c549d6e14d59ba90e333befa95ce51ef8914d973bab5b5378d53d4c5208e

Observation dd1e125a-4572-4e4e-b84a-48a5cf860a21 · outbound

This paper cites Vila: On pre-training for visual language models.

CoMemo: LVLMs Need Image Context with Image Memory Vila: On pre-training for visual language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.300577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.300577Z digest=sha256:fb05eadb17395c3331e1d693ffae02a4f0371a18e5b96f5c91180131c07adf91

Observation 0efdde03-aabb-4c5e-9981-273822210940 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.303815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.303815Z digest=sha256:2d17ac9a5ff61a95611c4d696dcc80a28587c6015609bf35c555cce55822a8c1

Observation 01704a5f-c125-49bc-881a-eb17e9f8b86c · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.307014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.307014Z digest=sha256:cd3b8feea9ed049395bd81afc04589b89f056800ff38c98f3e91dc17effafe17

Observation 99dea663-00ac-46ea-babb-7b35d882486d · outbound

This paper cites Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering.

CoMemo: LVLMs Need Image Context with Image Memory Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.518169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.310104Z digest=sha256:2956d77445fca315ad8981e60c5e6ba821ae4c9f29622e72ed493f5c7a4d93d7

Observation b189077f-4b5e-4840-ad8e-bea1bbe31d7d · outbound

This paper cites Casia online and offline chinese handwriting databases.

CoMemo: LVLMs Need Image Context with Image Memory Casia online and offline chinese handwriting databases

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.506657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.313225Z digest=sha256:8c2232e6953f2fb4ad1ab3348d14a6b504d2d27ece12c19a7fbc6ad6d1b93f57

Observation 946e01d2-f093-45ca-a824-c0bf5ccc3cc5 · outbound

This paper cites Visual spatial reasoning.

CoMemo: LVLMs Need Image Context with Image Memory Visual spatial reasoning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.495266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.316435Z digest=sha256:6cf821590df72c981ce9869a81228d66163ddae6a2bc00e6b0470211505df30a

Observation b07f0ca7-744d-4623-b1e8-e0f98c805451 · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

CoMemo: LVLMs Need Image Context with Image Memory Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.484152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.319716Z digest=sha256:a4302602984a8287becfa489d318508d9dd3e17bec8a01cc601efd559cb6cb83

Observation 17c71ad7-9f4f-4159-9ce9-1a6764bd694f · outbound

This paper cites MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning.

CoMemo: LVLMs Need Image Context with Image Memory MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.322925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.322925Z digest=sha256:f7aa986ff4e34afefaa0101306bb71013f71604468a1e1b58d6cbb985e89f619

Observation f00e6b2e-9063-4c1a-8cd1-2a7c6a92d1a5 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.326278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.326278Z digest=sha256:0f3a1814297b2a9606c20ebc75187ed8d5c35fec1f73286b581dc8560d1beef2

Observation 604930f5-0b88-4c59-95d0-1ee410d9de32 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.329367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.329367Z digest=sha256:5c019f7aed51000f4ee3a8cd8b9d053d3a1f90b4f2dc30f7d02c734180618fdb

Observation 8546b904-8a57-41ab-a652-42fed20649b7 · outbound

This paper cites F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P.

CoMemo: LVLMs Need Image Context with Image Memory F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.332532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.332532Z digest=sha256:b483490c2603ac359e073ade09523d575d96c0aeddeb86ea19fcd232bcfbed85

Observation 90241c51-b471-4207-a6b0-416a39b7b9e8 · outbound

This paper cites Paying more attention to image: A training-free method for alleviating hallucination in lvlms.

CoMemo: LVLMs Need Image Context with Image Memory Paying more attention to image: A training-free method for alleviating hallucination in lvlms

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.449278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.335865Z digest=sha256:696472bccbe1d851f3930ba2779448c76be60d88a2a1be1368f270b815a7691e

Observation 733505ed-4a4d-490f-9430-7783d7a73374 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pp.\ 216--233.

CoMemo: LVLMs Need Image Context with Image Memory Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pp.\ 216--233

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.339582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.339582Z digest=sha256:4285829d146c68d8bf40867a301bcaf37a790aea32b58abc8aefe5bd3d550220

Observation 5695cc77-06fc-4aca-a66e-e9a988521821 · outbound

This paper cites MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs.

CoMemo: LVLMs Need Image Context with Image Memory MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.346839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.346839Z digest=sha256:5d1900551b986b04eca389f8294e34a56d358cd26de8cba5906c4f54257c0bdb

Observation e3b3ff6b-8025-4c8a-98d7-14cdf5e688b2 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

CoMemo: LVLMs Need Image Context with Image Memory Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.350057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.350057Z digest=sha256:794b51789a425853da3e1f0d935cc55f50ac089d878d19ec919ad1a5b0c23aa8

Observation 6e92e147-407f-4e5b-8b8f-1e49ef1e8d61 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

CoMemo: LVLMs Need Image Context with Image Memory Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.429290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.353502Z digest=sha256:537d00700967f6e9950382dadfa2d22936e96c9ce1aa0ba38234eaed048dd066

Observation 4f30c604-bac2-41b0-9774-f0e2832b62cd · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

CoMemo: LVLMs Need Image Context with Image Memory Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.357233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.357233Z digest=sha256:c4d8d6595db985c7023529392e42beec3525747dda8412980116cccec0316037

Observation 156f6993-0a26-418b-bb6f-867c05b90e02 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

CoMemo: LVLMs Need Image Context with Image Memory MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.361051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.361051Z digest=sha256:efeb1f391b6ceb84ac39b400707badde855f28f173c65e980fdb81c7b07a37e7

Observation f661c154-e39b-42b0-980a-4f0457f3041a · outbound

This paper cites Deepart: Learning joint representations of visual arts.

CoMemo: LVLMs Need Image Context with Image Memory Deepart: Learning joint representations of visual arts

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.418177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.364523Z digest=sha256:4b9e7a85dfb197a5bcb1210f8b09c96aba6832fe2583f74a1c070f84d0d0782a

Observation c77c4483-ffa4-4447-b75b-c3042ff0f6a1 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

CoMemo: LVLMs Need Image Context with Image Memory Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.406702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.367657Z digest=sha256:ee6c299985b777e1ec5555be2345557b08c82c8e03cc4cc292cb988f900fc7c2

Observation dfa4350b-4f3c-46ca-9875-93bcd26166eb · outbound

This paper cites and Bunke, H.

CoMemo: LVLMs Need Image Context with Image Memory and Bunke, H

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.394019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.370785Z digest=sha256:ed0a20767a70e56af5bf244ff7f42a6578d5d50f8d2513c8b53bd32d78d1b202

Observation 18c6213e-1545-47cb-916a-ab4fe68274be · outbound

This paper cites L., Tan, J.

CoMemo: LVLMs Need Image Context with Image Memory L., Tan, J

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.382660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.374097Z digest=sha256:7a0f9f07ea1c468002149df9b97c60b518a687fe986a3cdc193d312588c39fc0

Observation 137e982f-9a0d-49a5-8a9e-134d0c8a0659 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

CoMemo: LVLMs Need Image Context with Image Memory ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.377251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.377251Z digest=sha256:9536d46d5b2cb6e55da1d4019f15fa5d38a558d5d6923527085989702884601b

Observation be2f4648-53c9-4e91-8ab0-fc059295805c · outbound

This paper cites UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning.

CoMemo: LVLMs Need Image Context with Image Memory UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.380654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.380654Z digest=sha256:cc85775c230d885cd5ee65dd9550f03ce5beb798415d8af639a398b66817a098

Observation e6bcc7c4-897b-429b-ad41-109aa8beff90 · outbound

This paper cites Infographicvqa.

CoMemo: LVLMs Need Image Context with Image Memory Infographicvqa

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.371284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.384119Z digest=sha256:3215940634a4b3809e7d8f34d7db9f42b5df832f974e67ecd49ee182f5e668cb

Observation 9ff00e35-d786-456c-8814-5f0078660c84 · outbound

This paper cites M., and Kumar, P.

CoMemo: LVLMs Need Image Context with Image Memory M., and Kumar, P

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.360029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.387472Z digest=sha256:deaa49db5e6481b63a7283f5471e02712225f456e491e5d2c83976863b81ee12

Observation 3ecfc09d-1ebb-4353-a622-1da18e7a13fe · outbound

This paper cites Opengvlab/internvl-chat-v1-2-sft-data, Jan 2024.

CoMemo: LVLMs Need Image Context with Image Memory Opengvlab/internvl-chat-v1-2-sft-data, Jan 2024

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.348665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.391046Z digest=sha256:cc314036a4efa2a4e07e53010413263208ea7095f5f7c19ed514ec0b041cc2c3

Observation 63479621-6e1a-4838-8694-93568d5fa6e7 · outbound

This paper cites Training language models to follow instructions with human feedback.

CoMemo: LVLMs Need Image Context with Image Memory Training language models to follow instructions with human feedback

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.394227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.394227Z digest=sha256:f5d1f1e23f4428eda00049fcd45b45a9cec929e1d23e82453393f5b0e44168f0

Observation 0c1e3acc-6270-4b13-aa75-e1b031b50306 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

CoMemo: LVLMs Need Image Context with Image Memory Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.397502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.397502Z digest=sha256:165de59df8c81d12504ad56fa92710637ec9ef59aceafae44e5844820673c8e5

Observation 01a410e3-8cf4-4fae-b973-5d36686e14b5 · outbound

This paper cites A., Wang, L., Cervantes, C.

CoMemo: LVLMs Need Image Context with Image Memory A., Wang, L., Cervantes, C

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.329398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.400986Z digest=sha256:226aca9e7ee5870de70de060337fdab82bc538ce1f722bf44ed81147444e96ba

Observation 1cc2ab6b-5754-4d27-8241-182a2418c82a · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

CoMemo: LVLMs Need Image Context with Image Memory Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.318852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.404242Z digest=sha256:3d9fc836d8aa603c95c154be91310c11e1dd191354700102ed20eca20bd12a59

Observation bb4b8522-8214-418c-a2ff-1c533289bb58 · outbound

This paper cites Laion coco: 600m synthetic captions from laion2b-en.

CoMemo: LVLMs Need Image Context with Image Memory Laion coco: 600m synthetic captions from laion2b-en

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.307704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.407507Z digest=sha256:3588855a53cf0234c7fb1d5ebacdf3d3e90af0373767f3d0b9e18340b5eabcf5

Observation 1b84a50c-ef90-4ba9-bdef-b3783f87e3ad · outbound

This paper cites Solving geometry problems: Combining text and diagram interpretation.

CoMemo: LVLMs Need Image Context with Image Memory Solving geometry problems: Combining text and diagram interpretation

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.296729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.410789Z digest=sha256:53030bd6ee550f73729230f5cb16afa75a34a6746f077bbff8565e9593780fdd

Observation fbe9a7b8-59c9-41b3-99d9-a4dfb01ad904 · outbound

This paper cites an unresolved cited work.

CoMemo: LVLMs Need Image Context with Image Memory Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:02:31.285582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.414053Z digest=sha256:a18fbe7b20f918494aa7cb8c0db7e0a0560f90c9d9629554fabde3e1a9343ea1

Observation 1a3ad93e-ed8e-4a09-9d05-0bf9b7515aec · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

CoMemo: LVLMs Need Image Context with Image Memory Objects365: A large-scale, high-quality dataset for object detection

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.274988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.417252Z digest=sha256:a47711a5e713dc202b5aea0ffebf65aebdbe16852c0cddffd909ff08eb8e1536

Observation 0a2e401d-22de-4408-b203-8ce576a0d7e5 · outbound

This paper cites Towards vqa models that can read.

CoMemo: LVLMs Need Image Context with Image Memory Towards vqa models that can read

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.264334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.420406Z digest=sha256:35a4875f8c8b315c66a98ba44e03890a1c31b886e1b49ac095ab2c2872fa5c6b

Observation be159f74-34f0-4efd-9ca0-9bfb09e11a71 · outbound

This paper cites Towards vqa models that can read.

CoMemo: LVLMs Need Image Context with Image Memory Towards vqa models that can read

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.252411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.423527Z digest=sha256:90a2452e342bc3b246bc7230a867ea56ced991e2d37bc1b2eef0bdeaefc9be37

Observation 3179f98c-0cf1-44ad-ab44-ee10b2d0dbce · outbound

This paper cites Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text.

CoMemo: LVLMs Need Image Context with Image Memory Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.241606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.426830Z digest=sha256:6df406df5fc38e41798234f41cd18844ef0e69ed416d37ec7a0b04db3386badf

Observation d554ce3f-0191-4126-aca7-caf3d993cddc · outbound

This paper cites MileBench: Benchmarking MLLMs in Long Context.

CoMemo: LVLMs Need Image Context with Image Memory MileBench: Benchmarking MLLMs in Long Context

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.430076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.430076Z digest=sha256:22a84f7b873a7d59ae6cb855a9c254ad2f1448937e73fca730a1d1b8fc1c7a43

Observation 15374b95-0c4e-410f-af9a-b0853b967627 · outbound

This paper cites Transformer roadmap: 2.

CoMemo: LVLMs Need Image Context with Image Memory Transformer roadmap: 2

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.229796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.433564Z digest=sha256:f08b78526fe0f670abd18f1ec1db17bb623ab26b19e8d90386328e9f05c7f710

Observation 9d98cf6c-e5e3-4c5c-980b-b5d1a4852d4f · outbound

This paper cites C., Han, J., Ding, E., Liu, J., Karatzas, D., et al.

CoMemo: LVLMs Need Image Context with Image Memory C., Han, J., Ding, E., Liu, J., Karatzas, D., et al

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.218537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.436705Z digest=sha256:a7a0f3e58dcc5b02ef7c0cdc488e2e6d5478b0a1786fd26f2ba2f1cfb81035bf

Observation 18ec7216-01ed-42d5-85ed-341a4c6cd833 · outbound

This paper cites Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy, 2024.

CoMemo: LVLMs Need Image Context with Image Memory Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy, 2024

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.207757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.439836Z digest=sha256:3d8824717d7fd39ea406414da37aea8e44ff0ae294b2bac1ff016232663001b4

Observation 295ed525-12db-4af6-bcb7-6dfc706f61ef · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

CoMemo: LVLMs Need Image Context with Image Memory Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.444417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.444417Z digest=sha256:54be00465943552b8fdb3bc92b298d606c4e0c45d2f7d75367710db45d17c478

Observation a1f5b464-6f14-461f-aedd-9d6cbb42dc60 · outbound

This paper cites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images.

CoMemo: LVLMs Need Image Context with Image Memory COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.447765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.447765Z digest=sha256:2bbc3b2ffc223975917ef57a1da84c9fd8c8eb66b5727dade50a7c865e9f2d08

Observation 9f93a81d-03ac-4199-8d57-73741ee9121c · outbound

This paper cites V3det: Vast vocabulary visual detection dataset.

CoMemo: LVLMs Need Image Context with Image Memory V3det: Vast vocabulary visual detection dataset

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.196589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.451246Z digest=sha256:5c5f40ebcf7f93eed5d17c43826abdb9c252de7013626e438f3ce1a344fa2421

Observation ae545fb5-9c31-4421-9b57-625e6aeba795 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

CoMemo: LVLMs Need Image Context with Image Memory Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.454664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.454664Z digest=sha256:4eb1b25c7ad28502ec2f1ae752234ba93329700229a6cdcca93c79fcbdc05339

Observation 6c69544c-f78e-4907-bb82-9e20ca304583 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CoMemo: LVLMs Need Image Context with Image Memory Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.458205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.458205Z digest=sha256:22c6b696602489e1939e3db5c0133f87492622d2d8ae5614fae41f77f6238f39

Observation 74ab545b-117c-4047-824a-57ee643cddd5 · outbound

This paper cites The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World.

CoMemo: LVLMs Need Image Context with Image Memory The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.461439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.461439Z digest=sha256:43e455b46841fe29c640b2074779ff1c7d665a4c57f7d3453adf056a28dd00ea

Observation 4f7dfd5e-5b7a-4f5b-90cb-ac18b352031d · outbound

This paper cites Needle In A Multimodal Haystack.

CoMemo: LVLMs Need Image Context with Image Memory Needle In A Multimodal Haystack

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.464875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.464875Z digest=sha256:0d57f91d2661c981d5b3bd7347c1179358cb156962fe0132dbdbec0615bcb797

Observation 60177022-5d53-47f2-baf6-075177bef644 · outbound

This paper cites C., Luo, C., Jin, L., Chan, C.

CoMemo: LVLMs Need Image Context with Image Memory C., Luo, C., Jin, L., Chan, C

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:02:31.184921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T06:02:30.468386Z digest=sha256:fcb8d91e7ca00bf14313aee0c1c05ec7d35e7ffcf3f334d0415fec3ab4fec620

Pith citing papers

Observation 141c552d-38db-4546-9c1b-cf0bfe83f149 · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs CoMemo: LVLMs Need Image Context with Image Memory

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:23.021630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:53:06.494640Z digest=sha256:44736b7d5db474dff15045c4c38745064246fc894eb88995c239cd3ad0ae6d9e

Observation a6c78e64-9adc-4225-94c0-35ecef142c6c · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs CoMemo: LVLMs Need Image Context with Image Memory

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:50:51.525657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:49:15.136031Z digest=sha256:b22e2378852ba04555ed20a485c0035c35a3c774f32cd2c6f79d7ad979660d03