Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:46:48.069465Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 7 inbound Pith citation observations for arXiv:2505.17670.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:46:48.069465Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T20:06:09.180245Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T12:46:14.734852Z
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b3ca604c-c72c-4282-8495-ee246efcc922 · outbound
Towards General Continuous Memory for Vision-Language Models Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 278f4724-6492-4490-b066-f9d925d37e1b · outbound
Towards General Continuous Memory for Vision-Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7dab02d-da0b-4b10-b0cc-2a045714eae9 · outbound
Towards General Continuous Memory for Vision-Language Models Language models show human-like content effects on reasoning tasks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9182548-547a-48f2-9993-f7c7ca9a5b69 · outbound
Towards General Continuous Memory for Vision-Language Models Large Language Models for Mathematical Reasoning: Progresses and Challenges
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c45a2d7-0451-47f5-9157-735f62d6c4ca · outbound
Towards General Continuous Memory for Vision-Language Models Chain-of-thought prompting elicits reasoning in large language models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76db0541-0b53-45f3-b104-aa73d6a6cdeb · outbound
Towards General Continuous Memory for Vision-Language Models A Survey on Large Language Models for Code Generation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19e41b4e-40f0-450f-bcd8-869d35eab58b · outbound
Towards General Continuous Memory for Vision-Language Models Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9360e58-4bc3-4631-9051-5391a6c118bc · outbound
Towards General Continuous Memory for Vision-Language Models Retrieval-augmented generation for knowledge-intensive nlp tasks
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a68b3edb-9eba-4b8d-ae60-3c7f152e9ada · outbound
Towards General Continuous Memory for Vision-Language Models Augmenting Language Models with Long-Term Memory
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 578c124b-2444-4e26-9170-fd8bad3a10bd · outbound
Towards General Continuous Memory for Vision-Language Models Realm: Retrieval-augmented language model pre-training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dc929dfb-907b-4d47-b768-f72481c90c48 · outbound
Towards General Continuous Memory for Vision-Language Models Qwen2.5-VL Technical Report
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1bfa9a7-30c5-4c65-87c5-1b36405c4fb8 · outbound
Towards General Continuous Memory for Vision-Language Models VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe30ac93-9f60-40c0-869a-46dda89423ae · outbound
Towards General Continuous Memory for Vision-Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52cb878f-a48f-48f0-a17a-ff3fde0353f9 · outbound
Towards General Continuous Memory for Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce541d28-98a3-46e5-9886-83af0d845f2f · outbound
Towards General Continuous Memory for Vision-Language Models A mathematical framework for transformer circuits.Transformer Circuits Thread, 2021.https://transformer-circuits.pub/2021/framework/index.html
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a3810cb5-551b-4efc-90e1-76fbca416c3e · outbound
Towards General Continuous Memory for Vision-Language Models In-context Learning and Induction Heads
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1fad45e-028b-4846-a9f7-d268c0862a81 · outbound
Towards General Continuous Memory for Vision-Language Models Training Large Language Models to Reason in a Continuous Latent Space
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f70f8f0-2107-458d-9953-dc94385df636 · outbound
Towards General Continuous Memory for Vision-Language Models Perceiver IO: A General Architecture for Structured Inputs & Outputs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41e3e058-dcf4-4e39-a23a-39d94734ca0d · outbound
Towards General Continuous Memory for Vision-Language Models Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, and Simon Lacoste-Julien
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ade8e379-f223-4f6c-8e2d-a94d7ea002a7 · outbound
Towards General Continuous Memory for Vision-Language Models The trade-offs of domain adaptation for neural language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 55695af3-c974-40c8-8108-c350d956c877 · outbound
Towards General Continuous Memory for Vision-Language Models Attention Is All You Need
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 962c04ef-2a93-492b-8a05-7393312cf113 · outbound
Towards General Continuous Memory for Vision-Language Models What Does BERT Look At? An Analysis of BERT's Attention
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37383aeb-aba5-439d-a240-f36929e4a534 · outbound
Towards General Continuous Memory for Vision-Language Models LoRA: Low-Rank Adaptation of Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b26f56e-ef01-4ef6-afc2-d7fe685e034b · outbound
Towards General Continuous Memory for Vision-Language Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8efb0f53-c2e0-43c8-9e89-beba86e99e88 · outbound
Towards General Continuous Memory for Vision-Language Models EchoSight: Advancing Visual-Language Models with Wiki Knowledge
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8109059b-35dd-4671-8911-22da5803b3ed · outbound
Towards General Continuous Memory for Vision-Language Models Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f41da4d5-2059-4b1f-bb89-b810023dae84 · outbound
Towards General Continuous Memory for Vision-Language Models RoRA-VLM: Robust Retrieval-Augmented Vision Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed55e627-01c3-4b98-bb44-a9d824170bfd · outbound
Towards General Continuous Memory for Vision-Language Models xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b46d3d36-0c01-4465-ad51-30287da5651b · outbound
Towards General Continuous Memory for Vision-Language Models KV-Distill: Nearly Lossless Learnable Context Compression for LLMs
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03a1e32b-abe3-466e-a411-0abe0a4039db · outbound
Towards General Continuous Memory for Vision-Language Models VoCo-LLaMA: Towards Vision Compression with Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eb46eb0-8e47-4268-81dc-38ff53e2fe95 · outbound
Towards General Continuous Memory for Vision-Language Models MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceb034a8-c31f-4dc9-a8c5-183750514798 · outbound
Towards General Continuous Memory for Vision-Language Models M+: Extending MemoryLLM with Scalable Long-Term Memory
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd50865a-ef74-4eca-b4bb-25b067d1503b · outbound
Towards General Continuous Memory for Vision-Language Models MemGPT: Towards LLMs as Operating Systems
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16a8baad-8059-40d3-af0d-6ecc6d78427f · outbound
Towards General Continuous Memory for Vision-Language Models Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c86d863a-91bf-48be-aebf-2140ac37ae56 · outbound
Towards General Continuous Memory for Vision-Language Models Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5013a54-ceaf-4073-bf67-3fb52dcb7ccc · outbound
Towards General Continuous Memory for Vision-Language Models A-okvqa: A benchmark for visual question answering using world knowledge
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3bae3da7-054f-4a19-89dc-c3541920b6fd · outbound
Towards General Continuous Memory for Vision-Language Models Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 84408c59-a160-45c6-b174-057ff4613ccb · outbound
Towards General Continuous Memory for Vision-Language Models Learning transferable visual models from natural language supervision
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a40a3f2-b76e-4de8-841e-312eea322af1 · outbound
Towards General Continuous Memory for Vision-Language Models Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5e49f585-c4bc-4e47-b17c-275a90d4551a · outbound
Towards General Continuous Memory for Vision-Language Models Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ea0a8ad6-fd27-4b25-a64c-1fe69f0e084d · outbound
Towards General Continuous Memory for Vision-Language Models Open-domain visual entity recognition: Towards recognizing millions of wikipedia entities
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53986ec4-e589-4e80-a516-5f5bbe5bfa7a · outbound
Towards General Continuous Memory for Vision-Language Models MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e49340-0ac3-4d5c-b74a-89d0f684452c · outbound
Towards General Continuous Memory for Vision-Language Models Viquae, a dataset for knowledge-based visual question answering about named entities
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9da75c24-78c1-4c9f-a009-e9e5cef09de6 · outbound
Towards General Continuous Memory for Vision-Language Models CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e1c4b10-3cb9-435e-a872-12f2ddef8860 · outbound
Towards General Continuous Memory for Vision-Language Models Gpt-4o technical report, 2024
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0c0b63b3-180b-4a61-a0f8-e443ebfa1fd4 · outbound
Towards General Continuous Memory for Vision-Language Models Visual Instruction Tuning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c17250c3-a35b-4426-8a90-8dfa58cf63ff · outbound
Towards General Continuous Memory for Vision-Language Models Llava-next: Open large multimodal models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6688bed2-a55f-492c-af4f-e69bfeb2c54f · outbound
Towards General Continuous Memory for Vision-Language Models InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cc38497-8a40-401c-ad73-95df2674128c · outbound
Towards General Continuous Memory for Vision-Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6adb41ec-9728-4a2b-ba4a-65269eeac399 · outbound
Towards General Continuous Memory for Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71eb6c5c-9535-4755-8397-d6678dd7535c · outbound
Towards General Continuous Memory for Vision-Language Models Qwen2.5 Technical Report
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51587dfb-240b-437c-a0b9-2ad31a711c0f · outbound
Towards General Continuous Memory for Vision-Language Models A simple framework for contrastive learning of visual representations
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 076259f0-863f-4afb-b3b4-2af3a8426ac0 · outbound
Towards General Continuous Memory for Vision-Language Models Mm1: methods, analysis and insights from multimodal llm pre-training
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3b92a26-f366-4a72-b07e-6a44d775a3ef · outbound
Towards General Continuous Memory for Vision-Language Models Ulip-2: Towards scalable multimodal pre-training for 3d understanding
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f310d7f9-4473-4ad9-a8d6-dedaed9b76d3 · outbound
Towards General Continuous Memory for Vision-Language Models Improved baselines with visual instruction tuning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc6558bf-75bb-4a7a-86d8-ed1b32ac341d · outbound
Towards General Continuous Memory for Vision-Language Models Learning to compress prompts with gist tokens
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 458aabc8-aea5-4415-a175-1c59dd54f241 · outbound
Towards General Continuous Memory for Vision-Language Models In-Context Former: Lightning-fast Compressing Context for Large Language Model
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2d1aa9f7-4054-450e-83c6-9d10d4a87332 · outbound
Towards General Continuous Memory for Vision-Language Models Adapting llms for efficient context processing through soft prompt compression
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation af99f633-c4e5-4982-91d7-ec761ca23a6e · outbound
Towards General Continuous Memory for Vision-Language Models The probabilistic relevance framework: Bm25 and beyond.F oundations and Trends in Information Retrieval, 3(4):333–389, 2009
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d2222c44-7745-4fec-99f5-5a4fc3d02b47 · inbound
Recurrence Meets Transformers for Universal Multimodal Retrieval Towards General Continuous Memory for Vision-Language Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6642783a-77c2-42f9-a7e3-0c1bf2e562c8 · inbound
Dual Latent Memory for Visual Multi-agent System Towards General Continuous Memory for Vision-Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7657eba7-d3e1-4569-ae44-7480842fb4a5 · inbound
The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook Towards General Continuous Memory for Vision-Language Models
Reference 243
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef4b830d-e506-4531-8d95-c8ed1bab1ac1 · inbound
Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning Towards General Continuous Memory for Vision-Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8ea63aaf-ab62-4406-ae02-ad1908a9dbfc · inbound
Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting Towards General Continuous Memory for Vision-Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 641cd9fc-40c1-475b-b9c1-50db348e4452 · inbound
MMAgent-R$^2$: Learning to Rerank and Reject for Agentic mRAG Towards General Continuous Memory for Vision-Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8fa89f1a-d668-4203-9b84-dd6f543a8df5 · inbound
Reason Before You Retrieve: Agentic Planning for Multi-modal RAG Towards General Continuous Memory for Vision-Language Models
Reference 162
Source-reported events for the cited work
Unavailable: canonical work link unavailable.