Pith. sign in

Paper Citation Record · LEDGER

Towards General Continuous Memory for Vision-Language Models

As of 10 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 7 inbound Pith citation observations for arXiv:2505.17670.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17670 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:46:48.069465Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:06:09.180245Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T12:46:14.734852Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact2
  • verified fuzzy14
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3ca604c-c72c-4282-8495-ee246efcc922 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

Towards General Continuous Memory for Vision-Language Models Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:41.872722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:41.872722Z digest=sha256:6365d9a3ee64d1422cd1a5a2c9b6d2cbd3253dbac383bbe7ea9e99126bbeeead

Observation 278f4724-6492-4490-b066-f9d925d37e1b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Towards General Continuous Memory for Vision-Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:41.988564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:41.988564Z digest=sha256:df358fe70c6594766f6a812c9d1f6e6da2a05f3076b1bfbb16a4df4a32064fc7

Observation c7dab02d-da0b-4b10-b0cc-2a045714eae9 · outbound

This paper cites Language models show human-like content effects on reasoning tasks.

Towards General Continuous Memory for Vision-Language Models Language models show human-like content effects on reasoning tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:42.111101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:42.111101Z digest=sha256:097d4a8b40ef8002c6c957c57ade66ba886c92365e0e513b47f7ab42e45515d3

Observation e9182548-547a-48f2-9993-f7c7ca9a5b69 · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

Towards General Continuous Memory for Vision-Language Models Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:42.184846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:42.184846Z digest=sha256:461683730c9bfb6ae6ec7676a673a8f552b5e8be5f2b9b6e3406dd3d152e8182

Observation 6c45a2d7-0451-47f5-9157-735f62d6c4ca · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Towards General Continuous Memory for Vision-Language Models Chain-of-thought prompting elicits reasoning in large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:42.337052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:42.337052Z digest=sha256:7d35fde074cb4c85cfc8af0dd5d059da57f6fe913724454e79f7be2ad44dc0c0

Observation 76db0541-0b53-45f3-b104-aa73d6a6cdeb · outbound

This paper cites A Survey on Large Language Models for Code Generation.

Towards General Continuous Memory for Vision-Language Models A Survey on Large Language Models for Code Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:42.431189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:42.431189Z digest=sha256:7d5c743a16eb552ee81bd5d5e29df49b79cd1fcab5b5a6dec667553ba8d42562

Observation 19e41b4e-40f0-450f-bcd8-869d35eab58b · outbound

This paper cites Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?.

Towards General Continuous Memory for Vision-Language Models Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:42.517303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:42.517303Z digest=sha256:edc463bf36677491e6f05ffeecebed24ab82fc466d63aa0d71a4eef46531a8c6

Observation f9360e58-4bc3-4631-9051-5391a6c118bc · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

Towards General Continuous Memory for Vision-Language Models Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:46:52.115639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:46:42.593852Z digest=sha256:99e7e348cec961d348fc33d393020c08c4a6d4bebfd779e4dfe8bcd01fc80737

Observation a68b3edb-9eba-4b8d-ae60-3c7f152e9ada · outbound

This paper cites Augmenting Language Models with Long-Term Memory.

Towards General Continuous Memory for Vision-Language Models Augmenting Language Models with Long-Term Memory

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:42.716548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:42.716548Z digest=sha256:1426f32aa1da39de4f905a8d129eef36a8d048406f97d57d75060b9931d9ef89

Observation 578c124b-2444-4e26-9170-fd8bad3a10bd · outbound

This paper cites Realm: Retrieval-augmented language model pre-training.

Towards General Continuous Memory for Vision-Language Models Realm: Retrieval-augmented language model pre-training

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:46:51.943219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:46:42.790901Z digest=sha256:110eccdf13fd0da695882f150b3fa2368762370343a1071d3c14c9ee1a34e486

Observation dc929dfb-907b-4d47-b768-f72481c90c48 · outbound

This paper cites Qwen2.5-VL Technical Report.

Towards General Continuous Memory for Vision-Language Models Qwen2.5-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:42.912983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:42.912983Z digest=sha256:c7a3484fa22ca00dca8622c031a25c06362f85ea050d06551482cc413534b385

Observation a1bfa9a7-30c5-4c65-87c5-1b36405c4fb8 · outbound

This paper cites VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.

Towards General Continuous Memory for Vision-Language Models VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:42.990513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:42.990513Z digest=sha256:75ef276a4228a93ebdecdb4ada96f7d4c9d7a54caaee66dfa1f6924823691906

Observation fe30ac93-9f60-40c0-869a-46dda89423ae · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Towards General Continuous Memory for Vision-Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:43.067006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:43.067006Z digest=sha256:41d037d37b90cd888abd8784433f08c4716e2d542a877b8a1ca659a54e7ba379

Observation 52cb878f-a48f-48f0-a17a-ff3fde0353f9 · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

Towards General Continuous Memory for Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:43.129145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:43.129145Z digest=sha256:b88b7dd5138789233915de26efb8a2b395027309ffa778f14d310838002c1403

Observation ce541d28-98a3-46e5-9886-83af0d845f2f · outbound

This paper cites A mathematical framework for transformer circuits.Transformer Circuits Thread, 2021.https://transformer-circuits.pub/2021/framework/index.html.

Towards General Continuous Memory for Vision-Language Models A mathematical framework for transformer circuits.Transformer Circuits Thread, 2021.https://transformer-circuits.pub/2021/framework/index.html

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:46:51.738013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:46:43.199452Z digest=sha256:1dfa066bee200518aceb9c8e465fa5e8cd7fdf9a6b7583fe1f570891313c7c49

Observation a3810cb5-551b-4efc-90e1-76fbca416c3e · outbound

This paper cites In-context Learning and Induction Heads.

Towards General Continuous Memory for Vision-Language Models In-context Learning and Induction Heads

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:43.352571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:43.352571Z digest=sha256:d0bf5aa34d6120be7445c9336767b2d3c60f46afdcd6eb6209c745dd8503b819

Observation c1fad45e-028b-4846-a9f7-d268c0862a81 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Towards General Continuous Memory for Vision-Language Models Training Large Language Models to Reason in a Continuous Latent Space

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:43.470015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:43.470015Z digest=sha256:3d10ca28227368f47e6142e67b0014a7d3c867cfaa4f6768383aeabbedb9aadf

Observation 5f70f8f0-2107-458d-9953-dc94385df636 · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

Towards General Continuous Memory for Vision-Language Models Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:43.619890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:43.619890Z digest=sha256:384858b76928cce4d944397b8da20802c87b582ea59e3ab7784b14ada9530d3f

Observation 41e3e058-dcf4-4e39-a23a-39d94734ca0d · outbound

This paper cites Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, and Simon Lacoste-Julien.

Towards General Continuous Memory for Vision-Language Models Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, and Simon Lacoste-Julien

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:46:51.527275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:46:43.739351Z digest=sha256:dcb4cce04dd8c839a38dd9cae12d3e4ca6edb62dce27f03b45a63507b5c0a301

Observation ade8e379-f223-4f6c-8e2d-a94d7ea002a7 · outbound

This paper cites The trade-offs of domain adaptation for neural language models.

Towards General Continuous Memory for Vision-Language Models The trade-offs of domain adaptation for neural language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:46:51.349093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:46:43.816405Z digest=sha256:7b44bdb77b6e64ee7a49d7b1633f52e1e45f752d8df295df47f1f4ca8ca41648

Observation 55695af3-c974-40c8-8108-c350d956c877 · outbound

This paper cites Attention Is All You Need.

Towards General Continuous Memory for Vision-Language Models Attention Is All You Need

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:43.886123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:43.886123Z digest=sha256:2d8c86e70449b5fb15865bde06521c307856698451eff8efee39b7a9f3e94ced

Observation 962c04ef-2a93-492b-8a05-7393312cf113 · outbound

This paper cites What Does BERT Look At? An Analysis of BERT's Attention.

Towards General Continuous Memory for Vision-Language Models What Does BERT Look At? An Analysis of BERT's Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:43.963486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:43.963486Z digest=sha256:a321e51891e59ec436dc2016e28589765ee037bfe578516e22e0f150c39ab299

Observation 37383aeb-aba5-439d-a240-f36929e4a534 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Towards General Continuous Memory for Vision-Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.075042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.075042Z digest=sha256:86d3ec31adecc3a260b3608279ce6efeba59113ed0661716908457a6331600bc

Observation 1b26f56e-ef01-4ef6-afc2-d7fe685e034b · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Towards General Continuous Memory for Vision-Language Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.189348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.189348Z digest=sha256:ea43915665417a349650b034139e4319993290445da8a98ea6477da284b50457

Observation 8efb0f53-c2e0-43c8-9e89-beba86e99e88 · outbound

This paper cites EchoSight: Advancing Visual-Language Models with Wiki Knowledge.

Towards General Continuous Memory for Vision-Language Models EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.260422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.260422Z digest=sha256:251190a0cd632a4847d4ab775900e71e863b676710e31f62ebd4e9912fd14d8e

Observation 8109059b-35dd-4671-8911-22da5803b3ed · outbound

This paper cites Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering.

Towards General Continuous Memory for Vision-Language Models Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.356720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.356720Z digest=sha256:9c1da26b9b1e939d8c8e1746d44a019fa9dfe2112689f6a12554c18261166467

Observation f41da4d5-2059-4b1f-bb89-b810023dae84 · outbound

This paper cites RoRA-VLM: Robust Retrieval-Augmented Vision Language Models.

Towards General Continuous Memory for Vision-Language Models RoRA-VLM: Robust Retrieval-Augmented Vision Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.456987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.456987Z digest=sha256:846770eabeb2de4a8ea6076ab0ec09a363fd01bdcfb8c53dcd46129101740c14

Observation ed55e627-01c3-4b98-bb44-a9d824170bfd · outbound

This paper cites xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token.

Towards General Continuous Memory for Vision-Language Models xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.553768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.553768Z digest=sha256:7dbde29be13aff5c740b85ed689391a946953608f0912e5ae541fa6dfd628078

Observation b46d3d36-0c01-4465-ad51-30287da5651b · outbound

This paper cites KV-Distill: Nearly Lossless Learnable Context Compression for LLMs.

Towards General Continuous Memory for Vision-Language Models KV-Distill: Nearly Lossless Learnable Context Compression for LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.661849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.661849Z digest=sha256:86615db66fa6b3964a6ba4a412916babe6f972da43d43a71e2154bbd787f74b3

Observation 03a1e32b-abe3-466e-a411-0abe0a4039db · outbound

This paper cites VoCo-LLaMA: Towards Vision Compression with Large Language Models.

Towards General Continuous Memory for Vision-Language Models VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.748456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.748456Z digest=sha256:4dd1f9f53837442ba3082b2848c44d12aeb342cf548182d23e7834d62d54a06f

Observation 9eb46eb0-8e47-4268-81dc-38ff53e2fe95 · outbound

This paper cites MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding.

Towards General Continuous Memory for Vision-Language Models MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.831913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.831913Z digest=sha256:a696ba91090eb99bd0613fc6cb6a9c43d24fea252866f5140bbdd67552a41a9c

Observation ceb034a8-c31f-4dc9-a8c5-183750514798 · outbound

This paper cites M+: Extending MemoryLLM with Scalable Long-Term Memory.

Towards General Continuous Memory for Vision-Language Models M+: Extending MemoryLLM with Scalable Long-Term Memory

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.941755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.941755Z digest=sha256:a84846fef790392e052456ec8d2bd8f2a5390ba4f9c9bb0f6c11c43cc2bf0097

Observation cd50865a-ef74-4eca-b4bb-25b067d1503b · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

Towards General Continuous Memory for Vision-Language Models MemGPT: Towards LLMs as Operating Systems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:45.067719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:45.067719Z digest=sha256:33dcd670e25936fefa1f861be16e7019609abf7038e6cd8582bc06274512dc1e

Observation 16a8baad-8059-40d3-af0d-6ecc6d78427f · outbound

This paper cites Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?.

Towards General Continuous Memory for Vision-Language Models Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:45.236100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:45.236100Z digest=sha256:7cf6df23a31df3fa496faaa9dcfecc5465f6da5fe7abf4243d567eff48a31ed3

Observation c86d863a-91bf-48be-aebf-2140ac37ae56 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Towards General Continuous Memory for Vision-Language Models Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:45.381572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:45.381572Z digest=sha256:daa4b5d98e4f65f4d4bf2fb70291df34da29bb992503fba6852df21c693df519

Observation c5013a54-ceaf-4073-bf67-3fb52dcb7ccc · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

Towards General Continuous Memory for Vision-Language Models A-okvqa: A benchmark for visual question answering using world knowledge

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:46:51.094243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:46:45.520086Z digest=sha256:3c6a783c8c6dd59c266485cd0c5af78b937d57915a12053ac7962f9247521c2e

Observation 3bae3da7-054f-4a19-89dc-c3541920b6fd · outbound

This paper cites Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs.

Towards General Continuous Memory for Vision-Language Models Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:46:48.624520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:46:45.642876Z digest=sha256:59c551e60bfaa4b4151d230e6c8daa7022d2387f2f1817a9eb64d1b657f3f2b0

Observation 84408c59-a160-45c6-b174-057ff4613ccb · outbound

This paper cites Learning transferable visual models from natural language supervision.

Towards General Continuous Memory for Vision-Language Models Learning transferable visual models from natural language supervision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:45.818280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:45.818280Z digest=sha256:33993adc6ab86917ea0c0271160fd1fabf67d345a9d9bd4b8d27370cf3125348

Observation 3a40a3f2-b76e-4de8-841e-312eea322af1 · outbound

This paper cites Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning.

Towards General Continuous Memory for Vision-Language Models Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:46:50.906780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:46:45.983414Z digest=sha256:4d26fa9cf55c8128773586aa67d81252932fb082bd4bc3bb238152781d56558e

Observation 5e49f585-c4bc-4e47-b17c-275a90d4551a · outbound

This paper cites Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories.

Towards General Continuous Memory for Vision-Language Models Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:46:50.748700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:46:46.175604Z digest=sha256:35c2177482812d75e038321d9bef7a8f6e79de97f7f8ec50319eb6be80dcd985

Observation ea0a8ad6-fd27-4b25-a64c-1fe69f0e084d · outbound

This paper cites Open-domain visual entity recognition: Towards recognizing millions of wikipedia entities.

Towards General Continuous Memory for Vision-Language Models Open-domain visual entity recognition: Towards recognizing millions of wikipedia entities

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:46.320220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:46.320220Z digest=sha256:684fae85c9c62927c1c0a430e5e8caf6e4143064a63864ff7f78f34504a91332

Observation 53986ec4-e589-4e80-a516-5f5bbe5bfa7a · outbound

This paper cites MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models.

Towards General Continuous Memory for Vision-Language Models MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:46.436410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:46.436410Z digest=sha256:515121745e0a9d5f874abd2083c41cc6fb194d93d38f12f5241d9113c15f5e08

Observation 45e49340-0ac3-4d5c-b74a-89d0f684452c · outbound

This paper cites Viquae, a dataset for knowledge-based visual question answering about named entities.

Towards General Continuous Memory for Vision-Language Models Viquae, a dataset for knowledge-based visual question answering about named entities

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:46:50.599296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:46:46.556914Z digest=sha256:d665dae6d1bd84ef1e429728cbdbfcfcee159f797142af4312bb966b24944dbf

Observation 9da75c24-78c1-4c9f-a009-e9e5cef09de6 · outbound

This paper cites CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark.

Towards General Continuous Memory for Vision-Language Models CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:46.670649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:46.670649Z digest=sha256:cca2d5569bcf6c6ab76ddc85c00dc7486ca7c09cfe22db4cb7c9a6dbdf20f344

Observation 8e1c4b10-3cb9-435e-a872-12f2ddef8860 · outbound

This paper cites Gpt-4o technical report, 2024.

Towards General Continuous Memory for Vision-Language Models Gpt-4o technical report, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:46:50.397738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:46:46.762361Z digest=sha256:8e2d06b4cdf45791ab3f96422191b9c0180cef06efeb9f9a2bb844eca700670c

Observation 0c0b63b3-180b-4a61-a0f8-e443ebfa1fd4 · outbound

This paper cites Visual Instruction Tuning.

Towards General Continuous Memory for Vision-Language Models Visual Instruction Tuning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:46.835689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:46.835689Z digest=sha256:172a7bd35295abdb2c72db141aef37ad1bb305f575bae4e395425a769c9a9974

Observation c17250c3-a35b-4426-8a90-8dfa58cf63ff · outbound

This paper cites Llava-next: Open large multimodal models.

Towards General Continuous Memory for Vision-Language Models Llava-next: Open large multimodal models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:46:50.187196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:46:46.953015Z digest=sha256:6facdfdba46fb82ee59c44ae675f90a8b97e2257f3a0fb3074888834b0402059

Observation 6688bed2-a55f-492c-af4f-e69bfeb2c54f · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Towards General Continuous Memory for Vision-Language Models InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:47.018889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:47.018889Z digest=sha256:6174cbc44e85ec1bc993a227d3dc75d2d6d80f0c23d4c38b24dd157c17a95577

Observation 0cc38497-8a40-401c-ad73-95df2674128c · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

Towards General Continuous Memory for Vision-Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:47.087644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:47.087644Z digest=sha256:58d0b3b07e9ec791e7e5c6c8f3635d81cc19355745c4de4b827b1bbde1911a67

Observation 6adb41ec-9728-4a2b-ba4a-65269eeac399 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Towards General Continuous Memory for Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:47.157556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:47.157556Z digest=sha256:697fa56b35a72bbd3e7cb5c91c1f0a465bef092b06a67279dd738a03a97333cf

Observation 71eb6c5c-9535-4755-8397-d6678dd7535c · outbound

This paper cites Qwen2.5 Technical Report.

Towards General Continuous Memory for Vision-Language Models Qwen2.5 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:47.255694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:47.255694Z digest=sha256:94420995623bced5f13a5105ffbb7e15534c22da472e23e6b4926a15c41a791d

Observation 51587dfb-240b-437c-a0b9-2ad31a711c0f · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Towards General Continuous Memory for Vision-Language Models A simple framework for contrastive learning of visual representations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:47.396550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:47.396550Z digest=sha256:d8d20a6373e8991327469e79f721ac595eb872e0fb5753c07972551a29c74f8f

Observation 076259f0-863f-4afb-b3b4-2af3a8426ac0 · outbound

This paper cites Mm1: methods, analysis and insights from multimodal llm pre-training.

Towards General Continuous Memory for Vision-Language Models Mm1: methods, analysis and insights from multimodal llm pre-training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:47.494110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:47.494110Z digest=sha256:df3bbc0f6fd2606a87e7fed666976ab37b1e23a5aaffe19c63b35dea67cdad00

Observation d3b92a26-f366-4a72-b07e-6a44d775a3ef · outbound

This paper cites Ulip-2: Towards scalable multimodal pre-training for 3d understanding.

Towards General Continuous Memory for Vision-Language Models Ulip-2: Towards scalable multimodal pre-training for 3d understanding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:46:49.929357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:46:47.577690Z digest=sha256:7cdccfbf8da8933093bc49f38a10d052984c359f9e92613a9798ca7c6b9e6af6

Observation f310d7f9-4473-4ad9-a8d6-dedaed9b76d3 · outbound

This paper cites Improved baselines with visual instruction tuning.

Towards General Continuous Memory for Vision-Language Models Improved baselines with visual instruction tuning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:47.682456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:47.682456Z digest=sha256:013e8d792bced0b18f3ae945ccb69f6e31e82456c72889b0bd8777357c537ef6

Observation dc6558bf-75bb-4a7a-86d8-ed1b32ac341d · outbound

This paper cites Learning to compress prompts with gist tokens.

Towards General Continuous Memory for Vision-Language Models Learning to compress prompts with gist tokens

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:47.790233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:47.790233Z digest=sha256:65f2fb1c756730c1d4d1cf714c7acc83a643a9e4170b6c8ecfbfdabc26bfa5c8

Observation 458aabc8-aea5-4415-a175-1c59dd54f241 · outbound

This paper cites In-Context Former: Lightning-fast Compressing Context for Large Language Model.

Towards General Continuous Memory for Vision-Language Models In-Context Former: Lightning-fast Compressing Context for Large Language Model

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:46:48.281292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:46:47.871762Z digest=sha256:64e91ccb12351848b789bb9513b1a010eb5b4145ba96ab671e0e81fe1b5f632d

Observation 2d1aa9f7-4054-450e-83c6-9d10d4a87332 · outbound

This paper cites Adapting llms for efficient context processing through soft prompt compression.

Towards General Continuous Memory for Vision-Language Models Adapting llms for efficient context processing through soft prompt compression

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:46:49.643233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:46:47.968827Z digest=sha256:5edf1480bb5784ecfbf53b7c0531a012485fb099b53325a31654815c3c20cf46

Observation af99f633-c4e5-4982-91d7-ec761ca23a6e · outbound

This paper cites The probabilistic relevance framework: Bm25 and beyond.F oundations and Trends in Information Retrieval, 3(4):333–389, 2009.

Towards General Continuous Memory for Vision-Language Models The probabilistic relevance framework: Bm25 and beyond.F oundations and Trends in Information Retrieval, 3(4):333–389, 2009

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:46:49.442737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:46:48.069465Z digest=sha256:c5972269dd1039d5927adb23df11f2dd703d7760d7732959906ffdb2337061d4

Pith citing papers

Observation d2222c44-7745-4fec-99f5-5a4fc3d02b47 · inbound

Recurrence Meets Transformers for Universal Multimodal Retrieval cites this paper.

Recurrence Meets Transformers for Universal Multimodal Retrieval Towards General Continuous Memory for Vision-Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.180245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.180245Z digest=sha256:09f10f94fd6b7ce1dbe74e7739f4daafece6dd5e3d75051bb37d5de58287c538

Observation 6642783a-77c2-42f9-a7e3-0c1bf2e562c8 · inbound

Dual Latent Memory for Visual Multi-agent System cites this paper.

Dual Latent Memory for Visual Multi-agent System Towards General Continuous Memory for Vision-Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T06:04:53.456406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:04:53.456406Z digest=sha256:d2fbb2c5e72b9f3f6ccd76e464171f1aac9ad5dc9262983a442febfed67e1421

Observation 7657eba7-d3e1-4569-ae44-7480842fb4a5 · inbound

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook cites this paper.

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook Towards General Continuous Memory for Vision-Language Models

Reference 243

Resolution
unresolved
no resolver link, observed 2026-07-13T14:03:01.974171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:03:01.974171Z digest=sha256:a82b014cebc5df76115ff5f2ecde85eedc9834ea3c08367260994b8389ba992e

Observation ef4b830d-e506-4531-8d95-c8ed1bab1ac1 · inbound

Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning cites this paper.

Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning Towards General Continuous Memory for Vision-Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T15:28:33.935057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T06:30:55.592334Z digest=sha256:5c09f6ef013d05940c14aaac4dcf81d27e4d499a1c2b3adab67960d1b1e8a9c1

Observation 8ea63aaf-ab62-4406-ae02-ad1908a9dbfc · inbound

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting cites this paper.

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting Towards General Continuous Memory for Vision-Language Models

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:17:17.729849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T19:13:44.015782Z digest=sha256:503d719eacea20ceb33b71b1c3517739727eca79fa12f56248dc2c25ca8face2

Observation 641cd9fc-40c1-475b-b9c1-50db348e4452 · inbound

MMAgent-R$^2$: Learning to Rerank and Reject for Agentic mRAG cites this paper.

MMAgent-R$^2$: Learning to Rerank and Reject for Agentic mRAG Towards General Continuous Memory for Vision-Language Models

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T12:46:14.736176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-09T12:37:07.906342Z digest=sha256:32cf8d9739c2c7f70908faeeb81ecac9158de8327844a289044a307f125e564c

Observation 8fa89f1a-d668-4203-9b84-dd6f543a8df5 · inbound

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG cites this paper.

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG Towards General Continuous Memory for Vision-Language Models

Reference 162

Resolution
unresolved
no resolver link, observed 2026-08-02T10:20:56.296609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:20:56.296609Z digest=sha256:461eaad12e1ab2883fe95e534ea6eae6fc0b22541a8a19fe5e1c2d27825cf5b1