Pith. sign in

Paper Citation Record · LEDGER

Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2401.10529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.10529 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:12:00.777919Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T19:13:40.522467Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 90f2d030-8bf9-4556-ab50-675d6e5521a9 · inbound

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding cites this paper.

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:09:30.446521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T01:09:30.360275Z digest=sha256:eece192044d1544fe8c1054388e39b5537ab6a703868cccc2d27388ab8889052

Observation 04cca4ad-d10f-4e47-966c-7749221d48db · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 118

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:20:36.548843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:24a34fea9a418359268e78762eaee84e2ac3ea4f36ab813fc3df48f52e64ebb3

Observation 41a960b5-32f6-4684-97c3-a849407cd8b5 · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:37.000862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:37.000862Z digest=sha256:7451285dd0a7c18e7c2f42463ef7d9d54843bc55571215f838915d76d5cb6e43

Observation f2f9a9dc-f0df-4599-93ea-7c7e563967cc · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 255

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.136717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:f1110611e2d7e7bbb41560837bca7561fc82ea3c878cf44a27f4df7d2da7fd0e

Observation 337d59fe-bf91-4a28-b192-0808d859ed20 · inbound

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models cites this paper.

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:42.682141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:42.682141Z digest=sha256:f0b2b4eab9eaf22456325dc6a9de6023cc79026f8ee9218ade2755ba4ece1d3f

Observation 2b2e239a-2581-4888-8c64-1235e3b8fe61 · inbound

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding cites this paper.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.355816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.355816Z digest=sha256:ae879096d4a81755a1fe1453199502c840d4073ad7eadcf0157cb6d1c9e3d421

Observation fec5fe3f-2d51-4bd4-a3e1-a1c86c94eb84 · inbound

MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs cites this paper.

MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T17:01:41.209908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:01:41.209908Z digest=sha256:d5d8f99deb6311b69bc61badb513e541f27cc49532794d743e7b77c1636a7a8c

Observation 9d8b620f-9c05-470f-8649-689d65c15818 · inbound

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks cites this paper.

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 157

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:00.777919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:12:00.777919Z digest=sha256:f1ef0b18d85e88bd40ce2db95f7c20ccfcb7898af1a9f747e91a5b1814858901

Observation c98dafb4-293c-4733-8528-d5619c45cda9 · inbound

Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models cites this paper.

Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:07.707953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:07.707953Z digest=sha256:68f4396ef8e83928ee0dd33270e9551a219795e169817b932fe173e14861a478

Observation 0798781e-ecbc-47a2-85d0-3f8cf615c646 · inbound

What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding cites this paper.

What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:18.668518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:18.668518Z digest=sha256:cfd05d18de6d8f62d239641b082293906247867830e9251313cab081593e5770

Observation 990122ee-1d85-491e-8966-27c4fad6d45e · inbound

Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images cites this paper.

Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:34.691445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:43:34.691445Z digest=sha256:174eb790ecf91930248bc3a8fd27b3c81b9c2e8b850d420f4da05237f6a3305a

Observation 7b28090b-ee75-4c24-8f16-201a27e47f49 · inbound

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning cites this paper.

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:22.137629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:22.137629Z digest=sha256:76a941657b6e87813385279439594c9dd5c6e9936424fcc0f1b38192b0f289f5

Observation 275553a3-2d36-4427-becc-762e4e7daf42 · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:09.464295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:09.464295Z digest=sha256:7148b2c8e29000da7a60bfbe328d7aaed4157f26e1f483bac8924bf0121b9308

Observation 63eefdf4-10f7-41b5-a24d-b1cf97df0815 · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.886467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.886467Z digest=sha256:8c230993a9330f7372533a69e2044dc4f31a6675a07169eb367f13fa0c3057db

Observation f98ebe7d-445a-4880-9b65-cad8f634549e · inbound

CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language Models cites this paper.

CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language Models Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T12:32:14.066097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:32:14.066097Z digest=sha256:1314030c936e70743d56288a0567dd2ab78611643ac3dccdde864f75e6d9c448

Observation 8b18b037-8b02-4f4e-86ca-244c91950496 · inbound

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings cites this paper.

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T23:35:18.926817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:35:18.926817Z digest=sha256:b087e0896d9b579616472b5bbe7d881dcd8a5a02c34188ca09721bfcda5d28c1

Observation a0c0d88c-c8f5-4176-9148-54c8e93ec1ad · inbound

Spatio-Temporal Grounding of Large Language Models from Perception Streams cites this paper.

Spatio-Temporal Grounding of Large Language Models from Perception Streams Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:26:02.674180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:10:45.837684Z digest=sha256:a597964a5a0d3a096852549b3d284fa88fc4aa772f34d659597a55ec4be1e020

Observation 2583d57f-3e94-417a-bda6-ee3a603071d1 · inbound

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory cites this paper.

SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:13:40.524351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:11:15.761831Z digest=sha256:bcfa4ecf3a1714bc3749f7b69c14cb59cdf562d844a8bc3cd0c58b5f8c314c40

Observation f5853c1d-4920-472a-b56f-c7429762a3f2 · inbound

Beyond Retrieval: Analytic Memory for Multimodal Agents cites this paper.

Beyond Retrieval: Analytic Memory for Multimodal Agents Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T07:06:05.783875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T07:06:05.783875Z digest=sha256:1452262c215a0840cf4aac94f2612fa4d45fa02610da1c39221e6eabcea9de83