Pith. sign in

Paper Citation Record · LEDGER

MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2404.05726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.05726 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:26:44.330632Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.683925Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 10ff7284-3e28-4fa0-9804-a77c502c2499 · inbound

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models cites this paper.

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:22:35.539507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T19:22:35.305220Z digest=sha256:07963adf465d1f0d6dbe74ec6d26b67bd015e88ca6df872516f7bd56e8624b44

Observation 1fda6659-f351-4b47-9145-27bc985f5b95 · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.557379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:5f381513e4f17c7bb80390356ce3ff7bd27456ab7c762a44bd9fcc8c6cb2045c

Observation e3558315-0c00-4d59-819c-9fbc06c7dbbb · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.429141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:34b72f121ba4d4cf162fa77caa572c0a59733557314f8d35b57efe098223cf64

Observation e91f77fe-e0a8-43dd-a17d-f75572f8d718 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.729626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:84ebe88ac3da97893a2a7dfdd23b0cd3602c20b30f03213d7694a36bca2b449d

Observation c39e18e0-8aa9-4e5f-a448-cd0660fd5507 · inbound

Neptune: The Long Orbit to Benchmarking Long Video Understanding cites this paper.

Neptune: The Long Orbit to Benchmarking Long Video Understanding MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T16:58:28.890785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:58:28.890785Z digest=sha256:3701515a2434b140242a4c8b3994151daab32ab38b8b6bcaaa2187e79542c4e5

Observation 938bcdb0-1865-49dc-8ce8-f2956f94d0db · inbound

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions cites this paper.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.151966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.151966Z digest=sha256:0cd8841b8edc95ddb3a089faad33604cfaf688913b50d8078a5f94c853906591

Observation 4788f1fe-466d-4b84-bd73-8f0cdac9ceef · inbound

ClassComet: Exploring and Designing AI-generated Danmaku in Educational Videos to Enhance Online Learning cites this paper.

ClassComet: Exploring and Designing AI-generated Danmaku in Educational Videos to Enhance Online Learning MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T10:26:44.330632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:26:44.330632Z digest=sha256:ae4991dbb64f873d76681407a5f42be49f76e6ad9bd82ab29b8288545f3ab3c4

Observation 9eb46eb0-8e47-4268-81dc-38ff53e2fe95 · inbound

Towards General Continuous Memory for Vision-Language Models cites this paper.

Towards General Continuous Memory for Vision-Language Models MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.831913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.831913Z digest=sha256:3a53b5dc05e8b981931ffee95d6e128239cb08d086ca2ecc0424f6e164c85f9e

Observation fdc0e26e-7bbc-413b-abce-bc2b993963ca · inbound

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing cites this paper.

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:25.772596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:25.772596Z digest=sha256:97ccf838ecea4f3fcb7635905514e633762cb6eb409c872f36ab5df7523f3908

Observation 8ecbc304-40b5-428f-8234-871d5a1d251c · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.857202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.857202Z digest=sha256:1374575c0f916d73d4d99d4c5822b902e09c29f7904010b2e5466f0acf7f6360

Observation 2800f197-06c6-49eb-bd7d-99a1b40740be · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.020202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.020202Z digest=sha256:7173ab388813aac042e17769ffabdae14150b4e5e0fdf515614ee34a6df0258f

Observation 604b0fd2-5bef-4b08-adb0-c9b68053db71 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.327746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:46bbebee0a1620e5760498dab3e3155342c3af9729d9b7709a66e9d19f746682

Observation bb49de16-191a-4ae0-9658-6af5cbc4c952 · inbound

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning cites this paper.

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:24.819686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-20T15:12:00.408851Z digest=sha256:f1d006a8ab7bc6b906f2306df6d382eb0d38d8401fa51dab8a03f6de4ada805c

Observation 6a8ce2a8-7760-465a-8253-1ddea61a5315 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 156

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.685514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:e5684748547ca41ee0e75f9845d791dd90c2e631be179840db34fd82276fcf58