Pith. sign in

Paper Citation Record · LEDGER

MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2404.05726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.05726 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:46:44.831913Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.683925Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 10ff7284-3e28-4fa0-9804-a77c502c2499 · inbound

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models cites this paper.

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:22:35.539507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T19:22:35.305220Z digest=sha256:978ecedc4b2340b9c00f85911ad15ea27c5805793cda3bc3f06a3e040fb3c328

Observation 1fda6659-f351-4b47-9145-27bc985f5b95 · inbound

MLVU: Benchmarking Multi-task Long Video Understanding cites this paper.

MLVU: Benchmarking Multi-task Long Video Understanding MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:55:26.557379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:55:26.333923Z digest=sha256:ec775740f2a072448e807b7eb6c62066973e031abf9e1a83e0e5acd64bd2141b

Observation e3558315-0c00-4d59-819c-9fbc06c7dbbb · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.429141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:8aae02f7d80e7a90853fe657d984970b0af53d2b066b816447faa4b17e03202f

Observation e91f77fe-e0a8-43dd-a17d-f75572f8d718 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.729626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:9ba24c7fef18580bbd9687cbfe4e45033e3c7fbc0ac038c624fcc4aef4069ba9

Observation 9eb46eb0-8e47-4268-81dc-38ff53e2fe95 · inbound

Towards General Continuous Memory for Vision-Language Models cites this paper.

Towards General Continuous Memory for Vision-Language Models MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:44.831913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:44.831913Z digest=sha256:1af93f1643fa232103fda20bf14dd6a3164e9b68c84271f41d22c97615280d5a

Observation fdc0e26e-7bbc-413b-abce-bc2b993963ca · inbound

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing cites this paper.

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:25.772596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:25.772596Z digest=sha256:bcee0912f4930b735e5ca9b3114a61cca0c98048b877c89fbed9b4ebd07c0640

Observation 8ecbc304-40b5-428f-8234-871d5a1d251c · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.857202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.857202Z digest=sha256:b711ae62d34dabbb8d63981196bafb1711cf0f7a4a82ae0bc6ae190138a73ba0

Observation 2800f197-06c6-49eb-bd7d-99a1b40740be · inbound

Task-Aware KV Compression For Cost-Effective Long Video Understanding cites this paper.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.020202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.020202Z digest=sha256:6e126f63390f45c29cbde683e58e7296c131a99c1816100c938cb93362d2fe91

Observation 604b0fd2-5bef-4b08-adb0-c9b68053db71 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.327746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:81eb2ef29faee5614e414b82817b139265e34c033de9c117db5c4b12f1383740

Observation bb49de16-191a-4ae0-9658-6af5cbc4c952 · inbound

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning cites this paper.

PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:24.819686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T15:12:00.408851Z digest=sha256:a7e04a11e6fd0d66cc63a7182a6041a5a5796a011da9b3896920dce0daf93983

Observation 6a8ce2a8-7760-465a-8253-1ddea61a5315 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 156

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.685514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:3c694eda45a5e2c442af942c1aa021aecea587f85864f5cf89610b9da3091f3f