Pith. sign in

Paper Citation Record · LEDGER

Slow-Fast Architecture for Video Multi-Modal Large Language Models

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2504.01328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.01328 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:32:17.814422Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T22:13:59.574945Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 89728ef8-fc4a-4e38-ae24-691b01dab7f8 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Slow-Fast Architecture for Video Multi-Modal Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.814422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.814422Z digest=sha256:d9fa05746af6f21b176336ef96010439c26a0aeb0c8446d0360be929fac1ee0b

Observation 772f157f-b4d0-4285-8ea3-bfe4546a9a97 · inbound

LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs cites this paper.

LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs Slow-Fast Architecture for Video Multi-Modal Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:22:10.097721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T03:20:42.684286Z digest=sha256:4ffb5c8a04c08f2196368d66f66c4eba88a6e55c88f6837c8890c2c4b33ab65f

Observation 9502d1c9-2655-4300-9f10-c4a927bcd63f · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Slow-Fast Architecture for Video Multi-Modal Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.576656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:76f469dc6331701af727c991302327c54fa04df533801bfdc862ff1bf46d1a85

Observation 0e986fbd-7f99-4b03-9579-cfff989aa9dd · inbound

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer cites this paper.

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer Slow-Fast Architecture for Video Multi-Modal Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:25.911484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T13:01:42.880738Z digest=sha256:9077d57876614869b8b7311d86617e3d30ad9d0f001239c851b322ea847343ef