Pith. sign in

Paper Citation Record · LEDGER

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs

As of 7 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 2 inbound Pith citation observations for arXiv:2508.21044.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21044 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:41:15.586140Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T17:05:52.493660Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T17:08:00.954550Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f29c4805-84c0-4a34-92c7-6508dc49d3ec · outbound

This paper cites PruneVid: Visual Token Pruning for Efficient Video Large Language Models.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.535604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.535604Z digest=sha256:b352a2ef68be0bbecf87e87c86cb53fa87bcdc4892626a0c02ec555fba39bd24

Observation ce4e96b2-8cef-4572-8cf7-50b4b1ef4c25 · outbound

This paper cites In European Conference on Computer Vision, 323–340.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs In European Conference on Computer Vision, 323–340

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.541099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.541099Z digest=sha256:e832d1534664d8de7b101a3a2c3faceacd76e6d332d2d2e5c996f3dd166313cb

Observation cbfd5e3d-863a-4930-afec-362e3a42d90c · outbound

This paper cites arXiv preprint arXiv:2403.15388.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs arXiv preprint arXiv:2403.15388

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.547335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.547335Z digest=sha256:977278e4d5c1e0b3b527de37cbc0fedff569784c6a8b94ed67b0e5167c37f4af

Observation 1fb9ee88-8002-412b-98db-60ff0724ad84 · outbound

This paper cites arXiv preprint arXiv:2505.21334.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs arXiv preprint arXiv:2505.21334

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.552875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.552875Z digest=sha256:fefc7dbc66147ae56846666c3410c4cd9627064598e4d0b8527aa2ef2130def5

Observation e844714e-1986-46ad-abc5-1084ac886157 · outbound

This paper cites arXiv preprint arXiv:2503.11187.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs arXiv preprint arXiv:2503.11187

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.559338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.559338Z digest=sha256:1a6a1e4811bba0e68650a05e12968a1a80619849adae9155f5da40bfa8fa9d76

Observation 6e288fbc-89bf-4447-b45b-1edfde484ffc · outbound

This paper cites LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.565494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.565494Z digest=sha256:453ff6ac6fd5c433227677e8372fddb5366baa9956bd8eba4164a05f2ce31d18

Observation ea89fa3a-6c78-4761-b489-6d1bd969cae8 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.570938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.570938Z digest=sha256:97e5e63128de82bd295cc4f87bf65f7d9a38b93cb8d258ac6e9828063b28d03a

Observation e26d286f-8408-4b14-bce5-80b80213c53a · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.575985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.575985Z digest=sha256:15603ec95a17609f30629318011eb9442a33c8e7c0e5f39aecc94133841aadbc

Observation f5871a72-838b-4606-a26c-86655c41fd12 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.580800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.580800Z digest=sha256:982457262e638a4b3fcb02b96c28b72765ea8a34b4253896ead6165b8f40bd34

Observation 3852654f-126a-4d63-bc82-d74c946be87f · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs MLVU: Benchmarking Multi-task Long Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.586140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.586140Z digest=sha256:e1616be8f55199de7005b34c9d57045be586e5110142d43306b69b814fd286f3

Observation 5901147d-987a-471d-a87d-180b646201cc · outbound

This paper cites arXiv preprint arXiv:2411.17686.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs arXiv preprint arXiv:2411.17686

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.530096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.530096Z digest=sha256:cc236a358ce32c39180a498da7373665e6262f7f0c086932756ba1e61ed8b0ff

Observation 09ded1e6-3c85-45f7-881a-d26649f03049 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T14:41:15.525180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:41:15.525180Z digest=sha256:32eeed51899541f87482a8ab037a9f713d258af1ea3a8b0a55d39bde04de214f

Pith citing papers

Observation 29f1ba24-1c86-4ce9-8209-8e21e967bea8 · inbound

DINO-VO: Learning Where to Focus for Enhanced State Estimation cites this paper.

DINO-VO: Learning Where to Focus for Enhanced State Estimation MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:00.956426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T17:05:52.493660Z digest=sha256:cc5848275ddab3455a91980afeb3d3ca3d7ff532aa60279e0a2cebf147ebe69f

Observation d356d0ca-6110-45a7-ab72-1ed6e6e22b76 · inbound

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs cites this paper.

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:49:48.956060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:43:44.921189Z digest=sha256:7bf6b7694d742e74094fe0107ce974ff89ebf38279c1820a3bec6c75bf176cb7