Pith. sign in

Paper Citation Record · LEDGER

Importance-Based Token Merging for Efficient Image and Video Generation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2411.16720.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16720 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:45:11.197151Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T03:22:01.146276Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 11c4e6df-aab0-4c74-b5df-e9e305935ad8 · inbound

RainFusion: Adaptive Video Generation Acceleration via Multi-Dimensional Visual Redundancy cites this paper.

RainFusion: Adaptive Video Generation Acceleration via Multi-Dimensional Visual Redundancy Importance-Based Token Merging for Efficient Image and Video Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:11.197151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:11.197151Z digest=sha256:62b1976220955a97482821e6e3ad1e4e9ef429afa22c88e3ca03fa01c15536d1

Observation 7ea4f880-8b1d-4c83-85bd-f708eff1ac1c · inbound

FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers cites this paper.

FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers Importance-Based Token Merging for Efficient Image and Video Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:45.998524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:45.998524Z digest=sha256:32950a13cb911edd491701d5ec26d29bcaf4f6200555187904d3c13d6857272c

Observation 9c34d7d6-1947-4e07-988b-d84d2cb26193 · inbound

Local Representative Token Guided Merging for Text-to-Image Generation cites this paper.

Local Representative Token Guided Merging for Text-to-Image Generation Importance-Based Token Merging for Efficient Image and Video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:46.341931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:46.341931Z digest=sha256:a9420fb349d0b2f63f7087f3161cb02066d7c425070c655941b95447a1fe3b14

Observation bfcff2a8-ad00-4100-90cc-b2dbd277db2f · inbound

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs cites this paper.

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs Importance-Based Token Merging for Efficient Image and Video Generation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.148940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T03:18:11.993413Z digest=sha256:534760f5e6f14ca4bb58be075479587fee17f67b9ff99aeb72d5cfd5fe8f4a2d

Observation 20c6bc3f-fa03-417d-bad6-6cffea2bbe56 · inbound

DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking cites this paper.

DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking Importance-Based Token Merging for Efficient Image and Video Generation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:10:05.844225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T15:10:00.146694Z digest=sha256:5af434e96ff575e1a544fa54015555896f997dbe71e88bcc6c90ad598db9ace2