Pith. sign in

Paper Citation Record · LEDGER

TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2409.01156.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.01156 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:08:09.248749Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T07:34:22.143661Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 52f69615-1a5e-44bc-b7fc-b0b73babfd78 · inbound

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models cites this paper.

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:09.248749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:08:09.248749Z digest=sha256:fac9df427bc8d5a2d634569db69517535b6ed92f18adf40cc57648affeab3334

Observation 20943e42-6839-4b62-ad48-600cbe32f003 · inbound

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval cites this paper.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.599960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.599960Z digest=sha256:42f559a46644d05608c3b91dab1aa7f1f3c4001ef17baa28329ba772827883f7

Observation c1ba48df-e473-4e59-bdf4-1306f72d7a79 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:31.751666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:31.751666Z digest=sha256:baf80a75cfa04f518ff903a3d185973634bd6958190959f99387ee6d235829a0

Observation 86d26e7a-4e49-4ef1-8d1a-ee965f131884 · inbound

FastVGGT: Training-Free Acceleration of Visual Geometry Transformer cites this paper.

FastVGGT: Training-Free Acceleration of Visual Geometry Transformer TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:36:06.973826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T23:36:06.870759Z digest=sha256:7ec48d822490c01ab6d38d55248482910b6f2b812957564d382b016d1714606a

Observation 7ed984dd-c281-4f1b-bf8f-54d493bc3784 · inbound

TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval cites this paper.

TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:22:45.895680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:22:45.895680Z digest=sha256:f381cd6545b311d6fac4f90b5bebc46105b8cd706c9b05e3e4914d42600c6efe

Observation 98511968-0902-473b-a2c6-0aeff1f6d256 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.518670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:70e0b1e1d3b303e943eda530478b0759a16a3d1106d543143f3417d637bbd282

Observation e6373ec4-0774-4ee1-8c9a-40bb5e22f222 · inbound

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models cites this paper.

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:26:27.050959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:25:21.621268Z digest=sha256:0d5aeffb66790809183b7d1d53160417071bce8daba8f9164b255424fd96f87e

Observation 4d8556ba-eb10-4551-832a-2fafc5d97a49 · inbound

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis cites this paper.

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:06:09.663380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T15:05:37.964883Z digest=sha256:147dc2884d168eb6b54e1a792f34603f786832242e5a452ebb1892e6b69847e7

Observation 405e9bd1-0db9-4bfc-8bf8-6a6e5c2d1341 · inbound

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs cites this paper.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:22.145236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:da8d04ecc3988d6ff88ecec272c13c9232f225ecb741093c1afe56622d78b266

Observation 54e0cb18-43af-407f-b79f-dc31dac84aca · inbound

STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation cites this paper.

STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-12T06:06:47.233814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:06:47.233814Z digest=sha256:fe7ad09f7795b5f45ccb69ef177b9a426d2d12f4c862429102cee03c2ccd338c

Observation fe4e68f3-7808-4951-9552-1dcae442deca · inbound

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation cites this paper.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T23:59:17.042425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:59:17.042425Z digest=sha256:1399de8b057e9b367af8e404e59bd8015d45d79d3c8778d724d8eb2c0ba8ce16

Observation bd2233ae-0764-4622-ba01-0bc71e99655c · inbound

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation cites this paper.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.573191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.573191Z digest=sha256:01fa80f643b9bcd82705b9cb059480e38148bfdc41edf3d074c8619e69adbf2a