Pith. sign in

Paper Citation Record · LEDGER

TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2310.19060.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.19060 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:47:27.306650Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T02:46:16.800490Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 157cfe73-768d-4293-9e5c-ccb9bf8d79bb · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:16.803157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:abccdc03f871f8c4261ddd0d88f5fd6fed30d6dc3e3ea66a6a60840d2b2dcdce

Observation 9742fac6-7bc2-4221-87ff-638d8aad58c7 · inbound

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding cites this paper.

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:53:33.691158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T13:53:33.585035Z digest=sha256:7574cc7227383d9c4aefe3c7a6b3e3636fa8e9e65c2e8ff9f33ff6d08d41c139

Observation 03bae0ac-1fbe-4342-b3e4-9420f2fece5d · inbound

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models cites this paper.

DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:40:23.713555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:40:23.713555Z digest=sha256:158c1d0013809bee4c99f563c98f45ad11ed5b528fd06805832e6b6846bdcf28

Observation 5410bbec-2005-4be8-ba32-49007412a2e4 · inbound

LinVT: Empower Your Image-level Large Language Model to Understand Videos cites this paper.

LinVT: Empower Your Image-level Large Language Model to Understand Videos TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.384673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.384673Z digest=sha256:66570c96516888f80dbdad64b27939cb0036d59a4198c99c3b1f0f9bdc8be193

Observation 562f1fdf-064f-4014-a903-d8a630884241 · inbound

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos cites this paper.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.306650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.306650Z digest=sha256:8d9e37f3fa39b4135c6a9f616e89f9d95ca643b4f6733dd434da967dd6857757

Observation 0ab99992-cdb6-4ac8-81d3-054af5781bc6 · inbound

Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis cites this paper.

Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:00.476474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:12:00.476474Z digest=sha256:bded9119ddb1fc5296e69b3621f9026353397f60dc7af8e78a61c303c33e0b72

Observation 948a81fc-d353-4121-be4f-5d3b970c6457 · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.578062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:523f555ef0e338f812c3f69c2e6ad073f19001b3b2a80478aadb610292d9e56f