Pith. sign in

Paper Citation Record · LEDGER

TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2106.11297.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.11297 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:37:07.148327Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:57.309052Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f5d6fb1a-b50a-43e6-b7e3-d0504daab814 · inbound

Florence: A New Foundation Model for Computer Vision cites this paper.

Florence: A New Foundation Model for Computer Vision TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:38:09.558580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T09:38:09.427509Z digest=sha256:a024b0a8a8ea88f592f06f1a310003a32dd3890e81c0ced379655f1f3173aa53

Observation 30516c5a-679c-47c5-afd6-0be2ed571833 · inbound

PaLM-E: An Embodied Multimodal Language Model cites this paper.

PaLM-E: An Embodied Multimodal Language Model TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:29:30.012482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:5874c39b983492e9374b6cd29887bf6aa3fe67dac099b7eb67570959c8b06544

Observation c6e191aa-4563-456b-9d38-200a2507f493 · inbound

Principles of Visual Tokens for Efficient Video Understanding cites this paper.

Principles of Visual Tokens for Efficient Video Understanding TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.781891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.781891Z digest=sha256:42df8f931dcf192e2f9f8332d99508bdfa0b0de442d31175a81c883e89985bca

Observation c39a0be2-f050-43b9-b048-64ff5f04c902 · inbound

Fast Vision Mamba: Pooling Spatial Dimensions for Accelerated Processing cites this paper.

Fast Vision Mamba: Pooling Spatial Dimensions for Accelerated Processing TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T18:28:48.407448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:28:48.407448Z digest=sha256:d468ac74e2645295da2422fbb54e8e4fcd76b25593af3660cc72209f321ba6d5

Observation 95b55e89-a69e-446c-a8f2-dedc30beab81 · inbound

Back to Fundamentals: Low-Level Visual Features Guided Progressive Token Pruning cites this paper.

Back to Fundamentals: Low-Level Visual Features Guided Progressive Token Pruning TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T10:37:07.148327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:37:07.148327Z digest=sha256:bac8852735d9e774f0b0c62d424be0fbd08a9d02eadf21bd85b4214ebf1098a6

Observation faa12944-ddd0-4c44-b1b4-c285d500fa91 · inbound

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory cites this paper.

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.864851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T12:54:31.765909Z digest=sha256:31bb010b53ffb587a85e3f01d79c86a90a5f7759864f672e4ed4f9bc083d481d

Observation 20eb1f78-df11-452c-8d85-d2e28be4e7f5 · inbound

Cross-Modal Dual-Causal Learning for Long-Term Action Recognition cites this paper.

Cross-Modal Dual-Causal Learning for Long-Term Action Recognition TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:06:10.126258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:06:10.126258Z digest=sha256:09848141321afd64e74da0b21b3ea0b4cd38d5edbacea97051c0092ba36de44a

Observation 43973908-3670-483a-aa6a-24b75984db3a · inbound

Lightweight Backbone Networks Only Require Adaptive Lightweight Self-Attention Mechanisms cites this paper.

Lightweight Backbone Networks Only Require Adaptive Lightweight Self-Attention Mechanisms TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T05:44:19.783785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:44:19.783785Z digest=sha256:4f1bcc7a761357f74140490287a1f529164f46924bbbeaf947b9be5f0e24d0c7

Observation fbe25348-7814-438d-8213-2927ef1591f0 · inbound

Rethinking Visual Autoregressive Sampling with Information-Grounding Guidance cites this paper.

Rethinking Visual Autoregressive Sampling with Information-Grounding Guidance TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:15.404908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:15.404908Z digest=sha256:0b0654d051434b755755a1184c550eac90adce2a956217a3cb7f4d9b2587eb7a

Observation 3f421aac-06a1-4aea-b0b3-ee1b5bee056f · inbound

A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models cites this paper.

A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T21:31:35.530259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:31:35.530259Z digest=sha256:3c47d80f4e9d894d389ba031b63218486e37d644aec460bb2c940bbb2f2f4911

Observation a4b86630-081a-4e87-b794-675d41d62a52 · inbound

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies cites this paper.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.496346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.496346Z digest=sha256:12986c72957cb0ff343cbc6d6e07c345158fe9e7453d4aa3a5ebf47751566cd2

Observation 165234c0-d7f8-4856-8725-9299041fa718 · inbound

VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning cites this paper.

VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T23:00:31.128534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:00:31.128534Z digest=sha256:2b49ceb31a48f77e9360a671744220756252c435d10376ace250fa06fa6f1fbe

Observation 4406e2e2-1d2a-4158-b0ea-ec86cfa3b9d9 · inbound

Why Training-Free Token Reduction Collapses: The Inherent Instability of Pairwise Scoring Signals cites this paper.

Why Training-Free Token Reduction Collapses: The Inherent Instability of Pairwise Scoring Signals TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:02:25.416558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T07:59:10.678150Z digest=sha256:6bc51a4db23fa8c2659e96bf73b31c860d86b8d45e0e7d220fd5bc0507d26390

Observation 2fef50d5-e49b-4462-81b9-bd9f7a21d282 · inbound

Head Similarity: Modeling Structured Whole-Head Appearance Beyond Face Recognition cites this paper.

Head Similarity: Modeling Structured Whole-Head Appearance Beyond Face Recognition TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:55:52.997875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T02:54:58.288120Z digest=sha256:9f7799d9258401f35202282060c805ba0307cc0bab6b59460e1d69275eb6c6e5

Observation febe5295-0f6b-4e26-9148-b2cca8ad8efd · inbound

FlowNar: Scalable Streaming Narration for Long-Form Videos cites this paper.

FlowNar: Scalable Streaming Narration for Long-Form Videos TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:12:34.739821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T19:08:36.655886Z digest=sha256:b64adbe1851f028676a6c6d12d982adadbecc80b537225206f0cc19de4c26551

Observation fa4617a4-e2da-43c2-a07c-c1e12f017da7 · inbound

See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs cites this paper.

See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:36:24.027110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T14:04:43.270155Z digest=sha256:c942c6e63596265ba1ebadc602b430eb76835b41246e91555fe5208039f1655d

Observation 48025284-28ab-4294-8957-8eccd88a7b8f · inbound

Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models cites this paper.

Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:39:29.530970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T17:53:38.503877Z digest=sha256:931f5bfc2a618c8e5f933a2ccd117324723bba2473c417db24d287c818922a68

Observation b1d9f042-20ba-47ca-b090-35e5a176ecdb · inbound

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding cites this paper.

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:57.310471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T00:26:33.306128Z digest=sha256:ed345fd90576483238d568413f01635dade9acf2333f236b6b97045846f50a82

Observation f721aee8-9c59-4278-b5f7-33c600d4a454 · inbound

DinoLink: A Token-Centric Representation Compression Framework for Bandwidth-Constrained Collaborative V2X Perception cites this paper.

DinoLink: A Token-Centric Representation Compression Framework for Bandwidth-Constrained Collaborative V2X Perception TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:49:57.573225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T01:19:37.860746Z digest=sha256:f66889b5b0ad1c1cacad1698f2d81cf691444d33c255daa4db893267ee00cacb

Observation b299ec41-e64c-4819-94c5-0e0e881431e2 · inbound

DinoLink: A Token-Centric Representation Compression Framework for Bandwidth-Constrained Collaborative V2X Perception cites this paper.

DinoLink: A Token-Centric Representation Compression Framework for Bandwidth-Constrained Collaborative V2X Perception TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:40.915278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T06:24:11.726818Z digest=sha256:524593ef9555b071a562376b1513b1ff4d322a3546d9284e98ca79500999104f

Observation e59ce50a-90e3-4220-a333-c8ae05a661d9 · inbound

ATS-ToDMA: Adaptive Token Selection and Token-Domain Multiple Access for Cross-Modal Semantic Communications cites this paper.

ATS-ToDMA: Adaptive Token Selection and Token-Domain Multiple Access for Cross-Modal Semantic Communications TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T01:55:58.109603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:55:58.109603Z digest=sha256:b5a4f12e392c829ad334a8d182d93246b427f9923627d0f5846c18120ce316e2

Observation 4d53f9a7-1533-4aa6-8abd-329d56327106 · inbound

Brain-Aligned Multi-Stream Video Transformers with Sparse Self-Selection cites this paper.

Brain-Aligned Multi-Stream Video Transformers with Sparse Self-Selection TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:32.351486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:32.351486Z digest=sha256:61ccfc1c66a509fa8915e690d6cefe289272a217d138b6e85bd5c273bc5ba0a6

Observation ca2b3d6d-c851-4b23-9220-5e69de98087c · inbound

Gated Spatial Redundancy Projection for Pathology Transformer Attentions cites this paper.

Gated Spatial Redundancy Projection for Pathology Transformer Attentions TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:16.292894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:16.292894Z digest=sha256:05a1914203259d4dca9a788d95f259e009368ed327c50593e93a2059775285d1

Observation d0bf40bf-a9ea-4d8f-bcc7-5237f5c2718f · inbound

Gated Spatial Redundancy Projection for Pathology Transformer Attentions cites this paper.

Gated Spatial Redundancy Projection for Pathology Transformer Attentions TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-14T04:42:11.262758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:42:11.262758Z digest=sha256:99e8411941a2cb78b8415d8db80828d9266b8dd7282cf58d66b4e4084c62156b

Observation e10f9fa0-2134-4f70-86b5-d7bf8db6fed0 · inbound

Dual Anchors, Do It Better: Hierarchical Group Merging for Zero-Shot Anomaly Detection cites this paper.

Dual Anchors, Do It Better: Hierarchical Group Merging for Zero-Shot Anomaly Detection TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:37.467311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:37.467311Z digest=sha256:7dea201067aa7c1b81646bf7cc500251b87d64fd741676849a8b5a05c915ed17