Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2106.11297.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T18:28:48.407448Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:39:57.309052Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation f5d6fb1a-b50a-43e6-b7e3-d0504daab814 · inbound
Florence: A New Foundation Model for Computer Vision TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 30516c5a-679c-47c5-afd6-0be2ed571833 · inbound
PaLM-E: An Embodied Multimodal Language Model TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c39a0be2-f050-43b9-b048-64ff5f04c902 · inbound
Fast Vision Mamba: Pooling Spatial Dimensions for Accelerated Processing TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faa12944-ddd0-4c44-b1b4-c285d500fa91 · inbound
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 20eb1f78-df11-452c-8d85-d2e28be4e7f5 · inbound
Cross-Modal Dual-Causal Learning for Long-Term Action Recognition TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43973908-3670-483a-aa6a-24b75984db3a · inbound
Lightweight Backbone Networks Only Require Adaptive Lightweight Self-Attention Mechanisms TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbe25348-7814-438d-8213-2927ef1591f0 · inbound
Rethinking Visual Autoregressive Sampling with Information-Grounding Guidance TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f421aac-06a1-4aea-b0b3-ee1b5bee056f · inbound
A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4b86630-081a-4e87-b794-675d41d62a52 · inbound
Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 165234c0-d7f8-4856-8725-9299041fa718 · inbound
VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4406e2e2-1d2a-4158-b0ea-ec86cfa3b9d9 · inbound
Why Training-Free Token Reduction Collapses: The Inherent Instability of Pairwise Scoring Signals TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2fef50d5-e49b-4462-81b9-bd9f7a21d282 · inbound
Head Similarity: Modeling Structured Whole-Head Appearance Beyond Face Recognition TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation febe5295-0f6b-4e26-9148-b2cca8ad8efd · inbound
FlowNar: Scalable Streaming Narration for Long-Form Videos TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fa4617a4-e2da-43c2-a07c-c1e12f017da7 · inbound
See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 48025284-28ab-4294-8957-8eccd88a7b8f · inbound
Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b1d9f042-20ba-47ca-b090-35e5a176ecdb · inbound
EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f721aee8-9c59-4278-b5f7-33c600d4a454 · inbound
DinoLink: A Token-Centric Representation Compression Framework for Bandwidth-Constrained Collaborative V2X Perception TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b299ec41-e64c-4819-94c5-0e0e881431e2 · inbound
DinoLink: A Token-Centric Representation Compression Framework for Bandwidth-Constrained Collaborative V2X Perception TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e59ce50a-90e3-4220-a333-c8ae05a661d9 · inbound
ATS-ToDMA: Adaptive Token Selection and Token-Domain Multiple Access for Cross-Modal Semantic Communications TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d53f9a7-1533-4aa6-8abd-329d56327106 · inbound
Brain-Aligned Multi-Stream Video Transformers with Sparse Self-Selection TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.