Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:01:29.963011Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 3 inbound Pith citation observations for arXiv:2506.20066.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:01:29.963011Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:29:38.951111Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T17:29:40.174642Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1e43f906-9452-4009-90bd-a78c62ddc673 · outbound
ToSA: Token Merging with Spatial Awareness Dinov2: Learning robust visual features without supervision,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 31346f8b-c886-440a-8a6c-0d45e01fb506 · outbound
ToSA: Token Merging with Spatial Awareness Learning transferable visual models from natural language supervision,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47d741a3-4730-4439-8a42-c0a638032c5f · outbound
ToSA: Token Merging with Spatial Awareness Sigmoid loss for language image pre-training,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a04e200e-2497-4f6e-b485-3f1c3baefbb5 · outbound
ToSA: Token Merging with Spatial Awareness An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72db5e3e-9c2b-4405-babd-6f01a8e834c0 · outbound
ToSA: Token Merging with Spatial Awareness Visual instruction tuning,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a940bb97-d4fd-4cf9-8501-0c57c3225c6c · outbound
ToSA: Token Merging with Spatial Awareness LLaVA-OneVision: Easy Visual Task Transfer
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e74f5825-bd81-4f70-a135-506546da5152 · outbound
ToSA: Token Merging with Spatial Awareness Efficientvit: Memory efficient vision transformer with cascaded group attention,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9bd3f93e-bb79-45c5-a42a-7ae4eccb9a80 · outbound
ToSA: Token Merging with Spatial Awareness Dynamicvit: Efficient vision transformers with dynamic token sparsification,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 04f5c40e-20c1-44a1-bbde-7afa92299075 · outbound
ToSA: Token Merging with Spatial Awareness A-vit: Adaptive tokens for efficient vision transformer,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8f6813e0-f0c8-4423-bb16-19a4b436e0ba · outbound
ToSA: Token Merging with Spatial Awareness TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4463c08-a9f5-423f-84a5-02b9c02edb06 · outbound
ToSA: Token Merging with Spatial Awareness Token pooling in vision transformers for image classification,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e64efe56-fc1a-4d30-bb07-24fd401a13f5 · outbound
ToSA: Token Merging with Spatial Awareness Zero-shot 3d question answering via voxel-based dynamic token compression,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7d9d162f-cc20-4684-bca7-c8051db30524 · outbound
ToSA: Token Merging with Spatial Awareness Token merging: Your ViT but faster,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 71ebff2d-17a0-4d8a-b8e2-2956bc460053 · outbound
ToSA: Token Merging with Spatial Awareness What do Vision Transformers Learn? A Visual Exploration
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae73add2-f905-443d-b7aa-abbab6e06216 · outbound
ToSA: Token Merging with Spatial Awareness SpatialBot: Precise Spatial Understanding with Vision Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adb71e1e-120d-46fa-8bf5-a8ff8a7e430e · outbound
ToSA: Token Merging with Spatial Awareness Vqa: Visual question answering,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5817ad76-faa0-4ab9-a8ca-94e0cb8a63ae · outbound
ToSA: Token Merging with Spatial Awareness Gqa: A new dataset for real- world visual reasoning and compositional question answering,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6bab8da7-a90f-4414-be8d-514a9a4a2bf9 · outbound
ToSA: Token Merging with Spatial Awareness Openeqa: Embodied question answering in the era of foundation models,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aa4213b9-0494-4f27-ba2b-1861cf231b88 · outbound
ToSA: Token Merging with Spatial Awareness Sp-vit: Learning 2d spatial priors for vision transformers,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d9f1d06b-b422-4d58-aba9-6c12e1392931 · outbound
ToSA: Token Merging with Spatial Awareness Evo-vit: Slow-fast token evolution for dynamic vision transformer,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3f03106a-f2a8-43c6-a505-631d9b8d6fa9 · outbound
ToSA: Token Merging with Spatial Awareness Not all patches are what you need: Expediting vision transformers via token reorganizations,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 718ac7db-d307-450e-a7d2-bd35898f18f9 · outbound
ToSA: Token Merging with Spatial Awareness PPT: Token Pruning and Pooling for Efficient Vision Transformers
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7efb4b4-3a43-4e95-a7fc-162937ae0998 · outbound
ToSA: Token Merging with Spatial Awareness An image is worth 1/2 tokens after layer 2: Plug-and-play inference ac- celeration for large vision-language models,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e2a3183e-3954-428f-8d4d-1c3037f0678d · outbound
ToSA: Token Merging with Spatial Awareness Sparsevlm: Visual token sparsification for efficient vision-language model inference,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 69f882be-bbcd-4864-9259-482ec2173ddb · outbound
ToSA: Token Merging with Spatial Awareness Spatialvlm: Endowing vision-language models with spatial reasoning capabilities,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f2b6a679-237c-449a-ba4d-f6d133ceca6a · outbound
ToSA: Token Merging with Spatial Awareness Spatialrgpt: Grounded spatial reasoning in vision-language models,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation af21c5cc-af7a-428d-8816-33de85b0785c · outbound
ToSA: Token Merging with Spatial Awareness Attention is all you need,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea6ef157-f9ed-4beb-8122-dba2c9bdb969 · outbound
ToSA: Token Merging with Spatial Awareness Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd4975ca-7cd5-41c2-a9d6-060bf6e5d0b1 · outbound
ToSA: Token Merging with Spatial Awareness Instructblip: Towards general-purpose vision- language models with instruction tuning,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 242f7156-e487-47c6-8819-3b72cb02be43 · outbound
ToSA: Token Merging with Spatial Awareness Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3e07ea0-fdbd-4e2b-8c53-1f464f6b9e09 · outbound
ToSA: Token Merging with Spatial Awareness Improved baselines with visual instruction tuning,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 015df59e-fd2d-49de-b961-fa434bf38c49 · outbound
ToSA: Token Merging with Spatial Awareness Llama-vid: An image is worth 2 tokens in large language models,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7563aa29-e8bf-48af-b677-46ce31a1056b · outbound
ToSA: Token Merging with Spatial Awareness Vila: On pre-training for visual language models,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8f2e361f-155a-49e1-8770-0abe81f21a22 · outbound
ToSA: Token Merging with Spatial Awareness Depth anything v2,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3bf9bb16-7765-40f5-b839-9587f6f7bfcb · outbound
ToSA: Token Merging with Spatial Awareness AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcb28fe4-92e5-4dba-ab9a-101a3e869823 · outbound
ToSA: Token Merging with Spatial Awareness Longvlm: Efficient long video understanding via large language models,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8f28d1d8-cdee-4d34-8e7b-83e03edac955 · outbound
ToSA: Token Merging with Spatial Awareness Video-chatgpt: Towards detailed video understanding via large vision and language models,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2b74cc36-57e4-4d1d-99bb-022d8b87d674 · outbound
ToSA: Token Merging with Spatial Awareness Video-llama: An instruction-tuned audio-visual language model for video understanding,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fab74acb-a524-4e80-a5b2-bc01b79f2c1e · outbound
ToSA: Token Merging with Spatial Awareness VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f94ba385-b0bb-4504-a4ab-9870b150d347 · outbound
ToSA: Token Merging with Spatial Awareness Chat-univi: Unified visual representation empowers large language models with image and video understanding,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 91eac91c-5b43-4613-918e-ed4b68643fa5 · inbound
Warehouse Spatial Question Answering with LLM Agent ToSA: Token Merging with Spatial Awareness
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1ff61093-2310-4b77-bb4b-06a867db39d9 · inbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation ToSA: Token Merging with Spatial Awareness
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 974bb5d5-894c-4430-9d4f-4be4e26c6499 · inbound
WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation ToSA: Token Merging with Spatial Awareness
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.