Pith. sign in

Paper Citation Record · LEDGER

ToSA: Token Merging with Spatial Awareness

As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 3 inbound Pith citation observations for arXiv:2506.20066.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20066 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:01:29.963011Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:29:38.951111Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:29:40.174642Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1e43f906-9452-4009-90bd-a78c62ddc673 · outbound

This paper cites Dinov2: Learning robust visual features without supervision,.

ToSA: Token Merging with Spatial Awareness Dinov2: Learning robust visual features without supervision,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.312138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.840335Z digest=sha256:5a3d20aba05ed8c35b47827fe7fd0ffa15e6f1d6116f3d03d2a8cc35356c4b95

Observation 31346f8b-c886-440a-8a6c-0d45e01fb506 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

ToSA: Token Merging with Spatial Awareness Learning transferable visual models from natural language supervision,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.843839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.843839Z digest=sha256:93e031cbc86efdb9ed0d31e07ffd9aa5be81b7e2418e31e307d00abc9684c26d

Observation 47d741a3-4730-4439-8a42-c0a638032c5f · outbound

This paper cites Sigmoid loss for language image pre-training,.

ToSA: Token Merging with Spatial Awareness Sigmoid loss for language image pre-training,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.296994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.848181Z digest=sha256:8a211847abd61b6de4ae94b840635e0d045b0515c1d3fed0b807279780ede3c3

Observation a04e200e-2497-4f6e-b485-3f1c3baefbb5 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

ToSA: Token Merging with Spatial Awareness An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.852111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.852111Z digest=sha256:9d16f1fddc85bfc132e60326139863f9d5b1a2d53a8546893b4a0b742766b881

Observation 72db5e3e-9c2b-4405-babd-6f01a8e834c0 · outbound

This paper cites Visual instruction tuning,.

ToSA: Token Merging with Spatial Awareness Visual instruction tuning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.855346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.855346Z digest=sha256:57d11c0695b2cb15bd8eca9c7c2c931d4ce91d2179116fd288cd46623593e813

Observation a940bb97-d4fd-4cf9-8501-0c57c3225c6c · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

ToSA: Token Merging with Spatial Awareness LLaVA-OneVision: Easy Visual Task Transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.858437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.858437Z digest=sha256:142299c15555a0838790b5ffe10a86a8d750a2cde2e15555446ef06158665bcd

Observation e74f5825-bd81-4f70-a135-506546da5152 · outbound

This paper cites Efficientvit: Memory efficient vision transformer with cascaded group attention,.

ToSA: Token Merging with Spatial Awareness Efficientvit: Memory efficient vision transformer with cascaded group attention,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.282019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.861924Z digest=sha256:8aa4cf5f4b02a8cdd026cbc012548d521e9c9cda4700ac2649a89bb34d6fc3dd

Observation 9bd3f93e-bb79-45c5-a42a-7ae4eccb9a80 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification,.

ToSA: Token Merging with Spatial Awareness Dynamicvit: Efficient vision transformers with dynamic token sparsification,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.272427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.864786Z digest=sha256:52ca4e44fab2e2b3e9de459c37bfdefc8f4199b37c6b0961c35960448f5551e6

Observation 04f5c40e-20c1-44a1-bbde-7afa92299075 · outbound

This paper cites A-vit: Adaptive tokens for efficient vision transformer,.

ToSA: Token Merging with Spatial Awareness A-vit: Adaptive tokens for efficient vision transformer,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.262824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.867828Z digest=sha256:8a824a555c27179b77e77c2acb16fdb2fb3e66447f08c869464105d7f5b1d383

Observation 8f6813e0-f0c8-4423-bb16-19a4b436e0ba · outbound

This paper cites TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action.

ToSA: Token Merging with Spatial Awareness TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.871387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.871387Z digest=sha256:968475057849a0d5e7d6f9a116786fadf618f0d73957617588303e132cb2fc7d

Observation b4463c08-a9f5-423f-84a5-02b9c02edb06 · outbound

This paper cites Token pooling in vision transformers for image classification,.

ToSA: Token Merging with Spatial Awareness Token pooling in vision transformers for image classification,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.253657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.874526Z digest=sha256:d72e565a261eb67470b00ebd9d7761f33f63388bfe6788aebc4de2bcd3ed1d05

Observation e64efe56-fc1a-4d30-bb07-24fd401a13f5 · outbound

This paper cites Zero-shot 3d question answering via voxel-based dynamic token compression,.

ToSA: Token Merging with Spatial Awareness Zero-shot 3d question answering via voxel-based dynamic token compression,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.244426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.877445Z digest=sha256:8ac5f80d269775c83733c7522c49f0a5de896d8749308d102dd18afcfc25515e

Observation 7d9d162f-cc20-4684-bca7-c8051db30524 · outbound

This paper cites Token merging: Your ViT but faster,.

ToSA: Token Merging with Spatial Awareness Token merging: Your ViT but faster,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.235279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.880629Z digest=sha256:fd7cdcbbb02d30b298c2b9055c0ef3b3e5be4d2c8516414245e0b415e789d345

Observation 71ebff2d-17a0-4d8a-b8e2-2956bc460053 · outbound

This paper cites What do Vision Transformers Learn? A Visual Exploration.

ToSA: Token Merging with Spatial Awareness What do Vision Transformers Learn? A Visual Exploration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.883551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.883551Z digest=sha256:4818c4958f6a75b3da55f4c5b38bf10342747f2794c2deb30c695a3cc74d8168

Observation ae73add2-f905-443d-b7aa-abbab6e06216 · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

ToSA: Token Merging with Spatial Awareness SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.887022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.887022Z digest=sha256:273c46bb3d1ae0424613e1b65c71f9fdaa642987c5ef95bcfec96ef21f7dfa8d

Observation adb71e1e-120d-46fa-8bf5-a8ff8a7e430e · outbound

This paper cites Vqa: Visual question answering,.

ToSA: Token Merging with Spatial Awareness Vqa: Visual question answering,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.890588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.890588Z digest=sha256:d6964c791cbf90ea0c704a76fb8e6d407e49dda93a58691f1e6feb12b7229a3f

Observation 5817ad76-faa0-4ab9-a8ca-94e0cb8a63ae · outbound

This paper cites Gqa: A new dataset for real- world visual reasoning and compositional question answering,.

ToSA: Token Merging with Spatial Awareness Gqa: A new dataset for real- world visual reasoning and compositional question answering,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.220159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.894119Z digest=sha256:c069da438403ba48d698dbf2984bdf2493cfb3c97be3a1497d9bfdf5da365cc6

Observation 6bab8da7-a90f-4414-be8d-514a9a4a2bf9 · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models,.

ToSA: Token Merging with Spatial Awareness Openeqa: Embodied question answering in the era of foundation models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.211596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.896953Z digest=sha256:b1d94a8de3712b7791259eb7550bf1e7f660aecfccb83112ef44d38c6a14ebb7

Observation aa4213b9-0494-4f27-ba2b-1861cf231b88 · outbound

This paper cites Sp-vit: Learning 2d spatial priors for vision transformers,.

ToSA: Token Merging with Spatial Awareness Sp-vit: Learning 2d spatial priors for vision transformers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.202142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.899813Z digest=sha256:4a4788a87a15f989009ef4ac0bc07ec9f6f525e2c88e60431f6ab09761d03cbd

Observation d9f1d06b-b422-4d58-aba9-6c12e1392931 · outbound

This paper cites Evo-vit: Slow-fast token evolution for dynamic vision transformer,.

ToSA: Token Merging with Spatial Awareness Evo-vit: Slow-fast token evolution for dynamic vision transformer,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.193380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.902805Z digest=sha256:45f5fc2d39025f3776b2fc037655a419092860ab9a97e65732d3cd2b9ce3072c

Observation 3f03106a-f2a8-43c6-a505-631d9b8d6fa9 · outbound

This paper cites Not all patches are what you need: Expediting vision transformers via token reorganizations,.

ToSA: Token Merging with Spatial Awareness Not all patches are what you need: Expediting vision transformers via token reorganizations,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.183813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.905853Z digest=sha256:40d62f95e8ef9b2a10b6576591e015fdd831f2091ed39c438a09198792696117

Observation 718ac7db-d307-450e-a7d2-bd35898f18f9 · outbound

This paper cites PPT: Token Pruning and Pooling for Efficient Vision Transformers.

ToSA: Token Merging with Spatial Awareness PPT: Token Pruning and Pooling for Efficient Vision Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.908663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.908663Z digest=sha256:0969ea812635ca48f19f5fe19a97944b169f397c8ca52d5c84a8a0d623857918

Observation d7efb4b4-3a43-4e95-a7fc-162937ae0998 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference ac- celeration for large vision-language models,.

ToSA: Token Merging with Spatial Awareness An image is worth 1/2 tokens after layer 2: Plug-and-play inference ac- celeration for large vision-language models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.174741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.911840Z digest=sha256:bf89926eff882d9f03f8de2238baae421e8a9dd72c3d143d8fbd37dc477ec04b

Observation e2a3183e-3954-428f-8d4d-1c3037f0678d · outbound

This paper cites Sparsevlm: Visual token sparsification for efficient vision-language model inference,.

ToSA: Token Merging with Spatial Awareness Sparsevlm: Visual token sparsification for efficient vision-language model inference,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.164160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.914750Z digest=sha256:1b7e207eaa9f81202cc87ab41447e8c77373acfcdbe16b384591f5401db1ddad

Observation 69f882be-bbcd-4864-9259-482ec2173ddb · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities,.

ToSA: Token Merging with Spatial Awareness Spatialvlm: Endowing vision-language models with spatial reasoning capabilities,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.154332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.917634Z digest=sha256:86b5b42a31549945bf6f71f7fe126ae051e4a59ba8bb7a05956211753c9d5a65

Observation f2b6a679-237c-449a-ba4d-f6d133ceca6a · outbound

This paper cites Spatialrgpt: Grounded spatial reasoning in vision-language models,.

ToSA: Token Merging with Spatial Awareness Spatialrgpt: Grounded spatial reasoning in vision-language models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.145170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.920675Z digest=sha256:cbff3e250ef43d4686e20b7e361d9c6c26f87abafdc6ac4d87c3a2f13f92d60b

Observation af21c5cc-af7a-428d-8816-33de85b0785c · outbound

This paper cites Attention is all you need,.

ToSA: Token Merging with Spatial Awareness Attention is all you need,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.923678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.923678Z digest=sha256:e6ca4441a7d89fed470ff793e329717b9be38b1962517193e526cf9302bfe9e1

Observation ea6ef157-f9ed-4beb-8122-dba2c9bdb969 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

ToSA: Token Merging with Spatial Awareness Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.926930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.926930Z digest=sha256:c8513b3238e68ded2e1580f44c5060b30ce75c8f908aa45c0afb73c0695b3783

Observation bd4975ca-7cd5-41c2-a9d6-060bf6e5d0b1 · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning,.

ToSA: Token Merging with Spatial Awareness Instructblip: Towards general-purpose vision- language models with instruction tuning,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.929843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.929843Z digest=sha256:319584bf6e986b8e8e1589910ea3965d01c766c233b710e2b227352e085ec0df

Observation 242f7156-e487-47c6-8819-3b72cb02be43 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

ToSA: Token Merging with Spatial Awareness Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.932806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.932806Z digest=sha256:19a5831ecfee74a9c441cb47ee5223e87e4dee18913bdec1c20e0a5c5487c9d3

Observation a3e07ea0-fdbd-4e2b-8c53-1f464f6b9e09 · outbound

This paper cites Improved baselines with visual instruction tuning,.

ToSA: Token Merging with Spatial Awareness Improved baselines with visual instruction tuning,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.936217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.936217Z digest=sha256:15f2d9fc0941ee1a1472f32de7e1eab747733c374425bae0d97555afb4ecbbb4

Observation 015df59e-fd2d-49de-b961-fa434bf38c49 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models,.

ToSA: Token Merging with Spatial Awareness Llama-vid: An image is worth 2 tokens in large language models,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.939078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.939078Z digest=sha256:6b4e6b5172659675050083c07762b9c8bea6ca2d45eb2a9e0f3adedea08dfc80

Observation 7563aa29-e8bf-48af-b677-46ce31a1056b · outbound

This paper cites Vila: On pre-training for visual language models,.

ToSA: Token Merging with Spatial Awareness Vila: On pre-training for visual language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.109792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.941985Z digest=sha256:2a7eebaa01a5828962f9ee13945c569f5cd1c42e5f90ebb4d4cb5eab54a2fcca

Observation 8f2e361f-155a-49e1-8770-0abe81f21a22 · outbound

This paper cites Depth anything v2,.

ToSA: Token Merging with Spatial Awareness Depth anything v2,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.100457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.944983Z digest=sha256:10d8239fa0224b4f6c095f3f122ae3a5c5462540eb1dd9d791f0180321c8e08a

Observation 3bf9bb16-7765-40f5-b839-9587f6f7bfcb · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

ToSA: Token Merging with Spatial Awareness AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.947783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.947783Z digest=sha256:e96635bc2e9caaad844179d5eb468f566e98df0e7903117bb660935ef185be16

Observation bcb28fe4-92e5-4dba-ab9a-101a3e869823 · outbound

This paper cites Longvlm: Efficient long video understanding via large language models,.

ToSA: Token Merging with Spatial Awareness Longvlm: Efficient long video understanding via large language models,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.091338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.951218Z digest=sha256:7b441d5f0efa48123af8041b507d541ca037361760e69beb1e250273f2f4df34

Observation 8f28d1d8-cdee-4d34-8e7b-83e03edac955 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models,.

ToSA: Token Merging with Spatial Awareness Video-chatgpt: Towards detailed video understanding via large vision and language models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.082347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.954230Z digest=sha256:d9efe1488412503e3edadd50b01283c0fdedb80249389d3459f33b6b955a341f

Observation 2b74cc36-57e4-4d1d-99bb-022d8b87d674 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding,.

ToSA: Token Merging with Spatial Awareness Video-llama: An instruction-tuned audio-visual language model for video understanding,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.072638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.957084Z digest=sha256:29cad794c51ce99171618fc534ee84aede61b93369328d4acc0f0381a9defd79

Observation fab74acb-a524-4e80-a5b2-bc01b79f2c1e · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

ToSA: Token Merging with Spatial Awareness VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.959967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.959967Z digest=sha256:515262e9df51682a9b14ddf92600edb9e3769154cadf35f34c5895d68f9da48b

Observation f94ba385-b0bb-4504-a4ab-9870b150d347 · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding,.

ToSA: Token Merging with Spatial Awareness Chat-univi: Unified visual representation empowers large language models with image and video understanding,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.063216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:01:29.963011Z digest=sha256:68955146a1e850843c57d51efc3d9129423f8705e3923079db40ba5f491c2993

Pith citing papers

Observation 91eac91c-5b43-4613-918e-ed4b68643fa5 · inbound

Warehouse Spatial Question Answering with LLM Agent cites this paper.

Warehouse Spatial Question Answering with LLM Agent ToSA: Token Merging with Spatial Awareness

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:29:40.213090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:29:38.951111Z digest=sha256:54ab70bc497445d9396fd6f9103e6c5be5f1031fc77ce874afa2372de031bfcc

Observation 1ff61093-2310-4b77-bb4b-06a867db39d9 · inbound

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation cites this paper.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation ToSA: Token Merging with Spatial Awareness

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T23:59:16.247451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:59:16.247451Z digest=sha256:c281db95c865b6b44479649a9504c9492ba9b927d3094f477d06893a7e71219f

Observation 974bb5d5-894c-4430-9d4f-4be4e26c6499 · inbound

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation cites this paper.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation ToSA: Token Merging with Spatial Awareness

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.518353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.518353Z digest=sha256:7be1ceddbebe8b4961bf12b47bbd0d4e4910a118746531eb27544b94a1f5d1ff