Pith. sign in

Paper Citation Record · LEDGER

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models

As of 23 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2608.01644.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01644 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:51.667788Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f0cac426-2ff1-4169-b70c-0e2061e4f20d · outbound

This paper cites Token Merging: Your ViT But Faster.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Token Merging: Your ViT But Faster

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.531752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.531752Z digest=sha256:40759237bcebe023d074e5a618cc7187bf8f3ff0232aeffbb626a166b8083cf4

Observation e95ad9f4-fe05-433b-9d6a-1e36fc9ffd25 · outbound

This paper cites DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.535123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.535123Z digest=sha256:3b7e4c3c4c51ad0109bd6db92403275fa82e0aef21e8bfaa82e3da54a1dd5de1

Observation a5070a4e-c3c1-4be8-9a28-ba259dfc7587 · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.538242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.538242Z digest=sha256:5fa3976db0ef80c1ad6197b95102b8e2fa6fd50858f38b5f2e12e204941c14e2

Observation 7938faef-9d75-4293-aa42-e02aab49098f · outbound

This paper cites Johnson and Joram Lindenstrauss , journal =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Johnson and Joram Lindenstrauss , journal =

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:52.644192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-04T23:46:51.541420Z digest=sha256:6ef24e2ae0a9d96129c04833b583d40ce289c98b706a37d22886a28e78097d00

Observation 6faa0741-c26d-4918-945f-2a3936bdff27 · outbound

This paper cites Database-friendly Random Projections:.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Database-friendly Random Projections:

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:52.636330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-04T23:46:51.544042Z digest=sha256:8bf5b2dabe0fefff72e6375732c69b63ed3f736c5ee50258e0362fcd8702c3cb

Observation 818cb237-c809-4edf-a4ac-13d5b653a7be · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.546718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.546718Z digest=sha256:9fa904bdba7ca1cc7101866dc96a0b657b13b7bf704b3245db377155f6a27856

Observation 2dad76de-74ca-4d25-993a-6a3554dad9ed · outbound

This paper cites DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.549583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.549583Z digest=sha256:87c6d7f5372d88e35cc63cc0ad6af2096f6c6c44b7d62d25b07d67f083ccef94

Observation a915563f-32d0-46ba-badb-4d8b2818cde9 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.552299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.552299Z digest=sha256:e791aa414d2cbd1c44213ac7c08dbf27b50c902da9cce6aa5b562478e364e86e

Observation 608e3e4a-790e-4780-a5b2-6aeec30f49b0 · outbound

This paper cites IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.554772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.554772Z digest=sha256:d73a6c3a7c88bba84f1dceab3a5b3a6599d47cb5630a0c386e90dcae894c6f7a

Observation 6037976b-5fc2-4d7c-a168-9c15f12ab2bd · outbound

This paper cites PruneVid: Visual Token Pruning for Efficient Video Large Language Models.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.557165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.557165Z digest=sha256:878f2e2b8d85dcc537fdb19ae15d6b8f9adf0566fee4b3571428fa4b4adc6d0a

Observation b88481f7-f192-46f0-b1f2-d96b29e93664 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.559912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.559912Z digest=sha256:349a6719950c4a827ba1ddcbc4527b43a23ed60554c49d6ff8e18a2aaa97007f

Observation 4829ecd5-8452-4bbb-8fbc-339d242d13cd · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.562582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.562582Z digest=sha256:c3a3dbd64f71b3a369a278e85a35a0a2fb5bcebc57b54d157df59ff7ad5086cd

Observation c4c00660-b607-42b9-a766-4b247d742a41 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Flamingo: a Visual Language Model for Few-Shot Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.565175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.565175Z digest=sha256:5ae0af48b8e38886d49302994fd3057fbaef4ca63a4782b7aa9050c5684f748e

Observation 029a34b8-1856-4299-8a32-a5378955ea5a · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.568021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.568021Z digest=sha256:749f867fba060c7658acedfe23289340fcdb5267838a576c9b5f4a5605e8698b

Observation e0bfad12-8a05-46ca-a1dc-50349f089c34 · outbound

This paper cites PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T23:46:52.450235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-04T23:46:51.570775Z digest=sha256:ad5fc5ed2764eb50f4d9fdfa099550159d72ba5c42d34ac107d73368fb925efe

Observation a3c4e44a-af0c-4238-9ed0-9d3ae8b95f91 · outbound

This paper cites LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.573559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.573559Z digest=sha256:4144070d2263f3d60263eac6c8b359a4e8b4394b3b7a2f02c4c538759700b959

Observation 9ba06075-77b6-45c8-97a1-15ca8e42cf37 · outbound

This paper cites InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.576340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.576340Z digest=sha256:8dff05def85d749d21e47343e3841fa815117bf6e890f8740bed79ac8074f74a

Observation 123d02dd-81ae-4c86-8b8b-172b6f77087c · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models International Conference on Learning Representations (ICLR) , year =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.580850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.580850Z digest=sha256:09164c3c5059a76c252a2bce0fe88d3c1b62071bb66fb4a3ef72ff45390dc488

Observation cb192944-990e-4e1c-b078-8a1b37e8223e · outbound

This paper cites Transactions on Machine Learning Research (TMLR) , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Transactions on Machine Learning Research (TMLR) , year =

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.583327Z digest=sha256:ea147d40907e9b8fd9bf36f765452594e674f7803d7364a695bdd55bff0b8617

Observation 666bc42d-4ead-47fb-92f7-0e53f0c16fcc · outbound

This paper cites arXiv preprint arXiv:2505.18227 , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models arXiv preprint arXiv:2505.18227 , year =

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.585864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.585864Z digest=sha256:0037a4866d7f90322220459038154d63944129c654fe16e69ee81fc893ca8d66

Observation 6e0885d0-099b-498c-9184-bfdcca6e8f75 · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.588231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.588231Z digest=sha256:ef3de4825d8530bc4c9f5df83a48597fc33e9160fe13498c8a7c36b085235963

Observation ff6fa96a-a337-43e2-b5ac-823cdd111a8b · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.590927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.590927Z digest=sha256:df79c8151a6866c54f05497c141deb35fb001c03c9b335ad63d3e7ec83697781

Observation 5826cc77-e5f4-404d-a2a2-9e5cf40a94de · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models TempCompass: Do Video LLMs Really Understand Videos?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.593878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.593878Z digest=sha256:eec385a8fc84390884025c88845ff5b37f3b5193ff9777c34fc52a27e43ceeac

Observation 95377740-8756-4f75-bff3-e15c3b62e03a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.596591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.596591Z digest=sha256:fceccfaef37b77f0e0e8281d7f5945e3f29b4917d4ce43d1c9ee123b195e11e4

Observation 1c803385-804f-4d59-9e6e-d27819625dd2 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.599549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.599549Z digest=sha256:dbe7886d636a0883dcd5eef6dbc05f26c49129b056f15e9540c0d2cf8156fae4

Observation 555bb139-5a0c-4ef6-8d12-e2ea525bb244 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.602320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.602320Z digest=sha256:6d52fe2270d9cbf25b5d634dd74cf4eb781023ed37d7f227413b7884cb54ae24

Observation 128b6588-ea48-4bbf-b3f9-f3f9cc72aba4 · outbound

This paper cites Qwen2.5-VL Technical Report.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Qwen2.5-VL Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.605205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.605205Z digest=sha256:30dd84de756269135daef626cd019190a6149328dc0d79e374428f468cf97270

Observation 7dc63e75-60e7-488f-827d-82080ef78b28 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.608318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.608318Z digest=sha256:38b6170cd4c8cb0f191481c529f631cf5cba62ce2c17ab7eb43d3030a45c64f2

Observation cec09a60-c794-45e3-8958-0ab05708261a · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.611444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.611444Z digest=sha256:13c227f4d849e1b78a891051fc5eacba67e4d34fd5c604c15bb6fba8e0da85fd

Observation 5c8d3dba-cc1a-4f75-89fc-61f7f0faaa68 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.614564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.614564Z digest=sha256:28538c916ddd598a02804826ca7e2fd5a96b1683e12c9328e116ffa0c5af69b7

Observation bebe9447-190d-4975-8d70-6088898bdcbf · outbound

This paper cites arXiv preprint arXiv:2510.16598 , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models arXiv preprint arXiv:2510.16598 , year =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.617550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.617550Z digest=sha256:32858d112b1c8ec1048f4871e2fac0846e9f5d111dee8f837ba82c19c618a85b

Observation 5571a5f0-8880-4e9d-8541-74175d42959c · outbound

This paper cites TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.620346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.620346Z digest=sha256:38a18f28ac9e9477b0d79119fd465f0cfaeba9748be1ae72872111535ccf246c

Observation 28dfb864-4a37-4541-a6c0-9505ef6324ae · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.623682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.623682Z digest=sha256:3cebbbcaf0563c419e107bf6d210757eae5c1afbd2987b848ed32c9b8b3fdbe5

Observation 444f74dd-a255-4114-908e-61f7da1616fc · outbound

This paper cites Scalable Diffusion Models with Transformers.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Scalable Diffusion Models with Transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.626661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.626661Z digest=sha256:0579cf9dcefeae36c4b2321d7b25c7ba3c04c9fe1745c0262aa82e5f79da8898

Observation ec101824-b7d8-4757-8953-90f3e2cc7deb · outbound

This paper cites Masked Autoencoders Are Scalable Vision Learners.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Masked Autoencoders Are Scalable Vision Learners

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.629660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.629660Z digest=sha256:4022294d1a3ae340c8a7b5d313476b866a6e896c42b2a93b7a053476f5954b3c

Observation beb3c69b-1797-4c66-853a-d40f3b182606 · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.632640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.632640Z digest=sha256:90ddbae9cc36193d4b2aa0d3e2551b8115cdbe5a9d98f53bc1244799b99ad013

Observation a7fe473c-2c51-41f6-9fe4-6c91d74d1e75 · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.635563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.635563Z digest=sha256:ef62baaf1ed7d6fa0cf07d13a27a182e4e0040231e43519a840ae0d330276d5a

Observation b87acf7c-1da2-4d2b-9cd4-3f9b021cd49e · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.638965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.638965Z digest=sha256:48a4135fa6b545ef75eba4c3f1dac574fceec3f5018d6a87f3fd5e3aa2cf0c1b

Observation 84d2da63-359b-4e4c-9943-42b6095c3fb3 · outbound

This paper cites DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.641946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.641946Z digest=sha256:d979de032c8abafec351e57215ed4bcb6814c24263c6e8b16536bf1b19177879

Observation 9db99812-1f46-48fa-884e-5eb4287f9270 · outbound

This paper cites Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.645057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.645057Z digest=sha256:126c0037c2717056ab3b6ce65f0179a846471f64ba5d1517a70dd865ef815a30

Observation 60ce3a48-5358-4e4a-b3e9-7352a90a19b5 · outbound

This paper cites DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.648153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.648153Z digest=sha256:b2f275fbdfe8436d90943d26ecee4133656956b87f9e70d1dad59dbdd0b4e701

Observation a20520da-add8-468a-a26e-38c53cced558 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.651354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.651354Z digest=sha256:6a09a254d2d999203ce2cfaa059b768be63a9e25f0506cd69a2b143ac6e72134

Observation d890eb34-3b8c-4d7c-8973-4ab572bb59d4 · outbound

This paper cites Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.654336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.654336Z digest=sha256:5b3a3c81553d84920b25f4168d8666f76f7f4ecc816a4b5e967d170df8e45692

Observation 1029c04a-e870-4c34-b21f-3c5564ab4b5e · outbound

This paper cites LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.657161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.657161Z digest=sha256:bd8a9e0251f770f2d9db9f5c813c38c93431d38f91baea5b77be4bc08d501291

Observation b4359393-9c7c-47fa-85a1-44bdbd4b6697 · outbound

This paper cites Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.659904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.659904Z digest=sha256:e7df5431c79d059cb9cf108caeedaf1b38607f19ed54b0f370be43b236dca181

Observation 11af31fa-86ea-4456-9369-545eae1505ae · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.662622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.662622Z digest=sha256:4f6207cd475d4076958234b012eec92edacbf2f54e8e25e43c41c281da7d98bc

Observation 8cbc6767-9f7f-43ff-b579-578a66df21aa · outbound

This paper cites 2024 , eprint=.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models 2024 , eprint=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.665137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.665137Z digest=sha256:906b26c9e29a5798c7f5eaf5dc36c59ccd9f31793a551daa1205ea6a0e70242e

Observation 8b6aee8f-232a-474b-8d9f-7d5811d96d24 · outbound

This paper cites 2026 , eprint=.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models 2026 , eprint=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.667788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.667788Z digest=sha256:571cffe7025c1e2af30a820345f8aa14834321115c74b515f07d875e2e12ef71

Pith citing papers

No inbound Pith citation observations are available.