Pith. sign in

Paper Citation Record · LEDGER

LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2501.03895.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03895 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:32:28.293861Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9da856d3-e288-4a0d-90b0-9d9b6ae8b686 · inbound

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs cites this paper.

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T00:57:32.564573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:57:32.564573Z digest=sha256:be298d10d76609aaf23daadaa52963faf3c4c5551550322a6ef74468b799a02b

Observation 86aeb704-3adc-4170-afc5-5eac48021193 · inbound

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler cites this paper.

TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T14:18:18.012399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:18:18.012399Z digest=sha256:456a145a3408124977099db63bf5a20a85807384dadc132ce907f781bc6bf402

Observation c5d1cee3-db0c-4fd8-b453-1aa778b5ed74 · inbound

VideoRoPE: What Makes for Good Video Rotary Position Embedding? cites this paper.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.222155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.222155Z digest=sha256:9f9bf8ea3c50fefef00fe4c9f3673dcbf97df2f18ea1f67fbe438fa0202f312e

Observation 80dec6b0-91d2-4194-a46b-e89c217429fe · inbound

IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs cites this paper.

IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:32:28.293861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:32:28.293861Z digest=sha256:8dbc5579e45c9cab161d6129ebed9e74eefd64fdf65444c189c87e337c02226b

Observation 5fa331de-f931-4698-b381-4bd57deec528 · inbound

DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs cites this paper.

DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T10:55:03.234346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:55:03.234346Z digest=sha256:2b8ea90105da18d28ebb570970b2c03bdcb2ca87a836020f9656d072276d4ef0

Observation 36dc7495-b72b-4676-9c79-17a83e1ba0e0 · inbound

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning cites this paper.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.226484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.226484Z digest=sha256:aab2df999198039534a7237a269c5adbec2713a3d905b64ada8fb982019654c1

Observation 4e812d10-ea7d-45f4-80f4-3e28783cfb9e · inbound

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models cites this paper.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:04.258730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:04.258730Z digest=sha256:29ce31e42d74909809ca3c4cb8df51efa2d1adc7a6714a1dd05d0e2c7687f574

Observation cb4cd5d7-3937-4748-ae83-c7fe93378ead · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.869569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.869569Z digest=sha256:4709cc991b1ae9e9fc0571e804019047b9ba45cd5529d97d3b61c34460e9a3d6

Observation d1f0c0cb-268c-402c-aad2-c20d039f92d4 · inbound

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding cites this paper.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.411061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.411061Z digest=sha256:e3274508c849d6936c37e23d9013e2221079bf57e922a6dbdfc9c40417fe576b

Observation 220f8bba-cf8e-446a-b0d3-556a1f5f6b26 · inbound

NoLoCo: No-all-reduce Low Communication Training Method for Large Models cites this paper.

NoLoCo: No-all-reduce Low Communication Training Method for Large Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:03.894658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:03.894658Z digest=sha256:2755567429f07730af2ba6d492f6371e4ce29e86135fcf484f03de3060a130b2

Observation a45288c7-d67a-44b0-81fd-85d7a70a5977 · inbound

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes cites this paper.

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T22:37:40.244937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:37:40.244937Z digest=sha256:79d267ecf1145045799e73286fe9b9ff1c6f380fe1959c79078e9994aae79e1e

Observation a9433110-79f7-4096-895e-e2f652542aaa · inbound

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning cites this paper.

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T18:39:07.591851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:39:07.591851Z digest=sha256:5c9f1b949f531d09c67e7e803344e82e07f809b0394608a2039d1b7ea6bcdbc7

Observation 806989ac-a5bb-43c2-aff6-6e3192fa1a0a · inbound

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models cites this paper.

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T05:51:16.098415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:51:16.098415Z digest=sha256:d9c151425f1476f09f81f5bc35173ba1ce6328b6bda8019ce65ed583f6c75719

Observation b7271e21-b9d0-4e84-8ef0-a93975e2687e · inbound

Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models cites this paper.

Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:40:43.060000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T22:40:39.892802Z digest=sha256:4bd73bd7ca079a9d31eeb53e129b9ec02f56ad1fd187f1756920869367f58b34

Observation be70dd87-0aee-4f07-a90a-33ab85d0a529 · inbound

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors cites this paper.

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T13:05:56.397948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:05:56.397948Z digest=sha256:88e54e2a3c841b75e5359a692bfe3a32ed8d43e23f416e06f4c9700224e8d803

Observation 20621b8b-84bf-4333-bd44-84af059d8666 · inbound

Synthetic Homes: A Multimodal Generative AI Pipeline for Residential Building Data Generation under Data Scarcity cites this paper.

Synthetic Homes: A Multimodal Generative AI Pipeline for Residential Building Data Generation under Data Scarcity LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:11:39.115233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T17:10:57.875842Z digest=sha256:40f99baf4560d42e8c4c61d6a13f6e658f2142bec37a95dcc3ba28b7cc398e43

Observation 4730ebf1-852b-481d-81f1-9ddd9edb92b7 · inbound

Synthetic Homes: A Multimodal Generative AI Pipeline for Residential Building Data Generation under Data Scarcity cites this paper.

Synthetic Homes: A Multimodal Generative AI Pipeline for Residential Building Data Generation under Data Scarcity LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T16:08:37.856081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:08:37.856081Z digest=sha256:77028f13939b45979af99b266c779b64a84c291e3e5b83375c831b5c390860a1

Observation 28e56081-00fd-499b-8177-d31081ce1074 · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:43.035260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:43.035260Z digest=sha256:0841e4c953d439cfa821a8c60f475e0f8e6ffe1ec58a89387e16da4b8835443b

Observation 4557b983-9211-4401-936e-7bf88f5b7569 · inbound

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation cites this paper.

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T12:02:02.320472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:02:02.320472Z digest=sha256:51e4b0f94ca862949d81d505c1ba91b4670c153d951d1359375ebcc276f6ad1b

Observation ba33d555-e37f-4f04-b8c6-f4f9b1180065 · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:59.229137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:59.229137Z digest=sha256:1ca02fa98d2b345f4dc3403d92782b17166d16fc64e8c5ea3df27953957135f8

Observation 1d3bd22d-f1d0-434f-a862-2d9b7bdc212a · inbound

Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models cites this paper.

Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:58.504522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:27:58.757680Z digest=sha256:9937de7f2b2f676f0abb5cb051cf897b92c79906fdfbedfa20390178c216f62c

Observation ee25deb0-089f-455c-adea-041508b539fc · inbound

Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models cites this paper.

Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.365574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T08:39:20.202572Z digest=sha256:d4ca37ca2df07171905bdc862e3228542528333d77cbfd8ae2ea8e60e88b67b0

Observation 5f1b2b21-8ab8-422c-9e15-cc1db7e8f5bc · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:03.736027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:c3ca1b890d937dcdc7b4cea3a98696ccd544f2e623f165e326c36ab60fa86f27

Observation 52b28f62-8889-4351-94fb-d11f64bf0b04 · inbound

Geometry-Guided 3D Visual Token Pruning for Video-Language Models cites this paper.

Geometry-Guided 3D Visual Token Pruning for Video-Language Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:51:09.946533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T05:49:38.346274Z digest=sha256:a1923050a9d3be5db46e1b7b4a1658ff5af5fca48988fb9d4f8b31e495105332

Observation 037ff41f-8d0c-4e46-98a5-f3a2270c6af9 · inbound

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading cites this paper.

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:06.496376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T16:10:22.588945Z digest=sha256:06ab8aa23e7e562ad00f2aee8958319f1b714882a1f2ba96205486d2270e9c0d

Observation bec6f267-091b-4cb4-a809-46d6344c8821 · inbound

LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute cites this paper.

LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:40:52.141344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T01:36:50.406353Z digest=sha256:8e52a64399d66b0f8fec8dd91db7531ca92afc830c06702abbf215d30277605e

Observation feb54c7e-bd0f-41d2-8b8f-9b63ec5c1470 · inbound

OProver: A Unified Framework for Agentic Formal Theorem Proving cites this paper.

OProver: A Unified Framework for Agentic Formal Theorem Proving LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:48:23.428962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T14:43:46.517807Z digest=sha256:200f2096f780c9df6030dd175ae9b10f83d8bad5604b48eb1d5c33b726a97a3b

Observation 171008aa-deaf-41a7-a9c0-d7ae860611bf · inbound

Focus-then-Context: Subject-Centric Progressive Visual Token Reduction for Vision-Language Models cites this paper.

Focus-then-Context: Subject-Centric Progressive Visual Token Reduction for Vision-Language Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:23:58.577019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-21T05:20:55.448430Z digest=sha256:2401f8df4aa35085ac2d33e91e357b5340c70a431b819eeec2aa1bb319945adb

Observation 690383a7-c803-4306-8b92-9682bf0bf7e8 · inbound

VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding cites this paper.

VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:34:02.610791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T22:24:20.787671Z digest=sha256:14cfbde8100d3d2fea2e934a8ba17de99f8ee69ffc315c43984a90f6d9856433

Observation c62617d9-5947-41e0-b401-21df000f508f · inbound

CIVIC: End-to-End Sequence Compactness for Efficient Vision-Language Models cites this paper.

CIVIC: End-to-End Sequence Compactness for Efficient Vision-Language Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:13:26.440208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T12:12:51.867760Z digest=sha256:c52914f06a9fe7c3bd21e342c577050dc5dc458904ce417738d69d9aac66fc7f

Observation a0720b38-b6e7-43a7-abf4-ca9f78a33327 · inbound

The Hidden Evolution of Disguised Visual Context inside the VLM cites this paper.

The Hidden Evolution of Disguised Visual Context inside the VLM LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:31.826417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T18:08:56.044278Z digest=sha256:aa47c28e4e76f54642a7e849b2fd98004b1308f20a4da4f2adcc6013e2b3940b

Observation c01c7c92-3aae-491b-a2ad-9db3abb7ac87 · inbound

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval cites this paper.

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T05:49:36.567425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T15:34:55.062016Z digest=sha256:3ceb278cfc5a72a583dbdfcd7cad3f427a7f3c12fccf609272bf00a45638adc1

Observation 85f39ace-49c0-4eb0-b2ad-36bfa1e9112e · inbound

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs cites this paper.

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.604793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T05:41:04.184461Z digest=sha256:84cb12386cf4f528d91364f8db3ffc67334abb25c118a598dc38318671fd17ed

Observation 500c4645-e0ec-476e-979a-18597bb1ea26 · inbound

EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage cites this paper.

EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 83

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:34:32.421404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T05:25:07.553526Z digest=sha256:5648495f4669282fa35c1488ddeb55d7bda6b623e2293a889fd13fa9083a1eb0

Observation af44867e-005c-483f-afe5-ce72a7bc4d89 · inbound

PDD-RRG: Posterior Diagnostic Decision for Study-level Radiology Report Generation cites this paper.

PDD-RRG: Posterior Diagnostic Decision for Study-level Radiology Report Generation LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:38.855362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T01:03:38.855362Z digest=sha256:66c989d3a08f9aef121284deefa89ae7c585ea092dc68f1ebef2063164bfee9a

Observation 100f7f4a-5ab0-4897-aeac-f15ce15c3ddc · inbound

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin cites this paper.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.274367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.274367Z digest=sha256:3f0da07eedb3cbc6029e9126ee1a8a969c637871e78752543e56c6d4d1c5eb25

Observation 1cde152f-f8de-45dd-8e0e-f00f4a94f582 · inbound

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning cites this paper.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.468007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.468007Z digest=sha256:defe5986f63bef9647151810ef7833ff2ec1837e65c42454eb453f5984de51ac