Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:34:16.703536Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2504.17892.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:34:16.703536Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:31:44.517145Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-11T21:31:45.066251Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4a78ded4-8a92-4aaa-a465-8944dcc5a272 · outbound
Token Sequence Compression for Efficient Multimodal Computing A multimodal architecture for ai agents, 2023
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a5d2f209-181e-496d-8490-615ba2603449 · outbound
Token Sequence Compression for Efficient Multimodal Computing Token merging: Your vit but faster
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 693a7eb5-bfbe-4cb5-81e3-e487fa364f70 · outbound
Token Sequence Compression for Efficient Multimodal Computing An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c43e1943-2d31-402e-8503-8685a083bf33 · outbound
Token Sequence Compression for Efficient Multimodal Computing MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea89dfbb-b98e-4bb9-913a-73c111eb7851 · outbound
Token Sequence Compression for Efficient Multimodal Computing Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c0443a45-2164-455c-b96d-6d119e3cdb53 · outbound
Token Sequence Compression for Efficient Multimodal Computing Perceiver: General Perception with Iterative Attention
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3092b459-a3eb-40c7-8e2b-4c055d4cefc8 · outbound
Token Sequence Compression for Efficient Multimodal Computing Towards efficient visual-language alignment of the q-former for visual reason- ing tasks
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 769df055-957c-4c70-9a3c-61711c810f86 · outbound
Token Sequence Compression for Efficient Multimodal Computing Llava-onevision: Easy visual task transfer
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9cc093b6-24cf-42f2-8872-ac28aa6964cf · outbound
Token Sequence Compression for Efficient Multimodal Computing Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b6c77503-6edc-415d-b963-a5a1caea6dfb · outbound
Token Sequence Compression for Efficient Multimodal Computing Evaluating Object Hallucination in Large Vision-Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef17429c-fcbc-4651-a657-d8d1282c74f3 · outbound
Token Sequence Compression for Efficient Multimodal Computing Vila: On pre-training for visual language models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3d034a8a-b008-4a08-9316-0f1a5407658d · outbound
Token Sequence Compression for Efficient Multimodal Computing Improved Baselines with Visual Instruction Tuning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b10226c-38a9-49f4-8ad8-231b41d9079d · outbound
Token Sequence Compression for Efficient Multimodal Computing Llava-next: Im- proved reasoning, ocr, and world knowledge
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 50708892-7149-437b-9553-899a08765bc6 · outbound
Token Sequence Compression for Efficient Multimodal Computing MMBench: Is Your Multi-modal Model an All-around Player?
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ade9eed7-f91b-4eab-8e3e-677dacce5f1b · outbound
Token Sequence Compression for Efficient Multimodal Computing Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 84364a5f-d43c-4bcb-95ec-5e1b3323eb15 · outbound
Token Sequence Compression for Efficient Multimodal Computing Learning transferable visual models from natural language supervision
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 169b392e-d44b-4f91-a208-eb8275377bae · outbound
Token Sequence Compression for Efficient Multimodal Computing Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2255e0cc-cfa5-4979-9417-87a98d688b15 · outbound
Token Sequence Compression for Efficient Multimodal Computing Towards vqa models that can read
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ed5fc4ea-c91a-443b-b807-3a5321940efe · outbound
Token Sequence Compression for Efficient Multimodal Computing Next-gpt: Any-to-any multimodal llm
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 72e1df33-5ca7-4d9e-8236-2a4417ee3bf8 · outbound
Token Sequence Compression for Efficient Multimodal Computing Visionzip: Longer is better but not necessary in vision language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 581969a4-9e02-4872-bedd-83577b5e6cc4 · outbound
Token Sequence Compression for Efficient Multimodal Computing X-vila: Cross-modality align- ment for large language model
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 736433ff-84ff-4c46-b937-5ef4e7766373 · outbound
Token Sequence Compression for Efficient Multimodal Computing Mm-vet: Evaluating large multimodal models for integrated capabilities
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1976f3d4-e1ff-4750-82be-1a0e20bd777b · outbound
Token Sequence Compression for Efficient Multimodal Computing LLM Inference Unveiled: Survey and Roofline Model Insights
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 980db933-c9f3-4f20-b082-57dd7fdb5880 · outbound
Token Sequence Compression for Efficient Multimodal Computing Sigmoid loss for language image pre-training
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c34eabbf-2f40-45e8-8169-338c768cda8b · outbound
Token Sequence Compression for Efficient Multimodal Computing SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a809c974-a1d4-4aa0-a0aa-b8b498adb9a9 · inbound
p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay Token Sequence Compression for Efficient Multimodal Computing
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.