Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:37:08.457373Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 1 inbound Pith citation observation for arXiv:2412.04317.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:37:08.457373Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T15:40:05.730181Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T22:16:15.637500Z
87 of 87 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b8a476c1-9b94-4994-8048-469c0fea17c3 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 630a6798-b520-4f35-bebb-085cdcd1846f · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Bottom-up and top-down attention for image captioning and visual question answering
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7153172c-ea37-4214-8c96-4f76675d4715 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Qwen Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bbea82d-0bf0-4f1b-98f7-05fb89259d53 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Gemma: Introducing new state-of-the-art open models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48487edc-c0c6-4e0c-9400-a6d7a23f3fca · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Language models are few-shot learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 997d8738-1bd1-4fd5-9852-7e9055b7c5f3 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Honeybee: Locality-enhanced projector for multimodal llm
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5ce3a98-1b80-4e69-8f83-431aa4261138 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2f7a125-3085-4c86-8b39-27d816b7fb04 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Lawrence Zit- nick
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25bd50ef-cdc9-4fcf-95db-4c4bdd09e788 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression PaLI: A Jointly-Scaled Multilingual Language-Image Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4da9f6f9-2f64-4a7e-bc1e-72e72a807aee · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression How far are we to gpt-4v? closing the gap to commercial multimodal models with open- source suites, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dffd154-db90-4c29-a3ad-c7e60474eb54 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cfabdcd-b567-4d75-9956-05574f714485 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9544fdc1-b620-420f-9a3f-199c7d42de2d · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Instructblip: Towards general- purpose vision-language models with instruction tuning,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 402e63c1-ce06-472d-b65b-62a8131ba07d · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mme: A compre- 9 hensive evaluation benchmark for multimodal large language models, 2024
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2e3c46bf-2041-457c-a0a3-71607bc6ddfc · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 527398e0-0bf9-48a6-a377-b88a7d565e53 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Minicpm: Un- veiling the potential of small language models with scalable training strategies, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation eab121c9-2d36-4c35-a170-f3587b65e6eb · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f0c000c-631d-492c-97a2-2c3ead4dec59 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Token merging for training- free semantic binding in text-to-image synthesis, 2024
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dc1890a4-bb9d-4e1a-bc26-0cbe64f298a1 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e149fc0b-b43f-464a-b106-f780aadeb0cb · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Phi-2: The surprising power of small language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5200a819-4edd-424e-bdc8-2efbd06283bb · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression In defense of grid features for visual question answering
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2f69eca5-7024-437d-b957-f3d20f9fb4be · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Contrast and classify: Training robust vqa models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 295cde2f-b1d4-433a-9d22-a4aa1aad5b48 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Referitgame: Referring to objects in pho- tographs of natural scenes
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 14175937-1d25-4def-9bdb-51590fdf9c75 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression A diagram is worth a dozen images
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19186a00-6f9d-426d-a3ce-03d6c2c877e0 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Seed-bench: Bench- marking multimodal large language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6fb39a02-0e10-4c36-befe-608ae9bd53d9 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f294ee1a-9caf-4144-95ce-b22068d88a12 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3339ae5-77eb-42b9-97b5-e22a2112aad6 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 64483dff-5527-4c2b-ba48-8dfb8a2a7b10 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0e2b9f2-4dee-4358-9153-54d800348272 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Evaluating object hallucination in large vision-language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation be65e829-2862-4674-a628-60aeec5efeb0 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ec3cd88-490d-46c8-b9b7-102a66d57b8f · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mon- key: Image resolution and text label are important things for large multi-modal models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1a3aeb31-8b46-455b-aa92-cc18366b8b84 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Improved baselines with visual instruction tuning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e5c29df2-6b30-4368-b898-6e8415d2ca4f · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b77bb695-f88d-4844-aa9e-2d08401f8d9f · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Visual instruction tuning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8e6abf79-d0d5-4d55-a27f-6f3591ac9396 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mmbench: Is your multi-modal model an all-around player?, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4effb140-fbb8-4d2c-943c-6b91af32192d · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Decoupled Weight Decay Regularization
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26d39eb5-e572-43be-ad1b-f553186cd52e · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab43a5b8-737a-4264-a66a-835219ebd03b · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72c5abab-e79f-4606-9c23-25d0eb12da07 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7327ee58-e4d1-4ef8-ac14-950aa62dff77 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression To- wards lightweight transformer via group-wise transforma- tion for vision-and-language tasks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f496d380-de84-48cd-9559-242a5d46327a · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Cheap and quick: Efficient vision- language instruction tuning for large language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f99125b1-42d1-4b03-b293-14644c5b97a5 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Moil: Momentum imita- tion learning for efficient vision-language adaptation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9fbdc75-f9fc-4890-b09e-5eb04d28568a · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Towards language-guided visual recog- nition via dynamic convolutions
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a0e638c8-4f5a-4ed1-ba4a-15faa67b0baf · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e27a5ca-7e74-46f2-ae78-61ad377c1b7f · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Chartqa: A benchmark for question answer- ing about charts with visual and logical reasoning, 2022
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b94b76ec-e7bd-4786-92b9-4028b6f44007 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0cc2e6e6-aab0-40db-b992-db6c012216f2 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression DINOv2: Learning Robust Visual Features without Supervision
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 208567de-d3f6-45da-9e74-a99002bd76a9 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Learning transferable visual models from natural language supervision, 2021
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f54f0555-8992-441c-8355-7a9e7a0297b2 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Learning transferable visual models from natural language supervi- sion
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f09e7d82-b271-4736-aea6-4599f8fcc571 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Imp: Highly Capable Large Multimodal Models for Mobile Devices
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 873ba2fe-e190-4711-860f-6bf6827ed107 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression When do we not need larger vision models?, 2024
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 16054fb2-0a0c-4f1b-8662-43a7cd5815a4 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60bbf9ad-f9a3-4482-b029-b3826a42ba27 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Towards vqa models that can read
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 32e1e7cc-c75c-48c7-bf92-4d50e1954a75 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Gemma: Open Models Based on Gemini Research and Technology
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92d80c96-c396-4a7f-8abc-a8b4742a2691 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef2dddbc-3ab3-413d-bbb2-889b9be9d0df · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4b708de-7375-4531-97d7-9e73da4f9c3e · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f271e797-ccf5-484f-8716-5a434963d2a9 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df808884-647b-4e16-bd72-646e0931afae · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Stacked attention networks for image question answering
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2bf59c01-1ee7-41e3-ad7f-a02e7a9590f1 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression mplug-owl: Modularization empowers large language mod- els with multimodality, 2024
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 55a8b314-8652-4f49-8c5e-a1fc978e722e · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92f936f9-cd1b-4c9a-9078-69b019d9c7ef · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mm-vet: Evaluating large multimodal models for integrated capabilities, 2023
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a8ab1367-712e-4e3e-ab80-7da9fa0e8e85 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Deep modular co-attention networks for visual question an- swering
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f72e0beb-c44f-48c2-96de-368aa8556536 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Tinygpt-v: Efficient multimodal large lan- guage model via small backbones, 2024
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation caa8085a-2a3c-4a56-a9cc-ff0e030ee692 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b2f69a39-5f6d-4fe1-94ea-6661d0005158 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Sigmoid loss for language image pre-training,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acf37657-d8f1-4fec-a07d-c07e8a932943 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0f53947-2e19-437d-93ad-4e44cb06403e · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Lmms- eval: Reality check on the evaluation of large multimodal models, 2024
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c07de0f-ffba-412b-b950-8ba9bbd313f2 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Vinvl: Revisiting visual representations in vision-language models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4536e021-a980-4ee2-8a99-b48b30bccffe · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression OPT: Open Pre-trained Transformer Language Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dd6a1d4-33c7-4ec0-94cd-52a422597aae · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Free vqa models from knowledge iner- tia by pairwise inconformity learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 39a79b3a-3f8d-4519-9f6c-dfb5d8b3ebd6 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Trar: Routing the attention spans in transformer for visual question answering
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 863d176b-39ce-4965-b82c-bc35082c741d · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6d31a72e-f5cf-467c-bda0-e859e0ea21ac · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mipha: A Comprehensive Overhaul of Multimodal Assistant with Small Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a1cb42e-0145-42bc-93d2-d71b52a7757a · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 044164c0-8c90-40f7-8768-cc4f0c7ebb24 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9803d311-a481-4c98-9916-a256a7ef60a0 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b8c945af-388a-4435-8980-7a542ac4e4a4 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e7484725-fc2b-4fdf-8647-f81f95fc8b3c · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5d5ce4cc-e1d8-4f57-97bb-d4ec7709e76a · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Therefore, the value of the square in the figure is 2
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 823a836a-6e37-43b4-bf38-f2db8912ba61 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression control center
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3ab4ddba-1331-4c01-9e1b-b3940d8881e0 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ad67f5a0-d0b2-4b9b-95da-b78400b582a0 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 275209cb-2bf7-4178-aac4-6d48a670f5a5 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9c492150-ab38-4202-8025-e9e1cf4e7473 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression RIGHT",
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6c5fd70e-35c4-484c-aaea-fffc2a260bd9 · outbound
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Decreased
Reference 2012
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a4d841cc-d0f7-44db-9149-201bb63d1284 · inbound
Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.