Pith. sign in

Paper Citation Record · LEDGER

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression

As of 12 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 1 inbound Pith citation observation for arXiv:2412.04317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04317 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:37:08.457373Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T15:40:05.730181Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:16:15.637500Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b8a476c1-9b94-4994-8048-469c0fea17c3 · outbound

This paper cites GPT-4 Technical Report.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.227561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.227561Z digest=sha256:22c059372aad8a5b52c56d2cb83336cd52575f044ddf92e48bd88845e1dc5709

Observation 630a6798-b520-4f35-bebb-085cdcd1846f · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Bottom-up and top-down attention for image captioning and visual question answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.231500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.231500Z digest=sha256:48a3cf199fe0e5759502fe95fed517e0c8913d16f6c1f383bc2a394f9a8e50f2

Observation 7153172c-ea37-4214-8c96-4f76675d4715 · outbound

This paper cites Qwen Technical Report.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.234380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.234380Z digest=sha256:acb222f168b4598a39b19b22fdf05f3a9f60e8bde3702ac6c51a241ce276a7db

Observation 1bbea82d-0bf0-4f1b-98f7-05fb89259d53 · outbound

This paper cites Gemma: Introducing new state-of-the-art open models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Gemma: Introducing new state-of-the-art open models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.237958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.237958Z digest=sha256:58b4f5b1e8c7b91465b3d7849cfc05c3b2bd579b437d00f2cc61be24937efb9a

Observation 48487edc-c0c6-4e0c-9400-a6d7a23f3fca · outbound

This paper cites Language models are few-shot learners.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.240691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.240691Z digest=sha256:3bfce833f0a399d41ed34c1e621848049ead48394c461b1bf68f5c5e967e175f

Observation 997d8738-1bd1-4fd5-9852-7e9055b7c5f3 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Honeybee: Locality-enhanced projector for multimodal llm

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.243430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.243430Z digest=sha256:48212d285963234f71c37adf8e540de6a5f5ee6bc7959b836e2989a4a09b7f7b

Observation c5ce3a98-1b80-4e69-8f83-431aa4261138 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.246133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.246133Z digest=sha256:c27f3de0126492dc254888cebe3ba1116717c9f1cdd662d9b1130feac7847425

Observation a2f7a125-3085-4c86-8b39-27d816b7fb04 · outbound

This paper cites Lawrence Zit- nick.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Lawrence Zit- nick

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.249780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.249780Z digest=sha256:44879cffb94ca76a61fe32d8eb62e25d487c3402fada68c5580822f9ab2fa368

Observation 25bd50ef-cdc9-4fcf-95db-4c4bdd09e788 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.252944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.252944Z digest=sha256:cd2b57f5f1592d1d163e196e199e216ea4275be0c5494242e05bb60ef6e9813a

Observation 4da9f6f9-2f64-4a7e-bc1e-72e72a807aee · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open- source suites, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression How far are we to gpt-4v? closing the gap to commercial multimodal models with open- source suites, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.256084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.256084Z digest=sha256:3a7258d28d7d9a80d2b524b1d86d9bdcc0ed86d013891edfc05dacae8b7c01f9

Observation 5dffd154-db90-4c29-a3ad-c7e60474eb54 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.258655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.258655Z digest=sha256:9b24a6d9080fbed67da039454b04b42299d2135ae6e775c791c41417123872a2

Observation 2cfabdcd-b567-4d75-9956-05574f714485 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.261489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.261489Z digest=sha256:55f61d6f57fe9256e68d7164bda8038842c2766af8494cabac1d3c07e35e1745

Observation 9544fdc1-b620-420f-9a3f-199c7d42de2d · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.264453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.264453Z digest=sha256:7559e3797b4d8681b2439ff09a1fad4bd3f526103a1762dff5498db63ea7b006

Observation 402e63c1-ce06-472d-b65b-62a8131ba07d · outbound

This paper cites Mme: A compre- 9 hensive evaluation benchmark for multimodal large language models, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mme: A compre- 9 hensive evaluation benchmark for multimodal large language models, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:09.007536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.267203Z digest=sha256:ae006f00d9c7843cae818f4608d58e715ce75d12891970106ec75882e2a8e9ea

Observation 2e3c46bf-2041-457c-a0a3-71607bc6ddfc · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:09.000272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.269857Z digest=sha256:ff9222fe538c316a96bf860a24a7ffbc8d0c406695e1b35ac419e7de0950b026

Observation 527398e0-0bf9-48a6-a377-b88a7d565e53 · outbound

This paper cites Minicpm: Un- veiling the potential of small language models with scalable training strategies, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Minicpm: Un- veiling the potential of small language models with scalable training strategies, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.992917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.272901Z digest=sha256:41393a469189bc8d307a0414f715b5b0097aa91a24bb15edf895cee5b023e605

Observation eab121c9-2d36-4c35-a170-f3587b65e6eb · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.275372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.275372Z digest=sha256:58a0b299631483f76e1b2783e2eea1d16f9ad464fcf42d6637b3fc20491b76b9

Observation 6f0c000c-631d-492c-97a2-2c3ead4dec59 · outbound

This paper cites Token merging for training- free semantic binding in text-to-image synthesis, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Token merging for training- free semantic binding in text-to-image synthesis, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.985333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.278335Z digest=sha256:cefb11b52dfc5fa2d33d00838750f8d96511c78c72704d67e4ded63b7389c729

Observation dc1890a4-bb9d-4e1a-bc26-0cbe64f298a1 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.977404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.280814Z digest=sha256:2789aa6a8470d7531e8a2319b296c5fd119c48051ac4782e2c9e55f05e36a40b

Observation e149fc0b-b43f-464a-b106-f780aadeb0cb · outbound

This paper cites Phi-2: The surprising power of small language models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Phi-2: The surprising power of small language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.969681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.283336Z digest=sha256:c2ccae7a148caed1301b45100c9187593979bef79d75e202bb86409ce35b5f19

Observation 5200a819-4edd-424e-bdc8-2efbd06283bb · outbound

This paper cites In defense of grid features for visual question answering.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression In defense of grid features for visual question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.962265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.285778Z digest=sha256:caf0b51e0862107ba1e0bc012153285a7fc3592e5214d065fd9e2e7f288b90f9

Observation 2f69eca5-7024-437d-b957-f3d20f9fb4be · outbound

This paper cites Contrast and classify: Training robust vqa models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Contrast and classify: Training robust vqa models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.955006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.288247Z digest=sha256:396cdcc759d43fcd81c0ffe8113e5d6c7d20396fe4c49d4c0d0a15e90aa3e93e

Observation 295cde2f-b1d4-433a-9d22-a4aa1aad5b48 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.947365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.290817Z digest=sha256:3c4c1557309680b7f9cdd1ad5be82ca6c0dcb9c31998a20a6c14d4dc9706dbed

Observation 14175937-1d25-4def-9bdb-51590fdf9c75 · outbound

This paper cites A diagram is worth a dozen images.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression A diagram is worth a dozen images

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.293615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.293615Z digest=sha256:b12b4d30065237340b23516e5b2b1a0ab22ad82f39d174b17e02044b4eb892dd

Observation 19186a00-6f9d-426d-a3ce-03d6c2c877e0 · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Seed-bench: Bench- marking multimodal large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.936399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.296242Z digest=sha256:6b1dbf759004db743ecc7320eee6691032b5e5521c32ddab449590262b7d2c25

Observation 6fb39a02-0e10-4c36-befe-608ae9bd53d9 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.299006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.299006Z digest=sha256:f29bb670d186a4e81d4de3a89a3bdd4bb7f196b8a10953aba8da56556178a862

Observation f294ee1a-9caf-4144-95ce-b22068d88a12 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.301741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.301741Z digest=sha256:71546cb8c3dc014096fd67e1f74e85d52c496c11faf4faeda8cc58ea8e4bdf17

Observation d3339ae5-77eb-42b9-97b5-e22a2112aad6 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.924912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.304194Z digest=sha256:255aaea1816b1c916acdf54cb1a7ca75161025816aac8a02354dd8578cc96dde

Observation 64483dff-5527-4c2b-ba48-8dfb8a2a7b10 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.306619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.306619Z digest=sha256:895be6987aab85c7f4ea1f189d2cf1d16f3de11775a84fef99295fe473a88734

Observation d0e2b9f2-4dee-4358-9153-54d800348272 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Evaluating object hallucination in large vision-language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.917574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.309221Z digest=sha256:ec236b98986300ea456e086543afefae3d8e12e2903bb0d42c9501ee3ca55a55

Observation be65e829-2862-4674-a628-60aeec5efeb0 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.311714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.311714Z digest=sha256:8d035feeedc9c45861c90296246ea63cde6297234d2c4057327c1a9c7bc4723e

Observation 6ec3cd88-490d-46c8-b9b7-102a66d57b8f · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.909721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.314473Z digest=sha256:b67f61f52b5c73d8a29c65cf69ff5aaaba0ff97070921a5e691dcd1cff4513e2

Observation 1a3aeb31-8b46-455b-aa92-cc18366b8b84 · outbound

This paper cites Improved baselines with visual instruction tuning.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Improved baselines with visual instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.902388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.317072Z digest=sha256:a594dcac5dc2719555d5d3fa5c7a5ef9633c6c3895421a704a1fb98b503189a5

Observation e5c29df2-6b30-4368-b898-6e8415d2ca4f · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.894198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.319771Z digest=sha256:afb24f12d3ca48539568bde23a9013f1c7ff93c9383f9daded87863b23c3b8ad

Observation b77bb695-f88d-4844-aa9e-2d08401f8d9f · outbound

This paper cites Visual instruction tuning.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Visual instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.885784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.322118Z digest=sha256:f712815ecc5f373e2522cf8d043506f6f21a7b901464446eb8e0756b4947a2a5

Observation 8e6abf79-d0d5-4d55-a27f-6f3591ac9396 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mmbench: Is your multi-modal model an all-around player?, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.877386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.324392Z digest=sha256:93a2953c04f9798c5b23f72da767c1a9859155f59ee00d53948ee488fc7bbd85

Observation 4effb140-fbb8-4d2c-943c-6b91af32192d · outbound

This paper cites Decoupled Weight Decay Regularization.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Decoupled Weight Decay Regularization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.326784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.326784Z digest=sha256:1697fe07e94c9a3eed86c8fda3a16d64d40efd99c6b38b86c81807a3e06b8c96

Observation 26d39eb5-e572-43be-ad1b-f553186cd52e · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.329705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.329705Z digest=sha256:53097c56538f5cfbe72e6c01413f7000c75acd1da9d68a243fe69ec6a9644a6e

Observation ab43a5b8-737a-4264-a66a-835219ebd03b · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.332592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.332592Z digest=sha256:4ab77dfa82b74a198b769f63cf487ef6d5815121fc46ee15d2ecfe5673ebd0d0

Observation 72c5abab-e79f-4606-9c23-25d0eb12da07 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.335366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.335366Z digest=sha256:d4d3d35051d1da817f08068f803c92d00c08267868475ae2de7c945537a96dd6

Observation 7327ee58-e4d1-4ef8-ac14-950aa62dff77 · outbound

This paper cites To- wards lightweight transformer via group-wise transforma- tion for vision-and-language tasks.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression To- wards lightweight transformer via group-wise transforma- tion for vision-and-language tasks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.865032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.338155Z digest=sha256:ed945f11c76ecc34dc37c45e74310274348e38b0a5951ea715609bbe222a6062

Observation f496d380-de84-48cd-9559-242a5d46327a · outbound

This paper cites Cheap and quick: Efficient vision- language instruction tuning for large language models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Cheap and quick: Efficient vision- language instruction tuning for large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.857341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.341165Z digest=sha256:5d9a405e38ea1c92ff5aeb50e486032854d2ad78edc7e15d11442f090fe704f7

Observation f99125b1-42d1-4b03-b293-14644c5b97a5 · outbound

This paper cites Moil: Momentum imita- tion learning for efficient vision-language adaptation.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Moil: Momentum imita- tion learning for efficient vision-language adaptation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.343738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.343738Z digest=sha256:9c7b86593dfb81c16da128d57cfd4537069ec3bd81161c67b23ddb971d41384b

Observation c9fbdc75-f9fc-4890-b09e-5eb04d28568a · outbound

This paper cites Towards language-guided visual recog- nition via dynamic convolutions.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Towards language-guided visual recog- nition via dynamic convolutions

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.845990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.346149Z digest=sha256:b8968d8d6e65d709b94f5d6b91c6c1e5e7d0d88da776dd4b291064427c998298

Observation a0e638c8-4f5a-4ed1-ba4a-15faa67b0baf · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.348453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.348453Z digest=sha256:15c4f9fa03208c2c29db3219e1eea7a75f8f7b12a5a20a2ba715f63d6ee5c842

Observation 8e27a5ca-7e74-46f2-ae78-61ad377c1b7f · outbound

This paper cites Chartqa: A benchmark for question answer- ing about charts with visual and logical reasoning, 2022.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Chartqa: A benchmark for question answer- ing about charts with visual and logical reasoning, 2022

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.838006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.351277Z digest=sha256:098718079ce8a155816820807170a6a841042a606f40c932c0d77ccc89a25447

Observation b94b76ec-e7bd-4786-92b9-4028b6f44007 · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.829927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.353618Z digest=sha256:87cf9fbf2208324e0afc21b4bbb169b4de4a91d53f0c7c47777ee6b62144d508

Observation 0cc2e6e6-aab0-40db-b992-db6c012216f2 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression DINOv2: Learning Robust Visual Features without Supervision

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.356012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.356012Z digest=sha256:200ead9d72f354cebb18724d6bf1faf8b38fe12b23a130eb41569e1e066bc496

Observation 208567de-d3f6-45da-9e74-a99002bd76a9 · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Learning transferable visual models from natural language supervision, 2021

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.358592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.358592Z digest=sha256:df2ebae95a36a6cadee0104e7689f94348e7be0070e1582ef281725f9173896f

Observation f54f0555-8992-441c-8355-7a9e7a0297b2 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Learning transferable visual models from natural language supervi- sion

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.361676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.361676Z digest=sha256:4ecbc19a51a83d4a1748cfeb155897bf59bea4896532d02062d1cfb721e66419

Observation f09e7d82-b271-4736-aea6-4599f8fcc571 · outbound

This paper cites Imp: Highly Capable Large Multimodal Models for Mobile Devices.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Imp: Highly Capable Large Multimodal Models for Mobile Devices

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.364232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.364232Z digest=sha256:ebe5b59b75d1ed94964a402eee1c46b08a406b4000b22c9b04172b9b183dd42c

Observation 873ba2fe-e190-4711-860f-6bf6827ed107 · outbound

This paper cites When do we not need larger vision models?, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression When do we not need larger vision models?, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.812922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.367027Z digest=sha256:2ea50507e2e27195ede63da8522f9068ee7889893c28269db6a24fff4ba84b1b

Observation 16054fb2-0a0c-4f1b-8662-43a7cd5815a4 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.369438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.369438Z digest=sha256:2474a3d96ed90a684e1db5327159d30b5a414a90c5b775bb28af685780503429

Observation 60bbf9ad-f9a3-4482-b029-b3826a42ba27 · outbound

This paper cites Towards vqa models that can read.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Towards vqa models that can read

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.805615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.372029Z digest=sha256:7f331972fdfe48a8a276e9cd27df1035fe44dfa3f5b9b39e353a66fcc2683b46

Observation 32e1e7cc-c75c-48c7-bf92-4d50e1954a75 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Gemma: Open Models Based on Gemini Research and Technology

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.374548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.374548Z digest=sha256:78c161fc393926412b920da41ef673a1349014d7852eb31c840e92f376f2f25c

Observation 92d80c96-c396-4a7f-8abc-a8b4742a2691 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.377247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.377247Z digest=sha256:fa0e43b3def8326d2e348cdfd330780f2adc78e37eb857e79a6d7c7b482401e8

Observation ef2dddbc-3ab3-413d-bbb2-889b9be9d0df · outbound

This paper cites Well-Read Students Learn Better: On the Importance of Pre-training Compact Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.379783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.379783Z digest=sha256:b8c73719670c8e6eb1f729489338eabcc3351eea5a44d58f7e7447cd280adf0a

Observation a4b708de-7375-4531-97d7-9e73da4f9c3e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.382822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.382822Z digest=sha256:f2fedd4acf24faf57a7125e4c7471538f5c592e27d29a8815d0c86e18a69ae0f

Observation f271e797-ccf5-484f-8716-5a434963d2a9 · outbound

This paper cites Show, Attend and Tell: Neural Image Caption Generation with Visual Attention.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Show, Attend and Tell: Neural Image Caption Generation with Visual Attention

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.385543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.385543Z digest=sha256:723a6413b498b84a19b7602c886e01b3616654e662fc341348b0e3464f4ef890

Observation df808884-647b-4e16-bd72-646e0931afae · outbound

This paper cites Stacked attention networks for image question answering.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Stacked attention networks for image question answering

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.798043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.388241Z digest=sha256:59bc28d8fc223c9660efa68d5ec0f230aac898472a3e274258ed7e40bf6f5638

Observation 2bf59c01-1ee7-41e3-ad7f-a02e7a9590f1 · outbound

This paper cites mplug-owl: Modularization empowers large language mod- els with multimodality, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression mplug-owl: Modularization empowers large language mod- els with multimodality, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.790234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.390580Z digest=sha256:cf02363596caa5bcbf9168853dc7053f64ea81dbd293b608304fe8339fd92264

Observation 55a8b314-8652-4f49-8c5e-a1fc978e722e · outbound

This paper cites Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.393098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.393098Z digest=sha256:dbbe46449cbefa66c11be65477b61a9b84a17cab754238ee3b0ed3f32ce0f23b

Observation 92f936f9-cd1b-4c9a-9078-69b019d9c7ef · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities, 2023.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mm-vet: Evaluating large multimodal models for integrated capabilities, 2023

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.782501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.395868Z digest=sha256:c597d743c671f68fa209573ce98c30d492956ddaa3702775362800c7038f8ba5

Observation a8ab1367-712e-4e3e-ab80-7da9fa0e8e85 · outbound

This paper cites Deep modular co-attention networks for visual question an- swering.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Deep modular co-attention networks for visual question an- swering

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.775215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.398429Z digest=sha256:e664baf43d340e331fb225bb135c198e65f78bec128a3db3886977b501a9858f

Observation f72e0beb-c44f-48c2-96de-368aa8556536 · outbound

This paper cites Tinygpt-v: Efficient multimodal large lan- guage model via small backbones, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Tinygpt-v: Efficient multimodal large lan- guage model via small backbones, 2024

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.767930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.401027Z digest=sha256:cd9381f3c4526b5c62f54951cdc522785d7f97da294016c740c0a42bb97f735d

Observation caa8085a-2a3c-4a56-a9cc-ff0e030ee692 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.760040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.403823Z digest=sha256:d303c482f4f7a797c35ebccb92d6865f64d8adb62c6cd5be5d305d9251978fc6

Observation b2f69a39-5f6d-4fe1-94ea-6661d0005158 · outbound

This paper cites Sigmoid loss for language image pre-training,.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Sigmoid loss for language image pre-training,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.407076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.407076Z digest=sha256:f7c1f3c928b3d94c85b88ecb08232f94b6e2612f5c5616648e72069b42854dbe

Observation acf37657-d8f1-4fec-a07d-c07e8a932943 · outbound

This paper cites MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.409547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.409547Z digest=sha256:11bac782234f02cafac2da34257ca973085ada28de2a8f9c14340dece6073e83

Observation b0f53947-2e19-437d-93ad-4e44cb06403e · outbound

This paper cites Lmms- eval: Reality check on the evaluation of large multimodal models, 2024.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Lmms- eval: Reality check on the evaluation of large multimodal models, 2024

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.412907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.412907Z digest=sha256:f58de68debd9153940bb6e40b4d79ef161677648ff00080bed85cd66240a7d5a

Observation 5c07de0f-ffba-412b-b950-8ba9bbd313f2 · outbound

This paper cites Vinvl: Revisiting visual representations in vision-language models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Vinvl: Revisiting visual representations in vision-language models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.415267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.415267Z digest=sha256:09f5cc653ac44985b85db0481873723fae68490fb38acad44a7827f6f16cbd05

Observation 4536e021-a980-4ee2-8a99-b48b30bccffe · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression OPT: Open Pre-trained Transformer Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.417720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.417720Z digest=sha256:3c0c35d2b854c20144c090021bffb85a65c437c5b0587e4bbf07b67c90e28ffa

Observation 4dd6a1d4-33c7-4ec0-94cd-52a422597aae · outbound

This paper cites Free vqa models from knowledge iner- tia by pairwise inconformity learning.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Free vqa models from knowledge iner- tia by pairwise inconformity learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.740215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.420493Z digest=sha256:66282c9c695b10a8ae933980b40de04b6cbb02b4ec845120909731becf89e872

Observation 39a79b3a-3f8d-4519-9f6c-dfb5d8b3ebd6 · outbound

This paper cites Trar: Routing the attention spans in transformer for visual question answering.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Trar: Routing the attention spans in transformer for visual question answering

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.732876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.422900Z digest=sha256:0dc2f327b6344be07d6ae6937c213899c82172aa0f8d79647fd04d63309aa603

Observation 863d176b-39ce-4965-b82c-bc35082c741d · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.725697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.425431Z digest=sha256:e2f81ca2eb5a9a75db86a8c8da6e2585273491a48c9ac3743e0bab4f6428e36d

Observation 6d31a72e-f5cf-467c-bda0-e859e0ea21ac · outbound

This paper cites Mipha: A Comprehensive Overhaul of Multimodal Assistant with Small Language Models.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Mipha: A Comprehensive Overhaul of Multimodal Assistant with Small Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:08.427822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:37:08.427822Z digest=sha256:7431b6d5131d03456d2b75446833e682f794a4a20a0a3f43479221627dcb749e

Observation 7a1cb42e-0145-42bc-93d2-d71b52a7757a · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.718218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.430767Z digest=sha256:dab17116631d0e5a2f239a32593cfad66616b4e3caf8cc1e35fbdd0f25faab20

Observation 044164c0-8c90-40f7-8768-cc4f0c7ebb24 · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.711080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.433171Z digest=sha256:40308a088841b15fa561477dec5c2e03f66c61b6d4c6219a6e399028344be88c

Observation 9803d311-a481-4c98-9916-a256a7ef60a0 · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.702909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.435461Z digest=sha256:5501d73f632423a9818ddfa1cddd6188328c6c4a89f444752d712dcd538622c5

Observation b8c945af-388a-4435-8980-7a542ac4e4a4 · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.695244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.437867Z digest=sha256:5b202a307a601b69ca849ac2a5b96b360bc3c08cca8a05543ba5af9a13f05fbb

Observation e7484725-fc2b-4fdf-8647-f81f95fc8b3c · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.686740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.440180Z digest=sha256:a2b76c30184922bbe5de54ddf7c5e4373e0b46ea31bd74f4813b7a3076659a16

Observation 5d5ce4cc-e1d8-4f57-97bb-d4ec7709e76a · outbound

This paper cites Therefore, the value of the square in the figure is 2.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Therefore, the value of the square in the figure is 2

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.679023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.442544Z digest=sha256:dfb6ea488f1215f4264dc0e49d615874cf194adfaa5d495a40d648aa47d21626

Observation 823a836a-6e37-43b4-bf38-f2db8912ba61 · outbound

This paper cites control center.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression control center

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.670988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.445034Z digest=sha256:31dc6883a55ea876e9d9b4eb668c66d95f8a5cb9aeb0e7077f23ba6d5b35c1a0

Observation 3ab4ddba-1331-4c01-9e1b-b3940d8881e0 · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.655902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.450057Z digest=sha256:5b62ed04b664bfb805026881dfcbf780a8127fd0038159dbb57d133e74d4d4bb

Observation ad67f5a0-d0b2-4b9b-95da-b78400b582a0 · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.648609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.452450Z digest=sha256:59215fccc094159d38941277baa78449960ece2f4ae523937b9d95e4f0379cec

Observation 275209cb-2bf7-4178-aac4-6d48a670f5a5 · outbound

This paper cites an unresolved cited work.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:37:08.641489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.454943Z digest=sha256:0bfbe222d5ad52a20b5aa19f9b2569f2738103cafd5238309381d00055aa0bba

Observation 9c492150-ab38-4202-8025-e9e1cf4e7473 · outbound

This paper cites RIGHT",.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression RIGHT",

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.633594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.457373Z digest=sha256:46c0345bcb3b520226ddacc6ed0e74d720283bc60a46032a7e831a7b23dba615

Observation 6c5fd70e-35c4-484c-aaea-fffc2a260bd9 · outbound

This paper cites Decreased.

FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression Decreased

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:37:08.663257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T21:37:08.447590Z digest=sha256:b9d9be9878c7eda18425545b70e3e8df7290957177c014a98186cf05345a853f

Pith citing papers

Observation a4d841cc-d0f7-44db-9149-201bb63d1284 · inbound

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning cites this paper.

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:16:15.639326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T15:40:05.730181Z digest=sha256:e409879f61024d48f3807fb5950b473ff09fae6f87775db0fac4299e888de6ab