Pith. sign in

Paper Citation Record · LEDGER

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning

As of 18 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 2 inbound Pith citation observations for arXiv:2505.11945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11945 v2

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:48:16.136813Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T07:14:08.479610Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:08:21.980388Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f62878cd-7277-4c8d-a45e-65f364e40591 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.822537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.621753Z digest=sha256:fa44cd38be6447b1d3974b8a2b3c16d17c66891b1a1801b2747d073aca32188d

Observation 2f9599b4-91f2-4f73-a3f4-4b4f9829b513 · outbound

This paper cites Vita: An efficient video-to-text algorithm using vlm for rag-based video analysis system.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vita: An efficient video-to-text algorithm using vlm for rag-based video analysis system

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.804808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.632476Z digest=sha256:3e359d7906c8351dfdeb93f2b0321a0a3cb494d25f5d2a8a86cd655f81458e22

Observation 0861b875-b5a0-4ac6-8897-3dfa2c7dd57c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.639409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.639409Z digest=sha256:650cc65de76e61b4601e5ef4a7e38e2d97f12c669a118a16afd9c03e0e3bf990

Observation 9b1e6f6c-b87d-48a1-8d1b-2ff46a275d0c · outbound

This paper cites Honeybee: Locality- enhanced projector for multimodal llm.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Honeybee: Locality- enhanced projector for multimodal llm

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.788084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.652965Z digest=sha256:96ea24eb41d76af4721ed021c73dd48f84a2a5b32ab6c18980dac80ec05fa678

Observation 5e5dd523-61d0-4079-9fc0-db871464aa34 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.771427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.662831Z digest=sha256:aeffbcfca31c63ef47cd9b083c773aacd2d65ba9adf68ff359a620ebce0799d1

Observation f9dee378-b90f-4403-a288-3f01828d9a88 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.674264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.674264Z digest=sha256:190c451809addfc8dc9ae6c184943aac2bddb0979decfb3e02db11c16a9713c0

Observation 6d2fa2d2-90f0-4c95-bf16-5145408b2b86 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.681307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.681307Z digest=sha256:c4305fc08728353158c0e4c940dccba081f8389224fba51ea906284f4a96c53e

Observation 3d3af01d-f0fb-4eab-bff0-c6108479f7a8 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.687604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.687604Z digest=sha256:afdc7954810a2e03328dfa29e66d7c39ee436174d16e8d94bff7a22d77c2038b

Observation 77f1ecc6-94c2-4125-a27b-15255b2174bf · outbound

This paper cites Don’t look twice: Faster video transformers with run-length tokenization.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Don’t look twice: Faster video transformers with run-length tokenization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.717428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.694497Z digest=sha256:ce0c0c6e0121be9da56a2f6f61699a249ffccd056c161cb0aef99bca04dd59ea

Observation d8f9a9cd-0d68-487f-b0a9-7502e8ef8f87 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.699140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.699140Z digest=sha256:c6a92e7a9f5d147b6126c7e2b9fcfc75b55d312c3397809a7d34d8f34a95699a

Observation 31b2899d-2898-49e0-aba3-f4e768f28024 · outbound

This paper cites Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.701206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.705318Z digest=sha256:2428ea18ddbf04da8d030eae01ebed4376b38948a314ffd3faa07c5ee71518ae

Observation 0df83c2a-7098-47dc-8e16-ab180eeb1650 · outbound

This paper cites TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.713164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.713164Z digest=sha256:2fac12b7315b3afc4bc145cb2b2fd1794a36178aaf5f355a97c7f2edd3c667f2

Observation 12e9f6d2-3058-4da3-91a6-df0fd278f016 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.683415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.718760Z digest=sha256:8630f13b6444a10c89519ea3486c6c0db17e8c4543af173064c0fe57174fd002

Observation 16cdcd33-faf0-496f-96e7-ee146fc3d012 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.724986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.724986Z digest=sha256:8c51ec25a6ca1810b39637b896f86a8f293908e8c0fb99fb5d54072005e01c20

Observation 40832420-5bc7-414a-96e0-626514df2af2 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Efficiently modeling long sequences with structured state spaces

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.666969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.732650Z digest=sha256:e7a1d6c42e6e11a890e7db3f3eef12fcbfc0935b3a0c7d90e4819ecb7c390bd9

Observation 98a84017-9be8-4346-9658-a016bab80d0a · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.647982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.739852Z digest=sha256:8ed8c81c6b18d814b9d3bcbd638b411a36209c046fe32ab9ec52f1478506e377

Observation 58d24812-16da-4f2f-8774-ff66ed7f39a4 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vizwiz grand challenge: Answering visual questions from blind people

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.631083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.746571Z digest=sha256:d87808cf9b08790d34480b61c9cf1824f5e3ec440bc2326eb3a95f9fd9b344f6

Observation 0cf33c95-b05d-45aa-a327-2103b752530d · outbound

This paper cites Bliva: A simple multimodal llm for better handling of text-rich visual questions.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Bliva: A simple multimodal llm for better handling of text-rich visual questions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.608997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.752539Z digest=sha256:2303a46dc4bc8345e936e1fdc2cfff99880939650b1f1bb52885eaf7a2657f7f

Observation 055fb796-142a-4db1-9318-de1e4de9a968 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.759340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.759340Z digest=sha256:b08b764bc9921bfd0b682acd899d9c2f8905867a2eb2be8c1e5dd5283d731a36

Observation d026d179-89cc-4830-9ce7-8dd29369ae1e · outbound

This paper cites Token compensator: Altering inference cost of vision transformer without re-tuning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Token compensator: Altering inference cost of vision transformer without re-tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.580654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.767540Z digest=sha256:aa610a869ad991846e114ae603e5880e23f8994ce3a48806a5b3b01979c048e8

Observation cee4edba-c37a-4e81-92ca-93dbacb28d72 · outbound

This paper cites Logicad: Explainable anomaly detection via vlm-based text feature extraction.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Logicad: Explainable anomaly detection via vlm-based text feature extraction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.561972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.773640Z digest=sha256:1a5e0b81ce41de79c7289b44dfc8b4ded93edcd37065342207e92a881669ed2f

Observation 56fef084-41e6-4492-88a2-c0fbfcbcb10a · outbound

This paper cites Llms meet vlms: Boost open vocabulary object detection with fine-grained descriptors.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llms meet vlms: Boost open vocabulary object detection with fine-grained descriptors

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.545350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.780005Z digest=sha256:67c828d6646f1a88d293c92ba5ea89ae2073a9810afde3614b56d3b93cb27a4f

Observation c870c746-a0ec-401c-8f2e-ef8df16d13c7 · outbound

This paper cites Mm-reasoner: A multi- modal knowledge-aware framework for knowledge-based visual question answering.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mm-reasoner: A multi- modal knowledge-aware framework for knowledge-based visual question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.528772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.786893Z digest=sha256:480c1f934c30b0b4d324920ab298c3dccc8926622b2d78e01c2fb3a4213640f2

Observation 40462433-8ad0-4bf1-a6ce-6d052dd65f80 · outbound

This paper cites Vlm-pl: Advanced pseudo labeling approach for class incremental object detection via vision-language model.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vlm-pl: Advanced pseudo labeling approach for class incremental object detection via vision-language model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.510182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.795802Z digest=sha256:e3ce58c7409ca530c6869da796dc0dc2583eb388a280c5611b8bb3e4046a71ad

Observation 48131032-88e0-498d-b4d1-dba3fbf890d3 · outbound

This paper cites Lookupvit: Compressing visual information to a limited number of tokens.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Lookupvit: Compressing visual information to a limited number of tokens

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.492474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.801917Z digest=sha256:e0c481d1ff395de07a6df677ba23f33bf2d82e7508b4a766ffec7cf482082952

Observation 16f1ee93-5146-4761-b480-e14b0c4e315a · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.808159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.808159Z digest=sha256:6180328ef3ff2a905cd30a457793d4cc9cbcc47383e92d2dc929b9ee7e0a8262

Observation 04c32cae-143e-4ffc-9aee-35d91ae85dbb · outbound

This paper cites Ez-hoi: Vlm adaptation via guided prompt learning for zero-shot hoi detection.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Ez-hoi: Vlm adaptation via guided prompt learning for zero-shot hoi detection

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.463194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.814482Z digest=sha256:5d39189100611673cde066e2642ea384683710c980f636b6539bba881481b19a

Observation 3601f0f1-856f-4326-8837-e95d082e1208 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.820689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.820689Z digest=sha256:e06ad64b121ee31491092124a5da0fa8397ef8a643d4db0a843601c005a7154d

Observation d080d537-2f39-4dff-b5bb-e43875d8d4e6 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.829103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.829103Z digest=sha256:07fbf225fb54b13593cc9800bb9f36f8334b744b93a35b422f0e27c33309d05f

Observation 832cf5ed-d8b9-492e-921a-8ff379890722 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.444494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.836127Z digest=sha256:d5915f9c9022d82a6b71e0252dfe3eff2b61aecd1aba341a24b3a7e647ae3de9

Observation f1754f53-f7dd-44a1-b20f-3ddb0d75739e · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.407845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.842487Z digest=sha256:10010b1ce440a8607597949fcdf974116cd6670667603e874747d6b64be7eeaa

Observation 82cf2356-dfab-4ab6-ac15-a58e88dd345e · outbound

This paper cites Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.850545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.850545Z digest=sha256:124a91ec5c4ed71de7249fafcf5013fbfa6025ff825f3731817a0abfedc8092c

Observation 8b7a0790-66f6-455b-a4f4-4dedf1e16902 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.856988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.856988Z digest=sha256:c6e9ab83b195a1744e92000537d26be4ef3bb6d75517ff3f9ae12c32a6b6a309

Observation bb0d0afe-d458-4f1f-8384-fc77cc1ea76f · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llama-vid: An image is worth 2 tokens in large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.322034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.863990Z digest=sha256:b9598a77019f1482208447e2b304b40a922075b4c0924309dd4da9f18fb80270

Observation 204187f5-f6e1-456a-8ead-79c436535559 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.872063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.872063Z digest=sha256:41d4e779f64882424565e1314921a2ec40ab60f516e7110d9a9059ba9d3b3049

Observation e10ef480-4df8-47ca-afc9-4380f5a04b49 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Evaluating object hallucination in large vision-language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.304591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.879546Z digest=sha256:b68327821c993ccef5b66ee7f7c53adb5a116abccb25a17396a8f596533d296d

Observation f4aab4f0-8be9-4c5d-aee7-460d243a8e03 · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Monkey: Image resolution and text label are important things for large multi-modal models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.288374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.885567Z digest=sha256:ac033a74c9303f981b0e9545e7aa8606506ef104a1ee24b4f129be00e9e4cb96

Observation c32f603e-a5f6-491e-9037-dc2f2dfe2e2a · outbound

This paper cites Vila: On pre-training for visual language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vila: On pre-training for visual language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.272213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.892662Z digest=sha256:bdc03d722980f5e7f9b30d9a2c06bb00bfe242b5c07ca2e07e9df413c88a6c9b

Observation 12e92281-0c0f-4df5-b138-a885811ccef5 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.899314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.899314Z digest=sha256:4fb023d91c9d04a00cabf7ebb708e6493294e7cb4854d93cc643e935fcce85ed

Observation 2e34a6ef-d5ad-49b4-9c7e-9fbcb4e6a523 · outbound

This paper cites Improved baselines with visual instruction tuning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Improved baselines with visual instruction tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.907057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.907057Z digest=sha256:0aa3a7fbc21a0212a92f836a33a6bc0767a20dcb9914490889fa52fc0e6635cf

Observation b195c883-cf1e-4be7-8117-1de4ed7fd9f6 · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge, 2024.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llavanext: Improved reasoning, ocr, and world knowledge, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.912723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.912723Z digest=sha256:4e65d4d51ca741fa23d0ae879c906ae9e5e20d8522dcbf1bc339dd36ee849667

Observation b99e44de-561b-47f8-bb64-419825bc6dc4 · outbound

This paper cites Visual instruction tuning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Visual instruction tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.917778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.917778Z digest=sha256:99dbcb7f29d2c456930365285a2bea3023402fe9287dfef217831c5ae7958bc4

Observation 2af17887-9523-4eef-99ba-52009acf4f1a · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.924724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.924724Z digest=sha256:39edd896cfdedc9ae4706a57c25c9533c880d0efb175c77b5b7a2be78e2c7bd8

Observation 53b9be14-69f3-4160-b451-211b1d53e57a · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233, 2024.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.932776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.932776Z digest=sha256:a0571c788bb72d4bd927dfd011b6d99a6ec7c2891a01d0669873af9e6b7ab18e

Observation f529d4dc-f324-43c3-9a9b-4adefaa2e53f · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.938326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.938326Z digest=sha256:3bdbe5e36630b48f74eeb0e9cf36a56a27fc2acb4f500bdd73ec894ce690d2dd

Observation 7d473a4a-7a10-416e-9a8a-d5879019c480 · outbound

This paper cites Questioning, answering, and captioning for zero-shot detailed image caption.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Questioning, answering, and captioning for zero-shot detailed image caption

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.203605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.944171Z digest=sha256:f379e8810fbef051d3638406bc1d1e08a9b11911e3a3cd257c98ab00a70a9d23

Observation 7ed2d803-7c77-41b0-a3a6-b31c97e84256 · outbound

This paper cites Does vlm classification benefit from llm description semantics? In AAAI, pages 5973–5981, 2025.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Does vlm classification benefit from llm description semantics? In AAAI, pages 5973–5981, 2025

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.183406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.950924Z digest=sha256:85b31a10d8208c452d7c94fd0ccfea76fb9c35efce265092400036a472f10b10

Observation 6b057137-b19c-434d-9e52-c74a4c018f8b · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.956976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.956976Z digest=sha256:fde810a3832f7fcf8c95848d010c0229ef97bff801cb4789629d83141f1f57af

Observation 33740278-dc71-4c6b-a779-fae282c58dc3 · outbound

This paper cites Infographicvqa.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Infographicvqa

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.162471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.962777Z digest=sha256:49f26032346bba9a0517be508a2f74ed05f84c639491dce621641414e73db17a

Observation 44c17a7c-f79c-437d-89f4-eb8960cce2f0 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Docvqa: A dataset for vqa on document images

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.142650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.968333Z digest=sha256:e7e09ca4f400cee190c2264b9662a5e55620ee6ac76bc8b6d9377bcde3df2833

Observation 404e1c35-cf10-4e80-b998-2b129e6b152d · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Ocr-vqa: Visual question answering by reading text in images

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.106857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.974801Z digest=sha256:f7a635f23d755f5b813673992e37c9a7c6abd609f08ec3d7a032ac1b76b7fd48

Observation a17bc5c5-59a7-4933-b169-1d602e105c20 · outbound

This paper cites X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.981135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.981135Z digest=sha256:77a7e766611970b492d4fdef6c6915b622ec330092b3b1327a1adc8b358ec083

Observation 4c172126-ab99-41b1-96c4-0fb76ffaa218 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Learning transferable visual models from natural language supervision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.988182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.988182Z digest=sha256:cf14e8a593bab4e8276aee4c7328fd70923f690d5c2c0194698d9baeb45e3a86

Observation 524c67da-e1ee-419d-8dca-9bef68068679 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.996657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.996657Z digest=sha256:ba4070fcd3b8338503fb28e4bd248db2c59464a47111c399bdfcd6b7f14e2503

Observation 2f748f79-0d26-435b-8d19-8eed9c027152 · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Textcaps: a dataset for image captioning with reading comprehension

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.005761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.005761Z digest=sha256:d8a34d6decf57d54c010fa97b8ed18fc7b317c2cc62dddf6277cad77b1548149

Observation d8c46163-dada-4f41-a7ff-6e278ad53d2a · outbound

This paper cites Towards vqa models that can read.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Towards vqa models that can read

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.014992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.014992Z digest=sha256:ac7a22aa0cdd074e843da7bff1f1999baa195be7b130c9126c6770805b537b3a

Observation cbc2e32c-88a0-4a0e-9659-ace1bc2b6246 · outbound

This paper cites Less is more: A simple yet effective token reduction method for efficient multi-modal llms.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Less is more: A simple yet effective token reduction method for efficient multi-modal llms

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.053996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.022616Z digest=sha256:9e9705e3960d01ce48453aafacd9cbc93dc5fb3630af41d1475210e870b2f7d4

Observation 77bb4be7-8de6-4b4c-9f86-11a44302791e · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.030070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.030070Z digest=sha256:1f22431ae8e9bc829ae5572787f56cc7bd7f36fc6f43e0d1470b2dbd7ac80a12

Observation e2c41836-a8eb-4564-bef9-77486e1a450e · outbound

This paper cites FastVLM: Efficient Vision Encoding for Vision Language Models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning FastVLM: Efficient Vision Encoding for Vision Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.039140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.039140Z digest=sha256:6166defa7aad3b73c6b7e77970afe46cb86fd12bdcf5864706a600aba523fba6

Observation 1d4f5561-1963-4097-8a5c-13ca50a36610 · outbound

This paper cites Marvelovd: Marrying object recognition and vision-language models for robust open-vocabulary object detection.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Marvelovd: Marrying object recognition and vision-language models for robust open-vocabulary object detection

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.023069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.050031Z digest=sha256:ff05fc9829bf77735fc58bf4c45c9006e801c52ae4cd57f19e36e12ca9cd9dc5

Observation 2a1795cc-d234-44ba-9259-e91f96f51404 · outbound

This paper cites Fashionvqa: A domain-specific visual question answering system.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Fashionvqa: A domain-specific visual question answering system

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.002072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.059553Z digest=sha256:92e1cfacd8b9e517bcdd1a5c2d383746868b5e439885eaebb9dd03735c01d6eb

Observation ae314ba2-7d11-4713-9bc4-f20418aa3413 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Cogvlm: Visual expert for pretrained language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.069768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.069768Z digest=sha256:4a173c5a75a0968dab0b9f237120e50938c4d3bc202defedd7cdcd7a528ea6f6

Observation e9ba78d9-d7b9-4fdc-9969-9a291cf2c333 · outbound

This paper cites Rl-vlm-f: reinforcement learning from vision language foundation model feedback.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Rl-vlm-f: reinforcement learning from vision language foundation model feedback

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.974689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.075698Z digest=sha256:baf8aeaf97a333b84496555e0053d57f4c1516b1152c9311d711b60a952e5588

Observation f8610864-5872-44be-8ce8-29b904cb20d5 · outbound

This paper cites Vary: Scaling up the vision vocabulary for large vision- language model.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vary: Scaling up the vision vocabulary for large vision- language model

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.956137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.081522Z digest=sha256:b84bf051a5ca604476831490ea48ee2d472e9cc1f949676fb386520580ee43a9

Observation 31167906-34d2-4242-b1e7-f9d258976af0 · outbound

This paper cites PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.088460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.088460Z digest=sha256:81066f9c36e31c89934bd0f761c9b655ca333d9af1102a78a89e06b58a9dc530

Observation 19fe73d8-5170-463b-81d9-29179927abc9 · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Visionzip: Longer is better but not necessary in vision language models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.094655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.094655Z digest=sha256:e87529313b4bd91391e00c5c67bf0b8c22699fa1cdb4297ee4ed59705c104a05

Observation 3dce44ee-3c80-4486-8fa8-017cc21a1f7b · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.099677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.099677Z digest=sha256:f378e967d7a9e2752405e0d2c07042493f7d7ab1319af730419fec33844d937b

Observation 39af3fc9-e525-42ce-b7e7-76b20d873fbd · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.930470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.104787Z digest=sha256:18e19ddd20eed7a0ff6df53284f4b5d8f4c3f98d251d1e217641ee93757eb8a6

Observation c5d5a9a8-76c3-467a-ba48-ed8a6b1fbd4f · outbound

This paper cites Good at captioning bad at counting: Benchmarking gpt-4v on earth observation data.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Good at captioning bad at counting: Benchmarking gpt-4v on earth observation data

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.911023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.110320Z digest=sha256:5efdeacf4583e32c3022e651c372e87b47f9e7e2277df257c83ba82d7b22fd92

Observation 5b262117-94e1-4ebc-a208-b54d27729888 · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.116057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.116057Z digest=sha256:3314c63c291783a22100a9d3f638e35ff0ecf046b600ba512599ba461948a2fe

Observation 0e4b296a-43fb-4a93-82f4-ad1d4a8dabd1 · outbound

This paper cites Llava-mini: Efficient image and video large multimodal models with one vision token.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llava-mini: Efficient image and video large multimodal models with one vision token

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.889049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.122911Z digest=sha256:612ac92d1c363948b312c5e6d8ab96ce98c1a4f57d39c5c4162f0b6469a34def

Observation ec64c60e-2bcd-4618-af58-83360ab3a170 · outbound

This paper cites Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.128579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.128579Z digest=sha256:e7d62e03b3a96c73c3c6c351377690c31228fe0957241adc91245e0414d90661

Observation bc6ac840-ab06-4b33-b0cd-7d4836c0c254 · outbound

This paper cites FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.136813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.136813Z digest=sha256:c0bc1afb849bc66c1849b8e314610b3c181ef4868989ed7f92bff838de2c31dd

Pith citing papers

Observation bbac5bfd-94e9-46e3-a2ec-9a76cecb0132 · inbound

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs cites this paper.

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.131202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T03:18:11.993413Z digest=sha256:7fe10a7460aaedbb7317db96b00cc9bea325dd8510a4a93a69069085553694c4

Observation 0403b79d-6273-4590-8527-cf80380382d0 · inbound

The Hidden Power of Scaling Factor in LoRA Optimization cites this paper.

The Hidden Power of Scaling Factor in LoRA Optimization Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:08:21.981648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T07:14:08.479610Z digest=sha256:7aeb6b8310b89fd6519ed8162dffd5f45af6f11b1c7a053871473690032bd4b0