Pith. sign in

Paper Citation Record · LEDGER

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning

As of 18 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 2 inbound Pith citation observations for arXiv:2505.11945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11945 v2

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:48:16.136813Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T07:14:08.479610Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:08:21.980388Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f62878cd-7277-4c8d-a45e-65f364e40591 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.822537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.621753Z digest=sha256:5abdb2ca669f2f7e85b1108e87834d02b11e065dd68090d9ee5a54ed21e612e5

Observation 2f9599b4-91f2-4f73-a3f4-4b4f9829b513 · outbound

This paper cites Vita: An efficient video-to-text algorithm using vlm for rag-based video analysis system.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vita: An efficient video-to-text algorithm using vlm for rag-based video analysis system

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.804808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.632476Z digest=sha256:cbf7602eefd29a7d1fa340e557880be8ab41cb31c9cf80cae66854c43b7e43ff

Observation 0861b875-b5a0-4ac6-8897-3dfa2c7dd57c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.639409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.639409Z digest=sha256:86371aa9bcb158a5e358f2991474633fd1bde5307d9608b83fb1297af1b11eda

Observation 9b1e6f6c-b87d-48a1-8d1b-2ff46a275d0c · outbound

This paper cites Honeybee: Locality- enhanced projector for multimodal llm.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Honeybee: Locality- enhanced projector for multimodal llm

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.788084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.652965Z digest=sha256:d85bdf84249400ce048c669bc70a15fa87b66b8908f5b066815a141a81dfa6d0

Observation 5e5dd523-61d0-4079-9fc0-db871464aa34 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.771427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.662831Z digest=sha256:7867f2b7783994c720135c116827fd1b3812e266fa7d34fdab602f9062ac697b

Observation f9dee378-b90f-4403-a288-3f01828d9a88 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.674264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.674264Z digest=sha256:61460f6dd119bcddcd02a7672960570041ccef47a8e7d8d3643aab35ca4806c7

Observation 6d2fa2d2-90f0-4c95-bf16-5145408b2b86 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.681307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.681307Z digest=sha256:b3f3e7f5c7c948f7ac228ddf582ff90fbd6893543320779569e32d4a6ee4c9d1

Observation 3d3af01d-f0fb-4eab-bff0-c6108479f7a8 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.687604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.687604Z digest=sha256:382042e7d53546e7e98d78b944119c8e764f8594027c538f92b1e272c7d6c105

Observation 77f1ecc6-94c2-4125-a27b-15255b2174bf · outbound

This paper cites Don’t look twice: Faster video transformers with run-length tokenization.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Don’t look twice: Faster video transformers with run-length tokenization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.717428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.694497Z digest=sha256:1821c6d7733c0cea99ce30b34e04288fd4de3a34f658557f89e328661b84e59f

Observation d8f9a9cd-0d68-487f-b0a9-7502e8ef8f87 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.699140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.699140Z digest=sha256:2debb934d1a86123721a49ff97308cbf616e1e61ea4f43420c864206b4188345

Observation 31b2899d-2898-49e0-aba3-f4e768f28024 · outbound

This paper cites Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.701206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.705318Z digest=sha256:e5b16b29dbd20fdc4b97e15f13b044ad9a324b6259c33c1e132fc0feebe7cab3

Observation 0df83c2a-7098-47dc-8e16-ab180eeb1650 · outbound

This paper cites TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning TC-LLaVA: Rethinking the Transfer from Image to Video Understanding with Temporal Considerations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.713164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.713164Z digest=sha256:c9fd853549c80b326a880b996c301fe1227c9de199dae777d886be8a4c832f37

Observation 12e9f6d2-3058-4da3-91a6-df0fd278f016 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.683415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.718760Z digest=sha256:2eeac68671ed3581aec3364b9357764c5979dd6456f107f334620a8b504ae837

Observation 16cdcd33-faf0-496f-96e7-ee146fc3d012 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.724986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.724986Z digest=sha256:00be2feda4b6529a1a8c5f51f2ff180f47cf2f631889ba79fd640ed1f0121426

Observation 40832420-5bc7-414a-96e0-626514df2af2 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Efficiently modeling long sequences with structured state spaces

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.666969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.732650Z digest=sha256:04b272eb7eb6ef41fe8222e653bdad3747d59878e71f8e1d2b5d1985c09c3220

Observation 98a84017-9be8-4346-9658-a016bab80d0a · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.647982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.739852Z digest=sha256:4aa224ba2a1b121b0d6b64974557477e6c69146b3ac94b984a1b6c29bf94aa35

Observation 58d24812-16da-4f2f-8774-ff66ed7f39a4 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vizwiz grand challenge: Answering visual questions from blind people

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.631083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.746571Z digest=sha256:0fe1c7275e7aa6f6ff79a810962a199f343bb15791fd5c68e1c2287aeefc984b

Observation 0cf33c95-b05d-45aa-a327-2103b752530d · outbound

This paper cites Bliva: A simple multimodal llm for better handling of text-rich visual questions.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Bliva: A simple multimodal llm for better handling of text-rich visual questions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.608997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.752539Z digest=sha256:3791e0fb43c1cf44dc3b7c2432a2abab7ed18eaf85c0486e79cd1774f19b7a06

Observation 055fb796-142a-4db1-9318-de1e4de9a968 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.759340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.759340Z digest=sha256:1016a5cb5e5cead2e874743c5161f39ec37cee0c84bcbf6f971a33bdc6a8d733

Observation d026d179-89cc-4830-9ce7-8dd29369ae1e · outbound

This paper cites Token compensator: Altering inference cost of vision transformer without re-tuning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Token compensator: Altering inference cost of vision transformer without re-tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.580654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.767540Z digest=sha256:2c412d890ab25de924a6f856de9190e2bdef4634ffe0b8ddfe9aad1b1a7d0e04

Observation cee4edba-c37a-4e81-92ca-93dbacb28d72 · outbound

This paper cites Logicad: Explainable anomaly detection via vlm-based text feature extraction.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Logicad: Explainable anomaly detection via vlm-based text feature extraction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.561972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.773640Z digest=sha256:90bae8d046103d44eb4741e21f3a6a9ceb6395d9350265bf8e45aadce521c9d5

Observation 56fef084-41e6-4492-88a2-c0fbfcbcb10a · outbound

This paper cites Llms meet vlms: Boost open vocabulary object detection with fine-grained descriptors.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llms meet vlms: Boost open vocabulary object detection with fine-grained descriptors

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.545350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.780005Z digest=sha256:d43a0699b343ea353e4cfc421a586284dc370d6d4dfaeea1a666a89c74262faf

Observation c870c746-a0ec-401c-8f2e-ef8df16d13c7 · outbound

This paper cites Mm-reasoner: A multi- modal knowledge-aware framework for knowledge-based visual question answering.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mm-reasoner: A multi- modal knowledge-aware framework for knowledge-based visual question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.528772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.786893Z digest=sha256:229a764b170cfbb89acb7ad35b30b3368abeb283040f6d15579f4886c6d738c3

Observation 40462433-8ad0-4bf1-a6ce-6d052dd65f80 · outbound

This paper cites Vlm-pl: Advanced pseudo labeling approach for class incremental object detection via vision-language model.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vlm-pl: Advanced pseudo labeling approach for class incremental object detection via vision-language model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.510182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.795802Z digest=sha256:965dbf3cc7339200a88aa00477e03732120adbb7bb33c30b8bd4443a935b789b

Observation 48131032-88e0-498d-b4d1-dba3fbf890d3 · outbound

This paper cites Lookupvit: Compressing visual information to a limited number of tokens.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Lookupvit: Compressing visual information to a limited number of tokens

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.492474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.801917Z digest=sha256:eb50532ccf0759c7311587256840535dc8d091fb14a1da49d81ef4e61763f2cb

Observation 16f1ee93-5146-4761-b480-e14b0c4e315a · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.808159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.808159Z digest=sha256:17c8da7ef6c3dcaee5c542efc1eda26083e4aed4df5f896f5cc6d220110f64ed

Observation 04c32cae-143e-4ffc-9aee-35d91ae85dbb · outbound

This paper cites Ez-hoi: Vlm adaptation via guided prompt learning for zero-shot hoi detection.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Ez-hoi: Vlm adaptation via guided prompt learning for zero-shot hoi detection

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.463194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.814482Z digest=sha256:fd25c25ae89faef6926201fdeb3b0be4ae7ba3d7fceafe96e447e497609d6653

Observation 3601f0f1-856f-4326-8837-e95d082e1208 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.820689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.820689Z digest=sha256:63d500556ca3c7a9f1c323594a36e098530061d967f1469ed44b442e0ed7a125

Observation d080d537-2f39-4dff-b5bb-e43875d8d4e6 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.829103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.829103Z digest=sha256:4d73bf0f152a9a98d7058616de566f8bc28190032ff8b7fd71ae36a3dcf56aed

Observation 832cf5ed-d8b9-492e-921a-8ff379890722 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.444494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.836127Z digest=sha256:7097872995fbd27b7a5085718fce8e3d5a36681e4fc40d48bcd4f0f53f468081

Observation f1754f53-f7dd-44a1-b20f-3ddb0d75739e · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.407845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.842487Z digest=sha256:df700a88050f7151a4d75149b4d1dea35def1cb6a1a26fd7d933b722eccc55ca

Observation 82cf2356-dfab-4ab6-ac15-a58e88dd345e · outbound

This paper cites Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.850545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.850545Z digest=sha256:1790127a35ab251a1cd3191a7b47129944190b2bedb5203aa941900dd3d86cbc

Observation 8b7a0790-66f6-455b-a4f4-4dedf1e16902 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.856988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.856988Z digest=sha256:e8dcdad8ab39bc71cf78b7666cd45d4932935c5dea79d4d01832f4514d7ce2a8

Observation bb0d0afe-d458-4f1f-8384-fc77cc1ea76f · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llama-vid: An image is worth 2 tokens in large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.322034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.863990Z digest=sha256:3ca4af6524a0af0fced52ecaa4cf10c17943b126939731694841b0d1768e764b

Observation 204187f5-f6e1-456a-8ead-79c436535559 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.872063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.872063Z digest=sha256:dafee3592539e7970e72bf82b276f8f685fb4f9349258db0ed2f2c2520598bc6

Observation e10ef480-4df8-47ca-afc9-4380f5a04b49 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Evaluating object hallucination in large vision-language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.304591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.879546Z digest=sha256:ecaef453ba918534cc8abe785e9279dd9f7875cb8bbb799d6b3a09a0ecccaac5

Observation f4aab4f0-8be9-4c5d-aee7-460d243a8e03 · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Monkey: Image resolution and text label are important things for large multi-modal models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.288374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.885567Z digest=sha256:03fe2880fcdde17eabb363b48d78e58c698b070eeb5ce41f6f0d358487ca4690

Observation c32f603e-a5f6-491e-9037-dc2f2dfe2e2a · outbound

This paper cites Vila: On pre-training for visual language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vila: On pre-training for visual language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.272213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.892662Z digest=sha256:1b5c54b1182a31371aeafa49e3c71831d96d4064b61d1889c7d6d7745adbb95f

Observation 12e92281-0c0f-4df5-b138-a885811ccef5 · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.899314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.899314Z digest=sha256:564caf1bc451c033ff0e1fa987ea2569bee295b8318597063a4e2947cb998c26

Observation 2e34a6ef-d5ad-49b4-9c7e-9fbcb4e6a523 · outbound

This paper cites Improved baselines with visual instruction tuning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Improved baselines with visual instruction tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.907057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.907057Z digest=sha256:73d2b500d194ed5e5c7bea3d75567ea90defce845835b9f69cb79edb4d0074e3

Observation b195c883-cf1e-4be7-8117-1de4ed7fd9f6 · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge, 2024.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llavanext: Improved reasoning, ocr, and world knowledge, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.912723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.912723Z digest=sha256:31c3382ec0e0c98d9c4c8a888223769afea77e9686cff82bc8ae6ea16fb9e283

Observation b99e44de-561b-47f8-bb64-419825bc6dc4 · outbound

This paper cites Visual instruction tuning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Visual instruction tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.917778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.917778Z digest=sha256:4ea79a8a5ca491288a20c1053be9caf019ba2843e0ed8624c6bacc594e3e3be4

Observation 2af17887-9523-4eef-99ba-52009acf4f1a · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.924724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.924724Z digest=sha256:00007caa0b089b19007a4c1c82196a7e0013540d9ab490b11855265953efb1b0

Observation 53b9be14-69f3-4160-b451-211b1d53e57a · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233, 2024.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.932776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.932776Z digest=sha256:2dc7a0c56567cfe0a5607e3c4c108661c6f8f0f078f49a9058855fb29ff320cb

Observation f529d4dc-f324-43c3-9a9b-4adefaa2e53f · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.938326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.938326Z digest=sha256:9d33fd4c152e43832e705f347b97865944c75f07afec95dbe7a1dbd581700ce3

Observation 7d473a4a-7a10-416e-9a8a-d5879019c480 · outbound

This paper cites Questioning, answering, and captioning for zero-shot detailed image caption.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Questioning, answering, and captioning for zero-shot detailed image caption

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.203605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.944171Z digest=sha256:ed0e3d0366d4a5e6f8d78427bde796a5ba58b6679b8cd3a68ce57176578d15ad

Observation 7ed2d803-7c77-41b0-a3a6-b31c97e84256 · outbound

This paper cites Does vlm classification benefit from llm description semantics? In AAAI, pages 5973–5981, 2025.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Does vlm classification benefit from llm description semantics? In AAAI, pages 5973–5981, 2025

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.183406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.950924Z digest=sha256:7fcc04cfb466cfb1ff00b94d6dcf098bc1e8424d3aad6934e45012dc4df168b8

Observation 6b057137-b19c-434d-9e52-c74a4c018f8b · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.956976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.956976Z digest=sha256:97dc77463908cefeee15140be7c70248b809842065067454ae908a139b637c87

Observation 33740278-dc71-4c6b-a779-fae282c58dc3 · outbound

This paper cites Infographicvqa.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Infographicvqa

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.162471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.962777Z digest=sha256:707f16340fa8547d99feaaa866d2595f793d3040579dcfefac0b37495aaa1b5b

Observation 44c17a7c-f79c-437d-89f4-eb8960cce2f0 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Docvqa: A dataset for vqa on document images

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.142650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.968333Z digest=sha256:4d4ea61d75b735682dd7ed389f8a8a17c3c3b1f3984d36db9a942b58c6701720

Observation 404e1c35-cf10-4e80-b998-2b129e6b152d · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Ocr-vqa: Visual question answering by reading text in images

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.106857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:15.974801Z digest=sha256:6eb72fd817060ca072ae95b5748bde1ed94961826bb82215da59dfe62009401c

Observation a17bc5c5-59a7-4933-b169-1d602e105c20 · outbound

This paper cites X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.981135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.981135Z digest=sha256:2fe0c71931bb6d01d1c49116735aaea039b57a0cabfade167c777715f600d32e

Observation 4c172126-ab99-41b1-96c4-0fb76ffaa218 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Learning transferable visual models from natural language supervision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.988182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.988182Z digest=sha256:0c7144226f474fe52b70ed18f1e472339a87a815258bce452bdc2dbf648129a7

Observation 524c67da-e1ee-419d-8dca-9bef68068679 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:15.996657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:15.996657Z digest=sha256:2e4c69d196dfe1de0c08d50f22bfc1063f00e45237f75c0555cee78145e31af2

Observation 2f748f79-0d26-435b-8d19-8eed9c027152 · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Textcaps: a dataset for image captioning with reading comprehension

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.005761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.005761Z digest=sha256:74d12ca153e25c24b75c341c51deec81046e713b830ea6f3bee37e192a84efdf

Observation d8c46163-dada-4f41-a7ff-6e278ad53d2a · outbound

This paper cites Towards vqa models that can read.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Towards vqa models that can read

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.014992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.014992Z digest=sha256:c03af3e36f5099879692fe11af5c89a36375bb27309f3194551e8ed08f5f6f4d

Observation cbc2e32c-88a0-4a0e-9659-ace1bc2b6246 · outbound

This paper cites Less is more: A simple yet effective token reduction method for efficient multi-modal llms.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Less is more: A simple yet effective token reduction method for efficient multi-modal llms

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.053996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.022616Z digest=sha256:c5c473e6733c9fc4a8c5f03f93634103801d488c2cf01baf85d97fcbe54229ba

Observation 77bb4be7-8de6-4b4c-9f86-11a44302791e · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.030070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.030070Z digest=sha256:a23d330d665e92ca1765b7a3fbe11048f99fc4355ac8742e8d35d62a94e1364e

Observation e2c41836-a8eb-4564-bef9-77486e1a450e · outbound

This paper cites FastVLM: Efficient Vision Encoding for Vision Language Models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning FastVLM: Efficient Vision Encoding for Vision Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.039140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.039140Z digest=sha256:dbb3fc9dc13174e963a4fb9d9789d062fc709aea6409ff93e512f2537e934dd9

Observation 1d4f5561-1963-4097-8a5c-13ca50a36610 · outbound

This paper cites Marvelovd: Marrying object recognition and vision-language models for robust open-vocabulary object detection.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Marvelovd: Marrying object recognition and vision-language models for robust open-vocabulary object detection

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.023069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.050031Z digest=sha256:8d3106f4654dc146e5a48ef4b3b363003cff91a4ed027eaf0292b785a7c13362

Observation 2a1795cc-d234-44ba-9259-e91f96f51404 · outbound

This paper cites Fashionvqa: A domain-specific visual question answering system.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Fashionvqa: A domain-specific visual question answering system

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:17.002072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.059553Z digest=sha256:a4d4b520a6c476c37c59d75ee5dceab414e44eea1e925f3595c620dd3a50839d

Observation ae314ba2-7d11-4713-9bc4-f20418aa3413 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Cogvlm: Visual expert for pretrained language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.069768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.069768Z digest=sha256:07f97db686280ba1be2d87021ac737d95f41f63eb353f1cc87b662a87b8e1855

Observation e9ba78d9-d7b9-4fdc-9969-9a291cf2c333 · outbound

This paper cites Rl-vlm-f: reinforcement learning from vision language foundation model feedback.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Rl-vlm-f: reinforcement learning from vision language foundation model feedback

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.974689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.075698Z digest=sha256:a50bad6459274c264200a91967f770afad773604b04d2437b90328a5badc4f9b

Observation f8610864-5872-44be-8ce8-29b904cb20d5 · outbound

This paper cites Vary: Scaling up the vision vocabulary for large vision- language model.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Vary: Scaling up the vision vocabulary for large vision- language model

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.956137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.081522Z digest=sha256:9532a8cb371f284322a23be1b4ebe7c09d27da74b6a9e4431503ac3a095e7e00

Observation 31167906-34d2-4242-b1e7-f9d258976af0 · outbound

This paper cites PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.088460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.088460Z digest=sha256:9e3ee8fdd8ca0bd6c91ba5f059b464f3146bed22dac5793080972645c9bad195

Observation 19fe73d8-5170-463b-81d9-29179927abc9 · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Visionzip: Longer is better but not necessary in vision language models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.094655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.094655Z digest=sha256:bf9dc57b4cf54e45d0ac038d80f0815f2290cb1a13fab1b16afd0716c225af3a

Observation 3dce44ee-3c80-4486-8fa8-017cc21a1f7b · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.099677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.099677Z digest=sha256:a2b1083ad0f0bbafe07aa4682134f8d0d3747cc206cb03d9cc1a6b66db1ef3ac

Observation 39af3fc9-e525-42ce-b7e7-76b20d873fbd · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.930470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.104787Z digest=sha256:84aac745961ca0a53baede1b3ebd2efa3f9a5e26051b87ad49aa462433ac46ad

Observation c5d5a9a8-76c3-467a-ba48-ed8a6b1fbd4f · outbound

This paper cites Good at captioning bad at counting: Benchmarking gpt-4v on earth observation data.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Good at captioning bad at counting: Benchmarking gpt-4v on earth observation data

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.911023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.110320Z digest=sha256:53406344770c58ccde4a6553a55244fd7a42e3480e174d1c52163f7a5b568eb3

Observation 5b262117-94e1-4ebc-a208-b54d27729888 · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.116057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.116057Z digest=sha256:2dcd1365b0d61c52305f638ec291abb43c76454409ddec18f557ffd94877d667

Observation 0e4b296a-43fb-4a93-82f4-ad1d4a8dabd1 · outbound

This paper cites Llava-mini: Efficient image and video large multimodal models with one vision token.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Llava-mini: Efficient image and video large multimodal models with one vision token

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:48:16.889049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:48:16.122911Z digest=sha256:8741c999170daad9c1637b983cec76643436b8094d2f8a762f224dd59bdfa5a5

Observation ec64c60e-2bcd-4618-af58-83360ab3a170 · outbound

This paper cites Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.128579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.128579Z digest=sha256:440a21035562db1c32635b39b48c0cea85475d20499c715ca69a02c10b3ff374

Observation bc6ac840-ab06-4b33-b0cd-7d4836c0c254 · outbound

This paper cites FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression.

Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:16.136813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:16.136813Z digest=sha256:8b77cdd1830723add449c97bfba55a28717092cba6c94eb7ccbe2258fb2a4d11

Pith citing papers

Observation bbac5bfd-94e9-46e3-a2ec-9a76cecb0132 · inbound

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs cites this paper.

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.131202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T03:18:11.993413Z digest=sha256:ec1f8dcd86ab2122e77a510696ac9c4eb38d6013c4579b36b0e3723ec6f57578

Observation 0403b79d-6273-4590-8527-cf80380382d0 · inbound

The Hidden Power of Scaling Factor in LoRA Optimization cites this paper.

The Hidden Power of Scaling Factor in LoRA Optimization Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:08:21.981648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T07:14:08.479610Z digest=sha256:3a6d048dd54112d3f648350248f4932eb5a96c3234357145970465064a32e9f4