Pith. sign in

Paper Citation Record · LEDGER

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts

As of 13 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 0 inbound Pith citation observations for arXiv:2411.13909.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13909 v2

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:51:51.871655Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

96 of 96 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved65
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 25e15dd4-fb5e-48d0-a5ec-8b91e0174180 · outbound

This paper cites Fuyu-8b: A multimodal architecture for ai agents.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Fuyu-8b: A multimodal architecture for ai agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.115101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.115101Z digest=sha256:1f7e0e97376d286f526e3b14a53b2db5fb22141de2123674e911eb3fa2d6efc1

Observation 2dfca88f-5ed2-48a6-909d-ccbd65904bf5 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.125826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.125826Z digest=sha256:8de0e4e1012234b6c7be8b2e77a55d203f3ee00b15ffcd83958cb3c1c93070e2

Observation 8be1ce38-cf7a-4685-921f-6b439e7b402a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.133619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.133619Z digest=sha256:748ee5744beefd4e44dce2440eed67a01bbd82a743e135dab8647f08fff2e872

Observation 16ea1d5e-ef55-4af3-ae88-158a8d85d017 · outbound

This paper cites Token Merging: Your ViT But Faster.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Token Merging: Your ViT But Faster

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.149219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.149219Z digest=sha256:197c9893c59e463a6c478b36702f1fe732d6a8d0e9665d1bdc523314b8299da8

Observation 01be99ee-480d-4502-abda-68e0d3eed005 · outbound

This paper cites Language Models are Few-Shot Learners.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.162624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.162624Z digest=sha256:79f44c49de60e111c6c34c3527fac8f81b1bef20e4aaa26cabfbc5623b07eb90

Observation d86304a0-35c0-4903-8796-e548f8ea350e · outbound

This paper cites Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.175846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.175846Z digest=sha256:d63480c86c20f347b3d88dc11e18edcb4dcd36f65acc50db109b16d49f5b9bea

Observation 7ea5e044-5a17-44a0-8c83-4e88c9eb8af6 · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.189481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.189481Z digest=sha256:f498bd763843cfdf2c45772692faa5ebaab95e9eb7795d75d5b840c792aff7ea

Observation 1d9af478-0dec-4101-8916-7bb93b3584eb · outbound

This paper cites Towards unifying medical vision-and-language pre-training via soft prompts, 2023.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Towards unifying medical vision-and-language pre-training via soft prompts, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.196581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.196581Z digest=sha256:c84ded551e56b14c395df171a60485d9632d04806873bdfa2d198a75dcf3d346

Observation 6cd7bae4-de4b-450e-9772-b8ebfb2a07c5 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Gonzalez, Ion Stoica, and Eric P

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.203642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.203642Z digest=sha256:41cf9f1cf1c6bbb0c062180b2a0f13a038fbb3aa4f04f203833c5c8fd89ce343

Observation 96b12ac5-9880-4e23-a978-7e31e4b2bda4 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.208742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.208742Z digest=sha256:6c8fce6afcdc8093c11eacb43f5baa968e81f676876cd60913003a8447a6a5c5

Observation d27889d3-50b6-45b6-ac7f-944f46061dd6 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.214972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.214972Z digest=sha256:b6a07318a497ad5885f66f962723287dce90cbdc2e9aa9c18c328ebe432fa6e3

Observation a0482907-b9f7-4175-9def-39199e954b99 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.221314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.221314Z digest=sha256:0607816a576c235fa4d9bcace8f890f7dac31c8b02caabdf81bf2a6d2d45c5af

Observation 52e81a30-a017-4365-be17-ac2af99121f1 · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.226824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.226824Z digest=sha256:c06bc73db4e88a4df45c5fde0d734fcc1cb973b6e610acf42535aa0fbc992b4b

Observation aed7bdbc-02b3-4e40-9283-6eee2677b121 · outbound

This paper cites The llama 3 herd of models, 2024.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts The llama 3 herd of models, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.233360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.233360Z digest=sha256:4d04f0c44886d27de982771125e2723c9c563fa5cefc2cbb689d117dc3b4b909

Observation 56ba8c92-f660-4a0f-89ee-b9ceb40c2164 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.240282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.240282Z digest=sha256:b74d0e712fe22485702fd4929c7b99580f6c57942f014b14f301cac5478c94f3

Observation ebf3e11b-eea7-4ca5-af0e-8dc4246c0da3 · outbound

This paper cites Domain adaptation via prompt learning.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Domain adaptation via prompt learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.251837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.251837Z digest=sha256:f372cbd29cfc09ca0351c9af30368406f6e30eb9c6b36c2a97fa177c22e1d9e9

Observation 78bc37ef-d666-4c01-a95c-01f59493b87b · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.260941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.260941Z digest=sha256:5c6b96fabf81b1186a73baeb828fa93c823d2c0bd68eca1fad2b1eb6214b0062

Observation a876b31c-c966-47f7-84f4-ebae66d9cb7e · outbound

This paper cites Openllama: An open reproduc- tion of llama, 2023.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Openllama: An open reproduc- tion of llama, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.267323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.267323Z digest=sha256:88b6e4590c4a66a51fe2d6f0c06873b2e7f83287d0054ca27de5806401392989

Observation 21eaa993-e0b6-4e4f-9de9-885d1ab33ca1 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.273091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.273091Z digest=sha256:e259db7bf9bf9fc7ba35cb707c3eee134af671d5122cacc72abfa51f030b0884

Observation 132734f6-3fda-4f3f-a0d1-198453b5212a · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Vizwiz grand challenge: Answering visual questions from blind people

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.285067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.285067Z digest=sha256:66cc0e02c8c78e66f31ad57599a59aaf175c279693acc5ca9e7d5d65e54e9f38

Observation 02b4bf6f-2fbb-47a7-97bd-7807a5d0dc9c · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.290677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.290677Z digest=sha256:697d2ccbf49e9b97952010b11b24db00109887562b53516290bb9122a0f28ca4

Observation ec195a06-b6df-4ebe-8992-51b49c28db46 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.297822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.297822Z digest=sha256:3731bb0077832c954860d0857299f81863f632fe7c5413bdba24f79b172d8dca

Observation 7ebef389-4205-4652-b6ca-8ca8bf1ca476 · outbound

This paper cites Vi- sual prompt tuning.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Vi- sual prompt tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.309409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.309409Z digest=sha256:87b5966e13c517693537ab4a786510a46ce43e9641c1cb4d0472df76ca28853a

Observation 81822f56-44ba-49f5-ab02-8fd91dfd962a · outbound

This paper cites an unresolved cited work.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.314393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.314393Z digest=sha256:09fa60f6926ae89d2f21547917e48fb5fe5e84935374f6583d5e17b2b95ff28a

Observation 7e954659-da04-4796-a11b-f9b7f15e460b · outbound

This paper cites From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.320880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.320880Z digest=sha256:d9532caa82d87bcd204d7ab8f88a885ed7e7d33ecb47b9b884fe3de0a4657974

Observation 3a9dfaba-2753-403c-9d9d-6fe224fdb951 · outbound

This paper cites Maple: Multi-modal prompt learning.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Maple: Multi-modal prompt learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.331412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.331412Z digest=sha256:bf13745af5d0dc7178f41fe79e6d22fd21963b379c3c4721822b7d3859a5776c

Observation 45f50271-4c28-4ee0-9f1d-159ba8212e0e · outbound

This paper cites Segment Anything.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Segment Anything

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.340574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.340574Z digest=sha256:d5aed3faa1549fb84857843bae28cb38f261eaee62f3783788e31d182ddedb99

Observation fff57757-6d4d-4a60-864b-924505ee6c53 · outbound

This paper cites Spvit: Enabling faster vision transformers via soft token pruning.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Spvit: Enabling faster vision transformers via soft token pruning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.346933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.346933Z digest=sha256:dbdc1d4ca580c0d35216766fb303ade0dacd2e0953105e39a1eb94d034ca5e24

Observation 6d29c858-c89c-41a0-a18d-8d6ab9b73095 · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts The power of scale for parameter-efficient prompt tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.353239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.353239Z digest=sha256:3b1869dd3c12fffa78febd59123c56cef225194b21445e2afff45ba1f9e62956

Observation e683a8cf-a4cf-4f82-b70a-9b1446688624 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.360879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.360879Z digest=sha256:cc3c4fa461d8d7639d73ce5f60f7a9464441c1d9526829901a23679aeae65738

Observation f3e2f510-242d-4244-ac39-afc68dc33698 · outbound

This paper cites Task-specific fine-tuning via variational information bottle- neck for weakly-supervised pathology whole slide image classification.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Task-specific fine-tuning via variational information bottle- neck for weakly-supervised pathology whole slide image classification

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:54.055635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.368671Z digest=sha256:50ce809d198082fcd385a774b89ad9c37477c99161e7e9d35844151fcc292854

Observation db50e5e4-7f4e-4c1e-ba3d-543fec037ddb · outbound

This paper cites Rethinking transformer for long contextual histopathology whole slide image analysis, 2024.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Rethinking transformer for long contextual histopathology whole slide image analysis, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:54.027824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.378182Z digest=sha256:db972636e098b259bf86109e5ebaf0c229daf9d6252ef0ba624b195635a9f3c0

Observation 896436c2-0078-4fcb-943f-bde17fb8f031 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.388053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.388053Z digest=sha256:dc35fc9d77ae023c31ad1d88aee349c752e793d2005697875c960433936934c3

Observation b0e6e73a-87ec-464d-81cf-0d72812d3c92 · outbound

This paper cites Tokenpacker: Efficient visual projector for multimodal llm, 2024.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Tokenpacker: Efficient visual projector for multimodal llm, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:54.000809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.394057Z digest=sha256:da3097dbb89cb8eff4d41240ac60e4bf7058b005e15051ed84bb151b9beca585

Observation 30fb327a-11ae-4523-ba7a-daa3d2ab37fa · outbound

This paper cites Prefix-tuning: Optimizing continuous prompts for generation.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Prefix-tuning: Optimizing continuous prompts for generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.967541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.399734Z digest=sha256:b0583b836186e6040ead06ee7fc628e4de78601d710bacc6f33f0e02b01efb6a

Observation e4302872-b085-425b-a8ed-fa7e91fb5fa5 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Evaluating Object Hallucination in Large Vision-Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.405396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.405396Z digest=sha256:79356496e77bb70eeeed8c79b0aa90b67f834b393bc6599623ca7af9bc6857ba

Observation 46b672a0-b583-4600-941d-361400c3489b · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.938850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.411371Z digest=sha256:c3bf3d37890b8cf0ab9031ca23f729a12697e9f96a7370e43175618dee3b8578

Observation 9545a833-4f95-49d7-8cb3-26e0a38b59da · outbound

This paper cites Not all patches are what you need: Expediting vision transformers via token reorganiza- tions.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Not all patches are what you need: Expediting vision transformers via token reorganiza- tions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.910850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.416544Z digest=sha256:c80b58d82df45ef0052e9d858a1cd36fca678fd866a171a89669c1e9391e85c8

Observation 7abf953f-ea58-4078-a81d-cbbf4112e3e7 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.423238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.423238Z digest=sha256:bec771b6c52330ef45f76fa466a9709542fc2df6be934c2eebc7e67c2c8504fc

Observation a7db231a-cc16-43ef-94bf-19048719c464 · outbound

This paper cites Vila: On pre-training for visual language models, 2023.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Vila: On pre-training for visual language models, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.429088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.429088Z digest=sha256:272979d04b3be6273b927f44efecad17cf7b1361c910896a15b2b911b7dac384

Observation b76987c9-cda7-40b8-b57e-4ff84bd9cc45 · outbound

This paper cites Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want,.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.439359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.439359Z digest=sha256:97fcbdea1bbb02e4d775a413eb7837d0e7193643315f0f3557bce8d646895d96

Observation 8acde3fc-cc86-4f71-9ed9-cb1763dcb9a5 · outbound

This paper cites Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.447029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.447029Z digest=sha256:7f139b5ccfffb695528a9ea0488663c7c280382458abc0fd012560378c7d2eed

Observation 8971a870-89dd-42c6-acc1-8a9dc523125d · outbound

This paper cites Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models, 2023.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models, 2023

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.855150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.455813Z digest=sha256:7f3b06a189046d45a5cbebfdbc169b28f266cd760ca7c65043b21b5d9a5d89df

Observation 3c85cca2-2749-4b40-b36d-b4b1dc9ffcf6 · outbound

This paper cites Visual instruction tuning.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Visual instruction tuning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.832375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.471072Z digest=sha256:0be44f2c5708c301efcb68f569d1779e039a9d149c09ae1f151ea2806b95787f

Observation 8347f6f0-5b3f-425a-97d5-d273a265eeff · outbound

This paper cites Improved baselines with visual instruction tuning.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Improved baselines with visual instruction tuning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.808615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.481122Z digest=sha256:f9b0091a9ea763cc89ad1a153a6885c176b41bbbebedcbdc41311ac335d25897

Observation 9f64413e-8685-476a-be50-3f3720fad9d9 · outbound

This paper cites Llava-1.6: Improved reasoning, ocr, and world knowledge, 2024.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Llava-1.6: Improved reasoning, ocr, and world knowledge, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.785240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.489726Z digest=sha256:166949d6182cfd98da2142278006e155bd014d9debf26de7fde8daa6d5779f1b

Observation 5bf72967-eb8f-42c2-9f6b-337257311c0f · outbound

This paper cites World model on million-length video and language with blockwise ringattention, 2024.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts World model on million-length video and language with blockwise ringattention, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.762595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.495850Z digest=sha256:e87104d283a03eda6085c8e40407f1252df49567ff0d573b0abfadf9a90e6153

Observation 2199b5a4-c859-45be-a4b7-7428330582c9 · outbound

This paper cites P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.735115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.502656Z digest=sha256:ab72ed46189b6342564c8711510f619af073446e8b152468a36ea035b8ef1eef

Observation 218681ba-83d3-4018-8ed1-ace6947eee90 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts MMBench: Is Your Multi-modal Model an All-around Player?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.512227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.512227Z digest=sha256:307a9b5e248488143bc6392e7c2a0852524e8d7b94b28e0661a8de40205a0df3

Observation a383fb7a-48a7-4f7e-bde7-ae9d8ba8a5b4 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.522927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.522927Z digest=sha256:b4fa06ab88ede52186ce9927f76a9ed707a689215cba76f1421dd74dc298078f

Observation aaf713b1-1088-4f31-a13a-5970acd6111f · outbound

This paper cites Learn to explain: Multimodal reasoning via 10 thought chains for science question answering.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Learn to explain: Multimodal reasoning via 10 thought chains for science question answering

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.710147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.529520Z digest=sha256:fba9ad4384421a99f8622b6a222004115bee61e7cbb7138a96eecb876a45dbbd

Observation 036e0362-408a-4ef8-b126-f2fd49856bf5 · outbound

This paper cites An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.535228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.535228Z digest=sha256:e7af90f361e6ffc659dc19b7fbfd8f1ca4930172327a54177b62992787b0b9ee

Observation 193a2832-4434-437b-a8e7-a55696fca248 · outbound

This paper cites Visual Perception by Large Language Model's Weights.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Visual Perception by Large Language Model's Weights

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.542580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.542580Z digest=sha256:60041dfb4a76dcc5b647525e6ec4ff6d22095bf90aa1d23d941b6e4ec8e325e7

Observation ad8e103d-077e-43ea-8636-f5f7d1e8be23 · outbound

This paper cites Token Pooling in Vision Transformers.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Token Pooling in Vision Transformers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.549984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.549984Z digest=sha256:9e76d7bb7c3d403b42b85635b6ba1ec064de6fc2f01e81c5e790c23de74603c2

Observation c1f9c4be-1a33-4d87-b263-3526922d29ac · outbound

This paper cites Chatgpt plugins.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Chatgpt plugins

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.685716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.558033Z digest=sha256:38a0b9ef8a6ae3bc100ee106b59c2e643c38f9def1ed26c1e3b3635c6e8d5dfa

Observation cd97e14f-003f-499e-94f0-530cd68d45ea · outbound

This paper cites Gpt-4v(ision) system card.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Gpt-4v(ision) system card

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.563801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.563801Z digest=sha256:7979f1575b057ecbb2cc83d522964083a33abd4a00ed7b939c30fecf05e9df8f

Observation 0a80067d-ff0b-4ec4-8da6-c99c998ff3d5 · outbound

This paper cites V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Rus- sell Howes, Po-Yao Huang, et al.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Rus- sell Howes, Po-Yao Huang, et al

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.644321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.570694Z digest=sha256:9e7e3167534ea3595b10c0345313a488346162b34d4c22f19755898f9b96bb61

Observation 6c73149d-2832-4f2b-a5ad-93c1871b829f · outbound

This paper cites Less is more: Pay less attention in vision transform- ers.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Less is more: Pay less attention in vision transform- ers

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.623986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.577940Z digest=sha256:6d10d5b6206469f7eeda76fceac7adc6a937dc95fb2621b62585d0b5061506b5

Observation 2280410b-3b1a-440a-8582-8be45c8eb50e · outbound

This paper cites Language models are unsu- pervised multitask learners.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Language models are unsu- pervised multitask learners

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.585147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.585147Z digest=sha256:02ff1cd3988c7576d15d66913efad33b447d461fb1f57449a746e3ea9cacb73d

Observation 043f2b9f-c29b-47ec-8e0b-c33c51fd2147 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Learn- ing transferable visual models from natural language super- vision

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.587710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.591793Z digest=sha256:3ab91b84a9908752eb6f5698eb8d00dd6fc475a4bd3977a788a6bf4825c26d4e

Observation 9ff59451-3750-4069-8ec2-b5bac5422d24 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Learning Transferable Visual Models From Natural Language Supervision

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.598489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.598489Z digest=sha256:d947fd2aaf4248766266a3507888b3c0ec6c93e95e4be18917541d230b76a5ea

Observation 1ae589ff-0e04-4017-91d4-2feb561732c9 · outbound

This paper cites Tokenlearner: Adaptive space-time tokenization for videos.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Tokenlearner: Adaptive space-time tokenization for videos

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.557328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.606248Z digest=sha256:0036586c98b099bedb20babdcda2e22aa525a7deb6521fc315dd9a200d22868d

Observation cfa4e5d7-c75a-4d76-a02d-0658841969f1 · outbound

This paper cites AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.615817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.615817Z digest=sha256:0a1ece201f6b608ca670444e983c820ed46fd62efa23d000c19d1f3b5da28026

Observation 4af9e36e-5d89-4cab-8991-79139306fd1e · outbound

This paper cites What does CLIP know about a red circle? Visual prompt engineering for VLMs.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.622598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.622598Z digest=sha256:42ebb12ec40946286ac4f0aa0b82867b884c77ce2e303464580d2ff0b6a3f732

Observation 9943d15d-6ed2-4702-8a41-1d39e9e9997d · outbound

This paper cites Unleashing the power of prompt-driven nu- cleus instance segmentation, 2024.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Unleashing the power of prompt-driven nu- cleus instance segmentation, 2024

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.535200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.629753Z digest=sha256:23bf7ba312787beee2544ac8bc95c6fba13ab83c1f5b309d8ed990a58f7baa22

Observation 30569455-0054-4bd0-80d9-6eabf4ce7fa0 · outbound

This paper cites Towards vqa models that can read.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Towards vqa models that can read

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.634993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.634993Z digest=sha256:95c4ef197f92d842c3db1b62d9f4449983bd74e7b49bcc5178bfd584ca53a704

Observation de667c41-506a-4b77-9f2a-55f0efb3212c · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.642361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.642361Z digest=sha256:d9e87a66df05df15ec91c1a027857cfa14f6ee25b2d58781cad03cfc1be34bc8

Observation b3406eeb-679a-4d6b-abfa-75d4aadba639 · outbound

This paper cites Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context, 2024.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context, 2024

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.649110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.649110Z digest=sha256:3804c36a333c690852507752b21540cb05548b6df0bfb3597738c3fce8b192a5

Observation 0c2dc4cc-e423-4ec9-91f3-2fad006c364d · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.656061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.656061Z digest=sha256:051d156657ea6c19ec61d4898d9395de6cca86bda53b01c759d51a61dfb90206

Observation 64db2029-d43b-4f97-bded-009ff49a11c7 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.469088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.665815Z digest=sha256:53836afd2dfe9ce7626fb6eceb05a3e114fe69b97624aadac20cfb4594b75794

Observation 1c36a860-840d-4904-a7a0-6f49d18cb409 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.671252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.671252Z digest=sha256:488a934491e93061d0c4e492a011eb42b05cfe977430760ae7e38dca992075e9

Observation 2036c52d-8c0d-4cda-8d27-2748e21d8784 · outbound

This paper cites Attention is all you need.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Attention is all you need

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.678359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.678359Z digest=sha256:4d8bf6fd6906979a74e3aff9d21977ea9245ee6c2d50378f671e33ec46b2c875

Observation e202018b-0e6c-4e9d-ae3a-b16738d7d2ee · outbound

This paper cites What Makes for Good Visual Tokenizers for Large Language Models?.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts What Makes for Good Visual Tokenizers for Large Language Models?

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.685067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.685067Z digest=sha256:88c3fb42d9cb09213c1b45b20b3632ac2dc7c828ddb3a3019a4026193e85a4f7

Observation a26334f5-8b8e-46e4-94c5-d7e68ed9429f · outbound

This paper cites Tarsier: Recipes for training and evaluating large video description models, 2024.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Tarsier: Recipes for training and evaluating large video description models, 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.428397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.697846Z digest=sha256:fe4f512f97882b1b25f9ee1e524c6544d6c0186a2cddd4f2ea6de60ba025702c

Observation db3febaa-6b80-4aa2-a1d9-92daa63ae9f0 · outbound

This paper cites Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.402379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.704154Z digest=sha256:92496e30fdf64430a3a01ef5d0dd5a2a841eabcaf15d0df2ccaa7d077223b6a2

Observation 06bd867a-6b0b-4965-8d69-050956f876dd · outbound

This paper cites Learning to prompt for con- tinual learning.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Learning to prompt for con- tinual learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.711559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.711559Z digest=sha256:bf5398c3a845fbfb29b81ced3f5a6883e28dfd2c23e2ddaa8a756a4deca6d136

Observation 96fc4e0b-9ae3-4075-8cde-94c63626d417 · outbound

This paper cites Mio: A foun- dation model on multimodal tokens, 2024.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Mio: A foun- dation model on multimodal tokens, 2024

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.360335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.718058Z digest=sha256:817ca2680373012cee02d29f4aae155edc53bd83e20b642e4655a1120341c93e

Observation 2d01d96d-148d-4bb8-85e8-565a1e4faf39 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.723519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.723519Z digest=sha256:56fdc260ea52eb2227b6b19f7981388d5a2388338cebf81060c3d44a789243b0

Observation 59ad4f3e-7d88-4f55-8152-b890cac004d4 · outbound

This paper cites Grok-1.5 vision preview.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Grok-1.5 vision preview

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.335965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.740454Z digest=sha256:6a9940396c9e386341bb419a538dfdb29b55c1fa9a69df4e9f19bef4ec82a7d9

Observation a05d27ac-d832-4f2d-8a5b-40941e86c5f4 · outbound

This paper cites C-pack: Packaged resources to advance general chi- nese embedding, 2023.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts C-pack: Packaged resources to advance general chi- nese embedding, 2023

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.306509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.750696Z digest=sha256:2d93cd87da7c6302a6ca7c7154a620fa6c7adf7c916106714f525bed90ad092d

Observation 88b0479f-34c2-4ba1-82ff-1d90556f228c · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.761561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.761561Z digest=sha256:e186a4ac08597529f947eb625d57bf5c3c3bcfba388c8c21f90565296316ccfb

Observation e7e61a9a-3b27-427e-909c-34c217e0147b · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.768960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.768960Z digest=sha256:0d610073e1ef436a371d40a0dab19928f619a625dfcb0af2a6ae9cf553447743

Observation 120c3296-73f3-4f4e-b0da-6bc8c5b4acfe · outbound

This paper cites Libra: Building decoupled vision system on large lan- guage models.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Libra: Building decoupled vision system on large lan- guage models

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.279187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.777203Z digest=sha256:13413d6eb7168c2c1820887153ecffe9b4bf41a09b072f21a01b5dcdcd69add0

Observation defe557f-ac3b-41f1-b193-4d7cc03f4cc3 · outbound

This paper cites Efficient model personalization in federated learning via client-specific prompt generation.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Efficient model personalization in federated learning via client-specific prompt generation

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.248758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.788220Z digest=sha256:a3da855bd0f6b589a61803c8d966009c3e125b29eaf4f0d64b7f5fb764626080

Observation 06b0a2f4-7a05-4924-86f1-0dfd06cf2d91 · outbound

This paper cites Minicpm-v: A gpt-4v level mllm on your phone, 2024.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Minicpm-v: A gpt-4v level mllm on your phone, 2024

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.224539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.794324Z digest=sha256:34ca64a5ec1f7b56e7282a0251caef000b6e114460abcc645d17298b96a592e2

Observation 9e1d771a-9a3f-48da-8d43-eb11fc489dc7 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.191202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.801180Z digest=sha256:a1f8e34b687958030126e67d53aa2eb51f14457d33bd03493b876de0f38213a0

Observation 4f6201c2-9065-4d2e-b0a6-958acf1c5ee5 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.809046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.809046Z digest=sha256:967d774209693adeb49ff82ee015ef8c41135ab32195656d390f8ca0c8b9ad10

Observation 9583c768-028d-405c-bb2c-4c0df17add0e · outbound

This paper cites Sigmoid loss for language image pre-training.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Sigmoid loss for language image pre-training

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.159672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.814869Z digest=sha256:150b4f241a7a38dff6ab4a0881064bbaa8eaf342d04523239165e630dd2a43a6

Observation 874eb5fc-de1b-4c35-9243-076f57beb7f7 · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.820564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.820564Z digest=sha256:6f22fc2972c8d8b698183ed06183d43ca143e57507deadf1a07d00b02bcd206c

Observation 035551e1-a0dc-4aa6-ad14-198223eb7ef5 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.827282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.827282Z digest=sha256:d205b253717215b6ba05eb4773b3e9b9c33c33890a0708454f763f950b538f2b

Observation 28f86b27-a50d-4d0d-87ce-21343dcc636a · outbound

This paper cites Long context transfer from language to vision, 2024.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Long context transfer from language to vision, 2024

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:51:53.128400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.833316Z digest=sha256:3c97b1914365e2352a2275d76676e238c7cc10f69186bdbc5e6f48864f19ab3c

Observation 9fe4f246-6e6f-4a33-a99c-7c132b2b21eb · outbound

This paper cites Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.844982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.844982Z digest=sha256:93d17b5bb1a2d373f157654f189cbe943e3221511ba6cf11e2cd91697e12f1d7

Observation 9f9955f7-6598-40b6-8644-766d416fc479 · outbound

This paper cites Learning to prompt for vision-language models.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Learning to prompt for vision-language models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.853839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.853839Z digest=sha256:fa77619f702cf80696ef7c9200b9b4bd0c7c73c4244aef6b0d7e4bcf5fe091a8

Observation f5f7f052-85c0-40d6-b03b-efe3baba85c9 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.860532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.860532Z digest=sha256:f9cb483f26c9b7c703711e7f07be21caf08033b04aa75a58d9d94f9e4208961f

Observation 316b5729-ad1f-47b6-9691-0148398ce56c · outbound

This paper cites an unresolved cited work.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:51:53.086621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.866589Z digest=sha256:467507d1f1c5a4b13192fa915f7d2551443a3ca81600894fed77e6193791a1b7

Observation cf9350ed-e91c-4afd-8cb4-03f0d90f6945 · outbound

This paper cites an unresolved cited work.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:51:53.062413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:51:51.871655Z digest=sha256:0ad55024cb256f5c73e9245dd7569bba0269496d2bf415e6c5e0ef34d90bd80b

Pith citing papers

No inbound Pith citation observations are available.