Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:51:51.871655Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 0 inbound Pith citation observations for arXiv:2411.13909.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:51:51.871655Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
96 of 96 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 25e15dd4-fb5e-48d0-a5ec-8b91e0174180 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Fuyu-8b: A multimodal architecture for ai agents
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dfca88f-5ed2-48a6-909d-ccbd65904bf5 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Flamingo: a Visual Language Model for Few-Shot Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8be1ce38-cf7a-4685-921f-6b439e7b402a · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16ea1d5e-ef55-4af3-ae88-158a8d85d017 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Token Merging: Your ViT But Faster
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01be99ee-480d-4502-abda-68e0d3eed005 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Language Models are Few-Shot Learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d86304a0-35c0-4903-8796-e548f8ea350e · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ea5e044-5a17-44a0-8c83-4e88c9eb8af6 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d9af478-0dec-4101-8916-7bb93b3584eb · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Towards unifying medical vision-and-language pre-training via soft prompts, 2023
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cd7bae4-de4b-450e-9772-b8ebfb2a07c5 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Gonzalez, Ion Stoica, and Eric P
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96b12ac5-9880-4e23-a978-7e31e4b2bda4 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d27889d3-50b6-45b6-ac7f-944f46061dd6 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0482907-b9f7-4175-9def-39199e954b99 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52e81a30-a017-4365-be17-ac2af99121f1 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts An image is worth 16x16 words: Trans- formers for image recognition at scale
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aed7bdbc-02b3-4e40-9283-6eee2677b121 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts The llama 3 herd of models, 2024
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56ba8c92-f660-4a0f-89ee-b9ceb40c2164 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebf3e11b-eea7-4ca5-af0e-8dc4246c0da3 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Domain adaptation via prompt learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78bc37ef-d666-4c01-a95c-01f59493b87b · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a876b31c-c966-47f7-84f4-ebae66d9cb7e · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Openllama: An open reproduc- tion of llama, 2023
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21eaa993-e0b6-4e4f-9de9-885d1ab33ca1 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 132734f6-3fda-4f3f-a0d1-198453b5212a · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Vizwiz grand challenge: Answering visual questions from blind people
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02b4bf6f-2fbb-47a7-97bd-7807a5d0dc9c · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec195a06-b6df-4ebe-8992-51b49c28db46 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ebef389-4205-4652-b6ca-8ca8bf1ca476 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Vi- sual prompt tuning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81822f56-44ba-49f5-ab02-8fd91dfd962a · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e954659-da04-4796-a11b-f9b7f15e460b · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a9dfaba-2753-403c-9d9d-6fe224fdb951 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Maple: Multi-modal prompt learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45f50271-4c28-4ee0-9f1d-159ba8212e0e · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Segment Anything
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fff57757-6d4d-4a60-864b-924505ee6c53 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Spvit: Enabling faster vision transformers via soft token pruning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d29c858-c89c-41a0-a18d-8d6ab9b73095 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts The power of scale for parameter-efficient prompt tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e683a8cf-a4cf-4f82-b70a-9b1446688624 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3e2f510-242d-4244-ac39-afc68dc33698 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Task-specific fine-tuning via variational information bottle- neck for weakly-supervised pathology whole slide image classification
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation db50e5e4-7f4e-4c1e-ba3d-543fec037ddb · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Rethinking transformer for long contextual histopathology whole slide image analysis, 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 896436c2-0078-4fcb-943f-bde17fb8f031 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e6e73a-87ec-464d-81cf-0d72812d3c92 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Tokenpacker: Efficient visual projector for multimodal llm, 2024
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 30fb327a-11ae-4523-ba7a-daa3d2ab37fa · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Prefix-tuning: Optimizing continuous prompts for generation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e4302872-b085-425b-a8ed-fa7e91fb5fa5 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Evaluating Object Hallucination in Large Vision-Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46b672a0-b583-4600-941d-361400c3489b · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Mon- key: Image resolution and text label are important things for large multi-modal models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9545a833-4f95-49d7-8cb3-26e0a38b59da · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Not all patches are what you need: Expediting vision transformers via token reorganiza- tions
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7abf953f-ea58-4078-a81d-cbbf4112e3e7 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7db231a-cc16-43ef-94bf-19048719c464 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Vila: On pre-training for visual language models, 2023
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b76987c9-cda7-40b8-b57e-4ff84bd9cc45 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8acde3fc-cc86-4f71-9ed9-cb1763dcb9a5 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8971a870-89dd-42c6-acc1-8a9dc523125d · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models, 2023
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3c85cca2-2749-4b40-b36d-b4b1dc9ffcf6 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Visual instruction tuning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8347f6f0-5b3f-425a-97d5-d273a265eeff · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Improved baselines with visual instruction tuning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9f64413e-8685-476a-be50-3f3720fad9d9 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Llava-1.6: Improved reasoning, ocr, and world knowledge, 2024
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5bf72967-eb8f-42c2-9f6b-337257311c0f · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts World model on million-length video and language with blockwise ringattention, 2024
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2199b5a4-c859-45be-a4b7-7428330582c9 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 218681ba-83d3-4018-8ed1-ace6947eee90 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts MMBench: Is Your Multi-modal Model an All-around Player?
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a383fb7a-48a7-4f7e-bde7-ae9d8ba8a5b4 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaf713b1-1088-4f31-a13a-5970acd6111f · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Learn to explain: Multimodal reasoning via 10 thought chains for science question answering
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 036e0362-408a-4ef8-b126-f2fd49856bf5 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 193a2832-4434-437b-a8e7-a55696fca248 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Visual Perception by Large Language Model's Weights
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad8e103d-077e-43ea-8636-f5f7d1e8be23 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Token Pooling in Vision Transformers
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1f9c4be-1a33-4d87-b263-3526922d29ac · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Chatgpt plugins
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cd97e14f-003f-499e-94f0-530cd68d45ea · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Gpt-4v(ision) system card
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a80067d-ff0b-4ec4-8da6-c99c998ff3d5 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Rus- sell Howes, Po-Yao Huang, et al
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6c73149d-2832-4f2b-a5ad-93c1871b829f · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Less is more: Pay less attention in vision transform- ers
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2280410b-3b1a-440a-8582-8be45c8eb50e · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Language models are unsu- pervised multitask learners
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 043f2b9f-c29b-47ec-8e0b-c33c51fd2147 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Learn- ing transferable visual models from natural language super- vision
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9ff59451-3750-4069-8ec2-b5bac5422d24 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Learning Transferable Visual Models From Natural Language Supervision
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ae589ff-0e04-4017-91d4-2feb561732c9 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Tokenlearner: Adaptive space-time tokenization for videos
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cfa4e5d7-c75a-4d76-a02d-0658841969f1 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4af9e36e-5d89-4cab-8991-79139306fd1e · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts What does CLIP know about a red circle? Visual prompt engineering for VLMs
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9943d15d-6ed2-4702-8a41-1d39e9e9997d · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Unleashing the power of prompt-driven nu- cleus instance segmentation, 2024
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 30569455-0054-4bd0-80d9-6eabf4ce7fa0 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Towards vqa models that can read
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de667c41-506a-4b77-9f2a-55f0efb3212c · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3406eeb-679a-4d6b-abfa-75d4aadba639 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context, 2024
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c2dc4cc-e423-4ec9-91f3-2fad006c364d · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64db2029-d43b-4f97-bded-009ff49a11c7 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1c36a860-840d-4904-a7a0-6f49d18cb409 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2036c52d-8c0d-4cda-8d27-2748e21d8784 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Attention is all you need
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e202018b-0e6c-4e9d-ae3a-b16738d7d2ee · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts What Makes for Good Visual Tokenizers for Large Language Models?
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a26334f5-8b8e-46e4-94c5-d7e68ed9429f · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Tarsier: Recipes for training and evaluating large video description models, 2024
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation db3febaa-6b80-4aa2-a1d9-92daa63ae9f0 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 06bd867a-6b0b-4965-8d69-050956f876dd · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Learning to prompt for con- tinual learning
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96fc4e0b-9ae3-4075-8cde-94c63626d417 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Mio: A foun- dation model on multimodal tokens, 2024
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2d01d96d-148d-4bb8-85e8-565a1e4faf39 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59ad4f3e-7d88-4f55-8152-b890cac004d4 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Grok-1.5 vision preview
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a05d27ac-d832-4f2d-8a5b-40941e86c5f4 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts C-pack: Packaged resources to advance general chi- nese embedding, 2023
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 88b0479f-34c2-4ba1-82ff-1d90556f228c · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7e61a9a-3b27-427e-909c-34c217e0147b · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 120c3296-73f3-4f4e-b0da-6bc8c5b4acfe · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Libra: Building decoupled vision system on large lan- guage models
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation defe557f-ac3b-41f1-b193-4d7cc03f4cc3 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Efficient model personalization in federated learning via client-specific prompt generation
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 06b0a2f4-7a05-4924-86f1-0dfd06cf2d91 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Minicpm-v: A gpt-4v level mllm on your phone, 2024
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9e1d771a-9a3f-48da-8d43-eb11fc489dc7 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4f6201c2-9065-4d2e-b0a6-958acf1c5ee5 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9583c768-028d-405c-bb2c-4c0df17add0e · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Sigmoid loss for language image pre-training
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 874eb5fc-de1b-4c35-9243-076f57beb7f7 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 035551e1-a0dc-4aa6-ad14-198223eb7ef5 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28f86b27-a50d-4d0d-87ce-21343dcc636a · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Long context transfer from language to vision, 2024
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9fe4f246-6e6f-4a33-a99c-7c132b2b21eb · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f9955f7-6598-40b6-8644-766d416fc479 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Learning to prompt for vision-language models
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5f7f052-85c0-40d6-b03b-efe3baba85c9 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 316b5729-ad1f-47b6-9691-0148398ce56c · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Unresolved cited work
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cf9350ed-e91c-4afd-8cb4-03f0d90f6945 · outbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts Unresolved cited work
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
No inbound Pith citation observations are available.