Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:42:13.979447Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 2 inbound Pith citation observations for arXiv:2506.16691.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:42:13.979447Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-26T18:08:56.044278Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T03:19:31.831734Z
73 of 73 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9aee4c1c-04f1-4f91-87fa-21060a6a7576 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ddc843a-52dd-4c47-bc13-381f8f0a0edb · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Flamingo: a Visual Language Model for Few-Shot Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02645f41-f1cb-41ec-b2a7-d7dea623420d · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Bottom-up and top-down attention for image captioning and visual question answering
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dc02497-c659-4e8c-8485-0fe1dc005226 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation The claude 3 model family: Opus, sonnet, haiku.Claude-3 Model Card, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6e64210-2b0a-4bfd-980a-64efcfea572c · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Vqa: Visual question answering
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c3e6796-1489-41cb-9281-7828da2465a2 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88557c66-ee15-4964-a694-9fe9e21332f9 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Layer Normalization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b14f4bd6-4b83-419e-a955-1f6fc8c0609e · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20014592-ddc7-4c9d-ac48-4e5e6e5bce6b · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Nltk: the natural language toolkit
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82b5761f-a244-4e85-a8c3-af36bc21f00c · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c0cd4ff-faab-4b10-8712-113e8b02756c · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7381ecb4-c8b7-408c-aeab-a652d13bdaa3 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86ed57a2-71bd-40c1-853d-80c33517ed53 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Uniter: Universal image-text representation learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5bee69ee-7818-46a8-9df8-f73dd6b562f5 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101, 2024
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 462315c5-df46-4cd5-902b-3e7e4a06829e · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df06f042-e510-4e5b-9b2e-e0738b1bc496 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79fbf471-1e39-4852-ba00-21bc3a6d6044 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Gonzalez, Ion Stoica, and Eric P
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d6832b0-410d-4b1f-92d6-4cc10e3b1c2e · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 449027db-9c5c-4f6d-8032-641bcdf7cb38 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13947f52-783b-454a-91fa-8130a6778368 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b2efe2f-aa52-4eb1-9131-c6e27d88536d · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95f8fe7e-d105-4ad8-81f1-b0a2a69cf37a · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e0b6d1c-7c87-40ec-9329-9ffcb2d1820f · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Vizwiz grand challenge: Answering visual questions from blind people
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78709f2e-cfab-4f01-bf8a-e7c4225d944f · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Measuring Massive Multitask Language Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48cc5443-fa87-494f-b4f2-7827e6dc9681 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae2c4f91-7e96-4078-b5a6-3d1d269e129a · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Revisiting visual question answering baselines
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12d634bf-4016-487f-8859-b9e3e94c7490 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation LLaVA-OneVision: Easy Visual Task Transfer
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce93ee01-fe91-472f-bcf2-f0a7db191003 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Seed-bench: Benchmarking multimodal large language models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 723cb2b3-701b-41ba-b3d1-7188ab3ff7da · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30b5d714-c6a1-48e9-9f0c-1694aef87db8 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b2980e-3966-49a0-8db7-1f6893e8df60 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Align before fuse: Vision and language representation learning with momentum distillation.Advances in neural information processing systems, 34:9694–9705, 2021
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97785b96-6907-414c-b42a-18ef2769ef47 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Videochat: Chat-centric video understanding, 2023
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2db0624-1235-4266-9b1c-cdaaacf9840a · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b47862b3-ec9f-4319-b291-e06a96870b66 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Llama-vid: An image is worth 2 tokens in large language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8622997f-673b-4e3f-a8c5-87bf0c700d2f · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Evaluating Object Hallucination in Large Vision-Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a231d075-a609-4905-b05a-9bab02935191 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b8cf93e-0509-4b35-8c94-f8154c1d8458 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce0f5d7b-3421-45af-b0a9-5c6c2810d65c · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 019916fc-4086-4e9e-ba18-06309506835d · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Improved baselines with visual instruction tuning, 2023
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 403461a0-177e-4689-a44c-3525b9891f31 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 297559bd-906f-46be-b760-ff7b5aa00b18 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdd77114-445b-4867-9373-2a5056bc7b38 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8519c36c-eef2-429e-aa49-9d5a12ae840e · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67b001ae-295d-45cf-906a-cafeb2cb178e · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a59232dd-a6e0-4986-affb-a787d45ad1d7 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7bf4ee0-a720-4031-a320-aec3d66ec9fa · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af78bd7e-562d-4997-bc73-5b3ee4c2cb8d · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Egoschema: A diagnostic benchmark for very long-form video language understanding.Advances in Neural Information Processing Systems, 36:46212–46244, 2023
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4a41ed7-6599-47f3-bd10-b788e8688ad2 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.Meta AI Blog
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6d8f519-7a59-4d74-9593-14d47ac90e97 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Learning transferable visual models from natural language supervision
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fdf2af3-937f-4bc1-ac7c-302b208475d2 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Language models are unsupervised multitask learners.OpenAI blog, 2019
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82f9ab2e-4bd1-4a08-917c-5cecb7d27d47 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Searching for Activation Functions
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7754035-5226-457d-a5d6-8fd86eda5e15 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation CinePile: A Long Video Question Answering Dataset and Benchmark
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18327225-f708-490a-a6ed-ddbf008e6d03 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Towards vqa models that can read
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63618e8b-195e-4ee8-969e-b82cdfb805f5 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Flops profiler, 2025
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 578a6673-3b6e-455d-8f3a-493a78cf7456 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Gemini: A Family of Highly Capable Multimodal Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb38303f-ae63-4e72-844c-9059286c1dbc · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Mlp- mixer: An all-mlp architecture for vision.Advances in neural information processing systems, 34:24261–24272, 2021
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab8e50e7-81d8-4492-a41f-7c3711c74420 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2debb5ab-3107-4b9d-bc1a-d9f60141ed15 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e808306e-e95a-40df-a2bf-51795d5012aa · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Patches Are All You Need?
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a476ad82-1e46-458e-a4f5-f3e88a671a9e · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Attention is all you need.Advances in neural information processing systems, 30, 2017
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdea2b54-0b7f-49e1-965a-c3998b9c5e04 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 556e51c2-2ca4-4449-9009-2ac112aa07e0 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Cogvlm: Visual expert for pretrained language models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4716d017-f258-4481-801d-4b9bdde49b55 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Qwen2 Technical Report
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90372d15-99a1-4b77-8d32-c05862fa7573 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation mplug-owl3: Towards long image-sequence understanding in multi-modal large language models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1843186c-721c-4ef5-9ade-c338527de335 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9467c2b3-921c-48c6-a38b-87433cedbbc4 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation CoCa: Contrastive Captioners are Image-Text Foundation Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2685d41-a199-451c-b56c-5dd04ca1235d · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Sigmoid loss for language image pre-training
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 601d00e5-a49a-4575-8f25-1aa2ee7a768e · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Root mean square layer normalization.Advances in Neural Information Processing Systems, 32, 2019
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de1d6570-519c-4776-876d-ecf848bd4c48 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed01f621-792c-4699-afb1-f54e0c1f7f6d · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35bc73c8-054a-4132-919a-d3a9d7e7a3e6 · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Long Context Transfer from Language to Vision
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 149c2511-5fae-4c5e-9958-47f2b52cd5fb · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation Wings: Learning Multimodal LLMs without Text-only Forgetting
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7abdeaac-5cb1-4a79-9983-d83daf5b6c7d · outbound
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation MLVU: Benchmarking Multi-task Long Video Understanding
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29312d64-c8fe-498a-bfda-894e49983ae1 · inbound
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6862ffb-5dcc-4899-848e-b4b0aa1d8cdb · inbound
The Hidden Evolution of Disguised Visual Context inside the VLM LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.