Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:53:23.309664Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 3 inbound Pith citation observations for arXiv:2504.19627.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:53:23.309664Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:55:17.207198Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-13T20:33:17.140335Z
77 of 77 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7cb46e91-eb48-41c7-92c8-bfa397199e3c · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d0a56a5-64e3-4ac9-9108-80443222eaad · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ed35da3-db7b-464c-8295-5a5d6a813147 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Instructblip: Towards general-purpose vision-language models with instruction tuning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a46b5166-1d76-433f-81d2-8e62119b7875 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbc81211-1750-476d-8e3f-b3e45b621b85 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Few-shot adversarial prompt learning on vision-language models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 60a87125-fb5b-40e8-8166-923489bf0d4d · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning A Survey on Vision-Language-Action Models for Embodied AI
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baa2b6f0-fd0b-41c6-8443-d1caefe4f41c · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5de1d7e7-ffa3-488e-a6d6-0523a92297a5 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 184d4616-edd7-4750-abc0-a424deda5c62 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Large (vision) language models for autonomous vehicles: Current trends and future directions
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 125a7038-da91-4d41-a57d-fb0b32f86c6c · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1772d2af-35bf-481c-8085-05657b88811c · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Visionzip: Longer is better but not necessary in vision language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26ca3679-d3c8-4375-a3f7-70c225e2e1b7 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Dynamic programming
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c42cde95-d287-4bba-95b6-27f6b109b870 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Learning transferable visual models from natural language supervision
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a075c2cb-a895-4cf8-9bc6-9aa595b1b141 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7d9b35f5-293f-426b-8548-5c0c2b443443 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Visual instruction tuning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4d6ef1b-882b-4b0b-87bc-3abfec4fbdcd · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43d2e9e3-b84b-48b8-8a2e-2ced21a70350 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Gqa: A new dataset for real-world visual reasoning and composi- tional question answering
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 02ad6099-06b8-4e66-92c6-7d783010e01c · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Vizwiz grand challenge: Answering visual questions from blind people
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6693ca22-460b-4a02-81ca-90555dfdac53 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8450252e-628c-4c77-8088-d01c3aa7a2fc · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Evaluating Object Hallucination in Large Vision-Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee254508-bd9d-41b0-91c0-40ef0f8af08a · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b74c2ad-5b6b-43bf-8c90-301577ca761d · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning MMBench: Is Your Multi-modal Model an All-around Player?
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc96e961-9e3e-4580-8239-753428c6f52c · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4323062e-8326-4835-8bdc-a0f6384cacdc · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba6e66c1-2206-4e77-b54b-0f8353a4c00c · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Towards vqa models that can read
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a874e4c-c043-42d0-b833-9300175f110c · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3a8dab6-9012-4923-9ab5-e5fc7f5b2c6b · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Referitgame: Referring to objects in photographs of natural scenes
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddfe012d-c71c-4bea-88d1-5f697d078454 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4676939-ae01-466d-b42d-5fe864d31a36 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Open-vocabulary detr with conditional matching
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5c55bf5d-da0f-46bd-99c0-70085a01d317 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Scene parsing through ade20k dataset
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f8696bc-7c0d-4266-b273-86b86cc46560 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c4945215-6adc-48d7-86dd-8a006d541023 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Video question answering via gradually refined attention over appearance and motion
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8cdabcd1-5c12-471b-a3ea-9b5ddd5cfc3d · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01aa8a78-92ef-4483-9c31-afa32bc6a364 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Lmms-eval: Reality check on the evaluation of large multimodal models, 2024
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 044bf9ca-59f7-4e07-be48-3c46853525af · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4950edca-f73c-49f4-8e3e-2c4f31b950b5 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Decoupled Weight Decay Regularization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 398a39d2-4839-4375-a88f-c193ae587840 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Least squares quantization in pcm
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 194cc1a1-ea3f-4f64-815e-13847a58483c · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 66aca493-689b-4178-a5f0-76f29e2a6b29 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Introducing idefics: An open reproduction of state-of-the-art visual language model
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3afc06f7-f63c-46b3-98ea-f3831861130c · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f7d9831-2036-445d-acc1-bd80ebf19868 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8166c230-2559-4498-8429-5b800357ffea · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 451a1e11-31fe-4c0f-83fc-853e607e4a7b · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3a7d57e-628c-45f2-ba7e-df93e34d1504 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Panoptic segmentation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a1974b4d-306f-4158-90b9-7e4745b1e338 · outbound
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 866ae55b-0eaf-4502-89f1-0875fa44cfcb · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Side adapter network for open-vocabulary semantic segmentation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 59ae4213-c261-4822-8aee-818be2a1d7a7 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b57d74d1-f250-4cce-987a-fce997094e50 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Coco-stuff: Thing and stuff classes in context
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5629e619-316f-4b4f-8320-a1d254c3f395 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Cat-seg: Cost aggregation for open-vocabulary semantic segmentation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 65787873-357b-40aa-86f2-0de5a299b0ab · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9d839f0-6ec9-4250-8664-a2eeb962113c · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Token Merging: Your ViT But Faster
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 827f4e3a-4921-4742-850b-2ed579237af7 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a996f7f8-2fca-4bad-abdc-7b3b162d9d90 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36dc7495-b72b-4676-9c79-17a83e1ba0e0 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2a58e83-7e7c-489e-a5e7-d95904f35d84 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Efficient large multi-modal models via visual context compression
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4b6e855d-f0e3-4462-937c-2e3a7a51e4be · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Matryoshka Query Transformer for Large Vision-Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c971ad23-f87b-4597-ab6f-bc2d6d00ccfe · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5618281-9002-4cab-8243-002686e3ef42 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d554a7df-a338-4878-9613-608e0b47a067 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Vary: Scaling up the vision vocabulary for large vision-language model
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 94c4da0a-b7a8-4151-9222-30312bdfe980 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Distilling large vision-language model with out-of-distribution generalizability
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 973a8408-6289-41b6-b827-0f546a63f4da · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Rs5m and georsclip: A large scale vision-language dataset and a large vision-language model for remote sensing
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1b76abfe-60ab-43d3-b62d-8f531ea618e4 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Visual In-Context Learning for Large Vision-Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2801aa98-1897-4871-94ba-0997556ae668 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Anomalygpt: Detecting industrial anomalies using large vision-language models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1a0b1ba4-5a23-418e-8f31-8278e8dcf93f · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Matryoshka query transformer for large vision-language models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 787510c3-31c8-4478-bb6e-001b07875047 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Pyramidclip: Hierarchical feature alignment for vision-language model pretraining
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3421e68b-421d-4b70-9fbe-95152307f7ad · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Matryoshka multimodal models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7ea3ea87-15dd-47b0-8c9a-030b198b93c0 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0515a80-6d45-40ff-90d7-e42ade2f7460 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning DreamLLM: Synergistic Multimodal Comprehension and Creation
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab44a47c-17d8-4fce-8e05-3e7632e2a063 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Vision-language models for vision tasks: A survey
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baacb898-2f08-4b8e-ac46-a2cabc026a14 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e123e2f9-c737-4629-be61-731edebd03e0 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning A Survey of Vision-Language Pre-Trained Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c852f6d2-61f0-4313-9e83-13b3a24245e0 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning A Survey on Hallucination in Large Vision-Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51a3d3c5-7bed-4174-98dd-978f5cb96ff2 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Vqa: Visual question answering
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 46e7a9cd-f51f-430d-ac1d-8d54fa3e8eb7 · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Vizwiz grand challenge: Answering visual questions from blind people
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 103d7a38-3bd6-4b96-ba23-4a2a22f5886b · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Cider: Consensus-based image description evaluation
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aee89f1-4bfa-4a48-890f-db5498a783ab · outbound
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Fully convolutional networks for semantic segmentation
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 38b7fd2e-bc50-48a5-89fb-25dcfb9270f2 · outbound
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0a048ce3-2f65-47d8-9e99-1a2d10ec34de · inbound
Towards Modality Generalization: A Benchmark and Prospective Analysis VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 439f993c-bf43-46df-ab72-58020bbae29a · inbound
CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 977f4470-facd-449e-89d1-1b2e24913411 · inbound
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.