Pith. sign in

Paper Citation Record · LEDGER

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning

As of 17 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 3 inbound Pith citation observations for arXiv:2504.19627.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19627 v2

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:53:23.309664Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:55:17.207198Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T20:33:17.140335Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7cb46e91-eb48-41c7-92c8-bfa397199e3c · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.013678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.013678Z digest=sha256:7593da610b43a48eb89023d3388feb73b3c588bf55bfe532766a2371b8f9974e

Observation 5d0a56a5-64e3-4ac9-9108-80443222eaad · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.018475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.018475Z digest=sha256:800a91b60cddaa23c69c56a6fd9aab1f8cb4594e7c008fcf10ee9ca61154680a

Observation 0ed35da3-db7b-464c-8295-5a5d6a813147 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.023000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.023000Z digest=sha256:7e0ac8135c9a9cdf0ee0b60d8438085fd987a15995022d166be33a6851bd25b0

Observation a46b5166-1d76-433f-81d2-8e62119b7875 · outbound

This paper cites DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.026924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.026924Z digest=sha256:2ea7a1b2fbe9666d842ab264ab35ffa8bbcebc34f41aca32f1de4ecc2c9012a1

Observation fbc81211-1750-476d-8e3f-b3e45b621b85 · outbound

This paper cites Few-shot adversarial prompt learning on vision-language models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Few-shot adversarial prompt learning on vision-language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.865303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.035776Z digest=sha256:73ffe1131f7df69e5755088ed620cec345e6a6ec6054259979117de2d0586480

Observation 60a87125-fb5b-40e8-8166-923489bf0d4d · outbound

This paper cites A Survey on Vision-Language-Action Models for Embodied AI.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning A Survey on Vision-Language-Action Models for Embodied AI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.043813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.043813Z digest=sha256:41ee706aff13c282ef7fce2f79037d57b293854f0d3cb941bea0ca78f8f3d639

Observation baa2b6f0-fd0b-41c6-8443-d1caefe4f41c · outbound

This paper cites EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.047792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.047792Z digest=sha256:35237df77094b6ea40d110c1ae125e2da258965e424a35dc6930ba297518fb11

Observation 5de1d7e7-ffa3-488e-a6d6-0523a92297a5 · outbound

This paper cites Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.051893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.051893Z digest=sha256:ff2979d87c74d655e65ef561cb4897c64680b07d440f56ee26e222a74e2c7e06

Observation 184d4616-edd7-4750-abc0-a424deda5c62 · outbound

This paper cites Large (vision) language models for autonomous vehicles: Current trends and future directions.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Large (vision) language models for autonomous vehicles: Current trends and future directions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.854388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.055858Z digest=sha256:71a6a2fe00d20535f46a411d1828fa4f46c1fbdeb803a37a2b0889714f284c70

Observation 125a7038-da91-4d41-a57d-fb0b32f86c6c · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.063240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.063240Z digest=sha256:cf1c15f9a509faf2251c90891888436fd8ced204148c3dc48090c0e1f1fe22be

Observation 1772d2af-35bf-481c-8085-05657b88811c · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Visionzip: Longer is better but not necessary in vision language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.067210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.067210Z digest=sha256:5614c421078c4fcc44437f27bb4971be3c10bb133b75367612a5298d26eb7df3

Observation 26ca3679-d3c8-4375-a3f7-70c225e2e1b7 · outbound

This paper cites Dynamic programming.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Dynamic programming

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.071287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.071287Z digest=sha256:eeb58fc508e692a4a46e35826a796c2d2623c7b4f626844a70e89ee625c059b5

Observation c42cde95-d287-4bba-95b6-27f6b109b870 · outbound

This paper cites Learning transferable visual models from natural language supervision.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Learning transferable visual models from natural language supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.075009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.075009Z digest=sha256:22a208addd9b6bea34d436f4dc9f4e695384c984ed622acd1312269dfae1df3e

Observation a075c2cb-a895-4cf8-9bc6-9aa595b1b141 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.829070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.078596Z digest=sha256:5a1578d90548a08dad14fb11ddc1d408b72a03f7dd8cb4bb7a535f2b78afa710

Observation 7d9b35f5-293f-426b-8548-5c0c2b443443 · outbound

This paper cites Visual instruction tuning.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Visual instruction tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.082634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.082634Z digest=sha256:48e24b74d2b0e8e07bfb1577882d6b7870a795440f7854b4c5a94e9465a5756a

Observation f4d6ef1b-882b-4b0b-87bc-3abfec4fbdcd · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.086330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.086330Z digest=sha256:1daa663ea9831ade404ef992172dee797d8f9e25f25d36b2f0692de2fd8644d3

Observation 43d2e9e3-b84b-48b8-8a2e-2ced21a70350 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and composi- tional question answering.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Gqa: A new dataset for real-world visual reasoning and composi- tional question answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.803885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.090053Z digest=sha256:d59d6bd622fbc3afd466be3fbf4d1ba1c3608ca4d954cb385be0962f19d80b56

Observation 02ad6099-06b8-4e66-92c6-7d783010e01c · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Vizwiz grand challenge: Answering visual questions from blind people

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.791621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.093515Z digest=sha256:11c98629876e5ec375701a7e4a359ad01fb32f1bfb50ad47d753d40c6d5813e0

Observation 6693ca22-460b-4a02-81ca-90555dfdac53 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.096850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.096850Z digest=sha256:36168b9c05986ba8b7844a9bb6a4f773d3fb822223aeb85ce4a624aebf6e399b

Observation 8450252e-628c-4c77-8088-d01c3aa7a2fc · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Evaluating Object Hallucination in Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.100232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.100232Z digest=sha256:f810d10665ee685d8afcbdab1b5e21d88164d370bd70499576ec1a08f85285d2

Observation ee254508-bd9d-41b0-91c0-40ef0f8af08a · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.103779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.103779Z digest=sha256:fcec55367dc9c28f57764120d8f12075731cb6df4c58c5d969bf49a62010cc58

Observation 4b74c2ad-5b6b-43bf-8c90-301577ca761d · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning MMBench: Is Your Multi-modal Model an All-around Player?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.107383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.107383Z digest=sha256:ddfba922957494783d5561ee6a3709a5799095252cff96a7c82c3cb0732969e3

Observation dc96e961-9e3e-4580-8239-753428c6f52c · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.111130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.111130Z digest=sha256:152b2500669b39d33eac88b96ab25bea2e1422656aeae92cc2c6c05a12b72468

Observation 4323062e-8326-4835-8bdc-a0f6384cacdc · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.114832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.114832Z digest=sha256:c06cc9840fbfe676f0c00f985c1b2cc47a3ad35710946dd3eda0df69ce73b193

Observation ba6e66c1-2206-4e77-b54b-0f8353a4c00c · outbound

This paper cites Towards vqa models that can read.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Towards vqa models that can read

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.118315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.118315Z digest=sha256:60a3c2ddcbe26a23e6823e5f5633a4493aece54c07b6af0c4207c81cfe381ebd

Observation 5a874e4c-c043-42d0-b833-9300175f110c · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.121996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.121996Z digest=sha256:7e3784aaf1932c02cd03360255e02e8b5258bb90b7aed19d07ab74bb021baff2

Observation d3a8dab6-9012-4923-9ab5-e5fc7f5b2c6b · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Referitgame: Referring to objects in photographs of natural scenes

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.125737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.125737Z digest=sha256:1ea6422e6b0772f5a6353c956cbca95a924a270b1ad47e49d4dfd369adf7ff15

Observation ddfe012d-c71c-4bea-88d1-5f697d078454 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.129212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.129212Z digest=sha256:d45e4e61647a4cb9ff058e419631c049169feec8fad5d44650b6e8d121e7f4a6

Observation a4676939-ae01-466d-b42d-5fe864d31a36 · outbound

This paper cites Open-vocabulary detr with conditional matching.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Open-vocabulary detr with conditional matching

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.697737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.132835Z digest=sha256:c0dc259e02a7695a33e1c3c19c246066b5f86209ed63c46c0ad4f1ede62d5907

Observation 5c55bf5d-da0f-46bd-99c0-70085a01d317 · outbound

This paper cites Scene parsing through ade20k dataset.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Scene parsing through ade20k dataset

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.136470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.136470Z digest=sha256:cc63663124d663d17d206c892504899344d60060436ceb8c9e5a51f8ce5f0445

Observation 6f8696bc-7c0d-4266-b273-86b86cc46560 · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.567173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.140146Z digest=sha256:674a5c2bcfb4a70be7f8a244f9e91be82dc272be36a460c84f84d976850a4474

Observation c4945215-6adc-48d7-86dd-8a006d541023 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Video question answering via gradually refined attention over appearance and motion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.422819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.143671Z digest=sha256:020e94f9b0e06b86a356445cee8a2977aadc142e98f4c5704f67e5c385f67d54

Observation 8cdabcd1-5c12-471b-a3ea-9b5ddd5cfc3d · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.147269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.147269Z digest=sha256:3ee99588ee3f30d75ccf02aa4c8f60eacb7e43ebcdb2b450445cab3c40352388

Observation 01aa8a78-92ef-4483-9c31-afa32bc6a364 · outbound

This paper cites Lmms-eval: Reality check on the evaluation of large multimodal models, 2024.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Lmms-eval: Reality check on the evaluation of large multimodal models, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.150698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.150698Z digest=sha256:d27cf7c79d9cd6dc2a5d97d19318637a5ed1de5ee1d3e216aec991d67ce48c3c

Observation 044bf9ca-59f7-4e07-be48-3c46853525af · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.153965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.153965Z digest=sha256:96f8e973f761faa3909e4837b0975021d4978f296a4b4163238ce117fd89ba70

Observation 4950edca-f73c-49f4-8e3e-2c4f31b950b5 · outbound

This paper cites Decoupled Weight Decay Regularization.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Decoupled Weight Decay Regularization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.157826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.157826Z digest=sha256:23d8be0d7d00a0ae3a1f01b20dcdd72effaa4f0cb5ac5714875c11e5e93f449f

Observation 398a39d2-4839-4375-a88f-c193ae587840 · outbound

This paper cites Least squares quantization in pcm.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Least squares quantization in pcm

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.390410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.161487Z digest=sha256:c81e2b4ca0278b25ea2252ebffb21abca25875c19c86b15be5b5f19cec1fd8ee

Observation 194cc1a1-ea3f-4f64-815e-13847a58483c · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.378541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.165006Z digest=sha256:b94ec7dcf6dea843af9b1439f7dad48bd2c9fc221c81101f7cb26e2cb01e4b2d

Observation 66aca493-689b-4178-a5f0-76f29e2a6b29 · outbound

This paper cites Introducing idefics: An open reproduction of state-of-the-art visual language model.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Introducing idefics: An open reproduction of state-of-the-art visual language model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.366329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.169026Z digest=sha256:1519b47b89d8792761977ef73d3dd1ad6829bdcb69271608cc4a6a2ede4d1ecb

Observation 3afc06f7-f63c-46b3-98ea-f3831861130c · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.172405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.172405Z digest=sha256:462baa70ce0208f4942f53ebc2094fa302c6ce466998cad899a81fb64944ddb7

Observation 4f7d9831-2036-445d-acc1-bd80ebf19868 · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.176102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.176102Z digest=sha256:8350182cb0ac7d93e33650a4ae5cd78eb7a74602d11f0df28c4a5b3deb021e65

Observation 8166c230-2559-4498-8429-5b800357ffea · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.180051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.180051Z digest=sha256:7e1c44448064205fdc76c802a219881de8494a4b777b85328c613204fc2fdcb3

Observation 451a1e11-31fe-4c0f-83fc-853e607e4a7b · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.183611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.183611Z digest=sha256:ca1e32a570ea2515750055fb1ecc803bf062817232f3fc0498b37143465ce7bd

Observation c3a7d57e-628c-45f2-ba7e-df93e34d1504 · outbound

This paper cites Panoptic segmentation.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Panoptic segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.276606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.187798Z digest=sha256:46cda40d90f08d77b1530f8dce2959c28655e4ccf54fbf2b8362fe69488a9024

Observation a1974b4d-306f-4158-90b9-7e4745b1e338 · outbound

This paper cites Mask r-cnn.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Mask r-cnn

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.139140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.191714Z digest=sha256:d97a05f894f8e7add5545f8cad009d8077b7574f7e19803797111d7c5df8158f

Observation 866ae55b-0eaf-4502-89f1-0875fa44cfcb · outbound

This paper cites Side adapter network for open-vocabulary semantic segmentation.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Side adapter network for open-vocabulary semantic segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:24.022353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.195408Z digest=sha256:15eed25ee952da274c7cd06a41e46ac1d553acc65589a60619673ac17d000566

Observation 59ae4213-c261-4822-8aee-818be2a1d7a7 · outbound

This paper cites F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.199481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.199481Z digest=sha256:c0b35c6807b33df353d93e3e860d9d4ad724c90370c409eff388ee0c0a95ebf6

Observation b57d74d1-f250-4cce-987a-fce997094e50 · outbound

This paper cites Coco-stuff: Thing and stuff classes in context.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Coco-stuff: Thing and stuff classes in context

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.994302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.203621Z digest=sha256:635447e89a4ea61f43baf648d996aa6a68ca1b3b0d4d94662a289251ef280500

Observation 5629e619-316f-4b4f-8320-a1d254c3f395 · outbound

This paper cites Cat-seg: Cost aggregation for open-vocabulary semantic segmentation.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Cat-seg: Cost aggregation for open-vocabulary semantic segmentation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.982575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.207203Z digest=sha256:6e30ca913e47e8c10231752b2d27d72f421e65b80767875f654a794d72c65b88

Observation 65787873-357b-40aa-86f2-0de5a299b0ab · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.210832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.210832Z digest=sha256:7fe4b58d8001d56b8824edd754595065c6a6faa425632248ecc0af5a655eec79

Observation a9d839f0-6ec9-4250-8664-a2eeb962113c · outbound

This paper cites Token Merging: Your ViT But Faster.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Token Merging: Your ViT But Faster

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.214627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.214627Z digest=sha256:e519af6e51bf6bc1c72eb34daa5fd079e32bb8b8a489e77a2bd1ba0b68849c5d

Observation 827f4e3a-4921-4742-850b-2ed579237af7 · outbound

This paper cites QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-16T05:53:23.431195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.218705Z digest=sha256:2cdc7d814d183ae5fd337f30b2ff7811c2122ea926c8ea588c6795554de08cd1

Observation a996f7f8-2fca-4bad-abdc-7b3b162d9d90 · outbound

This paper cites Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.222499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.222499Z digest=sha256:4c037ac2b6b4176798f441e87fafc3209b021356ff268cf468405bd8d209a8e2

Observation 36dc7495-b72b-4676-9c79-17a83e1ba0e0 · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.226484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.226484Z digest=sha256:4dfb68dcccc109e4c8e0927e6c9744c1eb846520ecc7b93c928e51538ca23c03

Observation d2a58e83-7e7c-489e-a5e7-d95904f35d84 · outbound

This paper cites Efficient large multi-modal models via visual context compression.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Efficient large multi-modal models via visual context compression

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.970400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.230431Z digest=sha256:bfa953b87318677bc28f0457283762fbed18040beab908cb4734e55889c74190

Observation 4b6e855d-f0e3-4462-937c-2e3a7a51e4be · outbound

This paper cites Matryoshka Query Transformer for Large Vision-Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Matryoshka Query Transformer for Large Vision-Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.233834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.233834Z digest=sha256:daebd546190c8c89d53388d004686dbd69823921245fdf6e3b92866d0ce1b929

Observation c971ad23-f87b-4597-ab6f-bc2d6d00ccfe · outbound

This paper cites MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.237184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.237184Z digest=sha256:9522f69ee6b100bb6e8ab239f3f7e4e19174630820fc442cfd582a287d1a000d

Observation a5618281-9002-4cab-8243-002686e3ef42 · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.240784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.240784Z digest=sha256:488b9fea479e02e67ae13c1a5ff5bd1c910ece0103b5488b5425a8c6d3fc4b8e

Observation d554a7df-a338-4878-9613-608e0b47a067 · outbound

This paper cites Vary: Scaling up the vision vocabulary for large vision-language model.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Vary: Scaling up the vision vocabulary for large vision-language model

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.955745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.244004Z digest=sha256:05cda21e4808f9f0cf542f4e70d5c7e9c89e49b6fd9ae514b0d684a13953ca43

Observation 94c4da0a-b7a8-4151-9222-30312bdfe980 · outbound

This paper cites Distilling large vision-language model with out-of-distribution generalizability.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Distilling large vision-language model with out-of-distribution generalizability

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.936213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.247576Z digest=sha256:c5fe2696529e4cd7769a13e5602111a80effe50ed8239b604678770129384604

Observation 973a8408-6289-41b6-b827-0f546a63f4da · outbound

This paper cites Rs5m and georsclip: A large scale vision-language dataset and a large vision-language model for remote sensing.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Rs5m and georsclip: A large scale vision-language dataset and a large vision-language model for remote sensing

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.924090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.251000Z digest=sha256:926f30388fba3a3274a477b224ae8bf88d24194bb5fa4dd435d1af9030839a30

Observation 1b76abfe-60ab-43d3-b62d-8f531ea618e4 · outbound

This paper cites Visual In-Context Learning for Large Vision-Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Visual In-Context Learning for Large Vision-Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.254383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.254383Z digest=sha256:5c0d2905dd596627ab6bbfbc388336411cb830e9f3152454ad71e7fe06b76bec

Observation 2801aa98-1897-4871-94ba-0997556ae668 · outbound

This paper cites Anomalygpt: Detecting industrial anomalies using large vision-language models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Anomalygpt: Detecting industrial anomalies using large vision-language models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.912006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.257827Z digest=sha256:a330778fd5076df270b55ed01bcbf552b866a6e1bffd96dda8b8e0c9fed3b017

Observation 1a0b1ba4-5a23-418e-8f31-8278e8dcf93f · outbound

This paper cites Matryoshka query transformer for large vision-language models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Matryoshka query transformer for large vision-language models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.900177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.261043Z digest=sha256:a2c6356dd21c128c46cbb8900b3165703d040ce216faaf1bcf8fa196a208c73d

Observation 787510c3-31c8-4478-bb6e-001b07875047 · outbound

This paper cites Pyramidclip: Hierarchical feature alignment for vision-language model pretraining.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Pyramidclip: Hierarchical feature alignment for vision-language model pretraining

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.889111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.263859Z digest=sha256:7f180aa5a46bb7b56a162c67c274e6b149c8953652bd334c0884823250fe27aa

Observation 3421e68b-421d-4b70-9fbe-95152307f7ad · outbound

This paper cites Matryoshka multimodal models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Matryoshka multimodal models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.877145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.267730Z digest=sha256:5a4cbacda81a6a601439c566f9e9bdd5206a1051607d8bbbea994e28dfc42fba

Observation 7ea3ea87-15dd-47b0-8c9a-030b198b93c0 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.271111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.271111Z digest=sha256:f02cf642c9e2150ad31d4aabb8d9821ae3a156d3e2c95e0ef726366d40241a91

Observation c0515a80-6d45-40ff-90d7-e42ade2f7460 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.274389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.274389Z digest=sha256:dd2cc976bc462b3bcb729fc419411dd0f9ceddcac41e97a3ccae5029885b494f

Observation ab44a47c-17d8-4fce-8e05-3e7632e2a063 · outbound

This paper cites Vision-language models for vision tasks: A survey.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Vision-language models for vision tasks: A survey

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.278220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.278220Z digest=sha256:2e219d02e4d2b832ddeb681e38dfc89a29375198eb41adc67f96ecd7c63b8abf

Observation baacb898-2f08-4b8e-ac46-a2cabc026a14 · outbound

This paper cites A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.281964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.281964Z digest=sha256:3d80562ce279c6b36bfeea346d1921e19e0b8461ecd682e314245b39a8ca1c0c

Observation e123e2f9-c737-4629-be61-731edebd03e0 · outbound

This paper cites A Survey of Vision-Language Pre-Trained Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning A Survey of Vision-Language Pre-Trained Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.286456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.286456Z digest=sha256:e1ffa1e5b4740dcad8ca70efcdc0cf8f21ad583c45cd3a32fcea4469dcb85f9d

Observation c852f6d2-61f0-4313-9e83-13b3a24245e0 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning A Survey on Hallucination in Large Vision-Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.290607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.290607Z digest=sha256:61ea980c3a35f2314bc44d152ce9c0c589b5077d885940d3617a4239591d15b6

Observation 51a3d3c5-7bed-4174-98dd-978f5cb96ff2 · outbound

This paper cites Vqa: Visual question answering.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Vqa: Visual question answering

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.855623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.294509Z digest=sha256:f70d36e3906d033d48532533db0877b888816a28e81f43fa4227e7426b598dcc

Observation 46e7a9cd-f51f-430d-ac1d-8d54fa3e8eb7 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Vizwiz grand challenge: Answering visual questions from blind people

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.298221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.298221Z digest=sha256:10e3920e4c98438a093549b7a407f4fcfb22f08e186681bfcd96e06402233f6f

Observation 103d7a38-3bd6-4b96-ba23-4a2a22f5886b · outbound

This paper cites Cider: Consensus-based image description evaluation.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Cider: Consensus-based image description evaluation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:23.301738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:53:23.301738Z digest=sha256:66821c5651063a28c19389219ee4fcfc26405d3c26d84048a6b7a5b98fb7e3fd

Observation 4aee89f1-4bfa-4a48-890f-db5498a783ab · outbound

This paper cites Fully convolutional networks for semantic segmentation.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning Fully convolutional networks for semantic segmentation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.831153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.305013Z digest=sha256:28dc71f304da6a6e3562854f328179b5745e9bfd51e280b0ef82594efe942b25

Observation 38b7fd2e-bc50-48a5-89fb-25dcfb9270f2 · outbound

This paper cites ℎ#"ℎ$"ℎ%.

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning ℎ#"ℎ$"ℎ%

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:53:23.819377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:53:23.309664Z digest=sha256:c69edc5ca809666cb4152ac75e4f4080c80beb6995b569f0a74dfa407e27a436

Pith citing papers

Observation 0a048ce3-2f65-47d8-9e99-1a2d10ec34de · inbound

Towards Modality Generalization: A Benchmark and Prospective Analysis cites this paper.

Towards Modality Generalization: A Benchmark and Prospective Analysis VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T04:55:17.207198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:55:17.207198Z digest=sha256:348f599153b799df5bc1529752eb0cc6bf9f9813a14a288f1f66e8fa4099a415

Observation 439f993c-bf43-46df-ab72-58020bbae29a · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:17.141525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:a7b3c394dd4432f248764d635c7b6b76ab23bdf8fa2eeade19de64d538e1fc22

Observation 977f4470-facd-449e-89d1-1b2e24913411 · inbound

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding cites this paper.

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:01.611015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T14:50:37.022338Z digest=sha256:6043d925584eda46a295c819c86fede17fdd25ed3d7d535be7bdf58c9d007a50