{"work":{"id":"44525076-312a-4259-b79c-134cd7eeb297","openalex_id":"https://openalex.org/W4382491206","doi":"10.48550/arxiv.2306.15195","arxiv_id":"2306.15195","raw_key":null,"title":"Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic","authors":null,"authors_text":"Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, Rui Zhao","year":2023,"venue":"cs.CV","abstract":"In human conversations, individuals can indicate relevant regions within a scene while addressing others. In turn, the other person can then respond by referring to specific regions if necessary. This natural referential ability in dialogue remains absent in current Multimodal Large Language Models (MLLMs). To fill this gap, this paper proposes an MLLM called Shikra, which can handle spatial coordinate inputs and outputs in natural language. Its architecture consists of a vision encoder, an alignment layer, and a LLM. It is designed to be straightforward and simple, without the need for extra vocabularies, position encoder, pre-/post-detection modules, or external plug-in models. All inputs and outputs are in natural language form. Referential dialogue is a superset of various vision-language (VL) tasks. Shikra can naturally handle location-related tasks like REC and PointQA, as well as conventional VL tasks such as Image Captioning and VQA. Experimental results showcase Shikra's promising performance. Furthermore, it enables numerous exciting applications, like providing mentioned objects' coordinates in chains of thoughts and comparing user-pointed regions similarities. Our code, model and dataset are accessed at https://github.com/shikras/shikra.","external_url":"https://arxiv.org/abs/2306.15195","cited_by_count":71,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2306.15195","created_at":"2026-05-09T06:20:38.492107+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic","render_title":"Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic"},"hub":{"state":{"work_id":"44525076-312a-4259-b79c-134cd7eeb297","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":107,"external_cited_by_count":71,"distinct_field_count":11,"first_pith_cited_at":"2023-05-05T17:59:46+00:00","last_pith_cited_at":"2026-07-07T17:58:33+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T10:49:25.727561+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":23},{"context_role":"baseline","n":4},{"context_role":"dataset","n":1}],"polarity_counts":[{"context_polarity":"background","n":22},{"context_polarity":"baseline","n":4},{"context_polarity":"support","n":1},{"context_polarity":"use_dataset","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic","claims":[{"claim_text":"In human conversations, individuals can indicate relevant regions within a scene while addressing others. In turn, the other person can then respond by referring to specific regions if necessary. This natural referential ability in dialogue remains absent in current Multimodal Large Language Models (MLLMs). To fill this gap, this paper proposes an MLLM called Shikra, which can handle spatial coordinate inputs and outputs in natural language. Its architecture consists of a vision encoder, an alignment layer, and a LLM. It is designed to be straightforward and simple, without the need for extra ","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"pabilities. Initially, models like Flamingo [3] and BLIP- 2 [33] aligned frozen vision encoders with LLMs for visual question answering. Subsequently, the LLaV A series [39- 41], MiniGPT-4 [96], and mPLUG-Owl series [86-88] in- troduced visual instruction tuning to improve instruction- following. Models such as VisionLLM [78], KOSMOS- 2 [57], Shikra [11], and the Qwen-VL series [5, 6, 77] en- hanced LMMs with visual grounding for tasks like region description and localization. InternVL [15] scal","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"research on MLLMs focuses on text content generation grounded in text prompts and image [20], [24]/video [25], [26]/audio [27]. Subsequent works have expanded the capa- bilities or the usage scenarios, including: (1) Better granular- ity support. Finer control on user prompts is developed to support specific regions through boxes [28] or a certain ob- ject through a click [29]. (2) Enhanced support on input and output modalities [30], [31], such as image, video, audio, and point cloud. Besides i","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"ception in MLLMs. MMStar argues that many existing evaluations overestimate multimodal competence because they admit language shortcuts or data leakage, while BLINK demonstrates that state- of-the-art MLLMs remain far below human performance on core perceptual tasks such as relative depth estimation, visual correspon- dence, and multi-view reasoning [2, 5, 12, 25, 37, 38]. More targeted studies further show that current LVLMs struggle to perceive basic geometric information and spatial relations","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"prompts to explicit visual tokens and feature-level fusion. This progression reflects a shift from language-side prompting to more grounded object modeling. Text as PromptA straightforward paradigm is to encode object cues as textualized spatial prompts and feed them into MLLMs through the language interface.Kosmos2[122] links text spans to image regions via location tokens,Shikra[ 24] treats spatial coordinates as natural-language inputs and outputs,Pink[201] represents region cues with textual","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"1 81.7 76.6 84.1 69.2 82.8 82.6 81.6 InternVL3.5-1B 85.4 89.7 80.2 77.7 85.5 69.5 81.9 81.6 81.4 InternVL3-2B [187] 89.8 92.6 86.4 84.0 89.2 76.5 87.6 87.2 86.7 InternVL3.5-2B 88.7 91.6 84.8 82.7 88.4 76.6 85.6 85.5 85.5 Qwen2.5-VL-3B [5] 89.1 91.7 84.0 82.4 88.0 74.1 85.2 85.7 85.0 InternVL3.5-4B 92.5 94.3 88.2 87.6 92.3 81.6 89.6 89.3 89.4 Shikra-7B [10] 87.0 90.6 80.2 81.6 87.4 72.1 82.3 82.2 82.9 CogVLM-Grounding [141] 92.8 94.8 89.0 88.7 92.9 83.4 89.8 90.8 90.3 Qwen2-VL-7B [138] 91.7 93.6 ","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Synthetic Image2Latex, Synthetic Handwritten OCR, Synthetic Infographic2Markdown KVQA [207], A-OKVQA [205], ViQuAE [123], iNaturalist2018 [237], MovieNet [95], ART500K [176], KonIQ-10K [91], IconQA [167], VisualMRC [225], ChemVLM Data [129], ScienceQA [165], AI2D [109],Knowledge TQA [110], Wikipedia-QA [81], Synthetic Multidisciplinary Knowledge / QA Objects365 [208], GRIT [278], RefCOCO [280], GPT4Gen-RD-BoxCoT [27], All-Seeing-V1 [251],Grounding All-Seeing-V2 [250], V3Det [243], TolokaVQA [236","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (22 contexts).","role_counts":[{"n":22,"context_role":"background"},{"n":4,"context_role":"baseline"},{"n":1,"context_role":"dataset"}]},"error":null,"updated_at":"2026-07-04T06:46:37.594433+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"dd80061e-c973-48f3-af57-2764aa5433c7","orcid":null,"display_name":"Keqin Chen"},{"id":"7b050f17-ee07-466c-b566-5602f325d055","orcid":null,"display_name":"Zhao Zhang"},{"id":"4f3ff8bc-2eef-4cea-9f7e-5689a5f89a39","orcid":null,"display_name":"Weili Zeng"},{"id":"7be9a633-1067-4c3f-b452-412339b25cf9","orcid":null,"display_name":"Richong Zhang"},{"id":"f594eee8-e9d2-41c5-9e75-8e58323846de","orcid":null,"display_name":"Feng Zhu"},{"id":"ffcd7d23-d2f6-4ae8-b9ac-f42695a19a5e","orcid":null,"display_name":"Rui Zhao"}]},"error":null,"updated_at":"2026-07-04T06:46:38.419497+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T18:20:10.243554+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":15},{"title":"Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond","work_id":"cbc2bb21-b6bb-46c0-80bf-107e195ffe10","shared_citers":15},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":12},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":12},{"title":"MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models","work_id":"806d2e73-71b3-4d56-87e0-39d571cc15d6","shared_citers":12},{"title":"MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models","work_id":"a7e3a737-e007-42bc-be89-c4d34c5ee071","shared_citers":11},{"title":"Kosmos-2: Grounding Multimodal Large Language Models to the World","work_id":"46e7f9e9-24c6-49af-b7d5-96159fa6f443","shared_citers":10},{"title":"LLaVA-OneVision: Easy Visual Task Transfer","work_id":"f5f2452b-f2a9-49ac-b38d-c76e18cdfe49","shared_citers":9},{"title":"Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context","work_id":"80e3e977-f1bb-4c83-8d0c-1ab0a0c5c3f1","shared_citers":8},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":8},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":8},{"title":"MMBench: Is Your Multi-modal Model an All-around Player?","work_id":"3b44943d-0f15-4228-9ac3-0e376f4f9ada","shared_citers":8},{"title":"arXiv preprint arXiv:2310.07704 (2023)","work_id":"0db2bf55-f649-4428-bcc9-688724b38e57","shared_citers":7},{"title":"Cogvlm: Visual expert for pretrained language models","work_id":"0d81fb99-dae6-46d2-8bed-c01dcbd7d7cf","shared_citers":7},{"title":"Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning","work_id":"ba8e8164-e47f-42d6-83ad-696cb57ee79a","shared_citers":7},{"title":"MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities","work_id":"7f3bac41-a0a5-4a7a-bfd2-526b616db745","shared_citers":7},{"title":"mplug-owl: Modularization empowers large lan- guage models with multimodality","work_id":"74a7deb6-48be-4132-9d35-882cc5870ebd","shared_citers":7},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":7},{"title":"BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models","work_id":"63d03f4d-15f4-4583-8286-913c19f02294","shared_citers":6},{"title":"Evaluating Object Hallucination in Large Vision-Language Models","work_id":"66d8ac3e-c134-4995-b528-550afa17586f","shared_citers":6},{"title":"Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling","work_id":"ee70bdc8-4656-4849-ada7-ce42a2278d70","shared_citers":6},{"title":"Ferret-v2: An improved baseline for referring and grounding with large language models","work_id":"c4fa27bd-d0f2-4bfc-985d-56cb8646c5f0","shared_citers":6},{"title":"Hallusionbench: An advanced diagnostic suite for entangled language halluci- nation & visual illusion in large vision-language models","work_id":"9e143dc0-9074-4760-b9db-ef0ac15dc3cc","shared_citers":6},{"title":"How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites","work_id":"3714835e-c5a6-4d7e-950c-be44670ed9e6","shared_citers":6}],"time_series":[{"n":5,"year":2023},{"n":6,"year":2024},{"n":2,"year":2025},{"n":21,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T18:20:01.602731+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T18:20:06.238375+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic","claims":[{"claim_text":"In human conversations, individuals can indicate relevant regions within a scene while addressing others. In turn, the other person can then respond by referring to specific regions if necessary. This natural referential ability in dialogue remains absent in current Multimodal Large Language Models (MLLMs). To fill this gap, this paper proposes an MLLM called Shikra, which can handle spatial coordinate inputs and outputs in natural language. Its architecture consists of a vision encoder, an alignment layer, and a LLM. It is designed to be straightforward and simple, without the need for extra ","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"pabilities. Initially, models like Flamingo [3] and BLIP- 2 [33] aligned frozen vision encoders with LLMs for visual question answering. Subsequently, the LLaV A series [39- 41], MiniGPT-4 [96], and mPLUG-Owl series [86-88] in- troduced visual instruction tuning to improve instruction- following. Models such as VisionLLM [78], KOSMOS- 2 [57], Shikra [11], and the Qwen-VL series [5, 6, 77] en- hanced LMMs with visual grounding for tasks like region description and localization. InternVL [15] scal","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"research on MLLMs focuses on text content generation grounded in text prompts and image [20], [24]/video [25], [26]/audio [27]. Subsequent works have expanded the capa- bilities or the usage scenarios, including: (1) Better granular- ity support. Finer control on user prompts is developed to support specific regions through boxes [28] or a certain ob- ject through a click [29]. (2) Enhanced support on input and output modalities [30], [31], such as image, video, audio, and point cloud. Besides i","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"ception in MLLMs. MMStar argues that many existing evaluations overestimate multimodal competence because they admit language shortcuts or data leakage, while BLINK demonstrates that state- of-the-art MLLMs remain far below human performance on core perceptual tasks such as relative depth estimation, visual correspon- dence, and multi-view reasoning [2, 5, 12, 25, 37, 38]. More targeted studies further show that current LVLMs struggle to perceive basic geometric information and spatial relations","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"prompts to explicit visual tokens and feature-level fusion. This progression reflects a shift from language-side prompting to more grounded object modeling. Text as PromptA straightforward paradigm is to encode object cues as textualized spatial prompts and feed them into MLLMs through the language interface.Kosmos2[122] links text spans to image regions via location tokens,Shikra[ 24] treats spatial coordinates as natural-language inputs and outputs,Pink[201] represents region cues with textual","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"1 81.7 76.6 84.1 69.2 82.8 82.6 81.6 InternVL3.5-1B 85.4 89.7 80.2 77.7 85.5 69.5 81.9 81.6 81.4 InternVL3-2B [187] 89.8 92.6 86.4 84.0 89.2 76.5 87.6 87.2 86.7 InternVL3.5-2B 88.7 91.6 84.8 82.7 88.4 76.6 85.6 85.5 85.5 Qwen2.5-VL-3B [5] 89.1 91.7 84.0 82.4 88.0 74.1 85.2 85.7 85.0 InternVL3.5-4B 92.5 94.3 88.2 87.6 92.3 81.6 89.6 89.3 89.4 Shikra-7B [10] 87.0 90.6 80.2 81.6 87.4 72.1 82.3 82.2 82.9 CogVLM-Grounding [141] 92.8 94.8 89.0 88.7 92.9 83.4 89.8 90.8 90.3 Qwen2-VL-7B [138] 91.7 93.6 ","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Synthetic Image2Latex, Synthetic Handwritten OCR, Synthetic Infographic2Markdown KVQA [207], A-OKVQA [205], ViQuAE [123], iNaturalist2018 [237], MovieNet [95], ART500K [176], KonIQ-10K [91], IconQA [167], VisualMRC [225], ChemVLM Data [129], ScienceQA [165], AI2D [109],Knowledge TQA [110], Wikipedia-QA [81], Synthetic Multidisciplinary Knowledge / QA Objects365 [208], GRIT [278], RefCOCO [280], GPT4Gen-RD-BoxCoT [27], All-Seeing-V1 [251],Grounding All-Seeing-V2 [250], V3Det [243], TolokaVQA [236","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (22 contexts).","role_counts":[{"n":22,"context_role":"background"},{"n":4,"context_role":"baseline"},{"n":1,"context_role":"dataset"}]},"error":null,"updated_at":"2026-07-04T06:46:37.589648+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic","claims":[{"claim_text":"In human conversations, individuals can indicate relevant regions within a scene while addressing others. In turn, the other person can then respond by referring to specific regions if necessary. This natural referential ability in dialogue remains absent in current Multimodal Large Language Models (MLLMs). To fill this gap, this paper proposes an MLLM called Shikra, which can handle spatial coordinate inputs and outputs in natural language. Its architecture consists of a vision encoder, an alignment layer, and a LLM. It is designed to be straightforward and simple, without the need for extra ","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T18:20:10.248593+00:00"}},"summary":{"title":"Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic","claims":[{"claim_text":"In human conversations, individuals can indicate relevant regions within a scene while addressing others. In turn, the other person can then respond by referring to specific regions if necessary. This natural referential ability in dialogue remains absent in current Multimodal Large Language Models (MLLMs). To fill this gap, this paper proposes an MLLM called Shikra, which can handle spatial coordinate inputs and outputs in natural language. Its architecture consists of a vision encoder, an alignment layer, and a LLM. It is designed to be straightforward and simple, without the need for extra ","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":15},{"title":"Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond","work_id":"cbc2bb21-b6bb-46c0-80bf-107e195ffe10","shared_citers":15},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":12},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":12},{"title":"MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models","work_id":"806d2e73-71b3-4d56-87e0-39d571cc15d6","shared_citers":12},{"title":"MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models","work_id":"a7e3a737-e007-42bc-be89-c4d34c5ee071","shared_citers":11},{"title":"Kosmos-2: Grounding Multimodal Large Language Models to the World","work_id":"46e7f9e9-24c6-49af-b7d5-96159fa6f443","shared_citers":10},{"title":"LLaVA-OneVision: Easy Visual Task Transfer","work_id":"f5f2452b-f2a9-49ac-b38d-c76e18cdfe49","shared_citers":9},{"title":"Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context","work_id":"80e3e977-f1bb-4c83-8d0c-1ab0a0c5c3f1","shared_citers":8},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":8},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":8},{"title":"MMBench: Is Your Multi-modal Model an All-around Player?","work_id":"3b44943d-0f15-4228-9ac3-0e376f4f9ada","shared_citers":8},{"title":"arXiv preprint arXiv:2310.07704 (2023)","work_id":"0db2bf55-f649-4428-bcc9-688724b38e57","shared_citers":7},{"title":"Cogvlm: Visual expert for pretrained language models","work_id":"0d81fb99-dae6-46d2-8bed-c01dcbd7d7cf","shared_citers":7},{"title":"Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning","work_id":"ba8e8164-e47f-42d6-83ad-696cb57ee79a","shared_citers":7},{"title":"MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities","work_id":"7f3bac41-a0a5-4a7a-bfd2-526b616db745","shared_citers":7},{"title":"mplug-owl: Modularization empowers large lan- guage models with multimodality","work_id":"74a7deb6-48be-4132-9d35-882cc5870ebd","shared_citers":7},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":7},{"title":"BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models","work_id":"63d03f4d-15f4-4583-8286-913c19f02294","shared_citers":6},{"title":"Evaluating Object Hallucination in Large Vision-Language Models","work_id":"66d8ac3e-c134-4995-b528-550afa17586f","shared_citers":6},{"title":"Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling","work_id":"ee70bdc8-4656-4849-ada7-ce42a2278d70","shared_citers":6},{"title":"Ferret-v2: An improved baseline for referring and grounding with large language models","work_id":"c4fa27bd-d0f2-4bfc-985d-56cb8646c5f0","shared_citers":6},{"title":"Hallusionbench: An advanced diagnostic suite for entangled language halluci- nation & visual illusion in large vision-language models","work_id":"9e143dc0-9074-4760-b9db-ef0ac15dc3cc","shared_citers":6},{"title":"How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites","work_id":"3714835e-c5a6-4d7e-950c-be44670ed9e6","shared_citers":6}],"time_series":[{"n":5,"year":2023},{"n":6,"year":2024},{"n":2,"year":2025},{"n":21,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"f594eee8-e9d2-41c5-9e75-8e58323846de","orcid":null,"display_name":"Feng Zhu","source":"manual","import_confidence":0.72},{"id":"dd80061e-c973-48f3-af57-2764aa5433c7","orcid":null,"display_name":"Keqin Chen","source":"manual","import_confidence":0.72},{"id":"7be9a633-1067-4c3f-b452-412339b25cf9","orcid":null,"display_name":"Richong Zhang","source":"manual","import_confidence":0.72},{"id":"ffcd7d23-d2f6-4ae8-b9ac-f42695a19a5e","orcid":null,"display_name":"Rui Zhao","source":"manual","import_confidence":0.72},{"id":"4f3ff8bc-2eef-4cea-9f7e-5689a5f89a39","orcid":null,"display_name":"Weili Zeng","source":"manual","import_confidence":0.72},{"id":"7b050f17-ee07-466c-b566-5602f325d055","orcid":null,"display_name":"Zhao Zhang","source":"manual","import_confidence":0.72}]}}