{"work":{"id":"b3d6fb46-4169-4a28-8f7e-2ca6774211da","openalex_id":"https://openalex.org/W1889081078","doi":"10.48550/arxiv.1504.00325","arxiv_id":"1504.00325","raw_key":null,"title":"Microsoft COCO Captions: Data Collection and Evaluation Server","authors":null,"authors_text":"Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollar","year":2015,"venue":"cs.CV","abstract":"In this paper we describe the Microsoft COCO Caption dataset and evaluation server. When completed, the dataset will contain over one and a half million captions describing over 330,000 images. For the training and validation images, five independent human generated captions will be provided. To ensure consistency in evaluation of automatic caption generation algorithms, an evaluation server is used. The evaluation server receives candidate captions and scores them using several popular metrics, including BLEU, METEOR, ROUGE and CIDEr. Instructions for using the evaluation server are provided.","external_url":"https://arxiv.org/abs/1504.00325","cited_by_count":1634,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"1504.00325","created_at":"2026-05-10T05:46:10.154706+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Microsoft COCO Captions: Data Collection and Evaluation Server","render_title":"Microsoft COCO Captions: Data Collection and Evaluation Server"},"hub":{"state":{"work_id":"b3d6fb46-4169-4a28-8f7e-2ca6774211da","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":108,"external_cited_by_count":1634,"distinct_field_count":10,"first_pith_cited_at":"2019-07-22T11:21:08+00:00","last_pith_cited_at":"2026-07-07T23:17:26+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T17:09:27.041533+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":15},{"context_role":"dataset","n":15}],"polarity_counts":[{"context_polarity":"background","n":16},{"context_polarity":"use_dataset","n":13},{"context_polarity":"unclear","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Microsoft COCO Captions: Data Collection and Evaluation Server","claims":[{"claim_text":"In this paper we describe the Microsoft COCO Caption dataset and evaluation server. When completed, the dataset will contain over one and a half million captions describing over 330,000 images. For the training and validation images, five independent human generated captions will be provided. To ensure consistency in evaluation of automatic caption generation algorithms, an evaluation server is used. The evaluation server receives candidate captions and scores them using several popular metrics, including BLEU, METEOR, ROUGE and CIDEr. Instructions for using the evaluation server are provided.","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"(3)What is the im- pact of the Conditional Ground-Truth Injection (CGI) mechanism on learning dynamics? 5.1 Experimental Setup Benchmark Datasets.To evaluate S-GRPO's adaptability and performance, we conduct experiments on three benchmark datasets spanning diverse domains: image captioning, geometric problem- solving, and medical X-ray diagnostics. • COCO Caption [6]: A widely used dataset for image captioning, containing 123,287 images. We use the original training and test splits of the datase","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"OXYECOM BENCH , indicating that insufﬁcient e-commerce-speciﬁc knowledge infusion can mute the advantages of advanced general-purpose models. 2 Related Work 2.1 General Multimodal Benchmarks The rapid progress of multimodal large language models (MLLMs) has been mirrored by a rich landscape of evaluation suites, ranging from early perception and single-hop reasoning benchmarks (VQA [8], COCO Captions [9], OK-VQA [10], TextVQA [11]) and comprehensive multi-task suites (MME [ 12], MMBench [ 13], M","claim_type":"background","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"sertion scores indicate that the highlighted regions have a stronger causal effect on the model's responses. We conduct experiments using Qwen2.5-VL-3B as the VLM, and compare our method with several baselines, in- cluding CAM [97], Grad-CAM [57], raw attention, atten- tion rollout [1], ATTN-LRP [2], and TAM [40]. The eval- uation is performed on three datasets, namely the COCO Caption dataset [13], GranDf [54], and OpenPSG [98]. For each dataset, we sample 1k images for evaluation. The re- sult","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"This approach allows us to benefit from readily-available image instruction data, creating a system capable of handling both images and videos with shared spatial perception and reasoning capacity. Stage1: Alignment. To strike a balance between training convergency and efficiency we introduce 25M vision-text pairs for one epoch of fine-tuning, The data consists 10M video-text pairs from WebVid-10M, and 15M image-text pairs from COCO Caption [ 6], Visual Genome [ 17], SBU Captions [31], CC3M [35]","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"High-Quality Bilingual Dataset Pre-training Dataset. The pre-training dataset utilized in our InternVL 1.5 encompasses a diverse range of pub- licly accessible sources. We provide an overview of these datasets in Table 1a. These datasets span multi- 4 task ratio dataset Laion-EN (en) [93], Laion-ZH (zh) [93], COYO (zh) [10],Captioning 53.9% GRIT (zh) [90], COCO (en) [17], TextCaps (en) [99] Objects365 (en&zh) [97], GRIT (en&zh) [90],Detection 5.2% All-Seeing (en&zh) [119] Wukong-OCR (zh) [29], L","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"bination with other NLP metrics [50, 60, 42, 65, 58, 43]. BLEU has become a standard metric in the field of NLP due to its efficient and fast calculation capabilities, especially for comparisons at the corpus level. However, its reliance on surface-level n-gram overlaps means it often fails to assess context or fluency effectively, leading to poor performance on individual sentences [116]. As a result, in tasks like image captioning, it is commonly paired with other techniques to address these l","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Microsoft COCO Captions: Data Collection and Evaluation Server because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (15 contexts).","role_counts":[{"n":15,"context_role":"background"},{"n":15,"context_role":"dataset"}]},"error":null,"updated_at":"2026-07-03T11:14:06.103519+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"f871a5ec-7966-46e3-b89d-8aa9f74186a1","orcid":null,"display_name":"Xinlei Chen"},{"id":"2228b505-7e88-4b74-9438-5b11d5648e2b","orcid":null,"display_name":"Hao Fang"},{"id":"00e2756c-d70b-49a9-b019-ad42b25d5fab","orcid":null,"display_name":"Tsung-Yi Lin"},{"id":"a3d49b33-1d94-41ad-9709-ce6f77da6d04","orcid":null,"display_name":"Ramakrishna Vedantam"},{"id":"a0d85cd0-cbd5-41ea-a62f-355c2ef0f812","orcid":null,"display_name":"Saurabh Gupta"},{"id":"2fbd3b65-8775-4dd0-9095-feb61c7c9159","orcid":null,"display_name":"Piotr Dollar"}]},"error":null,"updated_at":"2026-07-03T11:14:06.539393+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-16T13:18:57.417704+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":13},{"title":"Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond","work_id":"cbc2bb21-b6bb-46c0-80bf-107e195ffe10","shared_citers":13},{"title":"BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models","work_id":"63d03f4d-15f4-4583-8286-913c19f02294","shared_citers":11},{"title":"MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models","work_id":"806d2e73-71b3-4d56-87e0-39d571cc15d6","shared_citers":11},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":10},{"title":"MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models","work_id":"a7e3a737-e007-42bc-be89-c4d34c5ee071","shared_citers":10},{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"ed240a10-5b19-406c-baa5-30803f465785","shared_citers":9},{"title":"EVA-CLIP: Improved Training Techniques for CLIP at Scale","work_id":"0c16c250-fd0f-446a-bbb0-ea8dd0ba5ccd","shared_citers":8},{"title":"MMBench: Is Your Multi-modal Model an All-around Player?","work_id":"3b44943d-0f15-4228-9ac3-0e376f4f9ada","shared_citers":8},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":8},{"title":"Visual Instruction Tuning","work_id":"68be622d-a6dc-4a13-82de-e3054a3dc509","shared_citers":8},{"title":"Evaluating Object Hallucination in Large Vision-Language Models","work_id":"66d8ac3e-c134-4995-b528-550afa17586f","shared_citers":7},{"title":"LLaVA-OneVision: Easy Visual Task Transfer","work_id":"f5f2452b-f2a9-49ac-b38d-c76e18cdfe49","shared_citers":7},{"title":"Otter: A Multi-Modal Model with In-Context Instruction Tuning","work_id":"33cb3a7a-6091-48db-a246-802bbb055f43","shared_citers":7},{"title":"PaLM: Scaling Language Modeling with Pathways","work_id":"a94f3ef7-2c49-4445-93fe-6ec16aafd966","shared_citers":7},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":7},{"title":"SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension","work_id":"23881ff0-b851-474c-8712-90744cc07a3a","shared_citers":7},{"title":"ShareGPT4V: Improving Large Multi-Modal Models with Better Captions","work_id":"90e2b26a-3d27-4567-86b5-929b582a8034","shared_citers":7},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":6},{"title":"Florence: A New Foundation Model for Computer Vision","work_id":"99823072-36a8-4b10-9ef5-a7f91da74650","shared_citers":6},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":6},{"title":"Improved Baselines with Visual Instruction Tuning","work_id":"5baeaa33-5986-44a3-85a4-fcabd6fc1e8d","shared_citers":6},{"title":"Kosmos-2: Grounding Multimodal Large Language Models to the World","work_id":"46e7f9e9-24c6-49af-b7d5-96159fa6f443","shared_citers":6},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":6}],"time_series":[{"n":1,"year":2019},{"n":5,"year":2022},{"n":11,"year":2023},{"n":6,"year":2024},{"n":21,"year":2026}],"dependency_candidates":[{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"ZAYA1-VL-8B Technical Report","primary_cat":"cs.CV","context_text":"[92] Mustafa Shukor, Maxime Oquab, Ishan Misra, and Enrico Fini. Scaling laws for native multimodal models. arXiv preprint arXiv:2504.07951, 2025. [93] Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Doll'ar, and C. Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server.ArXiv, abs/1504.00325, 2015. [94] Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hy- pernymed, image alt-text dataset for automatic image captioning. In Iryna Gurevych and Yusuke Miyao, editors,Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556-2565, Melbourne, Australia,","citing_arxiv_id":"2605.08560"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"Empowering Video Translation using Multimodal Large Language Models","primary_cat":"cs.CV","context_text":"modeling for long-term temporal reasoning [186], and audio language modeling based on Whisper-like or speech- oriented frameworks [187]. Systems such as Flamingo-style 4 TABLE I Comparative evaluation on standard open-end zero-shot VideoQA benchmarks. This table shows the accuracy scores for all models on MSVD-QA, MSRVTT-QA, and ActivityNet-QA. Model MSVD-QA [170] MSRVTT-QA [171] ActivityNet-QA [172] FAVOR [173] 67.8 59.3 / Dolphin [174] 72.7 62.6 49.1 OneLLM [175] 56.5 53.8 / Video-SALMONN [176] 67.9 59.5 / AV-LLM [56] 67.3 53.7 47.2 GroundingGPT [63] 67.8 51.6 44.7 Macaw-LLM [177] 42.1 25.5 14.5 PandaGPT [178] 46.7 23.7 11.2 PG-Video-LLaVA [62] 64.1 51.6 39.9 VaQuitA [11] 74.6 68.6 48.8 LSTP [48] 71.4 57.","citing_arxiv_id":"2604.11283"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward","primary_cat":"cs.CV","context_text":"sertion scores indicate that the highlighted regions have a stronger causal effect on the model's responses. We conduct experiments using Qwen2.5-VL-3B as the VLM, and compare our method with several baselines, in- cluding CAM [97], Grad-CAM [57], raw attention, atten- tion rollout [1], ATTN-LRP [2], and TAM [40]. The eval- uation is performed on three datasets, namely the COCO Caption dataset [13], GranDf [54], and OpenPSG [98]. For each dataset, we sample 1k images for evaluation. The re- sults are shown in Table 1. Our method attains faithful- ness on par with, or better than, the baselines. For exam- ple, on COCO Captions, the deletion score is lower than the second-best approach by 5.46%, 4.77%, and 2.57% at 5%, 15%, and 30% pixel perturbation, respectively.","citing_arxiv_id":"2604.04500"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"Emu3: Next-Token Prediction is All You Need","primary_cat":"cs.CV","context_text":"tokenize images, text, and videos into a discrete space, and jointly train a single transformer from scratch on a mix of multimodal sequences. Emu3 achieves state-of-the-art performance compared to well-established task-specific models in gen- eration and perception tasks. Emu3 outperforms the flagship Stable Diffusion model,i.e., SDXL [66], in both the human evaluation and the public text-to-image benchmarks such as MSCOCO-30K [15], GenEval [26], T2I-CompBench [ 32], and DPG-Bench [ 31]. For vision-language understanding, Emu3 competes with the popular vision-language model, i.e., LLaV A-1.6 [56], on a series of public vision-language benchmarks, including SEED-Bench [45], RealWorldQA [91], OCRBench [59], etc. Emu3 is capable of generating videos. Unlike Sora [8] that employs the video diffusion model to","citing_arxiv_id":"2409.18869"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites","primary_cat":"cs.CV","context_text":"High-Quality Bilingual Dataset Pre-training Dataset. The pre-training dataset utilized in our InternVL 1.5 encompasses a diverse range of pub- licly accessible sources. We provide an overview of these datasets in Table 1a. These datasets span multi- 4 task ratio dataset Laion-EN (en) [93], Laion-ZH (zh) [93], COYO (zh) [10],Captioning 53.9% GRIT (zh) [90], COCO (en) [17], TextCaps (en) [99] Objects365 (en&zh) [97], GRIT (en&zh) [90],Detection 5.2% All-Seeing (en&zh) [119] Wukong-OCR (zh) [29], LaionCOCO-OCR (en) [94],OCR (large) 32.0% Common Crawl PDF (en&zh) MMC-Inst (en) [61], LSVT (zh) [105], ST-VQA (en) [9] RCTW-17 (zh) [98], ReCTs (zh) [137], ArT (en&zh) [19], SynthDoG (en&zh) [41], COCO-Text (en) [114], ChartQA (en) [81], CTW (zh) [134], DocVQA (en) [82],","citing_arxiv_id":"2404.16821"},{"n":1,"role":"dataset","polarity":"background","paper_title":"ShareGPT4V: Improving Large Multi-Modal Models with Better Captions","primary_cat":"cs.CV","context_text":"Inthedistance,treescanbeseen,addingatouchofnaturetothisman-madesetting.Theimageisasnapshotofeverydaylifeatatrainstation,capturingbothitsroutineoperationsanditsinherentcharm. (a) Comparison of Captions' Quality(b) Comparison of Performance Figure 1. (a) We showcase a comparison between the caption in our proposed ShareGPT4V dataset and those utilized by recent large multi-modal models (LMMs). Unlike COCO-Caption [7] involves brief human-made captions on the main subject. LLaV A-Instruct [31] combines human-made captions, bounding boxes, and GPT4 [39] to 'imagine' the image details, which leads to inevitable er- ror/hallucination description (marked in red). Our approach involves feeding carefully designed prompts along with images directly into the advanced GPT4-Vision [40] and the descriptions are more detailed and accurate (marked in blue).","citing_arxiv_id":"2311.12793"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"VideoChat: Chat-Centric Video Understanding","primary_cat":"cs.CV","context_text":"This approach allows us to benefit from readily-available image instruction data, creating a system capable of handling both images and videos with shared spatial perception and reasoning capacity. Stage1: Alignment. To strike a balance between training convergency and efficiency we introduce 25M vision-text pairs for one epoch of fine-tuning, The data consists 10M video-text pairs from WebVid-10M, and 15M image-text pairs from COCO Caption [ 6], Visual Genome [ 17], SBU Captions [31], CC3M [35] and CC12M [4]. The input prompts for LLMs are as followed: • \"###Human: <Video>video_embed</Video> video_instruction ###Assistant: \" • \"###Human: <Image>image_embed</Image> image_instruction ###Assistant: \" The video_embed and image_embed are the output from the token interface. Meanwhile the video_instruction and image_instruction provide concise video and image descriptions","citing_arxiv_id":"2305.06355"}]},"error":null,"updated_at":"2026-05-16T13:18:52.427760+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-16T13:18:55.522697+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Microsoft COCO Captions: Data Collection and Evaluation Server","claims":[{"claim_text":"In this paper we describe the Microsoft COCO Caption dataset and evaluation server. When completed, the dataset will contain over one and a half million captions describing over 330,000 images. For the training and validation images, five independent human generated captions will be provided. To ensure consistency in evaluation of automatic caption generation algorithms, an evaluation server is used. The evaluation server receives candidate captions and scores them using several popular metrics, including BLEU, METEOR, ROUGE and CIDEr. Instructions for using the evaluation server are provided.","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"(3)What is the im- pact of the Conditional Ground-Truth Injection (CGI) mechanism on learning dynamics? 5.1 Experimental Setup Benchmark Datasets.To evaluate S-GRPO's adaptability and performance, we conduct experiments on three benchmark datasets spanning diverse domains: image captioning, geometric problem- solving, and medical X-ray diagnostics. • COCO Caption [6]: A widely used dataset for image captioning, containing 123,287 images. We use the original training and test splits of the datase","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"OXYECOM BENCH , indicating that insufﬁcient e-commerce-speciﬁc knowledge infusion can mute the advantages of advanced general-purpose models. 2 Related Work 2.1 General Multimodal Benchmarks The rapid progress of multimodal large language models (MLLMs) has been mirrored by a rich landscape of evaluation suites, ranging from early perception and single-hop reasoning benchmarks (VQA [8], COCO Captions [9], OK-VQA [10], TextVQA [11]) and comprehensive multi-task suites (MME [ 12], MMBench [ 13], M","claim_type":"background","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"sertion scores indicate that the highlighted regions have a stronger causal effect on the model's responses. We conduct experiments using Qwen2.5-VL-3B as the VLM, and compare our method with several baselines, in- cluding CAM [97], Grad-CAM [57], raw attention, atten- tion rollout [1], ATTN-LRP [2], and TAM [40]. The eval- uation is performed on three datasets, namely the COCO Caption dataset [13], GranDf [54], and OpenPSG [98]. For each dataset, we sample 1k images for evaluation. The re- sult","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"This approach allows us to benefit from readily-available image instruction data, creating a system capable of handling both images and videos with shared spatial perception and reasoning capacity. Stage1: Alignment. To strike a balance between training convergency and efficiency we introduce 25M vision-text pairs for one epoch of fine-tuning, The data consists 10M video-text pairs from WebVid-10M, and 15M image-text pairs from COCO Caption [ 6], Visual Genome [ 17], SBU Captions [31], CC3M [35]","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"High-Quality Bilingual Dataset Pre-training Dataset. The pre-training dataset utilized in our InternVL 1.5 encompasses a diverse range of pub- licly accessible sources. We provide an overview of these datasets in Table 1a. These datasets span multi- 4 task ratio dataset Laion-EN (en) [93], Laion-ZH (zh) [93], COYO (zh) [10],Captioning 53.9% GRIT (zh) [90], COCO (en) [17], TextCaps (en) [99] Objects365 (en&zh) [97], GRIT (en&zh) [90],Detection 5.2% All-Seeing (en&zh) [119] Wukong-OCR (zh) [29], L","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"bination with other NLP metrics [50, 60, 42, 65, 58, 43]. BLEU has become a standard metric in the field of NLP due to its efficient and fast calculation capabilities, especially for comparisons at the corpus level. However, its reliance on surface-level n-gram overlaps means it often fails to assess context or fluency effectively, leading to poor performance on individual sentences [116]. As a result, in tasks like image captioning, it is commonly paired with other techniques to address these l","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Microsoft COCO Captions: Data Collection and Evaluation Server because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (15 contexts).","role_counts":[{"n":15,"context_role":"background"},{"n":15,"context_role":"dataset"}]},"error":null,"updated_at":"2026-07-03T11:14:06.105905+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Microsoft COCO Captions: Data Collection and Evaluation Server","claims":[{"claim_text":"In this paper we describe the Microsoft COCO Caption dataset and evaluation server. When completed, the dataset will contain over one and a half million captions describing over 330,000 images. For the training and validation images, five independent human generated captions will be provided. To ensure consistency in evaluation of automatic caption generation algorithms, an evaluation server is used. The evaluation server receives candidate captions and scores them using several popular metrics, including BLEU, METEOR, ROUGE and CIDEr. Instructions for using the evaluation server are provided.","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"sertion scores indicate that the highlighted regions have a stronger causal effect on the model's responses. We conduct experiments using Qwen2.5-VL-3B as the VLM, and compare our method with several baselines, in- cluding CAM [97], Grad-CAM [57], raw attention, atten- tion rollout [1], ATTN-LRP [2], and TAM [40]. The eval- uation is performed on three datasets, namely the COCO Caption dataset [13], GranDf [54], and OpenPSG [98]. For each dataset, we sample 1k images for evaluation. The re- sult","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"This approach allows us to benefit from readily-available image instruction data, creating a system capable of handling both images and videos with shared spatial perception and reasoning capacity. Stage1: Alignment. To strike a balance between training convergency and efficiency we introduce 25M vision-text pairs for one epoch of fine-tuning, The data consists 10M video-text pairs from WebVid-10M, and 15M image-text pairs from COCO Caption [ 6], Visual Genome [ 17], SBU Captions [31], CC3M [35]","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"High-Quality Bilingual Dataset Pre-training Dataset. The pre-training dataset utilized in our InternVL 1.5 encompasses a diverse range of pub- licly accessible sources. We provide an overview of these datasets in Table 1a. These datasets span multi- 4 task ratio dataset Laion-EN (en) [93], Laion-ZH (zh) [93], COYO (zh) [10],Captioning 53.9% GRIT (zh) [90], COCO (en) [17], TextCaps (en) [99] Objects365 (en&zh) [97], GRIT (en&zh) [90],Detection 5.2% All-Seeing (en&zh) [119] Wukong-OCR (zh) [29], L","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"[92] Mustafa Shukor, Maxime Oquab, Ishan Misra, and Enrico Fini. Scaling laws for native multimodal models. arXiv preprint arXiv:2504.07951, 2025. [93] Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Doll'ar, and C. Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server.ArXiv, abs/1504.00325, 2015. [94] Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hy- pernymed, image alt-text dataset for au","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Inthedistance,treescanbeseen,addingatouchofnaturetothisman-madesetting.Theimageisasnapshotofeverydaylifeatatrainstation,capturingbothitsroutineoperationsanditsinherentcharm. (a) Comparison of Captions' Quality(b) Comparison of Performance Figure 1. (a) We showcase a comparison between the caption in our proposed ShareGPT4V dataset and those utilized by recent large multi-modal models (LMMs). Unlike COCO-Caption [7] involves brief human-made captions on the main subject. LLaV A-Instruct [31] comb","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"modeling for long-term temporal reasoning [186], and audio language modeling based on Whisper-like or speech- oriented frameworks [187]. Systems such as Flamingo-style 4 TABLE I Comparative evaluation on standard open-end zero-shot VideoQA benchmarks. This table shows the accuracy scores for all models on MSVD-QA, MSRVTT-QA, and ActivityNet-QA. Model MSVD-QA [170] MSRVTT-QA [171] ActivityNet-QA [172] FAVOR [173] 67.8 59.3 / Dolphin [174] 72.7 62.6 49.1 OneLLM [175] 56.5 53.8 / Video-SALMONN [176","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Microsoft COCO Captions: Data Collection and Evaluation Server because it crossed a citation-hub threshold. Current citing contexts most often use it as dataset evidence (7 contexts).","role_counts":[{"n":7,"context_role":"dataset"},{"n":4,"context_role":"background"}]},"error":null,"updated_at":"2026-05-16T13:18:52.434023+00:00"}},"summary":{"title":"Microsoft COCO Captions: Data Collection and Evaluation Server","claims":[{"claim_text":"In this paper we describe the Microsoft COCO Caption dataset and evaluation server. When completed, the dataset will contain over one and a half million captions describing over 330,000 images. For the training and validation images, five independent human generated captions will be provided. To ensure consistency in evaluation of automatic caption generation algorithms, an evaluation server is used. The evaluation server receives candidate captions and scores them using several popular metrics, including BLEU, METEOR, ROUGE and CIDEr. Instructions for using the evaluation server are provided.","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"sertion scores indicate that the highlighted regions have a stronger causal effect on the model's responses. We conduct experiments using Qwen2.5-VL-3B as the VLM, and compare our method with several baselines, in- cluding CAM [97], Grad-CAM [57], raw attention, atten- tion rollout [1], ATTN-LRP [2], and TAM [40]. The eval- uation is performed on three datasets, namely the COCO Caption dataset [13], GranDf [54], and OpenPSG [98]. For each dataset, we sample 1k images for evaluation. The re- sult","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"This approach allows us to benefit from readily-available image instruction data, creating a system capable of handling both images and videos with shared spatial perception and reasoning capacity. Stage1: Alignment. To strike a balance between training convergency and efficiency we introduce 25M vision-text pairs for one epoch of fine-tuning, The data consists 10M video-text pairs from WebVid-10M, and 15M image-text pairs from COCO Caption [ 6], Visual Genome [ 17], SBU Captions [31], CC3M [35]","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"High-Quality Bilingual Dataset Pre-training Dataset. The pre-training dataset utilized in our InternVL 1.5 encompasses a diverse range of pub- licly accessible sources. We provide an overview of these datasets in Table 1a. These datasets span multi- 4 task ratio dataset Laion-EN (en) [93], Laion-ZH (zh) [93], COYO (zh) [10],Captioning 53.9% GRIT (zh) [90], COCO (en) [17], TextCaps (en) [99] Objects365 (en&zh) [97], GRIT (en&zh) [90],Detection 5.2% All-Seeing (en&zh) [119] Wukong-OCR (zh) [29], L","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"[92] Mustafa Shukor, Maxime Oquab, Ishan Misra, and Enrico Fini. Scaling laws for native multimodal models. arXiv preprint arXiv:2504.07951, 2025. [93] Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Doll'ar, and C. Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server.ArXiv, abs/1504.00325, 2015. [94] Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hy- pernymed, image alt-text dataset for au","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Inthedistance,treescanbeseen,addingatouchofnaturetothisman-madesetting.Theimageisasnapshotofeverydaylifeatatrainstation,capturingbothitsroutineoperationsanditsinherentcharm. (a) Comparison of Captions' Quality(b) Comparison of Performance Figure 1. (a) We showcase a comparison between the caption in our proposed ShareGPT4V dataset and those utilized by recent large multi-modal models (LMMs). Unlike COCO-Caption [7] involves brief human-made captions on the main subject. LLaV A-Instruct [31] comb","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"modeling for long-term temporal reasoning [186], and audio language modeling based on Whisper-like or speech- oriented frameworks [187]. Systems such as Flamingo-style 4 TABLE I Comparative evaluation on standard open-end zero-shot VideoQA benchmarks. This table shows the accuracy scores for all models on MSVD-QA, MSRVTT-QA, and ActivityNet-QA. Model MSVD-QA [170] MSRVTT-QA [171] ActivityNet-QA [172] FAVOR [173] 67.8 59.3 / Dolphin [174] 72.7 62.6 49.1 OneLLM [175] 56.5 53.8 / Video-SALMONN [176","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Microsoft COCO Captions: Data Collection and Evaluation Server because it crossed a citation-hub threshold. Current citing contexts most often use it as dataset evidence (7 contexts).","role_counts":[{"n":7,"context_role":"dataset"},{"n":4,"context_role":"background"}]},"graph":{"co_cited":[{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":13},{"title":"Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond","work_id":"cbc2bb21-b6bb-46c0-80bf-107e195ffe10","shared_citers":13},{"title":"BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models","work_id":"63d03f4d-15f4-4583-8286-913c19f02294","shared_citers":11},{"title":"MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models","work_id":"806d2e73-71b3-4d56-87e0-39d571cc15d6","shared_citers":11},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":10},{"title":"MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models","work_id":"a7e3a737-e007-42bc-be89-c4d34c5ee071","shared_citers":10},{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"ed240a10-5b19-406c-baa5-30803f465785","shared_citers":9},{"title":"EVA-CLIP: Improved Training Techniques for CLIP at Scale","work_id":"0c16c250-fd0f-446a-bbb0-ea8dd0ba5ccd","shared_citers":8},{"title":"MMBench: Is Your Multi-modal Model an All-around Player?","work_id":"3b44943d-0f15-4228-9ac3-0e376f4f9ada","shared_citers":8},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":8},{"title":"Visual Instruction Tuning","work_id":"68be622d-a6dc-4a13-82de-e3054a3dc509","shared_citers":8},{"title":"Evaluating Object Hallucination in Large Vision-Language Models","work_id":"66d8ac3e-c134-4995-b528-550afa17586f","shared_citers":7},{"title":"LLaVA-OneVision: Easy Visual Task Transfer","work_id":"f5f2452b-f2a9-49ac-b38d-c76e18cdfe49","shared_citers":7},{"title":"Otter: A Multi-Modal Model with In-Context Instruction Tuning","work_id":"33cb3a7a-6091-48db-a246-802bbb055f43","shared_citers":7},{"title":"PaLM: Scaling Language Modeling with Pathways","work_id":"a94f3ef7-2c49-4445-93fe-6ec16aafd966","shared_citers":7},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":7},{"title":"SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension","work_id":"23881ff0-b851-474c-8712-90744cc07a3a","shared_citers":7},{"title":"ShareGPT4V: Improving Large Multi-Modal Models with Better Captions","work_id":"90e2b26a-3d27-4567-86b5-929b582a8034","shared_citers":7},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":6},{"title":"Florence: A New Foundation Model for Computer Vision","work_id":"99823072-36a8-4b10-9ef5-a7f91da74650","shared_citers":6},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":6},{"title":"Improved Baselines with Visual Instruction Tuning","work_id":"5baeaa33-5986-44a3-85a4-fcabd6fc1e8d","shared_citers":6},{"title":"Kosmos-2: Grounding Multimodal Large Language Models to the World","work_id":"46e7f9e9-24c6-49af-b7d5-96159fa6f443","shared_citers":6},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":6}],"time_series":[{"n":1,"year":2019},{"n":5,"year":2022},{"n":11,"year":2023},{"n":6,"year":2024},{"n":21,"year":2026}],"dependency_candidates":[{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"ZAYA1-VL-8B Technical Report","primary_cat":"cs.CV","context_text":"[92] Mustafa Shukor, Maxime Oquab, Ishan Misra, and Enrico Fini. Scaling laws for native multimodal models. arXiv preprint arXiv:2504.07951, 2025. [93] Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Doll'ar, and C. Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server.ArXiv, abs/1504.00325, 2015. [94] Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hy- pernymed, image alt-text dataset for automatic image captioning. In Iryna Gurevych and Yusuke Miyao, editors,Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556-2565, Melbourne, Australia,","citing_arxiv_id":"2605.08560"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"Empowering Video Translation using Multimodal Large Language Models","primary_cat":"cs.CV","context_text":"modeling for long-term temporal reasoning [186], and audio language modeling based on Whisper-like or speech- oriented frameworks [187]. Systems such as Flamingo-style 4 TABLE I Comparative evaluation on standard open-end zero-shot VideoQA benchmarks. This table shows the accuracy scores for all models on MSVD-QA, MSRVTT-QA, and ActivityNet-QA. Model MSVD-QA [170] MSRVTT-QA [171] ActivityNet-QA [172] FAVOR [173] 67.8 59.3 / Dolphin [174] 72.7 62.6 49.1 OneLLM [175] 56.5 53.8 / Video-SALMONN [176] 67.9 59.5 / AV-LLM [56] 67.3 53.7 47.2 GroundingGPT [63] 67.8 51.6 44.7 Macaw-LLM [177] 42.1 25.5 14.5 PandaGPT [178] 46.7 23.7 11.2 PG-Video-LLaVA [62] 64.1 51.6 39.9 VaQuitA [11] 74.6 68.6 48.8 LSTP [48] 71.4 57.","citing_arxiv_id":"2604.11283"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward","primary_cat":"cs.CV","context_text":"sertion scores indicate that the highlighted regions have a stronger causal effect on the model's responses. We conduct experiments using Qwen2.5-VL-3B as the VLM, and compare our method with several baselines, in- cluding CAM [97], Grad-CAM [57], raw attention, atten- tion rollout [1], ATTN-LRP [2], and TAM [40]. The eval- uation is performed on three datasets, namely the COCO Caption dataset [13], GranDf [54], and OpenPSG [98]. For each dataset, we sample 1k images for evaluation. The re- sults are shown in Table 1. Our method attains faithful- ness on par with, or better than, the baselines. For exam- ple, on COCO Captions, the deletion score is lower than the second-best approach by 5.46%, 4.77%, and 2.57% at 5%, 15%, and 30% pixel perturbation, respectively.","citing_arxiv_id":"2604.04500"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"Emu3: Next-Token Prediction is All You Need","primary_cat":"cs.CV","context_text":"tokenize images, text, and videos into a discrete space, and jointly train a single transformer from scratch on a mix of multimodal sequences. Emu3 achieves state-of-the-art performance compared to well-established task-specific models in gen- eration and perception tasks. Emu3 outperforms the flagship Stable Diffusion model,i.e., SDXL [66], in both the human evaluation and the public text-to-image benchmarks such as MSCOCO-30K [15], GenEval [26], T2I-CompBench [ 32], and DPG-Bench [ 31]. For vision-language understanding, Emu3 competes with the popular vision-language model, i.e., LLaV A-1.6 [56], on a series of public vision-language benchmarks, including SEED-Bench [45], RealWorldQA [91], OCRBench [59], etc. Emu3 is capable of generating videos. Unlike Sora [8] that employs the video diffusion model to","citing_arxiv_id":"2409.18869"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites","primary_cat":"cs.CV","context_text":"High-Quality Bilingual Dataset Pre-training Dataset. The pre-training dataset utilized in our InternVL 1.5 encompasses a diverse range of pub- licly accessible sources. We provide an overview of these datasets in Table 1a. These datasets span multi- 4 task ratio dataset Laion-EN (en) [93], Laion-ZH (zh) [93], COYO (zh) [10],Captioning 53.9% GRIT (zh) [90], COCO (en) [17], TextCaps (en) [99] Objects365 (en&zh) [97], GRIT (en&zh) [90],Detection 5.2% All-Seeing (en&zh) [119] Wukong-OCR (zh) [29], LaionCOCO-OCR (en) [94],OCR (large) 32.0% Common Crawl PDF (en&zh) MMC-Inst (en) [61], LSVT (zh) [105], ST-VQA (en) [9] RCTW-17 (zh) [98], ReCTs (zh) [137], ArT (en&zh) [19], SynthDoG (en&zh) [41], COCO-Text (en) [114], ChartQA (en) [81], CTW (zh) [134], DocVQA (en) [82],","citing_arxiv_id":"2404.16821"},{"n":1,"role":"dataset","polarity":"background","paper_title":"ShareGPT4V: Improving Large Multi-Modal Models with Better Captions","primary_cat":"cs.CV","context_text":"Inthedistance,treescanbeseen,addingatouchofnaturetothisman-madesetting.Theimageisasnapshotofeverydaylifeatatrainstation,capturingbothitsroutineoperationsanditsinherentcharm. (a) Comparison of Captions' Quality(b) Comparison of Performance Figure 1. (a) We showcase a comparison between the caption in our proposed ShareGPT4V dataset and those utilized by recent large multi-modal models (LMMs). Unlike COCO-Caption [7] involves brief human-made captions on the main subject. LLaV A-Instruct [31] combines human-made captions, bounding boxes, and GPT4 [39] to 'imagine' the image details, which leads to inevitable er- ror/hallucination description (marked in red). Our approach involves feeding carefully designed prompts along with images directly into the advanced GPT4-Vision [40] and the descriptions are more detailed and accurate (marked in blue).","citing_arxiv_id":"2311.12793"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"VideoChat: Chat-Centric Video Understanding","primary_cat":"cs.CV","context_text":"This approach allows us to benefit from readily-available image instruction data, creating a system capable of handling both images and videos with shared spatial perception and reasoning capacity. Stage1: Alignment. To strike a balance between training convergency and efficiency we introduce 25M vision-text pairs for one epoch of fine-tuning, The data consists 10M video-text pairs from WebVid-10M, and 15M image-text pairs from COCO Caption [ 6], Visual Genome [ 17], SBU Captions [31], CC3M [35] and CC12M [4]. The input prompts for LLMs are as followed: • \"###Human: <Video>video_embed</Video> video_instruction ###Assistant: \" • \"###Human: <Image>image_embed</Image> image_instruction ###Assistant: \" The video_embed and image_embed are the output from the token interface. Meanwhile the video_instruction and image_instruction provide concise video and image descriptions","citing_arxiv_id":"2305.06355"}]},"authors":[{"id":"2228b505-7e88-4b74-9438-5b11d5648e2b","orcid":null,"display_name":"Hao Fang","source":"manual","import_confidence":0.72},{"id":"2fbd3b65-8775-4dd0-9095-feb61c7c9159","orcid":null,"display_name":"Piotr Dollar","source":"manual","import_confidence":0.72},{"id":"a3d49b33-1d94-41ad-9709-ce6f77da6d04","orcid":null,"display_name":"Ramakrishna Vedantam","source":"manual","import_confidence":0.72},{"id":"a0d85cd0-cbd5-41ea-a62f-355c2ef0f812","orcid":null,"display_name":"Saurabh Gupta","source":"manual","import_confidence":0.72},{"id":"00e2756c-d70b-49a9-b019-ad42b25d5fab","orcid":null,"display_name":"Tsung-Yi Lin","source":"manual","import_confidence":0.72},{"id":"f871a5ec-7966-46e3-b89d-8aa9f74186a1","orcid":null,"display_name":"Xinlei Chen","source":"manual","import_confidence":0.72}]}}