Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:49:28.151834Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 0 inbound Pith citation observations for arXiv:2505.08971.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:49:28.151834Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
96 of 96 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b4a69d45-0675-4a6f-9889-01abd5f98526 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Pixtral 12B
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eac79578-b33d-4d61-8055-54d5347c4c7a · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Flamingo: a Visual Language Model for Few-Shot Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e4b0177-ebec-4c49-bb28-d1620663d0b2 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MINT-1T: scaling open-source multimodal data by 10x: A multimodal dataset with one trillion tokens
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b7b35db-ab7d-4346-bb4a-ec2459e40c71 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf650b6e-bd5b-42eb-8b00-a0ef908fd3d7 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Language models are few-shot learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1748332-e4c4-4340-ab26-ebe6c57d6913 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Lundberg, Harsha Nori, Hamid Palangi, Marco T´ulio Ribeiro, and Yi Zhang
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 193d4a2f-c0a6-4676-bdd2-cc379ab3f6dc · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Sharegpt4v: Improving large multi-modal models with better captions
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e5337ba4-b7db-4866-a6ed-8045cbbe4d53 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d6176b4-2d18-4521-bfd1-43c0e6ddca68 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Sharegpt4video: Improving video understanding and generation with better captions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b900c1c4-cf50-4c71-a589-a6197200a49f · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training CompCap: Improving Multimodal Large Language Models with Composite Captions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3f217c6-9abe-4c43-ab6a-a8756bade8e0 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training DRESS: Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef02f1db-390f-4a24-b7e8-23c29d79ab26 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training ViStruct: Visual Structural Knowledge Extraction via Curriculum Guided Code-Vision Representation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 80f5c340-a661-4d8b-9aea-0324e68270d6 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Scaling Laws for Predicting Downstream Performance in LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a0b0cc-42e7-4957-8077-93ec3a6cb291 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79b26b88-b6ec-4573-8418-6c9a459c59cf · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f997c8d-10f6-4d55-aa56-1ef07d717643 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training VisionArena: 230K Real World User-VLM Conversations with Preference Labels
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 365ae205-0775-4f22-812c-fb6b119112ac · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb3e2227-14db-4f52-a364-5876d877d964 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training NVLM: Open Frontier-Class Multimodal LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a44695a5-20f7-4c90-9b42-69e27ac5eb05 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Unveiling encoder-free vision-language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82f4f70d-6151-44d4-a90c-b8a976f052a3 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training A survey of vision-language pre-trained models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2638719a-6663-4c36-8752-37b15a7b659b · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 318baf8d-a093-4204-af64-0384b002044f · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training The Llama 3 Herd of Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17d48c46-1b18-46bd-8c29-9480716cab80 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training On Pre-training of Multimodal Language Models Customized for Chart Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b3c34e9-828b-4093-9126-33a9a2bd0920 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce61c136-8d4b-4514-8c5e-204acd8d0cc5 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training On Pre-training of Multimodal Language Models Customized for Chart Understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d789315-dd49-4844-a28b-1c0980a06322 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Making llama SEE and draw with SEED tokenizer
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 261dcc71-7fd3-4ff6-a3db-4693e805c0ce · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e8ba83b-2c37-4ea6-813b-205a423b23d6 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training GneissWeb: Preparing High Quality Data for LLMs at Scale
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2c1f37b1-9fae-4d97-b5b4-a15a1b0d5d06 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90741552-0327-4828-a6bd-6b6511beaa7a · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Exploring the frontier of vision-language models: A survey of current methodologies and future directions
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02a9fbaa-accb-46e2-9ae0-697d8c48db7f · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7140819-8ced-4d7d-b98a-4f199578fae2 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e98f92c9-47e8-4a08-a328-0c1cdccc6de9 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd34ff79-fe5c-48bb-81c4-b443b111b954 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Scaling Laws for Neural Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1cc2334-1a52-4202-9562-d6522531ad34 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Compression Represents Intelligence Linearly
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f83a591-95c1-4821-8c14-0b91a272b4e2 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training GPT-4o System Card
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3172c1e-ea2e-428d-96a7-4e2d469dff55 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15b72efd-41e5-4a00-861f-fecff345027b · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Grounding language models to images for multimodal inputs and outputs
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 876999fe-af79-4977-ba31-ae9939d891db · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training RL with KL penalties is better viewed as Bayesian inference
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb459e6f-d094-4a22-b529-c4d5196e416d · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Datacomp-lm: In search of the next generation of training sets for language models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f610eba-bf88-417c-9484-6b7154780b3f · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Otter: A Multi-Modal Model with In-Context Instruction Tuning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0160033-8a5e-4766-9c5b-caf5b1c93210 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Seed-bench: Benchmarking multimodal large language models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7480b147-b897-4d8e-9dd4-b9379c5240ac · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Silkie: Preference Distillation for Large Visual Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d435741-251e-484a-a5fa-5145de9af93e · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83152920-85f9-464a-a8b2-9e820abe7f2d · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e389e1cd-a794-40d7-bfc2-84c10a3bb9d7 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64b974cf-03f0-4957-9ade-a54420a8ddf9 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Multimodal arxiv: A dataset for improving scientific comprehension of large vision-language models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e1943bf-0a4e-4196-943c-f2405ecb5dc8 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Visualbert: A simple and performant baseline for vision and language
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 030a7c39-0ea4-48c8-985e-e6137171ad2d · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f3ca946-0d59-4336-ab7c-e5c123507aef · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Evaluating Object Hallucination in Large Vision-Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a24ff640-5790-44ae-81e3-8638138857b3 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0b3e2b7-8bed-41e6-810d-ead96c0b64d5 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55e44184-c986-4471-9ab1-d201a7d5b174 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64af87a0-4391-48d9-a7b1-402833b212ec · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Microsoft coco: Common objects in context
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0f72c81f-683a-4b5a-8cb7-c7b016a21ec2 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Visual instruction tuning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c8d007ba-729f-4e2b-ba97-02809fd98630 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74826612-12bf-4a03-a56c-c979fb637aa1 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Improved Baselines with Visual Instruction Tuning
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25e0a611-7da3-4a9a-a5f3-78cf6b062332 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c4e6354-915d-4aab-ade1-b9e1fd6fde6e · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Diving into Self-Evolving Training for Multimodal Reasoning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abcf93ad-1c9a-4057-a401-855473f41316 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebb61c47-cd18-4434-9020-7256098dcdf9 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd2b56e1-36e0-4c13-ba33-2fda71e1fd25 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13a1a594-15ed-44fe-968a-0f57024e18d4 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96804bae-efba-4945-9a92-aa6c61c3e9ae · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 85f6e33c-f575-48ce-8f74-8ac8a5a68c97 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Learning transferable visual models from natural language supervision
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6cbddcc-b0f9-4c4f-87ed-ae55c6cbc287 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1d25b96-0c5f-40ef-85d1-08513a34091f · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Openomni: Large language models pivot zero-shot omnimodal alignment across language with real-time self-aware emotional speech synthesis
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb22e9f9-d6f7-4ba8-aab5-1e175efc7fce · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Policy opti- mization via importance sampling
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dbcd1a2-a988-47a4-a2d8-dd151e8fc9d5 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Laion- 5b: An open large-scale dataset for training next generation image-text models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9abe44e-0b76-4db6-ac12-ddaf0d0f5925 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Shaker, Salman H
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3879d285-9344-4d6b-bf72-5688cc28d5f0 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6bea31c-8c9a-4b51-8623-3ed4649d3763 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb89ef5c-69ad-403c-aafe-b27ca3640b4e · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Emu: Generative pretraining in multimodality
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5b5cbdee-8ca9-4930-91d5-bebdfe1845d3 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 360d5162-a319-46d7-8474-878065e5879a · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3da53c17-8322-42eb-b5d4-d3185b89ee71 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training PandaGPT: One Model To Instruction-Follow Them All
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ff4b12a-1022-4dd7-a36f-6792416b7a70 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2df2a628-0da3-4ae6-bef0-8615714bda71 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Aligning Large Multimodal Models with Factually Augmented RLHF
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84b0381f-ec30-40e2-8293-c9407eb28d2f · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3335f5c4-4051-491a-b313-b5ff7f2c1351 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Reconstructive Visual Instruction Tuning
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad5553e3-56c8-4a2e-89a1-af392a36dd7a · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d93d7d5-4d49-4725-bcc1-7909cbddcb40 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Scaling Pre-training to One Hundred Billion Data for Vision Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9749ed0-4144-48d1-ad61-70dc9e887f97 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d837735-db69-4a1c-a801-adf9a07ce6de · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5f5d2ea-f6fa-4feb-a59e-cef1e23e31cd · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Knowledge-augmented few-shot visual relation detection
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation eed933be-e5d8-4304-81af-4f50737727fe · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Qwen2.5 Technical Report
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 290d4613-fbd5-4354-b505-ccf4a1af24e1 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3c8e3e5-187a-4f61-b571-83731f8e9d7a · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Capsfusion: Rethinking image-text data at scale
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5713fdfb-d4ab-4fe3-8ad7-32ea69cf88df · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 048a03d7-ca78-4126-a150-c8fcb0b526ad · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Xing, Xiaodan Liang, and Zhiqiang Shen
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3aea2af3-136b-4e7b-94b9-a7548592bc4a · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Vinvl: Revisiting visual representations in vision-language models
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 443d7519-58aa-4878-8b9a-b7cf2ec0812a · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e80c6efc-31ac-41fb-aad6-ac2ca9b4bde8 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Minigpt-5: Interleaved vision-and-language generation via generative vokens
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0b3a4518-f49f-4699-99a1-4af729b51b37 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Minigpt-4: Enhanc- ing vision-language understanding with advanced large language models
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 87d92ebd-c559-40de-af3c-bd1e0548a290 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training Generalized decoding for pixel, image, and language
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d08c95db-92d6-44b6-a2db-cdbd44d42d01 · outbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training CompCap: Improving Multimodal Large Language Models with Composite Captions
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.