Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:37:50.854503Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 1 inbound Pith citation observation for arXiv:2505.24164.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:37:50.854503Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T11:13:16.940627Z
A source-named dated measurement, never combined with another source.
Source: cited_works
79 of 79 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a205b59b-d940-44ec-b20b-b1ce05113763 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b68f802-f117-4d20-9c57-9b2f6fc45502 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Qwen Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc5380aa-495b-403f-b824-c951af462fd2 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a354a0fe-2f23-4469-9fb5-6d57164bfc21 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MapQA: A Dataset for Question Answering on Choropleth Maps
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ba50aa3-3999-460d-a51b-f97816fe33cf · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c960d2c9-ada4-4dc6-b172-2fc19031cb0f · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4615098-7cfd-4dfc-abb7-03ccc8f23df0 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 910e8922-4bdd-4d7c-95a5-cd450e5e09d3 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models R1-v: Reinforcing super generalization ability in vision-language models with less than $3.https://github.com/Deep-Agent/R1-V, 2025
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad424351-a347-45cf-9118-0edb67921a5c · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82dd815f-cc4c-4aac-a85e-a3dd84af6992 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1e9d65b-6059-4ad6-be0b-c45f35d4a376 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Gonzalez, Ion Stoica, and Eric P
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3531540a-cfb7-48ce-8f1b-fe894f167470 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 271734de-2d98-44a3-b87c-25be0653fdd4 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d29e0a8c-db58-4aaa-aa98-07b0f9a8dac0 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Bert: Pre-training of deep bidirec- tional transformers for language understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b27fa32-f58d-4a23-b944-4dd3a2c638d1 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models On Path to Multimodal Generalist: General-Level and General-Bench
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa40b409-6c78-442b-84d4-fc1701f09a5a · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d9e2b31-4beb-472a-9052-bcc9a9d5f27c · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 701d8e6c-81b1-4d7a-85e4-20d51bce3fd2 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Cantor: Inspiring multimodal chain-of-thought of mllm.ACM MM, 2024
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba33b5e7-45dd-4ffc-b33e-d25f34b1d1ed · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0812ec8e-1fea-492a-9148-2110ca2e4d0d · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9848067d-40d2-44bd-a3a3-aac8c40b6c7f · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fa536d7-b889-42b8-afff-03f25b2f2e80 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models GPT-4o System Card
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1dc0482-7b39-4693-b682-578cc52782db · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 045da288-d805-43d5-872b-449c83ab6d74 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bacc674a-5048-48a6-8529-75e556c96760 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models FigureQA: An Annotated Figure Dataset for Visual Reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0140bf2a-a5ea-4d8a-a649-d9bf3e65de77 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models A diagram is worth a dozen images
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 080d880f-5ecb-4032-a0cc-3c6fb3a37e2f · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb4f7a2e-b358-47af-b11e-7b046b6e82c5 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ab70df7-2694-4d02-9eac-22b516c053c1 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 059e7857-0160-4348-a977-ab5a86d5b89e · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 311990ad-275b-4c1d-943d-b49c5c529f35 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models DeepSeek-V3 Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3d474fd-64fa-4b36-b4b2-e8e5f73f7bed · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Visual spatial reasoning.TACL, 2023
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8a42cc8-70a8-4716-91f5-9b8cd2cf7515 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Improved baselines with visual instruction tuning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 848b80bc-c73f-4377-93be-7c36b161b31b · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Visual instruction tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f937c7b9-8cfc-4f61-9ad7-50ae0c9f615b · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Mmbench: Is your multi-modal model an all-around player? In ECCV, 2024
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aec24713-eebd-41a3-a6c3-f838f33c3b54 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Ocrbench: on the hidden mystery of ocr in large multimodal models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49741e44-f24e-421a-9897-31100a89f9e4 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14856869-d85c-4176-9161-29c064ad3d38 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6a495537-6ea1-4de9-9765-6701e7bbcf4e · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a02b294-cac4-4826-b864-84262d5b913c · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Cheap and quick: Efficient vision-language instruction tuning for large language models.NeurIPS, 2023
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d0232be-36ae-4f25-a436-a166fe6a183e · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90b0c37c-04c1-4dc9-93d3-0e3233b263b0 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Video-rag: Visually-aligned retrieval-augmented long video comprehension.arXiv preprint arXiv:2411.13093, 2024
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a972f16f-ab27-413d-8770-72947fbea608 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db05e5a2-683a-422b-9098-eb214b87e391 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Generation and comprehension of unambiguous object descriptions
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6aa570e5-7757-4a4d-9507-a2c62a29d32a · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feb65134-920d-4bf9-8fbb-cd50e46492e0 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Infographicvqa
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb606e72-c086-4dac-bfbc-73847b3c9101 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Docvqa: A dataset for vqa on document images
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7e68bb9-1055-4b61-88d0-f5951e67213b · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d097eab-a669-4ed7-9c41-8431050b7f92 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Training language models to follow instructions with human feedback.NeurIPS, 2022
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e11f3118-abc1-47d5-8ca4-0dd272df0855 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Learning transferable visual models from natural language supervision
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0ff8c5d-9f1b-41e6-911a-855075ae6886 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Direct preference optimization: Your language model is secretly a reward model.NeurIPS, 2023
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 879f13ae-d0b9-4b5c-a844-38f843aaf72a · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 791d5a95-3c79-490a-b88c-e20176d1b599 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8abd1ea-e136-48d5-a244-bb3c22e22062 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Proximal Policy Optimization Algorithms
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0afd5ec-f21b-4d35-8aca-16b8b8b998d0 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75ff2723-9024-4e4a-883d-f23746d7593b · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2323e1fd-fc9b-4b47-9264-b649af1f9014 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Long-vita: Scaling large multi-modal models to 1 million tokens with leading short-context accuray.arXiv preprint arXiv:2502.05177, 2025
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 020874f4-2c31-4234-a8ee-10e3083fa6cc · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Gemini: A Family of Highly Capable Multimodal Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13ea1e57-dedf-46a7-90b6-3388db6d6b98 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61cfd93d-8772-489c-9a56-e7c2efee99cf · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e22b9b99-6fa5-4feb-9213-419fb69ecb85 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc7baed4-4b75-4a71-ad43-1f2c7c87d3b2 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b330a79-d238-4eb8-889c-aa84aefb3fce · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Qwen2.5-1M Technical Report
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9297dae2-6215-44be-8af7-ec6c6cb9706f · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 575783b0-dd21-4941-b5eb-464ca8c32aea · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ee38d23-7e1c-47d9-9d2f-37e72a156375 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20fefdf2-e833-4a4e-aee3-f6208b8e0f46 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7099775a-08ff-495f-82ff-6a60aa94f3bf · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7976af05-89f0-42e2-a8c6-20aa80e3b659 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Sigmoid loss for language image pre-training
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b77ae93-91de-49b5-8e52-1196ea0c1596 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb348a27-c885-4a0a-9667-9fb0ed96adda · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71428098-f6cc-4eec-83ea-f68eedcd226a · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5984fb9c-150c-4bbd-9f4c-cb0289374dcf · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5be7b6d3-807f-4516-908c-2ae4f0eb2d20 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Enhancing multimodal large language models complex reason via similarity computation.AAAI, 2024
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b8c0bd6-739c-4e66-bd74-6d7741264b85 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MultiHiertt: Numerical reasoning over multi hierarchical tabular and textual data
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 627beefb-2d62-4fc4-9721-e296d26b16a3 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MLVU: Benchmarking Multi-task Long Video Understanding
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1c18b3a-35d8-48a1-9128-c02701f41d19 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea093f35-2b5b-40ce-9134-c4dd340489b8 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9105b61a-f351-4ab7-be17-3628b1c0f762 · outbound
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Genimage: A million-scale benchmark for detecting ai-generated image.NeurIPS,
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47793b81-619c-4a58-94a6-a23602ca3b68 · inbound
Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.