Pith. sign in

Paper Citation Record · LEDGER

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

As of 7 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 1 inbound Pith citation observation for arXiv:2505.24164.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24164 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:37:50.854503Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T11:13:16.940627Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a205b59b-d940-44ec-b20b-b1ce05113763 · outbound

This paper cites GPT-4 Technical Report.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:42.783123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:42.783123Z digest=sha256:3b298ebbc4226ddc169107bbb10f3278c55af6fd7bb5e71d6644890170149131

Observation 5b68f802-f117-4d20-9c57-9b2f6fc45502 · outbound

This paper cites Qwen Technical Report.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:42.873209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:42.873209Z digest=sha256:6b501ae13918e4bdde02a8cdf88df1cd3f48a09b73c264ba66f6eb9f2cd26d49

Observation bc5380aa-495b-403f-b824-c951af462fd2 · outbound

This paper cites Qwen2.5-VL Technical Report.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:42.957030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:42.957030Z digest=sha256:448d4fea15bcb621a3fc43e63ce132a748e5b5669309aafe88b79a5296f33d2c

Observation a354a0fe-2f23-4469-9fb5-6d57164bfc21 · outbound

This paper cites MapQA: A Dataset for Question Answering on Choropleth Maps.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MapQA: A Dataset for Question Answering on Choropleth Maps

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:43.047242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:43.047242Z digest=sha256:062beb523e3f0db9654620d994e366630a02276da48069b53adf2f15cb54dcbf

Observation 0ba50aa3-3999-460d-a51b-f97816fe33cf · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:43.118516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:43.118516Z digest=sha256:3dd3ad95eb3895a565c184374af38bb9bc744de7f23223714c72f7a233e4c55f

Observation c960d2c9-ada4-4dc6-b172-2fc19031cb0f · outbound

This paper cites UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:43.218009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:43.218009Z digest=sha256:ff83ad6d3e9a7fc98fc430edf41c7a027c75085125bbd9e1cbf7e521aadb8499

Observation e4615098-7cfd-4dfc-abb7-03ccc8f23df0 · outbound

This paper cites GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:43.332355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:43.332355Z digest=sha256:2217dca0533d846b74c3a55d1fc1f472092acfe2bb8ff5f3b3fca6f1c351a196

Observation 910e8922-4bdd-4d7c-95a5-cd450e5e09d3 · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than $3.https://github.com/Deep-Agent/R1-V, 2025.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models R1-v: Reinforcing super generalization ability in vision-language models with less than $3.https://github.com/Deep-Agent/R1-V, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:56.305112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:43.437869Z digest=sha256:27a6e43259c34de0623ab3f3a4b09d9ba30bce3ffbb88cc36470ce894afc9a66

Observation ad424351-a347-45cf-9118-0edb67921a5c · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:43.536170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:43.536170Z digest=sha256:3b7b42d19c73de1150928f1a410d9c169aac9e2e6ffd3fcd1708ba09cf8abb64

Observation 82dd815f-cc4c-4aac-a85e-a3dd84af6992 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:43.644290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:43.644290Z digest=sha256:a1c2f33f857ddf26731e45256d178e48a9d0d7bc23e511c86e308d5156505e59

Observation d1e9d65b-6059-4ad6-be0b-c45f35d4a376 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Gonzalez, Ion Stoica, and Eric P

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:56.125085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:43.735687Z digest=sha256:17a9988a7901b8c4d6de625ebea0a0553ae01fdb85a732ebacf3ff3f37426fde

Observation 3531540a-cfb7-48ce-8f1b-fe894f167470 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:55.949527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:43.800014Z digest=sha256:4bc66038b7796a17d052d4a8de15c71912f63e13f873da7899a920cea92b4107

Observation 271734de-2d98-44a3-b87c-25be0653fdd4 · outbound

This paper cites Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:55.748517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:43.845422Z digest=sha256:9c319b660ad5c67b493eabdb628fe58ca56685fbb50d3b3c06164807bd8041d0

Observation d29e0a8c-db58-4aaa-aa98-07b0f9a8dac0 · outbound

This paper cites Bert: Pre-training of deep bidirec- tional transformers for language understanding.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Bert: Pre-training of deep bidirec- tional transformers for language understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:55.599809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:43.920685Z digest=sha256:8956b90cecad5a74023ef42537854ef056f14c1a34962479d8b35cd99ca0c1a7

Observation 1b27fa32-f58d-4a23-b944-4dd3a2c638d1 · outbound

This paper cites On Path to Multimodal Generalist: General-Level and General-Bench.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models On Path to Multimodal Generalist: General-Level and General-Bench

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:43.985549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:43.985549Z digest=sha256:a4e85d300df503aed62a7ce2302e00ce3fb01c1147e841ee64a8f2eb95bed3df

Observation fa40b409-6c78-442b-84d4-fc1701f09a5a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.072119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.072119Z digest=sha256:15b7741ee5444d2cd7b3e5066e6207e569914e13329ea6a43a0c4a97b779a3a0

Observation 2d9e2b31-4beb-472a-9052-bcc9a9d5f27c · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.110624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.110624Z digest=sha256:bf073769e9f761ef476c939f0b56e6f021f41c58f74a905e6454a42041ce856f

Observation 701d8e6c-81b1-4d7a-85e4-20d51bce3fd2 · outbound

This paper cites Cantor: Inspiring multimodal chain-of-thought of mllm.ACM MM, 2024.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Cantor: Inspiring multimodal chain-of-thought of mllm.ACM MM, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:55.458081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:44.208714Z digest=sha256:a394e6d61020e1f7117144db1fb595bacc38616510a4c104ed32846bbd299f4d

Observation ba33b5e7-45dd-4ffc-b33e-d25f34b1d1ed · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:55.289867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:44.295082Z digest=sha256:74d10683b763b716b55b98a7a942747b72b894b128c0340054b0aae939ead1c4

Observation 0812ec8e-1fea-492a-9148-2110ca2e4d0d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.393355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.393355Z digest=sha256:3c6b8b63d086caa35288b6c75f836243916818bf55076ae17db25580ebc9e093

Observation 9848067d-40d2-44bd-a3a3-aac8c40b6c7f · outbound

This paper cites Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.473998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.473998Z digest=sha256:25d6bdf9681f66775d5c08854d77195bf93b429162ff868f3217e6a0a390995a

Observation 3fa536d7-b889-42b8-afff-03f25b2f2e80 · outbound

This paper cites GPT-4o System Card.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models GPT-4o System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.516429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.516429Z digest=sha256:8635f8917b301d66bb3d4db61234d818a92f88a5ecd78994724390912d36ef6c

Observation d1dc0482-7b39-4693-b682-578cc52782db · outbound

This paper cites Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.565913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.565913Z digest=sha256:29b221a246eb1c5e33e41e49ef9725b5303f6a16ea9c910ea636f08a92cef746

Observation 045da288-d805-43d5-872b-449c83ab6d74 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:55.115246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:44.642885Z digest=sha256:be9117528223570c11cd0a33855b2cd00f161ce8eb7ab07c1f2e5dfb23882c8d

Observation bacc674a-5048-48a6-8529-75e556c96760 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.703110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.703110Z digest=sha256:c67e4a49d398ae0ca7d22bff7e350f1a3a43e2aaa6432544c2d9d3bdf1612db8

Observation 0140bf2a-a5ea-4d8a-a649-d9bf3e65de77 · outbound

This paper cites A diagram is worth a dozen images.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models A diagram is worth a dozen images

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:54.951630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:44.746200Z digest=sha256:2ffdf6b36664897b767b868ac212c971554c93a85ed8e2ab888b2c4d0706ccbe

Observation 080d880f-5ecb-4032-a0cc-3c6fb3a37e2f · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.784518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.784518Z digest=sha256:b0048608437473092f803413b7ec818f053269966042f8f7efc277ca5c0ad8c5

Observation fb4f7a2e-b358-47af-b11e-7b046b6e82c5 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.857093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.857093Z digest=sha256:02556296d3876ccc7a7d3c8eced67403e3fc01100542b44d4b2bfed31fea16c9

Observation 2ab70df7-2694-4d02-9eac-22b516c053c1 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:54.794908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:44.902929Z digest=sha256:775865ba431742d808357ed94935303f72ba262caf7da1b449e0ad3b24fd64af

Observation 059e7857-0160-4348-a977-ab5a86d5b89e · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.976247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.976247Z digest=sha256:d964f62f74c372722acb80d30baa7630f2ffe9af24451673eb29611dbb2c2f3f

Observation 311990ad-275b-4c1d-943d-b49c5c529f35 · outbound

This paper cites DeepSeek-V3 Technical Report.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models DeepSeek-V3 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:45.063410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:45.063410Z digest=sha256:9c89cf09fc246fa46ec38f2ecfa90fc5ee2c56d5114df60aedab16b9aa208ae8

Observation e3d474fd-64fa-4b36-b4b2-e8e5f73f7bed · outbound

This paper cites Visual spatial reasoning.TACL, 2023.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Visual spatial reasoning.TACL, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:54.625675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:45.149681Z digest=sha256:ad96ca4fcbd531ccf10b764852392d452f848fe7638e60b8b5ce6522b405ed4e

Observation f8a42cc8-70a8-4716-91f5-9b8cd2cf7515 · outbound

This paper cites Improved baselines with visual instruction tuning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Improved baselines with visual instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:54.441783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:45.152810Z digest=sha256:41bb1133771d779f64cd529182e61f92a51a31e88d5a7939c76e6996bb892437

Observation 848b80bc-c73f-4377-93be-7c36b161b31b · outbound

This paper cites Visual instruction tuning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Visual instruction tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:45.234152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:45.234152Z digest=sha256:963fa67fb76113bb0ba8d79007139b760e90ec28e6946084645f20408b2160ba

Observation f937c7b9-8cfc-4f61-9ad7-50ae0c9f615b · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In ECCV, 2024.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Mmbench: Is your multi-modal model an all-around player? In ECCV, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:54.262844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:45.396245Z digest=sha256:c6f1bb4304e96a8d1889f93245b2e0abdf9e5c902bc0c6c28c45ea2714fd7e2f

Observation aec24713-eebd-41a3-a6c3-f838f33c3b54 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:54.130873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:45.618007Z digest=sha256:bdac0680b7704ca57f41aab66168a7e8a2e3198d8c42d638aa4924b806f04861

Observation 49741e44-f24e-421a-9897-31100a89f9e4 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:45.787946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:45.787946Z digest=sha256:768b12562d501e463aac66e2087518ecca0bbaf478950639e0767d253dfb5d45

Observation 14856869-d85c-4176-9161-29c064ad3d38 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:53.983355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:45.960301Z digest=sha256:cd6f2da30ff8e7d2ad6a853ce96d84a8ce7e89a34af303701a1dde5f8074d018

Observation 6a495537-6ea1-4de9-9765-6701e7bbcf4e · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:53.808532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:46.128342Z digest=sha256:15f022b13e0f878150a63f273249497fc7b85a1b0d4b465b6ead3c50acf4d86d

Observation 4a02b294-cac4-4826-b864-84262d5b913c · outbound

This paper cites Cheap and quick: Efficient vision-language instruction tuning for large language models.NeurIPS, 2023.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Cheap and quick: Efficient vision-language instruction tuning for large language models.NeurIPS, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:53.634049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:46.363496Z digest=sha256:ddbc0ba22c1dd4f3d106c48c1ce3c247a652ec06ec98cbd9b525964195748a7b

Observation 7d0232be-36ae-4f25-a436-a166fe6a183e · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:46.509499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:46.509499Z digest=sha256:1311b59c028081ba9dbe8cd37402f0b65a0f14b81d8cb8242967e6d9cf471681

Observation 90b0c37c-04c1-4dc9-93d3-0e3233b263b0 · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension.arXiv preprint arXiv:2411.13093, 2024.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Video-rag: Visually-aligned retrieval-augmented long video comprehension.arXiv preprint arXiv:2411.13093, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:46.617047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:46.617047Z digest=sha256:727cada328a51b598d10e56524df7066238c7cd93d197dccb867abf672eb41dc

Observation a972f16f-ab27-413d-8770-72947fbea608 · outbound

This paper cites MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:46.790491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:46.790491Z digest=sha256:147b21cabdb0abeb6327e6d6b96cc9056324c3abefd3cd60c5157a83d5fc9793

Observation db05e5a2-683a-422b-9098-eb214b87e391 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Generation and comprehension of unambiguous object descriptions

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:53.474140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:46.908031Z digest=sha256:1649b645404c827886ceb4dd44ba1566d287117f45319dba03b355068ee3d63f

Observation 6aa570e5-7757-4a4d-9507-a2c62a29d32a · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:47.054610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:47.054610Z digest=sha256:3fa96965e288f3924f0c144598eff3aaaa1ce651a219dd702f3a857a4d5a2528

Observation feb65134-920d-4bf9-8fbb-cd50e46492e0 · outbound

This paper cites Infographicvqa.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Infographicvqa

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:53.318799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:47.219388Z digest=sha256:e6c06b4dfaeb47a0582ed05b1ba1a9e6db537b9db2052c1f535a8be4dac1a8fb

Observation eb606e72-c086-4dac-bfbc-73847b3c9101 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Docvqa: A dataset for vqa on document images

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:53.165749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:47.344366Z digest=sha256:0b3401463f58d5ae05ce73ce17449aed4537d7b65b5ee9383bc62a952d3c1491

Observation a7e68bb9-1055-4b61-88d0-f5951e67213b · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:47.537408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:47.537408Z digest=sha256:52543a1b3c03d59f83854677e0ca264276ec73df35fdc7c9a6b4d1f6d16ca40c

Observation 9d097eab-a669-4ed7-9c41-8431050b7f92 · outbound

This paper cites Training language models to follow instructions with human feedback.NeurIPS, 2022.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Training language models to follow instructions with human feedback.NeurIPS, 2022

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:52.920910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:47.706893Z digest=sha256:0702db6baf488cd94360affb2234349b705710491dcfd35a51398bebf542c7d2

Observation e11f3118-abc1-47d5-8ca4-0dd272df0855 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Learning transferable visual models from natural language supervision

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:47.865340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:47.865340Z digest=sha256:cc47d1ed8133cc2dbf1196858f48a23ea9be30b399bbd72ad47cd2dbb823dfca

Observation c0ff8c5d-9f1b-41e6-911a-855075ae6886 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.NeurIPS, 2023.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Direct preference optimization: Your language model is secretly a reward model.NeurIPS, 2023

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:52.695218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:48.030928Z digest=sha256:023098e11ae47b70782d9cce985c6e454860d14fe02797dd376823d490654166

Observation 879f13ae-d0b9-4b5c-a844-38f843aaf72a · outbound

This paper cites Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.199482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.199482Z digest=sha256:93beff2aabbdeb2f7b81592ca6376c9acfc674919a7feb60a09a5c82829b46cd

Observation 791d5a95-3c79-490a-b88c-e20176d1b599 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.318387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.318387Z digest=sha256:46f7d7037d06762e5ecce973f94094ce8b5e99217d31998261412f897d0d043d

Observation a8abd1ea-e136-48d5-a244-bb3c22e22062 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Proximal Policy Optimization Algorithms

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.437328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.437328Z digest=sha256:e26712e6c3f4ef55b29a4d8640f18649303071a0c6ec0bb2a39543bc0287c23b

Observation d0afd5ec-f21b-4d35-8aca-16b8b8b998d0 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.531778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.531778Z digest=sha256:fb93dee19911d24fcaef3310c711e9de1cc2df93a7557eddba66f44e4995bbe7

Observation 75ff2723-9024-4e4a-883d-f23746d7593b · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.622594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.622594Z digest=sha256:5acd1ffd6f3739b48a60616fb6e577eebb5041086a944f37f2fb04b8edd335dc

Observation 2323e1fd-fc9b-4b47-9264-b649af1f9014 · outbound

This paper cites Long-vita: Scaling large multi-modal models to 1 million tokens with leading short-context accuray.arXiv preprint arXiv:2502.05177, 2025.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Long-vita: Scaling large multi-modal models to 1 million tokens with leading short-context accuray.arXiv preprint arXiv:2502.05177, 2025

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.691734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.691734Z digest=sha256:3ea83d7b86d1015ca2f8d8efe58cc45b9d0869fd801ef35177c1193b4c497dcb

Observation 020874f4-2c31-4234-a8ee-10e3083fa6cc · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.782938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.782938Z digest=sha256:312be4044e45c8cfd452dd98520fce0e8547e34bbaca84c721edf6291ee9e77a

Observation 13ea1e57-dedf-46a7-90b6-3388db6d6b98 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.877517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.877517Z digest=sha256:d639e9a94bf09fe32731c5217af9bdffcfbd4a9700d8793a370155c60e2910f0

Observation 61cfd93d-8772-489c-9a56-e7c2efee99cf · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.974699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.974699Z digest=sha256:2ea188904e3f416f66c295b470eabf22c8cd65849683a5f49768ea88b12d07e7

Observation e22b9b99-6fa5-4feb-9213-419fb69ecb85 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:49.073537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:49.073537Z digest=sha256:17a9b16096cb99a523cbb71e669dc67217b7c67f504c6527158f5d74f7220859

Observation bc7baed4-4b75-4a71-ad43-1f2c7c87d3b2 · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:49.199326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:49.199326Z digest=sha256:4dbeb7ba53710f084f21f04449a9b8e21bfd87ab37ad9126d33cbd8c51ba3f2a

Observation 9b330a79-d238-4eb8-889c-aa84aefb3fce · outbound

This paper cites Qwen2.5-1M Technical Report.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Qwen2.5-1M Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:49.263799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:49.263799Z digest=sha256:d671229df08fdd10f4d36d6f1d88779da6669dcb40d49290e127875193ab6ae4

Observation 9297dae2-6215-44be-8af7-ec6c6cb9706f · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:49.357574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:49.357574Z digest=sha256:300fa2b60ecf29d085a75ff7ba65bd8ad507f1d834055c2ac35f4b67dc6da24b

Observation 575783b0-dd21-4941-b5eb-464ca8c32aea · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:49.453135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:49.453135Z digest=sha256:e602988f6000f3dd6b48acd75f2073f17d4dde2a22a6465e209f96dea33106d6

Observation 6ee38d23-7e1c-47d9-9d2f-37e72a156375 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:49.581401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:49.581401Z digest=sha256:0b0e2c18b26d0eec1caac6c60e40eb3dee596cdf57066a20cee1e92d39d2ca4b

Observation 20fefdf2-e833-4a4e-aee3-f6208b8e0f46 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:49.696896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:49.696896Z digest=sha256:29a61e4c39607e1dd9ac286034b55a5379dcbf8c15852ed823a54e0f00ad9ce4

Observation 7099775a-08ff-495f-82ff-6a60aa94f3bf · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:52.485583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:49.827448Z digest=sha256:5d89caff020c17b1ddef8ce74077b833897d8f519619854abe7d14f0d41ab0a9

Observation 7976af05-89f0-42e2-a8c6-20aa80e3b659 · outbound

This paper cites Sigmoid loss for language image pre-training.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Sigmoid loss for language image pre-training

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:52.301424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:49.936532Z digest=sha256:fe1f555761c1c503969c569cc73d54b65a0bb552aa04523381dbe0b9bc90f8b4

Observation 9b77ae93-91de-49b5-8e52-1196ea0c1596 · outbound

This paper cites Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.037158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.037158Z digest=sha256:ec33743b2eaf08145a0b215787c7c69e3e4b940eecb0e7d8751090c9ecfb6145

Observation cb348a27-c885-4a0a-9667-9fb0ed96adda · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.123928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.123928Z digest=sha256:8e2df08245816097baa45c906096391a48543d45abd726ea1540733fcd4f37d6

Observation 71428098-f6cc-4eec-83ea-f68eedcd226a · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:52.132720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:50.210810Z digest=sha256:64682269d86cea0df133e31fa7093ded47c91cb5ac82f03a123dc75210703517

Observation 5984fb9c-150c-4bbd-9f4c-cb0289374dcf · outbound

This paper cites Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.273762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.273762Z digest=sha256:88273ad3c7852de1089a329331428993ea28dc12713a1ef321792ecde694c055

Observation 5be7b6d3-807f-4516-908c-2ae4f0eb2d20 · outbound

This paper cites Enhancing multimodal large language models complex reason via similarity computation.AAAI, 2024.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Enhancing multimodal large language models complex reason via similarity computation.AAAI, 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:51.942715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:50.386378Z digest=sha256:af81199e7ff4c1d53c32f972387baa952bd64ccbdc95cb0dce19913ff506cfd7

Observation 3b8c0bd6-739c-4e66-bd74-6d7741264b85 · outbound

This paper cites MultiHiertt: Numerical reasoning over multi hierarchical tabular and textual data.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MultiHiertt: Numerical reasoning over multi hierarchical tabular and textual data

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:51.749703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:50.496263Z digest=sha256:4f9b7d3e864d827b3bda554534cbf92ceee0af5f1ff4c401c60a53f5ccae4a4a

Observation 627beefb-2d62-4fc4-9721-e296d26b16a3 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.564213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.564213Z digest=sha256:cec132463cb392e290a116055e0057994b560637594f3d7309438df1950a2783

Observation b1c18b3a-35d8-48a1-9128-c02701f41d19 · outbound

This paper cites Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.658662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.658662Z digest=sha256:4139a615468d5a32efc0a20dc105d4c45d196c365bdf86c7e257fb0b5ad60b83

Observation ea093f35-2b5b-40ce-9134-c4dd340489b8 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.749824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.749824Z digest=sha256:19ae34b6ecc49826171cb370f70e98b67d0ca673a9791da8f1fb59dc06c6c6d7

Observation 9105b61a-f351-4ab7-be17-3628b1c0f762 · outbound

This paper cites Genimage: A million-scale benchmark for detecting ai-generated image.NeurIPS,.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Genimage: A million-scale benchmark for detecting ai-generated image.NeurIPS,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:51.554229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:37:50.854503Z digest=sha256:901642f2b0d01ca4bc69d580331987a9d83406f9cea300fcd582fdab47f5be75

Pith citing papers

Observation 47793b81-619c-4a58-94a6-a23602ca3b68 · inbound

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding cites this paper.

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T11:13:16.940627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:13:16.940627Z digest=sha256:2614aa2c32a00986f38354704c29ca58a3765aaff881433e23edf632dc702c5f