Pith. sign in

Paper Citation Record · LEDGER

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

As of 23 August 2026, this Paper Citation Record lists 97 of 97 outbound references and 27 inbound Pith citation observations for arXiv:2502.10391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.10391 v1

Coverage vector

measured 97 of 97 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:23:50.744223Z

measured 124 of 124 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:28:02.977973Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:27:37.082400Z

Reference resolution

97 of 97 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved89
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f1b187c-2e91-4e11-8100-8651fc77a71f · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.523362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.523362Z digest=sha256:9b53e7e43f778c97c58faf4b73dc5b46cebe6101bac793554e04e1d223f8e3e1

Observation d2981123-6329-48b3-a413-4224712e0d6b · outbound

This paper cites Pixtral 12B.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Pixtral 12B

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.531232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.531232Z digest=sha256:290e8d6a3e28a3de1d053cbf7f35457eabc1f1c0adf9870e3a55a7cc144b8b3a

Observation 4f9838e1-e980-48ff-95a1-a0e08ad56315 · outbound

This paper cites Direct Preference Optimization with an Offset.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Direct Preference Optimization with an Offset

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.539127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.539127Z digest=sha256:87029cc4e35583bd72fed87176e577805f74db1f022f8270bdcfe72c4a7f109f

Observation 499c3639-92f4-4e37-a1e9-9ec70e78aeee · outbound

This paper cites Vqa: Visual question answering.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Vqa: Visual question answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.546058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.546058Z digest=sha256:61c98abbd217997d1d673830108171f7774d4028a0a6a5cfd323c5bd829ee77a

Observation ef901459-183a-457e-afde-9f2ece4289a4 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.552483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.552483Z digest=sha256:a668f51b2d1d6281821924c3f5156cfe202a91af00c43ee8f533db91227d74a2

Observation 2dcc7400-3811-4bfa-9905-b9f83fd47c1f · outbound

This paper cites TouchStone: Evaluating Vision-Language Models by Language Models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment TouchStone: Evaluating Vision-Language Models by Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.558553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.558553Z digest=sha256:a662400320b8135538551211a6d00bcdb1a921b179aebf2961975f231bb466f8

Observation b6300708-95dd-4364-98a6-de3baf853e30 · outbound

This paper cites VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.564656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.564656Z digest=sha256:eb7a1c931895fc827acc3460ffdc1b9081d3a307d81b3e78dba6cfd2987314c0

Observation a837664d-9ae9-4e18-8bcf-5152b23862da · outbound

This paper cites Language models are few-shot learners.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Language models are few-shot learners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.622783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.622783Z digest=sha256:403685e81b9dbad8c0786e26f1bf68f4b6bd0a0c778cdc356d73890aafdcaa23

Observation d1236c06-b979-4832-ba05-170e0a96013e · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.637480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.637480Z digest=sha256:1cb655a04962d21ac992349298ddd740587309fb3e11ba216403263d6f9a1724

Observation 8f907f7d-2050-48be-abc9-27c2a2d25941 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.679159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.679159Z digest=sha256:26e91773e01b3c12055a84afa0c68d1ad73a9ec0cf3d536f775a79df70666c8f

Observation ff0a1798-c851-430b-a8bf-b6b3fc6034c5 · outbound

This paper cites WebSRC: A Dataset for Web-Based Structural Reading Comprehension.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment WebSRC: A Dataset for Web-Based Structural Reading Comprehension

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.712859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.712859Z digest=sha256:68b188da2dfef896f01fdf9dee73b9acff7b9a74b4b52ed35cc8759c110dc630

Observation 1f52f174-8dd9-40e8-a87d-96932e6cb91d · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.745194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.745194Z digest=sha256:e6e9c5e0a71c4358008586f238200f42a91084edd035a50d906c72b42b5f1e6e

Observation e1dcd352-3243-4d99-8347-784d4cd096b5 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.779403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.779403Z digest=sha256:b433c78015ce71aca09286dd5cb2724b059372c44063831acf1cc2f49b966fe8

Observation e14d5402-2079-4dc0-97bb-ceb0a69e7e40 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.787756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.787756Z digest=sha256:ad511ed7b4987adac47a67e783922c94474fef0f1efdad0d09da27440c1fcff1

Observation 6e9ef6e6-ad4d-4ddd-ac35-445bbf735d62 · outbound

This paper cites Provably Robust DPO: Aligning Language Models with Noisy Feedback.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.795434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.795434Z digest=sha256:a368f878b95d2280f933860d83d1410fca53b939144a4bbd41a8f8aa6276ebb8

Observation c0ad6dc9-7477-4748-a754-7959c63eb6e4 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment NVLM: Open Frontier-Class Multimodal LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.802749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.802749Z digest=sha256:9a714ea88d1e4408f968dc3f10a9e3a6cf58187527faf7222e1f658cf1644580

Observation bf019ee6-6568-4fcc-b033-6447fa3d5023 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.810907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.810907Z digest=sha256:16daabbf7c1ada6c5953ca09985b21f8bc178f8724918bce2d35a79768f2db7e

Observation ace44fbc-4616-4b1d-8ea0-67038fb73abd · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.816591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.816591Z digest=sha256:c4a78f2a9d5d243cc110c4d3f3f12cb06023ff564f290f4d9c133120c4f2ef20

Observation 181ac02a-f004-420c-9145-ee9aa73774f4 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.823966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.823966Z digest=sha256:306d77778d28c87c7639f617cccedae0e071042350e88702bf472ce7fe615ec0

Observation 63fabce5-e990-4e70-aed4-1c56716c81ea · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.829925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.829925Z digest=sha256:95742aa2ce3589ef43af48bbbfe113a81453f2e9abc6cf7d3a747f3111f2826e

Observation 2eaab4af-c665-4c62-9349-65ff4b72a787 · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.836288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.836288Z digest=sha256:73531a4b5c62117024ee2de176ab695b3a42170a4537f5b580af8bd60cfb0f51

Observation c292084b-02c9-477f-b785-df96ca3191a2 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.843627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.843627Z digest=sha256:bfc349deb8f85b04f4aecc9282d5faed1a7bdc4f1f94eb7d46addd8e625eefb5

Observation e71ccea0-7fe7-42a3-88eb-843910d7a257 · outbound

This paper cites MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.849204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.849204Z digest=sha256:068cc9d5ab0c76f118f2ce50d051fa0afc9eaebbec39dfaa7570936d02aa6167

Observation 7d442750-fa52-491e-a38f-d1a8e0582475 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.886875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.886875Z digest=sha256:83f5e05f990911010f94d43296c590999abe36a9671f5294a1022c0c5bf7a4f6

Observation 6896c9b0-4cf1-422d-bf1f-0956f20e2c7d · outbound

This paper cites Infimm- eval: Complex open-ended reasoning evaluation for multi-modal large language models, 2023.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Infimm- eval: Complex open-ended reasoning evaluation for multi-modal large language models, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.893155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.893155Z digest=sha256:159affc22a97b0f854315d0d49490b031fcee11b40d35c371b1baf69861211b7

Observation 660fabdc-fa88-4af2-8055-43e8eb1228fb · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.899164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.899164Z digest=sha256:a94fab7b6ca45c8a8454c8304309b1f8a4b0a7dcea27e662a38b8bc16f207ee9

Observation 5cc59590-50e3-48f3-ad73-bba894f76022 · outbound

This paper cites VLSBench: Unveiling Visual Leakage in Multimodal Safety.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment VLSBench: Unveiling Visual Leakage in Multimodal Safety

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.905004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.905004Z digest=sha256:279f8be265c0aecca32313ce001e5c005ff00d4b9ed3c57eaf7d387989caa045

Observation 15a70143-3fe8-4ccc-b980-27fc3a716ce4 · outbound

This paper cites Mistral 7B.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Mistral 7B

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.910667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.910667Z digest=sha256:a8d7d214d9fce4fe40c2cb88cf8d185a9189f8930c842ae718fce10910f67fdb

Observation 52047ede-8d17-4efc-b77e-38bb3fb5d9be · outbound

This paper cites A diagram is worth a dozen images.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment A diagram is worth a dozen images

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.916456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.916456Z digest=sha256:0f7d7865c5f1544cd1405403fa4b065bcc35c6ba6f2058df406ccf2c6c27d5e0

Observation b855676b-334a-40de-9c4b-032a8a83edf5 · outbound

This paper cites Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.921302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.921302Z digest=sha256:54b0f51d839f59e545835e09f18c6de3cccbdc0bb636999c9ffbe8733a802094

Observation 93133f22-5bd8-43e2-8e97-4dbdc2bae508 · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.926724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.926724Z digest=sha256:badd0c7cb7a13124b402a283592b6ae4e6e4090861012e5813112fe38c7ed946

Observation 664a80c7-ddfa-489f-a597-752f7529b8da · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment LLaVA-OneVision: Easy Visual Task Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.932313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.932313Z digest=sha256:93c1833c1c1f4eb9c23bde08b64c0c16d1e8427a06f99b25ff678130cad9fa66

Observation 611c7e4c-0ec4-4ffb-99de-ecdc32dfe686 · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.937654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.937654Z digest=sha256:f3bd6118087bd2be5dcd5889e40954feed7ab3a6e95a4e01576f162b6593dc0b

Observation 29c7f356-9e0a-4313-a05a-78815ee61366 · outbound

This paper cites SEED-Bench-2: Benchmarking Multimodal Large Language Models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment SEED-Bench-2: Benchmarking Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.960160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.960160Z digest=sha256:5978dd28a02309a0c87376130f706d4c7c28d4fdd907be74ac6828938344f788

Observation d7dd9c67-2de6-4825-8f10-5c3a29621257 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:49.989643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:49.989643Z digest=sha256:d990e84d4f02566ac5ac5313e2fbe0ee41354074588bb04cc549ff738cfe1d93

Observation 07bc4a16-2f98-491a-9d28-29b500f400ec · outbound

This paper cites VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.023089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.023089Z digest=sha256:809609491fb9f7dd6f89826ebbaa750f11192c4842e4b2ae0ebf00c5dccf5880

Observation 1a02d2a3-5dbb-4a20-85be-16e009ade606 · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Silkie: Preference Distillation for Large Visual Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.053172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.053172Z digest=sha256:62d7a47e06b6f6798408d27db79993beefd01e2b42df70d1136f9d84d7a83551

Observation 6c7b39fa-091e-4f51-8c1b-4c81bdd9c969 · outbound

This paper cites Red Teaming Visual Language Models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Red Teaming Visual Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.078954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.078954Z digest=sha256:49a772c4ce16fa90e4d3fc2eed8f5c9d09d7ae11cde55c10570a39da6f05e529

Observation cc259b27-001f-40d1-a94d-c7a538e22db5 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.105233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.105233Z digest=sha256:cbf778c365a87716a2416a3a05c39bb40a3b80861ccd7e22f189c7e615b1ec45

Observation 307625ca-1c5d-45e0-99b1-1cbb7602ab7b · outbound

This paper cites Evaluating object hallucination in large vision-language models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Evaluating object hallucination in large vision-language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.113534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.113534Z digest=sha256:46222db3a3be7bde10c2d6323ba530c0c1c10e8c2052632d69a2fb0b61a6ff14

Observation 34a3a2f1-82cb-4415-b4fa-9ea99c0f61e3 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Evaluating Object Hallucination in Large Vision-Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.119255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.119255Z digest=sha256:b6cb44e764b250dacd56c5798d59caf0eb25327f8acf20016fbf32fc54fd9026

Observation f0cc6a6c-1c85-4340-914d-9337bc03a544 · outbound

This paper cites Mitigat- ing hallucination in large multi-modal models via robust instruction tuning.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Mitigat- ing hallucination in large multi-modal models via robust instruction tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.126001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.126001Z digest=sha256:e09b522fcd6df8e80eb7fa510c375daac4b2425dc81fae21934cce5dd625e303

Observation 54005270-c6d4-4418-a075-56be5888183f · outbound

This paper cites Visual Instruction Tuning.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Visual Instruction Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.131674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.131674Z digest=sha256:843fe1978efb6efad86fe846c9301d54bc7867a19722c71c62136408b4b9e2f5

Observation 0b0b0de4-b3d3-47d1-83ea-afb183c1f36d · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment MMBench: Is Your Multi-modal Model an All-around Player?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.140384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.140384Z digest=sha256:f745a8f97df16d823e8a74b06a38ab9f1ebbeff59908014a890c786739d6a8f4

Observation 3ba341f0-f496-4fb8-84c4-f4b44c4e0116 · outbound

This paper cites On the hidden mystery of ocr in large multimodal models, 2024.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment On the hidden mystery of ocr in large multimodal models, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.147293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.147293Z digest=sha256:e5c41b46891084a98211140008b26f36f071896ffcfe20aaf27366cad4ae45ed

Observation bd0373a9-ea0e-4aed-9ced-1b4573ec00b5 · outbound

This paper cites Deep learning face attributes in the wild.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Deep learning face attributes in the wild

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.158193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.158193Z digest=sha256:14c0ea1f3ffb4bb29e3555e36bee4d6bbdf8871f7a028250b62ac9e5a1fd48d4

Observation 9d58c1f4-8351-449d-adba-5a00f15124bd · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.166266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.166266Z digest=sha256:e0019f7b763f15c9f5ea04dccab5b293c752c7a8bf3d554bef4d7915fdc86508

Observation 497f175d-391a-4e59-ba2e-3abc07d35f41 · outbound

This paper cites Mathvista: Evaluating mathemati- cal reasoning of foundation models in visual contexts.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Mathvista: Evaluating mathemati- cal reasoning of foundation models in visual contexts

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.175244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.175244Z digest=sha256:2b93e9dad4eb1ef5293c94ef42dcf47068d12287a77436796f62b57717d1f5fe

Observation 806c8e65-cc4a-4347-88c4-7850e57ad0dd · outbound

This paper cites WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.181923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.181923Z digest=sha256:3c0c54b6417c10ccfdb2b3841538951a73acb6cc973dea5e2737c66382499c47

Observation 29f4eb57-c8e3-4128-94fd-c083cc3d04f7 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.187690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.187690Z digest=sha256:faea54047dfd108fb937b2f62be28ad23f54a48cdff08b018b147d764eda347d

Observation a5f2564d-012d-42b1-a222-5a08b7d1521d · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.200612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.200612Z digest=sha256:df1b43c8763b3b31fdff2a6d9640b76d6407bbe4111b41fdb765e76ee6fac4fd

Observation 2f099e37-bcc5-4b03-8d86-1c85b9f0d297 · outbound

This paper cites Infographicvqa.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Infographicvqa

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.206580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.206580Z digest=sha256:552beab3aa79d1eeff277b5d85ae09b3c19c1d59274b89af4972d416f9dae09a

Observation 50556b81-2ef0-404c-b359-2507d4d9174b · outbound

This paper cites Docvqa: A dataset for vqa on docu- ment images.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Docvqa: A dataset for vqa on docu- ment images

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.216767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.216767Z digest=sha256:771a9d4fc71cedd27818c8081a0fafd391aa77da4255169e8ec1131a4e60da77

Observation 43faba2b-9de1-401d-b4cb-5a5d719163de · outbound

This paper cites Gpt-4 technical report.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Gpt-4 technical report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.221793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.221793Z digest=sha256:0c8e38a76b96890a879a6408291dc3df36f75c4071102c55f1701c4bed8d917f

Observation a812eeb3-55c8-487c-ae51-de08261b7ab3 · outbound

This paper cites Towards vqa models that can read.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Towards vqa models that can read

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.230249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.230249Z digest=sha256:952624eec4e77c0ba77a2ffec31f7094a9db12833377b1a8e3ff506619ffb65c

Observation e956c9ac-fcff-4fb9-8157-1a9e98f75bfa · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.247240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.247240Z digest=sha256:ad7f8c421a1b493172541cf13a50a66f941008b65aae2dee727b3edd15d82a2f

Observation 2472ce56-6b13-4997-9b04-5c126aa820c4 · outbound

This paper cites Stanford alpaca: An instruction-following llama model, 2023.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Stanford alpaca: An instruction-following llama model, 2023

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.254953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.254953Z digest=sha256:1bae57ef070f455e13f33f3890c1383d0896793b383ad89777169216e085696f

Observation c3bb5cc7-1fc9-49da-ad03-d63dc6f87a74 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.291605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.291605Z digest=sha256:5655747e6137f63407c2d9c33c7165b8f5a972d03c865cc3bca4e98a5e82cf84

Observation 1fe8f71c-2019-44d6-80fb-a2c6a047ee68 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment LLaMA: Open and Efficient Foundation Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.322062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.322062Z digest=sha256:b5a2a1a7d580cd03634d09e91eeb00bfe6ac4e458c186113eb60d24da58511f4

Observation 743f2a8a-ae08-463b-8d0f-8a9b113cfb1d · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.351029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.351029Z digest=sha256:593fc9a71d8eb4525ec1a35c791dae88335aa01542a202e9ef1a774c01f407c5

Observation e2ead512-85c7-4613-9168-fcf320b02b7f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.380394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.380394Z digest=sha256:a62b705eccb9be3b5fac9f4b2f427a9e40b7dd6335e34f9eb80eb723658e6b72

Observation e8882991-f408-49fa-a646-9d5517477e26 · outbound

This paper cites $\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment $\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.415535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.415535Z digest=sha256:1c0f7457bf6054794951157cdfdd80d61ca9c17b5f7e8484454125d6ab09273a

Observation 650289be-35c1-41fc-b3ce-57969e223ca4 · outbound

This paper cites Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.428020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.428020Z digest=sha256:170d1002a6a1572ca92acc476f8db49b2a4e82d287d44f09314b185fcc59da3d

Observation 3c26313c-026e-4330-aa1a-b358d00c869b · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.433458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.433458Z digest=sha256:4c3fe24a5c333f66e8d68c6d2c871b7c8945f0b97844ddc738dcf6abc16bc17b

Observation 54fc93e1-c672-4bc6-bd01-f06f3c2cd15f · outbound

This paper cites ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.439693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.439693Z digest=sha256:bfe2a2ee5873707f2a95e75c82f2cb38208ceb5ca15014539670136ff1fe0709

Observation 867af584-15e8-4bc6-9321-a6d55ec46471 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.446156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.446156Z digest=sha256:4be0a851f5f64e59b32dda6517df254e721fc0cb6d61158812b021ea74d71af8

Observation 013cdb5e-9deb-4336-bb59-f5fe0855a149 · outbound

This paper cites Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi, 2024.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi, 2024

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.452017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.452017Z digest=sha256:8f4aa5de45cf2655ed6421db86d377b606151f30797c1cffe6cd78abb6950dc0

Observation c0652f30-b4bd-4230-acb9-e8430b4f3d87 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.459008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.459008Z digest=sha256:74f7094d363ece09c3da63c8aa6468eaccb54ea9dc73f4a8ec9b0cacdb8cec48

Observation a02aa576-5f30-4343-a073-97a557b4733c · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.467364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.467364Z digest=sha256:03efaf2c62b84a8485e71376ac6967c787f36174e38995a64d6b024af692d7fd

Observation e25afa4c-d9d2-4085-aeb8-7986d680207b · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabil- ities.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Mm-vet: Evaluating large multimodal models for integrated capabil- ities

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.473897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.473897Z digest=sha256:cbebf5a4657ac8330dc44e5f41c935ca0aefe7e44eac4b1b13025d091050acbb

Observation d4229abe-225d-4a77-996c-286beb3372a5 · outbound

This paper cites Self-Generated Critiques Boost Reward Modeling for Language Models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.481768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.481768Z digest=sha256:2b3989821f33d9f01b6cf56467d48ef1e5ee3b10dfac37e2aa4640170b23ae67

Observation 818779c1-2700-4ed9-b07f-b56ceb686f4b · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.488488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.488488Z digest=sha256:93773f3a79b1df5f5d9d94efe423124b681597e8e6ba6f566234aaf0ab6d0b6e

Observation 35e5ba8f-43a4-44c6-8c8c-ad4b8ed4eb98 · outbound

This paper cites Anyattack: Self-supervised generation of targeted adversarial attacks for vision-language models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Anyattack: Self-supervised generation of targeted adversarial attacks for vision-language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:53.060100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.494838Z digest=sha256:3b98958200453aae8c004a516c734d00b36c5fedcc2b9cf5c16f807f5e568586

Observation b973462c-04d2-4ef6-9c00-0bdbeb016073 · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.500362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.500362Z digest=sha256:9b72c993824b320a0645aa0ffa6923e2398901975982403bdb53e1dbc04eea8c

Observation ad2f3176-c300-4042-a77f-2367a292f710 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.508069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.508069Z digest=sha256:b2d8e9bfe85c71ab38d79c32711507d792fadc8a9db6cdddbed1e63d4a1c1347

Observation f8a1bd3f-8f1a-4e44-ac29-90882517e885 · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.516471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.516471Z digest=sha256:f1ec3b92689973dda2eb80c14729df2a41843756dfaf943f96e05bd3960d7d83

Observation da4ce379-8a1f-4957-b71b-6372f400c90f · outbound

This paper cites Debiasing Multimodal Large Language Models via Penalization of Language Priors.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Debiasing Multimodal Large Language Models via Penalization of Language Priors

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.524073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.524073Z digest=sha256:f91e3dafb0b062f5b5d39990150d75c670935ed5f2b92fee0191c96befe61dc5

Observation 243109b7-9c10-4e28-9420-bcd07a671154 · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.530539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.530539Z digest=sha256:4991d03367e68be96c5580b55a377c4c7bcac1a41ba8e5ef5a1fa4ef85e7338f

Observation 45ddd3ac-d608-4fad-8f74-2418c6c694fe · outbound

This paper cites Multitrust: A comprehensive benchmark towards trustworthy multimodal large language models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Multitrust: A comprehensive benchmark towards trustworthy multimodal large language models

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:53.039442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.541250Z digest=sha256:723eeab1cac38e44ee01d915f706a9e2edc02c28d68217cc64c017e9c5e93964

Observation 73da6a09-7f01-4c25-b393-89a6b3c6cfe2 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.551232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.551232Z digest=sha256:3ce3d7f4f1dd36caa7f16b9abed9258a4442c24b3d20b2e8fd432bc2cc350bc5

Observation bf4c39da-e164-4a99-8e60-4ee06a4b2c4c · outbound

This paper cites Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:50.561901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:50.561901Z digest=sha256:9d5db9ab9e32cf4f4d1e0a264419693a8eff21019820e5959361d48417588b30

Observation 7cddd6f3-8f5a-40a4-b218-d727338672f6 · outbound

This paper cites an unresolved cited work.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:23:53.018387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.591600Z digest=sha256:cb07cefd53c38a2b8ad38b21b3b86bd9c0d3017e2067ca1729271b4d4f811d89

Observation df2f2f70-22d7-4e94-90ac-4e278ccfd0d3 · outbound

This paper cites Minimize errors and misleading information in object relationship descriptions.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Minimize errors and misleading information in object relationship descriptions

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:52.996592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.614571Z digest=sha256:cc8d35a4075994bac82e53c32543f6988791792968cf9e6f6e89128473b298d6

Observation 0af65dce-5e38-4900-9fac-e1c181639057 · outbound

This paper cites an unresolved cited work.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:23:52.973227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.637416Z digest=sha256:e03f135d90db62eb9de16c5a71c1902cc31b5b34bb572a0e81ad1e5dec215787

Observation feaa45a4-8089-4844-9586-8c5f4934f993 · outbound

This paper cites Rating Scale: • Severely Inaccurate: Major errors in object descriptions, relationships, or attributes, or references to non-existent objects.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Rating Scale: • Severely Inaccurate: Major errors in object descriptions, relationships, or attributes, or references to non-existent objects

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:52.952674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.658532Z digest=sha256:99903ca07ca643d298c70dc2ace0ec0ca2344acaf1fdcfc4fc4dd1ab366c937e

Observation e9b2b42c-bec4-4786-b0d2-8fd8d079d772 · outbound

This paper cites an unresolved cited work.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:23:52.930281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.668279Z digest=sha256:6855ec5ab630cfd7ae765aedaf26d0bccf86b75fc3cfeb1a2450be1d72d561d1

Observation bba01542-4ba8-4b6f-b818-9f141c17f1cd · outbound

This paper cites an unresolved cited work.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work

Reference 90

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:23:52.905978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.674125Z digest=sha256:749e62cc53a6f8e8061d067444ce5313280e28b75cf725539687a2c3a76a11a4

Observation 641145cc-ad23-4261-b4ce-2e26fd23b2dd · outbound

This paper cites Rating Scale: • Not Helpful: The response does not address the user’s prompt, providing entirely irrelevant information.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Rating Scale: • Not Helpful: The response does not address the user’s prompt, providing entirely irrelevant information

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:52.881053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.680665Z digest=sha256:c92e5a6aaa2209fcc7cd1a63fdff7156044b68316b112ba9095d78e7fea90a4f

Observation 4e6f05a8-7ba5-4de2-9c92-90db14272085 · outbound

This paper cites an unresolved cited work.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:23:52.859324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.687335Z digest=sha256:bd0ae575f5f72f5d643a265c7111443521cb320a2b35e4316284d2e972cb8a25

Observation 4472717b-5d78-4143-bba7-a6b160ade966 · outbound

This paper cites an unresolved cited work.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:23:52.837576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.694405Z digest=sha256:8395ecb84a2eb34436148fe2cbc98a9a8b62b6aa564ffeb221689a2ab42ccb86

Observation 8a14ad96-75a4-4276-b9cf-fcb3b8ce1eb1 · outbound

This paper cites an unresolved cited work.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:23:52.817360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.699636Z digest=sha256:8e6bb0e73b2192a2cbfc83c9e0f5d6a30350fa1fc7e27ca578aa64bea5d47f37

Observation 3fa1a848-13a6-4ce2-9f24-e3cae8da8789 · outbound

This paper cites Rating Scale: • Highly Unsafe/Privacy Infringing/Unfair/Harmful: The response contains severely in- appropriate content that violates ethical standards.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Rating Scale: • Highly Unsafe/Privacy Infringing/Unfair/Harmful: The response contains severely in- appropriate content that violates ethical standards

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:52.795219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.708433Z digest=sha256:20cb5ab9b13c6a0e4e70aae9ca3ca7192117928f756038b64183deee21278758

Observation 5cd7bab8-0691-4bf4-8681-7493b6f3778d · outbound

This paper cites an unresolved cited work.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:23:52.670958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.716816Z digest=sha256:02f527af9b1b193bb30181d4f6589a06c2cac05688ea158806c8280a007256f5

Observation c8387972-2eaf-41bc-9039-9a889791bda9 · outbound

This paper cites an unresolved cited work.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:23:52.605802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.725674Z digest=sha256:95603ea6899183f3c7a9e05961cd4cccfb64b2da95dc353c8367c012dacc01d0

Observation f7f1d375-68cd-4e58-9466-016ab5c1d92c · outbound

This paper cites an unresolved cited work.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:23:52.582997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.731660Z digest=sha256:340f2e133387105f3d5434e3ca02c99813b80bf3047850512eb0d23c17180e30

Observation 345db11a-f047-4552-9a9a-186070e0c9d4 · outbound

This paper cites If a tie occurs, provide a negative example (for multiple-choice, offer an incorrect answer; for long text, modify the content to include erroneous information).

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment If a tie occurs, provide a negative example (for multiple-choice, offer an incorrect answer; for long text, modify the content to include erroneous information)

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:52.559238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.738265Z digest=sha256:c956e82c7469980c3c3676e17d20c5b58a03aa4341230b3acb6121ba091a9116

Observation ff238f26-110a-4a13-9526-d065d338e51a · outbound

This paper cites Please describe this image,.

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Please describe this image,

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:52.526170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T18:23:50.744223Z digest=sha256:f62cd58a8995ac5eea6118888c45c56b9fdc32f00d995d4f4d68326198870d37

Pith citing papers

Observation 6617a416-055b-4156-8a1b-15a070b0c255 · inbound

A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges cites this paper.

A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 266

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:23.962027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:17:23.962027Z digest=sha256:332f72a3554a7cbe3a42916f818a68be27fcb87f401566421337f756cc49ae12

Observation da679117-4f8e-4591-b954-03f590d5c904 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 283

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.493833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:6d7bcb109a1231c5494d7df2be1e9db4c1f2331d41d56620f15c435806336801

Observation f7f675fc-ef38-4d3a-9d70-8a0f7dac1cb9 · inbound

LAD-Reasoner: Tiny Multimodal Models are Good Reasoners for Logical Anomaly Detection cites this paper.

LAD-Reasoner: Tiny Multimodal Models are Good Reasoners for Logical Anomaly Detection MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T12:28:02.977973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:28:02.977973Z digest=sha256:37bf451d4c9c3ae24792eb72dd6b5478b3985e5ffe28342e2665042e6f04da50

Observation ae56dbab-7a2e-43bd-8dcc-caf3f8bb724f · inbound

Generative RLHF-V: Learning Principles from Multi-modal Human Preference cites this paper.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:55.253963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:55.253963Z digest=sha256:6eaf719249e473c6a9ecd204e20c176595e0256249ff1db75046839a67a276ee

Observation d5852424-7ed5-485c-ba4e-571f5d9b5900 · inbound

VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL cites this paper.

VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:19.926347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:19.926347Z digest=sha256:a9badbfa84fb9e049c030e1d6b58e923cd101068c30c0d98e95fb209afff7613

Observation af67a684-595c-4119-b025-89b1851ed955 · inbound

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking cites this paper.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.702270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.702270Z digest=sha256:35b743053047aba122845f3e2f967eb1fd80f1ffa5c9ee380ca0e987b5eb6f96

Observation 9b0ec985-fb2e-429e-bbeb-c23793eb36c8 · inbound

RACE-Align: Retrieval-Augmented and Chain-of-Thought Enhanced Preference Alignment for Large Language Models cites this paper.

RACE-Align: Retrieval-Augmented and Chain-of-Thought Enhanced Preference Alignment for Large Language Models MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:21:40.615790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:21:40.615790Z digest=sha256:0981b45b391709ad34765df4c13fe51f8a737bb10a96c4cd46808d957e40d5c7

Observation bf83dd4e-5a6b-4fe9-b9fc-de3de6805c90 · inbound

Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment cites this paper.

Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T11:22:50.342524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:22:50.342524Z digest=sha256:437822453d96c5ac2e54d5c166a66891e968f3a17da85e1e8a536fb8a6bc5cb9

Observation 7508d138-a168-4880-8c5b-b1f4a844fd42 · inbound

Activation Reward Models for Few-Shot Model Alignment cites this paper.

Activation Reward Models for Few-Shot Model Alignment MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.531667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.531667Z digest=sha256:fe1d07b4e80f60579646a5fc70b20570a40ea4e8c780328d9dc523c8d5031fef

Observation ad249ce9-44b4-4525-ad20-ccc919238823 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:54.913563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:54.913563Z digest=sha256:4d5134d8cfd6d4610139331f5cb76b78d8c7d6aa54ff931146573d0f03a3cc5f

Observation 9d7afc8e-51e1-4898-9593-28a73edb2c14 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.964837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.964837Z digest=sha256:c617bde5e0444d7777a295a2a5793b8cb50fac9725ec11d43aa853f91b3f6656

Observation 1e71b9a7-b5a1-432e-9845-65efc072eef0 · inbound

Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization cites this paper.

Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T23:11:32.187646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:11:32.187646Z digest=sha256:8abc168b1bead3dcd5b2c3c2a559137fa90c545af7529c6ef0dc151c879568ea

Observation 91815a4b-498f-4e7f-a4c8-c2af1f570f6b · inbound

NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation cites this paper.

NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:40:53.024462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T04:39:58.296388Z digest=sha256:bd51e01f7f154dbf4f395869b96905808f1103afef494043122a922b46203cb9

Observation 4fd40dad-e259-44af-8477-a5422eadcb23 · inbound

Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation cites this paper.

Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:51:03.556344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T15:17:39.344348Z digest=sha256:66186c8dce01e1bb169d6b839df41acd03999ea167ce2d5641a921514ed4d597

Observation 8d23e198-328e-4f1d-82ac-3dc0cd04cfd3 · inbound

Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts cites this paper.

Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.826529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T05:33:20.557526Z digest=sha256:1e787db0e6f204e54c88d40adb0cf0ae03ae909107dd717af2fb7236c92a752c

Observation dd93ae88-6690-4b52-b09e-642ad95b154e · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.467982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:763805723ec138ef8bc79b5d0e886b5f0a15c715b7a3b4b6c254882807d55685

Observation 6868191b-257a-47b9-9050-a49334d09a96 · inbound

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification cites this paper.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.331700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:9021b58e86235f305d4e18ace1771309cf43afbf2ed2e68f14982fbe11257747

Observation b80cfbf5-0edc-4e68-925f-24735fec3fbd · inbound

Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos cites this paper.

Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:28:11.975833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T10:26:30.661042Z digest=sha256:062986c91ad678edeff04176676c2d1cc5490a3cef8aba289a16d5ad8518f186

Observation 79bdc297-7bde-4070-af2b-c81a42a28719 · inbound

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning cites this paper.

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:36:10.480475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T06:34:57.483234Z digest=sha256:bffcb0983f54065b4389a635ac6b1468af940a63de070a207e2488f14ce17818

Observation b7625811-fe53-4eb1-88e4-c186deff03ab · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.935715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:b4c31548ac72165ba012e42f07b641d63678af2dc92d4052978aa8b0fb4e0a12

Observation dddedaac-32b3-492e-8e54-1c748ab6b628 · inbound

Constitutional On-Policy Safe Distillation cites this paper.

Constitutional On-Policy Safe Distillation MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:25.457078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T11:47:14.135793Z digest=sha256:eb680699c52289c5044c2cdbfef0ae205f0b2cd29a54cbe68dd89254b7f86bcc

Observation d6d659d7-c35c-4a1d-bed0-a7be6463fb99 · inbound

DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving cites this paper.

DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:27:25.547848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T18:58:26.701929Z digest=sha256:8e503df1eaf333224300947111cbcef7d16c3dc3f5f4691b949cdca0c9b25366

Observation 9fe2c2b3-6b3b-40da-ae05-615b3fd39c35 · inbound

Kwai Keye-VL-2.0 Technical Report cites this paper.

Kwai Keye-VL-2.0 Technical Report MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:37.083717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T13:53:10.352603Z digest=sha256:264c77fee270cbf8e5f1b960d2462649fc7b33413a0e8626055da634d12de294

Observation af9dee90-b69b-43f7-91f5-af013c095807 · inbound

Multimodal Reward Hacking in Reinforcement Learning cites this paper.

Multimodal Reward Hacking in Reinforcement Learning MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:a2fa0d32ab1873bc533995377eac026acddf8bc100895604ee741cf08d3146e5

Observation 63b262a9-8058-415e-821c-e6de659462c8 · inbound

Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs cites this paper.

Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T04:15:39.381986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:15:39.381986Z digest=sha256:f755591ff4c7456e1094565453af506eb2b5767d86b36e80e68570b2b25ad4ae

Observation 2bc85b1a-3960-4f20-bac0-ddfd98d4493f · inbound

The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal Guardrails cites this paper.

The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal Guardrails MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T00:20:22.395848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:20:22.395848Z digest=sha256:22c6a9ce95c9e60fa2d35310673967dd788d48121195e92d275e3815fcd4937e

Observation c10d224c-4731-4766-8304-f2d4adcf8da4 · inbound

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning cites this paper.

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T14:24:49.741420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:24:49.741420Z digest=sha256:b4f992bf978a33261c9ae7a08fa7a7ea1c9fa3acff9f35afafd96894547cbbfd