Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T18:23:50.744223Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 97 of 97 outbound references and 27 inbound Pith citation observations for arXiv:2502.10391.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T18:23:50.744223Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:28:02.977973Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T04:27:37.082400Z
97 of 97 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8f1b187c-2e91-4e11-8100-8651fc77a71f · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2981123-6329-48b3-a413-4224712e0d6b · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Pixtral 12B
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f9838e1-e980-48ff-95a1-a0e08ad56315 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Direct Preference Optimization with an Offset
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 499c3639-92f4-4e37-a1e9-9ec70e78aeee · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Vqa: Visual question answering
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef901459-183a-457e-afde-9f2ece4289a4 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dcc7400-3811-4bfa-9905-b9f83fd47c1f · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment TouchStone: Evaluating Vision-Language Models by Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6300708-95dd-4364-98a6-de3baf853e30 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a837664d-9ae9-4e18-8bcf-5152b23862da · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Language models are few-shot learners
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1236c06-b979-4832-ba05-170e0a96013e · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f907f7d-2050-48be-abc9-27c2a2d25941 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff0a1798-c851-430b-a8bf-b6b3fc6034c5 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment WebSRC: A Dataset for Web-Based Structural Reading Comprehension
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f52f174-8dd9-40e8-a87d-96932e6cb91d · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1dcd352-3243-4d99-8347-784d4cd096b5 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e14d5402-2079-4dc0-97bb-ceb0a69e7e40 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e9ef6e6-ad4d-4ddd-ac35-445bbf735d62 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0ad6dc9-7477-4748-a754-7959c63eb6e4 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment NVLM: Open Frontier-Class Multimodal LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf019ee6-6568-4fcc-b033-6447fa3d5023 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ace44fbc-4616-4b1d-8ea0-67038fb73abd · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 181ac02a-f004-420c-9145-ee9aa73774f4 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63fabce5-e990-4e70-aed4-1c56716c81ea · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eaab4af-c665-4c62-9349-65ff4b72a787 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c292084b-02c9-477f-b785-df96ca3191a2 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e71ccea0-7fe7-42a3-88eb-843910d7a257 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d442750-fa52-491e-a38f-d1a8e0582475 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment BLINK: Multimodal Large Language Models Can See but Not Perceive
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6896c9b0-4cf1-422d-bf1f-0956f20e2c7d · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Infimm- eval: Complex open-ended reasoning evaluation for multi-modal large language models, 2023
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 660fabdc-fa88-4af2-8055-43e8eb1228fb · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cc59590-50e3-48f3-ad73-bba894f76022 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment VLSBench: Unveiling Visual Leakage in Multimodal Safety
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15a70143-3fe8-4ccc-b980-27fc3a716ce4 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Mistral 7B
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52047ede-8d17-4efc-b77e-38bb3fb5d9be · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment A diagram is worth a dozen images
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b855676b-334a-40de-9c4b-032a8a83edf5 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93133f22-5bd8-43e2-8e97-4dbdc2bae508 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Otter: A Multi-Modal Model with In-Context Instruction Tuning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 664a80c7-ddfa-489f-a597-752f7529b8da · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment LLaVA-OneVision: Easy Visual Task Transfer
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 611c7e4c-0ec4-4ffb-99de-ecdc32dfe686 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29c7f356-9e0a-4313-a05a-78815ee61366 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment SEED-Bench-2: Benchmarking Multimodal Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7dd9c67-2de6-4825-8f10-5c3a29621257 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07bc4a16-2f98-491a-9d28-29b500f400ec · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a02d2a3-5dbb-4a20-85be-16e009ade606 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Silkie: Preference Distillation for Large Visual Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c7b39fa-091e-4f51-8c1b-4c81bdd9c969 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Red Teaming Visual Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc259b27-001f-40d1-a94d-c7a538e22db5 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 307625ca-1c5d-45e0-99b1-1cbb7602ab7b · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Evaluating object hallucination in large vision-language models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a3a2f1-82cb-4415-b4fa-9ea99c0f61e3 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Evaluating Object Hallucination in Large Vision-Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0cc6a6c-1c85-4340-914d-9337bc03a544 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Mitigat- ing hallucination in large multi-modal models via robust instruction tuning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54005270-c6d4-4418-a075-56be5888183f · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Visual Instruction Tuning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b0b0de4-b3d3-47d1-83ea-afb183c1f36d · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment MMBench: Is Your Multi-modal Model an All-around Player?
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ba341f0-f496-4fb8-84c4-f4b44c4e0116 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment On the hidden mystery of ocr in large multimodal models, 2024
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd0373a9-ea0e-4aed-9ced-1b4573ec00b5 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Deep learning face attributes in the wild
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d58c1f4-8351-449d-adba-5a00f15124bd · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 497f175d-391a-4e59-ba2e-3abc07d35f41 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Mathvista: Evaluating mathemati- cal reasoning of foundation models in visual contexts
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 806c8e65-cc4a-4347-88c4-7850e57ad0dd · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29f4eb57-c8e3-4128-94fd-c083cc3d04f7 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5f2564d-012d-42b1-a222-5a08b7d1521d · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f099e37-bcc5-4b03-8d86-1c85b9f0d297 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Infographicvqa
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50556b81-2ef0-404c-b359-2507d4d9174b · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Docvqa: A dataset for vqa on docu- ment images
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43faba2b-9de1-401d-b4cb-5a5d719163de · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Gpt-4 technical report
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a812eeb3-55c8-487c-ae51-de08261b7ab3 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Towards vqa models that can read
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e956c9ac-fcff-4fb9-8157-1a9e98f75bfa · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Aligning Large Multimodal Models with Factually Augmented RLHF
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2472ce56-6b13-4997-9b04-5c126aa820c4 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Stanford alpaca: An instruction-following llama model, 2023
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3bb5cc7-1fc9-49da-ad03-d63dc6f87a74 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fe8f71c-2019-44d6-80fb-a2c6a047ee68 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment LLaMA: Open and Efficient Foundation Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 743f2a8a-ae08-463b-8d0f-8a9b113cfb1d · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2ead512-85c7-4613-9168-fcf320b02b7f · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8882991-f408-49fa-a646-9d5517477e26 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment $\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 650289be-35c1-41fc-b3ce-57969e223ca4 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c26313c-026e-4330-aa1a-b358d00c869b · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment LLaVA-Critic: Learning to Evaluate Multimodal Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54fc93e1-c672-4bc6-bd01-f06f3c2cd15f · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 867af584-15e8-4bc6-9321-a6d55ec46471 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 013cdb5e-9deb-4336-bb59-f5fe0855a149 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Mmt-bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask agi, 2024
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0652f30-b4bd-4230-acb9-e8430b4f3d87 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a02aa576-5f30-4343-a073-97a557b4733c · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e25afa4c-d9d2-4085-aeb8-7986d680207b · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Mm-vet: Evaluating large multimodal models for integrated capabil- ities
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4229abe-225d-4a77-996c-286beb3372a5 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 818779c1-2700-4ed9-b07f-b56ceb686f4b · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35e5ba8f-43a4-44c6-8c8c-ad4b8ed4eb98 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Anyattack: Self-supervised generation of targeted adversarial attacks for vision-language models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b973462c-04d2-4ef6-9c00-0bdbeb016073 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad2f3176-c300-4042-a77f-2367a292f710 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8a1bd3f-8f1a-4e44-ac29-90882517e885 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da4ce379-8a1f-4957-b71b-6372f400c90f · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Debiasing Multimodal Large Language Models via Penalization of Language Priors
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 243109b7-9c10-4e28-9420-bcd07a671154 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45ddd3ac-d608-4fad-8f74-2418c6c694fe · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Multitrust: A comprehensive benchmark towards trustworthy multimodal large language models
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 73da6a09-7f01-4c25-b393-89a6b3c6cfe2 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf4c39da-e164-4a99-8e60-4ee06a4b2c4c · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cddd6f3-8f5a-40a4-b218-d727338672f6 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation df2f2f70-22d7-4e94-90ac-4e278ccfd0d3 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Minimize errors and misleading information in object relationship descriptions
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0af65dce-5e38-4900-9fac-e1c181639057 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation feaa45a4-8089-4844-9586-8c5f4934f993 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Rating Scale: • Severely Inaccurate: Major errors in object descriptions, relationships, or attributes, or references to non-existent objects
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e9b2b42c-bec4-4786-b0d2-8fd8d079d772 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bba01542-4ba8-4b6f-b818-9f141c17f1cd · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 641145cc-ad23-4261-b4ce-2e26fd23b2dd · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Rating Scale: • Not Helpful: The response does not address the user’s prompt, providing entirely irrelevant information
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4e6f05a8-7ba5-4de2-9c92-90db14272085 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4472717b-5d78-4143-bba7-a6b160ade966 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8a14ad96-75a4-4276-b9cf-fcb3b8ce1eb1 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3fa1a848-13a6-4ce2-9f24-e3cae8da8789 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Rating Scale: • Highly Unsafe/Privacy Infringing/Unfair/Harmful: The response contains severely in- appropriate content that violates ethical standards
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5cd7bab8-0691-4bf4-8681-7493b6f3778d · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c8387972-2eaf-41bc-9039-9a889791bda9 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f7f1d375-68cd-4e58-9466-016ab5c1d92c · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Unresolved cited work
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 345db11a-f047-4552-9a9a-186070e0c9d4 · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment If a tie occurs, provide a negative example (for multiple-choice, offer an incorrect answer; for long text, modify the content to include erroneous information)
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ff238f26-110a-4a13-9526-d065d338e51a · outbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Please describe this image,
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6617a416-055b-4156-8a1b-15a070b0c255 · inbound
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 266
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da679117-4f8e-4591-b954-03f590d5c904 · inbound
From System 1 to System 2: A Survey of Reasoning Large Language Models MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 283
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f7f675fc-ef38-4d3a-9d70-8a0f7dac1cb9 · inbound
LAD-Reasoner: Tiny Multimodal Models are Good Reasoners for Logical Anomaly Detection MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae56dbab-7a2e-43bd-8dcc-caf3f8bb724f · inbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5852424-7ed5-485c-ba4e-571f5d9b5900 · inbound
VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af67a684-595c-4119-b025-89b1851ed955 · inbound
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b0ec985-fb2e-429e-bbeb-c23793eb36c8 · inbound
RACE-Align: Retrieval-Augmented and Chain-of-Thought Enhanced Preference Alignment for Large Language Models MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf83dd4e-5a6b-4fe9-b9fc-de3de6805c90 · inbound
Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7508d138-a168-4880-8c5b-b1f4a844fd42 · inbound
Activation Reward Models for Few-Shot Model Alignment MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad249ce9-44b4-4525-ad20-ccc919238823 · inbound
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 133
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d7afc8e-51e1-4898-9593-28a73edb2c14 · inbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e71b9a7-b5a1-432e-9845-65efc072eef0 · inbound
Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91815a4b-498f-4e7f-a4c8-c2af1f570f6b · inbound
NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4fd40dad-e259-44af-8477-a5422eadcb23 · inbound
Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8d23e198-328e-4f1d-82ac-3dc0cd04cfd3 · inbound
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dd93ae88-6690-4b52-b09e-642ad95b154e · inbound
Building a Precise Video Language with Human-AI Oversight MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6868191b-257a-47b9-9050-a49334d09a96 · inbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b80cfbf5-0edc-4e68-925f-24735fec3fbd · inbound
Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 79bdc297-7bde-4070-af2b-c81a42a28719 · inbound
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b7625811-fe53-4eb1-88e4-c186deff03ab · inbound
Toward Native Multimodal Modeling: A Roadmap MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 143
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dddedaac-32b3-492e-8e54-1c748ab6b628 · inbound
Constitutional On-Policy Safe Distillation MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d6d659d7-c35c-4a1d-bed0-a7be6463fb99 · inbound
DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9fe2c2b3-6b3b-40da-ae05-615b3fd39c35 · inbound
Kwai Keye-VL-2.0 Technical Report MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation af9dee90-b69b-43f7-91f5-af013c095807 · inbound
Multimodal Reward Hacking in Reinforcement Learning MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63b262a9-8058-415e-821c-e6de659462c8 · inbound
Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bc85b1a-3960-4f20-bac0-ddfd98d4493f · inbound
The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal Guardrails MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c10d224c-4731-4766-8304-f2d4adcf8da4 · inbound
SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.