Pith. sign in

Paper Citation Record · LEDGER

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation

As of 13 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2412.04903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04903 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:15:45.548297Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T15:05:21.907878Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T15:08:02.074896Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c20a20e2-9235-4bb3-8daf-3390873d5016 · outbound

This paper cites GPT-4 Technical Report.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.192262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.192262Z digest=sha256:d73e31165db0541c7af2d136b2189d61432e026bf670eebc44a2b67ff7ab5b1e

Observation cfa5f7cc-b9ea-44ce-b6da-5432b3cd586a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.198801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.198801Z digest=sha256:7f58b578f1ee3bb09110fe8b4836598ae803087d70d7eb2777e936dd8d84efe1

Observation 6075f1ff-8257-44bd-b27f-71af6e04cae8 · outbound

This paper cites Introducing our multimodal models, 2023.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Introducing our multimodal models, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.204734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.204734Z digest=sha256:11608bce3b20bf8ce39465714aa0165da727575ac89d0019fbb8f490b38e0d09

Observation c4122eb2-7ede-4595-85b8-ff6065f0956c · outbound

This paper cites Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.210468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.210468Z digest=sha256:05db72a364c95ad7208751ecb8d455dcf5d51762d2f5dbccde50c5be61a7749f

Observation 7531db02-af9a-4a18-b63e-6703c14ef1b4 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.215944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.215944Z digest=sha256:d47f3632774d361b8ff446cb6a6a64d2ab74a769733c872195f1c8a44daa448c

Observation 7bfe13d4-debb-4035-b2ac-03525b209723 · outbound

This paper cites PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.221696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.221696Z digest=sha256:1f4e32fbede34028fbe4e4bf26f4a67e3582873bd62399062206db8111a7316c

Observation 064bd64b-b1a0-49fd-9e69-17caf8363ac1 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.228427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.228427Z digest=sha256:b86519240aa81d9ffed1366f6960d1380e60d08ffaafa822682b7e575b82187c

Observation afab7f37-2681-4a98-b2e9-bb74ac27b656 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.235243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.235243Z digest=sha256:26496f62e2562bf6933dc5976b0e7a4ada18a8514c796fed1dc23529b68a7787

Observation 900be8f2-8b35-4dfd-9cba-ace4b33dd605 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.241288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.241288Z digest=sha256:077e816c80b6c370c4111de6005362791ae8c89be8e53cf88a66dfbc54f36708

Observation 9d0e3942-1d76-4462-b1b4-731178172ab8 · outbound

This paper cites Enhancing Large Vision Language Models with Self-Training on Image Comprehension.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Enhancing Large Vision Language Models with Self-Training on Image Comprehension

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.246901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.246901Z digest=sha256:84f7bf9b10e5c8cd74bc5c1a212d58d8c7f10c0ded681e641ef40694a1ad6788

Observation 278c5dde-25c1-4f08-abba-8d176d3e2e4a · outbound

This paper cites What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.252677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.252677Z digest=sha256:53709e281fee90e4030a67284d7f9aab49ffdf8cc6079eebf35fd8d8e1177444

Observation 399e962d-1baf-4ad7-9d1d-5d50926b6b24 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.879880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.259223Z digest=sha256:11d06bb17ed2562f86c4836a6230ec88c9facf5848545e50eac44db603c18f70

Observation be302331-927a-466f-8a70-b0f5d69dee51 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.264403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.264403Z digest=sha256:dece29a60beb4a644db86e65144c496118f22146a92e8e8496113c0ce9bfb256

Observation 869f7c6d-08f2-4dd3-8008-560deda00da0 · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.270537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.270537Z digest=sha256:8ee2b77e72908e10b9331424a6156fcdfdaf86d93379dcd6dc04d028a0f86947

Observation 47b76b57-8d5f-4f02-a2b6-6ebd0eaad57c · outbound

This paper cites Efficient Multimodal Learning from Data-centric Perspective.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Efficient Multimodal Learning from Data-centric Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.276775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.276775Z digest=sha256:1bebdfa5001f95d25cd256b255dcf2123d5c27d00d6de96b1facbc71d44ea93b

Observation 14d1d624-dfba-4e20-94fb-6fb692d2dfdf · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation LoRA: Low-Rank Adaptation of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.282235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.282235Z digest=sha256:718a2f6c7205f7ade77d9bc1ea3df4b59e792d92d9825d0aa47195946ee3c598

Observation 173b546f-be86-4f94-a87c-7f56937a687a · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.288162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.288162Z digest=sha256:05616d1dac932bfcc7d3e5a83225bbf31a43ee6a0c25e9a6eb50ff28df40346d

Observation c0ca5f28-10e4-461c-91b1-241142baa4e9 · outbound

This paper cites LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.294591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.294591Z digest=sha256:90c21e520f8da6f8f4a5e30e23d9b57582a7f2749403608d06ec5cd3a4dd4998

Observation f8238795-5298-42c2-962a-696a6a35e620 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.300584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.300584Z digest=sha256:e54119ac29c6b1e9a7013f2512cc2748ddf0afeeff678700b9f3297fd7e93c5f

Observation 752bf7ef-9dcf-48ea-a232-c45f7efea877 · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Silkie: Preference Distillation for Large Visual Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.306173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.306173Z digest=sha256:ddcb0dfde61a46e85f46ae156e6495b1e4c4d3e0863d80ff90719189c3936021

Observation f2837536-35ff-452f-816f-8af47781f5a6 · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.311788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.311788Z digest=sha256:74e94541542e4421c155dd8de46126997d07f11ea930a80ebf9b20700d57055d

Observation 50dcb1f3-f1c9-4d1a-8f0d-c8ef8a333b65 · outbound

This paper cites Red Teaming Visual Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Red Teaming Visual Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.317341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.317341Z digest=sha256:f4382ed6196e82a0a0abb44a00fa346309bfe26cb90d69a27cad59122f9e9972

Observation 1a7c7d87-f8d3-4c01-9b31-ecfd53584c39 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Evaluating Object Hallucination in Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.322691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.322691Z digest=sha256:edba2579696c59507d066106dd00aa7fc3017765c2684c35e6fffbaca1cdcdfa

Observation ff5a77d4-9051-48a1-b6fe-b71348d67325 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.328671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.328671Z digest=sha256:7832c1dfc2a8cb5f1ebc6a87ce01e6c3ba4febc86f01fe45afdf8ddbf6602d7e

Observation a89266fe-bff0-40d3-badf-f235576c5d77 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Improved Baselines with Visual Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.334538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.334538Z digest=sha256:f5656d46d87e04ae9610c4aa119250e9b018fa331c6cc0c1f777e45f0767ba1f

Observation 969ef877-2a1a-4ec9-b1f6-bb1c97f6f50a · outbound

This paper cites Visual instruction tuning.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Visual instruction tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.341133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.341133Z digest=sha256:d8d051088837fdabd2f84a1aacf98a7fd632fda8b0ab4d9451b74621248695b4

Observation e2418008-9e24-4709-b8b6-79ba61956036 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation A Survey on Hallucination in Large Vision-Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.347819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.347819Z digest=sha256:b48738814b5b568f8ed77ee2915831c7bf1b3dd03013095a6395dd2df6ebdc19

Observation 689856cc-c6a7-4f5a-9617-d0f9a1bcdc0d · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.850399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.353023Z digest=sha256:dc5f7da2a09cb6c36068c247c346e72bc7e6cbc394e81c7a75c6b7fa4da5a668

Observation 9bc3d518-bdfe-4e85-a664-19494be0f3b7 · outbound

This paper cites Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.359223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.359223Z digest=sha256:e4071f0a4804853a4a61e719c43811920f777255ff39652bd0fcee4051155377

Observation 3c91929f-c41c-443e-bb96-11fdea8b1140 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.364993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.364993Z digest=sha256:51e02d4c7008949413954e0fd6f0f166f90aefa99a7a2308e9a56e76265ff617

Observation 5bcec186-4edc-48ce-a99a-56c5c1b823e5 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation DINOv2: Learning Robust Visual Features without Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.371287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.371287Z digest=sha256:a31cd2bcd36e381e90759da652c67dc7c91ef4c6cedd78baa540ec3abd1d9ed9

Observation a9ec3eef-ea7b-4fd0-b9a4-b9e280f06b30 · outbound

This paper cites Training language models to follow instructions with human feedback.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Training language models to follow instructions with human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.377436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.377436Z digest=sha256:0c22d8459cedf8f4738cfa91e9e21d748fc6fa660d121929e35a26d4cb3b1f02

Observation 539e37b4-d317-4ea6-8ea7-4cc0bc67d028 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Learning transferable visual models from natural language supervi- sion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.382777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.382777Z digest=sha256:35959433fa6ba5c2c5508f02667904f4ba48c34a817f3648c3bfced9e4979bd5

Observation 46b8b64a-4161-4b9b-b325-50fc5e150257 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Direct preference optimization: Your language model is secretly a reward model

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.798944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.388405Z digest=sha256:d8be8b2ea19146c30bb1ba1052f3ff32568b3f5f5a31b47c8141768148a9825e

Observation d66c0994-55d1-4555-9588-607c3c1e22e7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.393913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.393913Z digest=sha256:829759cc5a18d43857909930a65c1e3245d4b120161d581fb5da25a64f12863d

Observation 0a179c69-c0b6-4d63-8011-d0a23245e2af · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.399783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.399783Z digest=sha256:b7cefbc3b4c3a53fedc8ef503955671c2c08a8af3ca446642ed49341808bc6c4

Observation f9d311ea-0459-4edd-9201-2f7766722622 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Gemini: A Family of Highly Capable Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.405958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.405958Z digest=sha256:fd4723a1a4c13be9674532d05225515ab86a4db87a6914824c561fe77c826bb7

Observation ec515631-2aea-4330-9deb-950b9b20bb3e · outbound

This paper cites Eyes wide shut? exploring the 10 visual shortcomings of multimodal llms.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Eyes wide shut? exploring the 10 visual shortcomings of multimodal llms

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.780892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.411814Z digest=sha256:c8f1c7d731417834787375a5438e6dfd1614b7e07cd7ad2615a64713d0ba0674

Observation a8c18b14-d571-498e-94b0-8c673eaac303 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.416910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.416910Z digest=sha256:9e4690fa1c3c50e26c82ddc8dd08e3aa89b28acb7ac3bd2fb4280c54b7a74951

Observation b781e7cd-e719-4fb8-bdf5-c8ba61f9eb03 · outbound

This paper cites Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.422147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.422147Z digest=sha256:df55732438f2489740f7c597f2b3607f65c3e6c3b9f69bf2f06219c683103ea3

Observation b5dd3bf9-8510-42d6-8dd5-0c174b40ebe7 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Emu3: Next-Token Prediction is All You Need

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.427816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.427816Z digest=sha256:1327ae547cc6cc5f2f5377b4484a7bdba597a367b2872fb4e511bf68f7461285

Observation 90bd0078-f692-4549-91bd-85fda6748089 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.763059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.433882Z digest=sha256:614aeb655c0dee7a530b574d6f06f329c70e4a3962c4121f8f89dcdaa76d445a

Observation 673ab646-95d2-4282-8bbc-bc16281423d8 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.439011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.439011Z digest=sha256:99f429db7f4a7e6bbbe270880852a2f2317c9cbdc21cc9873891ef47a2c1ea0f

Observation f319dd1b-8512-41cd-9fe4-7f8eedb838b4 · outbound

This paper cites Vigor: Improving visual ground- ing of large vision language models with fine-grained reward modeling.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Vigor: Improving visual ground- ing of large vision language models with fine-grained reward modeling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.444343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.444343Z digest=sha256:ea2e3029327ef79bd2d51fb57f9a0392921f03020f04eb4aee44f41a2f6ddb6d

Observation 6e3e43e3-f35b-4de7-a02d-cb4cad39ea32 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.449491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.449491Z digest=sha256:a38390ceb2761e0862dd91b4b51f4b042f4e6b9564992863346679e428e1c933

Observation 1039a7d5-e991-4703-94d3-df29571e1b96 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional hu- man feedback.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional hu- man feedback

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.745492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.455553Z digest=sha256:1801c03d5bc8e6cb7913ceb7e4f840d4c03abee803f5aa175819ceceb13e86a4

Observation d9c0d487-ee6f-4416-a2bc-58167ce97727 · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.460464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.460464Z digest=sha256:35215d397cf576652cffca08e9cfa84ed34a4e65ce424387fb685e271aa85c15

Observation a3255b99-0504-4be1-867f-244dfc44a424 · outbound

This paper cites Self-Rewarding Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Self-Rewarding Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.466231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.466231Z digest=sha256:67bb52aaff43e685fc1b61236348226dd06092ec710e2879e8e55bfb64fcf97f

Observation 4c91f07c-3a13-41d5-aa0a-d4d3fd452b2e · outbound

This paper cites PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.471772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.471772Z digest=sha256:355b72405293447deb27b7cef672174383ad9b4169e69c7cf873e070adbdd05e

Observation 538b0ddb-5dbf-4130-ab77-8f763dc35550 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.477475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.477475Z digest=sha256:44e930aac242de565533dbc0685a2c8da74dadd77838785edae4190f4a1340d4

Observation eda6088e-7fec-415d-ba2a-21aaecb2a30b · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation SVIT: Scaling up Visual Instruction Tuning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.483079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.483079Z digest=sha256:6d223eb22f72852bc7e4801a36047bbd3e2a0de3c92ce88a874fe9214ea069b1

Observation 75cd0b7e-36f6-420e-b4de-136dfbbaaccc · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.488218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.488218Z digest=sha256:3be532ce5226d3c9bec26bcee3a39514b66a7b64427e5aef1e20f653ff9f1cec

Observation 7b241dee-3bf3-494f-83bf-bdb857a076f1 · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Calibrated Self-Rewarding Vision Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.494340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.494340Z digest=sha256:2052809b2c590b0d66e7d5633241fed67ef9a0e96faaa171ae96f56af0d70ca7

Observation 3ebe7466-a4a3-4a02-92d7-4503ba250d5b · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.500631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.500631Z digest=sha256:c03dc4c9b68d01f3f0597df3146f0909f7ce03e4b54d03efcf6a17388705ee3f

Observation 3ade1ccc-ebd7-4f56-9f90-2287aa58bc8e · outbound

This paper cites GPT-4o and our Critic model produce similar scores for responses, but they fail to identify the flaws in bad responses from the baseline LLA V A model.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation GPT-4o and our Critic model produce similar scores for responses, but they fail to identify the flaws in bad responses from the baseline LLA V A model

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.728234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.506876Z digest=sha256:11e998522c7d0c782f09b25d976f0a1b2da1fe871e9320c7690bdaca9e2dce8f

Observation 37566070-fac6-4220-bf83-d2c6b060ebaf · outbound

This paper cites As shown in Table 8, most of the experiment is con- ducted with prompts in rating style, apart from the ablation study presented in Section 5.3.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation As shown in Table 8, most of the experiment is con- ducted with prompts in rating style, apart from the ablation study presented in Section 5.3

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.710936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.513267Z digest=sha256:79285a80495212b86ab39dbeec1afb63415d84cb88cc8a163820cf53f5396047

Observation 9973f9c9-347f-40a6-8ac0-b9fa44c59baf · outbound

This paper cites The training details are shown in Table 2.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation The training details are shown in Table 2

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.691045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.519074Z digest=sha256:40329409a197fa3f5bc81c1c3e245ca316769948863ebe18d16ebc37e78db8f5

Observation 95650ce3-5617-4935-b03a-8d964506304d · outbound

This paper cites Using annotated preference data, one round of preference learning is conducted on LLaV A1.5.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Using annotated preference data, one round of preference learning is conducted on LLaV A1.5

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.670611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.524962Z digest=sha256:f4c37f4102e2ab1671b3ca94da5fa89fa88b6b4fb2f390b98ce0e2d9f73e0c28

Observation e4972ee4-a8ae-4bcb-b70a-11b9cb3e5809 · outbound

This paper cites an unresolved cited work.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:15:46.652699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.531083Z digest=sha256:34b7f8d5619cfc46890dc13493d1081b177e792061de4ce60bfaf97d77e2c6ae

Observation 9661aef2-010b-4c59-8e7b-598599019fca · outbound

This paper cites an unresolved cited work.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:15:46.636166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.537284Z digest=sha256:8b3a406a790b78a655dd4ffc39e78159b5e7707914922ec21f8f961a444cb729

Observation e55212cd-03c3-4adb-81a8-741fcfaca837 · outbound

This paper cites Here, we will show some examples between EACO and baseline LLaV A-v1.6- Mistral-7B in Table 9 and 10.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation Here, we will show some examples between EACO and baseline LLaV A-v1.6- Mistral-7B in Table 9 and 10

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.619883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.542436Z digest=sha256:f1ffd49bfad79944c18057264741d5487b14449c62a00d01010accd66a516f15

Observation d6e4cda5-e15e-4dcd-8cd7-63f1716d6fc7 · outbound

This paper cites score:⟨total points⟩.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation score:⟨total points⟩

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:15:46.601541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T21:15:45.548297Z digest=sha256:ecd1e422500d6da18d953fac4e86778da4441187aec07723339be7d2deb41533

Pith citing papers

Observation 822144ed-1b63-4e98-b397-83b8525ce130 · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:08:02.076951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:74f8b0f40325937ddc55a8f0eb0e2e1bb25f5f7b436e044186f3c275ca272bfb