Pith. sign in

Paper Citation Record · LEDGER

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization

As of 19 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 2 inbound Pith citation observations for arXiv:2504.15619.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15619 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:27:36.967307Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T09:32:15.287630Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T06:54:20.297814Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 093899e6-7095-4e10-b223-4ddb7bcf3d23 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.775609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.775609Z digest=sha256:25ca8bd89bb00020516803ddb5769bc1b69a1405269b427a0c57d9e7350afa9c

Observation 92150601-f9d1-4c86-bdd9-6a97ca3e59e7 · outbound

This paper cites Rank analysis of incomplete block designs: I.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Rank analysis of incomplete block designs: I

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.781271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.781271Z digest=sha256:a445c22984eadf4f3dccc9c376e5256d77effe1b4f86afde1f1b8d33b18ef1f7

Observation 719192e5-89df-4588-bac2-e43d57f8e384 · outbound

This paper cites End-to- end object detection with transformers.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization End-to- end object detection with transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.786299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.786299Z digest=sha256:a4a16aba228aed6cbb3727b68dad0281ddfa6c0c87a8d504a1de559696b44282

Observation d81c4bbe-ef1f-4e95-a20b-9f4be29fe3b7 · outbound

This paper cites On Softmax Direct Preference Optimization for Recommendation.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization On Softmax Direct Preference Optimization for Recommendation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.791252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.791252Z digest=sha256:a232a3c757c384c92010778b7e44792be5d9d1d2c673b1018e3d68d4fe207162

Observation bb9c2a54-d787-4175-94d9-5ca5817a5c73 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.795638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.795638Z digest=sha256:ab0d310b88d38c86687037acd7db957e6bb6a6fa7bb2d411285f0b8c72cc0839

Observation d04cb186-8fa6-46a8-80e5-ebef887297ef · outbound

This paper cites Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.799901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.799901Z digest=sha256:4e43a71306c22b3f892ff466b0dd4caf5228e6efe34fd6a76f408b082aeb716b

Observation 3a17869f-2166-42cf-8746-51fd6f8a1e80 · outbound

This paper cites The Llama 3 Herd of Models.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.804737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.804737Z digest=sha256:49dc72013696b55be25b1ca96fbdae53bb13b2012a8365b522c797ccc5d44476

Observation 95945a37-ba21-4a3d-baa5-fdab24ce670e · outbound

This paper cites Token pref- erence optimization with self-calibrated visual-anchored rewards for hallucination mitigation.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Token pref- erence optimization with self-calibrated visual-anchored rewards for hallucination mitigation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.809824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.809824Z digest=sha256:35c1b36c7a077ac186e0ac83d902e814bbd6e6ded9e79fe40d22f9133a21705a

Observation a6d22f44-da87-4636-b64c-3c08af0b715a · outbound

This paper cites Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.650484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.814777Z digest=sha256:f9118e580fc4908b5fe4353fbceade973be23879559691a4f57bf6e0c358b86b

Observation d8808e51-ce19-4561-ba05-82f0bd35a36f · outbound

This paper cites Vcoder: Ver- satile vision encoders for multimodal large language models.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Vcoder: Ver- satile vision encoders for multimodal large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.636582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.819185Z digest=sha256:cb5ce73206b5873ef230cd5c63f721cf62c63b096e4dabe548d4475429b9b0fd

Observation eed9db34-3573-4c79-9481-0aebebfa17fd · outbound

This paper cites Segment any- thing.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Segment any- thing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.823599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.823599Z digest=sha256:2ef52440a5393c95d35b56510e4c1b7a0db26962abecd1e43b568962ec15fcad

Observation 7a39659f-7ac2-4425-9003-a4517094f08f · outbound

This paper cites Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.613634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.828357Z digest=sha256:de37989467ac09d86d46997bcbd4e1d88b8ecd5b0e95895e8443920665ba3146

Observation 5a629912-dca5-4e7f-9a44-ec6495c14280 · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.600403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.832906Z digest=sha256:407ee4c2bac465be0c04ef27dd0d1e2926c93b9a11d886cf37088427b1983247

Observation 4934d308-8fd1-44a2-b4de-e18754222cd3 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.586286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.837390Z digest=sha256:78590d318dfde52743e3cee709ff8e5da0af832d78835c864aa8419f5aac74cd

Observation 82e1dca7-42f8-4e1b-a925-a5a10ba7aaf6 · outbound

This paper cites Visual instruction tuning.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Visual instruction tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.842748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.842748Z digest=sha256:efcc5a8f766cf10ebd5088964e70837f44a6dc6192ab11c2d5188d1091035dac

Observation 424ef249-6884-4c4d-884f-fb2f98b8c06c · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization A Survey on Hallucination in Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.846971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.846971Z digest=sha256:053079659c78c905971a2ab6e01c1fbc00543cd368fa545ef20ae7e787dfbf63

Observation 130569f9-1b16-4e9e-8e42-4c1bcff895e6 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.851574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.851574Z digest=sha256:c426eeaa581dee2e63d12145c1fa82d3909c3ff66edceb3c5e41740ffbd2fb93

Observation 5f4a199b-33f3-482a-8863-026fdecfdc87 · outbound

This paper cites DAMA: Data- and Model-aware Alignment of Multi-modal LLMs.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization DAMA: Data- and Model-aware Alignment of Multi-modal LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.855648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.855648Z digest=sha256:f9d692ba5f132e90c796e4858adc902029954a6e60679d36542f4a33b0684a78

Observation 125a8a78-1c1b-4aac-af6b-4f56387f5b13 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.859960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.859960Z digest=sha256:b50e389501bc4fef68577ababf4513c8cbddda2f66e854d680e4210fb745fc1f

Observation 23502d6e-736e-4c68-87de-ac85f7769eb8 · outbound

This paper cites Training language models to follow instructions with human feedback.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Training language models to follow instructions with human feedback

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.552804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.865016Z digest=sha256:cc8c88ed57fac6e191e69307665bc52864cabfd980cb8e4370833592376349f0

Observation adad47df-3b0a-4954-9301-9987268c5d02 · outbound

This paper cites The analysis of permutations.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization The analysis of permutations

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.538273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.868685Z digest=sha256:39a18154c0cc422c1fdfad94a62a53c36b412105988962381796b8beea7cd0d8

Observation dd965f93-ac32-4fd8-8348-4841393b0271 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Learning transferable visual models from natural language supervi- sion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.872489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.872489Z digest=sha256:5d9b98931f6c75a858ea50fceac44c0eec8d478aa31a4d93a85bab679eb32f92

Observation e26f40fc-0a66-49a1-8279-52db45feb442 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Direct preference optimization: Your language model is secretly a reward model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.515321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.876456Z digest=sha256:1bea913a6715dc2f7cf6e81d6bc98db36e025a25a3f476e6b1f927a16ec4e69c

Observation e19d75d0-46dc-46b5-b2e6-f4718afbb906 · outbound

This paper cites Object Hallucination in Image Captioning.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Object Hallucination in Image Captioning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.881296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.881296Z digest=sha256:7243495a622bf448763727da64ccb3a52f7f10a2f54ae23b26d185ff279a1a2b

Observation 3f03910b-e67d-4962-94b5-a4567d505901 · outbound

This paper cites Proximal Policy Optimization Algorithms.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Proximal Policy Optimization Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.885833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.885833Z digest=sha256:f222e9d84b61182065831521935d8eccb082efe9f8c8101caa647fb1416048c4

Observation 37f7f5e2-deb2-4bf3-a4a3-a7ad16d48696 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Objects365: A large-scale, high-quality dataset for object detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.501824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.890601Z digest=sha256:9a0d75acc1d1dbce6d3120abdc05a21cc1d1f35b6a3a3e298299d7ee51efa9a8

Observation 1214fcfc-d56e-4ed2-9e45-9a34c45751bf · outbound

This paper cites Aligning large multi- modal models with factually augmented rlhf.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Aligning large multi- modal models with factually augmented rlhf

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.487560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.895020Z digest=sha256:6d272eaf8a2e5a5185bb439a09b8b453e6904771e35243b89c567f8a80300b9d

Observation ee3e42fc-198b-4fbc-baa8-8b677aebf1ee · outbound

This paper cites Resolution-robust large mask inpainting with fourier convolutions.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Resolution-robust large mask inpainting with fourier convolutions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.473590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.899747Z digest=sha256:8c0d7f90057fc5d38b9d38c9e3aa48bf18e36dc704062e21c199fe146882664e

Observation c1c23691-359e-4934-86a2-7da63f6ed46c · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.459135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.904069Z digest=sha256:a900b4d95cbed81fc1e0d831d3a657dec120c7ecec4da7ffe627b95dfc7b1bd1

Observation 2d99a1ee-41d2-498b-9578-1562c7ca1bd3 · outbound

This paper cites mDPO: Conditional Preference Optimization for Multimodal Large Language Models.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization mDPO: Conditional Preference Optimization for Multimodal Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.908291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.908291Z digest=sha256:66a68eaaeebee2bf32cc21d8d526914426080431100ede8994d27e8fece591dc

Observation 7de03321-57d7-460a-8924-106a1a52e5c4 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.912762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.912762Z digest=sha256:09391be50c9d87e87fc45afcc41fda395a606711ab2a40c25093fa73fd5f1d46

Observation 5f2e13fd-6b91-40ca-b3e9-fb4316cd8b43 · outbound

This paper cites V-dpo: Mitigating hallucination in large vision language models via vision-guided direct preference optimization.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization V-dpo: Mitigating hallucination in large vision language models via vision-guided direct preference optimization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.445952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.917543Z digest=sha256:0bb73b3d6ff5b2d3287f088ee474f85f18a06c4ac6306f4d3e91a15b0a16dc39

Observation b1706a88-95c2-4dc8-91d8-c7a31db28773 · outbound

This paper cites Miti- gating object hallucination via concentric causal attention.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Miti- gating object hallucination via concentric causal attention

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.433369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.921589Z digest=sha256:b60dcd0b50e42a7fc087456f91bd62b9f3db0c2bacf8cfd6f61dc61199fd48dd

Observation 543e52e2-b383-42e1-b31c-521ecae710f5 · outbound

This paper cites Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.419164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.925145Z digest=sha256:983fab660fe36b8f8cabb661e54a4d17b6e4b8dbfa0971b608e4525d99105d86

Observation 1b7b53f2-0ecd-44f9-8737-94a9d76baf2e · outbound

This paper cites Gradient surgery for multi-task learning.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Gradient surgery for multi-task learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.403148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.928913Z digest=sha256:bd51c6df00daa0ffd7eaa405f4720ec3a4bf3bdd644359be6ee75ca4288f394d

Observation 6ac5f39a-6808-4e98-a437-78bec96e61a2 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional hu- man feedback.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional hu- man feedback

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.389025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.932671Z digest=sha256:6cd109c1d4f51059076cb03469a179f415ab08ed98c20413398189b3387ae5ed

Observation a5b3b0f2-d9a5-4cd6-b1e4-0139ba3583ec · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.936265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.936265Z digest=sha256:521c9b7a19d8260b6104310824ce786882b6108c1675cffbd4c3c3ebd30698d0

Observation 73ad28ea-13da-4df4-b350-6544a29966c6 · outbound

This paper cites Less is more: Mitigat- ing multimodal hallucination from an eos decision perspec- tive.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Less is more: Mitigat- ing multimodal hallucination from an eos decision perspec- tive

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.373451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.940132Z digest=sha256:1d1dc8d27b8b9ecc4471e60fea5d1a9b9a9df868c7c9e039c265abec961fc395

Observation 45faf425-0cfa-479f-8119-e2b3bd570cd4 · outbound

This paper cites Dino: Detr with improved denoising anchor boxes for end-to-end object de- tection.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Dino: Detr with improved denoising anchor boxes for end-to-end object de- tection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.359000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.943904Z digest=sha256:acceea98a781a6b4803f954af0b55e622d07219429a0f0f81d18a9378708c00b

Observation 01843850-9619-4c58-b617-ebe5e29f0e00 · outbound

This paper cites Automated multi-level preference for mllms.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Automated multi-level preference for mllms

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.344886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:27:36.948595Z digest=sha256:0f33fd2450f8075d38765f5bab5f6c4bd6f7a977d1cd98b3278f48a951ea1192

Observation 44b884ca-911e-4888-bdc4-d7fd8c394d63 · outbound

This paper cites Recognize Anything: A Strong Image Tagging Model.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Recognize Anything: A Strong Image Tagging Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.953580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.953580Z digest=sha256:bce97806793aa8d3ae9db80b0535a189cf7f60b588b00b1348b17623a1a6e485

Observation 692784a0-41d9-449d-ae1a-f18255bc8761 · outbound

This paper cites Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.958272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.958272Z digest=sha256:926c33249c64f47672b0be4ff5015dc405c72c3b6ac1e998e203511d0bf0faed

Observation e37eccb4-077b-490d-8871-adb8cb90c6ee · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.962572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.962572Z digest=sha256:a475add35ed090863de8b4b2b972121ed5c8892eccfbe0f6b4b19f445d76047d

Observation 332e1d13-7462-41bc-a4f0-d5d64fe41616 · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.967307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.967307Z digest=sha256:9c45edfd68ea0c64d7c5f917d00bd6ce7980c5054da0a73d31da2b3643120db6

Pith citing papers

Observation 96fb79cb-15f2-49a7-a315-a5403f276793 · inbound

Experience Augmented Policy Optimization for LLM Reasoning cites this paper.

Experience Augmented Policy Optimization for LLM Reasoning AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:20.299776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T06:52:55.368187Z digest=sha256:baa25c8308ca169ef324f71a373ef61a6265152d4f361072f38c855e463d3b29

Observation d05ecd30-59cf-4ae0-81dc-2c38046ca202 · inbound

Experience Augmented Policy Optimization for LLM Reasoning cites this paper.

Experience Augmented Policy Optimization for LLM Reasoning AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:15.287630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:15.287630Z digest=sha256:dd0e1c7f7ef4cef829bff5f88b650defebcb0340260eb66f9121f85a9a09abd0