Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:27:19.422402Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2506.17901.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:27:19.422402Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T04:55:07.037081Z
A source-named dated measurement, never combined with another source.
Source: cited_works
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 273b3927-c791-4d73-ae2b-d1bc481b4586 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa400c8d-d4ca-4635-95fa-5316f5e2134a · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32175eb1-23fc-440a-8e50-8f07e4a0e8bf · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b16d96-276a-42a5-9803-672a34bcd566 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs GPT-4 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d19aa91a-eb45-4ff2-ad7c-43d7846ad7cd · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Improved baselines with visual instruction tuning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca2a5fca-638c-4e0f-9852-7cb5c097dad5 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f82fde58-886e-4f91-9968-e168b67e6bee · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Gemini: A Family of Highly Capable Multimodal Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93ceceb7-6b52-404b-913f-734d14c999be · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs A Survey on Hallucination in Large Vision-Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02ba2f93-d8a2-4594-a6b1-07d36ce67313 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Hallucination of Multimodal Large Language Models: A Survey
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40e6b8fc-6543-44f4-93fa-f8dc726fa650 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe7b840-b758-452b-ad9d-111b7f59d2b0 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs The deluge of spurious correlations in big data
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd9ca390-6541-4b3a-9a8a-5272bd001bbe · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Spurious correlations in machine learning: A survey
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48703dd0-3329-485b-9045-e787be222265 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mitigating object hallucinations in large vision-language models through visual contrastive decoding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52ba5597-4e00-4d45-bd0d-02cfdc97634e · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Described object detection: Liberating object detection with flexible expressions
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b42d05f7-f51c-4955-a2e2-1cbf62a10090 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mattnet: Modular attention network for referring expression comprehension
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f7685f1-b46c-4210-9ca0-edab7be1a164 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs A fast and accurate one-stage approach to visual grounding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a588206f-1b12-4e05-bdeb-c69053877f28 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mdetr: Modulated detection for end-to-end multi-modal under- standing
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 45b28f3a-55ae-4190-ad07-926b025c1485 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3c8b111-2b48-48c6-ba44-67752411513b · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Grounded language-image pre-training
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18aadb9d-8643-4a34-9cc2-8673b052e883 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0827b9be-23cc-4579-8685-d2ad90942133 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Phrasecut: Language- based image segmentation in the wild
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a711b74c-fd5d-4930-b514-ea8bfed4ece3 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Advancing referring expression segmentation beyond single image
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a943a3ac-de6c-4415-8207-6933e1295116 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Language as queries for referring video object segmentation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 67ce0251-1079-440a-83eb-03fc44b7c557 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89fc4323-c67c-43ff-a2a0-cf1dfbfb5186 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9eb3e5b-e6f3-4f26-84dc-86866fc8b19a · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Llava-grounding: Grounded visual chat with large multimodal models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49669868-d3c4-4eed-8c4e-67a765c3624d · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c608b4a2-7148-4b59-8f83-29cd6fb8de7a · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Lisa: Reasoning segmentation via large language model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2df227a7-fe8c-42da-9e7c-5d795981635f · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Gsva: Generalized segmentation via multimodal large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d5202c9-d0da-4c77-b193-e2d61653cb87 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Glamm: Pixel grounding large multimodal model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eac4a933-149b-4c56-8259-00a0d42a7e3a · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Pixellm: Pixel reasoning with large multimodal model
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c6ec7099-6cf0-421d-883e-b59151bb6fe6 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Ground- hog: Grounding large language models to holistic segmentation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0038e964-8c30-4998-bc21-3b7d57f2bec4 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Woodpecker: Hallucination correction for multimodal large language models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ffdafa19-6a83-46f6-8203-d7442500db88 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d8a7ba5-f0f8-42b4-9f20-5a7d4346ee61 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e69ed7f-7f4b-442c-afca-f7e6eafabd88 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a788739c-d432-430e-ab17-4578f08fca09 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c07d63-eb7f-4cfe-b766-28ba79cd7a93 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 941699a5-a68d-4549-b759-f20cd207bae8 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0289201-04d8-4b11-abd7-52c0fd50dc5c · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83809ade-bc12-4f4c-a9f3-9d932fe0c3fa · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9266776b-4b02-4827-9325-d3ce05895f3b · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f50dc7-c9ee-46c5-a2bb-9acd531102e8 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mitigating fine-grained hallucination by fine-tuning large vision-language models with caption rewrites
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a4157b9-c626-4e58-ae53-43ce1a26264a · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a79bad53-b962-41ef-921b-072ee2ab4e2b · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c32024ec-ed48-4892-8f1a-7e24be6b056e · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Silkie: Preference Distillation for Large Visual Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e1ece5e-0ad4-4f77-827f-bbcf6daad7bb · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Detecting and preventing hallucinations in large vision language models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73579177-224d-48bb-b450-acc716d508dc · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Aligning Large Multimodal Models with Factually Augmented RLHF
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df92342c-82be-4b17-ab21-c721e7ac7c18 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Multimodal Chain-of-Thought Reasoning in Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 270c8a6e-257e-4f8a-88e7-8247fbfbe69f · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Lora: Low-rank adaptation of large language models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a16614ee-0a09-4571-a35c-29ce190a0800 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Segment anything
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6087d896-7570-4182-b910-32cb740d880a · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Decoupled Weight Decay Regularization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 721e8b60-a827-4b75-bf89-3337f723d851 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Haloquest: A visual hallucination dataset for advancing multimodal reasoning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8cc0add7-bea0-48fe-baad-d63090dd7ee1 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Evaluating Object Hallucination in Large Vision-Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c82111ec-6b2e-4b60-9ff0-465307f7a97b · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9220180b-39ae-4622-aa9e-460ed9cef1f2 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ca48f0e-6a07-4bc5-ad2c-1bd431309bfb · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs A survey of multimodel large language models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation be649093-e7c6-4ac7-bd39-c3a49928230a · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Referitgame: Referring to objects in photographs of natural scenes
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 517d6065-62a0-40aa-91b7-816888cdbd26 · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Empowering Segmentation Ability to Multi-modal Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d59454f-1ad6-49e5-9138-13c7cecfad0e · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0cc1135-b213-4530-bd64-d317a063904d · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6a2fb75-2e9a-4ac6-b8cc-131aba5b33ac · outbound
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aabe33ca-8363-46f4-84c2-e6edcf9dd92d · inbound
Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.