Pith. sign in

Paper Citation Record · LEDGER

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs

As of 7 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2506.17901.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17901 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:27:19.422402Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T04:55:07.037081Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 273b3927-c791-4d73-ae2b-d1bc481b4586 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.137072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.137072Z digest=sha256:a9bfcf6f39c5b6dc72671bd17bc293e46ae3a1d88095f5b6ebc1da7ebd195ec0

Observation aa400c8d-d4ca-4635-95fa-5316f5e2134a · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.142549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.142549Z digest=sha256:6bf826b36acee22458c12474d1db5fc0f7ff76b0dacf89768ca722c352f19fad

Observation 32175eb1-23fc-440a-8e50-8f07e4a0e8bf · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.147296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.147296Z digest=sha256:4a65a2d87b479d31078be1378a4bb8162644b4c07ae368aaf64275a1cde26bb4

Observation d4b16d96-276a-42a5-9803-672a34bcd566 · outbound

This paper cites GPT-4 Technical Report.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.153898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.153898Z digest=sha256:8978c9c65a609b186f7bc89ab2ca2c39d4ba1acc5f081297892cd153b4194604

Observation d19aa91a-eb45-4ff2-ad7c-43d7846ad7cd · outbound

This paper cites Improved baselines with visual instruction tuning.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Improved baselines with visual instruction tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.159315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.159315Z digest=sha256:bf9241049f25e61190e3a4f427fd412dff3a81930bf383fe36db7fe2f97778fc

Observation ca2a5fca-638c-4e0f-9852-7cb5c097dad5 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.164043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.164043Z digest=sha256:f6050dd6d9e4b4824308eac76f107ac87e9092739ef6cfaae5b66a6b9ec7d7d0

Observation f82fde58-886e-4f91-9968-e168b67e6bee · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.168676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.168676Z digest=sha256:b7dd876e5440722476a7b495f1d470c9205038d18c06f7be3b07905ca0c0d09c

Observation 93ceceb7-6b52-404b-913f-734d14c999be · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs A Survey on Hallucination in Large Vision-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.173386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.173386Z digest=sha256:9fb518698601653644ca98fc35ca62ad7e37c81da1d50e3104951bbadea7bc96

Observation 02ba2f93-d8a2-4594-a6b1-07d36ce67313 · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Hallucination of Multimodal Large Language Models: A Survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.178110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.178110Z digest=sha256:386465f9321dce21fc8f1515604e6939f6bce78bd362d88b2da312939ec00c42

Observation 40e6b8fc-6543-44f4-93fa-f8dc726fa650 · outbound

This paper cites Analyzing and Mitigating Object Hallucination in Large Vision-Language Models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Analyzing and Mitigating Object Hallucination in Large Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.182352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.182352Z digest=sha256:e523b63fb510a3f71c2f4da7ec530f0404576b682f8fe944ee57fd28e3388637

Observation bbe7b840-b758-452b-ad9d-111b7f59d2b0 · outbound

This paper cites The deluge of spurious correlations in big data.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs The deluge of spurious correlations in big data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.186869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.186869Z digest=sha256:901b194969be8c18aa61f7a2d81b6906cb9f4b112e7ededfb8d819d1433b73ab

Observation dd9ca390-6541-4b3a-9a8a-5272bd001bbe · outbound

This paper cites Spurious correlations in machine learning: A survey.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Spurious correlations in machine learning: A survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.191298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.191298Z digest=sha256:cadb79be2414316f64a07f10469c353adbfcfaadc2ea444a9f9fe5eb566cd330

Observation 48703dd0-3329-485b-9045-e787be222265 · outbound

This paper cites Mitigating object hallucinations in large vision-language models through visual contrastive decoding.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mitigating object hallucinations in large vision-language models through visual contrastive decoding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.248345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.196105Z digest=sha256:e993c532effe8bfa92a35ccc8648c0a54e841f676c83d3c301879472deebbe2d

Observation 52ba5597-4e00-4d45-bd0d-02cfdc97634e · outbound

This paper cites Described object detection: Liberating object detection with flexible expressions.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Described object detection: Liberating object detection with flexible expressions

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.233649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.200189Z digest=sha256:aef519ef8c6869a26259ece513d53298c83b67101e8af3da5230fcddbea8cc4b

Observation b42d05f7-f51c-4955-a2e2-1cbf62a10090 · outbound

This paper cites Mattnet: Modular attention network for referring expression comprehension.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mattnet: Modular attention network for referring expression comprehension

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.220220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.205034Z digest=sha256:8628a2555331db442552ea5d8314c8f82ce28386758a7d5e6cccbf5ecd3bebbc

Observation 9f7685f1-b46c-4210-9ca0-edab7be1a164 · outbound

This paper cites A fast and accurate one-stage approach to visual grounding.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs A fast and accurate one-stage approach to visual grounding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.207560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.209588Z digest=sha256:1dd454865d22f4c0b7b80feda73eb39b85ee8e21b66ce5b2196ef7f1105f90b7

Observation a588206f-1b12-4e05-bdeb-c69053877f28 · outbound

This paper cites Mdetr: Modulated detection for end-to-end multi-modal under- standing.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mdetr: Modulated detection for end-to-end multi-modal under- standing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.194646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.213891Z digest=sha256:b0a4bd8e5e8bb9b7ed9e8c860bd5fd9dddf12ee18084d72459e694b34a8c6fbf

Observation 45b28f3a-55ae-4190-ad07-926b025c1485 · outbound

This paper cites Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.181167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.217883Z digest=sha256:9b1c5f8a4e60ee896d9cce503f6e9e813d8260c0128c5283b3e65be3becc96cb

Observation d3c8b111-2b48-48c6-ba44-67752411513b · outbound

This paper cites Grounded language-image pre-training.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Grounded language-image pre-training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.222225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.222225Z digest=sha256:5380088e937e475b25807a0a45dc914beec57a2b3b970b46fdd805cfe62feeaa

Observation 18aadb9d-8643-4a34-9cc2-8673b052e883 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.226753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.226753Z digest=sha256:35fb792361ece05c531460f1efe425eef6885db7c7001590b004669c5227b6ec

Observation 0827b9be-23cc-4579-8685-d2ad90942133 · outbound

This paper cites Phrasecut: Language- based image segmentation in the wild.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Phrasecut: Language- based image segmentation in the wild

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.151548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.230979Z digest=sha256:a52ea97950cd63268414eac47e0a38b473b64f612ef6e89a14e2eda902cccea1

Observation a711b74c-fd5d-4930-b514-ea8bfed4ece3 · outbound

This paper cites Advancing referring expression segmentation beyond single image.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Advancing referring expression segmentation beyond single image

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.138869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.235705Z digest=sha256:4a4d5ce70a706cc626a459c253e46dd14833c3e4645ea74d74315145b81ab8c9

Observation a943a3ac-de6c-4415-8207-6933e1295116 · outbound

This paper cites Language as queries for referring video object segmentation.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Language as queries for referring video object segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.125731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.239704Z digest=sha256:050610d1cc8022ed0c0a563ea4adce0bdeb74ee91e0cf912a912e37a36e7a759

Observation 67ce0251-1079-440a-83eb-03fc44b7c557 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.244352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.244352Z digest=sha256:4e00db424670e3d55b2d96d31032bca9725a91bd4d913de7748530963bc16def

Observation 89fc4323-c67c-43ff-a2a0-cf1dfbfb5186 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.249169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.249169Z digest=sha256:8c59377bd20203a80610fdde4945e789219eaa8256c0790942b5951af2c1b35a

Observation e9eb3e5b-e6f3-4f26-84dc-86866fc8b19a · outbound

This paper cites Llava-grounding: Grounded visual chat with large multimodal models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Llava-grounding: Grounded visual chat with large multimodal models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.111553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.254479Z digest=sha256:ad7c69bd6703b8a4ee4384751563985bd3948406d447d305b8a99be0bf931986

Observation 49669868-d3c4-4eed-8c4e-67a765c3624d · outbound

This paper cites BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.259475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.259475Z digest=sha256:f769e51c3d7f73dbec1c0f75ecc4cff32c6856e5bd755e68a6c10c033376f97e

Observation c608b4a2-7148-4b59-8f83-29cd6fb8de7a · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Lisa: Reasoning segmentation via large language model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.264438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.264438Z digest=sha256:5e418ea60deed2cef2c0e3a6a2e19b0be533a82a934c527122979cfe736d76f1

Observation 2df227a7-fe8c-42da-9e7c-5d795981635f · outbound

This paper cites Gsva: Generalized segmentation via multimodal large language models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Gsva: Generalized segmentation via multimodal large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.085128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.268906Z digest=sha256:8bc444553bb1cb16af1df59d0ac3bbe81e437d5ec432a07903e5e5df32a9a115

Observation 3d5202c9-d0da-4c77-b193-e2d61653cb87 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Glamm: Pixel grounding large multimodal model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.274135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.274135Z digest=sha256:73fd7771fc812f20cd62004b64d7fc7c4fb2681c13df3d4fd608a0962f0bc398

Observation eac4a933-149b-4c56-8259-00a0d42a7e3a · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Pixellm: Pixel reasoning with large multimodal model

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.063536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.278494Z digest=sha256:8b0204303a161883ce59e328e328a5f4ea9d8246490697e8e2e9d6fbfc161dc3

Observation c6ec7099-6cf0-421d-883e-b59151bb6fe6 · outbound

This paper cites Ground- hog: Grounding large language models to holistic segmentation.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Ground- hog: Grounding large language models to holistic segmentation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.049650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.282891Z digest=sha256:5a8f5c927ba6cfe96245aede4fac82fa6e4900674be2fbb7110cbc34822a7f65

Observation 0038e964-8c30-4998-bc21-3b7d57f2bec4 · outbound

This paper cites Woodpecker: Hallucination correction for multimodal large language models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Woodpecker: Hallucination correction for multimodal large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.035923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.287227Z digest=sha256:e7b0cee56526f417d075ee8842f3aa18fe8d159a7d5230cc67efdf8c661eaf36

Observation ffdafa19-6a83-46f6-8203-d7442500db88 · outbound

This paper cites Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.291980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.291980Z digest=sha256:f8917ead72ac10695c82250ff27b657b377b5804d1cb0c6ced5ee3fbed0d42f8

Observation 2d8a7ba5-f0f8-42b4-9f20-5a7d4346ee61 · outbound

This paper cites Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:20.020444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.297724Z digest=sha256:1000fecae111901dea072de68828d32a4f4bb10c7b43c0e2da53b746691ed402

Observation 8e69ed7f-7f4b-442c-afca-f7e6eafabd88 · outbound

This paper cites HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.302520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.302520Z digest=sha256:d0fec219c28fa9784eb2e77a99342dc097b52db31e535ab7ab7bb4518f1cb06d

Observation a788739c-d432-430e-ab17-4578f08fca09 · outbound

This paper cites IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.308057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.308057Z digest=sha256:5a41c91da6026212d70c993f137f4cca97608eb66e9d6abcb365c7f4e35f3b72

Observation 26c07d63-eb7f-4cfe-b766-28ba79cd7a93 · outbound

This paper cites Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.313207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.313207Z digest=sha256:ea3215caa33f8ebe960847d587f79d2b9dce972b35b02ba4904c4593e9d5b7b1

Observation 941699a5-a68d-4549-b759-f20cd207bae8 · outbound

This paper cites Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.317381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.317381Z digest=sha256:23522ce91bb3152e3edd6bd6eb7199c7b24fb82836cf7fc23354f2b2a6a65cbe

Observation a0289201-04d8-4b11-abd7-52c0fd50dc5c · outbound

This paper cites MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.322266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.322266Z digest=sha256:bff0e019ee2f691f01aa8a5be97c5e771b877606005c58681f2020ad2ed223ea

Observation 83809ade-bc12-4f4c-a9f3-9d932fe0c3fa · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.327237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.327237Z digest=sha256:52f44f10212d475e49f99562ee1f5a99d65238166250fc02c9a16f8fded6dbe4

Observation 9266776b-4b02-4827-9325-d3ce05895f3b · outbound

This paper cites Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.331587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.331587Z digest=sha256:8a8728bc9a18d3a034e61f01badb27eea7166fd97c107ab76fd3016af74a7efd

Observation 44f50dc7-c9ee-46c5-a2bb-9acd531102e8 · outbound

This paper cites Mitigating fine-grained hallucination by fine-tuning large vision-language models with caption rewrites.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mitigating fine-grained hallucination by fine-tuning large vision-language models with caption rewrites

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:19.998872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.335920Z digest=sha256:fb6ee38aeb04bcb536a9e6a35d2ec77056c8f4b5eb1f7d1e504e5cb230aa57c9

Observation 5a4157b9-c626-4e58-ae53-43ce1a26264a · outbound

This paper cites Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.340811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.340811Z digest=sha256:c24715c959f6d2876b82bb5f6e1ed34870efef15dfdfc934910d3e957ad2a681

Observation a79bad53-b962-41ef-921b-072ee2ab4e2b · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.345738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.345738Z digest=sha256:e86846a5503b0506cd0d3813b0e362cd8942028b7d8225cb43a96ea51d3623ed

Observation c32024ec-ed48-4892-8f1a-7e24be6b056e · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Silkie: Preference Distillation for Large Visual Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.350818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.350818Z digest=sha256:7b2549b196446e8e5de2a7cea9f4bfa0637b94717609b8eb07e2a4a81a53ae80

Observation 2e1ece5e-0ad4-4f77-827f-bbcf6daad7bb · outbound

This paper cites Detecting and preventing hallucinations in large vision language models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Detecting and preventing hallucinations in large vision language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:19.985299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.355853Z digest=sha256:3cc25689aac1727951ef266df70c39bd3dfe46e5eda4cb8e8b0dc7a790485dce

Observation 73579177-224d-48bb-b450-acc716d508dc · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.360031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.360031Z digest=sha256:799bde960e2c8c0efac3767407aef90633e621f4ee3232f6b61662f62e2d1594

Observation df92342c-82be-4b17-ab21-c721e7ac7c18 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Multimodal Chain-of-Thought Reasoning in Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.364377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.364377Z digest=sha256:c1f2d2b5504a8f9381b7a5fac18d29d782cccd62479713364ebf25e43198fb77

Observation 270c8a6e-257e-4f8a-88e7-8247fbfbe69f · outbound

This paper cites Lora: Low-rank adaptation of large language models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Lora: Low-rank adaptation of large language models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.368667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.368667Z digest=sha256:6a58689ef58e34ba0aef52aa639e6023db2d28237a5d15562f4f94e2a316ed99

Observation a16614ee-0a09-4571-a35c-29ce190a0800 · outbound

This paper cites Segment anything.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Segment anything

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.373049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.373049Z digest=sha256:0d328cb17f391d10fba3a16d65a6e19835e6344ac5b88f15b66df85fe97d2250

Observation 6087d896-7570-4182-b910-32cb740d880a · outbound

This paper cites Decoupled Weight Decay Regularization.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Decoupled Weight Decay Regularization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.377682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.377682Z digest=sha256:721e65a0f78dda949d1951c6003133c4065a833631a89cf0ee670f69fa2726cd

Observation 721e8b60-a827-4b75-bf89-3337f723d851 · outbound

This paper cites Haloquest: A visual hallucination dataset for advancing multimodal reasoning.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Haloquest: A visual hallucination dataset for advancing multimodal reasoning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:19.954610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.381749Z digest=sha256:19df38099083922d450f08f76c9e7d1a5f008cf67fdd03a1053478e1314556eb

Observation 8cc0add7-bea0-48fe-baad-d63090dd7ee1 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Evaluating Object Hallucination in Large Vision-Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.385833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.385833Z digest=sha256:80a71fc4a591a3b6373f51996f8ceba5265c7b70d8beb1140698a50b90c26151

Observation c82111ec-6b2e-4b60-9ff0-465307f7a97b · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.390047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.390047Z digest=sha256:995816be10d7baaaa1d75ef47ca839c7a3ce1c9be393a34cffbe3f225f774125

Observation 9220180b-39ae-4622-aa9e-460ed9cef1f2 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.395003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.395003Z digest=sha256:a66863ed0ecb494b2700c9682b44b0b20057922dd2d58dc520ebc21e07a05cb5

Observation 4ca48f0e-6a07-4bc5-ad2c-1bd431309bfb · outbound

This paper cites A survey of multimodel large language models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs A survey of multimodel large language models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:19.923518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.399460Z digest=sha256:8460dfa0ff32d15624bb0a91aa54da4cfde51b59df8d17e36e265b868702fcc4

Observation be649093-e7c6-4ac7-bd39-c3a49928230a · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Referitgame: Referring to objects in photographs of natural scenes

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.403801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.403801Z digest=sha256:23bb3a566dcd57bd080e25e7f05077eb68d6d503fa8a016ad8b5db278c8f470b

Observation 517d6065-62a0-40aa-91b7-816888cdbd26 · outbound

This paper cites Empowering Segmentation Ability to Multi-modal Large Language Models.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Empowering Segmentation Ability to Multi-modal Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.407889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.407889Z digest=sha256:b13f02229405d5d43bf4149c63d592c7020c127f8cc1de0d7f082868cd74df81

Observation 5d59454f-1ad6-49e5-9138-13c7cecfad0e · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.413192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.413192Z digest=sha256:151adea4c82a0c2ba1e5b06d7a77207904f73bea7e57d876a72201d806fa3f60

Observation e0cc1135-b213-4530-bd64-d317a063904d · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.418235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.418235Z digest=sha256:496d328ca3d43c1d465924da62e8d41c136041c3093da392d7c7799009e7ebbb

Observation b6a2fb75-2e9a-4ac6-b8cc-131aba5b33ac · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:27:19.899780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:27:19.422402Z digest=sha256:f760699c518150ac3d2291f684825a2518abd90aa6dd666d74cc9c85dd7f2077

Pith citing papers

Observation aabe33ca-8363-46f4-84c2-e6edcf9dd92d · inbound

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding cites this paper.

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T04:55:07.037081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T04:55:07.037081Z digest=sha256:ee0face5780932cfac2fb856b30c51747ef65f4dd6906b7944b6ccf6e616a986