Pith. sign in

Paper Citation Record · LEDGER

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction

As of 10 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2502.00717.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00717 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:02:45.221151Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T08:14:36.819533Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b22dc06-eb47-44c9-9f03-282b443d89b8 · outbound

This paper cites Qwen Technical Report.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.050009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.050009Z digest=sha256:4a4dcad4113ccb12cfcbe5424b21a9d6929eb6abfeafa4e0c86f03da8c798728

Observation 343c750e-369e-41c8-9f7d-1aea43af12e6 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.055692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.055692Z digest=sha256:1de3bebe325fd01269d182fb98c61e01fec787fe1165518f08f3b738752e63d8

Observation ab325d68-16af-4cea-b535-77d1c8190fff · outbound

This paper cites Y., Bhiwandiwalla, A., Tseng, S.-Y., Olson, M.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Y., Bhiwandiwalla, A., Tseng, S.-Y., Olson, M

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.747062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.060876Z digest=sha256:71f2b9f991f4b2a33500976d7a4eb1ad003ff1cf0211187e53970dc04728b55c

Observation 9559d8da-160f-4045-9d0a-02a68d05ef0e · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Honeybee: Locality-enhanced projector for multimodal llm

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.067233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.067233Z digest=sha256:584fe00dedac25be6d0700a798805ca542a64d9db659ebede9329fcaecdd27ec

Observation 383ebc9f-4bf6-4de8-bc9d-85bb801cb0f0 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.723271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.072247Z digest=sha256:a4f6805aca1eac05adbacac6557facaf40db231d439236e64581e9d884683367

Observation f88f5eed-8fc8-4fd3-bb2b-bb5685368e1b · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.708557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.076964Z digest=sha256:15b25c3c1069374b8aa39bb4ed9f17f6ce7ba9673de322b2d576f97e1f7b2bee

Observation 89cc4221-5b42-401d-86b6-b2dbb14f532e · outbound

This paper cites E., Stoica, I., and Xing, E.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction E., Stoica, I., and Xing, E

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.082246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.082246Z digest=sha256:76e4cc21f2dad1ae376b9701b82d7a15e52619fd2b3a15864f480d81c6905712

Observation d6157353-10cd-43f9-86e8-79faf796d208 · outbound

This paper cites Vision transformers need registers.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Vision transformers need registers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.684270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.086650Z digest=sha256:543a8a10e5a1e5245c690de26da9122315283cdb3930a6ea2b46b94fdfd5c01f

Observation 5bf0cfb9-7f44-40ca-a3f1-b9f54ccac391 · outbound

This paper cites Intern LM - XC omposer2-4 KHD : A pioneering large vision-language model handling resolutions from 336 pixels to 4k HD.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Intern LM - XC omposer2-4 KHD : A pioneering large vision-language model handling resolutions from 336 pixels to 4k HD

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.670033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.091136Z digest=sha256:9778e011f056b446ceb96e3d2ee9c444bcb73f5070b014601c2edc4eefdc0764

Observation e372e6e3-31ba-4f6a-865b-bc629db98ea6 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.095827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.095827Z digest=sha256:071e8a3d24f289f91dff67bfedab8cb0391b6d99390a4f7ef0e9e26671797af6

Observation 2f08a5cf-e9aa-4df2-b351-654abf5b0a14 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.100605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.100605Z digest=sha256:0e887f049c2d193c2297bfaad3b9c25f2f1a85089974a0438c902c77ae013a86

Observation 95cfe3cb-7e6d-4a2f-baa9-b6167e64caf6 · outbound

This paper cites A., Ma, W.-C., and Krishna, R.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction A., Ma, W.-C., and Krishna, R

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.651387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.105324Z digest=sha256:44de125beedb1b268c1b168f9f76105df2616565465cfe6b520b2885bf057cf8

Observation 4ffa2056-e84d-400b-a949-dc8d7f4c3210 · outbound

This paper cites Detecting and preventing hallucinations in large vision language models.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Detecting and preventing hallucinations in large vision language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.109865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.109865Z digest=sha256:48327f8ad29710e7293de29b1ba4156858ba391a5220bccb977fc70627bd0d71

Observation 3a2aeaed-9f09-4a0f-a3b5-70f50e8d4156 · outbound

This paper cites J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.114157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.114157Z digest=sha256:87605db882f0cdf7513db5d527cbe06fc50ae06bfc842262102a7fb092207f6b

Observation 30b95d8e-1ec7-475b-afcf-3c95fb442e79 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Lisa: Reasoning segmentation via large language model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.118566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.118566Z digest=sha256:8fe7206c61053832cd0d9a65f75a78d081ed3db5024ac3b4136afdbb61e51c92

Observation 349682b0-1ac4-443d-82db-dc44b05e0421 · outbound

This paper cites Mitigating object hallucinations in large vision-language models through visual contrastive decoding.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Mitigating object hallucinations in large vision-language models through visual contrastive decoding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.122378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.122378Z digest=sha256:8f532b1e68f27207fdc9c187832ed633daa7d2bbedcf9cb55ca347b724903bd0

Observation 64779bfe-e615-4e27-a556-e84bd1fa8e30 · outbound

This paper cites L., Holtzman, A., Fried, D., Liang, P., Eisner, J., Hashimoto, T.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction L., Holtzman, A., Fried, D., Liang, P., Eisner, J., Hashimoto, T

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.599498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.126301Z digest=sha256:8799930d714eb29a16ec48e50189e51bca4a0f4f5d74d8b272e4ac7017e9bfad

Observation f62d742f-848e-496e-880e-a38e7393ee98 · outbound

This paper cites X., and Wen, J.-R.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction X., and Wen, J.-R

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.584380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.130189Z digest=sha256:b9fa1054f66a4ad958c1b90b91dec5a2e9e32066ce4f62375c7a2c31a5c6c385

Observation 781cc07a-de18-4ce9-95aa-a5efcc45a1b3 · outbound

This paper cites an unresolved cited work.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.134091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.134091Z digest=sha256:a5d36770e7f0fdca0ce4dd25a87343699d9c3e3cb17996c4ca4b0e4a21415c97

Observation 5672227d-54e7-428c-8ba6-7beafcbbd65c · outbound

This paper cites an unresolved cited work.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.138911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.138911Z digest=sha256:7060dba827fc6a0c2ec359c70b69c727eefd22acb683321e7f8145f03138534a

Observation 571167a3-04f2-4dbf-9284-baab24b883aa · outbound

This paper cites an unresolved cited work.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.142715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.142715Z digest=sha256:614007e05026a4eb8cf533d8651b3fcd62f904b6d78b93b34d3a15b5ca7408bf

Observation 97b06fc0-2d68-42f4-b18d-11fab6096eee · outbound

This paper cites an unresolved cited work.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.146514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.146514Z digest=sha256:6aec413896d9e69d0d873688a6994b716dc3d78fd2c53ede54d5584e189bf7bf

Observation 0252356e-de8d-423d-9efc-932b5513b6eb · outbound

This paper cites Paying more attention to image: A training-free method for alleviating hallucination in lvlms.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Paying more attention to image: A training-free method for alleviating hallucination in lvlms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.150444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.150444Z digest=sha256:06b114fcd0502cef45280663bb6d836a032b0e9f01afc6ce91f03790885c610b

Observation 08b1f8ac-079d-4cb0-9e07-ee0e7ba93298 · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.154459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.154459Z digest=sha256:53e4b816c72cde09d24045a33902e0cc90daea33750af12eacc1b4f4773b82af

Observation ebedfd82-17f5-42e2-aff5-020d954ec983 · outbound

This paper cites A., Burns, K., Darrell, T., and Saenko, K.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction A., Burns, K., Darrell, T., and Saenko, K

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.514090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.158411Z digest=sha256:edd680a301379b3a5e5fca28cae51e32c5a98e4221a52714386c3daae48f62d2

Observation ce2b2060-7f87-46b3-8445-9512446e5bbb · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction LLaMA: Open and Efficient Foundation Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.162801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.162801Z digest=sha256:2b94d42e577f93fa8f3f3e804da7b50fd579c7e0c21d01d157bf9f82df48b202

Observation d0aff9bc-1a78-4a8b-825f-341e3ee67787 · outbound

This paper cites Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.167476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.167476Z digest=sha256:c00864cc9b703d93acbe351b4c8de5787dcea005a3dc73c643fe266d4b597b31

Observation 6329c177-7233-409f-9149-ae22f4970fcc · outbound

This paper cites Dilu: A knowledge-driven approach to autonomous driving with large language models.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Dilu: A knowledge-driven approach to autonomous driving with large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.500108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.172458Z digest=sha256:980a2b3c50c9bffd0e28ee5f83d0f3e1a0f75295b9ed91bb45c89a4c4b8b1a22

Observation 369c3dfa-6660-49a0-befb-f956d42bfe32 · outbound

This paper cites Q-instruct: Improving low-level visual abilities for multi-modality foundation models.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Q-instruct: Improving low-level visual abilities for multi-modality foundation models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.485571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.176770Z digest=sha256:34c39d5a9ef3114f3e86af7e94abf3f46467a4e18b100099f9681c8f7e9f6b7c

Observation cf895cd9-62e4-4ca5-99a7-5fabac22c513 · outbound

This paper cites and Xie, S.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction and Xie, S

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.472927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.180929Z digest=sha256:b81cdda81427f23494ff9614d1b7feb0e240fa71887296feaf53034f7da0b52a

Observation e776ef32-e83d-4066-a204-ccbd657f4f72 · outbound

This paper cites D., and Potts, C.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction D., and Potts, C

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.458914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.185361Z digest=sha256:728d0c070ad80f63974afe982efba82148535c10b98660ef4802112b18b906c8

Observation e619e128-cb61-4748-b43d-8ae82552b336 · outbound

This paper cites Efficient streaming language models with attention sinks.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Efficient streaming language models with attention sinks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.444862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.189945Z digest=sha256:c23651cf14584dd69b881bc344ee6cc3d3230fa484c1aaf43be3c544aaf51be2

Observation e9bbe534-fe4a-4a91-a9c7-17caf404e538 · outbound

This paper cites Multi-modal concept alignment pre-training for generative medical visual question answering.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Multi-modal concept alignment pre-training for generative medical visual question answering

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.430732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.194609Z digest=sha256:a309edd15aa00e75d423761243380ad0904aa1cb7211d6b2d28c29dd47d3e2ee

Observation af8a30b9-e676-481c-b2c3-9729c9a16f13 · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.199090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.199090Z digest=sha256:52556bfc115b77d8063b78032ecd473d38c128a1ffe0119fddda5003b557808b

Observation eca925e7-d6ec-4a08-a5f2-101acc3a8cdc · outbound

This paper cites Attention prompting on image for large vision-language models.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Attention prompting on image for large vision-language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.416106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.203690Z digest=sha256:ede91acb5011c81ceb86728744f744b676c1bb7428819e961d831d13a1bd6f62

Observation 682319d0-9100-471a-960d-3840f09aa90a · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.401593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.208263Z digest=sha256:e8c719c138b5b5c35c926979ef17c279b6f32475d5fd51d0819791afe13158e2

Observation 6eccde3b-c143-49ff-b44d-ab87ed1cf2ea · outbound

This paper cites Llava-grounding: Grounded visual chat with large multimodal models.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction Llava-grounding: Grounded visual chat with large multimodal models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:02:45.386299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T18:02:45.212469Z digest=sha256:02a559f5be8ec2eca0cb26186be44c41384b09ab9bad4a6d152069ef7aa1e3c4

Observation c8558b2e-e214-4e40-ba7c-19c2b880d798 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.216603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.216603Z digest=sha256:f4b7f646cf2998fd7103429bc1c228db6d5a3c12cbbf9f97d8a82a306812b085

Observation 5cc824b5-e1ea-4247-9981-92659d3c69a5 · outbound

This paper cites write newline.

MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction write newline

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:45.221151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:45.221151Z digest=sha256:975ba0f429b1b075d068dd2bb41a88e045967c8f920f8d6ee3d46ae03583c36f

Pith citing papers

Observation fd86c048-156a-4a5a-8069-44b7abf78f3a · inbound

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation cites this paper.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.819533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.819533Z digest=sha256:6fc2a3e63b76c631feabebbfe3e31ab27dbc36cd8c6606a30dd4505db76114f2