Pith. sign in

Paper Citation Record · LEDGER

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach

As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 4 inbound Pith citation observations for arXiv:2411.17760.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17760 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:42:49.344975Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:18:40.043652Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T23:08:36.273358Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9664206d-b1d4-4869-8eb1-0818ffc363dc · outbound

This paper cites Pixtral 12B.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Pixtral 12B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.163096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.163096Z digest=sha256:00ac3e630df73fe04c4a5ba9a9d98cab56762780ae09e2a84d3219fca0dc2922

Observation 5e3ee981-f43b-485c-9ede-81c4fea061cf · outbound

This paper cites Understanding Alignment in Multimodal LLMs: A Comprehensive Study.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Understanding Alignment in Multimodal LLMs: A Comprehensive Study

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.168659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.168659Z digest=sha256:72840a21b5864e15ce0c5b8a3d47df6334904c173b66e62b949adb70413aa1d4

Observation 17e6d224-bede-4c49-b14f-71b440800297 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.173432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.173432Z digest=sha256:eb4edda22c0283be3e76dec201420a724d9b901c27e399993346793f6303919c

Observation 75bb9f34-627c-4ae2-815e-6857afc292fa · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.178130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.178130Z digest=sha256:c86575def6894f67d352cb139800bb13aaa415d85a4aed5b43950019463ca07e

Observation d7091b5f-b2e2-45a6-adeb-5ff4feaac8d8 · outbound

This paper cites Enhancing Large Vision Language Models with Self-Training on Image Comprehension.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Enhancing Large Vision Language Models with Self-Training on Image Comprehension

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.182518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.182518Z digest=sha256:36d490894ee9b19532c1331147616e6292c64fcaf01adaaaba2e85df8a7b1b30

Observation af9c54af-75df-4baf-a867-4d0c221a1384 · outbound

This paper cites The Llama 3 Herd of Models.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.187310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.187310Z digest=sha256:c29536798a9daaa8eb8d579d4cdb60681ac3c89d77ec3595057ae22820f0e81e

Observation c0df4a06-4e16-49d9-af38-74f31c16f0d1 · outbound

This paper cites Multi-modal hal- lucination control by visual information grounding.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Multi-modal hal- lucination control by visual information grounding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:50.097804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.191662Z digest=sha256:cddc57c8c0cbe9734d2cf6ba8afd26415ac6735ced08ca65b23920e6d59a601d

Observation a6a8cf59-d076-4253-8f86-fb6f5e809efa · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.195846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.195846Z digest=sha256:201737739788755d4d3b8a1a11cfee5e6e262f22df4687b5c8dd688cc3380be1

Observation 372d4f4a-0fa0-4afe-b7ee-558f26b3504b · outbound

This paper cites Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:50.082783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.200633Z digest=sha256:a34889a88184fa4693073aed2ccf24c228a3f23720e2f2bffeba6bfffd457327

Observation 33434ecd-d986-47a4-aa56-a16a269641ee · outbound

This paper cites Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:50.067522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.205414Z digest=sha256:1824e1e358c5199fee6136a1be842e66cb15af6b628934e68d51d7a7949532f0

Observation 379e3394-eb4e-4230-89de-79b16e85e071 · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Silkie: Preference Distillation for Large Visual Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.210147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.210147Z digest=sha256:8d449e4ee15d5781aa240c512e32de8b39077387efbf78aefc6286ab99c50dae

Observation fcd3d2ba-a673-4504-9a7c-c4e75ec25384 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.214533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.214533Z digest=sha256:af06832be8a052b6f5cae6d6a79d5f456be801241a4158a2b2cee9caaf7e736a

Observation 2f0b93c8-b5df-4f57-9bc7-14f2f304cf2e · outbound

This paper cites Improved baselines with visual instruction tuning.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Improved baselines with visual instruction tuning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:50.052599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.219391Z digest=sha256:8f7702cc86ea63a8a2c96e85a30b135238cae78c67bd568293fe6823b6a3386a

Observation 359f8dce-5ce2-4753-be9e-407e9a530613 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.223779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.223779Z digest=sha256:a3f1e1ccc6e43e9eaca1fc97483a7115123b8b171ad36da25cae732e3e3e0433

Observation 201c41a6-b09f-4247-875f-d7e42ef98740 · outbound

This paper cites Visual instruction tuning.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Visual instruction tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.228071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.228071Z digest=sha256:302f4417ac70689892869ad9d473fec5492ba4a23b833bbeb15d2c95128b909c

Observation fd6427e1-1030-4dc9-bea9-b48249d23e54 · outbound

This paper cites CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.232514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.232514Z digest=sha256:f40a20851374f02f676a61d889d37d36aefcc1af209b11e9d2108ec39941c67b

Observation 361d47bc-3e91-40a3-a0fc-e3bf17a30525 · outbound

This paper cites Training language models to follow instructions with human feedback.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Training language models to follow instructions with human feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.236969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.236969Z digest=sha256:21d8c4b6d4af119148ec78050c9b68d7c75fb3ba835fc620e6ee5bc2afb3a9bf

Observation fa9b8285-67da-410a-b1cd-18b65fffeb48 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Direct preference optimization: Your language model is secretly a reward model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:50.008652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.241298Z digest=sha256:b8537f2f4577586737d090c1853605605491d77641b4dab1bad1de9249496fe2

Observation a979d35e-c0be-4df4-bd71-34a8209f3ca7 · outbound

This paper cites Object Hallucination in Image Captioning.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Object Hallucination in Image Captioning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.246029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.246029Z digest=sha256:627fb83e0735d55ecb0e8309b4b69d5c1018005617d6dbf2aa2e73b9aa6e332d

Observation 660a6b1d-8f45-4379-a4f5-fc496e0a5cd4 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.250628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.250628Z digest=sha256:47e26d6314d36c861f2dc923a548fe3a1142c4b26384a4dd05a4054f942b0eea

Observation 021fd46d-24eb-4e5f-abac-454855afb5ad · outbound

This paper cites A Survey on Self-Evolution of Large Language Models.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach A Survey on Self-Evolution of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.255192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.255192Z digest=sha256:616f886d6ecc416f1d895117a61661c5e14de42d9ac98ac904b182971975c253

Observation afac81d5-79f2-44d9-baed-07cc2a12bcca · outbound

This paper cites Planbench: An extensible benchmark for evaluating large language mod- els on planning and reasoning about change.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Planbench: An extensible benchmark for evaluating large language mod- els on planning and reasoning about change

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:49.994252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.259767Z digest=sha256:aa0c3cbb27e4a375757f9c64d7524306dc0f819816f9033c31e33d656985917c

Observation b76cbf24-7d04-465b-ad59-d167cca69c75 · outbound

This paper cites LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.263920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.263920Z digest=sha256:efb3c42d7745ebe5063fb39c6c2d0d01828252606154782ffe28d44bd78ea06e

Observation edea6d44-21f1-4b73-93e6-bf0f9cdef135 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach CogVLM: Visual Expert for Pretrained Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.268335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.268335Z digest=sha256:b2df440bd403af91f326c459155fc04f4cf0f5d375967e0bec10100cbb94ef78

Observation e873e831-6705-4fb1-932d-1d76b66ed484 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional hu- man feedback.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional hu- man feedback

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:49.980201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.272884Z digest=sha256:8a4ef5cbb5cb9815e4c44d18d1e4b3d29cc0d952d20ee0f6b0abf0d5270731f1

Observation b75e6337-6f51-457c-8ca2-e82c3a73f64d · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.277290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.277290Z digest=sha256:b383923b7466c720cf682a4ec81438716fa01a0b0d96892fbc662e595ffa2a3f

Observation 7a023b58-55ce-49b4-a52a-f0c56adade20 · outbound

This paper cites Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.281803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.281803Z digest=sha256:2c83e2dde677fe22b99b55d3138de673cb8aaa2865e05e1c291ba58d85535dba

Observation fedf3290-a067-4888-9074-b4406b4c1ae6 · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.286239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.286239Z digest=sha256:17925c044caf014962d0d5cd57b853d0fdb48f5c85f7cf1619f0177e84704fb3

Observation 88cb4e6c-d4a0-49fd-b3cc-5a46da43d503 · outbound

This paper cites Analyzing and Mitigating Object Hallucination in Large Vision-Language Models.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Analyzing and Mitigating Object Hallucination in Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.290545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.290545Z digest=sha256:5e67caa7acdd76013d82ef231bd7e0674949af2e3650e2e5af46c63fcc9b038d

Observation 584df2d3-30ba-4c47-9cc3-f1d50f109cfa · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.295242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.295242Z digest=sha256:07f048042b9c4581d07a724119cc567a795a5d4e3fa643ed1b4f0ef096646339

Observation c3de1e84-38b5-4f68-92d3-cb91b1ff467f · outbound

This paper cites an unresolved cited work.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:42:49.965730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.299970Z digest=sha256:8bc0f55b4afedccf2e7188fba8e68def8a622eabf19d33bdb30df544e09753c5

Observation 8093b29e-f648-4049-a2ae-e032ca3521f3 · outbound

This paper cites - The salad is correctly noted as being part of the spread, but the caption could be more specific about the contents of the salad.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach - The salad is correctly noted as being part of the spread, but the caption could be more specific about the contents of the salad

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:49.951336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.304830Z digest=sha256:c8974e47ed748e029d7e7f0b3ac28f77c77563237b03865ab38915e4f0869e16

Observation f7d23ead-32f6-4401-8886-e3249bd6a60e · outbound

This paper cites However, it inaccurately refers to a wine glass; the image shows glasses of what appears to be a juice or iced tea rather than wine glasses.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach However, it inaccurately refers to a wine glass; the image shows glasses of what appears to be a juice or iced tea rather than wine glasses

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:49.937245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.309115Z digest=sha256:1a82d787c26222ba5023cc94e7e1231d6aaecff5f359f8393b045e0bd8220b11

Observation 4b3f4c28-64c5-45c3-9935-d37023f4c9c1 · outbound

This paper cites - The mention of a potted plant is inaccurate; while there is foliage in the background, it cannot be clearly identified as a potted plant.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach - The mention of a potted plant is inaccurate; while there is foliage in the background, it cannot be clearly identified as a potted plant

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:49.922953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.313325Z digest=sha256:b1e3c2dc30cf52b6e364b4fe607d9a0f06e464ea20f84a4022120eeafe08e013

Observation 2e97b3bb-d8b5-4eef-bf42-56d7646867ce · outbound

This paper cites The scene includes forks but not knives.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach The scene includes forks but not knives

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:49.908289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.317986Z digest=sha256:f9e361323a56e1abbb024b727ffaa0b77d77dce37aa88de2f9784c42c3062251

Observation f94029b2-2584-4ce3-a07b-e72cce4a1f55 · outbound

This paper cites an unresolved cited work.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:42:49.893274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.322769Z digest=sha256:b83a27080bd659ed3c319cb93b742c35aadd7ca5bc20cb283c4ec2c50f8e0c7b

Observation 37d3f887-7c41-4f0b-ba8a-803fc1a0bf16 · outbound

This paper cites an unresolved cited work.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:42:49.876922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.327257Z digest=sha256:d7e35ad0ccbb6190c660cafa5f35e69007c4ef5b387c91bd178eb3b81dc7e40c

Observation 8f869abc-499c-4d5b-b7f1-bf423f33c387 · outbound

This paper cites an unresolved cited work.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:42:49.862174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.331632Z digest=sha256:3a0e64724d34269a0ed10a720a76d752cd766ad8595e5178335512ccec7e925e

Observation aedf9a77-1b0b-4dd9-94a4-7e0014abfc5f · outbound

This paper cites It connects this diagram to the essay's content, although it does not specify what the diagram illustrates.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach It connects this diagram to the essay's content, although it does not specify what the diagram illustrates

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:49.846812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.335819Z digest=sha256:1bb0939975946039129fb2dd00acc43c664ca5f1eb87a450d575ea4f8df830fd

Observation 1185ff57-3593-4cce-ae4b-238674a23caf · outbound

This paper cites an unresolved cited work.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:42:49.831471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.340316Z digest=sha256:8ba4c6851e65e92e7ebc7450fbc3861c0a9a8dc832a9e2bde5c074b64e39ccfc

Observation 11ce3f01-c141-48c3-96f2-abd6d739b99f · outbound

This paper cites Overall, the caption effectively captures the primary elements of the image, including the handwriting, topic, and visual characteristics.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Overall, the caption effectively captures the primary elements of the image, including the handwriting, topic, and visual characteristics

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:49.816188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:42:49.344975Z digest=sha256:f92c829105714e2343eb4a4741d62a2dfdb3fe235c6efb8151ae30a6d9aefea5

Pith citing papers

Observation 0d0af034-c869-4ba6-b326-848eca0e2cd2 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.278054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:ecc760c7c9bbaf3531eb16e762ba3a0472dbe195f393838ba3e610742595e4c3

Observation 612bf3fb-f15e-4b05-85a7-cc4a788cbf1d · inbound

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model cites this paper.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.043652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.043652Z digest=sha256:da897aa54a27982c18916e33c33b21538484e2242a6a642db4b69c2696b2383a

Observation 700e3458-958d-4cc6-9c8f-fd54624f5400 · inbound

Visual Preference Optimization with Rubric Rewards cites this paper.

Visual Preference Optimization with Rubric Rewards Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:00.904493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:45:52.980881Z digest=sha256:0580cfc59935e2cdb1d783411a9fa8740384517ec3506db40644d78c1c656649

Observation 6218e471-fbd3-45c2-985f-6b97bc14c7d4 · inbound

CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs cites this paper.

CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T12:29:07.424531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T12:29:07.424531Z digest=sha256:d6eb2a2d9c3a7411cd325c5cdb9526ccaef69f07676de517b963eae500e7da13