Pith. sign in

Paper Citation Record · LEDGER

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach

As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 4 inbound Pith citation observations for arXiv:2411.17760.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17760 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:42:49.344975Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:18:40.043652Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T23:08:36.273358Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9664206d-b1d4-4869-8eb1-0818ffc363dc · outbound

This paper cites Pixtral 12B.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Pixtral 12B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.163096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.163096Z digest=sha256:00ac3e630df73fe04c4a5ba9a9d98cab56762780ae09e2a84d3219fca0dc2922

Observation 5e3ee981-f43b-485c-9ede-81c4fea061cf · outbound

This paper cites Understanding Alignment in Multimodal LLMs: A Comprehensive Study.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Understanding Alignment in Multimodal LLMs: A Comprehensive Study

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.168659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.168659Z digest=sha256:75b0c42ae61bebad142a28371a4fb6aa77fa77572be42dd74ba25c201bb18de8

Observation 17e6d224-bede-4c49-b14f-71b440800297 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.173432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.173432Z digest=sha256:eb4edda22c0283be3e76dec201420a724d9b901c27e399993346793f6303919c

Observation 75bb9f34-627c-4ae2-815e-6857afc292fa · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.178130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.178130Z digest=sha256:c86575def6894f67d352cb139800bb13aaa415d85a4aed5b43950019463ca07e

Observation d7091b5f-b2e2-45a6-adeb-5ff4feaac8d8 · outbound

This paper cites Enhancing Large Vision Language Models with Self-Training on Image Comprehension.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Enhancing Large Vision Language Models with Self-Training on Image Comprehension

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.182518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.182518Z digest=sha256:36d490894ee9b19532c1331147616e6292c64fcaf01adaaaba2e85df8a7b1b30

Observation af9c54af-75df-4baf-a867-4d0c221a1384 · outbound

This paper cites The Llama 3 Herd of Models.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.187310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.187310Z digest=sha256:c29536798a9daaa8eb8d579d4cdb60681ac3c89d77ec3595057ae22820f0e81e

Observation c0df4a06-4e16-49d9-af38-74f31c16f0d1 · outbound

This paper cites Multi-modal hal- lucination control by visual information grounding.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Multi-modal hal- lucination control by visual information grounding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:50.097804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.191662Z digest=sha256:c51af96146e105b792ebab233d90cdb0adee874e4e95ecf5d0d5599ab2053bb3

Observation a6a8cf59-d076-4253-8f86-fb6f5e809efa · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.195846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.195846Z digest=sha256:201737739788755d4d3b8a1a11cfee5e6e262f22df4687b5c8dd688cc3380be1

Observation 372d4f4a-0fa0-4afe-b7ee-558f26b3504b · outbound

This paper cites Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:50.082783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.200633Z digest=sha256:502e6cfa378f4ec9ec5ae560ab2fea66662439526b78682710758d490cf46d92

Observation 33434ecd-d986-47a4-aa56-a16a269641ee · outbound

This paper cites Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:50.067522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.205414Z digest=sha256:5ad72e61199f5ec6da00a1201e147d64484f9c9862307be0708b1f294c59ec0b

Observation 379e3394-eb4e-4230-89de-79b16e85e071 · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Silkie: Preference Distillation for Large Visual Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.210147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.210147Z digest=sha256:8d449e4ee15d5781aa240c512e32de8b39077387efbf78aefc6286ab99c50dae

Observation fcd3d2ba-a673-4504-9a7c-c4e75ec25384 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.214533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.214533Z digest=sha256:af06832be8a052b6f5cae6d6a79d5f456be801241a4158a2b2cee9caaf7e736a

Observation 2f0b93c8-b5df-4f57-9bc7-14f2f304cf2e · outbound

This paper cites Improved baselines with visual instruction tuning.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Improved baselines with visual instruction tuning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:50.052599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.219391Z digest=sha256:ac2234cb81ccb36ccaabad4bae55e21050f167109fcbfbdb822e3cec0bcdd7a7

Observation 359f8dce-5ce2-4753-be9e-407e9a530613 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.223779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.223779Z digest=sha256:a3f1e1ccc6e43e9eaca1fc97483a7115123b8b171ad36da25cae732e3e3e0433

Observation 201c41a6-b09f-4247-875f-d7e42ef98740 · outbound

This paper cites Visual instruction tuning.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Visual instruction tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.228071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.228071Z digest=sha256:302f4417ac70689892869ad9d473fec5492ba4a23b833bbeb15d2c95128b909c

Observation fd6427e1-1030-4dc9-bea9-b48249d23e54 · outbound

This paper cites CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.232514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.232514Z digest=sha256:f40a20851374f02f676a61d889d37d36aefcc1af209b11e9d2108ec39941c67b

Observation 361d47bc-3e91-40a3-a0fc-e3bf17a30525 · outbound

This paper cites Training language models to follow instructions with human feedback.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Training language models to follow instructions with human feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.236969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.236969Z digest=sha256:21d8c4b6d4af119148ec78050c9b68d7c75fb3ba835fc620e6ee5bc2afb3a9bf

Observation fa9b8285-67da-410a-b1cd-18b65fffeb48 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Direct preference optimization: Your language model is secretly a reward model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:50.008652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.241298Z digest=sha256:c209b32d3f33ada929d002af3253551ed738f9e6ca498b5b63278f56c6fe8606

Observation a979d35e-c0be-4df4-bd71-34a8209f3ca7 · outbound

This paper cites Object Hallucination in Image Captioning.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Object Hallucination in Image Captioning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.246029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.246029Z digest=sha256:627fb83e0735d55ecb0e8309b4b69d5c1018005617d6dbf2aa2e73b9aa6e332d

Observation 660a6b1d-8f45-4379-a4f5-fc496e0a5cd4 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.250628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.250628Z digest=sha256:47e26d6314d36c861f2dc923a548fe3a1142c4b26384a4dd05a4054f942b0eea

Observation 021fd46d-24eb-4e5f-abac-454855afb5ad · outbound

This paper cites A Survey on Self-Evolution of Large Language Models.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach A Survey on Self-Evolution of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.255192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.255192Z digest=sha256:616f886d6ecc416f1d895117a61661c5e14de42d9ac98ac904b182971975c253

Observation afac81d5-79f2-44d9-baed-07cc2a12bcca · outbound

This paper cites Planbench: An extensible benchmark for evaluating large language mod- els on planning and reasoning about change.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Planbench: An extensible benchmark for evaluating large language mod- els on planning and reasoning about change

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:49.994252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.259767Z digest=sha256:6a21678d8d5fd402c3ec7fda2e033875ef5e8dc50770638565657dd3dacf39e8

Observation b76cbf24-7d04-465b-ad59-d167cca69c75 · outbound

This paper cites LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.263920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.263920Z digest=sha256:efb3c42d7745ebe5063fb39c6c2d0d01828252606154782ffe28d44bd78ea06e

Observation edea6d44-21f1-4b73-93e6-bf0f9cdef135 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach CogVLM: Visual Expert for Pretrained Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.268335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.268335Z digest=sha256:b2df440bd403af91f326c459155fc04f4cf0f5d375967e0bec10100cbb94ef78

Observation e873e831-6705-4fb1-932d-1d76b66ed484 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional hu- man feedback.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional hu- man feedback

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:49.980201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.272884Z digest=sha256:d1d55bd95358ccc23b294a1f54fddc095531a82dd40f2566adde098df0ef8d18

Observation b75e6337-6f51-457c-8ca2-e82c3a73f64d · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.277290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.277290Z digest=sha256:b383923b7466c720cf682a4ec81438716fa01a0b0d96892fbc662e595ffa2a3f

Observation 7a023b58-55ce-49b4-a52a-f0c56adade20 · outbound

This paper cites Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.281803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.281803Z digest=sha256:2c83e2dde677fe22b99b55d3138de673cb8aaa2865e05e1c291ba58d85535dba

Observation fedf3290-a067-4888-9074-b4406b4c1ae6 · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.286239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.286239Z digest=sha256:17925c044caf014962d0d5cd57b853d0fdb48f5c85f7cf1619f0177e84704fb3

Observation 88cb4e6c-d4a0-49fd-b3cc-5a46da43d503 · outbound

This paper cites Analyzing and Mitigating Object Hallucination in Large Vision-Language Models.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Analyzing and Mitigating Object Hallucination in Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.290545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.290545Z digest=sha256:5e67caa7acdd76013d82ef231bd7e0674949af2e3650e2e5af46c63fcc9b038d

Observation 584df2d3-30ba-4c47-9cc3-f1d50f109cfa · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:42:49.295242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:42:49.295242Z digest=sha256:07f048042b9c4581d07a724119cc567a795a5d4e3fa643ed1b4f0ef096646339

Observation c3de1e84-38b5-4f68-92d3-cb91b1ff467f · outbound

This paper cites an unresolved cited work.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:42:49.965730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.299970Z digest=sha256:cfbb2ce6c1a50a6cbb1ae4c142420cb6d492925963c731d28a185433bd62e76e

Observation 8093b29e-f648-4049-a2ae-e032ca3521f3 · outbound

This paper cites - The salad is correctly noted as being part of the spread, but the caption could be more specific about the contents of the salad.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach - The salad is correctly noted as being part of the spread, but the caption could be more specific about the contents of the salad

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:49.951336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.304830Z digest=sha256:c882e77db6e4bd41f552cfc77775dbe2de894250aa89ff079d8c6d3417f45b34

Observation f7d23ead-32f6-4401-8886-e3249bd6a60e · outbound

This paper cites However, it inaccurately refers to a wine glass; the image shows glasses of what appears to be a juice or iced tea rather than wine glasses.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach However, it inaccurately refers to a wine glass; the image shows glasses of what appears to be a juice or iced tea rather than wine glasses

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:49.937245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.309115Z digest=sha256:c5772cc92ae7d06b01bbf986f079ec06cbb9308d9abf5a4a984a2a474e837f47

Observation 4b3f4c28-64c5-45c3-9935-d37023f4c9c1 · outbound

This paper cites - The mention of a potted plant is inaccurate; while there is foliage in the background, it cannot be clearly identified as a potted plant.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach - The mention of a potted plant is inaccurate; while there is foliage in the background, it cannot be clearly identified as a potted plant

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:49.922953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.313325Z digest=sha256:abb3928d28220bf8f29ce8a0845df066bfcce7a741ca851b8ac701d6b710b10d

Observation 2e97b3bb-d8b5-4eef-bf42-56d7646867ce · outbound

This paper cites The scene includes forks but not knives.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach The scene includes forks but not knives

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:49.908289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.317986Z digest=sha256:399c89d87e058e070b9f793e5fca62fc3fd38dee08e0fb8f3a780695723f842a

Observation f94029b2-2584-4ce3-a07b-e72cce4a1f55 · outbound

This paper cites an unresolved cited work.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:42:49.893274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.322769Z digest=sha256:41d539449cc80b3b29f8f7e5f3fa60825ece3b1691381bb6a4e314458f4eba0b

Observation 37d3f887-7c41-4f0b-ba8a-803fc1a0bf16 · outbound

This paper cites an unresolved cited work.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:42:49.876922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.327257Z digest=sha256:cbc5e86fc0250eb14dc852bd5fa2f3569daaadd7fe34479d7866b6253d6ba3bc

Observation 8f869abc-499c-4d5b-b7f1-bf423f33c387 · outbound

This paper cites an unresolved cited work.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:42:49.862174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.331632Z digest=sha256:be4b5b298580ba3f0f7fae8cbfb4158425f8c88f41cb1e1a181793e6838db939

Observation aedf9a77-1b0b-4dd9-94a4-7e0014abfc5f · outbound

This paper cites It connects this diagram to the essay's content, although it does not specify what the diagram illustrates.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach It connects this diagram to the essay's content, although it does not specify what the diagram illustrates

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:49.846812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.335819Z digest=sha256:2d757bc21339d1aeee2d9c90f9627d1f83fae10ff507c31644e867706bb5b7f7

Observation 1185ff57-3593-4cce-ae4b-238674a23caf · outbound

This paper cites an unresolved cited work.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:42:49.831471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.340316Z digest=sha256:5f98ba1e1d515fa30c5a16acd09e20775e817af659b3cabdc91450cb7032ae51

Observation 11ce3f01-c141-48c3-96f2-abd6d739b99f · outbound

This paper cites Overall, the caption effectively captures the primary elements of the image, including the handwriting, topic, and visual characteristics.

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach Overall, the caption effectively captures the primary elements of the image, including the handwriting, topic, and visual characteristics

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:42:49.816188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:42:49.344975Z digest=sha256:9ff15f9a0cf1d83e5c12f0a64210aeebba1f40434097ca78f69169fe31720bad

Pith citing papers

Observation 0d0af034-c869-4ba6-b326-848eca0e2cd2 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.278054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:a01af1d345aa912afe05abfacac4303fab0e9d185ab17604ce7e51450edf7cc3

Observation 612bf3fb-f15e-4b05-85a7-cc4a788cbf1d · inbound

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model cites this paper.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.043652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.043652Z digest=sha256:da897aa54a27982c18916e33c33b21538484e2242a6a642db4b69c2696b2383a

Observation 700e3458-958d-4cc6-9c8f-fd54624f5400 · inbound

Visual Preference Optimization with Rubric Rewards cites this paper.

Visual Preference Optimization with Rubric Rewards Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:00.904493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T15:45:52.980881Z digest=sha256:994ee6f77a1b3af09426ed79ec99f6fbd206beaacdb4cff4ff3c67e5ffc14f6a

Observation 6218e471-fbd3-45c2-985f-6b97bc14c7d4 · inbound

CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs cites this paper.

CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T12:29:07.424531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T12:29:07.424531Z digest=sha256:d6eb2a2d9c3a7411cd325c5cdb9526ccaef69f07676de517b963eae500e7da13