Pith. sign in

Paper Citation Record · LEDGER

Improving Large Vision and Language Models by Learning from a Panel of Peers

As of 7 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2509.01610.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01610 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:27:27.748564Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved62
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1f574b3-4004-49c8-8e3c-55cab892c658 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Improving Large Vision and Language Models by Learning from a Panel of Peers Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.253037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.253037Z digest=sha256:da08a451435606e4d30ede11b33ba4a6ccdeb9113583118cdb48ebd2960d238e

Observation 4c935171-7f80-46b8-8f73-3b7aace4c50b · outbound

This paper cites GPT-4 Technical Report.

Improving Large Vision and Language Models by Learning from a Panel of Peers GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.258508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.258508Z digest=sha256:432d153bdeae61b37fbff509d7506e09666df85a8b10dbf805d984ead1ff397e

Observation 9fd19040-31cd-4fd4-a9e3-c43fd6659dd0 · outbound

This paper cites The Llama 3 Herd of Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.263431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.263431Z digest=sha256:70936978b505ae7f2c1460ed41d59c0104963fa5e92b74e65d98996179bd0894

Observation 416f724d-3c68-4ef6-a538-9b84233ed729 · outbound

This paper cites Theoretical guarantees on the best-of-n alignment policy.

Improving Large Vision and Language Models by Learning from a Panel of Peers Theoretical guarantees on the best-of-n alignment policy

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.267919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.267919Z digest=sha256:6786a3d4cb54f33847f1472020758e96244145b45f46912a7137ef833b87e76c

Observation 320ed74c-f5d9-420b-9c91-687bf8a4bf98 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Improving Large Vision and Language Models by Learning from a Panel of Peers PaliGemma: A versatile 3B VLM for transfer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.272581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.272581Z digest=sha256:c234058b9ee12693778097a0222be3eeb322f5be65a946d9bfebfcdbe9995fd3

Observation 06e2aed1-d8a9-4e36-ab9f-129346445690 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Improving Large Vision and Language Models by Learning from a Panel of Peers Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.277331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.277331Z digest=sha256:836dbc6c0a7de9575638bb36946536c74047bc02e3fbc06398355302bf9cb1bb

Observation f1af862d-6277-4e88-9be9-a3154e63f00d · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.281965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.281965Z digest=sha256:c9020bdc557d2a30750b9d7f3ee4d7b80e34332d5fc7ba8116a2303b10d54f74

Observation 28d8a27c-bb4b-46e4-aaa9-8f61d7bf5d73 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Improving Large Vision and Language Models by Learning from a Panel of Peers How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.286826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.286826Z digest=sha256:4b68bbe826ba10a5368d6ebef8fa4ec66a060795244f86da43b38bf5f54557b9

Observation d9b9f837-3e73-4667-bc9c-7e7ea26dece4 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Improving Large Vision and Language Models by Learning from a Panel of Peers Gonzalez, Ion Stoica, and Eric P

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.291111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.291111Z digest=sha256:8d93ae46a0b0a156e72c17a69cd86214f20d8900f70a085a3b14721969bcfebf

Observation cf6b9be5-c95d-4ac1-9d4a-fa4e70c5cf4f · outbound

This paper cites Self-Improving Robust Preference Optimization.

Improving Large Vision and Language Models by Learning from a Panel of Peers Self-Improving Robust Preference Optimization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:27:28.478077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.295655Z digest=sha256:ad7168f01f45ad142408f5d30d8b2d53c5ecaab4c86898561fb2fb98653cabf7

Observation 7849c7cf-8ac5-4f82-b7c4-a15dd2047b5f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Improving Large Vision and Language Models by Learning from a Panel of Peers Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.300251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.300251Z digest=sha256:a87b2dd6941728ae1379ffd9815b24d31049d9a8315e7dd68bd8e2a558b44bea

Observation 25de84a8-6ce4-4abb-94c1-6dcbf63da87a · outbound

This paper cites Reward model ensembles help mitigate overop- timization.

Improving Large Vision and Language Models by Learning from a Panel of Peers Reward model ensembles help mitigate overop- timization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.304752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.304752Z digest=sha256:d7764a1842bc81d7b11208bf0f5b21e1c092b8042deb14b4d93eda524bf85843

Observation f2b6f9bb-88c2-47ae-98a6-2035402afea5 · outbound

This paper cites Nvlm: Open frontier-class multimodal llms.

Improving Large Vision and Language Models by Learning from a Panel of Peers Nvlm: Open frontier-class multimodal llms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.309716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.309716Z digest=sha256:f7f26d572fb29b18ba7d67835cd1f1b6f7b15ba666727fd4dd2b8fe74ad099a3

Observation 0103621a-185e-4f2d-b475-f7f8af9f2323 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Improving Large Vision and Language Models by Learning from a Panel of Peers Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.314082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.314082Z digest=sha256:55a514777a71cace1c50acee32e5380540cf7933a943e2074a35d3057b5b7947

Observation 90da2d0b-414b-43cc-878b-39ed1d204b76 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.318382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.318382Z digest=sha256:e46cb245684f3af9febaa55c956fa4c621cc710f799f2baa3d9af8a75723c987

Observation f011d6d3-4fcb-48eb-8931-1cdf418b6bb7 · outbound

This paper cites Enhancing Large Vision Language Models with Self-Training on Image Comprehension.

Improving Large Vision and Language Models by Learning from a Panel of Peers Enhancing Large Vision Language Models with Self-Training on Image Comprehension

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.323258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.323258Z digest=sha256:809a04f618712b9e8873605a38e6f470374b553d2f0c2e50c88019e3afe66daf

Observation 66190a49-3bc1-4536-b81d-6a3e48288695 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Improving Large Vision and Language Models by Learning from a Panel of Peers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.327711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.327711Z digest=sha256:33fa7725ddc2b4fbb1900350495f1af7feedac2672a81786f4ee0e28f411c1a9

Observation 42bc6ae6-ce22-4d6e-9970-e4b71f5ef03c · outbound

This paper cites VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.332321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.332321Z digest=sha256:caa0aff0478250228876578f9b15ec33b3fcec8b9cccaaf12bc6b8c259fb9c79

Observation b32c752a-098d-4cc0-a4ed-b31bd5e24d6c · outbound

This paper cites VILA$^2$: VILA Augmented VILA.

Improving Large Vision and Language Models by Learning from a Panel of Peers VILA$^2$: VILA Augmented VILA

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.336733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.336733Z digest=sha256:a3bf658f3d77eac565975e8cbe0347d3e36b8ff56a125ddbf56753772e70d442

Observation 02edb8e5-c7e7-422a-8a61-a909e22db5cd · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 20

Resolution
malformed identifier
no resolver link, observed 2026-08-05T12:27:27.341177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.341177Z digest=sha256:16cbc9bdb9e575325abb96ec7d68c8d867d00365dad8d2a081a856ab35afb3aa

Observation e7d8a105-9d58-482a-a151-9c16b8c07adb · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Improving Large Vision and Language Models by Learning from a Panel of Peers CogVLM2: Visual Language Models for Image and Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.345868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.345868Z digest=sha256:59435b84cc77b187d1f494569d82580d9141c1c7d73c6e2606e54042b5291cf9

Observation 34a2a7b2-1722-49e8-82e6-09870a987a03 · outbound

This paper cites Mistral 7B.

Improving Large Vision and Language Models by Learning from a Panel of Peers Mistral 7B

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.350403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.350403Z digest=sha256:400856a34df72519fa9055a27081c77a1bc33b6359240e7aae14ea09a0fc6b88

Observation 9c3c92b1-5766-418b-8d6c-ac768568c5ef · outbound

This paper cites A diagram is worth a dozen images.

Improving Large Vision and Language Models by Learning from a Panel of Peers A diagram is worth a dozen images

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.354977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.354977Z digest=sha256:d831f8b29884bce700c899fa524ba88678056a462a4b73e089078424747a0763

Observation acd8ce11-70b5-4725-8770-f498c30bbdb2 · outbound

This paper cites What matters when building vision-language models?.

Improving Large Vision and Language Models by Learning from a Panel of Peers What matters when building vision-language models?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.359046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.359046Z digest=sha256:c496736a966210f4ae60af3b8d1fa62f4c903d69c61dd7e36b3d334f63e7cf18

Observation 3662a7e8-8af8-48b9-8b10-acc81c44ad5d · outbound

This paper cites Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation.

Improving Large Vision and Language Models by Learning from a Panel of Peers Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.363208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.363208Z digest=sha256:7863725cbf6054ac7f969ec4c339a4f4a8fe16221a63e28ffe65828c98e8843a

Observation 07985a7d-1f8e-4115-98be-191d10668a3f · outbound

This paper cites Seed-bench: Benchmark- ing multimodal large language models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Seed-bench: Benchmark- ing multimodal large language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.367800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.367800Z digest=sha256:c917a0071f60eab39d3fda73a3b9ca013a2407d16389c46d3e7f25e0c2562e06

Observation dba96a56-05ab-4b2b-b413-3f53fbd6fc1d · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Improving Large Vision and Language Models by Learning from a Panel of Peers LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.371859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.371859Z digest=sha256:6f457088ec58f283eb034207d261522eb5df8dfc352b96c55576ef9c190cc601

Observation 9d6d39c8-6cdd-40ae-8751-8b9193767366 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation.

Improving Large Vision and Language Models by Learning from a Panel of Peers Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.375958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.375958Z digest=sha256:852302f503899b7e3cf6e9eae7e1b838652ecffb324ba63312f57255847b2857

Observation 82399be2-2ebe-412b-b0e7-3220ae1b890c · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Silkie: Preference Distillation for Large Visual Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.380128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.380128Z digest=sha256:237c10d3e0f5d9600199b73a55751e386f5b2501eecb55e9fc05fba74726e1d7

Observation c33c94a5-eab8-4461-85ea-8c3b8f5d4bc2 · outbound

This paper cites Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.384681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.384681Z digest=sha256:419930246acdf7e1eb40bcbe4abe65b96e9f43ba9f31f50e6ad7d0560e4b14c3

Observation 30ec6ac9-7ef6-4634-bebb-c6747c8d80e3 · outbound

This paper cites Evaluating object hallucination in large vision- language models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Evaluating object hallucination in large vision- language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.995747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.389624Z digest=sha256:50a911590e309b30245fdf080fac25e3d267aadcb4b243022b79e34338646cd7

Observation a7466534-95ef-460a-87be-f8b235512307 · outbound

This paper cites Improved baselines with visual instruction tuning.

Improving Large Vision and Language Models by Learning from a Panel of Peers Improved baselines with visual instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.982356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.393767Z digest=sha256:bc752342c04b329fbcb3cd96eca4efe37cd6f3dfddeee58a3c2585da1b0c2c0d

Observation 37bfb32d-4845-4131-9a94-ce197e6f21f5 · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

Improving Large Vision and Language Models by Learning from a Panel of Peers Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.967890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.397960Z digest=sha256:6accac5bbe0f7cadec91b9b0e86998c4abf5045575ca1fb46b662d359902fd3c

Observation 950d56e3-af88-4e43-8517-b06cfc55a828 · outbound

This paper cites Visual instruction tuning.

Improving Large Vision and Language Models by Learning from a Panel of Peers Visual instruction tuning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.953364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.402375Z digest=sha256:6604e1e4ecb7aa1501538f76af8dadef121f765a0e2cb7cab49ff5b17ad85e0a

Observation 2ed6002d-8d62-4da6-b39e-f97824231234 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Improving Large Vision and Language Models by Learning from a Panel of Peers MMBench: Is Your Multi-modal Model an All-around Player?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.406477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.406477Z digest=sha256:ff2fc4453f426d4603ddb38e8d72961de3c8eb85fa81dbdfe337acc6e76894a4

Observation 6e1762ba-a8fd-48c6-992c-3711ca12ccfd · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.411073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.411073Z digest=sha256:efa7fd406487cbd826c3f9e627cc32216e1a0fc758d40c507640efa4959acb73

Observation d20ad29c-87a5-4429-8f20-0c1e1d7690fe · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Improving Large Vision and Language Models by Learning from a Panel of Peers Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.940018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.415644Z digest=sha256:2c06a5fcbf3510ad3aeda9d67a64f6fc4a5d4057387e486d0a51335a67b43a57

Observation e97ef82c-565c-4adc-a8db-04befdf0f731 · outbound

This paper cites Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.

Improving Large Vision and Language Models by Learning from a Panel of Peers Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.926060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.420121Z digest=sha256:3d550bfa80c70801f5b04c4484ca794e02740f57f0a819868d41d2987b1f3c48

Observation 13178546-7ca2-40d5-bc30-fc30b9c5f40e · outbound

This paper cites WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences.

Improving Large Vision and Language Models by Learning from a Panel of Peers WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.424386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.424386Z digest=sha256:5b0d71d7fdd81dcc01c453026bc7d61946d4addedf6ceac8c34296b33a42e0e1

Observation f5f06566-9406-46e5-a889-f9733b351302 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Improving Large Vision and Language Models by Learning from a Panel of Peers Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.910736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.428752Z digest=sha256:44d636ebfdcde340621980ce8d0130f0392a140d8f8a50480e51b8a6e8c4313a

Observation 2e92be31-7c7f-4e0c-8035-366addc3ef44 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Improving Large Vision and Language Models by Learning from a Panel of Peers SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.432883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.432883Z digest=sha256:77b2f9e263c8eb33e63cd95c2b9c434f138f45660e4ac9729e150216ec5cbcab

Observation b41f54fc-d48d-402c-ba54-270dfcb7dbe9 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Improving Large Vision and Language Models by Learning from a Panel of Peers Ocr-vqa: Visual question answering by reading text in images

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.895584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.437976Z digest=sha256:2e2e5d10f9ef9e1ede49d4c3926404a588924fe5248ee137710205fd75c7f26b

Observation 3f687fb6-424b-45d1-b9a3-966342625502 · outbound

This paper cites Cognitive perspectives on peer learning.

Improving Large Vision and Language Models by Learning from a Panel of Peers Cognitive perspectives on peer learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.881934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.442099Z digest=sha256:94593803b7bfcb7a0344b9337144e4a6530bb3ce57bcfcc66f2bc977aeb8e91a

Observation 0326da7b-9272-4cf7-a834-aee208336164 · outbound

This paper cites Gpt-4o system card, 2024.

Improving Large Vision and Language Models by Learning from a Panel of Peers Gpt-4o system card, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.868024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.446058Z digest=sha256:1995507a47598ee60bc31366dd95424f29f724fbfb71a99fecd6410f5759b532

Observation 597baf8e-52f1-47be-b60e-8c8c75ca7ca0 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Improving Large Vision and Language Models by Learning from a Panel of Peers DINOv2: Learning Robust Visual Features without Supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.450137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.450137Z digest=sha256:c63d0d0f2b4aff234fb5e4b1746dc968122bbcc0af185ae5e7cf6cc6683407d3

Observation 63c5b924-f8ec-461d-a37e-57946919abf0 · outbound

This paper cites Im2text: Describing images using 1 million captioned photographs.

Improving Large Vision and Language Models by Learning from a Panel of Peers Im2text: Describing images using 1 million captioned photographs

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.853943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.454211Z digest=sha256:3e42957ac08f942f9dc1cd24c743dbc4e89df4f731d366caf90345ba98bb4442

Observation 75b0c464-63e7-407c-97fe-aa939aa3cdd9 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744,.

Improving Large Vision and Language Models by Learning from a Panel of Peers Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.458154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.458154Z digest=sha256:f4c794519dfcef50fa7a208f851ea100cf8152eae1ca997f6f34e07a12943a94

Observation 90f58df7-b687-4767-a59c-96d4a4f2d225 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Improving Large Vision and Language Models by Learning from a Panel of Peers Learning transferable visual models from natural language supervi- sion

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.830811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.462330Z digest=sha256:9ea9e2ded0877297d8253cbaa692ae1113fddabbadb1b83bdf72e899317a3719

Observation b85784c7-5972-4bb2-be94-8da74896a9ff · outbound

This paper cites Direct prefer- ence optimization: Your language model is secretly a reward model.

Improving Large Vision and Language Models by Learning from a Panel of Peers Direct prefer- ence optimization: Your language model is secretly a reward model

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.816705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.466445Z digest=sha256:8566aa66e885dd1217f37c4fc03ac9f7ad1fbf2529e5c40473d307f60610df20

Observation c7b418e8-b013-4630-a495-c252d4887b88 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Improving Large Vision and Language Models by Learning from a Panel of Peers LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.470827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.470827Z digest=sha256:06f50b7e48c1afbbdedeb8cca2716a559fd51ec6a4a21dae2181af7343f59dac

Observation 7399ffbf-fc36-446c-ad1e-8773d4ff093d · outbound

This paper cites Laion-5b: An open large-scale dataset for training next gen- eration image-text models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Laion-5b: An open large-scale dataset for training next gen- eration image-text models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.802576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.475190Z digest=sha256:b38b657ace80bb6414faafa6cd2fd4fec13a635ebdc04d07a82e1eb8f0d2ef43

Observation a79e8db8-136f-46f2-9a94-1aeb06a99d5e · outbound

This paper cites 10 Towards vqa models that can read.

Improving Large Vision and Language Models by Learning from a Panel of Peers 10 Towards vqa models that can read

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.788522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.479349Z digest=sha256:c6e7da0171bffdf38faa37b28c310f7aa87caf2d7ab32e717f3a8c3c94c53ff1

Observation 9ab19677-5ba0-4af5-8d41-9ecc26741426 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Improving Large Vision and Language Models by Learning from a Panel of Peers Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.483552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.483552Z digest=sha256:41ec17d4f9d3c17441c08783c0e533dcd91cfa152a344f5366cd40aeefb4da0e

Observation 67920b50-78f7-444c-84e6-d4870286071c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Gemini: A Family of Highly Capable Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.487953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.487953Z digest=sha256:827d5e13651408af5bbf829efd336221624dbadca4c17302066056d4f8a06895

Observation 37673a03-885b-4207-9ee3-6ed10e60fa56 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Improving Large Vision and Language Models by Learning from a Panel of Peers Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.492852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.492852Z digest=sha256:bd5ea8f009b950fdfa786020708af02d7498d76dbaa04efec0b54e38da66ed17

Observation bd83f7df-7bb9-49f0-ba36-e49354fd136e · outbound

This paper cites Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.498054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.498054Z digest=sha256:debb649487b886d833c08f5b26e353707cff8ef7b867cca379895722d947d73b

Observation 97f328e9-a648-4b06-ba86-243fe4c8799c · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Improving Large Vision and Language Models by Learning from a Panel of Peers To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.502470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.502470Z digest=sha256:feaa932ef78e6385b3ec3074daa8c6e4e9abeed86a7cca8f06b8a4762bf5e296

Observation ac0dddc9-b246-4a3f-adfb-aae0a8e5797f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Improving Large Vision and Language Models by Learning from a Panel of Peers Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.618880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.618880Z digest=sha256:423f546326d0fa9ec20ce449c0fa0db5da8036e46eda833e230c38091fefaaba

Observation 13502761-0064-481f-a9ac-e087347d6afa · outbound

This paper cites Self-Taught Evaluators.

Improving Large Vision and Language Models by Learning from a Panel of Peers Self-Taught Evaluators

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.624059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.624059Z digest=sha256:e3df99b26e17625c56a1b5c7fd34c7bf6299b8233ede27886584bfc2f6cc3b1c

Observation fe9fa07f-4118-41da-b60a-83beba9e54c7 · outbound

This paper cites Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement.

Improving Large Vision and Language Models by Learning from a Panel of Peers Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.628450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.628450Z digest=sha256:0a33fc69f7f8cd79fa0fb7394f96e71f913f61c1c94a560e7d32810a6be76563

Observation abf26adc-4c7d-4de0-b06e-55ef50970bdd · outbound

This paper cites HelpSteer2: Open-source dataset for training top-performing reward models.

Improving Large Vision and Language Models by Learning from a Panel of Peers HelpSteer2: Open-source dataset for training top-performing reward models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.632845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.632845Z digest=sha256:6729209ad7f843e886eabbf50dc12ed5cd2cdd74cee2f76b6adca4aed38b48d6

Observation 8cbc05ac-6824-4707-82ad-3e89cd1b6897 · outbound

This paper cites Helpsteer: Multi-attribute helpfulness dataset for steerlm.

Improving Large Vision and Language Models by Learning from a Panel of Peers Helpsteer: Multi-attribute helpfulness dataset for steerlm

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.774047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.637328Z digest=sha256:b1433a6a6fb2f3b80e82e458016ee82fa077b92331c2bf3e6f67eb8a2c0c9334

Observation 940e9c55-ee50-4923-8c90-8447152a6c2a · outbound

This paper cites Pytorch image models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Pytorch image models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.759248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.641868Z digest=sha256:16890f363115bf02c0e2e584025823e5b1ae0c7c40410ebd808ada8ccc527008

Observation bcd8c5eb-a312-4d5d-9586-bbe4bcfb2a80 · outbound

This paper cites Gpt-4v (ision) is a human-aligned evaluator for text-to-3d generation.

Improving Large Vision and Language Models by Learning from a Panel of Peers Gpt-4v (ision) is a human-aligned evaluator for text-to-3d generation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.743120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.646815Z digest=sha256:9228bd118051a2ce44348a5aa27114a360f510162fbbf46e6135a7e5b4dcc3ca

Observation 6cc5a664-6b16-4ddd-9aa2-fbd97503f7b5 · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Improving Large Vision and Language Models by Learning from a Panel of Peers Self-Play Preference Optimization for Language Model Alignment

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.650830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.650830Z digest=sha256:12e608a67996e36ea27171d391bbe08b81d8e7d373e8b70a92c157a87976fcfa

Observation c5fecd07-c731-45a4-85ac-1c5864a6c4b5 · outbound

This paper cites Grok-1.5 vision preview, 2024.

Improving Large Vision and Language Models by Learning from a Panel of Peers Grok-1.5 vision preview, 2024

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.728428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.655183Z digest=sha256:fef96364e5a0a735d36fbd4230f7a3b90d5dcd32f2ef40a1d9a7cbc35400b2a5

Observation 066bc149-5148-4573-8a15-86059f0bc9a7 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.660164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.660164Z digest=sha256:13f28ca66f869cbbe665dae8e68b61e5a67e5696901e66e53c00192bc9ed4c70

Observation baadd476-aa13-4853-a946-7bf01cdf7f50 · outbound

This paper cites The Perfect Blend: Redefining RLHF with Mixture of Judges.

Improving Large Vision and Language Models by Learning from a Panel of Peers The Perfect Blend: Redefining RLHF with Mixture of Judges

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.664545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.664545Z digest=sha256:2dd4b468ee183853c0010875afb965ba2f5dd7324d19755de5d4931d4537c121

Observation fb65bd6c-1ccb-4633-b0cf-d7c324804aec · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

Improving Large Vision and Language Models by Learning from a Panel of Peers xgen-mm (blip-3): A family of open large multimodal models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.668920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.668920Z digest=sha256:9b76d7c174b4846e58dd17464893949cb6a801cfd13639c2e63364658dc34ad9

Observation 83b74cf1-87ec-45fe-90d2-88c05ee447f7 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

Improving Large Vision and Language Models by Learning from a Panel of Peers Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.713157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.673178Z digest=sha256:7103a17229da017a444f5dfd3e6133260052a1f4e7457f838e2b65fe232c104c

Observation f2ef592e-4043-490e-bf7f-5fa2444b0a74 · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.

Improving Large Vision and Language Models by Learning from a Panel of Peers Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.677331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.677331Z digest=sha256:83c47569e6e03cee2b467c6aaa10d0d1e8e0ae47c991db04825b0a3ef03225bd

Observation 42ec46ea-b362-462d-9aa1-88f10452c4f3 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities.

Improving Large Vision and Language Models by Learning from a Panel of Peers Mm-vet: Evaluating large multimodal models for integrated capabilities

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.698235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.681561Z digest=sha256:5bc1ea0b22b977d05608b1d928a71f2b95f9f53c1b976576c8d11c3d2bfae4b7

Observation 8180939e-c5b8-47c3-9251-d259117020b4 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

Improving Large Vision and Language Models by Learning from a Panel of Peers Florence: A New Foundation Model for Computer Vision

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.685759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.685759Z digest=sha256:8f11e75c804b9503ddc1eff146d587744e73f25c51ca5719968bb8c490736a45

Observation 46ff7d88-df70-4e82-ba8b-5a1480e5df3a · outbound

This paper cites Self-Rewarding Language Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Self-Rewarding Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.690602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.690602Z digest=sha256:a8514f06384b5574f7edd86baf68bc61affb771720e07a02b40465d583004dd9

Observation c205a9dc-0742-4e1d-bede-ceea4f8e680e · outbound

This paper cites Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi.

Improving Large Vision and Language Models by Learning from a Panel of Peers Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.683747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.695166Z digest=sha256:62a904df38b8bd8f0cf71184f3c627e4f4f3cb7967cef8b8940c9262c43784d8

Observation 03f86c78-4642-404b-a7cd-beae779d7761 · outbound

This paper cites Sigmoid loss for language image pre-training.

Improving Large Vision and Language Models by Learning from a Panel of Peers Sigmoid loss for language image pre-training

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.699358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.699358Z digest=sha256:0ea8f77a43bfeed20f3df4e8319bd5febb345df942030dbe3cce89606182fc9e

Observation bdb9d7e0-39e1-48e3-879a-fda3f07ec52d · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

Improving Large Vision and Language Models by Learning from a Panel of Peers SVIT: Scaling up Visual Instruction Tuning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.703738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.703738Z digest=sha256:1abe695f7c5b16ac5c5a053f4a216006ca2843a7db1abbec8152aaaca87dbcb4

Observation b5d7a50d-8486-46ac-bcce-8160e888fd94 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Improving Large Vision and Language Models by Learning from a Panel of Peers PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.708238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.708238Z digest=sha256:6d1b35a36cf6aeaa33b95262d859d933e1696db34c819e198ce769feb67a2ac4

Observation 8b2d4026-ad5a-41cd-bf59-bca8ef8d7882 · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

Improving Large Vision and Language Models by Learning from a Panel of Peers Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.712961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.712961Z digest=sha256:b68efd1c9cc986ff2426efe0a7275295b2fc6fefc0cc44adc71ed7227681b8c0

Observation c5bb8057-b04a-42c5-bca4-1400387f4776 · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Calibrated Self-Rewarding Vision Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.717430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.717430Z digest=sha256:f7349794c5dad1c65bf837d8b56af1faf23bcd80c0fdb21ae07d6d47658e0ce1

Observation 2c1e9cbc-fa00-4d8b-821b-9d59a657ad41 · outbound

This paper cites Self-Supervised Visual Preference Alignment.

Improving Large Vision and Language Models by Learning from a Panel of Peers Self-Supervised Visual Preference Alignment

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.721849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.721849Z digest=sha256:37e8f859dda4e86bc14576f559c77c6eeb0cd86ebd5963d33abd6703978ec4ea

Observation b463cf89-7f55-479e-8d87-a6594ee11399 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Improving Large Vision and Language Models by Learning from a Panel of Peers Fine-Tuning Language Models from Human Preferences

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.726448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.726448Z digest=sha256:472b8186a278b3c3d57cddbe135c003ff6312ca0b5f1fa0150a8c10e09867121

Observation a5740f99-3e2c-4ca3-bb51-b46176fd490a · outbound

This paper cites an unresolved cited work.

Improving Large Vision and Language Models by Learning from a Panel of Peers Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:27:28.659736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.731433Z digest=sha256:3cb4ff81b75089af958afe66c60e76e70f6b05667e0e5e5f12a53828d530946a

Observation acdac01d-a280-4cea-a386-e0382703e33a · outbound

This paper cites an unresolved cited work.

Improving Large Vision and Language Models by Learning from a Panel of Peers Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:27:28.644530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.735735Z digest=sha256:a250cc596e806679ae2327427827d61a6bf5086aa3bd8e6140e75096a4a3ff25

Observation d2caee6f-78d4-4a12-b36b-5cee87651a6c · outbound

This paper cites an unresolved cited work.

Improving Large Vision and Language Models by Learning from a Panel of Peers Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:27:28.629984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.739945Z digest=sha256:9e0dc2f2bccdaf428f28f83fd32a5e747c28174367d5ddf99c0437f5ca672832

Observation 5c4fa842-9a44-4f81-b912-1ab012c1b2fc · outbound

This paper cites an unresolved cited work.

Improving Large Vision and Language Models by Learning from a Panel of Peers Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:27:28.615834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.744158Z digest=sha256:3bf085a330382054352ee490973175e70b118a0605ba88b029e36f4d9a32c2f2

Observation cc9c01e3-a089-4554-bcb7-abdc2d80a1c7 · outbound

This paper cites • Rating Guidelines: Models receive detailed explanations for scoring each dimension.

Improving Large Vision and Language Models by Learning from a Panel of Peers • Rating Guidelines: Models receive detailed explanations for scoring each dimension

Reference 87

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T12:27:28.601990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T12:27:27.748564Z digest=sha256:8be94e9927974ddc18d529e2fbc074d903f38d7105986fe364f82f3ba780caab

Pith citing papers

No inbound Pith citation observations are available.