Pith. sign in

Paper Citation Record · LEDGER

Improving Large Vision and Language Models by Learning from a Panel of Peers

As of 17 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2509.01610.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01610 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:27:27.748564Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved62
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1f574b3-4004-49c8-8e3c-55cab892c658 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Improving Large Vision and Language Models by Learning from a Panel of Peers Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.253037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.253037Z digest=sha256:73f0ce75cdd0db58d9274402ed985370679db7bbde6d55628589ba29732fdd82

Observation 4c935171-7f80-46b8-8f73-3b7aace4c50b · outbound

This paper cites GPT-4 Technical Report.

Improving Large Vision and Language Models by Learning from a Panel of Peers GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.258508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.258508Z digest=sha256:28ca0297711c8c8db7cd661c3bcfd1a42efbfcea02b809efbacf2fffd20c70bd

Observation 9fd19040-31cd-4fd4-a9e3-c43fd6659dd0 · outbound

This paper cites The Llama 3 Herd of Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.263431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.263431Z digest=sha256:0e91aee4d629b62a74bb81f905adc382f8e5d7bd9e0898ee3d75d84b3dbc3996

Observation 416f724d-3c68-4ef6-a538-9b84233ed729 · outbound

This paper cites Theoretical guarantees on the best-of-n alignment policy.

Improving Large Vision and Language Models by Learning from a Panel of Peers Theoretical guarantees on the best-of-n alignment policy

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.267919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.267919Z digest=sha256:bb9c793af568fe01704199c595ded281fdec2a62703b9040e9a12a67dba17572

Observation 320ed74c-f5d9-420b-9c91-687bf8a4bf98 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Improving Large Vision and Language Models by Learning from a Panel of Peers PaliGemma: A versatile 3B VLM for transfer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.272581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.272581Z digest=sha256:cb154dfe0b8cea534aade0ac6ac76c06ec234231176d004705208efbbe1e159d

Observation 06e2aed1-d8a9-4e36-ab9f-129346445690 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Improving Large Vision and Language Models by Learning from a Panel of Peers Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.277331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.277331Z digest=sha256:255196d3aa13d437a7abd2b027c42871b5a7e3930d94b2f4d0718a46aca8c56c

Observation f1af862d-6277-4e88-9be9-a3154e63f00d · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.281965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.281965Z digest=sha256:8308d6bff5486871a53d07dac4a23091ae9dc7c867ada92c37dc4376535260c3

Observation 28d8a27c-bb4b-46e4-aaa9-8f61d7bf5d73 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Improving Large Vision and Language Models by Learning from a Panel of Peers How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.286826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.286826Z digest=sha256:2559373934e9e98031aa9e0cd392c31d29740a2bfd0408cea86dedd5f068599b

Observation d9b9f837-3e73-4667-bc9c-7e7ea26dece4 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Improving Large Vision and Language Models by Learning from a Panel of Peers Gonzalez, Ion Stoica, and Eric P

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.291111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.291111Z digest=sha256:c1972f8d5927fb220102d969632a631c5888f4345114e725d1e63c5dc93e104d

Observation cf6b9be5-c95d-4ac1-9d4a-fa4e70c5cf4f · outbound

This paper cites Self-Improving Robust Preference Optimization.

Improving Large Vision and Language Models by Learning from a Panel of Peers Self-Improving Robust Preference Optimization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:27:28.478077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.295655Z digest=sha256:d51adef32d34df764bcc5805d6390011eb38f2287db3e039c2155a46a6729d61

Observation 7849c7cf-8ac5-4f82-b7c4-a15dd2047b5f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Improving Large Vision and Language Models by Learning from a Panel of Peers Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.300251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.300251Z digest=sha256:f42be830ff36678a6b9b4537a8fef098e1c6a9256e744094960f1b68e2f120b3

Observation 25de84a8-6ce4-4abb-94c1-6dcbf63da87a · outbound

This paper cites Reward model ensembles help mitigate overop- timization.

Improving Large Vision and Language Models by Learning from a Panel of Peers Reward model ensembles help mitigate overop- timization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.304752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.304752Z digest=sha256:878356e1409fd31231428203134691088e71271b84526f82d75ce8a32b8d403a

Observation f2b6f9bb-88c2-47ae-98a6-2035402afea5 · outbound

This paper cites Nvlm: Open frontier-class multimodal llms.

Improving Large Vision and Language Models by Learning from a Panel of Peers Nvlm: Open frontier-class multimodal llms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.309716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.309716Z digest=sha256:624c23ecd7e0efbbe807669ed1b7f8cc5225ede8768a57dfae393a4e7b9c4494

Observation 0103621a-185e-4f2d-b475-f7f8af9f2323 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Improving Large Vision and Language Models by Learning from a Panel of Peers Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.314082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.314082Z digest=sha256:692c1dff1f17675b0342787d4add1c71cf8b9a894d7d617e59a76e2ea40aa67e

Observation 90da2d0b-414b-43cc-878b-39ed1d204b76 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.318382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.318382Z digest=sha256:2f351079d33e8ef86783bfd42b444cd8d5afb29357996f634a1e142d1ccdb23b

Observation f011d6d3-4fcb-48eb-8931-1cdf418b6bb7 · outbound

This paper cites Enhancing Large Vision Language Models with Self-Training on Image Comprehension.

Improving Large Vision and Language Models by Learning from a Panel of Peers Enhancing Large Vision Language Models with Self-Training on Image Comprehension

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.323258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.323258Z digest=sha256:c8f3c9afc77fd75f50724096c3d98edfa5f87acc11043b22bb9d6a07511ffd40

Observation 66190a49-3bc1-4536-b81d-6a3e48288695 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Improving Large Vision and Language Models by Learning from a Panel of Peers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.327711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.327711Z digest=sha256:d65c0ff26a8685b9f3ec672698b6d506a818344ef0d7d967cc1cf17c2747ab18

Observation 42bc6ae6-ce22-4d6e-9970-e4b71f5ef03c · outbound

This paper cites VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.332321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.332321Z digest=sha256:e392594f4e4a0e739af4e5c77dd7c1a0c462f3bbc6c4b9eab8522ed45297e201

Observation b32c752a-098d-4cc0-a4ed-b31bd5e24d6c · outbound

This paper cites VILA$^2$: VILA Augmented VILA.

Improving Large Vision and Language Models by Learning from a Panel of Peers VILA$^2$: VILA Augmented VILA

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.336733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.336733Z digest=sha256:fabc4b099872751d94206e41c9637419d147be1341648494f56d9cb24da1f1cc

Observation 02edb8e5-c7e7-422a-8a61-a909e22db5cd · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 20

Resolution
malformed identifier
no resolver link, observed 2026-08-05T12:27:27.341177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.341177Z digest=sha256:f09d5976993490f9bb7469df4f22e761c31deac75acd50b995be1987acb2d7dc

Observation e7d8a105-9d58-482a-a151-9c16b8c07adb · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Improving Large Vision and Language Models by Learning from a Panel of Peers CogVLM2: Visual Language Models for Image and Video Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.345868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.345868Z digest=sha256:ab15f003714acdb8b1200c9ef4cff1b3d67ea6783bc96c2bda5c555f7d36a23e

Observation 34a2a7b2-1722-49e8-82e6-09870a987a03 · outbound

This paper cites Mistral 7B.

Improving Large Vision and Language Models by Learning from a Panel of Peers Mistral 7B

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.350403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.350403Z digest=sha256:786cf7ee9f3ecd48bab1e17dcd715c581dbf57f2ba383fbecac758ffa8ec37b3

Observation 9c3c92b1-5766-418b-8d6c-ac768568c5ef · outbound

This paper cites A diagram is worth a dozen images.

Improving Large Vision and Language Models by Learning from a Panel of Peers A diagram is worth a dozen images

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.354977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.354977Z digest=sha256:41da22ffcb94f570ae978b307d29055266907a00cecb945fe2e196abb695f75d

Observation acd8ce11-70b5-4725-8770-f498c30bbdb2 · outbound

This paper cites What matters when building vision-language models?.

Improving Large Vision and Language Models by Learning from a Panel of Peers What matters when building vision-language models?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.359046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.359046Z digest=sha256:3d39fb3bc3e48e2e0b8c041f635962fe7d6f63c5e3f7b87875ffdb9c324ffb0c

Observation 3662a7e8-8af8-48b9-8b10-acc81c44ad5d · outbound

This paper cites Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation.

Improving Large Vision and Language Models by Learning from a Panel of Peers Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.363208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.363208Z digest=sha256:d6b87ed1d6ca78497bfd6240e5e25c260f0d662ff3086c3c217573891ebd62e7

Observation 07985a7d-1f8e-4115-98be-191d10668a3f · outbound

This paper cites Seed-bench: Benchmark- ing multimodal large language models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Seed-bench: Benchmark- ing multimodal large language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.367800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.367800Z digest=sha256:b9ab35ba36900bb8bef874b14256e6e2e4a6b69f631be19bd38de573ac62c1f1

Observation dba96a56-05ab-4b2b-b413-3f53fbd6fc1d · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Improving Large Vision and Language Models by Learning from a Panel of Peers LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.371859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.371859Z digest=sha256:8cc55ea43e4c54e2be2099bfcdbefdfa44cfbda6634ee3361cdc5d79c941e52e

Observation 9d6d39c8-6cdd-40ae-8751-8b9193767366 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation.

Improving Large Vision and Language Models by Learning from a Panel of Peers Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.375958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.375958Z digest=sha256:2af1d1f371d3623f56a4d5727fb5b9a627d9ec18771a9210cfc43608e3740da4

Observation 82399be2-2ebe-412b-b0e7-3220ae1b890c · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Silkie: Preference Distillation for Large Visual Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.380128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.380128Z digest=sha256:f36afe98d648417a48733a46a2adb43f5983ba2ba023f2b963ee847241bf1693

Observation c33c94a5-eab8-4461-85ea-8c3b8f5d4bc2 · outbound

This paper cites Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.384681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.384681Z digest=sha256:f891df35680205fe71269ac00fcad0e793d002d11cc33a0894f0d6c6c61e9817

Observation 30ec6ac9-7ef6-4634-bebb-c6747c8d80e3 · outbound

This paper cites Evaluating object hallucination in large vision- language models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Evaluating object hallucination in large vision- language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.995747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.389624Z digest=sha256:a434cf3db4bb00b39cc3b0adfff46d4d76a8d9af85a54cda9423bcd450442ad7

Observation a7466534-95ef-460a-87be-f8b235512307 · outbound

This paper cites Improved baselines with visual instruction tuning.

Improving Large Vision and Language Models by Learning from a Panel of Peers Improved baselines with visual instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.982356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.393767Z digest=sha256:1a42baff11c84493579c4c45d52d126c28e9e211ffaf912891f6d94b74e5bc7d

Observation 37bfb32d-4845-4131-9a94-ce197e6f21f5 · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

Improving Large Vision and Language Models by Learning from a Panel of Peers Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.967890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.397960Z digest=sha256:e383198e48ed9b13cf25493e603fa1c30e4151e40a72b0d64bda9496bbb38a04

Observation 950d56e3-af88-4e43-8517-b06cfc55a828 · outbound

This paper cites Visual instruction tuning.

Improving Large Vision and Language Models by Learning from a Panel of Peers Visual instruction tuning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.953364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.402375Z digest=sha256:e8dafac7340036cc138c05411daf3182169dec3c198cae418db8b0f6f4346a50

Observation 2ed6002d-8d62-4da6-b39e-f97824231234 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Improving Large Vision and Language Models by Learning from a Panel of Peers MMBench: Is Your Multi-modal Model an All-around Player?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.406477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.406477Z digest=sha256:d1be43dd1bc7127619c02db2e52c713276d9c53478b66581263c293665b03657

Observation 6e1762ba-a8fd-48c6-992c-3711ca12ccfd · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.411073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.411073Z digest=sha256:aa5c4affbb0f22450ff08f20291df7fe0182e2e6c180ed579a003b2854caebbc

Observation d20ad29c-87a5-4429-8f20-0c1e1d7690fe · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Improving Large Vision and Language Models by Learning from a Panel of Peers Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.940018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.415644Z digest=sha256:bbb2e6bdc3b06b2114e6598979f1f64efa72ab23f6ed3da9c1eb3c56b4b760c9

Observation e97ef82c-565c-4adc-a8db-04befdf0f731 · outbound

This paper cites Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.

Improving Large Vision and Language Models by Learning from a Panel of Peers Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.926060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.420121Z digest=sha256:5fd9dd70dca00cabd1367f7ad8a9174a97151c2d1feab95068d98d168fba8e84

Observation 13178546-7ca2-40d5-bc30-fc30b9c5f40e · outbound

This paper cites WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences.

Improving Large Vision and Language Models by Learning from a Panel of Peers WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.424386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.424386Z digest=sha256:adc93231f99ff27cbc49bfefd437b3d55e245c1a2bb5f0e17613b5ebbf827ec5

Observation f5f06566-9406-46e5-a889-f9733b351302 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Improving Large Vision and Language Models by Learning from a Panel of Peers Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.910736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.428752Z digest=sha256:d769ed14e9269762f2c8731ad15eb5cf7926f404c0650c403828107dc74d7428

Observation 2e92be31-7c7f-4e0c-8035-366addc3ef44 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Improving Large Vision and Language Models by Learning from a Panel of Peers SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.432883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.432883Z digest=sha256:a81360c6a91be4efaf0fea9be6dc9892f527041aa766a40233c4bee5c0453064

Observation b41f54fc-d48d-402c-ba54-270dfcb7dbe9 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Improving Large Vision and Language Models by Learning from a Panel of Peers Ocr-vqa: Visual question answering by reading text in images

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.895584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.437976Z digest=sha256:505e8edd3fc907c8616b613ea4a36e13e3d6bd313ba7affec4be8d1a82258d05

Observation 3f687fb6-424b-45d1-b9a3-966342625502 · outbound

This paper cites Cognitive perspectives on peer learning.

Improving Large Vision and Language Models by Learning from a Panel of Peers Cognitive perspectives on peer learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.881934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.442099Z digest=sha256:abbbdc2c010c529e407880dcf431907b5970058835ffddac4a1f20dc39c481a3

Observation 0326da7b-9272-4cf7-a834-aee208336164 · outbound

This paper cites Gpt-4o system card, 2024.

Improving Large Vision and Language Models by Learning from a Panel of Peers Gpt-4o system card, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.868024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.446058Z digest=sha256:7de560e908cdd29f26999361e755bf569ba0ed7e721980ebddc771a07e3e794f

Observation 597baf8e-52f1-47be-b60e-8c8c75ca7ca0 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Improving Large Vision and Language Models by Learning from a Panel of Peers DINOv2: Learning Robust Visual Features without Supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.450137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.450137Z digest=sha256:10f0d9be1e4a24908c5469b8445ea4dc7a51e30166dc2b7d85e04759cdd26c95

Observation 63c5b924-f8ec-461d-a37e-57946919abf0 · outbound

This paper cites Im2text: Describing images using 1 million captioned photographs.

Improving Large Vision and Language Models by Learning from a Panel of Peers Im2text: Describing images using 1 million captioned photographs

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.853943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.454211Z digest=sha256:5a44c12de0941311c9e2a3538c74a51066639d4288116bbbfebaa88d6a28dc3c

Observation 75b0c464-63e7-407c-97fe-aa939aa3cdd9 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744,.

Improving Large Vision and Language Models by Learning from a Panel of Peers Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.458154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.458154Z digest=sha256:972aa0ad9df1ea66a5e9919cd1b9c68bf64935ede6a2bbf522e199076a348e9f

Observation 90f58df7-b687-4767-a59c-96d4a4f2d225 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Improving Large Vision and Language Models by Learning from a Panel of Peers Learning transferable visual models from natural language supervi- sion

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.830811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.462330Z digest=sha256:d4cbf7ed7b1f0af5b0ae26410b41f16ac213410ff62fc5800e4847e183fd927f

Observation b85784c7-5972-4bb2-be94-8da74896a9ff · outbound

This paper cites Direct prefer- ence optimization: Your language model is secretly a reward model.

Improving Large Vision and Language Models by Learning from a Panel of Peers Direct prefer- ence optimization: Your language model is secretly a reward model

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.816705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.466445Z digest=sha256:e4b04551cf1e7b6b4c928ed3fab9eb00a3b56b5e69a35aff6f9eda235678ed4f

Observation c7b418e8-b013-4630-a495-c252d4887b88 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Improving Large Vision and Language Models by Learning from a Panel of Peers LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.470827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.470827Z digest=sha256:809c9fea254b44e661048edac59f4a58092e42ed7cca0b1f26d3a8088134360e

Observation 7399ffbf-fc36-446c-ad1e-8773d4ff093d · outbound

This paper cites Laion-5b: An open large-scale dataset for training next gen- eration image-text models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Laion-5b: An open large-scale dataset for training next gen- eration image-text models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.802576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.475190Z digest=sha256:269e996350239762113f23a39c0578c76bb26841e4ad4cc43ede7fd765f4bf75

Observation a79e8db8-136f-46f2-9a94-1aeb06a99d5e · outbound

This paper cites 10 Towards vqa models that can read.

Improving Large Vision and Language Models by Learning from a Panel of Peers 10 Towards vqa models that can read

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.788522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.479349Z digest=sha256:b35762afe58ef369e5d46e78f0d481e7f6bd3bd1da7f2c7d5f592b062e02373e

Observation 9ab19677-5ba0-4af5-8d41-9ecc26741426 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Improving Large Vision and Language Models by Learning from a Panel of Peers Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.483552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.483552Z digest=sha256:f786129187f927b3bda535ce95172332f0b5e66edebb3c7854dc1cb93b672036

Observation 67920b50-78f7-444c-84e6-d4870286071c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Gemini: A Family of Highly Capable Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.487953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.487953Z digest=sha256:75bddd88a418bd33b2d858b421756462d0610e9e0ded7325e8d5ab63657ba694

Observation 37673a03-885b-4207-9ee3-6ed10e60fa56 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Improving Large Vision and Language Models by Learning from a Panel of Peers Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.492852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.492852Z digest=sha256:200cdf0af871b71ddb354fd4e07f28e29119a6aaa9efa445aff01c9445f484c9

Observation bd83f7df-7bb9-49f0-ba36-e49354fd136e · outbound

This paper cites Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.498054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.498054Z digest=sha256:c1a84a9db33f0ce43db51ae408afe74695caf16bbdcc8db6dfb9e71254870192

Observation 97f328e9-a648-4b06-ba86-243fe4c8799c · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Improving Large Vision and Language Models by Learning from a Panel of Peers To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.502470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.502470Z digest=sha256:53df938017ff258242ec6161c53b6b0a2bb682778876b14d5b78ce4ba92d4e3e

Observation ac0dddc9-b246-4a3f-adfb-aae0a8e5797f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Improving Large Vision and Language Models by Learning from a Panel of Peers Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.618880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.618880Z digest=sha256:43a2ac6287f229c85b10625b4e545889400ba6c8fe33865f7ef49a4d5387c6ea

Observation 13502761-0064-481f-a9ac-e087347d6afa · outbound

This paper cites Self-Taught Evaluators.

Improving Large Vision and Language Models by Learning from a Panel of Peers Self-Taught Evaluators

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.624059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.624059Z digest=sha256:48cb993f088501fcf57965ae12348321eda4b0c2f72e237107ed0480dffb42ad

Observation fe9fa07f-4118-41da-b60a-83beba9e54c7 · outbound

This paper cites Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement.

Improving Large Vision and Language Models by Learning from a Panel of Peers Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.628450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.628450Z digest=sha256:5d2f1a74b41dcc3e091045b1eafd12c175fcc1a62b0bb55af51cdf3e63b1f515

Observation abf26adc-4c7d-4de0-b06e-55ef50970bdd · outbound

This paper cites HelpSteer2: Open-source dataset for training top-performing reward models.

Improving Large Vision and Language Models by Learning from a Panel of Peers HelpSteer2: Open-source dataset for training top-performing reward models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.632845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.632845Z digest=sha256:f327a79ef0953d2a2f93f9e1c0f9c6f397c5cf502ce2f7db3e7fffd54efcf888

Observation 8cbc05ac-6824-4707-82ad-3e89cd1b6897 · outbound

This paper cites Helpsteer: Multi-attribute helpfulness dataset for steerlm.

Improving Large Vision and Language Models by Learning from a Panel of Peers Helpsteer: Multi-attribute helpfulness dataset for steerlm

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.774047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.637328Z digest=sha256:3886e13da9664ced373664ced5c2468c384a2c590ae0b0d8f1bd42b59de23165

Observation 940e9c55-ee50-4923-8c90-8447152a6c2a · outbound

This paper cites Pytorch image models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Pytorch image models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.759248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.641868Z digest=sha256:f66091b4a0a1ab25643bbf7b7ac5d45608623b7f8451018d45be78134bac1e98

Observation bcd8c5eb-a312-4d5d-9586-bbe4bcfb2a80 · outbound

This paper cites Gpt-4v (ision) is a human-aligned evaluator for text-to-3d generation.

Improving Large Vision and Language Models by Learning from a Panel of Peers Gpt-4v (ision) is a human-aligned evaluator for text-to-3d generation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.743120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.646815Z digest=sha256:f7e9e50bcd5283491f2370c797ac26dfeee782f98e858a18573e460c1934f7af

Observation 6cc5a664-6b16-4ddd-9aa2-fbd97503f7b5 · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Improving Large Vision and Language Models by Learning from a Panel of Peers Self-Play Preference Optimization for Language Model Alignment

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.650830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.650830Z digest=sha256:fdf7629b0872ac296e1aa9c2e1887ef2ee1ce822ef1304c8debed27dbe1712d4

Observation c5fecd07-c731-45a4-85ac-1c5864a6c4b5 · outbound

This paper cites Grok-1.5 vision preview, 2024.

Improving Large Vision and Language Models by Learning from a Panel of Peers Grok-1.5 vision preview, 2024

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.728428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.655183Z digest=sha256:0400f3bf48ef81627113a998528e5ef2ca27f2e066114bab1e6e4b4f3cb05c14

Observation 066bc149-5148-4573-8a15-86059f0bc9a7 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.660164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.660164Z digest=sha256:f55e8a9d680b5b1fe5be7e27073d4b20b754f59b9d676fcb9e455a8ac5fd01d8

Observation baadd476-aa13-4853-a946-7bf01cdf7f50 · outbound

This paper cites The Perfect Blend: Redefining RLHF with Mixture of Judges.

Improving Large Vision and Language Models by Learning from a Panel of Peers The Perfect Blend: Redefining RLHF with Mixture of Judges

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.664545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.664545Z digest=sha256:08f96cd2ca60908c4854ed98866e92ce97bb63842327d1731826cc91a98967a2

Observation fb65bd6c-1ccb-4633-b0cf-d7c324804aec · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

Improving Large Vision and Language Models by Learning from a Panel of Peers xgen-mm (blip-3): A family of open large multimodal models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.668920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.668920Z digest=sha256:3d36877dacf67ecacb007dc2808f9feab077dff44730dd6b8ec44f691114a272

Observation 83b74cf1-87ec-45fe-90d2-88c05ee447f7 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

Improving Large Vision and Language Models by Learning from a Panel of Peers Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.713157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.673178Z digest=sha256:41095c29e7d94f33cf0feafda2ad640d5176fc24d3b781406a7b2e813e8d1108

Observation f2ef592e-4043-490e-bf7f-5fa2444b0a74 · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.

Improving Large Vision and Language Models by Learning from a Panel of Peers Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.677331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.677331Z digest=sha256:0b6f7b5fe53332bca48ad2f213da14f00de656b78420d3428672f6d7bdedf85f

Observation 42ec46ea-b362-462d-9aa1-88f10452c4f3 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities.

Improving Large Vision and Language Models by Learning from a Panel of Peers Mm-vet: Evaluating large multimodal models for integrated capabilities

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.698235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.681561Z digest=sha256:c8b009e08190c9e877421502f4cf5d823570168462b1e9d1a3311a9b3c8b1626

Observation 8180939e-c5b8-47c3-9251-d259117020b4 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

Improving Large Vision and Language Models by Learning from a Panel of Peers Florence: A New Foundation Model for Computer Vision

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.685759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.685759Z digest=sha256:9b7bf345bf119bb100a2bb4a129c92f4ecdc8a9929509ca69d3fefd0513386bf

Observation 46ff7d88-df70-4e82-ba8b-5a1480e5df3a · outbound

This paper cites Self-Rewarding Language Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Self-Rewarding Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.690602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.690602Z digest=sha256:d8cde842293220da454a2764305b5b0ff0d4c0bcc24abeb6aab0266660328ff6

Observation c205a9dc-0742-4e1d-bede-ceea4f8e680e · outbound

This paper cites Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi.

Improving Large Vision and Language Models by Learning from a Panel of Peers Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:27:28.683747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.695166Z digest=sha256:74953e38901187a04d88c3ff0d74aa662979c52e17aa7951646752c8c87edf90

Observation 03f86c78-4642-404b-a7cd-beae779d7761 · outbound

This paper cites Sigmoid loss for language image pre-training.

Improving Large Vision and Language Models by Learning from a Panel of Peers Sigmoid loss for language image pre-training

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.699358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.699358Z digest=sha256:e4e2a26481d7d3149469ed7ca222d1c171bcdda2f9358cbd889bc2523c8620a6

Observation bdb9d7e0-39e1-48e3-879a-fda3f07ec52d · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

Improving Large Vision and Language Models by Learning from a Panel of Peers SVIT: Scaling up Visual Instruction Tuning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.703738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.703738Z digest=sha256:6cbc5329b97f727a99808cf3eafbe36e1076bceb8b253ac8fad5079cb703da0f

Observation b5d7a50d-8486-46ac-bcce-8160e888fd94 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Improving Large Vision and Language Models by Learning from a Panel of Peers PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.708238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.708238Z digest=sha256:5bac2dcc37828bb27a92f5bc276d6cf5c7cf8a98f4b20ee70387c0cf97897b3e

Observation 8b2d4026-ad5a-41cd-bf59-bca8ef8d7882 · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

Improving Large Vision and Language Models by Learning from a Panel of Peers Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.712961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.712961Z digest=sha256:d663297d38ef51c970c8a5a1c4518421bcfdd09a6119a84d16fc56aed3607eef

Observation c5bb8057-b04a-42c5-bca4-1400387f4776 · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

Improving Large Vision and Language Models by Learning from a Panel of Peers Calibrated Self-Rewarding Vision Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.717430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.717430Z digest=sha256:dd7aeec4e8af9fce78f7c410313c19aa0ed8dd219e1ca89fe988d89a934ce90b

Observation 2c1e9cbc-fa00-4d8b-821b-9d59a657ad41 · outbound

This paper cites Self-Supervised Visual Preference Alignment.

Improving Large Vision and Language Models by Learning from a Panel of Peers Self-Supervised Visual Preference Alignment

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.721849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.721849Z digest=sha256:08a634e09d35e62b523daa3464ff33bb9f902f80e07c5aeeb603a1af569d18ac

Observation b463cf89-7f55-479e-8d87-a6594ee11399 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Improving Large Vision and Language Models by Learning from a Panel of Peers Fine-Tuning Language Models from Human Preferences

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.726448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.726448Z digest=sha256:c7302946bca8424165b14f8d8527bf49e6c5970044f32feee562336021f0e9f6

Observation a5740f99-3e2c-4ca3-bb51-b46176fd490a · outbound

This paper cites an unresolved cited work.

Improving Large Vision and Language Models by Learning from a Panel of Peers Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:27:28.659736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.731433Z digest=sha256:08a4a5d9ceced958bef0d6efb7f70d862d592c936b4c6c1dc09b8abc7c697e33

Observation acdac01d-a280-4cea-a386-e0382703e33a · outbound

This paper cites an unresolved cited work.

Improving Large Vision and Language Models by Learning from a Panel of Peers Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:27:28.644530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.735735Z digest=sha256:1b9400aed009168b375904d8211b28c02d702476d12fff2c8f473785f3932f72

Observation d2caee6f-78d4-4a12-b36b-5cee87651a6c · outbound

This paper cites an unresolved cited work.

Improving Large Vision and Language Models by Learning from a Panel of Peers Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:27:28.629984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.739945Z digest=sha256:9aa89124ef7188bd4fe84d738898d60c88756f89f2fb1a27278866c184d42911

Observation 5c4fa842-9a44-4f81-b912-1ab012c1b2fc · outbound

This paper cites an unresolved cited work.

Improving Large Vision and Language Models by Learning from a Panel of Peers Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:27:28.615834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.744158Z digest=sha256:d73bc8e4050cb020effa0cd0841e81791e41f32a2501aa50269c05e039e90423

Observation cc9c01e3-a089-4554-bcb7-abdc2d80a1c7 · outbound

This paper cites • Rating Guidelines: Models receive detailed explanations for scoring each dimension.

Improving Large Vision and Language Models by Learning from a Panel of Peers • Rating Guidelines: Models receive detailed explanations for scoring each dimension

Reference 87

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T12:27:28.601990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T12:27:27.748564Z digest=sha256:31ef7c4a93dc7cb020e35acd32782c6cc07c65079286efc65fce573eb2bcf18a

Pith citing papers

No inbound Pith citation observations are available.