Pith. sign in

Paper Citation Record · LEDGER

Continual SFT Matches Multimodal RLHF with Negative Supervision

As of 13 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2411.14797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14797 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:58:43.424797Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 216260a6-88c1-4992-b5d7-7eb67303aaa8 · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond.

Continual SFT Matches Multimodal RLHF with Negative Supervision Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.874415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:42.711922Z digest=sha256:86ae923b835e978ee8eaeaf0d4c39d8d03d04a72622d2b812c17485f9c059573

Observation dd7b815f-e7c0-4a0d-b68a-e89f0ae32aad · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Continual SFT Matches Multimodal RLHF with Negative Supervision ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.716829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.716829Z digest=sha256:624274f248fff3da50d0982ad414b7a04f2b8d5a1855b31a5adba90b5f6f5ec0

Observation 3228bcd9-5962-426c-a75d-b0cfe8b77cfc · outbound

This paper cites Self-play fine-tuning converts weak language models to strong language models.

Continual SFT Matches Multimodal RLHF with Negative Supervision Self-play fine-tuning converts weak language models to strong language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.796661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:42.721830Z digest=sha256:5625931779df4a476038a56a833074a356e060a93404bc0732ea716eecd02766

Observation b4e405fd-7802-465a-b135-fa84d92fe807 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Continual SFT Matches Multimodal RLHF with Negative Supervision How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.726082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.726082Z digest=sha256:01ac78f596e4ffd0cc0d98c35e36ce01f6fb0a35ad5d07f2711cd6afb55ddf3d

Observation d7adcc2f-a2a2-4e3f-be14-bc214a52c496 · outbound

This paper cites Scaling instruction- finetuned language models.

Continual SFT Matches Multimodal RLHF with Negative Supervision Scaling instruction- finetuned language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.731110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.731110Z digest=sha256:774684585f9f5d8b75a560af36fd232c581a149cef98bddf14796dcd84ad3ab1

Observation d7bd5977-be98-4522-9546-98d7487d143d · outbound

This paper cites DreamLLM: Synergistic multimodal com- prehension and creation.

Continual SFT Matches Multimodal RLHF with Negative Supervision DreamLLM: Synergistic multimodal com- prehension and creation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.676044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:42.735944Z digest=sha256:a3c0cb1c6d0f2dbc5a51db0467272b23e72ce3e33e189aafeb9babea7b073548

Observation 6fb41568-fb5b-43c1-93ad-7c7ab2d28a3e · outbound

This paper cites Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective.

Continual SFT Matches Multimodal RLHF with Negative Supervision Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.740552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.740552Z digest=sha256:2a13bae031cc4910a346ccbd58554e91652967fc909633daf65c9f52e9c4da22

Observation 60aff75b-b0e4-404a-9576-bd5797658670 · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering.

Continual SFT Matches Multimodal RLHF with Negative Supervision GQA: A new dataset for real-world visual reasoning and compositional question answering

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.592976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:42.781559Z digest=sha256:bf23735d5291bf9c9fe90033cb039a01abaf861dea72d63f6292722a4f8f1434

Observation c9148c5a-5ee2-4f7d-b2e3-1c9ce62084fb · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Continual SFT Matches Multimodal RLHF with Negative Supervision Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.863703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.863703Z digest=sha256:2685a0ccc7ac3f3c2cebc23fafc5cbd6c96c6cb1a37391fecc5930851a8a5ec8

Observation afb2b299-70bc-4c87-a29e-5d634b8226a5 · outbound

This paper cites Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models.

Continual SFT Matches Multimodal RLHF with Negative Supervision Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.961623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.961623Z digest=sha256:025a23f1a8ed47151ebd9e603139025391749f295ae26c9e6c3b22225936b3d5

Observation 67336aec-9136-451a-a5fc-e9949679a539 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Continual SFT Matches Multimodal RLHF with Negative Supervision Evaluating object hallucination in large vision-language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.966247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.966247Z digest=sha256:1e0b271f712182543bff4cf3b02bf51a0cf0d5f5dc91e76958c756a981eb51a8

Observation ca51b88d-cb91-4471-a5e1-e991a12a02ff · outbound

This paper cites Microsoft coco: Common objects in context.

Continual SFT Matches Multimodal RLHF with Negative Supervision Microsoft coco: Common objects in context

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.544074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:42.970537Z digest=sha256:01b884bec70130038e7f5f8da353ca5d0853a5bd8f9aaa43c7bd9ba6d9dd76f8

Observation af73828a-48da-4410-ab60-3ae07928687a · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Continual SFT Matches Multimodal RLHF with Negative Supervision Improved Baselines with Visual Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.975249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.975249Z digest=sha256:e73d158d087c7565957ae116dbac345536ce09cbbbdbfc11bd583109e82e3c37

Observation 6780e922-21a2-4532-b766-7fd447b1d05d · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Continual SFT Matches Multimodal RLHF with Negative Supervision Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.980350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.980350Z digest=sha256:696f15f5f7fc30fe3267e10db3d8ad5d91bdbd537592f1f91a073a47e2ee6f89

Observation f2f0c405-731c-449b-9866-dd977f7744f6 · outbound

This paper cites Visual instruction tuning.

Continual SFT Matches Multimodal RLHF with Negative Supervision Visual instruction tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.518424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:42.984357Z digest=sha256:20dd0136beaa4e5f9ec4eebc45fa5c36f0c81298ea7bbb6a47c8aa94a5aa8d46

Observation d472d0a2-332d-4488-a741-661c935fcb61 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Continual SFT Matches Multimodal RLHF with Negative Supervision MMBench: Is Your Multi-modal Model an All-around Player?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.989080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.989080Z digest=sha256:69b77eaff8c40bc7c880512fdcf0305d10263120a1d2219930bcbe2bd3bf4311

Observation a30925e4-dab7-415a-a62f-7d9533e6f59d · outbound

This paper cites Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer.

Continual SFT Matches Multimodal RLHF with Negative Supervision Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.993918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.993918Z digest=sha256:4a0bc9bab9ced519b9a9bc249124db5862689b68bee7c47ba9bfff284b39e960

Observation 4a149c74-bc1f-48bd-bcf7-914f3eab9554 · outbound

This paper cites Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment.

Continual SFT Matches Multimodal RLHF with Negative Supervision Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:42.998808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:42.998808Z digest=sha256:caa53240c897df2846427b396b7f2068bd9dcaa710cbd1c0c24164ed2306f572

Observation 5c36a4d7-0de0-47d1-8af0-179aa8c5939c · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Continual SFT Matches Multimodal RLHF with Negative Supervision Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.005337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.005337Z digest=sha256:ccfee7e5a930b6fb73945b07d18f0b1dd9d3cded4872f64e72891246d039fc88

Observation 54b27ab6-d0a3-4c2e-8615-17d4a70fdf94 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Continual SFT Matches Multimodal RLHF with Negative Supervision Ocr-vqa: Visual question answering by reading text in images

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.009885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.009885Z digest=sha256:54d9ff1aab9fe0034b8807d21fd9cad5b2e2d1e888fc28987bef3e4f41b3b535

Observation 0073ea85-db72-4aca-bd51-d7b67566ff5c · outbound

This paper cites Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization.

Continual SFT Matches Multimodal RLHF with Negative Supervision Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.014298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.014298Z digest=sha256:f402d36b4b6bcb0264b8205edc0e3e44d8ec059a2a8820e3895a21fa5e864246

Observation bc6b0de0-f74a-44f4-a4ec-85938236add7 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Continual SFT Matches Multimodal RLHF with Negative Supervision Direct preference optimization: Your language model is secretly a reward model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.482432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:43.018997Z digest=sha256:75d1baca8c444cbad166b0cd04ac58d5c919b1c8254b8448dd3680ff95b458b3

Observation b85c0898-0943-4897-8bac-4a5da5a97b5c · outbound

This paper cites Object hallucination in image cap- tioning.

Continual SFT Matches Multimodal RLHF with Negative Supervision Object hallucination in image cap- tioning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.023292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.023292Z digest=sha256:bc2056f36714d7fbb2275dfbc4603196e17bbdc9d61f75e361ae7c6641939253

Observation c428b956-4f77-4182-9193-6431105f8112 · outbound

This paper cites Multitask prompted training enables zero-shot task generalization.

Continual SFT Matches Multimodal RLHF with Negative Supervision Multitask prompted training enables zero-shot task generalization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.455332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:43.061051Z digest=sha256:c11a067cf308e10e80ff5ce737035f1008e10aeced9ca54df8dd41ec19c759b2

Observation bacfa353-f4bd-4eb7-8b4c-ab861d80a813 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Continual SFT Matches Multimodal RLHF with Negative Supervision Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.084563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.084563Z digest=sha256:35a5ba7f8e021341638de251f75911a5c449ec5c3a60aabf187de718a32bb6f1

Observation e495b535-fcf1-4180-a3d4-a1768ff997bc · outbound

This paper cites Towards vqa models that can read.

Continual SFT Matches Multimodal RLHF with Negative Supervision Towards vqa models that can read

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.438925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:43.090981Z digest=sha256:e80fc32a150920a8eb6f323f4f60a852eddc232315c7d8cfdffb985e54ee4d64

Observation 9360bab2-3464-492d-95d0-f3c01d58014a · outbound

This paper cites Improving Multi-modal Large Language Model through Boosting Vision Capabilities.

Continual SFT Matches Multimodal RLHF with Negative Supervision Improving Multi-modal Large Language Model through Boosting Vision Capabilities

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:58:43.881132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:43.096094Z digest=sha256:dd683ac29b13a189a6db1889c3896077508a03d3c50db8896e3499d3b995efc0

Observation 0628616f-3067-4613-8c82-ca2b18bf89e4 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Continual SFT Matches Multimodal RLHF with Negative Supervision Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.100804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.100804Z digest=sha256:75fd01710324b8ad0f7ecda643cf439739b3fd974adcbab1b017862052759eb8

Observation a0a32663-3ec7-471b-90af-7515e6269bf9 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Continual SFT Matches Multimodal RLHF with Negative Supervision CogVLM: Visual Expert for Pretrained Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.106019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.106019Z digest=sha256:6bba03d675a65dc399f77c960f10b14d12ce220ceee0bf5ffa6ebb7a3146515c

Observation fda87008-9453-4592-8f8a-c60a3b90b76f · outbound

This paper cites Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement.

Continual SFT Matches Multimodal RLHF with Negative Supervision Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.111187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.111187Z digest=sha256:e442e8d9e1db0b94ddb5afafe4d9c8541ebe73b30e25f1fd30846058ce8fa6fc

Observation 85960aff-4a04-42e4-9bca-69303c6d15c0 · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

Continual SFT Matches Multimodal RLHF with Negative Supervision Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.116549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.116549Z digest=sha256:1a7b3f948b79d36ba16dddd29c3da06744928f5f3a64840b3a6364030246c889

Observation ba68516c-f71c-4423-bcd3-ffdb9bd47ce9 · outbound

This paper cites Finetuned language models are zero-shot learn- ers.

Continual SFT Matches Multimodal RLHF with Negative Supervision Finetuned language models are zero-shot learn- ers

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.360844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:43.121904Z digest=sha256:26ab59e617ea2385970690aaf9dd3f4835bbe1c93eb4111cdcaf5429b987481c

Observation e6e2386f-5a81-483e-8f08-7af1aca2430c · outbound

This paper cites an unresolved cited work.

Continual SFT Matches Multimodal RLHF with Negative Supervision Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.127468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.127468Z digest=sha256:4cd2dfe84ff67eb1cf4cfed6c0bec7023a2d3890ae119e5a11a45fff8262f6aa

Observation 2d472179-5ac1-45cd-a058-aaaf1c45f832 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

Continual SFT Matches Multimodal RLHF with Negative Supervision Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.132410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.132410Z digest=sha256:1623e9717c133ab3aadff850f1aa9382747bbf5092e3071dc3c3c07aa7154ec0

Observation 9c9d3f91-9754-4836-83aa-61b750e7b016 · outbound

This paper cites Is dpo superior to ppo for llm alignment? a comprehensive study.

Continual SFT Matches Multimodal RLHF with Negative Supervision Is dpo superior to ppo for llm alignment? a comprehensive study

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.258968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:43.137030Z digest=sha256:7765d2a08c486f0ed92f0b3db61913043a2a33119d30382eeeaf1f54e66c3da3

Observation f9cbea60-90fa-4966-b69b-2fd59088d728 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Continual SFT Matches Multimodal RLHF with Negative Supervision MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.142427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.142427Z digest=sha256:f4f44ac57e19de8737229dfa2e63a0508186314c442ce79f869f766c16596d75

Observation add161f7-45d0-4945-9a34-7804831dca39 · outbound

This paper cites Token-level direct prefer- ence optimization.

Continual SFT Matches Multimodal RLHF with Negative Supervision Token-level direct prefer- ence optimization

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.192399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:43.147270Z digest=sha256:588ebd1a1cd86f1eabe909e82ee7a473e13e122f8b72c2f97eb0f023e285756a

Observation 18c45ca8-5129-42e9-9460-eeec3061e06f · outbound

This paper cites ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning.

Continual SFT Matches Multimodal RLHF with Negative Supervision ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.151627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.151627Z digest=sha256:760a4f4e478566f720d981b5fcc0f45b657d2dd7334ede39deac0a2703a1a361

Observation ad2f2b85-ab4b-4fc2-8a3a-d7efb8b3cadf · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

Continual SFT Matches Multimodal RLHF with Negative Supervision Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.156556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.156556Z digest=sha256:518a945e52e9412754e53a1b16e33d8c3a5caf1757b3c138ed4bd69e55f323cf

Observation 117b7452-cfa3-4742-9915-59877dfd51ff · outbound

This paper cites Lima: Less is more for alignment.

Continual SFT Matches Multimodal RLHF with Negative Supervision Lima: Less is more for alignment

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.177098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:43.204931Z digest=sha256:b8b2e47d022019dc64f103f6649b84e36142ee740905bf541a19ef522299e006

Observation e6314292-d25f-44ef-ba1e-b1b1035789f1 · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

Continual SFT Matches Multimodal RLHF with Negative Supervision Calibrated Self-Rewarding Vision Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.287932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.287932Z digest=sha256:7306210809f85dfcadd0a614a7413ce98d1f67a160a87ca86c04dfbfbc30cde1

Observation 32679ab4-f065-4208-9c03-054790116fd6 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Continual SFT Matches Multimodal RLHF with Negative Supervision MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.351600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.351600Z digest=sha256:b8849a5b42cecb930d91d0844b93966d2ac8b57a9e2339d7d856e0aacc89cc4d

Observation 36e81391-2f15-4795-b1f8-4a5ccb92b691 · outbound

This paper cites Multi-label self- supervised learning with scene images.

Continual SFT Matches Multimodal RLHF with Negative Supervision Multi-label self- supervised learning with scene images

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.159954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:43.404328Z digest=sha256:96696f563c4034478bf3fabed7fcbf8db6d61919c25dda113aaafae50de2f25d

Observation 3e34811d-7d7c-42f2-9bb1-df6772bee8f3 · outbound

This paper cites Quantized feature distillation for network quantization.

Continual SFT Matches Multimodal RLHF with Negative Supervision Quantized feature distillation for network quantization

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.143565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:43.409800Z digest=sha256:b9fb6cbe3b3e458ba8510d992a3e56a0f09032820671bf0db61e6ad353b83d63

Observation d8ef4bab-6bf2-4feb-ac9e-bff7cffd3016 · outbound

This paper cites Self-Supervised Visual Preference Alignment.

Continual SFT Matches Multimodal RLHF with Negative Supervision Self-Supervised Visual Preference Alignment

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.414905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.414905Z digest=sha256:527858bbbc0e3630e0bad14ba6a079a6dc34a8474f0736331a3378793fba6ea4

Observation 44b228c8-8180-46e1-9277-8d27eca911cc · outbound

This paper cites Llava-phi: Efficient multi-modal assistant with small language model.

Continual SFT Matches Multimodal RLHF with Negative Supervision Llava-phi: Efficient multi-modal assistant with small language model

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:58:44.127762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:58:43.420092Z digest=sha256:67208d1e93e2dc160a0a36c55807583cd75c3b7e3dc0335a537f4ebfb3010632

Observation 52625be0-87db-4cb8-a6ca-fd175383b8a7 · outbound

This paper cites Multi: Multimodal understanding leaderboard with text and images.

Continual SFT Matches Multimodal RLHF with Negative Supervision Multi: Multimodal understanding leaderboard with text and images

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:43.424797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:58:43.424797Z digest=sha256:17b876ffe634dca15cf65b04aa32826471778403809f4ecfaf0d1600697a82ac

Pith citing papers

No inbound Pith citation observations are available.