Pith. sign in

Paper Citation Record · LEDGER

Controlling Multimodal LLMs via Reward-guided Decoding

As of 15 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2508.11616.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.11616 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:52:50.394161Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact2
  • verified fuzzy28
  • unresolved28
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1a4b941-605e-47a9-ab80-e1429cdd34d7 · outbound

This paper cites GPT-4 Technical Report.

Controlling Multimodal LLMs via Reward-guided Decoding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.154994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.154994Z digest=sha256:1a70ab6b6affd47f7cb40f7b186af91a4a5cc578dcff8937f69fb0621b7944d9

Observation f1ba0cee-9f31-4e8b-a0fc-4456da44a26f · outbound

This paper cites Understanding Alignment in Multimodal LLMs: A Comprehensive Study.

Controlling Multimodal LLMs via Reward-guided Decoding Understanding Alignment in Multimodal LLMs: A Comprehensive Study

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.159792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.159792Z digest=sha256:cf38e4f182ff79e014e918d0c60ed1387a4adfc1a94ff91553617822068d075d

Observation 9195a9d8-747a-4bec-887e-5f8ea3d54aba · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

Controlling Multimodal LLMs via Reward-guided Decoding Hallucination of Multimodal Large Language Models: A Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.164694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.164694Z digest=sha256:0ccab2eef1073d5620044f36bddbf4b2c762f9f1cd3d32e815ffd014dcfc8350

Observation c45988db-9f13-4a82-8f9f-bfc81ab5dda1 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Controlling Multimodal LLMs via Reward-guided Decoding PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.169620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.169620Z digest=sha256:86ce85f32f75e6ec2c4f41b32385f66dd5828122e86016c96389550008cfa500

Observation 37c161bf-3b75-408d-9fc4-df691ab59ba6 · outbound

This paper cites An Introduction to Vision-Language Modeling.

Controlling Multimodal LLMs via Reward-guided Decoding An Introduction to Vision-Language Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.173749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.173749Z digest=sha256:0936f26f63cf554344eb83e8386947cd034cf338ca3481534a1096b37a901c98

Observation 2e014888-65c9-4a83-a39c-9005d33696d5 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Controlling Multimodal LLMs via Reward-guided Decoding Rank analysis of incomplete block designs: I

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.289356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.178060Z digest=sha256:bfb66a9fde8f75a4d61d2bd055a1f513ad6e4105f07b31306c2c4672c1c40148

Observation ef37e90e-3dbf-4593-9d25-0db8535f987b · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Controlling Multimodal LLMs via Reward-guided Decoding Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.181895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.181895Z digest=sha256:3802b29c0ce162484462c3763a01c5bc90ec8acdda1f0ec2cddc663873b6326b

Observation 334f6ea5-faa9-4989-9a8a-1ab7459ff0ae · outbound

This paper cites Language Models are Few-Shot Learners.

Controlling Multimodal LLMs via Reward-guided Decoding Language Models are Few-Shot Learners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.186179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.186179Z digest=sha256:1641539d7886ae5e966392bc7ae32a109d5c6dac716fe05c91c5b59682df7a36

Observation 46ac2cda-50aa-4279-9c0f-b634d3061524 · outbound

This paper cites The (r) evolution of multi- modal large language models: A survey.

Controlling Multimodal LLMs via Reward-guided Decoding The (r) evolution of multi- modal large language models: A survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.189706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.189706Z digest=sha256:30fc5c8a0a53dd9f0303e094acf808cafddbdce751aec046ee868ccae0ef9b78

Observation 439aac8e-d30a-4625-ab01-fee4a7b80ec8 · outbound

This paper cites End-to- end object detection with transformers.

Controlling Multimodal LLMs via Reward-guided Decoding End-to- end object detection with transformers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.277739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.193277Z digest=sha256:35149265f4b4a8de068bc5b9c1f290b01f93540e791b3a31a00018316968797d

Observation 0d560a61-1267-4238-8b4c-a8e4fc62b192 · outbound

This paper cites Plug and play language models: A simple approach to controlled text generation.

Controlling Multimodal LLMs via Reward-guided Decoding Plug and play language models: A simple approach to controlled text generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.266435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.197122Z digest=sha256:6cffefedd41c29a8c2f13e083a301a78983a853a959255ede5214515a63c0821

Observation d48f62ed-f031-4991-b33f-ff70f1455e3b · outbound

This paper cites Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding.

Controlling Multimodal LLMs via Reward-guided Decoding Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.201207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.201207Z digest=sha256:24a7c5c4786c844de5db3b78c0ffe8894931e51783de7336f0a173b29870b3e1

Observation a0243cfc-04f9-4ed4-85e3-df8eba7f9d7a · outbound

This paper cites Reward-augmented decod- ing: Efficient controlled text generation with a unidirectional reward model.

Controlling Multimodal LLMs via Reward-guided Decoding Reward-augmented decod- ing: Efficient controlled text generation with a unidirectional reward model

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.254427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.205481Z digest=sha256:b3395fe709cc07001854d9fd34fffa9cf27517af8d50ee11d346d6ad30b62a84

Observation d9812ea6-4b95-47ed-be5f-1e9a817c7c31 · outbound

This paper cites The Llama 3 Herd of Models.

Controlling Multimodal LLMs via Reward-guided Decoding The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.209103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.209103Z digest=sha256:5c414f52ddd1e51ab439cc68d5c6a78850280746f58d2ec7a06a94df48a86646

Observation 1ff4a560-5341-427f-a45f-3b3a180ba014 · outbound

This paper cites Multi-modal hal- lucination control by visual information grounding.

Controlling Multimodal LLMs via Reward-guided Decoding Multi-modal hal- lucination control by visual information grounding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.243172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.214263Z digest=sha256:668d2eb1de5e5127b3ef65ab8ed2772429b645f5d4b3aa0e74ea3123128cd4dc

Observation 76f764cb-197a-4e02-8bac-ad017610ec3a · outbound

This paper cites Value Augmented Sampling for Language Model Alignment and Personalization.

Controlling Multimodal LLMs via Reward-guided Decoding Value Augmented Sampling for Language Model Alignment and Personalization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.217899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.217899Z digest=sha256:7275f8c78582f345aa35a3752ea63aa1fae635af4fa8f4122a500bb9dcca5d12

Observation bd8a1b6f-a981-41f8-bc3d-56d9f6a80c33 · outbound

This paper cites Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.

Controlling Multimodal LLMs via Reward-guided Decoding Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.231390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.222037Z digest=sha256:e28f9982d18a832c33f321538b2adfed2dbf28f24cff4a959bf89a0de6301edc

Observation 3e8ce828-3c3f-4b84-93ee-523f77de5faf · outbound

This paper cites Lora: Low- rank adaptation of large language models.

Controlling Multimodal LLMs via Reward-guided Decoding Lora: Low- rank adaptation of large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.220609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.225488Z digest=sha256:38b72529a0103fcfea2e2480fddff6857bbb875c802fe13ba5f1837edfedf09e

Observation 826fd636-5466-42e6-addf-efbb1a1c091e · outbound

This paper cites Args: Alignment as reward-guided search.

Controlling Multimodal LLMs via Reward-guided Decoding Args: Alignment as reward-guided search

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.210805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.229886Z digest=sha256:aba8c0dc05b3c2d37765da795d7d6808bd8ce2b17a5f391431dc0673b4142b5c

Observation b3ec0fdf-f59d-440d-9450-2477156f23d6 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Controlling Multimodal LLMs via Reward-guided Decoding RewardBench: Evaluating Reward Models for Language Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.233278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.233278Z digest=sha256:4de1892b5d783c664368f9d5cd0c98e48cc4fbd9efff5579fa30850f8898ae2c

Observation 372d28b3-8d50-408a-8274-a70465e5fb1d · outbound

This paper cites Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding.

Controlling Multimodal LLMs via Reward-guided Decoding Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.200259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.237336Z digest=sha256:0e778bfe80ac1b5bcb318263cf77cd1f3db675d4aaca7bf90dd17a7abc7dc448

Observation 19bba8e2-9a60-4c24-98e9-d6aaad167b33 · outbound

This paper cites Sequential monte carlo steering of large lan- guage models using probabilistic programs.

Controlling Multimodal LLMs via Reward-guided Decoding Sequential monte carlo steering of large lan- guage models using probabilistic programs

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.190272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.240846Z digest=sha256:d77b08e57ed7e5dd7c635661e19416ea3b7bc6e71279d45e26cdab293dce43b2

Observation ef43729f-048b-4a7a-89d6-9bb851ee4dc9 · outbound

This paper cites Cascade reward sampling for efficient decoding-time align- ment.

Controlling Multimodal LLMs via Reward-guided Decoding Cascade reward sampling for efficient decoding-time align- ment

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.177822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.245204Z digest=sha256:3f8955b74056223dd42f4bd12d0a83c885ca3388beb2109399310446896ee077

Observation 58053729-0009-4ac4-bf4f-2ff1c2de74c5 · outbound

This paper cites Vlfeedback: A large-scale ai feedback dataset for large vision-language models alignment.

Controlling Multimodal LLMs via Reward-guided Decoding Vlfeedback: A large-scale ai feedback dataset for large vision-language models alignment

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.166165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.248803Z digest=sha256:c262844d62b1c539b51045d9089af15c250fc9cf43df8468c3e5a13d5e85de91

Observation 67a17eee-0953-42c8-9389-3abfde6e5385 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Controlling Multimodal LLMs via Reward-guided Decoding Evaluating object hallucination in large vision-language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.153056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.252401Z digest=sha256:cdf05873f0e11cf6f9d8f95f0dc3124734d0482e0423e48e35a55a47b3371b5a

Observation 6597e2da-aee1-4f69-a328-ff6636a2ded5 · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

Controlling Multimodal LLMs via Reward-guided Decoding Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.140851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.256024Z digest=sha256:53ca05453df019894b8b90d2131852ee7a13cb8948e4327e684bd3554ead169a

Observation 29aaa85c-6b44-4b2f-87b7-622186dbcf1f · outbound

This paper cites Improved baselines with visual instruction tuning.

Controlling Multimodal LLMs via Reward-guided Decoding Improved baselines with visual instruction tuning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.128879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.259630Z digest=sha256:05dba337e0b23cd54fc5fbf6e9911178f37e3a92636460c8d0f8db7fccad0dcc

Observation 04e659d9-92b8-4c54-b0fa-96b3b42aafcd · outbound

This paper cites Don’t throw away your value model! generating more preferable text with value-guided monte-carlo tree search decoding.

Controlling Multimodal LLMs via Reward-guided Decoding Don’t throw away your value model! generating more preferable text with value-guided monte-carlo tree search decoding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.115433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.263735Z digest=sha256:2eaada7b68e50c908d8a3e2370aa488a9f334ed2ec85657a8995116da8b4c39d

Observation f205440f-a614-4cb6-a324-07957d941101 · outbound

This paper cites Investigating and mitigating object hallucinations in pretrained vision-language (clip) models.

Controlling Multimodal LLMs via Reward-guided Decoding Investigating and mitigating object hallucinations in pretrained vision-language (clip) models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.103132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.267400Z digest=sha256:3d9bf2da1013da438af63e95950beff0a96fb4ada7d2225f6c25465c596c37d2

Observation 02a7cfff-3c27-4b07-ad5b-22851ff919eb · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

Controlling Multimodal LLMs via Reward-guided Decoding SmolVLM: Redefining small and efficient multimodal models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.271460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.271460Z digest=sha256:71118e46385d90b366ad4aad6f0e730adf546a8418b9c4536d4f62eefd6cebfd

Observation 61c74f47-441a-413f-af26-9950f4aa96e3 · outbound

This paper cites Scaling open-vocabulary object detection.

Controlling Multimodal LLMs via Reward-guided Decoding Scaling open-vocabulary object detection

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.090677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.276395Z digest=sha256:093dad0f115ea098faec884b8f8bd4746d84fda29ab44c4d64bebba6da8b14fa

Observation 942d4a95-69c1-4301-a756-35226a71d2fa · outbound

This paper cites Controlled decod- ing from language models.

Controlling Multimodal LLMs via Reward-guided Decoding Controlled decod- ing from language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.079046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.280460Z digest=sha256:b77497cd612557472a76175ac897949dec6fe2b60abcd9c0b4fdb61e50cd59da

Observation 22c9f9d3-3275-46a8-8896-5c8e209054f7 · outbound

This paper cites Training language models to follow instructions with human feedback.

Controlling Multimodal LLMs via Reward-guided Decoding Training language models to follow instructions with human feedback

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.067203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.285603Z digest=sha256:b38cb684995105cf788112894caf9a807e111c323208b10cd1aa05ed660834c9

Observation 53fa1001-81ea-4ce1-a5ef-a22c76750746 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Controlling Multimodal LLMs via Reward-guided Decoding Learning transferable visual models from natural language supervi- sion

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.290201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.290201Z digest=sha256:408af69ab84d025c50cdad292115f79d597ab603104c5fe9648141adb873cc73

Observation a7dcd835-ec49-40c6-82ba-5f5a96b0f0f8 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Controlling Multimodal LLMs via Reward-guided Decoding Direct preference optimization: Your language model is secretly a reward model

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.049641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.295525Z digest=sha256:8c06aea79c2999adfb2cb188326c0a1c7b839a0790eca28736e0fb313dedb2f9

Observation 0f17a343-f98b-446c-878e-faeae2c5b3f3 · outbound

This paper cites A critical look at to- kenwise reward-guided text generation.

Controlling Multimodal LLMs via Reward-guided Decoding A critical look at to- kenwise reward-guided text generation

Reference 36

Resolution
verified exact
raw_fallback, observed 2026-08-05T19:52:50.735262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.299628Z digest=sha256:24581e4a3c6ffb4ddca94ac114aa111a179232a63b3e04f66e249732fab99363

Observation 0dfb7e8f-4626-4524-b70f-0fed2ea3b02b · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks.

Controlling Multimodal LLMs via Reward-guided Decoding Sentence-bert: Sentence embeddings using siamese bert-networks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.038539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.305083Z digest=sha256:4eafc7bcd4b69291005fa8a11333645e0a6672f36efe0ca65466f13232d46542

Observation 8fb41234-61de-4a1d-be41-294783ee2b63 · outbound

This paper cites Object hallucination in image cap- tioning.

Controlling Multimodal LLMs via Reward-guided Decoding Object hallucination in image cap- tioning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.026452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.309903Z digest=sha256:657f6d4025c3a7b69c983c0c4500b10fa0a014820ff518367b66f1fd65269089

Observation 38a8d27a-7220-42b2-932b-d2eb8fd01b5d · outbound

This paper cites Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment.

Controlling Multimodal LLMs via Reward-guided Decoding Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.314209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.314209Z digest=sha256:03f07656a46e4150bdaa386a3e083f219cabe543230ae60ed9e442af934442eb

Observation a8e5d578-55d5-4073-9049-6297fb93ae26 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Controlling Multimodal LLMs via Reward-guided Decoding Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.318279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.318279Z digest=sha256:05e1a18b10fd573ff97ebbc6df9c21f902eb20acd7aa49a9b990ffb93ca7aaf0

Observation ad339d84-e493-4831-89cc-7096e72b5e41 · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

Controlling Multimodal LLMs via Reward-guided Decoding PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.323102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.323102Z digest=sha256:40e2a303ad96cfcbd40a94156e6d2d5cf752b96d35c478681a1507282a5c027a

Observation bccd353f-0b69-4cef-9847-1bc0e1faafd3 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Controlling Multimodal LLMs via Reward-guided Decoding Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.327059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.327059Z digest=sha256:2041a8b12d9081383acdc12657f47caf812210ef2b4c5d0f2eef613bf1efcbd1

Observation e80dd9b1-5646-4091-99f2-c0e139688bce · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Controlling Multimodal LLMs via Reward-guided Decoding Gemini: A Family of Highly Capable Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.331059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.331059Z digest=sha256:1ff99bf1037f86b6fc32499ff4f7b2a782068edc7e7bf9a5ba9561dd9e8103ab

Observation b7c51513-c48c-4e4b-b0b8-a817735f8afd · outbound

This paper cites Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training.

Controlling Multimodal LLMs via Reward-guided Decoding Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:52:50.609153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.335004Z digest=sha256:7dbbca4f2c33a38f707e9523142c94b095c9bd26abd6071d07e3818d64f9e1dd

Observation a82b1c07-6d07-4831-ad0d-e46326757cc6 · outbound

This paper cites mDPO: Conditional Preference Optimization for Multimodal Large Language Models.

Controlling Multimodal LLMs via Reward-guided Decoding mDPO: Conditional Preference Optimization for Multimodal Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.338843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.338843Z digest=sha256:7686066d6dbff3e6d50cc53b0faaf342a0ad8afb2922795b433f8b197b9297fa

Observation f5bee5c2-4487-4ccc-a420-ecd5e506f822 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

Controlling Multimodal LLMs via Reward-guided Decoding AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.342468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.342468Z digest=sha256:923112fe6294781fc32310dd586989a27a8e5279fd8630c2b43b457908590ce7

Observation e1387abe-36ad-4a24-8474-a8aa06df1608 · outbound

This paper cites Fudge: Controlled text genera- tion with future discriminators.

Controlling Multimodal LLMs via Reward-guided Decoding Fudge: Controlled text genera- tion with future discriminators

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:51.013349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.346126Z digest=sha256:3504cc30b708fd92364c60c626663a45bf13fe2f921b9032880415fbd9543ee8

Observation 1dadae72-8f9d-4a43-98a3-363fc9f4b752 · outbound

This paper cites Woodpecker: Hallucination Correction for Multimodal Large Language Models.

Controlling Multimodal LLMs via Reward-guided Decoding Woodpecker: Hallucination Correction for Multimodal Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.351452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.351452Z digest=sha256:b786764c4652d1f83b5614cfc38f8a63953d3da25548bb75151d0addbddcb84e

Observation 84cbdba1-2f16-4ba6-b72f-98472b1082cb · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional hu- man feedback.

Controlling Multimodal LLMs via Reward-guided Decoding Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional hu- man feedback

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:50.999855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.355624Z digest=sha256:3b5f91daec87daa8d8facf3bf628a6f000d48a31ace1ea6612e924b31a50ebe0

Observation ae856375-f6d8-44ce-8fe2-1619c0597096 · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.

Controlling Multimodal LLMs via Reward-guided Decoding Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.359220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.359220Z digest=sha256:a33a677529c843f4570efab0a13e52c040b59eb3e369d62327a363f65c9e3437

Observation 2b6617e6-dee4-468c-9fce-02b3789331cf · outbound

This paper cites Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective.

Controlling Multimodal LLMs via Reward-guided Decoding Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.363102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.363102Z digest=sha256:68def8d9deb2215928a4171be4d22747c1a2645d7502a3300b78ca13c6738271

Observation a10d5500-02b9-4cc8-b43e-a2fbc8b451a7 · outbound

This paper cites Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models.

Controlling Multimodal LLMs via Reward-guided Decoding Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.367509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.367509Z digest=sha256:2cee48c76107306ea7fa2771e50b3dd6811cc5c61dcf2fb9cdf993f8bfe8f42a

Observation 2fbbeb2a-fc82-43fe-b537-59b5b329929d · outbound

This paper cites Multimodal chain-of-thought reasoning in language models.

Controlling Multimodal LLMs via Reward-guided Decoding Multimodal chain-of-thought reasoning in language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:50.984480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.372458Z digest=sha256:966054482facd931f86acf2dccb087cfde22501b59fe4eea5c9096e621cd1616

Observation 9a6e7ddc-b304-4ad9-ba1e-0ff10c7e9035 · outbound

This paper cites Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance.

Controlling Multimodal LLMs via Reward-guided Decoding Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.376122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.376122Z digest=sha256:d4a78f0c5c8f7b4f21c2fd1f0f81e04d498fae6a198901825f9d49ae739137dd

Observation 641aad0b-c7e6-4738-bd4b-3560e69da43c · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

Controlling Multimodal LLMs via Reward-guided Decoding Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.379966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.379966Z digest=sha256:7701a5e208d2026f9e6e2e76a5a83bb3c693804e2f130765e1a1890d0381d879

Observation fa86e7c7-b053-4781-b60f-33adc81f7be3 · outbound

This paper cites Aligning modalities in vision large lan- guage models via preference fine-tuning.

Controlling Multimodal LLMs via Reward-guided Decoding Aligning modalities in vision large lan- guage models via preference fine-tuning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:50.972910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.383630Z digest=sha256:f62c8c08cd00519d7d9498bbc8acc976ec06a6b693c119c569c440e129c21288

Observation 5369d77b-72fb-4a3e-b16e-483daf65cdb6 · outbound

This paper cites Analyzing and mitigating object hallucination in large vision-language models.

Controlling Multimodal LLMs via Reward-guided Decoding Analyzing and mitigating object hallucination in large vision-language models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:52:50.961526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.386798Z digest=sha256:9b8c755f08dd1a6c1dd2aaefb641bfcff65b3f8e841ed3118e5d431843e3768a

Observation d23d70d0-a0c0-4148-8f1a-b3ad230b8e11 · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

Controlling Multimodal LLMs via Reward-guided Decoding Calibrated Self-Rewarding Vision Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T19:52:50.390184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:52:50.390184Z digest=sha256:0d88ba83c67a47c849ce1a6430a6c4df27f4cce0e70e6192baa60254cc7409c7

Observation 9c058615-4b94-457e-9c0b-7ae63f453588 · outbound

This paper cites Describe this image in detail.

Controlling Multimodal LLMs via Reward-guided Decoding Describe this image in detail

Reference 59

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T19:52:50.950962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T19:52:50.394161Z digest=sha256:72b6fb5b6556814f7c7c267de8597f6373ffe4f7216c9ad494394ef3cc12375d

Pith citing papers

No inbound Pith citation observations are available.