Pith. sign in

Paper Citation Record · LEDGER

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models

As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2507.04976.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04976 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:41:39.062407Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8dc3bbba-9988-4504-9af2-da643588d748 · outbound

This paper cites Figure 14:Prompt used for Evaluation:(a) Evaluation prompt for answerable dataset, and (b) Evaluation prompt for our unanswerable dataset.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Figure 14:Prompt used for Evaluation:(a) Evaluation prompt for answerable dataset, and (b) Evaluation prompt for our unanswerable dataset

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:40.468104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:41:38.870093Z digest=sha256:6640dba6e5a5dcc2721743ad86f3fc88b32ee02b640201747f729536f2ea0b29

Observation 02f5da86-7927-41f2-a283-6efff1c09946 · outbound

This paper cites an unresolved cited work.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:41:41.346617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:41:37.065756Z digest=sha256:8ce31e6ab61dbfebc8ae52cb9d2c23150d5842bb673f1c20c421bc2d7c701dcc

Observation 57943758-8e5e-41ea-94d9-3e882fba7ed4 · outbound

This paper cites A.3 ETHICSSTATEMENT In our study, we utilize Large Language Models (LLM) to generate our UVQA dataset and evaluate video-LLMs, which may result in unintended outcomes.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models A.3 ETHICSSTATEMENT In our study, we utilize Large Language Models (LLM) to generate our UVQA dataset and evaluate video-LLMs, which may result in unintended outcomes

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:40.705794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:41:38.744180Z digest=sha256:261327127aeaab0c6c5044b3f98e908490256cff1d673b883af6d9989c5a14dd

Observation 85137c54-8704-444c-bb7d-3a88e310606d · outbound

This paper cites The Llama 3 Herd of Models.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.255667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.255667Z digest=sha256:0c9c1d9d25e4375f63b8f2cb6e74931c437e7e37b10e4cbc8b8ee104dbb5a674

Observation 73f65394-be23-4b3f-a8e5-b7a101e5b96d · outbound

This paper cites UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.335716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.335716Z digest=sha256:15475ae5000d2c5c41fd14bfbde5abcc67451bd86ce91e545f2ffa0c4387420c

Observation 62189781-c344-4e24-b858-45b5902f6ba8 · outbound

This paper cites URL https://aclanthology.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models URL https://aclanthology

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:41.204900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:41:37.912431Z digest=sha256:6f2e74251287f38af769bf6aed212195c056a6f727bf5ea732eacb164797f827

Observation 00f23829-0ca4-4a92-8137-c43ac7228a64 · outbound

This paper cites Language Models Can See: Plugging Visual Controls in Text Generation.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Language Models Can See: Plugging Visual Controls in Text Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.021957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.021957Z digest=sha256:9f751ee2bb239ba9976e98a934e62080abcadcddc040bd680cd41da53414e5e0

Observation 9118d1cc-2d16-46f2-b112-37ed7af3a42a · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.043123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.043123Z digest=sha256:96cdd1b0b35921b549f2795909d05c73013169457a32f9b0dd6fba24a76d6a88

Observation 0347a6c8-223a-4acd-985c-78e758eeced3 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.101205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.101205Z digest=sha256:0f8dc9f2d0e0363e2ef05287014738befb88671fde29f93bf08433f94a28ec16

Observation 0683f6b1-37ae-4a44-a710-5083f293a2a1 · outbound

This paper cites Know Your Limits: A Survey of Abstention in Large Language Models.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Know Your Limits: A Survey of Abstention in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.196093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.196093Z digest=sha256:ba59508410aade379be2ae64c06da8b04b3e91835fa78a67e20511cb08f690c4

Observation 8881582b-3d4b-45c7-9a09-7e66de6adf72 · outbound

This paper cites Reliable visual question answering: Abstain rather than answer incorrectly.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Reliable visual question answering: Abstain rather than answer incorrectly

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:41.101947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:41:38.247478Z digest=sha256:fce224ad2e198ae52e83dfff559a7552e6e593e6e5a6c34d8077e4a9be650771

Observation 94a70e9d-42ba-4c91-8346-677d493d6595 · outbound

This paper cites TLCR: Token-level continuous reward for fine-grained reinforcement learning from human feedback.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models TLCR: Token-level continuous reward for fine-grained reinforcement learning from human feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.377164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.377164Z digest=sha256:5b0e877f7066bbfcf54ebf5c28dd22eb923c0570ad07a4d9944d2fb0f07992af

Observation 161abb5c-963a-480d-9f2d-f8cc0305dab7 · outbound

This paper cites doi: 10.18653/v1/2022.emnlp-main.280.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models doi: 10.18653/v1/2022.emnlp-main.280

Reference 19

Resolution
verified exact
doi, observed 2026-08-06T19:41:39.212989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:41:38.442554Z digest=sha256:c48116f769510b737c75556847513118ce4a0497d1f12db3a335bcaacbe488ce

Observation 6eb2d644-0901-4600-9cfb-d1dd8dfff21a · outbound

This paper cites doi: 10.18653/v1/2023.findings-emnlp.797.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models doi: 10.18653/v1/2023.findings-emnlp.797

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.481770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.481770Z digest=sha256:30acd115a8d4c4a119463f4c3c706391f2c67aa836e954f1f97e24ee8909c841

Observation 80aecfca-840e-43ab-83a4-dc31abeba65e · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 22

Resolution
malformed identifier
no resolver link, observed 2026-08-06T19:41:38.650060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.650060Z digest=sha256:ec56ae0514f36d6054efa1aca6cacd451b02fcff888cf255be2a7c3da73b2543

Observation f6e78513-1439-4059-a12c-58f917db02c8 · outbound

This paper cites an unresolved cited work.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:41:40.862940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:41:38.693004Z digest=sha256:c817e3ae917eb1fed59ba44783cce89f39851f4a49ecb9e3a1a49f4be9b46621

Observation c83c0aad-ef3c-41b9-846e-4d605e0a9699 · outbound

This paper cites (2023); Li et al.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models (2023); Li et al

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:40.579369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:41:38.838760Z digest=sha256:84da196b2560d79e82918d3b084d3531199fbcd4af4e8cca1670d159e638c9c9

Observation 062e422e-1d8d-48b5-9331-3c8c8a40b6d8 · outbound

This paper cites an unresolved cited work.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:41:40.309993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:41:38.914823Z digest=sha256:177f6a2187c90c62d92cbdd04e70047ee7c5a2d36c9ebf86bd635bc24fb3f5e0

Observation 2b3a4891-e017-4a31-ac0b-74358f2b3dec · outbound

This paper cites If the question cannot be answered using the video content, state that it is unanswerable and provide a reason.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models If the question cannot be answered using the video content, state that it is unanswerable and provide a reason

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:40.088734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:41:39.013143Z digest=sha256:438c85657e118836786e124ec82dd42730fb20bd31c7e61724875fdf7f7f3f79

Observation d5262a23-4d2f-48bc-a386-d802029116e7 · outbound

This paper cites 5All annotators have TOEFL iBT scores above 100 and hold at least a bachelor’s degree.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models 5All annotators have TOEFL iBT scores above 100 and hold at least a bachelor’s degree

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:39.948586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:41:39.062407Z digest=sha256:6793baebfc09567c58b4322f82f82b68683c6f35f16449d8bbffd5b9d3e0c10b

Observation df799896-7081-41b3-a7e3-1a1f51c2dac9 · outbound

This paper cites Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David A.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David A

Reference 2008

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T19:41:39.608405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:41:37.587634Z digest=sha256:c9dee1c8feb40042d4c4fcfe3671b3a5296ac7a7ced39c4805f0e0fedc54242d

Observation 6d3a0d12-eb22-46e0-990d-cc9176b83382 · outbound

This paper cites Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.666437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.666437Z digest=sha256:249e2307e16386e251320e800638ba2653e0c9587d4f36080024944301a11103

Observation 75009a38-1cf7-40b2-9eeb-61476426c4a0 · outbound

This paper cites Alignment for Honesty.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Alignment for Honesty

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.287739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.287739Z digest=sha256:84b4b7a16d0acaefb78be070f40c7404767e28dc5c87fab72247d963753becc7

Observation 29e42f01-c0a9-4a27-b57b-40a538265d7d · outbound

This paper cites Mapping Images to Scene Graphs with Permutation-Invariant Structured Prediction.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Mapping Images to Scene Graphs with Permutation-Invariant Structured Prediction

Reference 2018

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:41:39.743993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:41:37.410814Z digest=sha256:829af0b5815724b1bc1c444907b33fd874569a4fd6d94f4f52cb6a34f73afec0

Observation e4ea3966-e96b-4e0e-93ab-c2b5f76afa5a · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video understanding.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Video-LLaMA: An instruction-tuned audio-visual language model for video understanding

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:40.964592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:41:38.533379Z digest=sha256:99a74672fdbbdf564ed9dc928a5641f1671914a3f89f4536deec2d276f4e0ea8

Observation 6d9a7dca-3d00-45b4-b331-d8a61213e178 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.146585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.146585Z digest=sha256:516444a140a4bcbfdcbb53faa171c2fc0741a1cdd144e4b12843105e6fce02a3

Observation 0087c442-8ec5-4d59-879e-a0b2f32f4ba8 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.813687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.813687Z digest=sha256:c0dd0a7e0fb492d03cd8274c004de7f074be36967c5e78307515b55b5641b077

Observation 08062c73-ad7b-42cc-8184-5b9fbc832a40 · outbound

This paper cites doi: 10.18653/v1/2023.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models doi: 10.18653/v1/2023

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.497820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.497820Z digest=sha256:20de21d9dfab70cf7469389ae61f21349d8b9cf47362256c02151dc1a954ccae

Observation 9242563d-580b-4c2c-9616-b4cf89155b5e · outbound

This paper cites PaLM 2 Technical Report.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models PaLM 2 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.015458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.015458Z digest=sha256:f80f4633f55c6919454df9fb2cdcb2222c12829e75556b415174d15af808ab07

Pith citing papers

No inbound Pith citation observations are available.