Pith. sign in

Paper Citation Record · LEDGER

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward

As of 12 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2608.06930.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06930 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:13:44.573450Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact2
  • verified fuzzy26
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4eaed5f7-1f2a-4ff8-8f8f-92924844fd08 · outbound

This paper cites Avocado: An audiovisual video captioner driven by temporal orchestration.arXiv preprint arXiv:2510.10395, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Avocado: An audiovisual video captioner driven by temporal orchestration.arXiv preprint arXiv:2510.10395, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.001303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.001303Z digest=sha256:1d45103fcc0afb125e7203ad66eeb3caedf9076624b09bea1b9b5557c3976096

Observation b695bdef-a788-4514-949d-8065a0a98e1f · outbound

This paper cites Ugc-videocaptioner: An omni ugc video detail caption model and new benchmarks.arXiv preprint arXiv:2507.11336, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Ugc-videocaptioner: An omni ugc video detail caption model and new benchmarks.arXiv preprint arXiv:2507.11336, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.012944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.012944Z digest=sha256:945d6f7b7fbb8683fe143f20ae422a1cd79f512387b8754631591b0b5f8b3ddc

Observation ad1b155c-49ef-45af-b7b2-78dea5511969 · outbound

This paper cites Daily-omni: Towards audio-visual reasoning with temporal alignment across modalities.arXiv preprint arXiv:2505.17862, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Daily-omni: Towards audio-visual reasoning with temporal alignment across modalities.arXiv preprint arXiv:2505.17862, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.020242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.020242Z digest=sha256:62c9c02f34256224f996c56091236a04f6fcfebec7f305f08d8e1522bc90601b

Observation f454879b-db8d-4894-b6a7-98084bef545c · outbound

This paper cites VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.027540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.027540Z digest=sha256:34e764f36a14f006697b04302103fe67c41188bd2aeee1f1f907089e68bf452a

Observation 8f6e4ae1-519f-4569-be74-61cf0eed85a9 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37: 19472–19495, 2024.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37: 19472–19495, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.162755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.035846Z digest=sha256:42997c6ab38a9ffb2c1a1ecf2343d6edfe02b757b8ce6bdef2506cb068f4bb84

Observation 55a344b0-2078-444f-9837-3581bfcb57db · outbound

This paper cites Mavors: Multi-granularity video representation for multimodal large language model.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mavors: Multi-granularity video representation for multimodal large language model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.043255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.043255Z digest=sha256:8b9f0ebb5bf8e3c2d6a66c6e0be0046d7574bc449902a4e513cd53566b2b2fae

Observation 5c73a9e0-e967-446c-a679-d61cfd231a08 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.051144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.051144Z digest=sha256:e5c1e81f7b88343c402974603f6238ceb71f96ba6fa307725e59923adab48dff

Observation 90db9274-bbf3-46ad-8514-45a615392547 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.059127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.059127Z digest=sha256:08f52558a960bbdde815cef7ee152021a19f19606856aca7c3bf4789b4a5a957

Observation b84cf9f9-f405-441b-9a80-a8a0db9a6853 · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.066103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.066103Z digest=sha256:7d0371dd4e4820b6b59092c76b70b6332e7cae0ee131abd2c7cecc6fbf4876fb

Observation ac750a63-35b8-467a-9019-cfa9d9a1d905 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.074176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.074176Z digest=sha256:c4d75d17f41aa199e88a9dee04f292aaeb6c832457fba40e5828703c59efe3dc

Observation 727dae58-6d8b-4478-bb13-886fcd349312 · outbound

This paper cites Advancing high-resolution video-language representation with large-scale video transcriptions.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Advancing high-resolution video-language representation with large-scale video transcriptions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.108399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.080659Z digest=sha256:c33a636a46bd09c6fafdc87100a2c2f688119baef0cfb64615044968b2753cc1

Observation 6995e3e9-3e81-4806-af86-eef40d176bdb · outbound

This paper cites video-salmonn 2: Captioning-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward video-salmonn 2: Captioning-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.087226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.087226Z digest=sha256:395e7c5e18dc98792c85566fb9ca60fd9ad0e2add6ef16dc970f2f5a1b51282e

Observation 7f752d0b-4aad-460d-9aba-c565c4d88bb7 · outbound

This paper cites AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.093438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.093438Z digest=sha256:4fda253977079e6a7fbea84d76c7fd12ee5256d295432c4908c547d990effc5c

Observation 70ff3cbf-52e2-4513-b5d7-e931862558b8 · outbound

This paper cites Longvale: Vision-audio-language-event benchmark towards time-aware omni-modal perception of long videos.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Longvale: Vision-audio-language-event benchmark towards time-aware omni-modal perception of long videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.086591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.099582Z digest=sha256:c806aba899b3175fed2883cc059a1b68cfdb32be4b93b065acfbb6df9d73ca46

Observation 8db4405e-fa54-4f54-95d9-ff440f26235c · outbound

This paper cites Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.106889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.106889Z digest=sha256:2faa68cc96b5efa3e42b4efc87d1793cd1f09f454ad636f7abd53124a7f1433a

Observation a402971b-163d-41f1-bf6a-5779e77e4a33 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.113320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.113320Z digest=sha256:30d5500d251c3bfd7f54f6eab71582a11a5d0333d31aa690460fb916e7783955

Observation fafedc81-20fb-4608-872c-4b5035d3bde9 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.120212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.120212Z digest=sha256:2af3b60982c06a317cea935fd0da92108118ab95379cfbd4fa12645b3ccb8ac5

Observation 87b8dd1b-90a8-4b20-bc27-6bb85e24b4a2 · outbound

This paper cites Watch, listen, and describe: Globally and locally aligned cross-modal attentions for video captioning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Watch, listen, and describe: Globally and locally aligned cross-modal attentions for video captioning

Reference 18

Resolution
verified exact
doi, observed 2026-08-10T18:13:44.635267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.126805Z digest=sha256:ef776e3dc2ccc1f5f0a0618e666085cd8f0cd943fdd8114106cdbffdc1428709

Observation e10c871d-faef-49c8-bac8-ad08c41aec11 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.133036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.133036Z digest=sha256:e17cf41d43013291fd5f4e9dcbda0d0eb9a34893e4b6226968e2d1bd29036b96

Observation 3a17029f-7952-4eac-a5f8-50f77b513d43 · outbound

This paper cites Qwen2.5-VL Technical Report.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Qwen2.5-VL Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.140160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.140160Z digest=sha256:79a37f4ad33cb4e2a63839f13b72f9592feb6708f4ababb81fe0d1cd7061ad61

Observation 2d47876c-b084-4857-a211-94c6eba7760b · outbound

This paper cites Glave-cap: Global-local aligned video captioning with vision expert integration.arXiv preprint arXiv:2509.11360, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Glave-cap: Global-local aligned video captioning with vision expert integration.arXiv preprint arXiv:2509.11360, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.146443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.146443Z digest=sha256:e86ab8f24a8d091c658bd6972ede06b968c45633714b25cea7694da6fcbed408

Observation f8f1a29e-aef0-4ac8-9db6-e8e868898a3a · outbound

This paper cites Qwen2.5-Omni Technical Report.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Qwen2.5-Omni Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.151110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.151110Z digest=sha256:41c0dc4d8255c0dcf4b9a46ad82b687bf2a113dd2d46af91bf2d4a5e11b6eed4

Observation ead60c4d-46c1-4950-977f-efa478de054f · outbound

This paper cites OmniCaptioner: One Captioner to Rule Them All.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward OmniCaptioner: One Captioner to Rule Them All

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.157674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.157674Z digest=sha256:c6f88b931c520755c4b2ebbf6b8625ca49636d14bb773d6e86a91b2a6f16a1a9

Observation d1c38d65-50d3-4629-b59f-5b52c5a203e4 · outbound

This paper cites Omni-captioner: Data pipeline, models, and benchmark for omni detailed perception.arXiv preprint arXiv:2510.12720, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Omni-captioner: Data pipeline, models, and benchmark for omni detailed perception.arXiv preprint arXiv:2510.12720, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.163332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.163332Z digest=sha256:0b768fc756eabb68fad1a241a13405e1bd5075bf9a7f3511add38c67830327b8

Observation aab17aee-273f-41a8-9111-9ea95809ed00 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.169838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.169838Z digest=sha256:f7b766396832776372fe6cea8adf6001790bdff35ccbab3b76c81cbcd1ab30c5

Observation 91ab4e8d-8701-43e7-a6f1-180c2e0142ef · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.176917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.176917Z digest=sha256:8faf56a634b2dd01fcdc6af3cf1743525c1f56a3c511d697427e16c5c2ca2aaf

Observation 1dfeba4d-990f-4466-961f-7991cdc82a8b · outbound

This paper cites AdaTooler-V: Adaptive Tool-Use for Images and Videos.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward AdaTooler-V: Adaptive Tool-Use for Images and Videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.183615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.183615Z digest=sha256:32143f195320a45d14c4c9e13e6c6af63e36b54c3e28799cb6d6730208321b28

Observation 0e486e27-48ba-4008-87db-37a63dcf1e18 · outbound

This paper cites Deepvideo-r1: Video rein- forcement fine-tuning via difficulty-aware regressive grpo.arXiv preprint arXiv:2506.07464, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Deepvideo-r1: Video rein- forcement fine-tuning via difficulty-aware regressive grpo.arXiv preprint arXiv:2506.07464, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.190682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.190682Z digest=sha256:a3de8b94601a2725cc340f2abeb40590c8e004bf308885ea4319a71c45514d92

Observation 89671b3f-be56-4556-bfb1-d94ce96e8412 · outbound

This paper cites Editthinker: Unlocking iterative reasoning for any image editor.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Editthinker: Unlocking iterative reasoning for any image editor

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.197366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.197366Z digest=sha256:074d1c35b718a057e60ea9388a542f6f6856009d985653c2fd56393f2604f899

Observation 7e3a51c8-22b5-4d13-a606-cc7a0e510022 · outbound

This paper cites Sophiavl-r1: Reinforcing mllms reasoning with thinking reward.arXiv preprint arXiv:2505.17018, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Sophiavl-r1: Reinforcing mllms reasoning with thinking reward.arXiv preprint arXiv:2505.17018, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.204027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.204027Z digest=sha256:0014932710f20eaa52ac9169e6aa58cb92875eaf09cfafb74321b47468cfa306

Observation 3ec3c7c0-87f4-447b-885a-edc9ca334b31 · outbound

This paper cites Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.209331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.209331Z digest=sha256:587ca2c187682c64d78b379daf1d17b1ce65fed9de49d9305bb3b70ffbc9b419

Observation cf7d5b55-0444-40c4-aacb-a628a9681f73 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.215349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.215349Z digest=sha256:7bef6eec6cd19719685d01aeb543b3c4c518aa8384757db246ce04447acd06c5

Observation 65539722-4660-4024-be93-223a12fa8d76 · outbound

This paper cites OneThinker: All-in-one Reasoning Model for Image and Video.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward OneThinker: All-in-one Reasoning Model for Image and Video

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.221576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.221576Z digest=sha256:45bab50e5f4c14b7acac7cb618d445f2255a1dfe867e312a170c16c0d16ce6d2

Observation 2eb3994f-034a-49ca-a654-13259944ce9d · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.228889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.228889Z digest=sha256:e5ce62fd71887da054000d39236613fe56bc5f3fa2cfe5d81cec3d16696506e5

Observation 53fcffc3-aa5f-49b5-ac92-20dd6ee28176 · outbound

This paper cites VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.235558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.235558Z digest=sha256:fc5406a2b12013cbfe42fb8e8253e073f7f1fd6e7de246e7c4279aa5c23e7a81

Observation d031b374-8983-4f7e-a325-cc8741da424e · outbound

This paper cites Exploring the role of audio in video captioning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Exploring the role of audio in video captioning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.067484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.241601Z digest=sha256:cd087868fdf3f2d7017dc96990f024698019907d431efb243182c01c37153e49

Observation 81e8b0ea-9029-44e0-8c6e-06c0df6bdacb · outbound

This paper cites Hybrid transformers for music source separation.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Hybrid transformers for music source separation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.049145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.247461Z digest=sha256:68780d19f9c3c99d9262c973d324dee565fba166e78e26df13dffc2e1afb7e6c

Observation cc0ed116-12d4-4b94-929f-bb14edf82074 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Audio-visual event localization in unconstrained videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.255613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.255613Z digest=sha256:57e8a1a71ff7c2ac58e73c8be38d9e3565e585127cd5840408eaa536a8bfb53e

Observation f3944060-20a1-4e43-ac79-d894f5e3ef54 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Vggsound: A large-scale audio-visual dataset

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.018344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.261995Z digest=sha256:a1015b3fcb9c226c571d0bbb6261fbc862eb7f7f9cc2ba0c45f914d936193baa

Observation 3ca62a8c-3c7f-419f-b375-76f43901be93 · outbound

This paper cites Condensed movies: Story based retrieval with contextual embeddings.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Condensed movies: Story based retrieval with contextual embeddings

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.268021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.268021Z digest=sha256:c26dda7df29f69acced354fa93b8aa576f71808566b343dd843d16dd7e691524

Observation a4d9dd4f-6e26-4315-92f0-725f21550be2 · outbound

This paper cites Avqa: A dataset for audio-visual question answering on videos.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Avqa: A dataset for audio-visual question answering on videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.276994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.276994Z digest=sha256:d3ad806a210e71d5852c87a0b636e6eb187707db41e201f2cdd4af423dfe0efd

Observation fee2fdf0-5625-4f0b-a12f-fb4d2404f25a · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Movienet: A holistic dataset for movie understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.978977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.284058Z digest=sha256:c7dab26a0795803b0e4cc159ebb88f7d74c1243d90004dde18831eea858ed1d9

Observation c4d13c8f-2283-41b8-ba93-d06bd766d3e0 · outbound

This paper cites A dataset for movie description.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward A dataset for movie description

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.962720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.292387Z digest=sha256:4d312b2098f12a34a3f8fe697da9347215e6356bb6ced3881326ab8a16e7ac23

Observation 09d3fbeb-410e-4146-9ec5-ad05891eb23f · outbound

This paper cites Auroracap: Efficient, performant video detailed captioning and a new benchmark.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Auroracap: Efficient, performant video detailed captioning and a new benchmark

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.946632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.300777Z digest=sha256:40d985520e4ea4b1363ac932b08c938a888828fef1218ae4712fef95283e2632

Observation cdd33149-e7cc-4cc7-89a3-72bc5fc4f766 · outbound

This paper cites Time-r1: Post-training large vision language model for temporal video grounding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Time-r1: Post-training large vision language model for temporal video grounding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.929111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.313655Z digest=sha256:133ca5cc2da9bd46ff5e12a4b551680a311c5f82a592886417a3ca7fd66760fd

Observation 41f95256-3d3a-4c1f-a884-73c828d10848 · outbound

This paper cites Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:13:44.864443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.320002Z digest=sha256:f07f28234101a49bc36c2ec0c25771ac399580051187f474f881551c039e04fa

Observation d6ea50e2-06d0-4b15-981a-95988c63ed04 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.329322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.329322Z digest=sha256:f804f2da093bd3a516dcf4cadc39dcfe69bcce9f8904331aec2f14f95a5e1790

Observation 324bbfe1-68ca-488a-968a-de6c43b9bfcf · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1 (2):3, 2022.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Lora: Low-rank adaptation of large language models.ICLR, 1 (2):3, 2022

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.336856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.336856Z digest=sha256:d1e493547376963c1b4ee33f5d604d7180d9e4fd12d749afa74c80e5ed8f4617

Observation fa408bef-362d-41bb-b75f-70b2b82f91e1 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Efficient memory management for large language model serving with pagedattention

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.342904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.342904Z digest=sha256:a04a8695327735f68f95e34c98c2a90592bb084798a552bee7e016ddf42cc797

Observation 455e53d5-8b79-462a-8263-e33b9bceb9df · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Vbench: Comprehensive benchmark suite for video generative models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.369645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.369645Z digest=sha256:c42e4726245b44add8b49fa45a91cb538e368bdf272a385e95a02147fc54b3ad

Observation 367cf444-21a9-45a3-bae8-ce99e94b6814 · outbound

This paper cites VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.376026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.376026Z digest=sha256:c65279797352a54fc29ca6b5a51dc921a75eff3049c04d29de566ae7f2240828

Observation 1074ca18-4c48-469d-a37d-db46b91920bc · outbound

This paper cites HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.385618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.385618Z digest=sha256:797b4511a02fcac6e321f270a298f444f5044063da9a80487df392d41248a5e6

Observation 8aeb4fe0-9bc3-4f5c-9714-ed4492a388df · outbound

This paper cites ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.392418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.392418Z digest=sha256:918ff4715684d8d4d52d292783acbf2c2632a4d87e2e4f867be059104700aa36

Observation d59819bb-9364-46b4-a338-0f3746d8b9a2 · outbound

This paper cites Minicpm-o 2.6: A gpt-4o level mllm for vision, speech, and multimodal live stream- ing on your phone.https://github.com/OpenBMB/MiniCPM-V, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Minicpm-o 2.6: A gpt-4o level mllm for vision, speech, and multimodal live stream- ing on your phone.https://github.com/OpenBMB/MiniCPM-V, 2025

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.876062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.399890Z digest=sha256:16acc9d3066ac35082cf30a3a05ac3fed4306770c65a4e6476ad8740263928d2

Observation dc9de645-8519-4180-9fae-51dfd07e6618 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.406841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.406841Z digest=sha256:50ad4677771356701d86cc771da8313af891a9ec1bfee856c82b4144c6d3fc12

Observation be280df9-dcdb-4e6d-9846-4a4242ff3b46 · outbound

This paper cites Self- critical sequence training for image captioning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Self- critical sequence training for image captioning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.414464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.414464Z digest=sha256:aef1678c89185d125e53debb88f071355410158148487db0c7119af6aee3aded

Observation c49af00d-fb79-4ee8-ac0e-6cd6f366ea28 · outbound

This paper cites id": "sample id.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward id": "sample id

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.848800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.422217Z digest=sha256:5d67f43592088f1a1ab494c7f92fbe5cfd7ebf22ee961614605158f2812c3b23

Observation da44a066-d0cc-4715-a367-9fcfba7c58c0 · outbound

This paper cites Use the exact terminology found in the text (e.g., if the caption says ”shatters”, use ”shatters”, not ”breaks”).

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Use the exact terminology found in the text (e.g., if the caption says ”shatters”, use ”shatters”, not ”breaks”)

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.829527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.429120Z digest=sha256:335dac4ebd70ec0ad8f778a8b81a42fc4dadc26e93b07acea36f9e086e5cefa2

Observation 31b2fb4b-b1be-4066-88df-0ed9d35bbe6f · outbound

This paper cites an unresolved cited work.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:13:46.809850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.441284Z digest=sha256:67568ef9c96bcf57ca12bc198bf8353481131428435d05da17f6f11d428ceccf

Observation 06ff5225-4b67-4407-a027-13511abd61a4 · outbound

This paper cites Start directly with the answer content.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Start directly with the answer content

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.788962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.450979Z digest=sha256:345f646d5a1d6b19b67075a2c09ab27c39efcf831bcc5ce6fa3a15bc02442f00

Observation fa9669e6-ff5d-45ad-bb6c-b38be7b63981 · outbound

This paper cites strengths.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward strengths

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.764825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.457246Z digest=sha256:6ecb65720b2e0743eef616db63b766e4e6a199ec77045098a5cbb5f75ff8740d

Observation 4a20e370-b5df-4187-8a70-e124cda1f609 · outbound

This paper cites an unresolved cited work.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:13:46.742112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.464484Z digest=sha256:e384c2864e466112e8c4199ac34bc1923b28204ec0500ef05f51783d696a4597

Observation bee123fd-afc5-43ed-bb1d-2bd4514bc827 · outbound

This paper cites evaluation.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward evaluation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.717454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.473015Z digest=sha256:16d3843913f6c938fb25d9bae2f06c98dc1f6d14db7426cd095977b5d571b1c5

Observation 6f54366b-2833-4bff-a4bb-38ad54d5fcc9 · outbound

This paper cites Do not add any artistic interpretation, subjective analysis, or infer any character’s internal 23 thoughts, emotions, or intentions beyond what is explicitly visible or audible.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Do not add any artistic interpretation, subjective analysis, or infer any character’s internal 23 thoughts, emotions, or intentions beyond what is explicitly visible or audible

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.692726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.480978Z digest=sha256:44ae64f8dd530ec98cca1c2dd1eedd566adca11d8a75d24a7f5c729ab2a419b1

Observation 90a1999f-d774-4123-a162-e1bbae85ba49 · outbound

This paper cites While Character A is speaking, Character B simultaneously turns their head.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward While Character A is speaking, Character B simultaneously turns their head

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.668462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.486950Z digest=sha256:9e1a0881f6096e7a5502e34109a3ae3a5125406e42afe71d24add757f4e08f13

Observation 36b615f7-5d65-48b6-b411-84544350326a · outbound

This paper cites He closed the door,.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward He closed the door,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.648814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.494599Z digest=sha256:8377a32d27c174c022cbbb86f43abe1b859ce5d7ff03b6bda5994e2692d8b2d1

Observation a39cd8a8-82c3-4890-ab62-d46deb18892b · outbound

This paper cites The camera cuts to a close-up of her face,.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward The camera cuts to a close-up of her face,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.628650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.502462Z digest=sha256:590e51eb489f656c2149dcd323593c05a8859256a99fd43ed4a89dbfe3d868b1

Observation 52b76ef2-21c0-4633-a1fe-87058085ec5f · outbound

This paper cites an unresolved cited work.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:13:46.608847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.516642Z digest=sha256:b53d8bd4493bcec3920fa9627a39c8abf9a8e70f836d44f3f2f5a97e3d501bf3

Observation e8eebfdd-7bb2-4176-9ed5-f3a690bb0f80 · outbound

This paper cites Do not use bullet points, headings, or line breaks within the narrative.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Do not use bullet points, headings, or line breaks within the narrative

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.577670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.523799Z digest=sha256:bb51401be7b62cac172db7e9bc480e95609d50313ed0aa3c779529915b51e399

Observation 70b54954-8d96-4140-823a-14fdffbda128 · outbound

This paper cites His voice, thick with sarcasm, says.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward His voice, thick with sarcasm, says

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.553912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.531066Z digest=sha256:997feee7dd1f3266d1ecb8ad46658c53df42a3840e304508b1fa210a89884566

Observation 90c315f3-909b-46c3-93ad-2d8bb8a5336c · outbound

This paper cites thud" as a book hits the table, a.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward thud" as a book hits the table, a

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.526057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.543586Z digest=sha256:540175deef0b76e9888b821c2f0e7cd266c4cadc06813dbd0c4ba6ffac373312

Observation 64b2eb73-aa40-48d3-8df0-360d4fdca671 · outbound

This paper cites While Person A’s hand reaches for the glass, Person B’s eyes dart to the side, and the music swells.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward While Person A’s hand reaches for the glass, Person B’s eyes dart to the side, and the music swells

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.499941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.557179Z digest=sha256:d9dc9bd41cfaa70520f3dc207625e4d9220355e9ef688d799f4d9c4b896d9bed

Observation 0092d044-e526-435f-8cb0-7f3df378f2fb · outbound

This paper cites Avoid summarizing.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Avoid summarizing

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.471133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.563308Z digest=sha256:e74f31798bacaeb7a11fa86bd6616426b3fc72e1d2926c058c2962951b4b3fd4

Observation 8457dc9e-0c4b-4dcb-8fe2-d16d504a83a2 · outbound

This paper cites strengths.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward strengths

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.445294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T18:13:44.573450Z digest=sha256:4c834df499b564290cc9a00f43cc9b46e3fff1b824fb6f35f98c1f64ba3ab944

Observation 271c8407-3349-4ac9-8db6-390573dd5ff2 · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.362349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.362349Z digest=sha256:fb3d91a361d3cdf382e0f6c998aa3efb372f04de0a8315e96182cec19a376c6e

Pith citing papers

No inbound Pith citation observations are available.