Pith. sign in

Paper Citation Record · LEDGER

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward

As of 12 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2608.06930.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06930 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:13:44.573450Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact2
  • verified fuzzy26
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4eaed5f7-1f2a-4ff8-8f8f-92924844fd08 · outbound

This paper cites Avocado: An audiovisual video captioner driven by temporal orchestration.arXiv preprint arXiv:2510.10395, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Avocado: An audiovisual video captioner driven by temporal orchestration.arXiv preprint arXiv:2510.10395, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.001303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.001303Z digest=sha256:6de44b0f55b113be19afaa5f22b9cda51a58d9923905f02cf1d8a4b5e1882dc5

Observation b695bdef-a788-4514-949d-8065a0a98e1f · outbound

This paper cites Ugc-videocaptioner: An omni ugc video detail caption model and new benchmarks.arXiv preprint arXiv:2507.11336, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Ugc-videocaptioner: An omni ugc video detail caption model and new benchmarks.arXiv preprint arXiv:2507.11336, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.012944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.012944Z digest=sha256:4dd8c43ca1923fe141f8c1a041efaf35bdefaab99de85312d88db62b209dfe4c

Observation ad1b155c-49ef-45af-b7b2-78dea5511969 · outbound

This paper cites Daily-omni: Towards audio-visual reasoning with temporal alignment across modalities.arXiv preprint arXiv:2505.17862, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Daily-omni: Towards audio-visual reasoning with temporal alignment across modalities.arXiv preprint arXiv:2505.17862, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.020242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.020242Z digest=sha256:67b6bc13d141b9c46c62dbcdfd32a8a58978cc1c9033635a2aa5f2d4e3f6ad90

Observation f454879b-db8d-4894-b6a7-98084bef545c · outbound

This paper cites VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.027540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.027540Z digest=sha256:df9b484434679679dec0ac7ab58f59f9ffdd42d702c14475fbd95909711eb546

Observation 8f6e4ae1-519f-4569-be74-61cf0eed85a9 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37: 19472–19495, 2024.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37: 19472–19495, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.162755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.035846Z digest=sha256:8cc6dce069fb958408bc68c660e332f9deadf0b25f84bb62ba577d60dcc68575

Observation 55a344b0-2078-444f-9837-3581bfcb57db · outbound

This paper cites Mavors: Multi-granularity video representation for multimodal large language model.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mavors: Multi-granularity video representation for multimodal large language model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.043255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.043255Z digest=sha256:c197a8c942c9e9cf29c8c34ccd3eb78d04278efff1531ba94d8ef4febcf6fced

Observation 5c73a9e0-e967-446c-a679-d61cfd231a08 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.051144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.051144Z digest=sha256:d52989144f1ecf60c957ee4a4e051d3ab47ed944dc1495c61c4ab5deb2a30b9b

Observation 90db9274-bbf3-46ad-8514-45a615392547 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.059127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.059127Z digest=sha256:a9e8c4eb0378ce5285e2669d43a650e1af9c4dabbb15444e1ac7d752a1f0badc

Observation b84cf9f9-f405-441b-9a80-a8a0db9a6853 · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.066103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.066103Z digest=sha256:7ada787be14939b080745ae8c8becd685118ac1e10915e11805459fce86b612e

Observation ac750a63-35b8-467a-9019-cfa9d9a1d905 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.074176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.074176Z digest=sha256:a4098a3720766ac117eb1d5aea913bfaf7d1661da2f09a6e28f87c354d533320

Observation 727dae58-6d8b-4478-bb13-886fcd349312 · outbound

This paper cites Advancing high-resolution video-language representation with large-scale video transcriptions.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Advancing high-resolution video-language representation with large-scale video transcriptions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.108399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.080659Z digest=sha256:b9438c4c7f4b8365e402877acba33401540dc89fc9d5ab41ea3eb506d5b97658

Observation 6995e3e9-3e81-4806-af86-eef40d176bdb · outbound

This paper cites video-salmonn 2: Captioning-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward video-salmonn 2: Captioning-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.087226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.087226Z digest=sha256:41dece32c6e09b8e7991a097b8593822eba6f3e5e59155ce8f3753afc1448d0a

Observation 7f752d0b-4aad-460d-9aba-c565c4d88bb7 · outbound

This paper cites AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.093438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.093438Z digest=sha256:42c7db22e55bc02622ba56fa5905b3d5ddc323c280bbebf54f0cec543da8aece

Observation 70ff3cbf-52e2-4513-b5d7-e931862558b8 · outbound

This paper cites Longvale: Vision-audio-language-event benchmark towards time-aware omni-modal perception of long videos.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Longvale: Vision-audio-language-event benchmark towards time-aware omni-modal perception of long videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.086591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.099582Z digest=sha256:88d430e5e8ed7ab5a41d9f3b480378f2cbf3838eb61f2b8a1f795c71bfe7d04f

Observation 8db4405e-fa54-4f54-95d9-ff440f26235c · outbound

This paper cites Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.106889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.106889Z digest=sha256:1c1cc3e44c040e42d9b6eccca60ac4f379fd21fc54496997f2a7e38ea052e617

Observation a402971b-163d-41f1-bf6a-5779e77e4a33 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.113320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.113320Z digest=sha256:b6b9c6f015713c8ae222318383125f7cc640f757d2621327f4ba1ef50d71f0bf

Observation fafedc81-20fb-4608-872c-4b5035d3bde9 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.120212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.120212Z digest=sha256:c763e812eeebe93a22f9f2783ac44e9ef5be9263486a163926b76eb5213cd899

Observation 87b8dd1b-90a8-4b20-bc27-6bb85e24b4a2 · outbound

This paper cites Watch, listen, and describe: Globally and locally aligned cross-modal attentions for video captioning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Watch, listen, and describe: Globally and locally aligned cross-modal attentions for video captioning

Reference 18

Resolution
verified exact
doi, observed 2026-08-10T18:13:44.635267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.126805Z digest=sha256:bfce08568e8965ca8ef8aba5c7180214601f43a701618606d60c0b6c7c7e27fa

Observation e10c871d-faef-49c8-bac8-ad08c41aec11 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.133036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.133036Z digest=sha256:227237245934e29b733bca357ab08b87b080935d9e4cbd874bb9a03e8da18627

Observation 3a17029f-7952-4eac-a5f8-50f77b513d43 · outbound

This paper cites Qwen2.5-VL Technical Report.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Qwen2.5-VL Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.140160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.140160Z digest=sha256:61df8a2971fe07adb1ed4a64ac2e2b6a81f7bca0161be68aa736cfeec307a8a0

Observation 2d47876c-b084-4857-a211-94c6eba7760b · outbound

This paper cites Glave-cap: Global-local aligned video captioning with vision expert integration.arXiv preprint arXiv:2509.11360, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Glave-cap: Global-local aligned video captioning with vision expert integration.arXiv preprint arXiv:2509.11360, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.146443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.146443Z digest=sha256:21641f2eb2dd3635c3e7ca2c21e6593cf2bb9d3ec2b9ed0838f4f1af6e341f6d

Observation f8f1a29e-aef0-4ac8-9db6-e8e868898a3a · outbound

This paper cites Qwen2.5-Omni Technical Report.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Qwen2.5-Omni Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.151110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.151110Z digest=sha256:dedf753d497dd6a9d25314119089c410acde23d94858399f7760e910ae5daa03

Observation ead60c4d-46c1-4950-977f-efa478de054f · outbound

This paper cites OmniCaptioner: One Captioner to Rule Them All.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward OmniCaptioner: One Captioner to Rule Them All

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.157674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.157674Z digest=sha256:6ffdc38656c892dd0a54c6a0ee9dd73d95efd6b003f8b19d5045a7fe8784de2a

Observation d1c38d65-50d3-4629-b59f-5b52c5a203e4 · outbound

This paper cites Omni-captioner: Data pipeline, models, and benchmark for omni detailed perception.arXiv preprint arXiv:2510.12720, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Omni-captioner: Data pipeline, models, and benchmark for omni detailed perception.arXiv preprint arXiv:2510.12720, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.163332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.163332Z digest=sha256:179612d3afd8a7709381135b085b316acddbdd4020b5e99cec107064975fae3a

Observation aab17aee-273f-41a8-9111-9ea95809ed00 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.169838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.169838Z digest=sha256:a56ec44a8610ef68345ffd9cd3aee8e91d37387a6f7cfed0b59f21a1b68f3d73

Observation 91ab4e8d-8701-43e7-a6f1-180c2e0142ef · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.176917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.176917Z digest=sha256:8404faa641781dcccb6eb59ef3f66d3f009afc1da7744df86f28de647da6d7fe

Observation 1dfeba4d-990f-4466-961f-7991cdc82a8b · outbound

This paper cites AdaTooler-V: Adaptive Tool-Use for Images and Videos.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward AdaTooler-V: Adaptive Tool-Use for Images and Videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.183615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.183615Z digest=sha256:5a2b887a5b7df073cf17798bf6c921d105cd943b37d13e3f7468460e713ab45d

Observation 0e486e27-48ba-4008-87db-37a63dcf1e18 · outbound

This paper cites Deepvideo-r1: Video rein- forcement fine-tuning via difficulty-aware regressive grpo.arXiv preprint arXiv:2506.07464, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Deepvideo-r1: Video rein- forcement fine-tuning via difficulty-aware regressive grpo.arXiv preprint arXiv:2506.07464, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.190682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.190682Z digest=sha256:076948e3789c48d13e55a1c542b2843ebef44233b52a873637a0735ae6f67249

Observation 89671b3f-be56-4556-bfb1-d94ce96e8412 · outbound

This paper cites Editthinker: Unlocking iterative reasoning for any image editor.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Editthinker: Unlocking iterative reasoning for any image editor

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.197366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.197366Z digest=sha256:deac305ebee29d2efd6c6e45f836f4b7e72b9b07a3539328c29787d5352e4f3b

Observation 7e3a51c8-22b5-4d13-a606-cc7a0e510022 · outbound

This paper cites Sophiavl-r1: Reinforcing mllms reasoning with thinking reward.arXiv preprint arXiv:2505.17018, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Sophiavl-r1: Reinforcing mllms reasoning with thinking reward.arXiv preprint arXiv:2505.17018, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.204027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.204027Z digest=sha256:ad07a33e1dfe4497fb85aecc7eab79da42e2c180d5b691192d13954dea9575b9

Observation 3ec3c7c0-87f4-447b-885a-edc9ca334b31 · outbound

This paper cites Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.209331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.209331Z digest=sha256:8c0da3e0e6b8944facbeee10851dc4cab330727bdb8ad6246b7596d84f3ba7c4

Observation cf7d5b55-0444-40c4-aacb-a628a9681f73 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.215349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.215349Z digest=sha256:e97c6723e561f577f0693d743d18e6b5bcbe52d0eff64212e30e6a48d886828c

Observation 65539722-4660-4024-be93-223a12fa8d76 · outbound

This paper cites OneThinker: All-in-one Reasoning Model for Image and Video.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward OneThinker: All-in-one Reasoning Model for Image and Video

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.221576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.221576Z digest=sha256:157f872fae45c319d8e366979b8c0dfd6c4fd7a3def70c5d99a6db0298f9cfe2

Observation 2eb3994f-034a-49ca-a654-13259944ce9d · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.228889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.228889Z digest=sha256:6be244b7c9f53f5cc5edddd1ded555f990bd433ba09273a67dc8fee0befd8c0f

Observation 53fcffc3-aa5f-49b5-ac92-20dd6ee28176 · outbound

This paper cites VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.235558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.235558Z digest=sha256:1f0ca04e5d39f0741546e5c3c96203bb4da71998b6d0f6b182b5068d4cd300e7

Observation d031b374-8983-4f7e-a325-cc8741da424e · outbound

This paper cites Exploring the role of audio in video captioning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Exploring the role of audio in video captioning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.067484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.241601Z digest=sha256:4dce1714426fc12495b18920eb09fad58406c351e64692b058b5ac7e3c826f62

Observation 81e8b0ea-9029-44e0-8c6e-06c0df6bdacb · outbound

This paper cites Hybrid transformers for music source separation.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Hybrid transformers for music source separation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.049145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.247461Z digest=sha256:c6214cad1d4d11e947d995c385a9a0cfba1591d37a7bfbfe438e6fef4a600a22

Observation cc0ed116-12d4-4b94-929f-bb14edf82074 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Audio-visual event localization in unconstrained videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.255613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.255613Z digest=sha256:a9b9e4f6ec8da37a5aceb1b72ca910445af3d68b7710b62bc768e99dc6f3b574

Observation f3944060-20a1-4e43-ac79-d894f5e3ef54 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Vggsound: A large-scale audio-visual dataset

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.018344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.261995Z digest=sha256:bcf342bca4e502e7fa7bd09719dcae055e31c667cb83facb88fbfa12074dbd3e

Observation 3ca62a8c-3c7f-419f-b375-76f43901be93 · outbound

This paper cites Condensed movies: Story based retrieval with contextual embeddings.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Condensed movies: Story based retrieval with contextual embeddings

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.268021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.268021Z digest=sha256:a8373fc5c9b2c28f4a1ffafd0e75a3037fe90eb630d5880809f1cab1672deb57

Observation a4d9dd4f-6e26-4315-92f0-725f21550be2 · outbound

This paper cites Avqa: A dataset for audio-visual question answering on videos.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Avqa: A dataset for audio-visual question answering on videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.276994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.276994Z digest=sha256:56f2e94a04da274946a338a30fe7d2616613a86523762da1311405e4c0a5c777

Observation fee2fdf0-5625-4f0b-a12f-fb4d2404f25a · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Movienet: A holistic dataset for movie understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.978977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.284058Z digest=sha256:69d98272d449f44b116b1569cacce13cf03f16198d09cb67e58c9912d5983dae

Observation c4d13c8f-2283-41b8-ba93-d06bd766d3e0 · outbound

This paper cites A dataset for movie description.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward A dataset for movie description

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.962720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.292387Z digest=sha256:b72678f87d087a85ee1c6e306fec5459fb495191b9cfbb4d547cb878fdc9ff9b

Observation 09d3fbeb-410e-4146-9ec5-ad05891eb23f · outbound

This paper cites Auroracap: Efficient, performant video detailed captioning and a new benchmark.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Auroracap: Efficient, performant video detailed captioning and a new benchmark

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.946632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.300777Z digest=sha256:3509531ddef8952644570d55154ff43dcca1a5c6336d4e332127a638aa3c7b91

Observation cdd33149-e7cc-4cc7-89a3-72bc5fc4f766 · outbound

This paper cites Time-r1: Post-training large vision language model for temporal video grounding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Time-r1: Post-training large vision language model for temporal video grounding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.929111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.313655Z digest=sha256:d186cfeaecca7d5ee69e895eee86df9a5c5676175689e4385c239237451cfb91

Observation 41f95256-3d3a-4c1f-a884-73c828d10848 · outbound

This paper cites Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:13:44.864443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.320002Z digest=sha256:601fc751c2920d92ff3f0baabca69ddcfef4d5e628357a10bec2cca76fd3371b

Observation d6ea50e2-06d0-4b15-981a-95988c63ed04 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.329322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.329322Z digest=sha256:683f6c1704438aece27cf80cfb470b934038de97bc4e5ef70a6bf04f037c490d

Observation 324bbfe1-68ca-488a-968a-de6c43b9bfcf · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1 (2):3, 2022.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Lora: Low-rank adaptation of large language models.ICLR, 1 (2):3, 2022

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.336856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.336856Z digest=sha256:76309f1aa5a26f44a578fc99b4d6aaa2b8db488566cf640653d3c0caee20cfb2

Observation fa408bef-362d-41bb-b75f-70b2b82f91e1 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Efficient memory management for large language model serving with pagedattention

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.342904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.342904Z digest=sha256:8d6e50b2a5d24d55d8273b7cff43dd32b71113c883d028e994ab2b5f43422e6f

Observation 455e53d5-8b79-462a-8263-e33b9bceb9df · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Vbench: Comprehensive benchmark suite for video generative models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.369645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.369645Z digest=sha256:419d5ca325e01961a1b5a2fea6f24a81702b8047dea90dd2e4bde9cadc73e555

Observation 367cf444-21a9-45a3-bae8-ce99e94b6814 · outbound

This paper cites VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.376026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.376026Z digest=sha256:c310e3aab3460f8600638290f62a925333a105feb3e4a16e001cfd0946eeae40

Observation 1074ca18-4c48-469d-a37d-db46b91920bc · outbound

This paper cites HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.385618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.385618Z digest=sha256:f96b8eaebc916055c9c5cf562c8f5835e2a4070f01973a9fea9d1b1d2a2fd43d

Observation 8aeb4fe0-9bc3-4f5c-9714-ed4492a388df · outbound

This paper cites ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.392418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.392418Z digest=sha256:0d8ed1b9a8f19107e31e88f88940d03af3aed2525755e8c487c576ef063d62f1

Observation d59819bb-9364-46b4-a338-0f3746d8b9a2 · outbound

This paper cites Minicpm-o 2.6: A gpt-4o level mllm for vision, speech, and multimodal live stream- ing on your phone.https://github.com/OpenBMB/MiniCPM-V, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Minicpm-o 2.6: A gpt-4o level mllm for vision, speech, and multimodal live stream- ing on your phone.https://github.com/OpenBMB/MiniCPM-V, 2025

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.876062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.399890Z digest=sha256:300c90d67d74e1c8b04c7c45833b283c2c0463f0e0c5297a11409747a5b7dff6

Observation dc9de645-8519-4180-9fae-51dfd07e6618 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.406841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.406841Z digest=sha256:2778b259bb30801a7e81c69b337584cb0335a7fa25137ee464bfc6b5868a0311

Observation be280df9-dcdb-4e6d-9846-4a4242ff3b46 · outbound

This paper cites Self- critical sequence training for image captioning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Self- critical sequence training for image captioning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.414464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.414464Z digest=sha256:d1ff7263a7843b594039c75a64f098f70886075e26c9ca56734121904d4f6b7d

Observation c49af00d-fb79-4ee8-ac0e-6cd6f366ea28 · outbound

This paper cites id": "sample id.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward id": "sample id

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.848800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.422217Z digest=sha256:4b2a5fc90e2b8721621284bc8606774e5571b620b306ceffd7f2637f8f25b612

Observation da44a066-d0cc-4715-a367-9fcfba7c58c0 · outbound

This paper cites Use the exact terminology found in the text (e.g., if the caption says ”shatters”, use ”shatters”, not ”breaks”).

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Use the exact terminology found in the text (e.g., if the caption says ”shatters”, use ”shatters”, not ”breaks”)

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.829527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.429120Z digest=sha256:2ef31b5705c71c756cdea392f51630118903627091b0fbf55e82ed61b935c9f4

Observation 31b2fb4b-b1be-4066-88df-0ed9d35bbe6f · outbound

This paper cites an unresolved cited work.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:13:46.809850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.441284Z digest=sha256:7a20b3ea8dd046a276967aa7fac1e4bfe814be2b8840cc479898f8ef7ddf54cf

Observation 06ff5225-4b67-4407-a027-13511abd61a4 · outbound

This paper cites Start directly with the answer content.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Start directly with the answer content

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.788962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.450979Z digest=sha256:ae7da5635fab5dd4368c7912d4982b920b94abdc56d06fcb2c61e06a0ca08d98

Observation fa9669e6-ff5d-45ad-bb6c-b38be7b63981 · outbound

This paper cites strengths.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward strengths

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.764825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.457246Z digest=sha256:db77ce3aa0011c62f5a4fbd3c69b3c80c5d6b485d9d474735293f1e0894cdbe0

Observation 4a20e370-b5df-4187-8a70-e124cda1f609 · outbound

This paper cites an unresolved cited work.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:13:46.742112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.464484Z digest=sha256:260929f83571cac7eae8480c343d86a5005d28238e4e16bedb5a246914a6eddf

Observation bee123fd-afc5-43ed-bb1d-2bd4514bc827 · outbound

This paper cites evaluation.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward evaluation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.717454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.473015Z digest=sha256:767a56c529bde1e78da8c339a561965b2947d05c7eb2f4735c3983700fdd2a51

Observation 6f54366b-2833-4bff-a4bb-38ad54d5fcc9 · outbound

This paper cites Do not add any artistic interpretation, subjective analysis, or infer any character’s internal 23 thoughts, emotions, or intentions beyond what is explicitly visible or audible.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Do not add any artistic interpretation, subjective analysis, or infer any character’s internal 23 thoughts, emotions, or intentions beyond what is explicitly visible or audible

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.692726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.480978Z digest=sha256:2ec55462d12161727a24833dc7f1bcd8d6b67312b921d0d077133bdd0a588a5f

Observation 90a1999f-d774-4123-a162-e1bbae85ba49 · outbound

This paper cites While Character A is speaking, Character B simultaneously turns their head.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward While Character A is speaking, Character B simultaneously turns their head

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.668462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.486950Z digest=sha256:6165772e791c35a499c5b0d3660dcec4d50aa27f584aef4667d90cf2afd35b4b

Observation 36b615f7-5d65-48b6-b411-84544350326a · outbound

This paper cites He closed the door,.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward He closed the door,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.648814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.494599Z digest=sha256:8e6d3a55d6d69ebd898239668aa7ca9289cf43d322c80e4444445937dab03a75

Observation a39cd8a8-82c3-4890-ab62-d46deb18892b · outbound

This paper cites The camera cuts to a close-up of her face,.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward The camera cuts to a close-up of her face,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.628650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.502462Z digest=sha256:a13af697a7b567ac140db6c5aab9742fdbc675ddae49560554189c2994b984f3

Observation 52b76ef2-21c0-4633-a1fe-87058085ec5f · outbound

This paper cites an unresolved cited work.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:13:46.608847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.516642Z digest=sha256:1fbb3ee2078affd073e28b1ba206d2165ea06300534fddc7926865e5c8c65bed

Observation e8eebfdd-7bb2-4176-9ed5-f3a690bb0f80 · outbound

This paper cites Do not use bullet points, headings, or line breaks within the narrative.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Do not use bullet points, headings, or line breaks within the narrative

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.577670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.523799Z digest=sha256:f4a407f0509c34f32bccd521e8deed811c1ff1dbbbea2e599cced72007221dd0

Observation 70b54954-8d96-4140-823a-14fdffbda128 · outbound

This paper cites His voice, thick with sarcasm, says.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward His voice, thick with sarcasm, says

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.553912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.531066Z digest=sha256:e90c911eae81ce121e25ae7ea2b92d9048042be70b65ec75e652eaeaf55bd10e

Observation 90c315f3-909b-46c3-93ad-2d8bb8a5336c · outbound

This paper cites thud" as a book hits the table, a.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward thud" as a book hits the table, a

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.526057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.543586Z digest=sha256:df7d407876e8cee36da9ca86e15713171d7c3282bf2db94fbf64315040db836a

Observation 64b2eb73-aa40-48d3-8df0-360d4fdca671 · outbound

This paper cites While Person A’s hand reaches for the glass, Person B’s eyes dart to the side, and the music swells.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward While Person A’s hand reaches for the glass, Person B’s eyes dart to the side, and the music swells

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.499941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.557179Z digest=sha256:19b48507136b9f2a139414a1811566c0f4257e92196b397852542dfa8b603510

Observation 0092d044-e526-435f-8cb0-7f3df378f2fb · outbound

This paper cites Avoid summarizing.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Avoid summarizing

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.471133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.563308Z digest=sha256:171c0abba0bad688be5f68235e3d0da3ccff637bb3c09f7c41e9a44993e2180f

Observation 8457dc9e-0c4b-4dcb-8fe2-d16d504a83a2 · outbound

This paper cites strengths.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward strengths

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.445294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T18:13:44.573450Z digest=sha256:a2b92ea08fdedbab6088a861f097b9899deb9d0d63bb475f292b256905ef0478

Observation 271c8407-3349-4ac9-8db6-390573dd5ff2 · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.362349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.362349Z digest=sha256:3e0358c628fd32c37ef3f75fc2cd9d2312afc8d5bdb9d628f41eee84558802e0

Pith citing papers

No inbound Pith citation observations are available.