Pith. sign in

Paper Citation Record · LEDGER

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward

As of 22 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2608.06930.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06930 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:13:44.573450Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact2
  • verified fuzzy26
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4eaed5f7-1f2a-4ff8-8f8f-92924844fd08 · outbound

This paper cites Avocado: An audiovisual video captioner driven by temporal orchestration.arXiv preprint arXiv:2510.10395, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Avocado: An audiovisual video captioner driven by temporal orchestration.arXiv preprint arXiv:2510.10395, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.001303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.001303Z digest=sha256:cd46652290a2766b7bd18df2c402c94d0190fa6ada438d39dfd731a50bb52c43

Observation b695bdef-a788-4514-949d-8065a0a98e1f · outbound

This paper cites Ugc-videocaptioner: An omni ugc video detail caption model and new benchmarks.arXiv preprint arXiv:2507.11336, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Ugc-videocaptioner: An omni ugc video detail caption model and new benchmarks.arXiv preprint arXiv:2507.11336, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.012944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.012944Z digest=sha256:340d21e1b11b6dee5a667a768a6fd83a23251867e702d3e79688b217d2042712

Observation ad1b155c-49ef-45af-b7b2-78dea5511969 · outbound

This paper cites Daily-omni: Towards audio-visual reasoning with temporal alignment across modalities.arXiv preprint arXiv:2505.17862, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Daily-omni: Towards audio-visual reasoning with temporal alignment across modalities.arXiv preprint arXiv:2505.17862, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.020242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.020242Z digest=sha256:9d619891c94e0a23126d45f3d58c80898087bf60bf5c288d9dc507cc5af3f1b6

Observation f454879b-db8d-4894-b6a7-98084bef545c · outbound

This paper cites VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.027540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.027540Z digest=sha256:7175137d86f5b645e4ed1b316267d35d314115721478e7ea562fac3a955086cd

Observation 8f6e4ae1-519f-4569-be74-61cf0eed85a9 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37: 19472–19495, 2024.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37: 19472–19495, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.162755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.035846Z digest=sha256:d88a44ad91424c0ffc2962d600dce259b733fe4755ba3d8a3e275256d030a570

Observation 55a344b0-2078-444f-9837-3581bfcb57db · outbound

This paper cites Mavors: Multi-granularity video representation for multimodal large language model.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mavors: Multi-granularity video representation for multimodal large language model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.043255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.043255Z digest=sha256:e9ac5f91e88bc67fdceed85617eeaa8279ccab3c0415b4ff1ea0efdbb3c7b1cf

Observation 5c73a9e0-e967-446c-a679-d61cfd231a08 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.051144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.051144Z digest=sha256:7e03f158b5a16eeee6442eba36dea7b3a123ea47679af30239a2595a0bac6995

Observation 90db9274-bbf3-46ad-8514-45a615392547 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.059127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.059127Z digest=sha256:32781bea327ceae0452f5832f90ea76322d1ac4a71fb6e3e27ab359736b763aa

Observation b84cf9f9-f405-441b-9a80-a8a0db9a6853 · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.066103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.066103Z digest=sha256:b6a410145a3507590d37dce6390fb27d530df187abcd8e6b31933bc6ca186617

Observation ac750a63-35b8-467a-9019-cfa9d9a1d905 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.074176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.074176Z digest=sha256:3cfc8d7105077be37c9c30dfc2ea25c01d426be395852dd3bdae80a38c74c1a1

Observation 727dae58-6d8b-4478-bb13-886fcd349312 · outbound

This paper cites Advancing high-resolution video-language representation with large-scale video transcriptions.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Advancing high-resolution video-language representation with large-scale video transcriptions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.108399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.080659Z digest=sha256:57b037bbc7e101ab92014b57b60f115b83def8dd55f7f9fe45e7f1a99d51f739

Observation 6995e3e9-3e81-4806-af86-eef40d176bdb · outbound

This paper cites video-salmonn 2: Captioning-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward video-salmonn 2: Captioning-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.087226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.087226Z digest=sha256:9d012beb83c641ab58036e0bc5049aa0953a8f44170332923b7909747aa44e50

Observation 7f752d0b-4aad-460d-9aba-c565c4d88bb7 · outbound

This paper cites AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.093438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.093438Z digest=sha256:9abdb51dcc8dea88e1d856309b363814a012566bed32ce230d9e05a57b37bcdf

Observation 70ff3cbf-52e2-4513-b5d7-e931862558b8 · outbound

This paper cites Longvale: Vision-audio-language-event benchmark towards time-aware omni-modal perception of long videos.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Longvale: Vision-audio-language-event benchmark towards time-aware omni-modal perception of long videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.086591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.099582Z digest=sha256:f515503afaa786c20c108fe1d7699c150b92f54a340577e2bc2af364090ea082

Observation 8db4405e-fa54-4f54-95d9-ff440f26235c · outbound

This paper cites Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.106889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.106889Z digest=sha256:77e4167f611b7649802a8137f738231d9365c0552c8c17d7d04c7a0dd6a7e060

Observation a402971b-163d-41f1-bf6a-5779e77e4a33 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.113320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.113320Z digest=sha256:74a96f1e54f6c833fc1570c6c33a5584ca178d58305bde16a46dd28282084a6c

Observation fafedc81-20fb-4608-872c-4b5035d3bde9 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.120212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.120212Z digest=sha256:ded66671fa7806fe6ffd537475d1af66ef9d9a7fd03686a9e059f6444835d574

Observation 87b8dd1b-90a8-4b20-bc27-6bb85e24b4a2 · outbound

This paper cites Watch, listen, and describe: Globally and locally aligned cross-modal attentions for video captioning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Watch, listen, and describe: Globally and locally aligned cross-modal attentions for video captioning

Reference 18

Resolution
verified exact
doi, observed 2026-08-10T18:13:44.635267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.126805Z digest=sha256:51c7a1c1e33d830e6be6c3952c47dac1fb4bce879bf7b2e2b0c4be768f723b54

Observation e10c871d-faef-49c8-bac8-ad08c41aec11 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.133036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.133036Z digest=sha256:a58a5cc396e295e059d6b1738b4f22b15a46ec53f1e49f0e95cdcf2eb4fe2890

Observation 3a17029f-7952-4eac-a5f8-50f77b513d43 · outbound

This paper cites Qwen2.5-VL Technical Report.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Qwen2.5-VL Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.140160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.140160Z digest=sha256:3b157e3db6691b3317e7267baf66b8af1588a5c518635adf586f24bc6dd4d84e

Observation 2d47876c-b084-4857-a211-94c6eba7760b · outbound

This paper cites Glave-cap: Global-local aligned video captioning with vision expert integration.arXiv preprint arXiv:2509.11360, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Glave-cap: Global-local aligned video captioning with vision expert integration.arXiv preprint arXiv:2509.11360, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.146443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.146443Z digest=sha256:5f611cd5c408e91e662c556487b3f99b108ffd0a8a284d88fb83f6ee279ec568

Observation f8f1a29e-aef0-4ac8-9db6-e8e868898a3a · outbound

This paper cites Qwen2.5-Omni Technical Report.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Qwen2.5-Omni Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.151110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.151110Z digest=sha256:3a5afd1ad6c274a14a96bda22fa033f3dbdfa6213da6d985c7b8596a7cea025a

Observation ead60c4d-46c1-4950-977f-efa478de054f · outbound

This paper cites OmniCaptioner: One Captioner to Rule Them All.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward OmniCaptioner: One Captioner to Rule Them All

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.157674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.157674Z digest=sha256:ed2861eb78f8cd313903c0bb75167ffdfdd91e7b3ef750230af18341d0074787

Observation d1c38d65-50d3-4629-b59f-5b52c5a203e4 · outbound

This paper cites Omni-captioner: Data pipeline, models, and benchmark for omni detailed perception.arXiv preprint arXiv:2510.12720, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Omni-captioner: Data pipeline, models, and benchmark for omni detailed perception.arXiv preprint arXiv:2510.12720, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.163332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.163332Z digest=sha256:305b8bcc90f25ab569303e29d354e423fb38c3f86b252d74a3621238ae96daba

Observation aab17aee-273f-41a8-9111-9ea95809ed00 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.169838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.169838Z digest=sha256:f47a04ccb4cb645571c463661837c01dca3c6a6b8765c3f71ea76516f99aeaba

Observation 91ab4e8d-8701-43e7-a6f1-180c2e0142ef · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.176917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.176917Z digest=sha256:b7f5687ef7355f2fee4eac5efffe6b939f094001e4eb89d83df85079a8fbbc0a

Observation 1dfeba4d-990f-4466-961f-7991cdc82a8b · outbound

This paper cites AdaTooler-V: Adaptive Tool-Use for Images and Videos.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward AdaTooler-V: Adaptive Tool-Use for Images and Videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.183615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.183615Z digest=sha256:702fba58b68e47748f10be305b760bf49fdc7dffb3374ea8fc179d56c8a75b1e

Observation 0e486e27-48ba-4008-87db-37a63dcf1e18 · outbound

This paper cites Deepvideo-r1: Video rein- forcement fine-tuning via difficulty-aware regressive grpo.arXiv preprint arXiv:2506.07464, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Deepvideo-r1: Video rein- forcement fine-tuning via difficulty-aware regressive grpo.arXiv preprint arXiv:2506.07464, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.190682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.190682Z digest=sha256:f0e2de9f1f279a9ca6dec13bf1bd5e5c001245efd6c5780d95091e70b0d483f7

Observation 89671b3f-be56-4556-bfb1-d94ce96e8412 · outbound

This paper cites Editthinker: Unlocking iterative reasoning for any image editor.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Editthinker: Unlocking iterative reasoning for any image editor

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.197366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.197366Z digest=sha256:354cc9105c83d5271a8b78b15231dd4136d22741ce2db7070cf9b0f77fd7cb15

Observation 7e3a51c8-22b5-4d13-a606-cc7a0e510022 · outbound

This paper cites Sophiavl-r1: Reinforcing mllms reasoning with thinking reward.arXiv preprint arXiv:2505.17018, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Sophiavl-r1: Reinforcing mllms reasoning with thinking reward.arXiv preprint arXiv:2505.17018, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.204027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.204027Z digest=sha256:26b22ef96f3a02fb2135114f278033f6ebff0b300ebf9903d84037ba6c789054

Observation 3ec3c7c0-87f4-447b-885a-edc9ca334b31 · outbound

This paper cites Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.209331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.209331Z digest=sha256:ca4c14426b28b31ec49d1f1ed43a22ea8c08e9c2ee3d55468154170b19dfe6a8

Observation cf7d5b55-0444-40c4-aacb-a628a9681f73 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.215349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.215349Z digest=sha256:c3dc27ce38c32628be036b4e32184080b7d15963f2ae1b48ff8713eba425d3ee

Observation 65539722-4660-4024-be93-223a12fa8d76 · outbound

This paper cites OneThinker: All-in-one Reasoning Model for Image and Video.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward OneThinker: All-in-one Reasoning Model for Image and Video

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.221576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.221576Z digest=sha256:ac76962022d08cc994185d7069dd2ff674c609f265b07ff4d2c01aa730a342fe

Observation 2eb3994f-034a-49ca-a654-13259944ce9d · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.228889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.228889Z digest=sha256:7f3a3bed79632b6e7a1bbaffb39eb8023e6a94d86cb2e4b2dedd58dd2ffe523e

Observation 53fcffc3-aa5f-49b5-ac92-20dd6ee28176 · outbound

This paper cites VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.235558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.235558Z digest=sha256:19cb4f8d2659703df2027a0d5df8234aa19a7f7473872aa5d98f85840702f094

Observation d031b374-8983-4f7e-a325-cc8741da424e · outbound

This paper cites Exploring the role of audio in video captioning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Exploring the role of audio in video captioning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.067484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.241601Z digest=sha256:69fb108c7f86216694e29329f7e7262da4b5a2526f3728aaa11222d0f2969f33

Observation 81e8b0ea-9029-44e0-8c6e-06c0df6bdacb · outbound

This paper cites Hybrid transformers for music source separation.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Hybrid transformers for music source separation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.049145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.247461Z digest=sha256:978b7809ddd4b0bad7cc193b2ac595730ef460dd6c7aab13e97c93acc676412f

Observation cc0ed116-12d4-4b94-929f-bb14edf82074 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Audio-visual event localization in unconstrained videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.255613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.255613Z digest=sha256:646d3f51e6f211f1f04e8dea1f6d224ddfae7ff08aef25a9ac8aa108dcb1adad

Observation f3944060-20a1-4e43-ac79-d894f5e3ef54 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Vggsound: A large-scale audio-visual dataset

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:47.018344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.261995Z digest=sha256:bd9f8b2de413d2dd27a3e4bf3044fd6863b51de06b2f7f9e62b2fdeccb4613c0

Observation 3ca62a8c-3c7f-419f-b375-76f43901be93 · outbound

This paper cites Condensed movies: Story based retrieval with contextual embeddings.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Condensed movies: Story based retrieval with contextual embeddings

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.268021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.268021Z digest=sha256:645f78b2f075fdf3c69894695713e48f198037090efb36eb94c10b4c3032e5d4

Observation a4d9dd4f-6e26-4315-92f0-725f21550be2 · outbound

This paper cites Avqa: A dataset for audio-visual question answering on videos.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Avqa: A dataset for audio-visual question answering on videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.276994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.276994Z digest=sha256:ed27d8d2b879cc4e0450da9ae657e047d752984f2ee42fc38755ab985001cf60

Observation fee2fdf0-5625-4f0b-a12f-fb4d2404f25a · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Movienet: A holistic dataset for movie understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.978977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.284058Z digest=sha256:6bf1948c011b7a1cee44c73f11ddb6281c2f44fc9d87ff3537c000156453271b

Observation c4d13c8f-2283-41b8-ba93-d06bd766d3e0 · outbound

This paper cites A dataset for movie description.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward A dataset for movie description

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.962720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.292387Z digest=sha256:9000ce7ee0fdf9bc888e8210b78e94613c8deb4010df93e8f9429a20f18fd788

Observation 09d3fbeb-410e-4146-9ec5-ad05891eb23f · outbound

This paper cites Auroracap: Efficient, performant video detailed captioning and a new benchmark.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Auroracap: Efficient, performant video detailed captioning and a new benchmark

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.946632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.300777Z digest=sha256:13e76790bb02c2badb161d237976cc3f4885138bae403fc7756ce0e533cadc7b

Observation cdd33149-e7cc-4cc7-89a3-72bc5fc4f766 · outbound

This paper cites Time-r1: Post-training large vision language model for temporal video grounding.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Time-r1: Post-training large vision language model for temporal video grounding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.929111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.313655Z digest=sha256:6f39beb4fb2ea955373c819f56581bb117fcaee3bb0888d1cf087b95595bf67c

Observation 41f95256-3d3a-4c1f-a884-73c828d10848 · outbound

This paper cites Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:13:44.864443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.320002Z digest=sha256:dacf112936b0bfa54a2c5297e0fa4a9ac9e8a08ad01b6fa5eddae1735fffe1e1

Observation d6ea50e2-06d0-4b15-981a-95988c63ed04 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.329322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.329322Z digest=sha256:c8e58eb227d7624a0a8ea50c9f7fd9e22f57618bfc421185b6bf286e3cb6f377

Observation 324bbfe1-68ca-488a-968a-de6c43b9bfcf · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1 (2):3, 2022.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Lora: Low-rank adaptation of large language models.ICLR, 1 (2):3, 2022

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.336856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.336856Z digest=sha256:0a9ef02e7fe5dfa2fca6818b1751ee1d58db5033e0f4933c1212978279cde714

Observation fa408bef-362d-41bb-b75f-70b2b82f91e1 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Efficient memory management for large language model serving with pagedattention

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.342904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.342904Z digest=sha256:46e2809bfe92be9342a9bfbe86cfa9505d96edc2b8386abe568b08b1606da4d2

Observation 455e53d5-8b79-462a-8263-e33b9bceb9df · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Vbench: Comprehensive benchmark suite for video generative models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.369645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.369645Z digest=sha256:0e6c570008bd2afbc29a3ed90388f345bcb9d044044706ed67e8006fd5a4b853

Observation 367cf444-21a9-45a3-bae8-ce99e94b6814 · outbound

This paper cites VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.376026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.376026Z digest=sha256:f4f46101cbf51d3c85e002f2a3fa0afa017d16a5a7ead23d1df4e6b1510b3c59

Observation 1074ca18-4c48-469d-a37d-db46b91920bc · outbound

This paper cites HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.385618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.385618Z digest=sha256:8857ca0f2b34b1e2ca84ba68e9878cb1f059a6b22f946d55d4ec5c56194a22ba

Observation 8aeb4fe0-9bc3-4f5c-9714-ed4492a388df · outbound

This paper cites ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.392418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.392418Z digest=sha256:19af5d4d874eccf4ab671b865eb3258a5a26935864ca6847d82f03aad32b8387

Observation d59819bb-9364-46b4-a338-0f3746d8b9a2 · outbound

This paper cites Minicpm-o 2.6: A gpt-4o level mllm for vision, speech, and multimodal live stream- ing on your phone.https://github.com/OpenBMB/MiniCPM-V, 2025.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Minicpm-o 2.6: A gpt-4o level mllm for vision, speech, and multimodal live stream- ing on your phone.https://github.com/OpenBMB/MiniCPM-V, 2025

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.876062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.399890Z digest=sha256:be8c1cb28374218b442d23cfc2a820eb24a53ffd0ca7a26d37182b5d53fa46d4

Observation dc9de645-8519-4180-9fae-51dfd07e6618 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.406841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.406841Z digest=sha256:95c48f230c2a291b20aa6bc3790587045c95be78f818aa478eee2c0538cfb05d

Observation be280df9-dcdb-4e6d-9846-4a4242ff3b46 · outbound

This paper cites Self- critical sequence training for image captioning.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Self- critical sequence training for image captioning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.414464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.414464Z digest=sha256:735219ded0207626de8fde72a918042fe02445427336800c0da4fef398b0c5bd

Observation c49af00d-fb79-4ee8-ac0e-6cd6f366ea28 · outbound

This paper cites id": "sample id.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward id": "sample id

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.848800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.422217Z digest=sha256:9fa0719ca126933596a33e4bbef0ea600ca719312d91ab540facd8c64a6911ee

Observation da44a066-d0cc-4715-a367-9fcfba7c58c0 · outbound

This paper cites Use the exact terminology found in the text (e.g., if the caption says ”shatters”, use ”shatters”, not ”breaks”).

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Use the exact terminology found in the text (e.g., if the caption says ”shatters”, use ”shatters”, not ”breaks”)

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.829527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.429120Z digest=sha256:c606f49f8a2ed6adf4ee4f01d024983cdd7cdd07505f82f7355146d7f954a82c

Observation 31b2fb4b-b1be-4066-88df-0ed9d35bbe6f · outbound

This paper cites an unresolved cited work.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:13:46.809850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.441284Z digest=sha256:7ba973f48e70be1084719bfc460a9584d0856693fe00324f0755b68a92968a63

Observation 06ff5225-4b67-4407-a027-13511abd61a4 · outbound

This paper cites Start directly with the answer content.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Start directly with the answer content

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.788962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.450979Z digest=sha256:969a80339f423d8a4e9c67243334a16171d3517cb3b670e6b179001e85d42ff0

Observation fa9669e6-ff5d-45ad-bb6c-b38be7b63981 · outbound

This paper cites strengths.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward strengths

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.764825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.457246Z digest=sha256:7cdfb1666a73bdf041bc997c451a69b02360955d8668a575adec79f2cd736483

Observation 4a20e370-b5df-4187-8a70-e124cda1f609 · outbound

This paper cites an unresolved cited work.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:13:46.742112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.464484Z digest=sha256:31373fba72c14404749291def97d3769b4b15772f1691d34ab74899e20f598a3

Observation bee123fd-afc5-43ed-bb1d-2bd4514bc827 · outbound

This paper cites evaluation.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward evaluation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.717454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.473015Z digest=sha256:9c771538cb95efb52e6434a78064ddcc8e1f511dcca3aee0754cde5b4f1fa9b0

Observation 6f54366b-2833-4bff-a4bb-38ad54d5fcc9 · outbound

This paper cites Do not add any artistic interpretation, subjective analysis, or infer any character’s internal 23 thoughts, emotions, or intentions beyond what is explicitly visible or audible.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Do not add any artistic interpretation, subjective analysis, or infer any character’s internal 23 thoughts, emotions, or intentions beyond what is explicitly visible or audible

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.692726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.480978Z digest=sha256:67f591f0df6981718c05835a60677bef82708816f5b6d7ad8c9fc17c450b2b6a

Observation 90a1999f-d774-4123-a162-e1bbae85ba49 · outbound

This paper cites While Character A is speaking, Character B simultaneously turns their head.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward While Character A is speaking, Character B simultaneously turns their head

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.668462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.486950Z digest=sha256:2157f46d15846e5277efaa5341af7b43e09788757de60d05034eb1a81442bcfb

Observation 36b615f7-5d65-48b6-b411-84544350326a · outbound

This paper cites He closed the door,.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward He closed the door,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.648814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.494599Z digest=sha256:b5a31fe00d7cbf3c885c385c4c6653400b0cda2e38c772760da5d490fb3c3a1a

Observation a39cd8a8-82c3-4890-ab62-d46deb18892b · outbound

This paper cites The camera cuts to a close-up of her face,.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward The camera cuts to a close-up of her face,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.628650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.502462Z digest=sha256:e2a6fdc2b6fdee6a105196aba8a9d99bed9a55f4ce136baa8f9103fa67b76e59

Observation 52b76ef2-21c0-4633-a1fe-87058085ec5f · outbound

This paper cites an unresolved cited work.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:13:46.608847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.516642Z digest=sha256:5ad018ad28c883bdc3bb6cc5a3cfea00acc572e2aec1b50480dd717ce3807072

Observation e8eebfdd-7bb2-4176-9ed5-f3a690bb0f80 · outbound

This paper cites Do not use bullet points, headings, or line breaks within the narrative.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Do not use bullet points, headings, or line breaks within the narrative

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.577670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.523799Z digest=sha256:086624afa8f9f06aca549559e7d073954eb93bd196970533524d0c0a8646e94a

Observation 70b54954-8d96-4140-823a-14fdffbda128 · outbound

This paper cites His voice, thick with sarcasm, says.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward His voice, thick with sarcasm, says

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.553912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.531066Z digest=sha256:ebbbb2dbf5f16f1c83fc36d731618fc2e57ebfe1bb4a70df92b399ae15ad43b1

Observation 90c315f3-909b-46c3-93ad-2d8bb8a5336c · outbound

This paper cites thud" as a book hits the table, a.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward thud" as a book hits the table, a

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.526057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.543586Z digest=sha256:32a5f4a85b3f4dad81eaa0ea4f3a49bb831671fe9b1de5942abc0ca5314a870b

Observation 64b2eb73-aa40-48d3-8df0-360d4fdca671 · outbound

This paper cites While Person A’s hand reaches for the glass, Person B’s eyes dart to the side, and the music swells.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward While Person A’s hand reaches for the glass, Person B’s eyes dart to the side, and the music swells

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.499941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.557179Z digest=sha256:228a7319f9c248ef0c50ff8633c2a456f3eb37c97cd3bf00a225936a9867a907

Observation 0092d044-e526-435f-8cb0-7f3df378f2fb · outbound

This paper cites Avoid summarizing.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Avoid summarizing

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.471133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.563308Z digest=sha256:e8340d6451d08ff2b2401aa4343699f4b61c50fabdd58db5a5600d89d359d638

Observation 8457dc9e-0c4b-4dcb-8fe2-d16d504a83a2 · outbound

This paper cites strengths.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward strengths

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:13:46.445294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T18:13:44.573450Z digest=sha256:46e7d37332823e9aeeea47bc6f239fa8c3e52f6e0286d0416591f659faac210c

Observation 271c8407-3349-4ac9-8db6-390573dd5ff2 · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.362349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.362349Z digest=sha256:5b2df4e8bb71adc90e5e90619918a32527d9d13a067c71663b58a7dd9fbf36ba

Pith citing papers

No inbound Pith citation observations are available.