Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:13:44.573450Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2608.06930.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:13:44.573450Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
75 of 75 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4eaed5f7-1f2a-4ff8-8f8f-92924844fd08 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Avocado: An audiovisual video captioner driven by temporal orchestration.arXiv preprint arXiv:2510.10395, 2025
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b695bdef-a788-4514-949d-8065a0a98e1f · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Ugc-videocaptioner: An omni ugc video detail caption model and new benchmarks.arXiv preprint arXiv:2507.11336, 2025
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad1b155c-49ef-45af-b7b2-78dea5511969 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Daily-omni: Towards audio-visual reasoning with temporal alignment across modalities.arXiv preprint arXiv:2505.17862, 2025
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f454879b-db8d-4894-b6a7-98084bef545c · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f6e4ae1-519f-4569-be74-61cf0eed85a9 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37: 19472–19495, 2024
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 55a344b0-2078-444f-9837-3581bfcb57db · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mavors: Multi-granularity video representation for multimodal large language model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c73a9e0-e967-446c-a679-d61cfd231a08 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90db9274-bbf3-46ad-8514-45a615392547 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b84cf9f9-f405-441b-9a80-a8a0db9a6853 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac750a63-35b8-467a-9019-cfa9d9a1d905 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 727dae58-6d8b-4478-bb13-886fcd349312 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Advancing high-resolution video-language representation with large-scale video transcriptions
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6995e3e9-3e81-4806-af86-eef40d176bdb · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward video-salmonn 2: Captioning-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220, 2025
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f752d0b-4aad-460d-9aba-c565c4d88bb7 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70ff3cbf-52e2-4513-b5d7-e931862558b8 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Longvale: Vision-audio-language-event benchmark towards time-aware omni-modal perception of long videos
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8db4405e-fa54-4f54-95d9-ff440f26235c · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a402971b-163d-41f1-bf6a-5779e77e4a33 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fafedc81-20fb-4608-872c-4b5035d3bde9 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87b8dd1b-90a8-4b20-bc27-6bb85e24b4a2 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Watch, listen, and describe: Globally and locally aligned cross-modal attentions for video captioning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e10c871d-faef-49c8-bac8-ad08c41aec11 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a17029f-7952-4eac-a5f8-50f77b513d43 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Qwen2.5-VL Technical Report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d47876c-b084-4857-a211-94c6eba7760b · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Glave-cap: Global-local aligned video captioning with vision expert integration.arXiv preprint arXiv:2509.11360, 2025
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8f1a29e-aef0-4ac8-9db6-e8e868898a3a · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Qwen2.5-Omni Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ead60c4d-46c1-4950-977f-efa478de054f · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward OmniCaptioner: One Captioner to Rule Them All
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1c38d65-50d3-4629-b59f-5b52c5a203e4 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Omni-captioner: Data pipeline, models, and benchmark for omni detailed perception.arXiv preprint arXiv:2510.12720, 2025
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aab17aee-273f-41a8-9111-9ea95809ed00 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91ab4e8d-8701-43e7-a6f1-180c2e0142ef · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dfeba4d-990f-4466-961f-7991cdc82a8b · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward AdaTooler-V: Adaptive Tool-Use for Images and Videos
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e486e27-48ba-4008-87db-37a63dcf1e18 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Deepvideo-r1: Video rein- forcement fine-tuning via difficulty-aware regressive grpo.arXiv preprint arXiv:2506.07464, 2025
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89671b3f-be56-4556-bfb1-d94ce96e8412 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Editthinker: Unlocking iterative reasoning for any image editor
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e3a51c8-22b5-4d13-a606-cc7a0e510022 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Sophiavl-r1: Reinforcing mllms reasoning with thinking reward.arXiv preprint arXiv:2505.17018, 2025
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ec3c7c0-87f4-447b-885a-edc9ca334b31 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf7d5b55-0444-40c4-aacb-a628a9681f73 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65539722-4660-4024-be93-223a12fa8d76 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward OneThinker: All-in-one Reasoning Model for Image and Video
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb3994f-034a-49ca-a654-13259944ce9d · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53fcffc3-aa5f-49b5-ac92-20dd6ee28176 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d031b374-8983-4f7e-a325-cc8741da424e · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Exploring the role of audio in video captioning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 81e8b0ea-9029-44e0-8c6e-06c0df6bdacb · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Hybrid transformers for music source separation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cc0ed116-12d4-4b94-929f-bb14edf82074 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Audio-visual event localization in unconstrained videos
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3944060-20a1-4e43-ac79-d894f5e3ef54 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Vggsound: A large-scale audio-visual dataset
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3ca62a8c-3c7f-419f-b375-76f43901be93 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Condensed movies: Story based retrieval with contextual embeddings
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4d9dd4f-6e26-4315-92f0-725f21550be2 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Avqa: A dataset for audio-visual question answering on videos
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fee2fdf0-5625-4f0b-a12f-fb4d2404f25a · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Movienet: A holistic dataset for movie understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c4d13c8f-2283-41b8-ba93-d06bd766d3e0 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward A dataset for movie description
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 09d3fbeb-410e-4146-9ec5-ad05891eb23f · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Auroracap: Efficient, performant video detailed captioning and a new benchmark
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cdd33149-e7cc-4cc7-89a3-72bc5fc4f766 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Time-r1: Post-training large vision language model for temporal video grounding
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 41f95256-3d3a-4c1f-a884-73c828d10848 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d6ea50e2-06d0-4b15-981a-95988c63ed04 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 324bbfe1-68ca-488a-968a-de6c43b9bfcf · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Lora: Low-rank adaptation of large language models.ICLR, 1 (2):3, 2022
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa408bef-362d-41bb-b75f-70b2b82f91e1 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Efficient memory management for large language model serving with pagedattention
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 455e53d5-8b79-462a-8263-e33b9bceb9df · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Vbench: Comprehensive benchmark suite for video generative models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 367cf444-21a9-45a3-bae8-ce99e94b6814 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1074ca18-4c48-469d-a37d-db46b91920bc · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aeb4fe0-9bc3-4f5c-9714-ed4492a388df · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d59819bb-9364-46b4-a338-0f3746d8b9a2 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Minicpm-o 2.6: A gpt-4o level mllm for vision, speech, and multimodal live stream- ing on your phone.https://github.com/OpenBMB/MiniCPM-V, 2025
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dc9de645-8519-4180-9fae-51dfd07e6618 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be280df9-dcdb-4e6d-9846-4a4242ff3b46 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Self- critical sequence training for image captioning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c49af00d-fb79-4ee8-ac0e-6cd6f366ea28 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward id": "sample id
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation da44a066-d0cc-4715-a367-9fcfba7c58c0 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Use the exact terminology found in the text (e.g., if the caption says ”shatters”, use ”shatters”, not ”breaks”)
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 31b2fb4b-b1be-4066-88df-0ed9d35bbe6f · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 06ff5225-4b67-4407-a027-13511abd61a4 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Start directly with the answer content
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fa9669e6-ff5d-45ad-bb6c-b38be7b63981 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward strengths
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4a20e370-b5df-4187-8a70-e124cda1f609 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bee123fd-afc5-43ed-bb1d-2bd4514bc827 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward evaluation
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6f54366b-2833-4bff-a4bb-38ad54d5fcc9 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Do not add any artistic interpretation, subjective analysis, or infer any character’s internal 23 thoughts, emotions, or intentions beyond what is explicitly visible or audible
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 90a1999f-d774-4123-a162-e1bbae85ba49 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward While Character A is speaking, Character B simultaneously turns their head
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 36b615f7-5d65-48b6-b411-84544350326a · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward He closed the door,
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a39cd8a8-82c3-4890-ab62-d46deb18892b · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward The camera cuts to a close-up of her face,
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 52b76ef2-21c0-4633-a1fe-87058085ec5f · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e8eebfdd-7bb2-4176-9ed5-f3a690bb0f80 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Do not use bullet points, headings, or line breaks within the narrative
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 70b54954-8d96-4140-823a-14fdffbda128 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward His voice, thick with sarcasm, says
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 90c315f3-909b-46c3-93ad-2d8bb8a5336c · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward thud" as a book hits the table, a
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 64b2eb73-aa40-48d3-8df0-360d4fdca671 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward While Person A’s hand reaches for the glass, Person B’s eyes dart to the side, and the music swells
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0092d044-e526-435f-8cb0-7f3df378f2fb · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Avoid summarizing
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8457dc9e-0c4b-4dcb-8fe2-d16d504a83a2 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward strengths
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 271c8407-3349-4ac9-8db6-390573dd5ff2 · outbound
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.