Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T10:02:03.322619Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2607.20389.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T10:02:03.322619Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c8393da9-816f-444a-a576-4e81edcb2b57 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd3d746c-0d31-4308-b2fd-fd6163347d76 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6ac68c0-9b2a-436d-a235-09b1308c211f · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Activitynet: A large- scale video benchmark for human activity understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a2ab20c-32cd-4207-afa5-4298ace80d12 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7ca5978-b6db-44e8-a69b-e6aeb2321a94 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e0cd9cf-368b-4956-b544-b5bbd374e380 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Vidbridge-r1: Bridging QA and captioning for RL-based video understanding models with intermediate proxy tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddef46af-9b02-46dd-962e-ae0c6a8e9869 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2793345b-030b-4ce0-b2ca-e126fe1c59c6 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e773c642-522a-4375-8026-1b9b52077a05 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60e1131a-e19f-4b5c-8192-c0af7254f2dd · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Gemini 3 Flash: frontier intelligence built for speed
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94bdbe3d-c78a-48d8-9813-755d914c131b · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d82e4985-eb75-4f10-b1a4-a9ed2fb4d011 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception VIDEOP2R: Video Understanding from Perception to Reasoning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da513b2d-6f92-4f17-b154-02f62604a93d · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Kimi-VL Technical Report
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0b71761-28b1-4e0c-94a8-e8ab46ebfcca · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0afa628-5378-4971-a9fa-dec576fd2943 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 138fd2ab-267f-4afa-8eab-76f6fbbcd211 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Video-llava: Learning united visual representation by alignment before projection
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03a3248e-05ee-4294-9b22-3a62ec7e38fd · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception MiMo-VL technical report, 2025
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32461e25-5782-49b6-a010-1cc41b265953 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53729e25-2423-4d2d-9905-c59511a9fde8 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a389b9d5-8ab3-4146-9ff9-23a00909d5e7 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Qwen3.5: Towards native multimodal agents, 2026
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe942059-c849-4b64-9137-101848421199 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a40ff84-6af1-4376-960f-a9988d8681c8 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Qwen3 technical report, 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c0f1f7c-4481-4f16-9324-f6b1129e406d · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Gemini 3.1 pro: A smarter model for your most complex tasks, 2026
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bef3798-b3c7-487e-941a-c8e92139deae · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Reft: Reasoning with reinforced fine-tuning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 922b3bc3-e978-43ae-b0ff-2453d87e706f · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f4661c7-a4c4-4aeb-8791-76af60bd4357 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5cf8e40-d2b5-4a4b-acff-aa37c7d95ec2 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6a1186f-f177-43b6-91e8-07412c4b0314 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 403360a5-3898-4e15-b88c-8971218457de · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0a9f074-7e9a-4d3e-b4f4-83053fe8869b · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Videochat-r1.5: Visual test-time scaling to reinforce multimodal reasoning by iterative perception
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7deff422-86ba-4f42-a199-15919089ae58 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de40401e-a909-464c-817d-d64cb75620bf · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Timelens: Rethinking video temporal grounding with multimodal llms.CoRR, abs/2512.14698, 2025
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aa51a07-7303-412e-9152-bd644c668234 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception perception
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f72a2dfa-011d-49dc-a372-c0d2efa55601 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 516049fe-0bb4-49ef-a319-813a14a8e89d · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception These extracted textual contents define which entities and events must appear in the perception trace
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46c5d169-24b9-4bae-85e2-2b1f2ae60a82 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception 14 Figure 5: Constructed training example from Caption-Anchored Perception Data Construction
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20d487d8-a1ea-47bf-9f0b-2becda677cf1 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception objects". - The value of
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70a5bb5e-573c-4fb4-95cb-3cb8e7635d20 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception - The same real-world instance MUST keep the same id across frames
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec7877d2-57b9-4db1-8847-142a7023c16f · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e5937c-4035-4979-9a5f-42c853934163 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception - bbox_2d: [xmin, ymin, xmax, ymax] in relative pixel coordinates, range [0, 1000]
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eefe8180-7335-4187-a559-1eab0c4047e6 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06653c42-c7b0-4e5d-8f5c-a112c2d4dcfd · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8b9f358-5f34-4e56-be07-028c06531626 · outbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception objects":[{
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.