Pith. sign in

Paper Citation Record · LEDGER

PercepCap: Video Captioner with Structured Spatio-Temporal Perception

As of 11 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2607.20389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.20389 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T10:02:03.322619Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c8393da9-816f-444a-a576-4e81edcb2b57 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.126290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.126290Z digest=sha256:38410ef518cd49b7e132f52eb00621522eb7a6559e80b186e99b54b43c30e65a

Observation fd3d746c-0d31-4308-b2fd-fd6163347d76 · outbound

This paper cites Qwen2.5-VL Technical Report.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.131618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.131618Z digest=sha256:60f31e67f04bf1eccaa9dd4e18d28bd06034d814a81bddb320df7acf941d5026

Observation e6ac68c0-9b2a-436d-a235-09b1308c211f · outbound

This paper cites Activitynet: A large- scale video benchmark for human activity understanding.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Activitynet: A large- scale video benchmark for human activity understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.136729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.136729Z digest=sha256:b280d6d931f6a091526efad06f5e2d7e95f327fb796b8b7430a545bcca0d53ed

Observation 4a2ab20c-32cd-4207-afa5-4298ace80d12 · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.141536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.141536Z digest=sha256:2ab6534590ddb0b0b2d7b9999e60b58f0e35b4330793a5253f3461ab6b50b1d7

Observation d7ca5978-b6db-44e8-a69b-e6aeb2321a94 · outbound

This paper cites VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.146589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.146589Z digest=sha256:5acabc47eca520f38036583e110e97212cc5917ec594e8f78ffccb64b55552c8

Observation 5e0cd9cf-368b-4956-b544-b5bbd374e380 · outbound

This paper cites Vidbridge-r1: Bridging QA and captioning for RL-based video understanding models with intermediate proxy tasks.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Vidbridge-r1: Bridging QA and captioning for RL-based video understanding models with intermediate proxy tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.151348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.151348Z digest=sha256:4289778296b6cbd04775ee816d3403fb4fa7728c37f2772dcd8425fb5b5032b1

Observation ddef46af-9b02-46dd-962e-ae0c6a8e9869 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.156349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.156349Z digest=sha256:a92ddb6c73a494bae73c53916afa7eb68603372749669a2388a772af1d64f240

Observation 2793345b-030b-4ce0-b2ca-e126fe1c59c6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.161003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.161003Z digest=sha256:775b37e3d3fa1cb71970dd6dac905dc0080d20769ff093ba5056b5a89af71ee6

Observation e773c642-522a-4375-8026-1b9b52077a05 · outbound

This paper cites ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.165633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.165633Z digest=sha256:0b0929279a945bc2b718f549b3ab31ec755c1482c6000ed939d4a18e2d430602

Observation 60e1131a-e19f-4b5c-8192-c0af7254f2dd · outbound

This paper cites Gemini 3 Flash: frontier intelligence built for speed.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Gemini 3 Flash: frontier intelligence built for speed

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.170486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.170486Z digest=sha256:c6aed3698f9c6b99393061949a3dda7c6d85caf6d2561b19ce4e383ebfdeab49

Observation 94bdbe3d-c78a-48d8-9813-755d914c131b · outbound

This paper cites MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.174754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.174754Z digest=sha256:c5b1e1ea20f6b601cd6d696ed9f2505a2046d44981cb9efd2a2a0465d44539f1

Observation d82e4985-eb75-4f10-b1a4-a9ed2fb4d011 · outbound

This paper cites VIDEOP2R: Video Understanding from Perception to Reasoning.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception VIDEOP2R: Video Understanding from Perception to Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.179245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.179245Z digest=sha256:ed9bfc2e800973c6cb7efb355269c8da876ff2237da2420185b33777c9e6afa0

Observation da513b2d-6f92-4f17-b154-02f62604a93d · outbound

This paper cites Kimi-VL Technical Report.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Kimi-VL Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.183827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.183827Z digest=sha256:cf5954b5a413fcac666bb8a56af3572e7c13df9a5cfcb543d271a62512f0db7a

Observation e0b71761-28b1-4e0c-94a8-e8ab46ebfcca · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.188685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.188685Z digest=sha256:a1c6c9568eb7aa483c0a1055c4e190faf838ea74685c3a3cf777b2931dc0ee19

Observation d0afa628-5378-4971-a9fa-dec576fd2943 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.193696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.193696Z digest=sha256:42678303ed502f81d0b25efe0a2281fedcff47af7738c81862ad76c2a986b159

Observation 138fd2ab-267f-4afa-8eab-76f6fbbcd211 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Video-llava: Learning united visual representation by alignment before projection

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.198541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.198541Z digest=sha256:c9cd633da600a2fbd99c2d1e2ad9ab830d1b9556b26db9409d92b88edd101f59

Observation 03a3248e-05ee-4294-9b22-3a62ec7e38fd · outbound

This paper cites MiMo-VL technical report, 2025.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception MiMo-VL technical report, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.203392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.203392Z digest=sha256:fcf8de0d785c0fe813aa9694f98074cec1a6ff73d30975f48ce351589984cb9f

Observation 32461e25-5782-49b6-a010-1cc41b265953 · outbound

This paper cites VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.207694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.207694Z digest=sha256:c2fa4c1cec6a613f9343a74585c43d58a810c3363e418721e27eb8d910492b6f

Observation 53729e25-2423-4d2d-9905-c59511a9fde8 · outbound

This paper cites an unresolved cited work.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.212359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.212359Z digest=sha256:cb7dfdd9ab8df192f29d16c73260fd6bc2507fd733139356a2fd088fca0ffebc

Observation a389b9d5-8ab3-4146-9ff9-23a00909d5e7 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, 2026.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Qwen3.5: Towards native multimodal agents, 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.216944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.216944Z digest=sha256:ff32b1d55e1ccada9f82be30a3c41ec183ca9e4ad45c2a87a41768beb54a0b93

Observation fe942059-c849-4b64-9137-101848421199 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.221848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.221848Z digest=sha256:04f4af3a3ddbde369c1ee141773c54eae4b63acdddc9be14650af70b336b09fb

Observation 4a40ff84-6af1-4376-960f-a9988d8681c8 · outbound

This paper cites Qwen3 technical report, 2025.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Qwen3 technical report, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.226653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.226653Z digest=sha256:1bdfd4fd79d3e498c79ace679582f3f6c3f976b1d70ce117637a1bc7b41f01ca

Observation 2c0f1f7c-4481-4f16-9324-f6b1129e406d · outbound

This paper cites Gemini 3.1 pro: A smarter model for your most complex tasks, 2026.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Gemini 3.1 pro: A smarter model for your most complex tasks, 2026

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.230903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.230903Z digest=sha256:5601a50a6f4c3ecec694909cff67c5c46e93cb128069550013d499047102930b

Observation 4bef3798-b3c7-487e-941a-c8e92139deae · outbound

This paper cites Reft: Reasoning with reinforced fine-tuning.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Reft: Reasoning with reinforced fine-tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.235083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.235083Z digest=sha256:04b9813fc10df82a87f1f9de749bae4f97b12d8a4fc9aa68a2ec0785a757dc29

Observation 922b3bc3-e978-43ae-b0ff-2453d87e706f · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.239135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.239135Z digest=sha256:a104a07690428b40fcf72db13619e9fb5d7e7f519c36676edb3520581cb61b07

Observation 0f4661c7-a4c4-4aeb-8791-76af60bd4357 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.243631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.243631Z digest=sha256:24bd6186368732070c35dbaec8c118f5a6f4c2845c8aa77b20f35a1c4b8361ba

Observation a5cf8e40-d2b5-4a4b-acff-aa37c7d95ec2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.248142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.248142Z digest=sha256:81811cbf963a542e50436212edce9b89e79b8909640d7fac2a990f964a7301d8

Observation f6a1186f-f177-43b6-91e8-07412c4b0314 · outbound

This paper cites YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.252658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.252658Z digest=sha256:4aad8dd259250426bad35b3451018ea4404c3dee2b93005c4127115abc2bf7a0

Observation 403360a5-3898-4e15-b88c-8971218457de · outbound

This paper cites CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.257559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.257559Z digest=sha256:54f14c5c033e250a566d9fe4b616e922ef1f1f23c95a073de5f323ff9cf943f9

Observation d0a9f074-7e9a-4d3e-b4f4-83053fe8869b · outbound

This paper cites Videochat-r1.5: Visual test-time scaling to reinforce multimodal reasoning by iterative perception.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Videochat-r1.5: Visual test-time scaling to reinforce multimodal reasoning by iterative perception

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.262030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.262030Z digest=sha256:a46d1bde68a356ac8998c923c003db5966538252ea4eae6f30b9c2863aba5916

Observation 7deff422-86ba-4f42-a199-15919089ae58 · outbound

This paper cites TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.266407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.266407Z digest=sha256:f6f61f9c39d748bcb40802947651a6b8682c56c7508785588125a95d3b49ea74

Observation de40401e-a909-464c-817d-d64cb75620bf · outbound

This paper cites Timelens: Rethinking video temporal grounding with multimodal llms.CoRR, abs/2512.14698, 2025.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Timelens: Rethinking video temporal grounding with multimodal llms.CoRR, abs/2512.14698, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.271785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.271785Z digest=sha256:2de7f366187c4dc08143051541bcfe7e77ad0e358b33871f48be399738c8539b

Observation 4aa51a07-7303-412e-9152-bd644c668234 · outbound

This paper cites perception.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception perception

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.276162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.276162Z digest=sha256:86ab99707ed6b2205e5903fabc1310d7f0d541f356d0d104841a6c1476c1c309

Observation f72a2dfa-011d-49dc-a372-c0d2efa55601 · outbound

This paper cites an unresolved cited work.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.280832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.280832Z digest=sha256:a0fae0c57482227aac26a685f110b59099da2475e0ce9d859e852a0c15b17635

Observation 516049fe-0bb4-49ef-a319-813a14a8e89d · outbound

This paper cites These extracted textual contents define which entities and events must appear in the perception trace.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception These extracted textual contents define which entities and events must appear in the perception trace

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.285330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.285330Z digest=sha256:b8e475486444d9fff1db498129c2542b43fdf08da00c807c0e792d976bee7650

Observation 46c5d169-24b9-4bae-85e2-2b1f2ae60a82 · outbound

This paper cites 14 Figure 5: Constructed training example from Caption-Anchored Perception Data Construction.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception 14 Figure 5: Constructed training example from Caption-Anchored Perception Data Construction

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.290370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.290370Z digest=sha256:731448426b0516f94999ca96a250e4d26c28ae76f979924a1f7f0d06f84ee2e6

Observation 20d487d8-a1ea-47bf-9f0b-2becda677cf1 · outbound

This paper cites objects". - The value of.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception objects". - The value of

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.294926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.294926Z digest=sha256:972557f3a8c3f31e8b9b319f7585a591a5a447f698d6ae33ba6e7f11f8096074

Observation 70a5bb5e-573c-4fb4-95cb-3cb8e7635d20 · outbound

This paper cites - The same real-world instance MUST keep the same id across frames.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception - The same real-world instance MUST keep the same id across frames

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.299431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.299431Z digest=sha256:b4e62b73fd636678b8f70dedfe95b558ef3109a7fa6651c0dc019446e2e761f0

Observation ec7877d2-57b9-4db1-8847-142a7023c16f · outbound

This paper cites an unresolved cited work.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.304452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.304452Z digest=sha256:bf97b5c264f0a973850ce5b1eaa70c496e7cf1e7ddece0e5d9d5c9c1a403f878

Observation b0e5937c-4035-4979-9a5f-42c853934163 · outbound

This paper cites - bbox_2d: [xmin, ymin, xmax, ymax] in relative pixel coordinates, range [0, 1000].

PercepCap: Video Captioner with Structured Spatio-Temporal Perception - bbox_2d: [xmin, ymin, xmax, ymax] in relative pixel coordinates, range [0, 1000]

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.309760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.309760Z digest=sha256:6c76bd3f3e86200e3fcb6c7c07a9fd11961c2b513f1c6b65bb29a16c6a65f881

Observation eefe8180-7335-4187-a559-1eab0c4047e6 · outbound

This paper cites an unresolved cited work.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.313844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.313844Z digest=sha256:b9915c61528c8510481a156280a3140a08edf1b09007546d23e5bed18facac11

Observation 06653c42-c7b0-4e5d-8f5c-a112c2d4dcfd · outbound

This paper cites an unresolved cited work.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.318125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.318125Z digest=sha256:f987e2b8fb0b5de510febc213361530c3d3f55799ecb3edffc737ce4f5ef3bf1

Observation c8b9f358-5f34-4e56-be07-028c06531626 · outbound

This paper cites objects":[{.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception objects":[{

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.322619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.322619Z digest=sha256:d991f97e0af48353b21f8c714fc49b4b1365eb8b6ba60b7b0ed667653b9bc20f

Pith citing papers

No inbound Pith citation observations are available.