Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:49:30.743524Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2506.16082.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:49:30.743524Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fa5e75c7-e623-44cf-a38d-39cffd799f59 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Swinbert: End-to-end transformers with sparse attention for video captioning,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1d89fd1c-5d43-41f3-b082-3be8c3e7a0bd · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Univl: A unified video and language pre-training model for multimodal understanding and generation,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 44fa5252-8e30-4483-9f19-d2bcfec2807d · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning End-to-end generative pretraining for multimodal video captioning,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e4831189-774a-4eac-8d97-886958a8f233 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Memory- attended recurrent network for video captioning,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 113ddb04-1b16-413b-99b6-452427e883e1 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Sports video captioning via attentive motion representation and group relationship modeling,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d43601c3-4508-4875-bfd2-808817b94359 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Reconstruction network for video captioning,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 10703529-50cb-490d-9161-a037c9997068 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Concept-aware video captioning: Describing videos with effective prior information,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ae159af9-be8a-438c-b1f6-4b7a4d07989c · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Dense- captioning events in videos,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cfeb5a98-db2d-42d5-9589-5adad7828359 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Jointly localizing and describing events for dense video captioning,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1822c6d2-42aa-4df6-b921-e077a0235f16 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning End-to- end dense video captioning with parallel decoding,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 11f72725-4a3f-4343-a00f-0edd60a4472f · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 542d302d-477c-4189-895d-7c8ed67132f5 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Cap4video: What can auxiliary captions do for text-video retrieval?
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d891ad30-fab8-40a4-8db5-83a0afcbc4d0 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Multi-event video-text retrieval,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f2c9c0eb-2240-4a09-8218-19fdfc163885 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Exploiting unlabeled videos for video-text retrieval via pseudo-supervised learning,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4e662bfc-e59e-4372-9249-c8be4ba23f18 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Video recap: Recursive captioning of hour-long videos,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5699f10b-f5c1-46ad-b6f2-758d831f54c2 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Vidchapters- 7m: Video chapters at scale,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0bc1ec0c-a38e-4e8d-b366-121891cef46e · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Hierarchical representation network with auxiliary tasks for video captioning and video question answering,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ad283bb7-0429-447c-b561-e24c5efdff8a · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Multi-modal dense video captioning,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ab34a21b-64be-463e-a64b-1a9239f2c474 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning A Better Use of Audio-Visual Cues: Dense Video Captioning with Bi-modal Transformer
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b4428d7-f918-40f0-af3a-9538660fe4e1 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Hierarchical context encoding for events caption- ing in videos,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3aed2096-8ddd-4045-a9a1-361f48bb4193 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning End-to-end object detection with transformers,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 794f98ca-fcfe-46b4-94b1-bf96aafb570d · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Do you remember? dense video captioning with cross-modal memory retrieval,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f70b8490-cd78-443e-bf6e-3592beb5a3a6 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Parallel pathway dense video captioning with deformable transformer,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2a23eb09-4550-4e63-901e-be803f5120dc · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0f88833e-93cf-45b4-90b9-7002b9add345 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Towards automatic learning of procedures from web instructional videos,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ca198214-842b-4c2c-9d69-00f1bbf8c895 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Bidirectional attentive fusion with context gating for dense video captioning,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 67355e09-8a00-4481-bef1-1ce32a7536e0 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Sketch, ground, and refine: Top-down dense video captioning,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5ed51ade-4dbb-4174-b9d4-1445d517539f · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning iPerceive: Applying Common-Sense Reasoning to Multi-Modal Dense Video Captioning and Video Question Answering
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 790c75d0-3028-4619-9d21-93b99435fccb · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Towards bridging event captioner and sentence localizer for weakly supervised dense event captioning,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ec21ce32-6afd-4ff0-b1c6-82e02b299266 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Streamlined dense video captioning,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 31f8ffc7-98bc-439b-9cf4-32cdbbc7a81d · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Watch, listen and tell: Multi- modal weakly supervised dense event captioning,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ef6ff805-c600-4479-8e8e-ce78bfbff472 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Dense procedure captioning in narrated instructional videos,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a99fc2d2-df09-4f3c-8b8a-7e48293598ba · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Retrieval- augmented generation for knowledge-intensive nlp tasks,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01aac300-6314-4204-b877-fb2c2bf0ad17 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Streaming dense video captioning,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f24b0380-e52d-42a3-9f68-cff204ddeec8 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8abfb3a-ee01-41e9-aca8-7db98dba97f2 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Learning texture transformer network for image super-resolution,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f755325a-cf61-4d1f-8217-8c1c20993a7f · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Siamese-detr for generic multi-object tracking,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 655c5caf-df19-4c6f-82d8-1c05d1e651d4 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Depth anything: Unleashing the power of large-scale unlabeled data,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06e03206-93a2-498f-9b16-467cad8297a0 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Spectralgpt: Spectral remote sensing foun- dation model,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f31e8ee-f887-4a7a-96eb-47f4b62d8179 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Segment anything in medical images,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e83cc329-7a9f-4831-a30e-7d562858aaaa · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Point to set similarity based deep feature learning for person re-identification,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2a888963-af32-49ba-a627-5475c9bec640 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Visual-linguistic feature align- ment with semantic and kinematic guidance for referring multi-object tracking,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ca1d6cf8-70dc-426e-840d-56823da30168 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Explainability enhanced object detection transformer with feature disentanglement,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0992a42d-65ad-4aba-a10c-c5cd0f1dfdc8 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Deformable DETR: Deformable Transformers for End-to-End Object Detection
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9095b9ae-19f1-417b-93fe-41701a75a88f · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Fast temporal activity proposals for efficient detection of human actions in untrimmed videos,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 99bd5ad5-0eee-4749-8580-f048a6992829 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Turn tap: Temporal unit regression network for temporal action proposals,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2b2f46a3-1035-4d4e-a0ba-a3755a11b930 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Daps: Deep action proposals for action understanding,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3964fccf-844e-4393-8b9e-5d074e022e84 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Bmn: Boundary-matching network for temporal action proposal generation,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6a62d0aa-f909-4ba9-adb5-d9e489ea7341 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Bsn: Boundary sensitive network for temporal action proposal generation,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fff87b8a-cc00-4133-a777-d22d4d091316 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Temporal action detection with structured segment networks,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e99f9f09-0088-4c18-ab25-7fc4378b8618 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Bert: Pre-training of deep bidirectional transformers for language understanding,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26838943-bc32-48b3-a35b-11a34589037b · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Learning transferable visual models from natural language supervision,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2378daf-4e9e-4f59-8b20-ba5aae863ffe · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Attention is all you need,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dad2060a-0920-499d-8779-5c23d59dc7e8 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning The hungarian method for the assignment problem,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a2d00ab-e957-48be-85e3-7dd8d9569679 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Generalized intersection over union: A metric and a loss for bounding box regression,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28fb2958-0a4d-4bef-b95f-b53ab4c74fca · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Focal loss for dense object detection,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02b0983d-d264-4de7-a0a4-78a13078d2d4 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0abc37a-0e0c-421d-b03e-a92d76fd7b81 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning VideoChat: Chat-Centric Video Understanding
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf0a8937-3623-44e7-970f-abd7ac383a08 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Timechat: A time-sensitive multimodal large language model for long video understanding,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b96baa8c-f98f-4b68-aac6-179d53c9f81e · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning End-to-end Dense Video Captioning as Sequence Generation
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation add7b680-b287-4f8a-8d02-c4296cecaa0a · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Event-centric hier- archical representation for dense video captioning,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b5151edd-cabd-4bd3-8084-49f77d5a5779 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning End-to-end dense video captioning with masked transformer,
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 739b80c0-5372-437d-932c-593f7a4026f5 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Dense-Captioning Events in Videos: SYSU Submission to ActivityNet Challenge 2020
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93559e2d-a349-4524-9cea-c70b6032fe40 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Cider: Consensus- based image description evaluation,
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ef41eb8-1756-4316-bd3d-1e91b39e8633 · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Bleu: a method for automatic evaluation of machine translation,
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ddaf0f5-115d-49e5-8bf2-e35f1bde38af · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Meteor: An automatic metric for mt evalua- tion with improved correlation with human judgments,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d775db5-68d8-4dba-8e81-96629aa209ef · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Soda: Story oriented dense video captioning evaluation framework,
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 560e4972-7cc4-490a-b7fa-6460964bbfeb · outbound
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Move forward and tell: A progressive generator of video descriptions,
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
No inbound Pith citation observations are available.