Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:06:05.171159Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2411.14688.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:06:05.171159Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
77 of 77 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9fc1bf43-b827-4330-835c-6a2b45ab28c6 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Flamingo: a visual language model for few-shot learning,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58eda270-95ed-4593-814f-882abfb3470e · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c5f0066a-1638-40ab-93fa-266c263e0cce · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1281e890-ec4b-4755-ae78-099640dadade · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning BEiT: BERT Pre-Training of Image Transformers
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f418b2a-b1fa-4535-bbab-388e3ab1876d · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Recur- rent memory transformer
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4b1e360d-70a1-4761-97a5-f603fb9458dc · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Quo vadis, action recognition? a new model and the kinetics dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0933075a-1a6f-4b96-9f7e-6a0273406282 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning PaLI-X: On Scaling up a Multilingual Vision and Language Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61cdd851-19ad-4708-b862-1df2747700ec · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning PaLI: A jointly-scaled multilingual language- image model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 322233ca-9b62-4e1c-8271-0d32a8f6c89e · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning VideoOFA: Two-Stage Pre-Training for Video-to-Text Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 476c7f7b-ebdf-4ebc-b6aa-c61fb20c7ffa · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Uniter: Universal image-text representation learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation db2c8653-8ef1-4f0a-8d31-e135ea8b7ea5 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning TALLFormer: Temporal Action Localization with a Long-memory Transformer
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 62d8344d-51fd-4c6b-bae4-521dc7e62b3d · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Monotonic chunkwise attention
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2f2ed543-2ca6-41ab-bcd2-41e67c5c429a · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Vision Transformers Need Registers
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b6b29fb-8f0a-4264-899a-69afab6caf9c · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning An empirical study of training end-to-end vision-and-language transformers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 64219baf-12c1-41bb-b565-5477d750100b · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Violet: End-to-end video-language transformers with masked visual-token mod- eling
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8dd6f2c7-edaa-4208-9707-51c3bd6387fb · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Soda: Story oriented dense video captioning evaluation framework
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf836ebb-e58f-4baf-9c41-2d409e04da11 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answer- ing
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 779e2d54-5c99-4f82-aec2-59951c501850 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning The” something something” video database for learning and evaluating visual common sense
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cb1b91ff-dee4-45ba-bdee-9f5c0a287456 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Videollm: Modeling video sequence with large language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eb4d5fbf-0614-410f-b635-255f322addb1 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Multimodal pretraining for dense video cap- tioning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e534e57a-ec11-4924-91e7-549b3c17c99b · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning A better use of audio-visual cues: Dense video captioning with bi-modal transformer
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 28f5fd1d-56f8-48f4-a619-f5a5ca7e567b · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Long movie clip classification with state-space video models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2ebbb894-1c9c-47ea-a25e-cc09b23d77d6 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Perceiver: General perception with iterative attention, 2021
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 866920ec-e48c-4654-a1d0-197ef9477bd4 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Scaling up visual and vision-language representation learning with noisy text supervision
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 634c3a2a-8ba3-4870-b3e7-b5d1ada9c606 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning The Kinetics Human Action Video Dataset
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b2d3f33-9c37-433f-b9a0-d5471995577a · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Dense-captioning events in videos
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ab86b914-033e-4fe5-86ec-c876dbe04c34 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning MaMMUT: A simple architecture for joint learning for mul- timodal tasks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 93a06bfb-34e4-4d4f-81f5-7c577724b57e · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Selvaraju, Akhilesh D
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2073b075-6195-4090-a649-687ad2ea368c · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16497643-3609-4795-953f-103946ee1173 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Unmasked teacher: Towards training-efficient video foundation models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ccea7bdf-7641-42e0-b7ec-5256e04f62a9 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Oscar: Object-semantics aligned pre-training for vision-language tasks
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 259d49b3-f1ee-41eb-aaf1-d1695b76edfe · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Eclipse: Efficient long-range video retrieval using sight and sound
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 508faf4f-05af-4693-85a3-1d318c14ad52 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4a015e24-5651-49eb-89cf-2299cc5cf88a · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e37f6b3-5420-48b1-8ef4-4ef8f2d6b586 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 544c7b6c-8643-477f-b6c9-ee214541a8b9 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f99164ad-09ba-48b6-9d4f-52c978b8cdf9 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e7c291c4-96bf-4303-ba3f-4eaafe0b1ad6 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Moments in time dataset: one million videos for event understanding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 269d64be-8381-4c1f-bdfe-c09523415829 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Re- thinking video vits: Sparse video tubes for joint image and video learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6df7cbe3-23f8-48ea-8e27-a3e6b204789d · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Dynamic pretraining of vision-language models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4c1d06c5-315f-4441-b52c-d889c3de41f7 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Mirasol3B: A multi- modal autoregressive model for time-aligned and contextual modalities
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 70723b90-532a-4496-baa8-fe1dbfba2780 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Timechat: A time-sensitive multimodallarge language model for long video understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eecb69be-a250-46cb-a0e7-46e88aa337b4 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Ryoo, AJ Piergiovanni, Anurag Arnab, Mostafa Dehghani, and Anelia Angelova
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 682ac631-1b37-4e47-9e54-d74537cbd27b · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Ryoo, Keerthana Gopalakrishnan, Kumara Ka- hatapitiya, Ted Xiao, Kanishka Rao, Austin Stone, Yao Lu, Julian Ibarz, and Anurag Arnab
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fe32cb33-f5a8-42cc-9296-cf38edf05b5b · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Tridet: Temporal action detection with relative boundary modeling
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6873270e-a51b-4777-aa9c-5a4191ca1ef4 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Flava: A foundational language and vision alignment model
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc54361f-70fe-4445-813e-2264bf0e17bc · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Ucf101: A dataset of 101 human action classes from videos in the wild
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 26e32ad5-2ac1-43f4-ad44-734d59ebf866 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Long-form video-language pre- training with multimodal temporal contrastive learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 72a9ef8c-a6be-41f9-87f5-904215e960a3 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Lxmert: Learning cross- modality encoder representations from transformers
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eb9d6011-49f0-4598-a980-1aae07ec971d · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning CLIP4Caption: CLIP for Video Caption
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation abaf2d17-1bd9-45b9-b3fa-65160eba14cb · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Cider: Consensus-based image description evalua- tion
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84909b79-6f9a-438a-9985-b6e0e2125196 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Bidirectional attentive fusion with context gating for dense video captioning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ecb8d9d6-4849-4fc7-a288-f1d99e3102e3 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning GIT: A Generative Image-to-text Transformer for Vision and Language
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25f47987-40a1-46cb-8ca5-451b566e3853 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Omnivid: A generative framework for universal video understanding
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ad699342-6b75-4888-a562-47e8d54a4a59 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2844a8a3-3fc3-485b-b257-5efe1d665c05 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning End-to-end dense video captioning with parallel decoding
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3866e6e4-4cd8-4a15-8fd0-c16be4a0a892 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77202398-63d6-4180-bcfc-77319f3070f1 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c79ed25-29e0-4b4f-a0ea-a4ffcf36494a · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dedb013e-e3c7-44ee-9654-c3856e996084 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Vl-bert: Pre-training of generic visual- linguistic representations
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3a2d0625-250b-4b1d-a5ab-ee791bbe4902 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Towards long-form video understanding
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0ffc5d5d-b2d5-41a5-abc9-b57b2ba4d9fb · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Dibs: Enhanc- ing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 484716e0-2545-429b-98de-c9ffdd87a54e · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f11a6134-c2e5-4735-9bcf-4d2ba70a9829 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ea51806-a878-4d44-a798-21e8c1530627 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f72021f7-070e-4aaa-8326-44b87f8ca744 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Coca: Contrastive captioners are image-text foundation models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5448c75-05e4-445e-9a98-1004fab46069 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Hierarchical video-moment retrieval and step-captioning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7cd6d66b-6d22-4bc3-8426-b728a86d9794 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Mer- lot: Multimodal neural script knowledge models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3c8c9704-693c-4510-a8dd-0692858bcee5 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Actionformer: Lo- calizing moments of actions with transformers
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3faccc33-6290-4dae-ad00-4adfd538343c · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning VinVL: Revisiting Visual Representations in Vision-Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 113ed232-ae2d-4a3c-bdd6-b7a0a89bc616 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Unifying event detec- tion and captioning as sequence generation via pre-training
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 786bfcf7-53f3-4260-9b5d-ef716e770842 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Open-ended long-form video question answering via hierarchical convolutional self-attention networks
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d6f167a0-1b80-45b5-8251-af197a0a4cbf · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Towards automatic learning of procedures from web instructional videos
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 01ea6556-418e-491e-94a0-f39784588f12 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning End-to-end dense video captioning with masked transformer
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4b0e473b-4958-4e01-be53-f7d53ff8bc8b · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Streaming dense video captioning
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 981ea771-ddc3-4beb-b52f-4a1c4a08eb2e · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Towards Understanding Sample Variance in Visually Grounded Language Generation: Evaluations and Observations
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6919d268-601d-4346-ab9b-df06ec865e03 · outbound
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Thapliyal, William Yang Wang, and Radu Soricut
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
No inbound Pith citation observations are available.