Pith. sign in

Paper Citation Record · LEDGER

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning

As of 13 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2411.14688.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14688 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:06:05.171159Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact4
  • verified fuzzy47
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9fc1bf43-b827-4330-835c-6a2b45ab28c6 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Flamingo: a visual language model for few-shot learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.828160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.828160Z digest=sha256:5e924649d4b5d00d03533de92afe842422226c07d13e5b2dbc31f86b04d1269c

Observation 58eda270-95ed-4593-814f-882abfb3470e · outbound

This paper cites an unresolved cited work.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Unresolved cited work

Reference 2

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T15:06:06.426126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.833355Z digest=sha256:5d26a49885fd04d089927bdfe75423cb2b9e5dc713bba4a6b82d835f3324f9e9

Observation c5f0066a-1638-40ab-93fa-266c263e0cce · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.837965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.837965Z digest=sha256:64040ee530961086474cddfc68a7fce8653ba9a94c3a8bf10837b053bda936ac

Observation 1281e890-ec4b-4755-ae78-099640dadade · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning BEiT: BERT Pre-Training of Image Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.842945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.842945Z digest=sha256:54f4bac3515bbb950eff83a4f414e959ec6f471fbf50af32ace29eacce9af280

Observation 3f418b2a-b1fa-4535-bbab-388e3ab1876d · outbound

This paper cites Recur- rent memory transformer.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Recur- rent memory transformer

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.402403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.847643Z digest=sha256:6d69b7e78eae4631253f458f1a911eec8223f5e310ace212eb3eb427a8b61180

Observation 4b1e360d-70a1-4761-97a5-f603fb9458dc · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Quo vadis, action recognition? a new model and the kinetics dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.852261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.852261Z digest=sha256:d73cb26076bb70624b38e13f6cfdebf8533d0ee01cc348e4cd651cad462a53dc

Observation 0933075a-1a6f-4b96-9f7e-6a0273406282 · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.856994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.856994Z digest=sha256:771380ae743dc975c2da290633c4409fa00113ebdf90516f84139ea550cb0bf8

Observation 61cdd851-19ad-4708-b862-1df2747700ec · outbound

This paper cites PaLI: A jointly-scaled multilingual language- image model.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning PaLI: A jointly-scaled multilingual language- image model

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.378746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.861858Z digest=sha256:a251c3865c02181e21b27ee9984edf3031af17bee89f60b17201ce5baed1275c

Observation 322233ca-9b62-4e1c-8271-0d32a8f6c89e · outbound

This paper cites VideoOFA: Two-Stage Pre-Training for Video-to-Text Generation.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning VideoOFA: Two-Stage Pre-Training for Video-to-Text Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.866401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.866401Z digest=sha256:e178856c05e3a604941d4516602352952077745893276627bb137d4c27f7381e

Observation 476c7f7b-ebdf-4ebc-b6aa-c61fb20c7ffa · outbound

This paper cites Uniter: Universal image-text representation learning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Uniter: Universal image-text representation learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.364399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.871346Z digest=sha256:95399f62e839e01fa10c8aeb21fbf29db013288aeba875090c76b720afef9325

Observation db2c8653-8ef1-4f0a-8d31-e135ea8b7ea5 · outbound

This paper cites TALLFormer: Temporal Action Localization with a Long-memory Transformer.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning TALLFormer: Temporal Action Localization with a Long-memory Transformer

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:06:05.606074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.875901Z digest=sha256:9c6f4f3f7a1dd0c58550c69899743b978422dce4814ff2d212b0dddbd8d95290

Observation 62d8344d-51fd-4c6b-bae4-521dc7e62b3d · outbound

This paper cites Monotonic chunkwise attention.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Monotonic chunkwise attention

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.349731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.881088Z digest=sha256:9a92196b5c749deede4358a2be5b10a97cbd639397704a730d1f7ea202b46933

Observation 2f2ed543-2ca6-41ab-bcd2-41e67c5c429a · outbound

This paper cites Vision Transformers Need Registers.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Vision Transformers Need Registers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.885366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.885366Z digest=sha256:907a2b811e69eefa64f0e76390e804c5912f64b99fed1d20e334f79844383ecb

Observation 5b6b29fb-8f0a-4264-899a-69afab6caf9c · outbound

This paper cites An empirical study of training end-to-end vision-and-language transformers.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning An empirical study of training end-to-end vision-and-language transformers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.335076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.890245Z digest=sha256:aef3fda4241709128e65758909d54bbbce92b30fe06e86a4cc7a2df9e0cb416a

Observation 64219baf-12c1-41bb-b565-5477d750100b · outbound

This paper cites Violet: End-to-end video-language transformers with masked visual-token mod- eling.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Violet: End-to-end video-language transformers with masked visual-token mod- eling

Reference 15

Resolution
verified exact
raw_fallback, observed 2026-08-12T15:06:05.568804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.894546Z digest=sha256:4fcb993f1b6d3357bc1f2f6e1c7c778d4b140132f93b3aaef8043e2d487cd16d

Observation 8dd6f2c7-edaa-4208-9707-51c3bd6387fb · outbound

This paper cites Soda: Story oriented dense video captioning evaluation framework.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Soda: Story oriented dense video captioning evaluation framework

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.898868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.898868Z digest=sha256:989c0911a0f25bc5654116119039c360744b2bcf54658f300f4b0edb8218abe3

Observation bf836ebb-e58f-4baf-9c41-2d409e04da11 · outbound

This paper cites Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answer- ing.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answer- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.310920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.903286Z digest=sha256:baf6ff919825dc8e9eda6cba1e583aeac512dd839a0aa97eabc3b9ebcb0aef0c

Observation 779e2d54-5c99-4f82-aec2-59951c501850 · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning The” something something” video database for learning and evaluating visual common sense

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.296440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.907497Z digest=sha256:65f1354267d3fed4aff1fa8655f6ac3c7e679eba2bb399255d3a6265a819fe42

Observation cb1b91ff-dee4-45ba-bdee-9f5c0a287456 · outbound

This paper cites Videollm: Modeling video sequence with large language models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Videollm: Modeling video sequence with large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.282074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.911726Z digest=sha256:d1316209dc913325f3a04cc6a470a299b603002c7b0273146ccef9c4cdf2158f

Observation eb4d5fbf-0614-410f-b635-255f322addb1 · outbound

This paper cites Multimodal pretraining for dense video cap- tioning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Multimodal pretraining for dense video cap- tioning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.267858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.916036Z digest=sha256:27843cfb0bdedea73a545a74042493d227d9e147cba6df94a214942b3fe8532e

Observation e534e57a-ec11-4924-91e7-549b3c17c99b · outbound

This paper cites A better use of audio-visual cues: Dense video captioning with bi-modal transformer.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning A better use of audio-visual cues: Dense video captioning with bi-modal transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.253553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.920357Z digest=sha256:42c891fa9e1d67a8fabfa2c6e1dac68d7970041ad7906321f7140f4559da9ccb

Observation 28f5fd1d-56f8-48f4-a619-f5a5ca7e567b · outbound

This paper cites Long movie clip classification with state-space video models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Long movie clip classification with state-space video models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.238572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.924749Z digest=sha256:a550eac8f9ce841f86938228b3bdf7bd45a3164eaefa36b90ef6d04ce29c1542

Observation 2ebbb894-1c9c-47ea-a25e-cc09b23d77d6 · outbound

This paper cites Perceiver: General perception with iterative attention, 2021.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Perceiver: General perception with iterative attention, 2021

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.223759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.929930Z digest=sha256:f1c2ecef9759198715deff42ac9fce3dd6f810fcbce1a345d5817998f31ff7e8

Observation 866920ec-e48c-4654-a1d0-197ef9477bd4 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Scaling up visual and vision-language representation learning with noisy text supervision

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.207617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.934248Z digest=sha256:c7ea29788137c2f2b699530acf7b21cebb5d9206278f9cee729b305a5b750d5f

Observation 634c3a2a-8ba3-4870-b3e7-b5d1ada9c606 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning The Kinetics Human Action Video Dataset

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.938660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.938660Z digest=sha256:504e2bd9f40f011aebe9a1ab270d75a2c95c394325ec10730534cc0207f367c2

Observation 7b2d3f33-9c37-433f-b9a0-d5471995577a · outbound

This paper cites Dense-captioning events in videos.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Dense-captioning events in videos

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.192583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.943485Z digest=sha256:854505dd57bb311ee389cb4be078f494cf04f60587ad35301e04b2d02a9ef1a7

Observation ab86b914-033e-4fe5-86ec-c876dbe04c34 · outbound

This paper cites MaMMUT: A simple architecture for joint learning for mul- timodal tasks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning MaMMUT: A simple architecture for joint learning for mul- timodal tasks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.177370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.948145Z digest=sha256:0458064781f2f99be184929651d45e2646d74ef96670f47fc94e071a0c9dac5a

Observation 93a06bfb-34e4-4d4f-81f5-7c577724b57e · outbound

This paper cites Selvaraju, Akhilesh D.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Selvaraju, Akhilesh D

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.162380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.952500Z digest=sha256:255fc2ce75a60841609ac977441bda128d6e060dc901d5c6d5f8ef23bd367bc4

Observation 2073b075-6195-4090-a649-687ad2ea368c · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.956799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.956799Z digest=sha256:95cacba07df74e3a7f751901379d5ae5477463d67cf995b50aebae3edae427dc

Observation 16497643-3609-4795-953f-103946ee1173 · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Unmasked teacher: Towards training-efficient video foundation models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.146858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.961415Z digest=sha256:4282475f59a957ad76fe876781c9d3da0e900e4e50f3327729f2ba95966e2dbb

Observation ccea7bdf-7641-42e0-b7ec-5256e04f62a9 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.965854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.965854Z digest=sha256:e70361ad7945db361b308fd489dbfeea542aaf804b5f44c0eb9d2a19c153684f

Observation 259d49b3-f1ee-41eb-aaf1-d1695b76edfe · outbound

This paper cites Eclipse: Efficient long-range video retrieval using sight and sound.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Eclipse: Efficient long-range video retrieval using sight and sound

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.123020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.970155Z digest=sha256:dc5d4e325f61c80723f1ef2ba04cdf850a3f6940b8d3bdbf405d9b95760e60df

Observation 508faf4f-05af-4693-85a3-1d318c14ad52 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.108611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.974651Z digest=sha256:a10b9c56d2e016fb36c35bec73dd0f15606222053c24591f09be356f4908e3c8

Observation 4a015e24-5651-49eb-89cf-2299cc5cf88a · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.978924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.978924Z digest=sha256:51d9e9aa84c3ab19765fd4b2221856e077cce228de07fd40b78a0220716659cd

Observation 2e37f6b3-5420-48b1-8ef4-4ef8f2d6b586 · outbound

This paper cites UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.983697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.983697Z digest=sha256:8f1c40481682af6a2862ad958e60419927d9c4aa937b7ba1bcf20bea07ca16e1

Observation 544c7b6c-8643-477f-b6c9-ee214541a8b9 · outbound

This paper cites CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.988340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.988340Z digest=sha256:048e7391319ced9ed09a841e46ffbdae8b140a6baf424f718c6fac62848a3b43

Observation f99164ad-09ba-48b6-9d4f-52c978b8cdf9 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.094585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.992820Z digest=sha256:45146739c638ae8aba60485a7dba02bd9aed2e7d802879c377563e52c293d5da

Observation e7c291c4-96bf-4303-ba3f-4eaafe0b1ad6 · outbound

This paper cites Moments in time dataset: one million videos for event understanding.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Moments in time dataset: one million videos for event understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.080134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:04.997164Z digest=sha256:67727acfc25b29cc5900cf91a2fc727e1395995d6e840ea25a29210c3b52efb6

Observation 269d64be-8381-4c1f-bdfe-c09523415829 · outbound

This paper cites Re- thinking video vits: Sparse video tubes for joint image and video learning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Re- thinking video vits: Sparse video tubes for joint image and video learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.066168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.001535Z digest=sha256:25337b9ec56693bffc36af43b2664f4d2c1d90802e6377c80def134ea68cf4e0

Observation 6df7cbe3-23f8-48ea-8e27-a3e6b204789d · outbound

This paper cites Dynamic pretraining of vision-language models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Dynamic pretraining of vision-language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.050986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.005634Z digest=sha256:688ba9067c04354b4c6142b9370538759e8aac1f71e93fc9308e277c38e5db6f

Observation 4c1d06c5-315f-4441-b52c-d889c3de41f7 · outbound

This paper cites Mirasol3B: A multi- modal autoregressive model for time-aligned and contextual modalities.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Mirasol3B: A multi- modal autoregressive model for time-aligned and contextual modalities

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.035370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.010128Z digest=sha256:fe77f78085d9028f88fa7ce1d26310ed06fd20cbc9f84ba3de5e40d6a5fb96ce

Observation 70723b90-532a-4496-baa8-fe1dbfba2780 · outbound

This paper cites Timechat: A time-sensitive multimodallarge language model for long video understanding.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Timechat: A time-sensitive multimodallarge language model for long video understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.020522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.014494Z digest=sha256:de633d2d324057235d1d030e38de5a8af21647e539990abd35c895b49fe0a18a

Observation eecb69be-a250-46cb-a0e7-46e88aa337b4 · outbound

This paper cites Ryoo, AJ Piergiovanni, Anurag Arnab, Mostafa Dehghani, and Anelia Angelova.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Ryoo, AJ Piergiovanni, Anurag Arnab, Mostafa Dehghani, and Anelia Angelova

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.005351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.019001Z digest=sha256:c1a6296364783c02781ade28e74f57f0f74725cb2b88acbf45b953ad38cc61fe

Observation 682ac631-1b37-4e47-9e54-d74537cbd27b · outbound

This paper cites Ryoo, Keerthana Gopalakrishnan, Kumara Ka- hatapitiya, Ted Xiao, Kanishka Rao, Austin Stone, Yao Lu, Julian Ibarz, and Anurag Arnab.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Ryoo, Keerthana Gopalakrishnan, Kumara Ka- hatapitiya, Ted Xiao, Kanishka Rao, Austin Stone, Yao Lu, Julian Ibarz, and Anurag Arnab

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.990104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.023173Z digest=sha256:d578fb8744eb1480c0270bd893d30fb49e7d214d8dafdf24e94ca8bf165a0901

Observation fe32cb33-f5a8-42cc-9296-cf38edf05b5b · outbound

This paper cites Tridet: Temporal action detection with relative boundary modeling.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Tridet: Temporal action detection with relative boundary modeling

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.975563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.027594Z digest=sha256:0393fedcf0533324cf7d3b4e33cdc96bbf4c671c08f53a23370c991e7383e77c

Observation 6873270e-a51b-4777-aa9c-5a4191ca1ef4 · outbound

This paper cites Flava: A foundational language and vision alignment model.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Flava: A foundational language and vision alignment model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.032048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.032048Z digest=sha256:62b0ec289925be72cb78b96e825278ad9df484f3a4f0cd7666aeb1a6f00072f7

Observation cc54361f-70fe-4445-813e-2264bf0e17bc · outbound

This paper cites Ucf101: A dataset of 101 human action classes from videos in the wild.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Ucf101: A dataset of 101 human action classes from videos in the wild

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.952057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.036535Z digest=sha256:2e504317c9b19c5a59e4181217e11514b7a745e537496cb7f8c76faf76696b0d

Observation 26e32ad5-2ac1-43f4-ad44-734d59ebf866 · outbound

This paper cites Long-form video-language pre- training with multimodal temporal contrastive learning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Long-form video-language pre- training with multimodal temporal contrastive learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.937238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.040934Z digest=sha256:33cc0f8cde11ed5d5e36c86d1858a408d2206355611ea67212474a0f3dcce585

Observation 72a9ef8c-a6be-41f9-87f5-904215e960a3 · outbound

This paper cites Lxmert: Learning cross- modality encoder representations from transformers.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Lxmert: Learning cross- modality encoder representations from transformers

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.922326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.045382Z digest=sha256:625f1d284bccac38ce56b8008a7c256eedac2f602f838fbbfcb8307a5e07ea35

Observation eb9d6011-49f0-4598-a980-1aae07ec971d · outbound

This paper cites CLIP4Caption: CLIP for Video Caption.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning CLIP4Caption: CLIP for Video Caption

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:06:05.353751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.049820Z digest=sha256:fd6bbcd1190eb7bec33cc0edb56674939103ec3235ac89900770295fba72b05e

Observation abaf2d17-1bd9-45b9-b3fa-65160eba14cb · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Cider: Consensus-based image description evalua- tion

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.054676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.054676Z digest=sha256:02a1470982debd322c3e4980262fac7d15d764deef5fff1280738ae2eeb9399c

Observation 84909b79-6f9a-438a-9985-b6e0e2125196 · outbound

This paper cites Bidirectional attentive fusion with context gating for dense video captioning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Bidirectional attentive fusion with context gating for dense video captioning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.897753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.058944Z digest=sha256:8b96a2fe8efb0d5f0da255230ff5a0426ce8db6631e0b5f96460307e34d25c2f

Observation ecb8d9d6-4849-4fc7-a288-f1d99e3102e3 · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.063232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.063232Z digest=sha256:1058a9e290adccf5aeb7da14209698d7682c7d288e8f8602efcc9acd98814308

Observation 25f47987-40a1-46cb-8ca5-451b566e3853 · outbound

This paper cites Omnivid: A generative framework for universal video understanding.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Omnivid: A generative framework for universal video understanding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.883020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.067933Z digest=sha256:b866dd303cedba337009bec0a3aba0b6ff8f2f665e48dd9bc43f20fda7579640

Observation ad699342-6b75-4888-a562-47e8d54a4a59 · outbound

This paper cites OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.072176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.072176Z digest=sha256:dd89e5bfe946e80f97e117737882a708ad2dfc7a864d67da349986153e14c704

Observation 2844a8a3-3fc3-485b-b257-5efe1d665c05 · outbound

This paper cites End-to-end dense video captioning with parallel decoding.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning End-to-end dense video captioning with parallel decoding

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.867617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.076799Z digest=sha256:1f47a54aaceae7f90a90ab04473cb65f02f0b62262bb2a699af771ce14170a7e

Observation 3866e6e4-4cd8-4a15-8fd0-c16be4a0a892 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.081141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.081141Z digest=sha256:528f673dd3b8504c61209433ae4d69a81c5232268944a38e2151bbcab38a0149

Observation 77202398-63d6-4180-bcfc-77319f3070f1 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.085716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.085716Z digest=sha256:0d1228487bfbdc688af146fa881fcff06af15b5ab09eb907808671f576d9d3a8

Observation 7c79ed25-29e0-4b4f-a0ea-a4ffcf36494a · outbound

This paper cites SimVLM: Simple Visual Language Model Pretraining with Weak Supervision.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.090407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.090407Z digest=sha256:8563ea22456064b884967b67987d35304862b43f83426bdbae8885bdc2948917

Observation dedb013e-e3c7-44ee-9654-c3856e996084 · outbound

This paper cites Vl-bert: Pre-training of generic visual- linguistic representations.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Vl-bert: Pre-training of generic visual- linguistic representations

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.852537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.095077Z digest=sha256:27d5890f32ece8ebd5c832ea25dfe786fbdda36a8bc63bd804ab06351e9d3b67

Observation 3a2d0625-250b-4b1d-a5ab-ee791bbe4902 · outbound

This paper cites Towards long-form video understanding.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Towards long-form video understanding

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.838141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.099227Z digest=sha256:c13cf8927dd234fa0b9dbd5b5a4ec48e77acd11b2a3bfcc22fc29d34b60162fb

Observation 0ffc5d5d-b2d5-41a5-abc9-b57b2ba4d9fb · outbound

This paper cites Dibs: Enhanc- ing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Dibs: Enhanc- ing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.823791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.103524Z digest=sha256:e66df4aed5bb1eccb1fad883b0c8e957c7d3ad1c6648639c688489f24890b3d0

Observation 484716e0-2545-429b-98de-c9ffdd87a54e · outbound

This paper cites mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.108091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.108091Z digest=sha256:163fd9fb4db2b4e73ce0b12594cfcb4d678fc08c7a1616a41d92b05127abdaac

Observation f11a6134-c2e5-4735-9bcf-4d2ba70a9829 · outbound

This paper cites VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.112688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.112688Z digest=sha256:f839c1943591d07a294df94170b6e2225601c914e73513976a09bac082dc1495

Observation 4ea51806-a878-4d44-a798-21e8c1530627 · outbound

This paper cites Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.808297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.117231Z digest=sha256:bab0c1f149170a32b6d9d871af0f5e3361ddc5210cffe03d63989a1cafe0fb6f

Observation f72021f7-070e-4aaa-8326-44b87f8ca744 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Coca: Contrastive captioners are image-text foundation models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.121604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.121604Z digest=sha256:57d5712bc07b3504368461817c8cdef03059e237f52b39130f52324685e29fe5

Observation c5448c75-05e4-445e-9a98-1004fab46069 · outbound

This paper cites Hierarchical video-moment retrieval and step-captioning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Hierarchical video-moment retrieval and step-captioning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.783130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.126063Z digest=sha256:fe63e961f13fcf39d491f9de648e3a2cf01406c3d89a9e82ee55bcdbe7f5a131

Observation 7cd6d66b-6d22-4bc3-8426-b728a86d9794 · outbound

This paper cites Mer- lot: Multimodal neural script knowledge models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Mer- lot: Multimodal neural script knowledge models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.768846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.130393Z digest=sha256:662de1c2ce3be024222152524e67670fbd7b66529d3e465cdbf4695a38193688

Observation 3c8c9704-693c-4510-a8dd-0692858bcee5 · outbound

This paper cites Actionformer: Lo- calizing moments of actions with transformers.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Actionformer: Lo- calizing moments of actions with transformers

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.754096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.134742Z digest=sha256:a571cb4f379e465f6c06e6b038ad00109b9f96acb751ccc0208f93c9c5bdb323

Observation 3faccc33-6290-4dae-ad00-4adfd538343c · outbound

This paper cites VinVL: Revisiting Visual Representations in Vision-Language Models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning VinVL: Revisiting Visual Representations in Vision-Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.139120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.139120Z digest=sha256:ddf6b1f8a0fecad49d7abe6e2b86a95851a56f1baa270cc441180ff759a68d00

Observation 113ed232-ae2d-4a3c-bdd6-b7a0a89bc616 · outbound

This paper cites Unifying event detec- tion and captioning as sequence generation via pre-training.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Unifying event detec- tion and captioning as sequence generation via pre-training

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.738616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.143718Z digest=sha256:974d2addf6a56a10f58bb359de7f9f2d7a9f6270463290a326bf4b61a4c27275

Observation 786bfcf7-53f3-4260-9b5d-ef716e770842 · outbound

This paper cites Open-ended long-form video question answering via hierarchical convolutional self-attention networks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Open-ended long-form video question answering via hierarchical convolutional self-attention networks

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.722709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.148081Z digest=sha256:440c7639bec028348b01571838c3d170632cc905e3ae94389e16647b152cf031

Observation d6f167a0-1b80-45b5-8251-af197a0a4cbf · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Towards automatic learning of procedures from web instructional videos

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.708283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.152827Z digest=sha256:53acf8e63b51fd1406c6fe70eaba1661bf47eb6324a620aa3d04eab1a3668d4a

Observation 01ea6556-418e-491e-94a0-f39784588f12 · outbound

This paper cites End-to-end dense video captioning with masked transformer.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning End-to-end dense video captioning with masked transformer

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.693993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.157215Z digest=sha256:c1fcee63d8080bb6d5b6f91dc3f534cfa6c6b28da59bc967f89b1cbf3bf6c737

Observation 4b0e473b-4958-4e01-be53-f7d53ff8bc8b · outbound

This paper cites Streaming dense video captioning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Streaming dense video captioning

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.679149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.161886Z digest=sha256:7d7f810487f946988a3f6d8dc32e2e6bc8a64db9d4b3068a48a555e038543580

Observation 981ea771-ddc3-4beb-b52f-4a1c4a08eb2e · outbound

This paper cites Towards Understanding Sample Variance in Visually Grounded Language Generation: Evaluations and Observations.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Towards Understanding Sample Variance in Visually Grounded Language Generation: Evaluations and Observations

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:06:05.215806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.166374Z digest=sha256:0fbf2ac7861a23c95407c9bbbdfb9ba1fbeb8408e3e537f479a6a50d8c1d59d1

Observation 6919d268-601d-4346-ab9b-df06ec865e03 · outbound

This paper cites Thapliyal, William Yang Wang, and Radu Soricut.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Thapliyal, William Yang Wang, and Radu Soricut

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.664229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:06:05.171159Z digest=sha256:646df4577ae5a71a58d5aa937dd4a90b517c83e6681e6e56624d61a9a5a7b770

Pith citing papers

No inbound Pith citation observations are available.