Pith. sign in

Paper Citation Record · LEDGER

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning

As of 14 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2411.14688.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14688 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:06:05.171159Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact4
  • verified fuzzy47
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9fc1bf43-b827-4330-835c-6a2b45ab28c6 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Flamingo: a visual language model for few-shot learning,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.828160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.828160Z digest=sha256:5bf38c7da27da6e60e69ef2285104e7e3022c7c759ffcf8e9d58c8ee6467b54f

Observation 58eda270-95ed-4593-814f-882abfb3470e · outbound

This paper cites an unresolved cited work.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Unresolved cited work

Reference 2

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T15:06:06.426126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.833355Z digest=sha256:273bacd9ebb6d3d7873ef18bea6d81a156df8c7c27677cc5be6b6a4d97e6c7a8

Observation c5f0066a-1638-40ab-93fa-266c263e0cce · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.837965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.837965Z digest=sha256:551645ee2be8e5f75a449ea71f6f41d391992de1ecaa611711599baa87b32208

Observation 1281e890-ec4b-4755-ae78-099640dadade · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning BEiT: BERT Pre-Training of Image Transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.842945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.842945Z digest=sha256:4641dab1e3adc5f6f28a49489d636af4d3c047721ab851e54a924137d6980421

Observation 3f418b2a-b1fa-4535-bbab-388e3ab1876d · outbound

This paper cites Recur- rent memory transformer.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Recur- rent memory transformer

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.402403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.847643Z digest=sha256:fd1a13833f161cdb18bcef720dacb8b628d2e47137c2473cea75bf818868435c

Observation 4b1e360d-70a1-4761-97a5-f603fb9458dc · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Quo vadis, action recognition? a new model and the kinetics dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.852261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.852261Z digest=sha256:8f79ee64b2b6ef137379066fa6e7f823c9d7a8853ac9c76f655c393458bd9b68

Observation 0933075a-1a6f-4b96-9f7e-6a0273406282 · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.856994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.856994Z digest=sha256:c8952342077a07b44053f14005c696b2a4dcf44ba9eaa774c4080454d323b945

Observation 61cdd851-19ad-4708-b862-1df2747700ec · outbound

This paper cites PaLI: A jointly-scaled multilingual language- image model.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning PaLI: A jointly-scaled multilingual language- image model

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.378746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.861858Z digest=sha256:050adea5d37d6f1035ce90daeea895a468cb22640b90032faddbd89809a419ca

Observation 322233ca-9b62-4e1c-8271-0d32a8f6c89e · outbound

This paper cites VideoOFA: Two-Stage Pre-Training for Video-to-Text Generation.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning VideoOFA: Two-Stage Pre-Training for Video-to-Text Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.866401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.866401Z digest=sha256:74f695f56b3b3c50b61d375bfca6c5bda8a7e45e338e781269d5e60ad4babc0b

Observation 476c7f7b-ebdf-4ebc-b6aa-c61fb20c7ffa · outbound

This paper cites Uniter: Universal image-text representation learning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Uniter: Universal image-text representation learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.364399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.871346Z digest=sha256:ef017e735f66ae694b2c3fbe76cf8e296b9ee93222d050fee6fbd89945e93c54

Observation db2c8653-8ef1-4f0a-8d31-e135ea8b7ea5 · outbound

This paper cites TALLFormer: Temporal Action Localization with a Long-memory Transformer.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning TALLFormer: Temporal Action Localization with a Long-memory Transformer

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:06:05.606074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.875901Z digest=sha256:c61233778bc19f4a97474f7f188063aabc85b614124e8a932622c2a8ebd193a8

Observation 62d8344d-51fd-4c6b-bae4-521dc7e62b3d · outbound

This paper cites Monotonic chunkwise attention.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Monotonic chunkwise attention

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.349731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.881088Z digest=sha256:367c8c92aa16bb5d77e7be94aa5ad472e92c91f1da3408dd0317d882695c889b

Observation 2f2ed543-2ca6-41ab-bcd2-41e67c5c429a · outbound

This paper cites Vision Transformers Need Registers.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Vision Transformers Need Registers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.885366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.885366Z digest=sha256:59363c5b2ff3c7765e8e653e67f29e79e86557a1ad06a1aa77a1e1aefb18f29c

Observation 5b6b29fb-8f0a-4264-899a-69afab6caf9c · outbound

This paper cites An empirical study of training end-to-end vision-and-language transformers.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning An empirical study of training end-to-end vision-and-language transformers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.335076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.890245Z digest=sha256:63c54fcb77ce295114990e4bd7d3b5ebc0594a653e3d01697d50a5f4766cae45

Observation 64219baf-12c1-41bb-b565-5477d750100b · outbound

This paper cites Violet: End-to-end video-language transformers with masked visual-token mod- eling.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Violet: End-to-end video-language transformers with masked visual-token mod- eling

Reference 15

Resolution
verified exact
raw_fallback, observed 2026-08-12T15:06:05.568804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.894546Z digest=sha256:5ac42e72cdcd3c2359ba706bacca884a5b822c2339e3523e9f1d29cadac4b234

Observation 8dd6f2c7-edaa-4208-9707-51c3bd6387fb · outbound

This paper cites Soda: Story oriented dense video captioning evaluation framework.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Soda: Story oriented dense video captioning evaluation framework

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.898868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.898868Z digest=sha256:452f5f06747a4920a52cddf3c9d41946dffb8ee9524b2e54a5c8f8a65e152850

Observation bf836ebb-e58f-4baf-9c41-2d409e04da11 · outbound

This paper cites Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answer- ing.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answer- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.310920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.903286Z digest=sha256:3191c4b1b7400593bacaac610136d5558a10703f2199b2e3bad58feec47faea7

Observation 779e2d54-5c99-4f82-aec2-59951c501850 · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning The” something something” video database for learning and evaluating visual common sense

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.296440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.907497Z digest=sha256:0de1294d7d5e0b2d86719e3eaa48a4a8fcdcf2fd9bc2d538d6e2deaeeec4b513

Observation cb1b91ff-dee4-45ba-bdee-9f5c0a287456 · outbound

This paper cites Videollm: Modeling video sequence with large language models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Videollm: Modeling video sequence with large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.282074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.911726Z digest=sha256:1828e31ba49a090e6c3affeb913683d594659c3bb17b684d6a14ba4d5df85f6d

Observation eb4d5fbf-0614-410f-b635-255f322addb1 · outbound

This paper cites Multimodal pretraining for dense video cap- tioning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Multimodal pretraining for dense video cap- tioning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.267858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.916036Z digest=sha256:226e5d7c759e91e61e47581e348989cec736657c8076f35bca85c610769ce67a

Observation e534e57a-ec11-4924-91e7-549b3c17c99b · outbound

This paper cites A better use of audio-visual cues: Dense video captioning with bi-modal transformer.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning A better use of audio-visual cues: Dense video captioning with bi-modal transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.253553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.920357Z digest=sha256:3e77e40d0c80521e7839dcb64f40f4907972520a05da44e1cdda5f4c645422eb

Observation 28f5fd1d-56f8-48f4-a619-f5a5ca7e567b · outbound

This paper cites Long movie clip classification with state-space video models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Long movie clip classification with state-space video models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.238572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.924749Z digest=sha256:dd0327d658c750128bde60e9db1f782d263ce83c7429cb3cb764a771f182f4cb

Observation 2ebbb894-1c9c-47ea-a25e-cc09b23d77d6 · outbound

This paper cites Perceiver: General perception with iterative attention, 2021.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Perceiver: General perception with iterative attention, 2021

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.223759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.929930Z digest=sha256:b6060fbeb5e8ee044aaaf6b1caebd2d5fc122a39b5f76bfe6a8a689209a1241f

Observation 866920ec-e48c-4654-a1d0-197ef9477bd4 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Scaling up visual and vision-language representation learning with noisy text supervision

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.207617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.934248Z digest=sha256:b155242f2067538b520b419a8f5a6834c1d6737d5084da32b8433080afe64558

Observation 634c3a2a-8ba3-4870-b3e7-b5d1ada9c606 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning The Kinetics Human Action Video Dataset

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.938660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.938660Z digest=sha256:c08c44a47470ad72a0a3352060bfd38ef6ee14893983c05ce8c106d7afb47082

Observation 7b2d3f33-9c37-433f-b9a0-d5471995577a · outbound

This paper cites Dense-captioning events in videos.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Dense-captioning events in videos

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.192583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.943485Z digest=sha256:7ebe3053c20c66121958d191b98259b094878e2aadac7f445efb43a22b85c64b

Observation ab86b914-033e-4fe5-86ec-c876dbe04c34 · outbound

This paper cites MaMMUT: A simple architecture for joint learning for mul- timodal tasks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning MaMMUT: A simple architecture for joint learning for mul- timodal tasks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.177370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.948145Z digest=sha256:7b566572153ddcafc253e4b257ce5bf50a08d4fea6dd9f2730dc555c04623b87

Observation 93a06bfb-34e4-4d4f-81f5-7c577724b57e · outbound

This paper cites Selvaraju, Akhilesh D.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Selvaraju, Akhilesh D

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.162380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.952500Z digest=sha256:af78fc79fc6af89ff40cc145bd37d8b72d7e736c0799e1fae3a00b62f1bdabf9

Observation 2073b075-6195-4090-a649-687ad2ea368c · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.956799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.956799Z digest=sha256:e392684de28dd14de71bfd5e8ba870f5f35640b6caef8d3d1cd22bb290bd570e

Observation 16497643-3609-4795-953f-103946ee1173 · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Unmasked teacher: Towards training-efficient video foundation models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.146858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.961415Z digest=sha256:403d526688eacba95f054f527783f95b4984f999e3ccb66525eb1e25dcae5d49

Observation ccea7bdf-7641-42e0-b7ec-5256e04f62a9 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.965854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.965854Z digest=sha256:95c339d2c83145a3497de673dc434b4270710bfaa89e5edbe2916806ef389a62

Observation 259d49b3-f1ee-41eb-aaf1-d1695b76edfe · outbound

This paper cites Eclipse: Efficient long-range video retrieval using sight and sound.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Eclipse: Efficient long-range video retrieval using sight and sound

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.123020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.970155Z digest=sha256:c8126b7271a1db2e04645a65a83a46513219e03a19a51a2281f66fb955ca335b

Observation 508faf4f-05af-4693-85a3-1d318c14ad52 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.108611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.974651Z digest=sha256:cd5f19c076b35f15a010ec373b11285d92d5da66222681c480fe8a1a09ef9299

Observation 4a015e24-5651-49eb-89cf-2299cc5cf88a · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.978924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.978924Z digest=sha256:56bb3093f89b118ac7b87941238500841ea13e2b7eef387a4a41948266e1089b

Observation 2e37f6b3-5420-48b1-8ef4-4ef8f2d6b586 · outbound

This paper cites UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.983697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.983697Z digest=sha256:12e2a701e1b7679509c596ac58b4d744f7acc9d5b344638997b506aa0dc759a2

Observation 544c7b6c-8643-477f-b6c9-ee214541a8b9 · outbound

This paper cites CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:04.988340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:04.988340Z digest=sha256:cf8d195b4bf7f2d6243a79855b5ba8af48de4aa324c4576a3d5d04826a233dda

Observation f99164ad-09ba-48b6-9d4f-52c978b8cdf9 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.094585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.992820Z digest=sha256:34bbe354f54d739aaedc354f85f54600cbf29d26852556c206e481af341a3a69

Observation e7c291c4-96bf-4303-ba3f-4eaafe0b1ad6 · outbound

This paper cites Moments in time dataset: one million videos for event understanding.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Moments in time dataset: one million videos for event understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.080134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:04.997164Z digest=sha256:f4c0166c6dc8e2fa0521084c2999eab3de2f1f682e041f6a29f97ee81c61894b

Observation 269d64be-8381-4c1f-bdfe-c09523415829 · outbound

This paper cites Re- thinking video vits: Sparse video tubes for joint image and video learning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Re- thinking video vits: Sparse video tubes for joint image and video learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.066168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.001535Z digest=sha256:273da59c9daef501cff182b2013dcc8a9aa7d4845cf795619adf379432657c2e

Observation 6df7cbe3-23f8-48ea-8e27-a3e6b204789d · outbound

This paper cites Dynamic pretraining of vision-language models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Dynamic pretraining of vision-language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.050986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.005634Z digest=sha256:9b78ee4ce63a08683f713dafc3f85ff01e2217ed4a98d4e15a6238d494875046

Observation 4c1d06c5-315f-4441-b52c-d889c3de41f7 · outbound

This paper cites Mirasol3B: A multi- modal autoregressive model for time-aligned and contextual modalities.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Mirasol3B: A multi- modal autoregressive model for time-aligned and contextual modalities

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.035370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.010128Z digest=sha256:1c9c81db9948bd0e7c7a20d3b8966d05a605cdd37713872fdade077e57962877

Observation 70723b90-532a-4496-baa8-fe1dbfba2780 · outbound

This paper cites Timechat: A time-sensitive multimodallarge language model for long video understanding.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Timechat: A time-sensitive multimodallarge language model for long video understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.020522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.014494Z digest=sha256:f7d0689608518b971477606a5884b8bd9c60e3279289f63c53e6cb6f59ade720

Observation eecb69be-a250-46cb-a0e7-46e88aa337b4 · outbound

This paper cites Ryoo, AJ Piergiovanni, Anurag Arnab, Mostafa Dehghani, and Anelia Angelova.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Ryoo, AJ Piergiovanni, Anurag Arnab, Mostafa Dehghani, and Anelia Angelova

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:06.005351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.019001Z digest=sha256:41ad1236f39447acb410233354603dd3444bb65dae09c0cdc0ed793112602aa2

Observation 682ac631-1b37-4e47-9e54-d74537cbd27b · outbound

This paper cites Ryoo, Keerthana Gopalakrishnan, Kumara Ka- hatapitiya, Ted Xiao, Kanishka Rao, Austin Stone, Yao Lu, Julian Ibarz, and Anurag Arnab.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Ryoo, Keerthana Gopalakrishnan, Kumara Ka- hatapitiya, Ted Xiao, Kanishka Rao, Austin Stone, Yao Lu, Julian Ibarz, and Anurag Arnab

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.990104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.023173Z digest=sha256:24494f8ae8227900adb6b6817fb5f7b690822a132722f86e2f51cf340ae75c7a

Observation fe32cb33-f5a8-42cc-9296-cf38edf05b5b · outbound

This paper cites Tridet: Temporal action detection with relative boundary modeling.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Tridet: Temporal action detection with relative boundary modeling

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.975563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.027594Z digest=sha256:e5fddfbbb09a4fe88036aecca4617ccf974be404f78161d961c225737e86aa8a

Observation 6873270e-a51b-4777-aa9c-5a4191ca1ef4 · outbound

This paper cites Flava: A foundational language and vision alignment model.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Flava: A foundational language and vision alignment model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.032048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.032048Z digest=sha256:7dc3fef25dc83f1cfda010fff64c274abbfe7ef7626ff2cd6effed2b8e1b6412

Observation cc54361f-70fe-4445-813e-2264bf0e17bc · outbound

This paper cites Ucf101: A dataset of 101 human action classes from videos in the wild.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Ucf101: A dataset of 101 human action classes from videos in the wild

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.952057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.036535Z digest=sha256:8cb93c070d57be7569099b50ea5ce7307b3ecc87e5349b8c4040e306addbf7c9

Observation 26e32ad5-2ac1-43f4-ad44-734d59ebf866 · outbound

This paper cites Long-form video-language pre- training with multimodal temporal contrastive learning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Long-form video-language pre- training with multimodal temporal contrastive learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.937238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.040934Z digest=sha256:999e51b2709b2c74d3c13b83ed7c5e3cc615d96608c1c76528240aee7fb7f144

Observation 72a9ef8c-a6be-41f9-87f5-904215e960a3 · outbound

This paper cites Lxmert: Learning cross- modality encoder representations from transformers.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Lxmert: Learning cross- modality encoder representations from transformers

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.922326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.045382Z digest=sha256:17a42024e7f0b8cf6648f2ce53cff0128f38a2fa3b2518442c886f0f2c8e0c2a

Observation eb9d6011-49f0-4598-a980-1aae07ec971d · outbound

This paper cites CLIP4Caption: CLIP for Video Caption.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning CLIP4Caption: CLIP for Video Caption

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:06:05.353751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.049820Z digest=sha256:26b7b447ce5c19610658eda460bcad4fffda7f51188ebab6f9d52c20ca51c6d7

Observation abaf2d17-1bd9-45b9-b3fa-65160eba14cb · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Cider: Consensus-based image description evalua- tion

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.054676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.054676Z digest=sha256:260f940c16ea9aa64c6d47de9e96daeecffaaaddfc12bbaff9d9ff047eb1fae1

Observation 84909b79-6f9a-438a-9985-b6e0e2125196 · outbound

This paper cites Bidirectional attentive fusion with context gating for dense video captioning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Bidirectional attentive fusion with context gating for dense video captioning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.897753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.058944Z digest=sha256:4bce1b87680211978f74cf141114701e969182c9d795d53a5de7c692d5c74bd1

Observation ecb8d9d6-4849-4fc7-a288-f1d99e3102e3 · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.063232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.063232Z digest=sha256:7ad65116852052ce65a898e12122f5ec66578f10110c5b055910afa21c2a9768

Observation 25f47987-40a1-46cb-8ca5-451b566e3853 · outbound

This paper cites Omnivid: A generative framework for universal video understanding.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Omnivid: A generative framework for universal video understanding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.883020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.067933Z digest=sha256:533825eabf56bafabb7249433acbda9214d8691396dccd7b268e96e30dbaccc3

Observation ad699342-6b75-4888-a562-47e8d54a4a59 · outbound

This paper cites OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.072176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.072176Z digest=sha256:5971f382a5c521f2fa970f364c37443cda09207b22296de24e07003fcdfd0c69

Observation 2844a8a3-3fc3-485b-b257-5efe1d665c05 · outbound

This paper cites End-to-end dense video captioning with parallel decoding.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning End-to-end dense video captioning with parallel decoding

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.867617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.076799Z digest=sha256:5f300cadb41831d0eac6d029644dde489656a54cd1df5615a516e0f87ddd8143

Observation 3866e6e4-4cd8-4a15-8fd0-c16be4a0a892 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.081141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.081141Z digest=sha256:8c266822b3ca7b72d3f8e426805d9d092761fad959dfe91c2871bd854b521ba9

Observation 77202398-63d6-4180-bcfc-77319f3070f1 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.085716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.085716Z digest=sha256:53247276382f7585d1a9764507dda0e93d39a2639b886c8bdfe37bbb7ca3cab5

Observation 7c79ed25-29e0-4b4f-a0ea-a4ffcf36494a · outbound

This paper cites SimVLM: Simple Visual Language Model Pretraining with Weak Supervision.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.090407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.090407Z digest=sha256:cd6f84fd950cebb87231212a5c254c3fa97bd3797d55514a846bb90d86dc23ba

Observation dedb013e-e3c7-44ee-9654-c3856e996084 · outbound

This paper cites Vl-bert: Pre-training of generic visual- linguistic representations.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Vl-bert: Pre-training of generic visual- linguistic representations

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.852537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.095077Z digest=sha256:8d32347f125632c857e66c6f6c01c23bcbfae44ca3538e2df568fb4d420afa0d

Observation 3a2d0625-250b-4b1d-a5ab-ee791bbe4902 · outbound

This paper cites Towards long-form video understanding.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Towards long-form video understanding

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.838141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.099227Z digest=sha256:4ac4ba404f8c1dee668a6d7acfab40d692eeac4d60b067e90c087e82f79b6e7d

Observation 0ffc5d5d-b2d5-41a5-abc9-b57b2ba4d9fb · outbound

This paper cites Dibs: Enhanc- ing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Dibs: Enhanc- ing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.823791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.103524Z digest=sha256:86f065ac7ee453069149b813d9f94cd7dc21123e97aac26c23442e15812cb19e

Observation 484716e0-2545-429b-98de-c9ffdd87a54e · outbound

This paper cites mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.108091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.108091Z digest=sha256:fc7a5a2285936631663c0d6ae8b08c245ac4edbac46929750d5225c73062f9fb

Observation f11a6134-c2e5-4735-9bcf-4d2ba70a9829 · outbound

This paper cites VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.112688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.112688Z digest=sha256:bee76ae2e6bc3b2f56c0957088cda8b7554e6278dc724e345feb3b6fb54ad30c

Observation 4ea51806-a878-4d44-a798-21e8c1530627 · outbound

This paper cites Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.808297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.117231Z digest=sha256:1f220f1898efd125f829515cacb64f9cd0cf1633c935e2343eba2dad60feed5c

Observation f72021f7-070e-4aaa-8326-44b87f8ca744 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Coca: Contrastive captioners are image-text foundation models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.121604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.121604Z digest=sha256:42bf549088e0ee4d64ac7f39b00039cd4a49888b8ee53db10231cdadf5a88073

Observation c5448c75-05e4-445e-9a98-1004fab46069 · outbound

This paper cites Hierarchical video-moment retrieval and step-captioning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Hierarchical video-moment retrieval and step-captioning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.783130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.126063Z digest=sha256:0e2d19b617c31be61ac21e4ce97411b87de8e5e4846c4a4417e6d15fb0afd0eb

Observation 7cd6d66b-6d22-4bc3-8426-b728a86d9794 · outbound

This paper cites Mer- lot: Multimodal neural script knowledge models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Mer- lot: Multimodal neural script knowledge models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.768846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.130393Z digest=sha256:50673f7a4b8693050e7427c60fe1007e6cc65f617d48c8182148d588af196ede

Observation 3c8c9704-693c-4510-a8dd-0692858bcee5 · outbound

This paper cites Actionformer: Lo- calizing moments of actions with transformers.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Actionformer: Lo- calizing moments of actions with transformers

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.754096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.134742Z digest=sha256:b4134e971f4a7727e56577290225e541d72a9524026584653b0adc2f91cd82dd

Observation 3faccc33-6290-4dae-ad00-4adfd538343c · outbound

This paper cites VinVL: Revisiting Visual Representations in Vision-Language Models.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning VinVL: Revisiting Visual Representations in Vision-Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.139120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.139120Z digest=sha256:ec0f8d0e7f922b78926f7890a63ac22512711d2ef722e587c00f90f00067f79a

Observation 113ed232-ae2d-4a3c-bdd6-b7a0a89bc616 · outbound

This paper cites Unifying event detec- tion and captioning as sequence generation via pre-training.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Unifying event detec- tion and captioning as sequence generation via pre-training

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.738616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.143718Z digest=sha256:2e63dec3356a112be39258db008d7447e1f421b55f66c650ab5cb68629d4b905

Observation 786bfcf7-53f3-4260-9b5d-ef716e770842 · outbound

This paper cites Open-ended long-form video question answering via hierarchical convolutional self-attention networks.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Open-ended long-form video question answering via hierarchical convolutional self-attention networks

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.722709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.148081Z digest=sha256:ed81b6e72a0e56ecb2d494743ca2ba2b7895935af77cfe3e6e6b0abd588388af

Observation d6f167a0-1b80-45b5-8251-af197a0a4cbf · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Towards automatic learning of procedures from web instructional videos

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.708283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.152827Z digest=sha256:cebc5d72fa4916ecdb9e387f14857098bd9fa890798d767cc6cfe97af1f563eb

Observation 01ea6556-418e-491e-94a0-f39784588f12 · outbound

This paper cites End-to-end dense video captioning with masked transformer.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning End-to-end dense video captioning with masked transformer

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.693993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.157215Z digest=sha256:b5241c25ac36451f096183b2ed68117923eb69d16fe5f8927bdc67931ae4284b

Observation 4b0e473b-4958-4e01-be53-f7d53ff8bc8b · outbound

This paper cites Streaming dense video captioning.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Streaming dense video captioning

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.679149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.161886Z digest=sha256:a3ebdc1247d08636d580cbbc347b15b2dd42669c2da81b4c2ac4969a6a6e5329

Observation 981ea771-ddc3-4beb-b52f-4a1c4a08eb2e · outbound

This paper cites Towards Understanding Sample Variance in Visually Grounded Language Generation: Evaluations and Observations.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Towards Understanding Sample Variance in Visually Grounded Language Generation: Evaluations and Observations

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:06:05.215806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.166374Z digest=sha256:bdaa8079e414855d6d58b6e7b926b1ef624e3bcc535c0f8861d43ca6dead2179

Observation 6919d268-601d-4346-ab9b-df06ec865e03 · outbound

This paper cites Thapliyal, William Yang Wang, and Radu Soricut.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning Thapliyal, William Yang Wang, and Radu Soricut

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:06:05.664229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:06:05.171159Z digest=sha256:c1fc627d5b2ec9ebad668f01c449a5f36e6b3c54baaa341ebacdf3fc0be8e613

Pith citing papers

No inbound Pith citation observations are available.