Pith. sign in

Paper Citation Record · LEDGER

DistinctAD: Distinctive Audio Description Generation in Contexts

As of 13 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 0 inbound Pith citation observations for arXiv:2411.18180.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18180 v1

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:30:20.484216Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

88 of 88 outbound references displayed

  • verified exact2
  • verified fuzzy56
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02369d82-26c8-4448-8328-7213f3562022 · outbound

This paper cites https://github.

DistinctAD: Distinctive Audio Description Generation in Contexts https://github

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.027395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.027395Z digest=sha256:6669c4b5ca04b1714881dc093d59ac5f1abb807b5151b32b41302371840c78ed

Observation eb38ada4-d5fa-40e8-a7e9-e286e73ee22b · outbound

This paper cites GPT-4 Technical Report.

DistinctAD: Distinctive Audio Description Generation in Contexts GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.033060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.033060Z digest=sha256:33cdd41852ae5c13037edf74de05dd8fcb89da033da825e7792cb963828959da

Observation 4fb03a4b-c28f-42c8-a36f-240869edfd83 · outbound

This paper cites Llama 3 model card.

DistinctAD: Distinctive Audio Description Generation in Contexts Llama 3 model card

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.038505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.038505Z digest=sha256:2ca9bb6c2214e2a98ae93b03f6aedad98b317577190cee074212b2be1bedf17e

Observation d7808cea-a78d-4ccf-9e44-d48708a49cd3 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

DistinctAD: Distinctive Audio Description Generation in Contexts Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.043031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.043031Z digest=sha256:654f2f48808caa6c0f180831ffb6ddcca0fe165d10e2b4cf6ef079de00082364

Observation f17910fa-f342-4a37-a905-be84e39eea45 · outbound

This paper cites Spice: Semantic propositional image cap- tion evaluation.

DistinctAD: Distinctive Audio Description Generation in Contexts Spice: Semantic propositional image cap- tion evaluation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.048062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.048062Z digest=sha256:435c371f74abd1a70d04351a72ca9ceb72b1ad490414a6c35a293d8b18010af2

Observation 3fb462c2-41ef-4a96-a49e-6d361bfa05e6 · outbound

This paper cites Condensed movies: Story based retrieval with con- textual embeddings.

DistinctAD: Distinctive Audio Description Generation in Contexts Condensed movies: Story based retrieval with con- textual embeddings

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.053508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.053508Z digest=sha256:a19bffc1082b6589c99ba4ec3167f0f5cc0da886463a5345900a975cdbbdf6cc

Observation 9e3b12e7-a13d-421b-9677-d07840be8bf3 · outbound

This paper cites WhisperX: Time-Accurate Speech Transcription of Long-Form Audio.

DistinctAD: Distinctive Audio Description Generation in Contexts WhisperX: Time-Accurate Speech Transcription of Long-Form Audio

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.058751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.058751Z digest=sha256:2ebc98455fa3619334296311aebd60d2198d3a997c49cdab51ac8ac69a655fb7

Observation 076c5170-ea3d-43ec-a676-3442d7555ced · outbound

This paper cites Livedescribe: can amateur describers create high-quality audio description? Journal of Visual Impairment & Blindness, 106(3):154–165,.

DistinctAD: Distinctive Audio Description Generation in Contexts Livedescribe: can amateur describers create high-quality audio description? Journal of Visual Impairment & Blindness, 106(3):154–165,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.064337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.064337Z digest=sha256:770d9566425db4268499dc8e2b088beb8ec5baa260c6cc231397cdb3112ebc39

Observation c7af109e-deea-41ea-9e70-80cbe62e5a6e · outbound

This paper cites End-to-end speaker seg- mentation for overlap-aware resegmentation.

DistinctAD: Distinctive Audio Description Generation in Contexts End-to-end speaker seg- mentation for overlap-aware resegmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.069791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.069791Z digest=sha256:cdb9a5d2c2bc3a175d758977530898c9dad80a85abdc9d5d61d7c63a49f5cc1f

Observation 3f495fa8-dcb0-40bb-95ca-2b6b659f319d · outbound

This paper cites Pyannote.

DistinctAD: Distinctive Audio Description Generation in Contexts Pyannote

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.075439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.075439Z digest=sha256:3a853495e9e844f839e6d6f377ae6523ce8b44e585f4af8a940122e3899b8d8c

Observation 343971bb-4cce-468e-9f84-5981aadac8a6 · outbound

This paper cites Groupcap: Group-based image captioning with structured relevance and diversity constraints.

DistinctAD: Distinctive Audio Description Generation in Contexts Groupcap: Group-based image captioning with structured relevance and diversity constraints

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.751384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.080873Z digest=sha256:adb5ed74d49b060d25839f31391af26a77b30c4e4f0857f9228066986a39a765

Observation c4d87e84-4245-486e-b5e9-a29d8bc02f6b · outbound

This paper cites Towards bridging event captioner and sentence localizer for weakly supervised dense event captioning.

DistinctAD: Distinctive Audio Description Generation in Contexts Towards bridging event captioner and sentence localizer for weakly supervised dense event captioning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.736099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.086187Z digest=sha256:bddf4f734dd384e03995b41d5bcd394ef8939d66e7463b915bde763203029a5e

Observation febee165-4f08-4e3a-a0a8-c454cf5c36c7 · outbound

This paper cites LLM-AD: Large Language Model based Audio Description System.

DistinctAD: Distinctive Audio Description Generation in Contexts LLM-AD: Large Language Model based Audio Description System

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.092465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.092465Z digest=sha256:ea2b580d99f66da0c501d839d117ba8aa8bcd960cdaa1f37bc60b30e26ac561a

Observation 95a61deb-f6e9-4959-b44c-01c58659e772 · outbound

This paper cites The ides of march.

DistinctAD: Distinctive Audio Description Generation in Contexts The ides of march

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.720897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.097964Z digest=sha256:e1f9d7bac3a335d56a497dd276ade8859a29c3397977e3a4fd0c763b6213227a

Observation 7d9787e5-9fbd-43d6-bab4-ecd5d9b3e214 · outbound

This paper cites Contrastive learning for image cap- tioning.

DistinctAD: Distinctive Audio Description Generation in Contexts Contrastive learning for image cap- tioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.706567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.103431Z digest=sha256:16f2deb20a308c5f4498f0d8521ab9da1a64780fdfde7d82581c15a5b9c2543e

Observation 43ab6fd3-6bb2-4d83-a571-6a99b8739eb2 · outbound

This paper cites Maximum likelihood from incomplete data via the em al- gorithm.

DistinctAD: Distinctive Audio Description Generation in Contexts Maximum likelihood from incomplete data via the em al- gorithm

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.691399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.108757Z digest=sha256:5d0598b879ddb06ffd4bba271884f7bca9c04423e5c6fd76dc5617b70956caf8

Observation 62983dfc-f032-44d7-bc9d-696460ca8dd3 · outbound

This paper cites Sketch, ground, and refine: Top-down dense video caption- ing.

DistinctAD: Distinctive Audio Description Generation in Contexts Sketch, ground, and refine: Top-down dense video caption- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.676885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.113915Z digest=sha256:c159b6223edc5d1b92e1a994cb3e4042752510f8e197d24ddb5bd72904db40a2

Observation 6b948ab6-4590-4931-9d80-9582089132b1 · outbound

This paper cites Describing differences in image sets with natural language.

DistinctAD: Distinctive Audio Description Generation in Contexts Describing differences in image sets with natural language

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.662408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.119583Z digest=sha256:6a3db663aac2e17985c56ec6039dc8e87a3280ee3390fa9021f1c9332ac8d9ca

Observation f8138907-0ea8-4c08-a8a8-2d55280fd1c7 · outbound

This paper cites An introduction to audio description: A prac- tical guide, 2016.

DistinctAD: Distinctive Audio Description Generation in Contexts An introduction to audio description: A prac- tical guide, 2016

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.647408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.125338Z digest=sha256:040b79fdb842d45d069f2dc77399cf137240f73b6c78a144d1c2549030ad209d

Observation 2200baab-6ac4-4870-9f70-d4dffe0e3e65 · outbound

This paper cites Autoad: Movie description in context.

DistinctAD: Distinctive Audio Description Generation in Contexts Autoad: Movie description in context

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.632668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.130735Z digest=sha256:5fbb9afb181387c8eb5b12dad7a5c99acaa42e8f22cdf0c63501602df3a7cb73

Observation 2321d48f-f5e8-4f29-a8d1-8b8d0623b92a · outbound

This paper cites Autoad ii: The sequel-who, when, and what in movie audio description.

DistinctAD: Distinctive Audio Description Generation in Contexts Autoad ii: The sequel-who, when, and what in movie audio description

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.618232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.135896Z digest=sha256:e310a55dab5fad5f6bd7784814ae1fe871871fe79e1b428c39d5719be31aea05

Observation 53171eea-5a48-4554-aa1d-239d3fa6737a · outbound

This paper cites Autoad iii: The prequel-back to the pixels.

DistinctAD: Distinctive Audio Description Generation in Contexts Autoad iii: The prequel-back to the pixels

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.602198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.140758Z digest=sha256:79fe4a112a2619a1ad444d6d834c324fba8d6434e0d3ed2345de5641cfe24b4d

Observation 33a2fcf4-4b2a-4ea7-8218-a032e2be1d59 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

DistinctAD: Distinctive Audio Description Generation in Contexts LoRA: Low-Rank Adaptation of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.145887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.145887Z digest=sha256:44a3d0c4a86f62d7a3d8c3f31c465cdb9e31b653c1e6d7d32cbe2158468ac772

Observation c9c2f81e-e134-450c-8a47-d2076cd99df9 · outbound

This paper cites A better use of audio-visual cues: Dense video captioning with bi-modal transformer.

DistinctAD: Distinctive Audio Description Generation in Contexts A better use of audio-visual cues: Dense video captioning with bi-modal transformer

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.586799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.150980Z digest=sha256:42cdb8d8db52ca70b2ea8839a2dbf126fc44d975fd985d5e88498eda1fe2e790

Observation 97f9ed30-2ff2-4ae7-bf0c-6ddf308c2093 · outbound

This paper cites Multi-modal dense video captioning.

DistinctAD: Distinctive Audio Description Generation in Contexts Multi-modal dense video captioning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.570936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.155823Z digest=sha256:c9e811ee1c2878fc3bc29e20d7b472cd72a2aa6bf0538fb8dac3d2048072bae8

Observation 97b1dc59-f90d-4f59-84b6-71e80a5e9d0c · outbound

This paper cites Expectation- maximization contrastive learning for compact video-and- language representations.

DistinctAD: Distinctive Audio Description Generation in Contexts Expectation- maximization contrastive learning for compact video-and- language representations

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.555158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.161137Z digest=sha256:b7c0ad2c91fbf1bc7178d82d280fdb7f56b9cbb647212150bfb79a322623fdb4

Observation e8d4ac44-a001-468e-a96b-75e5294b8943 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

DistinctAD: Distinctive Audio Description Generation in Contexts Adam: A Method for Stochastic Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.167142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.167142Z digest=sha256:4ab53717ea3668d33fcde9c590411660ebb6374f1f9343bd24b2e32a83026376

Observation 6de8d026-4d86-4e88-81b4-03bda2214312 · outbound

This paper cites Dense-captioning events in videos.

DistinctAD: Distinctive Audio Description Generation in Contexts Dense-captioning events in videos

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.537697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.172088Z digest=sha256:af7314a1ac5b962bbb6c54cb30843db5713bcb66acb873e40fe605e70311a4e7

Observation 2a8f33be-f36f-4294-be23-3a669671dc14 · outbound

This paper cites Tvqa: Localized, compositional video question answering.

DistinctAD: Distinctive Audio Description Generation in Contexts Tvqa: Localized, compositional video question answering

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.519937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.177778Z digest=sha256:129e25aefe7a371c77d0f854cc3a2afe4965a22a94bd810cc7999eafd9f4566f

Observation 27a36bf9-0566-4db7-abb9-5eb73b72a5f8 · outbound

This paper cites Deep dive: How audio description benefits ev- eryone, 2021.

DistinctAD: Distinctive Audio Description Generation in Contexts Deep dive: How audio description benefits ev- eryone, 2021

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.503708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.182692Z digest=sha256:97776f8cc75543d2460704b512839d69c07ac794b3cb8941d7ea7d3c51778909

Observation 23c28d39-e743-4b41-85f4-e50ea3378951 · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

DistinctAD: Distinctive Audio Description Generation in Contexts MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.188334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.188334Z digest=sha256:583370cfb5de32739b04c2cb77025e72121d22f123bd6bc1f668200bf7e4ead9

Observation 8b2cd985-cdd6-45d3-91ab-68ea8e942daf · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

DistinctAD: Distinctive Audio Description Generation in Contexts Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.193344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.193344Z digest=sha256:e7c9e4271a8c6cac0d9285c043ff14db1a2392941e8ca36f2fb5bad85a1878b7

Observation ab03d617-461f-43e0-b03a-772e6a86c436 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

DistinctAD: Distinctive Audio Description Generation in Contexts BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.199014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.199014Z digest=sha256:47949500648b96b30d25f53fa315644193c3ca130efd07aed9d8cde93a9fb759

Observation 9b1eea9f-6c29-4b92-9086-f7b5d762f49c · outbound

This paper cites Expectation-maximization attention net- works for semantic segmentation.

DistinctAD: Distinctive Audio Description Generation in Contexts Expectation-maximization attention net- works for semantic segmentation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.487462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.204541Z digest=sha256:0e931503e845e31f5e506d76a3f843c44f3f36d1d3f54f0b65e58916db2c39d4

Observation 3d624013-759a-42c4-a3e0-b7ef25019bf5 · outbound

This paper cites Jointly localizing and describing events for dense video captioning.

DistinctAD: Distinctive Audio Description Generation in Contexts Jointly localizing and describing events for dense video captioning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.471370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.209995Z digest=sha256:d41858f01d6468dea62a014a4f0d4e8d26b910f6cd83395ebaf2909c959b1ce4

Observation c74c845f-9177-4355-ac63-301279123e91 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

DistinctAD: Distinctive Audio Description Generation in Contexts Rouge: A package for automatic evaluation of summaries

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.214883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.214883Z digest=sha256:fdb12e10e882bf99fccbff0138b9b3d78ca28edfd5fc1093f332bbc3f7caa1d0

Observation 3b1b0724-63d7-4730-be95-1ac1a0ced0bd · outbound

This paper cites Swinbert: End-to-end transformers with sparse attention for video cap- tioning.

DistinctAD: Distinctive Audio Description Generation in Contexts Swinbert: End-to-end transformers with sparse attention for video cap- tioning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.444961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.221004Z digest=sha256:fb08a378bf19f43b30d126c55dc3e54dfd33fabed17e1bfb393d911495ba885f

Observation a3e0faf7-77be-4fb1-bb49-14965fc28da5 · outbound

This paper cites MM-VID: Advancing Video Understanding with GPT-4V(ision).

DistinctAD: Distinctive Audio Description Generation in Contexts MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.227041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.227041Z digest=sha256:d3ca984485d248392b2b2f226f9bd061e1e040a34b7374d0a6815a2b0a362045

Observation 22704919-9201-4f2c-85f9-42851797e616 · outbound

This paper cites Learning Video Context as Interleaved Multimodal Sequences.

DistinctAD: Distinctive Audio Description Generation in Contexts Learning Video Context as Interleaved Multimodal Sequences

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:30:20.692183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.233251Z digest=sha256:8c5fa1b29a15f29cebe8453e99dc0061f8ac55ea37f013e79e1d9b090c471101

Observation 435a6317-223f-4459-a81e-719b83a7ced4 · outbound

This paper cites Swem: Towards real- time video object segmentation with sequential weighted expectation-maximization.

DistinctAD: Distinctive Audio Description Generation in Contexts Swem: Towards real- time video object segmentation with sequential weighted expectation-maximization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.428781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.240064Z digest=sha256:fd180bd7f95e8c028537e68adc8a72710bf546043cbc835ad2e10d3c774c2758

Observation 7651313a-096d-4a23-beef-af6b83729484 · outbound

This paper cites Visual instruction tuning.

DistinctAD: Distinctive Audio Description Generation in Contexts Visual instruction tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.245406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.245406Z digest=sha256:544401869fba50fae03bba2c50a990e57aa79b746f282890ee0cbf8bf697bcd6

Observation b28c05f9-66b4-4dc8-85f7-c5c1c551cb50 · outbound

This paper cites Show, tell and discriminate: Image captioning by self-retrieval with partially labeled data.

DistinctAD: Distinctive Audio Description Generation in Contexts Show, tell and discriminate: Image captioning by self-retrieval with partially labeled data

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.403443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.250581Z digest=sha256:8e28c1549116bdd21bee3e1d623aadbc7fd22642ed9f0f30e434f952d8c91153

Observation 0aca757f-273d-446c-88ab-8bd49d920e18 · outbound

This paper cites Decoupled Weight Decay Regularization.

DistinctAD: Distinctive Audio Description Generation in Contexts Decoupled Weight Decay Regularization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.255802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.255802Z digest=sha256:75937a9792b79d3fa2293a9401a87d0dcd1100028ef2e84ceadef4f7baa986de

Observation a5e8bd5c-3fb9-47de-be4a-89969db9545b · outbound

This paper cites UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation.

DistinctAD: Distinctive Audio Description Generation in Contexts UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.261336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.261336Z digest=sha256:86e2d1ed2a83d671221eeda2ecb1c451ce6a7b2bad4136b5b08276474a0850eb

Observation 3189e427-a8c2-4978-b120-15cb89919346 · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.

DistinctAD: Distinctive Audio Description Generation in Contexts Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.388583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.266455Z digest=sha256:799ff421ba860cd363dbc2ad4d6384987ed615df26df3ab185848cb27cdf3065

Observation ea0e8003-9d20-4679-beb0-cca94e0b46df · outbound

This paper cites Discriminability objective for training de- scriptive captions.

DistinctAD: Distinctive Audio Description Generation in Contexts Discriminability objective for training de- scriptive captions

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.374327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.271181Z digest=sha256:e10d2c971af0b193cfbb3a47fd2dda71d228c3f9d197436e2b10607bc1c01aa5

Observation cb499339-aa4e-41f3-a53b-8d8e206e7376 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

DistinctAD: Distinctive Audio Description Generation in Contexts Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.358684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.276717Z digest=sha256:f3cbbd5fd3887ace59068aec730b81a605f074cf8523a1c66960b52a1982b48d

Observation 01e3ac2e-6a00-49b3-854c-e17bb462a0f1 · outbound

This paper cites End-to-end learning of visual representations from uncurated instruc- tional videos.

DistinctAD: Distinctive Audio Description Generation in Contexts End-to-end learning of visual representations from uncurated instruc- tional videos

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.342856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.281897Z digest=sha256:ec0ffa30656048f615b277fd0bb9bb2646717b1eb7ddf401f6656935da12900d

Observation 190f2db0-2770-4c52-99eb-25cd897095f7 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

DistinctAD: Distinctive Audio Description Generation in Contexts ClipCap: CLIP Prefix for Image Captioning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.286884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.286884Z digest=sha256:403d155504fc842d93ba19508d6bab396b264d397014c85c3ed8f0baac1ad724

Observation 86ba81e9-2cf4-4c67-be3a-5de3a0cfffdf · outbound

This paper cites Streamlined dense video captioning.

DistinctAD: Distinctive Audio Description Generation in Contexts Streamlined dense video captioning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.327232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.292010Z digest=sha256:32c2dbb3b257de629f32e411434c89740b5980e9d873417e03288fb274cf44d4

Observation 425d1255-1895-4219-bb4c-41052895f79c · outbound

This paper cites Text-Only Training for Image Captioning using Noise-Injected CLIP.

DistinctAD: Distinctive Audio Description Generation in Contexts Text-Only Training for Image Captioning using Noise-Injected CLIP

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.298071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.298071Z digest=sha256:a4ca501921090ea8b69014d5571cc2e7c474077daea2cf936596c9b493ba37df

Observation 3f725d40-a227-43e7-8e0d-210d9c0fb16c · outbound

This paper cites Gpt-4v(ision) system card.

DistinctAD: Distinctive Audio Description Generation in Contexts Gpt-4v(ision) system card

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.312297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.303224Z digest=sha256:792e67b5443b12a914cf9449d6cbd623d5362bb017ac57f08e2495ff4df7d880

Observation 73155be7-3987-4e50-ae7c-0f046f603097 · outbound

This paper cites Rescribe: Authoring and automatically editing audio descriptions.

DistinctAD: Distinctive Audio Description Generation in Contexts Rescribe: Authoring and automatically editing audio descriptions

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.297653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.308564Z digest=sha256:28c3c0e061ebce16863bc576d5eec605b4aff2420f9e9e8c7b47c6253c242f31

Observation 3d5f1b69-8075-4b51-9a4a-cc5f5cdb035e · outbound

This paper cites Gains and losses of watching audio described films for sighted viewers.

DistinctAD: Distinctive Audio Description Generation in Contexts Gains and losses of watching audio described films for sighted viewers

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.282629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.314508Z digest=sha256:24aed0a8d5813fb60055fd3493d946d3b63bfe6d81bce3b49d3e9061fec6c3c0

Observation d906e9e9-9010-472d-bbd9-eaedd8380162 · outbound

This paper cites Micap: A unified model for identity- aware movie descriptions.

DistinctAD: Distinctive Audio Description Generation in Contexts Micap: A unified model for identity- aware movie descriptions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.267702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.319563Z digest=sha256:d9151ee61218cb0496ed1a3b0166351a3c21fa4428287e9788552af01f878365

Observation 32fceb71-f90f-4330-bdf7-51aaf1b5c16a · outbound

This paper cites Language models are unsu- pervised multitask learners.

DistinctAD: Distinctive Audio Description Generation in Contexts Language models are unsu- pervised multitask learners

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.252571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.324457Z digest=sha256:a9d19ff9238c2ea74f12533a7b5a793d7a53e258b56ef7315b23bfbf1897111a

Observation e7fc0d86-6aa8-4923-b001-1811bc3f9e2c · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

DistinctAD: Distinctive Audio Description Generation in Contexts Learn- ing transferable visual models from natural language super- vision

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.235740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.329717Z digest=sha256:d8412057bbba3b62b5735efc26db740614afc4ac840549347d913e5def5ce842

Observation a8379797-23ed-4c49-ad6f-5e07ed01258c · outbound

This paper cites Watch, listen and tell: Multi-modal weakly supervised dense event captioning.

DistinctAD: Distinctive Audio Description Generation in Contexts Watch, listen and tell: Multi-modal weakly supervised dense event captioning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.219693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.335191Z digest=sha256:cd94915d4fc6cf3efe33aa6199a9568a927d71e95a30da33e07283446c14047f

Observation 41155eae-52c5-4b66-a002-c42f19129dd4 · outbound

This paper cites A dataset for movie description.

DistinctAD: Distinctive Audio Description Generation in Contexts A dataset for movie description

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.203727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.340416Z digest=sha256:97966f63a19aa1c0a7eb6a075fb2038cb3f771e5a0a3cff7bc91dddd7eb4322b

Observation 38e4d5a2-dcd2-47ee-b8a7-baaecf622c46 · outbound

This paper cites Movie description.

DistinctAD: Distinctive Audio Description Generation in Contexts Movie description

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.188477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.346187Z digest=sha256:6456b91c58224d46d87ce9441662a5730ca38c19aec3da442b628f8746489a75

Observation 76c7bb5c-52b7-45d5-855c-a215bda12685 · outbound

This paper cites End-to-end generative pretraining for mul- timodal video captioning.

DistinctAD: Distinctive Audio Description Generation in Contexts End-to-end generative pretraining for mul- timodal video captioning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.172595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.351379Z digest=sha256:7d4e11a39e31f3f02ae7ec540e1faba20aff3713b0d6333d309c9026314c8137

Observation a120a116-b88a-4257-8667-6ca8672067df · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

DistinctAD: Distinctive Audio Description Generation in Contexts Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.155988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.356250Z digest=sha256:9d6538f18c154f3aa0b332b91b6e6c463ac9c06ddfd1bea66cd340d349740eaa

Observation 08aeaa94-4788-42df-8cbb-9919bb0cb5a3 · outbound

This paper cites Weakly super- vised dense video captioning.

DistinctAD: Distinctive Audio Description Generation in Contexts Weakly super- vised dense video captioning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.139712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.361192Z digest=sha256:06f8cb92dda75f435dd23d2886df020cd3e24c021509ee50aae4ecd7df11baff

Observation 8f133d76-ee14-4dcf-b1ef-ae42d6434d7a · outbound

This paper cites Dense procedure captioning in narrated instructional videos.

DistinctAD: Distinctive Audio Description Generation in Contexts Dense procedure captioning in narrated instructional videos

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.124630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.365803Z digest=sha256:dba333a9b2b88f1563ec2d24efcded707b5f00b244ef20c319f934df1e170e6a

Observation 20c7e090-9998-4ffa-9abb-d6adec1a1b95 · outbound

This paper cites What does clip know about a red circle? visual prompt engineering for vlms.

DistinctAD: Distinctive Audio Description Generation in Contexts What does clip know about a red circle? visual prompt engineering for vlms

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.108216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.370645Z digest=sha256:371a1296190cbf15fd09574f0011d21b1f43053fed4b48f58a987115dcb600ad

Observation 77fa4c38-5a2e-4f7a-9df3-b99f0cb27b5e · outbound

This paper cites Audio description: The visual made verbal.

DistinctAD: Distinctive Audio Description Generation in Contexts Audio description: The visual made verbal

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.091809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.375474Z digest=sha256:51a107822440f82d331ddff60712cb329efc461717904a2dd8d99e9346f1ae1d

Observation 53f767b4-5613-419d-a45e-db061579556f · outbound

This paper cites Mad: A scalable dataset for language grounding in videos from movie audio descriptions.

DistinctAD: Distinctive Audio Description Generation in Contexts Mad: A scalable dataset for language grounding in videos from movie audio descriptions

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.074749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.381315Z digest=sha256:273a8fc0c4c4ae992d3c5f96693c20ae8d4a2339a0d40ea2a6be06d8a9ebf6c8

Observation 3a248ff5-3b61-4870-a9f7-97a836214593 · outbound

This paper cites Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research.

DistinctAD: Distinctive Audio Description Generation in Contexts Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.386515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.386515Z digest=sha256:d003681227a11b8ea66c751b6c2affe5526172d2218ec762a2249471a183f73b

Observation ae4f3e86-be4d-4912-ac37-7b24bae5ef75 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DistinctAD: Distinctive Audio Description Generation in Contexts Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.392394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.392394Z digest=sha256:78a14e2fcd9e50da4099654c8bd6f5e59a50c7b097668b9db1f883c2d8f1870e

Observation bca22f9a-8009-45ab-ac50-c0965efe2d57 · outbound

This paper cites Visualizing data using t-sne.

DistinctAD: Distinctive Audio Description Generation in Contexts Visualizing data using t-sne

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.397108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.397108Z digest=sha256:b53b12508766ff9f0bdd6ec8a3a9ff379abee83624770fc5cf93eb7add4a45df

Observation 169aabef-6ee4-4c1a-93a9-d4ba8292e72b · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

DistinctAD: Distinctive Audio Description Generation in Contexts Cider: Consensus-based image description evalua- tion

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.402833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.402833Z digest=sha256:db1a917613774d6b3de3af4bafa6c32a2126420eddae54093df0ee466d50c363

Observation ad533852-845d-4a22-b577-df1b28341f93 · outbound

This paper cites Joint optimization for cooperative image captioning.

DistinctAD: Distinctive Audio Description Generation in Contexts Joint optimization for cooperative image captioning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.036154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.407650Z digest=sha256:faff4c50fd3c8f3e07ef7695ca5a809dd1698872f164474a7754d1cc850b1556

Observation 75122a52-1e50-4ba9-a7b9-883bc6bcc160 · outbound

This paper cites Contextual AD Narration with Interleaved Multimodal Sequence.

DistinctAD: Distinctive Audio Description Generation in Contexts Contextual AD Narration with Interleaved Multimodal Sequence

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:30:20.579596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.412490Z digest=sha256:159bab14e16ff36bb3e6f7c412da1448d73a70da68019372062f9aa853e015c9

Observation c5b510ca-7152-4e1d-9b6a-2f092ffd6b15 · outbound

This paper cites Bidirectional attentive fusion with context gating for dense video captioning.

DistinctAD: Distinctive Audio Description Generation in Contexts Bidirectional attentive fusion with context gating for dense video captioning

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.018153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.418074Z digest=sha256:633b5c72199aaf29d43abde0b2de65a870ccff8cc82e809f01a2995a94b772c6

Observation 1dfdf37d-0993-4582-aaaf-97c44303a3b8 · outbound

This paper cites Compare and reweight: Distinctive image caption- ing using similar images sets.

DistinctAD: Distinctive Audio Description Generation in Contexts Compare and reweight: Distinctive image caption- ing using similar images sets

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:21.001138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.422918Z digest=sha256:b06adc8fc7f80db7383b44c2c2ec45985dce52c547378e1e73dcfcc023e69a43

Observation e6e2cb71-af4a-4e98-b6d2-e9d03477f90c · outbound

This paper cites Group-based distinctive image captioning with mem- ory attention.

DistinctAD: Distinctive Audio Description Generation in Contexts Group-based distinctive image captioning with mem- ory attention

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:20.985825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.427477Z digest=sha256:4f5375218549e4722bf79e0215749c83d1ce1374ed11a659aede32fef7af1975

Observation d05c025b-76ef-482a-a9e0-e7e8875308b5 · outbound

This paper cites On distinctive image captioning via comparing and reweighting.

DistinctAD: Distinctive Audio Description Generation in Contexts On distinctive image captioning via comparing and reweighting

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:20.970502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.431963Z digest=sha256:cecfa56c0210ef94458b77d47ea625aa81a1d9b355cb45c6ebe7c707b79a0967

Observation 52319beb-22ff-48b0-89a0-36799877cec7 · outbound

This paper cites Event-centric hierarchical representation for dense video captioning.

DistinctAD: Distinctive Audio Description Generation in Contexts Event-centric hierarchical representation for dense video captioning

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:20.955309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.436224Z digest=sha256:45742467b4e4cc6b9cf2c9d0dc4cc28d1d2b5dfc053071f1695f17285d83a0a2

Observation daf0ef99-e1bc-444a-86c1-acb21e8e4bb2 · outbound

This paper cites End-to-end dense video captioning with parallel decoding.

DistinctAD: Distinctive Audio Description Generation in Contexts End-to-end dense video captioning with parallel decoding

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:20.940335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.440887Z digest=sha256:da3c7569ed3bccfa4f121e3fa95f583bd7c6448c764bc42681c9fcbe1c28d1cb

Observation 970f8517-cb71-4a38-b45b-3fd2b3771cf3 · outbound

This paper cites Non-local neural networks.

DistinctAD: Distinctive Audio Description Generation in Contexts Non-local neural networks

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:20.924408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.446232Z digest=sha256:0e81edb24c112a66be86a72c9621516ba21ea1af5e22959d4600b4b21c3e904b

Observation 9a51e9ef-70a4-4971-b81c-0b827c920e80 · outbound

This paper cites AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description.

DistinctAD: Distinctive Audio Description Generation in Contexts AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.450704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.450704Z digest=sha256:0cbd3f419841243ccb459c64a759951d7faaa1d8c37cbb9a29cb70f8f8b77e10

Observation 8933c6c2-56e0-40d4-8412-55e8f33a91f1 · outbound

This paper cites Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning.

DistinctAD: Distinctive Audio Description Generation in Contexts Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:20.906219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.455460Z digest=sha256:bfff00154124c46d1e093775e803905058f49c919a273553040a63d75fb2fcaa

Observation 51d7a8be-751b-4b9b-affd-f1ca63fdc1f4 · outbound

This paper cites Image difference cap- tioning with pre-training and contrastive learning.

DistinctAD: Distinctive Audio Description Generation in Contexts Image difference cap- tioning with pre-training and contrastive learning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:20.891378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.460341Z digest=sha256:e51d0153f3dcedceff1eebeb2fb6239eb536860d86a86c249a386479dbef3723

Observation 2e8cb047-05c8-4101-b77a-fb540698f292 · outbound

This paper cites Videoblip, 2023.

DistinctAD: Distinctive Audio Description Generation in Contexts Videoblip, 2023

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:20.876416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.465130Z digest=sha256:e1193183ba7c6031cd7a5703a6c59ff2cb23cf27b52c430246625d9e17f1e703

Observation 83fa0344-cd93-42ba-840e-628f90421e7d · outbound

This paper cites Mm-narrator: Narrating long-form videos with multimodal in-context learning.

DistinctAD: Distinctive Audio Description Generation in Contexts Mm-narrator: Narrating long-form videos with multimodal in-context learning

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:20.860956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.469477Z digest=sha256:7d3a03c74cdf0b2f3710ed8ef0cfc80f28d3a7f8ecac91aa13db32f11d876794

Observation d9e78c4b-1d56-43b2-b1e1-f52f0a56bce6 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

DistinctAD: Distinctive Audio Description Generation in Contexts Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.473950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.473950Z digest=sha256:ff6fae7bd5f724bc082b4434faf942a9a08cba05cfbcb08ea09505ab5c2a2fe6

Observation b66d9701-a6e4-4fea-9071-e244ac9cb545 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

DistinctAD: Distinctive Audio Description Generation in Contexts BERTScore: Evaluating Text Generation with BERT

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:20.479562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:20.479562Z digest=sha256:34d465132c4a7116a334d9c2f6cbd2e82815a7a6151129ccdebe6745b26c8a94

Observation d7d91cf0-d474-4d6a-8f0b-d2fda08ed880 · outbound

This paper cites nuns” mistakenly appears in (d). AutoAD-II tends to gen- erates similar AD words, e.g. “furrowed brow.

DistinctAD: Distinctive Audio Description Generation in Contexts nuns” mistakenly appears in (d). AutoAD-II tends to gen- erates similar AD words, e.g. “furrowed brow

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:30:20.844965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:30:20.484216Z digest=sha256:921a3df969ef3d595d2f5a77aea24e55fff89830b3628869e3c581ac59e4801f

Pith citing papers

No inbound Pith citation observations are available.