Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:30:20.484216Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 0 inbound Pith citation observations for arXiv:2411.18180.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:30:20.484216Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
88 of 88 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 02369d82-26c8-4448-8328-7213f3562022 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts https://github
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb38ada4-d5fa-40e8-a7e9-e286e73ee22b · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fb03a4b-c28f-42c8-a36f-240869edfd83 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Llama 3 model card
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7808cea-a78d-4ccf-9e44-d48708a49cd3 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Flamingo: a visual language model for few-shot learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f17910fa-f342-4a37-a905-be84e39eea45 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Spice: Semantic propositional image cap- tion evaluation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fb462c2-41ef-4a96-a49e-6d361bfa05e6 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Condensed movies: Story based retrieval with con- textual embeddings
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e3b12e7-a13d-421b-9677-d07840be8bf3 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts WhisperX: Time-Accurate Speech Transcription of Long-Form Audio
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 076c5170-ea3d-43ec-a676-3442d7555ced · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Livedescribe: can amateur describers create high-quality audio description? Journal of Visual Impairment & Blindness, 106(3):154–165,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7af109e-deea-41ea-9e70-80cbe62e5a6e · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts End-to-end speaker seg- mentation for overlap-aware resegmentation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f495fa8-dcb0-40bb-95ca-2b6b659f319d · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Pyannote
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 343971bb-4cce-468e-9f84-5981aadac8a6 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Groupcap: Group-based image captioning with structured relevance and diversity constraints
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c4d87e84-4245-486e-b5e9-a29d8bc02f6b · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Towards bridging event captioner and sentence localizer for weakly supervised dense event captioning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation febee165-4f08-4e3a-a0a8-c454cf5c36c7 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts LLM-AD: Large Language Model based Audio Description System
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a61deb-f6e9-4959-b44c-01c58659e772 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts The ides of march
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7d9787e5-9fbd-43d6-bab4-ecd5d9b3e214 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Contrastive learning for image cap- tioning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 43ab6fd3-6bb2-4d83-a571-6a99b8739eb2 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Maximum likelihood from incomplete data via the em al- gorithm
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 62983dfc-f032-44d7-bc9d-696460ca8dd3 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Sketch, ground, and refine: Top-down dense video caption- ing
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6b948ab6-4590-4931-9d80-9582089132b1 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Describing differences in image sets with natural language
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f8138907-0ea8-4c08-a8a8-2d55280fd1c7 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts An introduction to audio description: A prac- tical guide, 2016
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2200baab-6ac4-4870-9f70-d4dffe0e3e65 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Autoad: Movie description in context
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2321d48f-f5e8-4f29-a8d1-8b8d0623b92a · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Autoad ii: The sequel-who, when, and what in movie audio description
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 53171eea-5a48-4554-aa1d-239d3fa6737a · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Autoad iii: The prequel-back to the pixels
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 33a2fcf4-4b2a-4ea7-8218-a032e2be1d59 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts LoRA: Low-Rank Adaptation of Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9c2f81e-e134-450c-8a47-d2076cd99df9 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts A better use of audio-visual cues: Dense video captioning with bi-modal transformer
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 97f9ed30-2ff2-4ae7-bf0c-6ddf308c2093 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Multi-modal dense video captioning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 97b1dc59-f90d-4f59-84b6-71e80a5e9d0c · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Expectation- maximization contrastive learning for compact video-and- language representations
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e8d4ac44-a001-468e-a96b-75e5294b8943 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Adam: A Method for Stochastic Optimization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6de8d026-4d86-4e88-81b4-03bda2214312 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Dense-captioning events in videos
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2a8f33be-f36f-4294-be23-3a669671dc14 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Tvqa: Localized, compositional video question answering
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 27a36bf9-0566-4db7-abb9-5eb73b72a5f8 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Deep dive: How audio description benefits ev- eryone, 2021
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 23c28d39-e743-4b41-85f4-e50ea3378951 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b2cd985-cdd6-45d3-91ab-68ea8e942daf · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Otter: A Multi-Modal Model with In-Context Instruction Tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab03d617-461f-43e0-b03a-772e6a86c436 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b1eea9f-6c29-4b92-9086-f7b5d762f49c · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Expectation-maximization attention net- works for semantic segmentation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3d624013-759a-42c4-a3e0-b7ef25019bf5 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Jointly localizing and describing events for dense video captioning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c74c845f-9177-4355-ac63-301279123e91 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Rouge: A package for automatic evaluation of summaries
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b1b0724-63d7-4730-be95-1ac1a0ced0bd · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Swinbert: End-to-end transformers with sparse attention for video cap- tioning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a3e0faf7-77be-4fb1-bb49-14965fc28da5 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts MM-VID: Advancing Video Understanding with GPT-4V(ision)
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22704919-9201-4f2c-85f9-42851797e616 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Learning Video Context as Interleaved Multimodal Sequences
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 435a6317-223f-4459-a81e-719b83a7ced4 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Swem: Towards real- time video object segmentation with sequential weighted expectation-maximization
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7651313a-096d-4a23-beef-af6b83729484 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Visual instruction tuning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b28c05f9-66b4-4dc8-85f7-c5c1c551cb50 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Show, tell and discriminate: Image captioning by self-retrieval with partially labeled data
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0aca757f-273d-446c-88ab-8bd49d920e18 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Decoupled Weight Decay Regularization
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5e8bd5c-3fb9-47de-be4a-89969db9545b · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3189e427-a8c2-4978-b120-15cb89919346 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ea0e8003-9d20-4679-beb0-cca94e0b46df · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Discriminability objective for training de- scriptive captions
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cb499339-aa4e-41f3-a53b-8d8e206e7376 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 01e3ac2e-6a00-49b3-854c-e17bb462a0f1 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts End-to-end learning of visual representations from uncurated instruc- tional videos
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 190f2db0-2770-4c52-99eb-25cd897095f7 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts ClipCap: CLIP Prefix for Image Captioning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86ba81e9-2cf4-4c67-be3a-5de3a0cfffdf · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Streamlined dense video captioning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 425d1255-1895-4219-bb4c-41052895f79c · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Text-Only Training for Image Captioning using Noise-Injected CLIP
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f725d40-a227-43e7-8e0d-210d9c0fb16c · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Gpt-4v(ision) system card
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 73155be7-3987-4e50-ae7c-0f046f603097 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Rescribe: Authoring and automatically editing audio descriptions
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3d5f1b69-8075-4b51-9a4a-cc5f5cdb035e · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Gains and losses of watching audio described films for sighted viewers
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d906e9e9-9010-472d-bbd9-eaedd8380162 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Micap: A unified model for identity- aware movie descriptions
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 32fceb71-f90f-4330-bdf7-51aaf1b5c16a · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Language models are unsu- pervised multitask learners
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e7fc0d86-6aa8-4923-b001-1811bc3f9e2c · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Learn- ing transferable visual models from natural language super- vision
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a8379797-23ed-4c49-ad6f-5e07ed01258c · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Watch, listen and tell: Multi-modal weakly supervised dense event captioning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 41155eae-52c5-4b66-a002-c42f19129dd4 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts A dataset for movie description
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 38e4d5a2-dcd2-47ee-b8a7-baaecf622c46 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Movie description
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 76c7bb5c-52b7-45d5-855c-a215bda12685 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts End-to-end generative pretraining for mul- timodal video captioning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a120a116-b88a-4257-8667-6ca8672067df · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 08aeaa94-4788-42df-8cbb-9919bb0cb5a3 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Weakly super- vised dense video captioning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8f133d76-ee14-4dcf-b1ef-ae42d6434d7a · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Dense procedure captioning in narrated instructional videos
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 20c7e090-9998-4ffa-9abb-d6adec1a1b95 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts What does clip know about a red circle? visual prompt engineering for vlms
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 77fa4c38-5a2e-4f7a-9df3-b99f0cb27b5e · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Audio description: The visual made verbal
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 53f767b4-5613-419d-a45e-db061579556f · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Mad: A scalable dataset for language grounding in videos from movie audio descriptions
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3a248ff5-3b61-4870-a9f7-97a836214593 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae4f3e86-be4d-4912-ac37-7b24bae5ef75 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bca22f9a-8009-45ab-ac50-c0965efe2d57 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Visualizing data using t-sne
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 169aabef-6ee4-4c1a-93a9-d4ba8292e72b · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Cider: Consensus-based image description evalua- tion
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad533852-845d-4a22-b577-df1b28341f93 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Joint optimization for cooperative image captioning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 75122a52-1e50-4ba9-a7b9-883bc6bcc160 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Contextual AD Narration with Interleaved Multimodal Sequence
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c5b510ca-7152-4e1d-9b6a-2f092ffd6b15 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Bidirectional attentive fusion with context gating for dense video captioning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1dfdf37d-0993-4582-aaaf-97c44303a3b8 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Compare and reweight: Distinctive image caption- ing using similar images sets
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e6e2cb71-af4a-4e98-b6d2-e9d03477f90c · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Group-based distinctive image captioning with mem- ory attention
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d05c025b-76ef-482a-a9e0-e7e8875308b5 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts On distinctive image captioning via comparing and reweighting
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 52319beb-22ff-48b0-89a0-36799877cec7 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Event-centric hierarchical representation for dense video captioning
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation daf0ef99-e1bc-444a-86c1-acb21e8e4bb2 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts End-to-end dense video captioning with parallel decoding
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 970f8517-cb71-4a38-b45b-3fd2b3771cf3 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Non-local neural networks
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9a51e9ef-70a4-4971-b81c-0b827c920e80 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8933c6c2-56e0-40d4-8412-55e8f33a91f1 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 51d7a8be-751b-4b9b-affd-f1ca63fdc1f4 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Image difference cap- tioning with pre-training and contrastive learning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2e8cb047-05c8-4101-b77a-fb540698f292 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Videoblip, 2023
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 83fa0344-cd93-42ba-840e-628f90421e7d · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Mm-narrator: Narrating long-form videos with multimodal in-context learning
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d9e78c4b-1d56-43b2-b1e1-f52f0a56bce6 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b66d9701-a6e4-4fea-9071-e244ac9cb545 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts BERTScore: Evaluating Text Generation with BERT
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7d91cf0-d474-4d6a-8f0b-d2fda08ed880 · outbound
DistinctAD: Distinctive Audio Description Generation in Contexts nuns” mistakenly appears in (d). AutoAD-II tends to gen- erates similar AD words, e.g. “furrowed brow
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
No inbound Pith citation observations are available.