Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T11:21:49.358294Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:1908.09317.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T11:21:49.358294Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7d68a81e-813b-4fc0-8937-6e11956cb70d · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Spice: Semantic propositional image cap- tion evaluation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5af452b7-3d16-462c-879f-85ce8327d493 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Guided open vocabulary image captioning with constrained beam search
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 97bd61fd-f0ce-4a27-be01-0d1a96643300 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Partially-supervised image captioning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cd622e6e-eb10-4e62-9da2-049a44506978 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Bottom-up and top-down attention for image captioning and visual question answering
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 465e539e-eaee-4d12-b385-cb964f062c7c · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Women also snowboard: Over- coming bias in captioning models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cd20f5a7-7f92-487b-9e7c-18306e0943e9 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Deep compositional captioning: Describing novel ob- ject categories without paired training data
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 37608bf8-7d38-4074-bf61-dad8e8ef08f0 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Lawrence Zitnick, and Devi Parikh
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f0696281-d196-425f-ae58-76177ceb549c · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Adversarial text generation via feature-mover’s distance
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a6a209ca-8435-4475-8d61-c1b6eb3e5543 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Show, adapt and tell: Adversarial training of cross-domain image cap- tioner
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 37b91ab8-d4c5-4820-a2d9-9baf2e576db4 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14fb92d0-7223-468e-b608-9afa051ee48e · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings To- wards diverse and natural image descriptions via a condi- tional GAN
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a296666d-260e-4bad-a15c-60fbf54f8b2f · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Visual dialog
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9b6c5ffe-db60-4622-ad3a-ac17afd8e0ba · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Meteor universal: Lan- guage specific translation evaluation for any target language
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 96ddcc7f-eeb6-4429-9539-e8ece88393e6 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Long-term recurrent convolutional net- works for visual recognition and description
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 174b51ff-a2df-4233-a99e-92a9490c6535 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Adversarial Feature Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d557ae01-89a5-4f1c-8b2f-3ea5e75fa703 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings VSE++: Improving Visual-Semantic Embeddings with Hard Negatives
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aea4e3e0-1b75-4341-b3a1-40ca4affc4dc · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings From captions to vi- sual concepts and back
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4866dddf-2f31-4059-b0bd-2177e7d0a151 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Unsupervised Image Captioning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 79f3acfb-a3a4-4087-bf21-c112bc25ec89 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings StyleNet: Generating attractive visual captions with styles
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ce139965-5b8c-46f6-ad4e-724719239783 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Un- paired image captioning by language pivoting
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5487240a-7548-4ed3-9600-ac34efdb04ac · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Improved training of wasserstein gans
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 711de645-aeae-4efe-b938-683a37a07f0c · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings MSCap: Multi-style image captioning with un- paired stylized text
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7d9848f4-4871-4f17-aedb-d955dd1e28a6 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings VizWiz grand challenge: Answering visual questions from blind people
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f1255e8a-a030-44fd-bf18-fc2e7ac68918 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Deep residual learning for image recognition
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7e9902d-867a-47e4-acbc-3be5f12ab9bd · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Speed/accuracy trade-offs for modern convolutional object detectors
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7412c2ab-6660-4bc3-b97e-1cbfe57af8b9 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings DenseCap: Fully convolutional localization networks for dense caption- ing
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 13c9505c-4874-4445-b13b-17d8a512eed3 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Deep visual-semantic align- ments for generating image descriptions
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e921cc49-5e6e-4c28-bfa2-e04e0d2a58bd · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Adam: A Method for Stochastic Optimization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7256ea6d-0897-4078-a360-1a693a8a1913 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edbb9db4-cad3-4413-ad80-ab692e553d16 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings OpenImages: A public dataset for large-scale multi-label and multi-class image classification
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 140f1fa0-6291-4298-aac8-482c66e61aae · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 71e19b4c-db59-4d3a-8318-29c1d03ef217 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings From word embeddings to document distances
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e639d1fd-85ff-4d48-91aa-73dce18fb0a8 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3547e4ff-82bd-4d82-913a-d8777c58e2ce · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Unsupervised machine translation using monolingual corpora only
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 87046abb-bc66-47a2-ad59-99ee07ee2c10 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Phrase-based & neural unsupervised machine translation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 75b16d00-50a4-4480-94a6-2fd4d930ec3a · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Generating Diverse and Accurate Visual Captions by Comparative Adversarial Learning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1966479c-1b24-4bb0-9faf-d738c6c39e50 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Rouge: A package for automatic evaluation of summaries
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 11b42c84-79c1-4e7b-ac00-ae6e040add50 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Microsoft COCO: Common objects in context
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d0480356-7346-47b4-9fcf-19c9de0e11dc · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Teaching machines to describe images via natural language feedback
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5b174ceb-e1e2-44d7-bef1-d72524f35b61 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Improved image captioning via policy gra- dient optimization of spider
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4a589ba-ea53-45d1-b4cd-00b88a060639 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Knowing when to look: Adaptive attention via a visual sen- tinel for image captioning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5297d2d2-0ce3-4cd1-a37d-29822afa472b · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Neural baby talk
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4646014e-da54-4b69-b8b8-d63ddf7abc43 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Visualizing data using t-SNE
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 61847bfc-ef07-40d7-99c1-324254d67622 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings The stan- ford CoreNLP natural language processing toolkit
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c8eec48e-6ba1-4a5d-ad17-e8f0b04ea120 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Learning like a child: Fast novel visual concept learning from sentence descriptions of images
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation caaacc1c-8c40-455b-8dde-0f7985c42798 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Sem- Style: Learning to generate stylised image captions using unaligned text
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 081451b6-48d3-42cd-8ff9-c1a9c6202780 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Jointly modeling embedding and translation to bridge video and language
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 210a1364-a49a-4038-a1dd-46f43fbb0cb8 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings GloVe: Global vectors for word representation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9d83399a-4384-45d3-8d09-8f5d91b25f1f · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b2437259-aa6f-40e0-a48e-16a13e1e9d77 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Self-critical sequence training for image captioning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 08d40f3f-c859-47e4-af41-b5c0fd0e97c8 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings ImageNet large scale visual recognition challenge
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0fd873de-d811-4278-bf7c-5eb2202282fb · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8d405e70-3988-446d-8598-da7516652da1 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Speaking the same language: Matching machine to human captions by adversarial train- ing
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 21de770f-2a64-44db-be90-6b321b68c7ea · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Deforming autoencoders: Unsupervised disentangling of shape and ap- pearance
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 24d7acb7-651e-404a-b054-6fb0e294b2af · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Engaging image captioning via per- sonality
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eb488959-56e3-4534-83f7-3e5b4a2a5dcc · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Towards text generation with adversarially learned neural outlines
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3b62f948-9c4f-425c-87a7-e44dc90b407d · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Sequence to sequence learning with neural networks
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2ca248f7-0463-4f3b-be81-6aca77e5b0f4 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Cider: Consensus-based image description evalua- tion
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e69ab5a-de3b-4ad6-abcc-de758b06f2a1 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Captioning images with diverse objects
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation add7915b-047b-4bb5-9fb3-b86bc1d53de2 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Show and tell: A neural image caption gen- erator
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation aba98efa-5dc5-4e67-ac51-3ba198649935 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Diverse and accurate image description using a variational auto-encoder with an additive gaussian encoding space
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f563f42d-6c22-4729-a433-657eb34d8d69 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Automatic alt-text: Computer-generated image de- scriptions for blind users on a social network service
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ae702d51-7242-4aeb-a8c3-04d1e674e811 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Show, attend and tell: Neural image caption gen- eration with visual attention
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3672b3b3-97c3-459d-8afe-c45b03a8cb18 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Review networks for caption gen- eration
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0f0f9291-e594-40ce-a1ad-0e53b8936924 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Incorpo- rating copying mechanism in image captioning for learn- ing novel objects
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d2256260-1f4b-4f0b-ac04-a156ac8daaca · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Explor- ing visual relationship for image captioning
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fcb4a1cd-f968-4957-9b8c-652f58199050 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Boosting image captioning with attributes
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0aee144c-dc01-4e1f-a6ad-c9d1b1f0af0d · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Image captioning with semantic attention
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 168df444-f91c-4e67-b062-2d06e4af929c · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Dual learning for cross-domain image captioning
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 51603e1c-a047-49ea-af49-27d042c7c6c5 · outbound
Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Unpaired image-to-image translation using cycle- consistent adversarial networks
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
No inbound Pith citation observations are available.