Pith. sign in

Paper Citation Record · LEDGER

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings

As of 18 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:1908.09317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.09317 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T11:21:49.358294Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact2
  • verified fuzzy58
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7d68a81e-813b-4fc0-8937-6e11956cb70d · outbound

This paper cites Spice: Semantic propositional image cap- tion evaluation.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Spice: Semantic propositional image cap- tion evaluation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.395385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.050800Z digest=sha256:a4a9ab0a7963d16053d5152b9e460f4c9edfe8ce0c8d59e8f3179c4089f3b399

Observation 5af452b7-3d16-462c-879f-85ce8327d493 · outbound

This paper cites Guided open vocabulary image captioning with constrained beam search.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Guided open vocabulary image captioning with constrained beam search

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.381155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.055957Z digest=sha256:9e4cfe110d402c3ea26b23009eae5b77acf9aa4d12a643696e16cfbc38e69785

Observation 97bd61fd-f0ce-4a27-be01-0d1a96643300 · outbound

This paper cites Partially-supervised image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Partially-supervised image captioning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.366747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.060608Z digest=sha256:42527e4b09e03ae9b1ef160b90534b2b8351137d4a2eb82038dd188550c58d04

Observation cd622e6e-eb10-4e62-9da2-049a44506978 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Bottom-up and top-down attention for image captioning and visual question answering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.351769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.065362Z digest=sha256:1037dace137e361fea6f33d0c4087ce0f83841263a248c952b88678cc6635c35

Observation 465e539e-eaee-4d12-b385-cb964f062c7c · outbound

This paper cites Women also snowboard: Over- coming bias in captioning models.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Women also snowboard: Over- coming bias in captioning models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.336902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.069958Z digest=sha256:dff25d85061cb100d3945044148933da41ef0cba1c4aaf6ec48ad83297604a8a

Observation cd20f5a7-7f92-487b-9e7c-18306e0943e9 · outbound

This paper cites Deep compositional captioning: Describing novel ob- ject categories without paired training data.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Deep compositional captioning: Describing novel ob- ject categories without paired training data

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.323001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.074636Z digest=sha256:fd5fadd3e7e2860250c7000ae24c0a60447433a520aa76a506ecfb7c7b6b02ae

Observation 37608bf8-7d38-4074-bf61-dad8e8ef08f0 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Lawrence Zitnick, and Devi Parikh

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.308732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.079192Z digest=sha256:dea2a9c59f6516e6701ce82fa124388c8415da98cdf94309ac8666751852fdc3

Observation f0696281-d196-425f-ae58-76177ceb549c · outbound

This paper cites Adversarial text generation via feature-mover’s distance.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Adversarial text generation via feature-mover’s distance

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.294420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.084024Z digest=sha256:26611883365aa67a34a0e3a4037be9c4a2548271f1ebcc63d831b2fc77c90424

Observation a6a209ca-8435-4475-8d61-c1b6eb3e5543 · outbound

This paper cites Show, adapt and tell: Adversarial training of cross-domain image cap- tioner.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Show, adapt and tell: Adversarial training of cross-domain image cap- tioner

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.280462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.088273Z digest=sha256:b7811264ed5ba67478e5d5b6e4e01f889a692ec004875a8b5710fa69aa568150

Observation 37b91ab8-d4c5-4820-a2d9-9baf2e576db4 · outbound

This paper cites Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.092525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.092525Z digest=sha256:7e91a486ef38be81c751f973593b8310ac14e99429cc63f77017bbc68fc0ce3a

Observation 14fb92d0-7223-468e-b608-9afa051ee48e · outbound

This paper cites To- wards diverse and natural image descriptions via a condi- tional GAN.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings To- wards diverse and natural image descriptions via a condi- tional GAN

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.266341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.097343Z digest=sha256:78361a8dbb218df2d1414cfb74c56bd915776a0f2e32877703337228d8b8b900

Observation a296666d-260e-4bad-a15c-60fbf54f8b2f · outbound

This paper cites Visual dialog.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Visual dialog

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.251995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.102402Z digest=sha256:c4338f133512f287a984788b49ae38992381de82951038033a07866f4a9dc827

Observation 9b6c5ffe-db60-4622-ad3a-ac17afd8e0ba · outbound

This paper cites Meteor universal: Lan- guage specific translation evaluation for any target language.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Meteor universal: Lan- guage specific translation evaluation for any target language

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.238471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.106878Z digest=sha256:ecdbe2b19318724e2e3aa0b48dc2968432d648d6a9b439ccd063fbb7c912e975

Observation 96ddcc7f-eeb6-4429-9539-e8ece88393e6 · outbound

This paper cites Long-term recurrent convolutional net- works for visual recognition and description.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Long-term recurrent convolutional net- works for visual recognition and description

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.224589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.111212Z digest=sha256:76cf4f59fad0ed47e73936b14521c5e9c009504ebe50e23d772ddaf688d1ac57

Observation 174b51ff-a2df-4233-a99e-92a9490c6535 · outbound

This paper cites Adversarial Feature Learning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Adversarial Feature Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.115213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.115213Z digest=sha256:ac5bcdcdce935a58a9e72a80dd405901bb40e3cdf7d4e2ef8ba1f24209af5cc1

Observation d557ae01-89a5-4f1c-8b2f-3ea5e75fa703 · outbound

This paper cites VSE++: Improving Visual-Semantic Embeddings with Hard Negatives.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings VSE++: Improving Visual-Semantic Embeddings with Hard Negatives

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.119851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.119851Z digest=sha256:4399026a601006578ca0a61415bcd6a904039695e87a25fa21176e4b3f2b0260

Observation aea4e3e0-1b75-4341-b3a1-40ca4affc4dc · outbound

This paper cites From captions to vi- sual concepts and back.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings From captions to vi- sual concepts and back

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.209802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.124345Z digest=sha256:3b4775e2ed5ca82f6a76067e2600f06e873afdf492034872e782ce4561760fe3

Observation 4866dddf-2f31-4059-b0bd-2177e7d0a151 · outbound

This paper cites Unsupervised Image Captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Unsupervised Image Captioning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-14T11:21:49.464577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.128637Z digest=sha256:b5c3dd49e9ec7b553a98501e22f0eb24d02c3f5348a05ec21f91414848c30c65

Observation 79f3acfb-a3a4-4087-bf21-c112bc25ec89 · outbound

This paper cites StyleNet: Generating attractive visual captions with styles.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings StyleNet: Generating attractive visual captions with styles

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.194138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.133156Z digest=sha256:fe8c80a56410c0ec3f6a4887a720ae8dcf60a3cb9e0ef01a0892c1078af6053e

Observation ce139965-5b8c-46f6-ad4e-724719239783 · outbound

This paper cites Un- paired image captioning by language pivoting.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Un- paired image captioning by language pivoting

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.179739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.138236Z digest=sha256:2ba555183bd1ab365646c5de6dc2feedf097f1dbc31b0c0f863114eece7871ae

Observation 5487240a-7548-4ed3-9600-ac34efdb04ac · outbound

This paper cites Improved training of wasserstein gans.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Improved training of wasserstein gans

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.165547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.142747Z digest=sha256:9c0f6ed391f0fb11bebf56f3f2a497a7c8b7920555cad3dc1f73796ccd4046ab

Observation 711de645-aeae-4efe-b938-683a37a07f0c · outbound

This paper cites MSCap: Multi-style image captioning with un- paired stylized text.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings MSCap: Multi-style image captioning with un- paired stylized text

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.151076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.147579Z digest=sha256:1cb14baad8ea155a02a22f684441ec34337cd5b5d2df4d0cd7a3126236464011

Observation 7d9848f4-4871-4f17-aedb-d955dd1e28a6 · outbound

This paper cites VizWiz grand challenge: Answering visual questions from blind people.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings VizWiz grand challenge: Answering visual questions from blind people

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.134384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.151969Z digest=sha256:664b1592e3c2de5ec5c5083a1916342c34895d9d2eaf98cb63fad45f7f6fd3e6

Observation f1255e8a-a030-44fd-bf18-fc2e7ac68918 · outbound

This paper cites Deep residual learning for image recognition.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Deep residual learning for image recognition

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.156222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.156222Z digest=sha256:b6bc5a79f63a875499794aec7b46f4456d50bb8156c3d33d29f65a6842fb0fb1

Observation f7e9902d-867a-47e4-acbc-3be5f12ab9bd · outbound

This paper cites Speed/accuracy trade-offs for modern convolutional object detectors.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Speed/accuracy trade-offs for modern convolutional object detectors

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.109732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.160518Z digest=sha256:0f79cba89eb2d58b4b672c224ee06f7814178d84eb7e7ed821f2f85b3eb71043

Observation 7412c2ab-6660-4bc3-b97e-1cbfe57af8b9 · outbound

This paper cites DenseCap: Fully convolutional localization networks for dense caption- ing.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings DenseCap: Fully convolutional localization networks for dense caption- ing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.094595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.164737Z digest=sha256:ab6820f3e2d88dbe135fa7b26c4fb026910d849667c31cf6871378fc69a4e2b6

Observation 13c9505c-4874-4445-b13b-17d8a512eed3 · outbound

This paper cites Deep visual-semantic align- ments for generating image descriptions.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Deep visual-semantic align- ments for generating image descriptions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.080284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.168914Z digest=sha256:2e3156a15d13abdf492bf2e69c1cf47b8f0c8e69a1b3ce1b0f869534b4901e10

Observation e921cc49-5e6e-4c28-bfa2-e04e0d2a58bd · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Adam: A Method for Stochastic Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.173388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.173388Z digest=sha256:26ad83978ff42e7e2f6c8903d2bf9389d9d74a6ac31b03742ba0d26ca340b92d

Observation 7256ea6d-0897-4078-a360-1a693a8a1913 · outbound

This paper cites Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.177786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.177786Z digest=sha256:e24826ab0144ee9a4c795e927abf843b49fa3335d473b94f13d56de266f294d8

Observation edbb9db4-cad3-4413-ad80-ab692e553d16 · outbound

This paper cites OpenImages: A public dataset for large-scale multi-label and multi-class image classification.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings OpenImages: A public dataset for large-scale multi-label and multi-class image classification

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.065623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.182853Z digest=sha256:658a8b9c3d04394d468b8288fac02992bc99d65075717dcd27118ebca151b3d2

Observation 140f1fa0-6291-4298-aac8-482c66e61aae · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.050852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.187321Z digest=sha256:e56cc397c27131d1901a8d0f245fec8666599f53a3ac1965d6b356a5a3ff7bf1

Observation 71e19b4c-db59-4d3a-8318-29c1d03ef217 · outbound

This paper cites From word embeddings to document distances.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings From word embeddings to document distances

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.035102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.191699Z digest=sha256:2fcdafbd1d5090547e200ee05922c0ff64cb0154fc2a87d508d9a2e1d24c36e2

Observation e639d1fd-85ff-4d48-91aa-73dce18fb0a8 · outbound

This paper cites The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.195941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.195941Z digest=sha256:a98b7fc4f990fcfddd5431d6150238c0b48288b5332c1768ce8223a54624e306

Observation 3547e4ff-82bd-4d82-913a-d8777c58e2ce · outbound

This paper cites Unsupervised machine translation using monolingual corpora only.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Unsupervised machine translation using monolingual corpora only

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.019131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.200592Z digest=sha256:1e5f8adf7f4105ee308f115568b1179a3c04e8c7ba89a03aa4a2b618cdae2f52

Observation 87046abb-bc66-47a2-ad59-99ee07ee2c10 · outbound

This paper cites Phrase-based & neural unsupervised machine translation.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Phrase-based & neural unsupervised machine translation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:50.003670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.204978Z digest=sha256:90c556ad51b6f318bbe69542b307fde3710fc08750371dac294cb5aaefefa2b0

Observation 75b16d00-50a4-4480-94a6-2fd4d930ec3a · outbound

This paper cites Generating Diverse and Accurate Visual Captions by Comparative Adversarial Learning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Generating Diverse and Accurate Visual Captions by Comparative Adversarial Learning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-14T11:21:49.400024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.209513Z digest=sha256:5f66b39e633c7aefca04ddf92c99b8312ee9b0915cd5be072a96e8c4d577a02c

Observation 1966479c-1b24-4bb0-9faf-d738c6c39e50 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Rouge: A package for automatic evaluation of summaries

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.989330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.214205Z digest=sha256:98ba01e46a2977be49599088c20abc45134cb78dd4db40cf8a4c349cf37f8d5b

Observation 11b42c84-79c1-4e7b-ac00-ae6e040add50 · outbound

This paper cites Microsoft COCO: Common objects in context.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Microsoft COCO: Common objects in context

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.974238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.218844Z digest=sha256:85a4abfbc9daecad1863d0ca6c7587133e4fe76da6b40de4d577c66aeac04f72

Observation d0480356-7346-47b4-9fcf-19c9de0e11dc · outbound

This paper cites Teaching machines to describe images via natural language feedback.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Teaching machines to describe images via natural language feedback

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.958902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.223234Z digest=sha256:2bf88a37a7b49fdf266f9abb1ce3b7f650be04d8048eaa283ad324194e48236a

Observation 5b174ceb-e1e2-44d7-bef1-d72524f35b61 · outbound

This paper cites Improved image captioning via policy gra- dient optimization of spider.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Improved image captioning via policy gra- dient optimization of spider

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.228016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.228016Z digest=sha256:27e6b75219d28cd3737d48c89876f86a224902a08b19b4ca37bbf0e86ecd6f38

Observation a4a589ba-ea53-45d1-b4cd-00b88a060639 · outbound

This paper cites Knowing when to look: Adaptive attention via a visual sen- tinel for image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Knowing when to look: Adaptive attention via a visual sen- tinel for image captioning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.232199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.232199Z digest=sha256:c16d4f865f2a5834577415f2d575d4e7a5c228e115056173ed7a8019f5f6a7c4

Observation 5297d2d2-0ce3-4cd1-a37d-29822afa472b · outbound

This paper cites Neural baby talk.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Neural baby talk

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.925602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.236314Z digest=sha256:9910c3f557c2dd610e6d6daba6f2d0b10dc8cab37ee4e87f164eb8e0e9af4eee

Observation 4646014e-da54-4b69-b8b8-d63ddf7abc43 · outbound

This paper cites Visualizing data using t-SNE.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Visualizing data using t-SNE

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.911058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.240435Z digest=sha256:7fcee5f92f833c6b102a61ed5456f5064bf79a74b4f582eeaa389f70a3198fef

Observation 61847bfc-ef07-40d7-99c1-324254d67622 · outbound

This paper cites The stan- ford CoreNLP natural language processing toolkit.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings The stan- ford CoreNLP natural language processing toolkit

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.896734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.245049Z digest=sha256:ad08012ad678e691e5d6342ae3cb9df01bb0052b263a85cc6b1de452a4b4251e

Observation c8eec48e-6ba1-4a5d-ad17-e8f0b04ea120 · outbound

This paper cites Learning like a child: Fast novel visual concept learning from sentence descriptions of images.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Learning like a child: Fast novel visual concept learning from sentence descriptions of images

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.883100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.249691Z digest=sha256:2a44fca57ff1c9d6d554b65ebb13ac5cf0cead8d537bab4bc7177896fa3d2c7f

Observation caaacc1c-8c40-455b-8dde-0f7985c42798 · outbound

This paper cites Sem- Style: Learning to generate stylised image captions using unaligned text.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Sem- Style: Learning to generate stylised image captions using unaligned text

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.868624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.253865Z digest=sha256:7e42527f0d724e7d3b872c3a3c5692282c0200643b9c69240705124548eef591

Observation 081451b6-48d3-42cd-8ff9-c1a9c6202780 · outbound

This paper cites Jointly modeling embedding and translation to bridge video and language.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Jointly modeling embedding and translation to bridge video and language

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.853953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.257982Z digest=sha256:02fa2c33655d769ead700328221b4e407e0e6377b6512b2600c26064969c7647

Observation 210a1364-a49a-4038-a1dd-46f43fbb0cb8 · outbound

This paper cites GloVe: Global vectors for word representation.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings GloVe: Global vectors for word representation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.839595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.262233Z digest=sha256:63bff2d02da96312ba65308eeef654fe18d8b6265218a3afb3309258f0b25743

Observation 9d83399a-4384-45d3-8d09-8f5d91b25f1f · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.825483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.267346Z digest=sha256:08e8f88615d501cabc82874fd19298055e6f1e992f74133b0ff8d1e781924a2a

Observation b2437259-aa6f-40e0-a48e-16a13e1e9d77 · outbound

This paper cites Self-critical sequence training for image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Self-critical sequence training for image captioning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.810801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.271747Z digest=sha256:b7d1d6156117c4e27e062d009feb83c6b1ec9a18852a1f88130dfc32b587f3b5

Observation 08d40f3f-c859-47e4-af41-b5c0fd0e97c8 · outbound

This paper cites ImageNet large scale visual recognition challenge.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings ImageNet large scale visual recognition challenge

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.796526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.276815Z digest=sha256:e1c640c89af1a5d730e0a81493e872a51df6409fb650a2619d592295f37bd7df

Observation 0fd873de-d811-4278-bf7c-5eb2202282fb · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.781359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.281797Z digest=sha256:7dc12652dd73280b1a88fc4790885c7c83cfc61ddf24ca72bc60760e2a67d67d

Observation 8d405e70-3988-446d-8598-da7516652da1 · outbound

This paper cites Speaking the same language: Matching machine to human captions by adversarial train- ing.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Speaking the same language: Matching machine to human captions by adversarial train- ing

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.764925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.286027Z digest=sha256:0742ec577905c4307916f9879921f527ec00c2129c82f90e474c77cc2804d7ee

Observation 21de770f-2a64-44db-be90-6b321b68c7ea · outbound

This paper cites Deforming autoencoders: Unsupervised disentangling of shape and ap- pearance.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Deforming autoencoders: Unsupervised disentangling of shape and ap- pearance

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.750435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.290120Z digest=sha256:321a9f2ce7150e8a971f3f80cbd15b30b7cf912bf84ca8f87c2cd41376dd06be

Observation 24d7acb7-651e-404a-b054-6fb0e294b2af · outbound

This paper cites Engaging image captioning via per- sonality.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Engaging image captioning via per- sonality

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.736319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.294410Z digest=sha256:a9751074dc9d04a20f164297a4646f4240c5078c391fb3cce3dfd4fa2d23e2db

Observation eb488959-56e3-4534-83f7-3e5b4a2a5dcc · outbound

This paper cites Towards text generation with adversarially learned neural outlines.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Towards text generation with adversarially learned neural outlines

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.722240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.298686Z digest=sha256:bdd62612b9b95090003342d30a371257608a6e1e93a2d18704c4faeb2f432ba8

Observation 3b62f948-9c4f-425c-87a7-e44dc90b407d · outbound

This paper cites Sequence to sequence learning with neural networks.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Sequence to sequence learning with neural networks

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.707847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.303007Z digest=sha256:cf3554cd40a5b3f3aabc4d9a6c2ae69492e4bef9f6315e4c744d6683627ba190

Observation 2ca248f7-0463-4f3b-be81-6aca77e5b0f4 · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Cider: Consensus-based image description evalua- tion

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-14T11:21:49.307627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:21:49.307627Z digest=sha256:f4df4962961ba1e3571819d3f54edb10961dc1e73f5e3a08da41fe649a8e6711

Observation 1e69ab5a-de3b-4ad6-abcc-de758b06f2a1 · outbound

This paper cites Captioning images with diverse objects.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Captioning images with diverse objects

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.684154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.311808Z digest=sha256:29d2a77138332eb0096d38f74ce4b6db58274ecc37ea6d458eca323bf6652770

Observation add7915b-047b-4bb5-9fb3-b86bc1d53de2 · outbound

This paper cites Show and tell: A neural image caption gen- erator.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Show and tell: A neural image caption gen- erator

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.669643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.316206Z digest=sha256:ed76e7c6013cf13e213270eabcd15b1e77dea7be247fe015c96bea513f696554

Observation aba98efa-5dc5-4e67-ac51-3ba198649935 · outbound

This paper cites Diverse and accurate image description using a variational auto-encoder with an additive gaussian encoding space.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Diverse and accurate image description using a variational auto-encoder with an additive gaussian encoding space

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.655861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.320325Z digest=sha256:cae5d4749935b5d81c8907684af3bd90bc1536cb6aecb8b0ca4695c4ceb3e808

Observation f563f42d-6c22-4729-a433-657eb34d8d69 · outbound

This paper cites Automatic alt-text: Computer-generated image de- scriptions for blind users on a social network service.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Automatic alt-text: Computer-generated image de- scriptions for blind users on a social network service

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.641250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.324482Z digest=sha256:9acbf2ec8d95c8b8998da886a049b321663095e064e290450024262f2fc76953

Observation ae702d51-7242-4aeb-a8c3-04d1e674e811 · outbound

This paper cites Show, attend and tell: Neural image caption gen- eration with visual attention.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Show, attend and tell: Neural image caption gen- eration with visual attention

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.626520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.328768Z digest=sha256:ea51fb64dc7cb9b8c008c43be437975633d2ab42255ec9aa975c88a3ce082606

Observation 3672b3b3-97c3-459d-8afe-c45b03a8cb18 · outbound

This paper cites Review networks for caption gen- eration.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Review networks for caption gen- eration

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.612477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.332844Z digest=sha256:e63a479267e1c68932dbb189b71e7a7eaf42c94dfc7a2116036aa4bee9536f1d

Observation 0f0f9291-e594-40ce-a1ad-0e53b8936924 · outbound

This paper cites Incorpo- rating copying mechanism in image captioning for learn- ing novel objects.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Incorpo- rating copying mechanism in image captioning for learn- ing novel objects

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.597659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.337040Z digest=sha256:158f2bac225e543c9dad3106e5f63718a6c6a7a8bfa7fd8f015a51c41f4c8354

Observation d2256260-1f4b-4f0b-ac04-a156ac8daaca · outbound

This paper cites Explor- ing visual relationship for image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Explor- ing visual relationship for image captioning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.580762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.341293Z digest=sha256:1b88afcc9d312e76e7b1539f3e73ee9b364606fdf9670b03bc98f9bd99f086ec

Observation fcb4a1cd-f968-4957-9b8c-652f58199050 · outbound

This paper cites Boosting image captioning with attributes.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Boosting image captioning with attributes

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.565477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.345606Z digest=sha256:4ee4d6fa68f636ef01c2ec9d13446b58b1ee8b729ea6eb36c170609a82e9f656

Observation 0aee144c-dc01-4e1f-a6ad-c9d1b1f0af0d · outbound

This paper cites Image captioning with semantic attention.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Image captioning with semantic attention

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.550752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.349967Z digest=sha256:e339d0fe5c766266af730381b61b473b71c62fc39313b787e35a97f737b03846

Observation 168df444-f91c-4e67-b062-2d06e4af929c · outbound

This paper cites Dual learning for cross-domain image captioning.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Dual learning for cross-domain image captioning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.536687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.354148Z digest=sha256:938ff89c28e291f37547e434de3e5c0f0773d5ae5e328b53cc0a896af3215c5e

Observation 51603e1c-a047-49ea-af49-27d042c7c6c5 · outbound

This paper cites Unpaired image-to-image translation using cycle- consistent adversarial networks.

Towards Unsupervised Image Captioning with Shared Multimodal Embeddings Unpaired image-to-image translation using cycle- consistent adversarial networks

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:21:49.522176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T11:21:49.358294Z digest=sha256:b918570ae08822ded7d80f0b3820fa2a4d022853f9ec422531715600aa3698af

Pith citing papers

No inbound Pith citation observations are available.