Pith. sign in

Paper Citation Record · LEDGER

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck

As of 16 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:1908.07094.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.07094 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:32:40.579862Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 35ad6b23-c40c-48e7-a66f-bda27a2c4e9d · outbound

This paper cites Deep speech 2: End-to-end speech recognition in english and mandarin.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Deep speech 2: End-to-end speech recognition in english and mandarin

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.728991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.187215Z digest=sha256:23636608c6953be0cdfc3a54639710315eeb210504f6b02472cd49f5c7d7b9be

Observation cbaadc98-e1c7-4c7d-928c-a793e9a2c4de · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Bottom-up and top-down attention for image captioning and visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.208110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.208110Z digest=sha256:f1312a1f523406947849996687945caac4baaec82753d54bfe44470f8c4d37d6

Observation 1509ed93-aed0-40c6-9850-3f3386e0f534 · outbound

This paper cites Deep voice: Real-time neural text-to-speech.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Deep voice: Real-time neural text-to-speech

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.703952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.214283Z digest=sha256:09588ae49e84c3e06a60c9e441ec502cef43b43d046353f846265f3d3602004d

Observation 9ea4c8e1-891d-4262-988b-de226798bd6d · outbound

This paper cites One-sided unsupervised do- main mapping.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck One-sided unsupervised do- main mapping

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.689013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.222855Z digest=sha256:a248c65cab7f349aacf45fb385dcabe018cdc8c20c91273a2cf6ba3bd9f50857

Observation 15406ef7-2d88-4131-9483-ecf74b271bdb · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.233991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.233991Z digest=sha256:2bae69e204e972d18ab901c5fea737d5523f3fb0137af3d1958d58c1eb4f309e

Observation 0918578d-903d-408a-9083-37f831003fcd · outbound

This paper cites Weiss, Kanishka Rao, Katya Gonina, Navdeep Jaitly, Bo Li, Jan Chorowski, and Michiel Bacchiani.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Weiss, Kanishka Rao, Katya Gonina, Navdeep Jaitly, Bo Li, Jan Chorowski, and Michiel Bacchiani

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.676458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.239835Z digest=sha256:e8d89287909977e9e5821a88956fd27dce1f760f75555030d45305832889267c

Observation 6d75fb6c-4c44-49e7-91e8-b631717e85e8 · outbound

This paper cites Learning phrase representations using rnn encoder-decoder for statistical machine translation.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Learning phrase representations using rnn encoder-decoder for statistical machine translation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.662912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.244801Z digest=sha256:2e8a7501ef2eb1353b6581e622b2fcb30866a1e2a6c8703ce1e66003cc8a2b9c

Observation 0dba30e2-25b4-461e-88c4-5c7022836145 · outbound

This paper cites StarGAN: Unified gener- ative adversarial networks for multi-domain image-to-image translation.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck StarGAN: Unified gener- ative adversarial networks for multi-domain image-to-image translation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.644638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.249480Z digest=sha256:29b957009f1a51e1ba34b03310a2728f5d2fe4f22f888ce00056b7815f65a4bf

Observation 0fbf77f8-2d26-4c74-84bb-23eca167d1c3 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:41.629458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.256473Z digest=sha256:d53fe4d0d7f19ca52f27a480fa1d09c6b75b6884b121b6aaff90209f41599b2b

Observation ab4d28ca-1ed6-4844-88a4-38219b608731 · outbound

This paper cites Unsupervised domain adaptation by backpropagation.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unsupervised domain adaptation by backpropagation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.614632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.263495Z digest=sha256:13905f05eba8319a0739f796546e65e378db33c2aecf0f185ddfda141e2d9e4d

Observation 3f08e268-2703-4f73-8cf7-420252166d2a · outbound

This paper cites Auditory-visual inte- gration during multimodal object recognition in humans: a behavioral and electrophysiological study.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Auditory-visual inte- gration during multimodal object recognition in humans: a behavioral and electrophysiological study

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.596764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.269378Z digest=sha256:c5b6512809230d7928b11035919100074fce30483fceed91a0b9125b38dfb704

Observation f545e8c4-795e-48d1-bde4-d3a0b34a6ae5 · outbound

This paper cites Deep voice 2: Multi-speaker neural text-to-speech.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Deep voice 2: Multi-speaker neural text-to-speech

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.581352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.275156Z digest=sha256:1423ad5d254c1054c7025939263ba5c91c89776a6ae0e37c39c0ac8bc6fc5efa

Observation a463a38d-d7be-409b-9e47-46c3bc12fa0b · outbound

This paper cites Generative adversarial nets.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Generative adversarial nets

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.565415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.280725Z digest=sha256:b64e6e0f78396b98741583e55c74b21da28ed81051b974ffb09cabb52ea77ada

Observation 55867ffc-d869-4696-b7f0-c23500615042 · outbound

This paper cites Griffin and Jae Lim.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Griffin and Jae Lim

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.542863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.287396Z digest=sha256:a9c9bf49c8c9aa03f153e2a3ec276c94516b1402aea36c572e3769531a26e3d4

Observation a9f517cf-adea-4265-83b5-cf604f0182e0 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:41.524456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.291930Z digest=sha256:71c37687f2172fb0d84e6218e2f348b9cdb33803a3f74c3d7d254b7aa10ae6b0

Observation a901f9a3-bdd6-4f8a-9474-2796c9971ffa · outbound

This paper cites Deep neural networks for acoustic modeling in speech recognition.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Deep neural networks for acoustic modeling in speech recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.507327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.299041Z digest=sha256:40900262f4a62440ba6f1b316894777280ded550d6af35d7000ea7039deb12a5

Observation ec1e21c5-7697-4666-a19b-21c451dc9a46 · outbound

This paper cites Huang, Z.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Huang, Z

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.492668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.304045Z digest=sha256:f8507b1464d3cf20b19b7a227eff684c25f4dcb53967673a0182e3676a1e2f49

Observation 4f98a6bf-714b-4324-b9c3-e4446dd196aa · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:41.477805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.312903Z digest=sha256:0e975746835a594ac3b04374876ab6d0e9e0bc5ed17ea164c755643b7a04d07e

Observation 33c9fec7-9ce1-4584-99ab-329778a5c0b0 · outbound

This paper cites Recurrent fusion network for image captioning.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Recurrent fusion network for image captioning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.463953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.318633Z digest=sha256:58de01c9fac31e24df7abea7741d415c09d579ad22effb39111e440c36f2845b

Observation 6522e939-ca0e-4114-b621-468754b8b2c3 · outbound

This paper cites Learning to discover cross-domain relations with generative adversarial networks.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Learning to discover cross-domain relations with generative adversarial networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.447931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.323490Z digest=sha256:32749c62fc91e1faf3c2b15a324ab953e26bc9c550798d406360e3c930fd8c77

Observation 3dbaa0dc-3b94-4030-8570-69154d0a83a1 · outbound

This paper cites Adam: A method for stochastic optimization.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Adam: A method for stochastic optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.328415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.328415Z digest=sha256:318077a773f03fa046398102b9dd17bc4a1eb944da9b107b4b070bde69a2097c

Observation 39a0947e-dbaf-4c35-85a7-b63450f74f29 · outbound

This paper cites Auto-encoding varia- tional bayes.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Auto-encoding varia- tional bayes

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.416403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.334491Z digest=sha256:2f69c650b78b70a11cf8c33e4c17229b8c0623973ddbe7e6c8d271263c7cfed3

Observation d394991e-db9c-4945-9809-0f3aeb1ff483 · outbound

This paper cites Shamma, Michael S.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Shamma, Michael S

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.398799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.339109Z digest=sha256:49d23f7d06ded1ee2d35af9690caa041275b13cba9401364ca70fcdff083836d

Observation 0df4bacb-40aa-4a6b-841f-f526c8ad1899 · outbound

This paper cites Letter-Based Speech Recognition with Gated ConvNets.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Letter-Based Speech Recognition with Gated ConvNets

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.343688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.343688Z digest=sha256:68fa01e53f7c218310188ecf7782eb467bfa296e33f82116e115b8022ade9736

Observation 97f37b93-b26e-48fc-9836-b7d6f98cb6f2 · outbound

This paper cites DA-GAN: instance-level image.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck DA-GAN: instance-level image

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.378509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.349055Z digest=sha256:ff54e2546c30422fba53c5acb78d447742f9539be0393f5319dad236097a71f1

Observation a0dc435e-e9af-4f45-a0f4-04a77d8939ee · outbound

This paper cites Neural TTS styl- ization with adversarial and collaborative games.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Neural TTS styl- ization with adversarial and collaborative games

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.363667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.354032Z digest=sha256:f19f5d910702d158158a206a0a7b678d50d4e8f47e62c81f11f9b47501fe1c50

Observation 6c2909de-9943-465a-b9a9-84e0aa189c2b · outbound

This paper cites Semstyle: Learning to generate stylised image captions using unaligned text.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Semstyle: Learning to generate stylised image captions using unaligned text

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.347881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.359162Z digest=sha256:0e803e2eafc65fca320da3d66fa7552e319b50683f1d69ba5e9bb61297e5b134

Observation e963026b-5cd3-49a3-9fd2-e22378dd0b3b · outbound

This paper cites Deep multi-scale video prediction beyond mean square error.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Deep multi-scale video prediction beyond mean square error

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.334089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.363608Z digest=sha256:ba2f65228dbf0519913bf49f41be32441a64a3e31d2c01fd0cc3aea771fcc4d2

Observation 246b088b-86e9-476a-83d3-30e418bd7a41 · outbound

This paper cites Fitting new speakers based on a short untranscribed sample.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Fitting new speakers based on a short untranscribed sample

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.320167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.368936Z digest=sha256:93387be44cbf0e8247e04a0ed483792c32a76af7dd4c4174198cbf6babbfafdd

Observation 183b6c45-042c-45ee-8f6d-f2a6a96a0ba2 · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Librispeech: An ASR corpus based on public domain audio books

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.304579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.373929Z digest=sha256:afef2adbb855f0ed95fda219155a6b28091f7763e00a6616bb2767b442fed3bf

Observation 9cf086df-c52c-485a-9e60-b0c2df823191 · outbound

This paper cites Attend to you: Personalized image captioning with context sequence memory networks.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Attend to you: Personalized image captioning with context sequence memory networks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.289510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.379814Z digest=sha256:7b136608e8447c38620047c74d8f56a869a4bfd3faf41d1bd5395b9459c002d6

Observation f73743cb-664a-4c60-b154-e764c072554d · outbound

This paper cites Beyond sensory im- ages: Object-based representation in the human ventral path- way.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Beyond sensory im- ages: Object-based representation in the human ventral path- way

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.274190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.384718Z digest=sha256:6318db50cfa9e0d80b441e3a93eb7e6bebbaf9df246814b0dce5922b6e328235

Observation 2ebac211-0c85-4b2a-a2db-75e0951216f2 · outbound

This paper cites Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.258001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.393548Z digest=sha256:a42ef1e4d97e41f80f250b9286131577f353b4a9998da66cb44dbdeab43fdec3

Observation 81148e27-f988-4720-b030-3d5de3180dc6 · outbound

This paper cites Generative ad- versarial text to image synthesis.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Generative ad- versarial text to image synthesis

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.237503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.398847Z digest=sha256:9ca05ebd670e4e992d6d72ddf060250e79867fd111dc8817150ac5642447c6dd

Observation 1e437235-6436-4bd5-921f-496d3aa71b31 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.221378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.403456Z digest=sha256:f96bac7e0be3daedf3f0a4ccf7ca672661b82818e696b961b8f5fa5abddb91a9

Observation cc08e4ad-b405-4ca7-a066-d8804c3554de · outbound

This paper cites Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.409378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.409378Z digest=sha256:021b7f34904c5a9ba2fb518d6cce26719bffb80ec31431b35327ea97211db2ac

Observation 39f2086d-8a96-4a69-b83e-0abc87399f1f · outbound

This paper cites Char2wav: End-to-end speech synthesis.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Char2wav: End-to-end speech synthesis

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.202586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.415356Z digest=sha256:0fad9675dc9a5980f7d333fc26af0ccb684114852730590a125b93b12df17dcb

Observation 15590278-f90c-4c7a-b371-9e1ad76a8ddb · outbound

This paper cites Szegedy, Wei Liu, Yangqing Jia, P.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Szegedy, Wei Liu, Yangqing Jia, P

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.188088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.420662Z digest=sha256:8e214b7191986335fb3fe545f5f8861268ae9106cb1a9907054bb88fdfd065a0

Observation 7494df59-bec1-44c6-9b46-d2080ae75d06 · outbound

This paper cites Unsupervised cross-domain image generation.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unsupervised cross-domain image generation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.170542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.425757Z digest=sha256:e2c2e7e47f0f53bbbde690c4a7be2aa5a439ea709eec2c7a3666dd2e66318b0b

Observation b60e1574-6217-40e3-ac05-4b017bf2481b · outbound

This paper cites V oiceloop: V oice fitting and synthesis via a phonolog- ical loop.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck V oiceloop: V oice fitting and synthesis via a phonolog- ical loop

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.154400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.430623Z digest=sha256:b7801059a5eb161554c03c70e84ec5e1ff86b8ae5e1bb2369dafbfc4d960f15d

Observation d91550cd-0fe8-416f-953a-1b62f9056112 · outbound

This paper cites The information bottleneck method.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck The information bottleneck method

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.139967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.436106Z digest=sha256:f049e849adc6fa728cd1c3c875880e5174c6e1d7d0331d9b03d05facbd5af2d5

Observation b5488199-086e-4b76-b792-c0391b49f077 · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck WaveNet: A Generative Model for Raw Audio

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.440702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.440702Z digest=sha256:b6645d2cc318f6dc3c45d75122461448dd961cdf21b4fa803d5ad6316bc7f791

Observation beaf1d4a-2160-4e1e-99a8-cc27cecad1aa · outbound

This paper cites Attention is all you need.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Attention is all you need

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.445811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.445811Z digest=sha256:35547bf5a0bd2099954c404f59e885992afaa009f03c014de72946c4f0e9848b

Observation c148cdd4-3077-4afe-9cab-f11e86832d9f · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.116749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.450090Z digest=sha256:45be3075db34311e57e86115d79595a99b35578227cb4bb7f1647ea91ef1cbc7

Observation 27f4cf11-4028-4c02-9d5b-7aadf127df87 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:41.098764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.455622Z digest=sha256:5a5f6228a1521d3edce6560e94f9d67a5427404af4d93cc8c892a838688be0b2

Observation 2d9e1973-0cce-46d7-b994-8c91366a3021 · outbound

This paper cites Show and tell: A neural image caption gen- erator.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Show and tell: A neural image caption gen- erator

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.079532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.460412Z digest=sha256:7dc1c38d156b733b48def96026fc36f346826e99747b3934db7077b480a7627e

Observation 0689af5f-ddea-4cec-bea7-52183f9196c5 · outbound

This paper cites Tacotron: Towards end- to-end speech synthesis.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Tacotron: Towards end- to-end speech synthesis

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.058522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.465080Z digest=sha256:d958f261f8f25d1477d638dc447b390a820c6027e71f0bd906fd2029ea4f3d9a

Observation 7b0e0399-1fdc-4ee2-9582-09cbbd2aeea6 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:41.042059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.469470Z digest=sha256:d6168c48c87966d232366c9b7757afca0833aa1a8b154a293b16396d4572cf66

Observation f35ff7ab-c1c2-438b-8ad2-39c01ecbca5b · outbound

This paper cites Memory networks.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Memory networks

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.024937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.473965Z digest=sha256:11a7344f0cf0532267a7e78cc5e686c9a8166eb3e9498b5c607668d3f881e3a0

Observation 68744106-0ce7-4a5a-9bdf-fece81af9708 · outbound

This paper cites Courville, Ruslan Salakhutdinov, Richard S.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Courville, Ruslan Salakhutdinov, Richard S

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.478934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.478934Z digest=sha256:dd3c0a72278650cfe645274ac5ef790e8ab49937d2181ea7716e53a06a0bd09c

Observation aeedcc49-4e5b-4512-ade5-3389ee9adb75 · outbound

This paper cites Attngan: Fine- grained text to image generation with attentional generative adversarial networks.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Attngan: Fine- grained text to image generation with attentional generative adversarial networks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:40.989008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.482808Z digest=sha256:de8ac261cb9892bf3ba6202f4f654c045441d5ba8100f7bea0bf5ba11f40f9be

Observation 3e2215f9-f6e5-4012-870b-d75b57324378 · outbound

This paper cites Image captioning with semantic attention.CVPR, 2017.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Image captioning with semantic attention.CVPR, 2017

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:40.970995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.488513Z digest=sha256:4dab2be00fc932e62833b30d8b0a5dbf8471c27aef5b4180a9b0646db2f034a5

Observation a0e63788-9d66-41a1-a2ca-70f492b48c75 · outbound

This paper cites Stackgan: Text to photo-realistic image synthesis with stacked genera- tive adversarial networks.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Stackgan: Text to photo-realistic image synthesis with stacked genera- tive adversarial networks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:40.948778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.492506Z digest=sha256:bb7b6853d16b5140162dc9fc7a0fb17424a8466e395a532d5d53484319a0171e

Observation 60d31258-403f-42cf-9b6f-f52dcfab64e9 · outbound

This paper cites Visual to sound: Generating natural sound for videos in the wild.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Visual to sound: Generating natural sound for videos in the wild

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.496609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.496609Z digest=sha256:8d9634cad7d8c5b6ebb05486dda6a318ac1925471b108b18822bbe501346d55c

Observation da4a2775-032b-41c5-93ea-9445644bcae2 · outbound

This paper cites Im- proving end-to-end speech recognition with policy learning.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Im- proving end-to-end speech recognition with policy learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:40.921821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.502211Z digest=sha256:7ed2ba6338850cbe0bb7003303fc5cc6902a727d656231fd7d147ea8c09e0c4b

Observation a3c5a992-6bd7-492b-b704-36bc6a53f641 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.899587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.507103Z digest=sha256:39a77b8960e0540b34f8796937a302c792b3fc7c0d3e93e1b19a586f61b50483

Observation ec905cbd-4e7b-4139-94d4-c13339b689f3 · outbound

This paper cites Efros, Oliver Wang, and Eli Shechtman.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Efros, Oliver Wang, and Eli Shechtman

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:40.884162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.512601Z digest=sha256:e0ec27beb2076a341ef98ff5e914d36b15a4d6e6f8c24e2f3f8c8c7745ee64ac

Observation 1a4ac184-6260-4cac-baf4-fb788c5abeac · outbound

This paper cites We encour- age the readers to refer to Figure 3 and Figure 4 of our main paper when reading this section.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck We encour- age the readers to refer to Figure 3 and Figure 4 of our main paper when reading this section

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:40.866425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.517413Z digest=sha256:ea4ffad093adf2a76e6fae31a320420bf082a9e9a4ce9f2eea1b163ba80c2ab3

Observation 5399a283-5229-4460-afb4-13fe1a6708d2 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.849508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.523658Z digest=sha256:21183c89e193307485b91dc4a242833f34f323645bfb16bb27891e1da12a4520

Observation 1cb707da-a64a-4932-b58d-57eddc55bec4 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.834024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.532088Z digest=sha256:a1029554b54f9e8f7417f413798d49b480a570163d237ba3f6b8027023995d77

Observation 43abc4ef-1f57-4c11-bafb-174ade716b24 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.816482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.537200Z digest=sha256:62d92690d451ffe42b0dc9a0b2178f366b4cc9a8a53877212a7dcaa44d646e57

Observation 224ff9f3-0ebb-4843-8926-383cad96ea42 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.800184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.542135Z digest=sha256:993f2ef98459f89f30e193e6253487d29df007aed954d4ed81ae9ff63a9d456e

Observation c04a12fe-06cb-46ad-b544-7e871c475c59 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.784457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.547781Z digest=sha256:cc90aadbc25cea4e628642f6a70f0bd1d3dd72428bbaaecc321340f56890497b

Observation c79de29b-7e42-4cd9-a4e7-5369cc87ce5a · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.766624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.553877Z digest=sha256:2f7c3fda714bde390004c960c1ab607bb5d1b10d8abae250f464e2ebfc9e8fb6

Observation e3f706e1-9dc3-4c1a-a3c8-f80baa5dbd8a · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.735273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.564751Z digest=sha256:c1f26409d42059bc34622db59bc51d7e54ce3f721e83b30bae7fc026cbd67fce

Observation e61b71e8-6b17-4bbf-81e9-5fab85ad1eb1 · outbound

This paper cites dog” and “zebra.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck dog” and “zebra

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:40.712826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.569987Z digest=sha256:5f1b864852e9b1d5abca6fc6a258f1a20f4c507dae8699177328958091157169

Observation 33f8c0bb-e942-47d4-b291-eabd291d05b6 · outbound

This paper cites bottleneck.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck bottleneck

Reference 70

Resolution
malformed identifier
raw_fallback, observed 2026-08-14T12:32:40.696611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.579862Z digest=sha256:28ae16dd40578bf059624a22412a2dab7f8b5e7be809b7944c76a60eb648625d

Observation 7c2b0122-2aba-4a76-a0b5-7d5f60ae7a6a · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 256

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.751501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.559055Z digest=sha256:f3e98424558fedb1774ab2caa82d5edcb5104bbd687de79c0368a7a8c9ba1fc1

Pith citing papers

No inbound Pith citation observations are available.