Pith. sign in

Paper Citation Record · LEDGER

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck

As of 16 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:1908.07094.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.07094 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:32:40.579862Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 35ad6b23-c40c-48e7-a66f-bda27a2c4e9d · outbound

This paper cites Deep speech 2: End-to-end speech recognition in english and mandarin.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Deep speech 2: End-to-end speech recognition in english and mandarin

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.728991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.187215Z digest=sha256:76f00d53b7a561c62cb6e242717672bab09e6dcec8b8d653d447c7794d002240

Observation cbaadc98-e1c7-4c7d-928c-a793e9a2c4de · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Bottom-up and top-down attention for image captioning and visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.208110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.208110Z digest=sha256:a736097555784189840c21f3b887423cd9a7f3fbb2ec6ef8f959322e05ffad95

Observation 1509ed93-aed0-40c6-9850-3f3386e0f534 · outbound

This paper cites Deep voice: Real-time neural text-to-speech.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Deep voice: Real-time neural text-to-speech

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.703952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.214283Z digest=sha256:e83b845d5a7aefd904955d7fe21890f842bdf3489f9299d4e2f07a802cc4e8d4

Observation 9ea4c8e1-891d-4262-988b-de226798bd6d · outbound

This paper cites One-sided unsupervised do- main mapping.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck One-sided unsupervised do- main mapping

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.689013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.222855Z digest=sha256:c43cef059ca512019e2e2132d12f64e1c3bcfba60aebd0098e613b2732811bea

Observation 15406ef7-2d88-4131-9483-ecf74b271bdb · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.233991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.233991Z digest=sha256:b2217aa3fae91de836f1c6f98ec3399bc2a86947d07273b95758adc96c400343

Observation 0918578d-903d-408a-9083-37f831003fcd · outbound

This paper cites Weiss, Kanishka Rao, Katya Gonina, Navdeep Jaitly, Bo Li, Jan Chorowski, and Michiel Bacchiani.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Weiss, Kanishka Rao, Katya Gonina, Navdeep Jaitly, Bo Li, Jan Chorowski, and Michiel Bacchiani

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.676458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.239835Z digest=sha256:fe578e4b86b180cc0391c7dd058ebe44fe4a7c15bdfdc9925321907078424097

Observation 6d75fb6c-4c44-49e7-91e8-b631717e85e8 · outbound

This paper cites Learning phrase representations using rnn encoder-decoder for statistical machine translation.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Learning phrase representations using rnn encoder-decoder for statistical machine translation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.662912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.244801Z digest=sha256:3d0d607a7c1939bbff0505175a2df78c244a9add23fac570587b1895b962807a

Observation 0dba30e2-25b4-461e-88c4-5c7022836145 · outbound

This paper cites StarGAN: Unified gener- ative adversarial networks for multi-domain image-to-image translation.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck StarGAN: Unified gener- ative adversarial networks for multi-domain image-to-image translation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.644638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.249480Z digest=sha256:1977e8b3bbabb1d90bdf314b2cb474efb1c80a780439697e7404a6f5f89abcb3

Observation 0fbf77f8-2d26-4c74-84bb-23eca167d1c3 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:41.629458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.256473Z digest=sha256:8e2d1ead7e760eaa8ce5723804e9d00fc50d566f153c23bfd58b44a4cdf7b138

Observation ab4d28ca-1ed6-4844-88a4-38219b608731 · outbound

This paper cites Unsupervised domain adaptation by backpropagation.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unsupervised domain adaptation by backpropagation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.614632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.263495Z digest=sha256:178224c9371f60460ee9ae1a04609973320a49fa47852ba339e1926b298ee65f

Observation 3f08e268-2703-4f73-8cf7-420252166d2a · outbound

This paper cites Auditory-visual inte- gration during multimodal object recognition in humans: a behavioral and electrophysiological study.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Auditory-visual inte- gration during multimodal object recognition in humans: a behavioral and electrophysiological study

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.596764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.269378Z digest=sha256:3e58c6f2b3b52a1f1d6ad4ad9e03c1fb49ce630dbcf4e146b32a67ee2114754b

Observation f545e8c4-795e-48d1-bde4-d3a0b34a6ae5 · outbound

This paper cites Deep voice 2: Multi-speaker neural text-to-speech.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Deep voice 2: Multi-speaker neural text-to-speech

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.581352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.275156Z digest=sha256:6fb9ec060049cfd4d3d01417a6a2ec21fec9a8c9f83ce0f0bb8297dc32c6a681

Observation a463a38d-d7be-409b-9e47-46c3bc12fa0b · outbound

This paper cites Generative adversarial nets.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Generative adversarial nets

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.565415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.280725Z digest=sha256:9ab2cab5e912934b237b914bc00254a2f01085f14127886c7c045c7377cca948

Observation 55867ffc-d869-4696-b7f0-c23500615042 · outbound

This paper cites Griffin and Jae Lim.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Griffin and Jae Lim

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.542863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.287396Z digest=sha256:a35dfabc46a2a349cecc0bd334d6d2fbb180c466ef749fa04205f9e7c02d2172

Observation a9f517cf-adea-4265-83b5-cf604f0182e0 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:41.524456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.291930Z digest=sha256:0312e7e4efa6c72757f9d9448ba2b1cc38c6ccc9a54bf5d5900f1f7ab3a41132

Observation a901f9a3-bdd6-4f8a-9474-2796c9971ffa · outbound

This paper cites Deep neural networks for acoustic modeling in speech recognition.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Deep neural networks for acoustic modeling in speech recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.507327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.299041Z digest=sha256:586211b02ed4aba2c252a6f76fd653d7987468f3f629694907567ab9dde42995

Observation ec1e21c5-7697-4666-a19b-21c451dc9a46 · outbound

This paper cites Huang, Z.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Huang, Z

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.492668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.304045Z digest=sha256:20527061ae71019b102e6b094267a6179d02aa55f6581d67b929f98db38ee046

Observation 4f98a6bf-714b-4324-b9c3-e4446dd196aa · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:41.477805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.312903Z digest=sha256:1d867dd9261b7206541df8f5f7385ec1b481acbc743040ff7cf0d41dfcaf49aa

Observation 33c9fec7-9ce1-4584-99ab-329778a5c0b0 · outbound

This paper cites Recurrent fusion network for image captioning.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Recurrent fusion network for image captioning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.463953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.318633Z digest=sha256:d8a0947fe3771f8e432a9ae6bf3dcc6f34e8a42e645f4328c46666784915e53a

Observation 6522e939-ca0e-4114-b621-468754b8b2c3 · outbound

This paper cites Learning to discover cross-domain relations with generative adversarial networks.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Learning to discover cross-domain relations with generative adversarial networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.447931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.323490Z digest=sha256:3ccf2d36040535f79369d2435d583b01931f7025047cfc8a828c20ae5d556339

Observation 3dbaa0dc-3b94-4030-8570-69154d0a83a1 · outbound

This paper cites Adam: A method for stochastic optimization.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Adam: A method for stochastic optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.328415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.328415Z digest=sha256:10a81756b4f042671187c1b5dae2e06acc7d8dfc3c8f07d4de6e442d4be057af

Observation 39a0947e-dbaf-4c35-85a7-b63450f74f29 · outbound

This paper cites Auto-encoding varia- tional bayes.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Auto-encoding varia- tional bayes

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.416403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.334491Z digest=sha256:e8167796894e023f143fadb19140df22502d5005b20cc8a67019dddece245618

Observation d394991e-db9c-4945-9809-0f3aeb1ff483 · outbound

This paper cites Shamma, Michael S.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Shamma, Michael S

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.398799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.339109Z digest=sha256:323d96a9114c308713ba7721a0a4d95da9ef334f689e586d0d65d0d1503702fa

Observation 0df4bacb-40aa-4a6b-841f-f526c8ad1899 · outbound

This paper cites Letter-Based Speech Recognition with Gated ConvNets.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Letter-Based Speech Recognition with Gated ConvNets

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.343688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.343688Z digest=sha256:150660101b9f7820943dbede3c2fe52d7fc48b6b7bd29ac8869d78353d4ab79b

Observation 97f37b93-b26e-48fc-9836-b7d6f98cb6f2 · outbound

This paper cites DA-GAN: instance-level image.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck DA-GAN: instance-level image

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.378509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.349055Z digest=sha256:392723266ba3e523070e78b952d3f897bb152e48ea973f558523ad6fbf4d0705

Observation a0dc435e-e9af-4f45-a0f4-04a77d8939ee · outbound

This paper cites Neural TTS styl- ization with adversarial and collaborative games.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Neural TTS styl- ization with adversarial and collaborative games

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.363667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.354032Z digest=sha256:d90db1e34938c89bfcfe79bf4fb9f9dd78db22c395b7226e1984e43b6c665dd1

Observation 6c2909de-9943-465a-b9a9-84e0aa189c2b · outbound

This paper cites Semstyle: Learning to generate stylised image captions using unaligned text.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Semstyle: Learning to generate stylised image captions using unaligned text

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.347881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.359162Z digest=sha256:48ef83ff27f8365ce60c19d432125c1156d5fcb75f8694f34dc6cf74aeee7825

Observation e963026b-5cd3-49a3-9fd2-e22378dd0b3b · outbound

This paper cites Deep multi-scale video prediction beyond mean square error.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Deep multi-scale video prediction beyond mean square error

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.334089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.363608Z digest=sha256:f50b5184392eb2513db9eb4091a5da5fa7bd609e2155c239a6a19987c0f67b57

Observation 246b088b-86e9-476a-83d3-30e418bd7a41 · outbound

This paper cites Fitting new speakers based on a short untranscribed sample.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Fitting new speakers based on a short untranscribed sample

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.320167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.368936Z digest=sha256:ec0bb8213de79a7d0eb46e22c3df377564feba69efb615a9029cae136f48242a

Observation 183b6c45-042c-45ee-8f6d-f2a6a96a0ba2 · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Librispeech: An ASR corpus based on public domain audio books

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.304579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.373929Z digest=sha256:3749bbf2aa496233c1677e3434dcce60421e5ebfaaddc616a15b36054ac6e137

Observation 9cf086df-c52c-485a-9e60-b0c2df823191 · outbound

This paper cites Attend to you: Personalized image captioning with context sequence memory networks.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Attend to you: Personalized image captioning with context sequence memory networks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.289510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.379814Z digest=sha256:52a0901e8dda9503cc6ed55dc5870b40f336aef3aa85f9b2439ed5d9984ce24a

Observation f73743cb-664a-4c60-b154-e764c072554d · outbound

This paper cites Beyond sensory im- ages: Object-based representation in the human ventral path- way.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Beyond sensory im- ages: Object-based representation in the human ventral path- way

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.274190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.384718Z digest=sha256:7c354cbaffc93c0512c20d1dd06001b3ff28e29d7b9d8340ee17b213a61f39e8

Observation 2ebac211-0c85-4b2a-a2db-75e0951216f2 · outbound

This paper cites Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.258001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.393548Z digest=sha256:bdc52825619a760210ac565566b9c3d9bc65ab060f6dc3049bdebc4fe4b56d9c

Observation 81148e27-f988-4720-b030-3d5de3180dc6 · outbound

This paper cites Generative ad- versarial text to image synthesis.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Generative ad- versarial text to image synthesis

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.237503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.398847Z digest=sha256:77aaaa001f4111156ca00baa07225e24ef68de1009a4aafb96db64b3c1d7f0c7

Observation 1e437235-6436-4bd5-921f-496d3aa71b31 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.221378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.403456Z digest=sha256:23cfd779290f56411a417c9cf229f985a55d96587adee221d7f37e3c1db58f3e

Observation cc08e4ad-b405-4ca7-a066-d8804c3554de · outbound

This paper cites Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.409378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.409378Z digest=sha256:9e56386144216175f814d77a102a0190bccea1f7115358bdb292cf4999d67601

Observation 39f2086d-8a96-4a69-b83e-0abc87399f1f · outbound

This paper cites Char2wav: End-to-end speech synthesis.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Char2wav: End-to-end speech synthesis

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.202586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.415356Z digest=sha256:57d3003142bb02be64cba2bbb8d5216d943df9e418c1c440eb9d8f42bf33a7c1

Observation 15590278-f90c-4c7a-b371-9e1ad76a8ddb · outbound

This paper cites Szegedy, Wei Liu, Yangqing Jia, P.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Szegedy, Wei Liu, Yangqing Jia, P

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.188088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.420662Z digest=sha256:355e23fb0cf7914afa1c07a8932ef24cd22d26f51aad182b2b74afa2605866f4

Observation 7494df59-bec1-44c6-9b46-d2080ae75d06 · outbound

This paper cites Unsupervised cross-domain image generation.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unsupervised cross-domain image generation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.170542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.425757Z digest=sha256:0767706a3b9edb863df3e57259bbbf441c0bdfc26c7d3f85078bd245861d07ce

Observation b60e1574-6217-40e3-ac05-4b017bf2481b · outbound

This paper cites V oiceloop: V oice fitting and synthesis via a phonolog- ical loop.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck V oiceloop: V oice fitting and synthesis via a phonolog- ical loop

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.154400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.430623Z digest=sha256:a581dc9479c2880b1672e7d6c89921829ab5bffd2d5874e52b0d2de48fc5c815

Observation d91550cd-0fe8-416f-953a-1b62f9056112 · outbound

This paper cites The information bottleneck method.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck The information bottleneck method

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.139967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.436106Z digest=sha256:d7c06420836ef3897d5c9e999ff2cf13a98b90bd3418018f75ff335535c9f3a4

Observation b5488199-086e-4b76-b792-c0391b49f077 · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck WaveNet: A Generative Model for Raw Audio

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.440702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.440702Z digest=sha256:7a7c02d69bd8b2f00694ca02279ffaf48bf9f8ffd5be5b90f03fa1c729200533

Observation beaf1d4a-2160-4e1e-99a8-cc27cecad1aa · outbound

This paper cites Attention is all you need.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Attention is all you need

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.445811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.445811Z digest=sha256:0f45702689492b758d1ebb6dc9a0549651b71cabb96e92476ac7ae59ffef32ef

Observation c148cdd4-3077-4afe-9cab-f11e86832d9f · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.116749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.450090Z digest=sha256:bedef170eb246548954817674eff31aa57e87507034f8ac7063b6bfcaebde4e6

Observation 27f4cf11-4028-4c02-9d5b-7aadf127df87 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:41.098764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.455622Z digest=sha256:4d0ba8d619f9f7200c7646d9b86b38da6a0c38ff00e3d5001d891dacd98518c5

Observation 2d9e1973-0cce-46d7-b994-8c91366a3021 · outbound

This paper cites Show and tell: A neural image caption gen- erator.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Show and tell: A neural image caption gen- erator

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.079532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.460412Z digest=sha256:566c0028389d30f6646ba54c90d4d1bb010d4abec7138450e5831e2fd887e904

Observation 0689af5f-ddea-4cec-bea7-52183f9196c5 · outbound

This paper cites Tacotron: Towards end- to-end speech synthesis.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Tacotron: Towards end- to-end speech synthesis

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.058522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.465080Z digest=sha256:ea1901f3fec3abf6976f34d2765a768050114d4ba54e2a3eb8be64f327f5889f

Observation 7b0e0399-1fdc-4ee2-9582-09cbbd2aeea6 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:41.042059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.469470Z digest=sha256:ee28e14d44ec88bbc7a1cbc40bdc75e502978542c223583dfca9e8df2c5a7fe3

Observation f35ff7ab-c1c2-438b-8ad2-39c01ecbca5b · outbound

This paper cites Memory networks.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Memory networks

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:41.024937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.473965Z digest=sha256:7c3855ce6601190425f962c4bc98078b9ea09ca9b193443aac26f1281dd15ec7

Observation 68744106-0ce7-4a5a-9bdf-fece81af9708 · outbound

This paper cites Courville, Ruslan Salakhutdinov, Richard S.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Courville, Ruslan Salakhutdinov, Richard S

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.478934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.478934Z digest=sha256:012ef67eb55670f033ce1db8df89222ee4b721c6bdaa4a4fc74896cb5a105517

Observation aeedcc49-4e5b-4512-ade5-3389ee9adb75 · outbound

This paper cites Attngan: Fine- grained text to image generation with attentional generative adversarial networks.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Attngan: Fine- grained text to image generation with attentional generative adversarial networks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:40.989008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.482808Z digest=sha256:6ba1329f0ce59bf6bb735006192f6f057a3a9e607464042d4cb1c2db310fb06e

Observation 3e2215f9-f6e5-4012-870b-d75b57324378 · outbound

This paper cites Image captioning with semantic attention.CVPR, 2017.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Image captioning with semantic attention.CVPR, 2017

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:40.970995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.488513Z digest=sha256:6fc0536a5aa9d1769b4e32bac472c8e654067133c1fb2860eea4139f106c57ea

Observation a0e63788-9d66-41a1-a2ca-70f492b48c75 · outbound

This paper cites Stackgan: Text to photo-realistic image synthesis with stacked genera- tive adversarial networks.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Stackgan: Text to photo-realistic image synthesis with stacked genera- tive adversarial networks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:40.948778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.492506Z digest=sha256:5154f15b1678431f2d5dc608b21910df805c74758487430becd7c3cbd99b18ec

Observation 60d31258-403f-42cf-9b6f-f52dcfab64e9 · outbound

This paper cites Visual to sound: Generating natural sound for videos in the wild.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Visual to sound: Generating natural sound for videos in the wild

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.496609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.496609Z digest=sha256:c53fad3adf5e900c235965008c9766ac5330bc5ad7d9cf875a529783c12c18e7

Observation da4a2775-032b-41c5-93ea-9445644bcae2 · outbound

This paper cites Im- proving end-to-end speech recognition with policy learning.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Im- proving end-to-end speech recognition with policy learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:40.921821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.502211Z digest=sha256:e23234fab8302ebc51a818dba65a8ed9b7e9fbf3de88cd42faf61bf62f7247f9

Observation a3c5a992-6bd7-492b-b704-36bc6a53f641 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.899587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.507103Z digest=sha256:e833e8352ff119fafc9c371a25842ee2bbb644ac71856c4620135a784ae23fe6

Observation ec905cbd-4e7b-4139-94d4-c13339b689f3 · outbound

This paper cites Efros, Oliver Wang, and Eli Shechtman.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Efros, Oliver Wang, and Eli Shechtman

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:40.884162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.512601Z digest=sha256:128d12d11125c577de1581650e7bb872c46c738fd396c4580b19f9a84f692610

Observation 1a4ac184-6260-4cac-baf4-fb788c5abeac · outbound

This paper cites We encour- age the readers to refer to Figure 3 and Figure 4 of our main paper when reading this section.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck We encour- age the readers to refer to Figure 3 and Figure 4 of our main paper when reading this section

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:40.866425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.517413Z digest=sha256:b70a0543422b01e4b49ca7bd3585dbaf8cc040fab4c63702ca7dc8f880846fc6

Observation 5399a283-5229-4460-afb4-13fe1a6708d2 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.849508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.523658Z digest=sha256:750e49a3b9a339e1fe545b66fadb1dcd8dbabab86b12bc53eeb1e9028a86a7de

Observation 1cb707da-a64a-4932-b58d-57eddc55bec4 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.834024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.532088Z digest=sha256:21b6d7273bdc84a7b5e958f9b6be91e0a5529811c6fdee8758bd61904eb87f9a

Observation 43abc4ef-1f57-4c11-bafb-174ade716b24 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.816482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.537200Z digest=sha256:9fd1bb895f88174b088a67268b7cfe6c11f134018fbd7c1d84b6906224aca835

Observation 224ff9f3-0ebb-4843-8926-383cad96ea42 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.800184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.542135Z digest=sha256:92a1343bc54c3460f85111101a4f767114b7273c127337c0fb60e82ccf0e044c

Observation c04a12fe-06cb-46ad-b544-7e871c475c59 · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.784457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.547781Z digest=sha256:174d7f3fd43a69f8f4e50b2985a8a2f440b35c484d3683947d4c3f174c00597f

Observation c79de29b-7e42-4cd9-a4e7-5369cc87ce5a · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.766624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.553877Z digest=sha256:a94c06f7b3f32d8f00a379fc85da29471e3227e0c203b45cc62cfe8d076879a1

Observation e3f706e1-9dc3-4c1a-a3c8-f80baa5dbd8a · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.735273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.564751Z digest=sha256:1330cc3a52fdc3db5b0eebff27aba6230471bebcbd0e40e626d49c59e86eaae0

Observation e61b71e8-6b17-4bbf-81e9-5fab85ad1eb1 · outbound

This paper cites dog” and “zebra.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck dog” and “zebra

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:32:40.712826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.569987Z digest=sha256:21abf936dc9dd922566db46fa4c9d5ecb110ba77abe6cee3dd311a1f12c51f03

Observation 33f8c0bb-e942-47d4-b291-eabd291d05b6 · outbound

This paper cites bottleneck.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck bottleneck

Reference 70

Resolution
malformed identifier
raw_fallback, observed 2026-08-14T12:32:40.696611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.579862Z digest=sha256:e6f4d883d7cce36fd923fc8a9f1c8acdb136988a2346e40d45a0223b9cb9d225

Observation 7c2b0122-2aba-4a76-a0b5-7d5f60ae7a6a · outbound

This paper cites an unresolved cited work.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Unresolved cited work

Reference 256

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:32:40.751501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T12:32:40.559055Z digest=sha256:6c7314be794b04d33498f0e72a9913e4f15c5b833c111368961a365bfa6aba5e

Pith citing papers

No inbound Pith citation observations are available.