Pith. sign in

Paper Citation Record · LEDGER

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning

As of 22 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 2 inbound Pith citation observations for arXiv:1908.08529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.08529 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T11:42:45.434477Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:46:41.404184Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T19:25:35.415589Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 600ac2fd-9328-479b-bec9-bd6b0d5fd3ed · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Bottom-up and top-down attention for image captioning and visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.866305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.298423Z digest=sha256:4f8b146cfb726f982372be150e8ed49bdd5b4f7ce6cb257bbb991f127cd693ca

Observation 0ad8e2ad-3a55-44c7-b033-575038a10565 · outbound

This paper cites an unresolved cited work.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:42:45.856501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.302638Z digest=sha256:877b3f69b0178ab8ca92b06a8ba653cbf157e29c71042242b14080de8418772f

Observation 70c4358a-c372-47a0-8a9e-28ea019f5223 · outbound

This paper cites Blei, and Michael I.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Blei, and Michael I

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.846996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.306085Z digest=sha256:0e7d30edea1b8c739c77e6ab8c6b4c92b5e7352f245aad211f88587f9268a089

Observation 704f534c-a71a-4169-897a-cefa146331b5 · outbound

This paper cites an unresolved cited work.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:42:45.837080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.309709Z digest=sha256:070781f760f95d72b5a75d7c5e01aab8f85950070d0e70d67c43942b7b6cff28

Observation 4a0c29a6-a3c4-4087-adee-03385d10adcb · outbound

This paper cites One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T11:42:45.313161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:42:45.313161Z digest=sha256:36b6f871194537cb461846b3c67fb8f6451679cfe09b3b4cfbfe47fcbf44e586

Observation d61e6592-34db-4a61-a9b0-3495c4ca64ce · outbound

This paper cites Mind’s eye: A recur- rent visual representation for image caption generation.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Mind’s eye: A recur- rent visual representation for image caption generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.827491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.317033Z digest=sha256:04a70930b57afe22e85e6f6dde60a0484e943ecbbe5e26b5243136969ee816f9

Observation 2246525e-459b-4cae-b88b-ef0159ef17eb · outbound

This paper cites De- scribing multimedia content using attention-based encoder- decoder networks.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning De- scribing multimedia content using attention-based encoder- decoder networks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.817531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.320524Z digest=sha256:ab4904e1b1bc45c64b2af331fb5075dd34cc9f2e20cad7d019f18aa4b96878bc

Observation b4f6355a-7844-4741-a3d2-b227c64d4941 · outbound

This paper cites A recurrent la- tent variable model for sequential data.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning A recurrent la- tent variable model for sequential data

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.807537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.323872Z digest=sha256:e58185f93fe9b8e9d97fdded41f46a32adba45288a3cdcfebacf7b68e8a158d3

Observation 352c2f3d-4396-4aac-a944-2564741392dc · outbound

This paper cites To- wards diverse and natural image descriptions via a condi- tional gan.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning To- wards diverse and natural image descriptions via a condi- tional gan

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.797929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.327378Z digest=sha256:765faf8ad0462951df2f0473ae4d6e45860ee282c37cf5c043d42eb201e0aaba

Observation 5afc8ada-e634-471b-a51d-8f7d43781268 · outbound

This paper cites Schwing, and David Forsyth.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Schwing, and David Forsyth

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.788332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.330461Z digest=sha256:c6437d063a4a0d097f510b36f6f180d93926ed1b2549491f7ba33bcccbfa7c00

Observation d9039ba8-40fd-4808-9adf-520e7e4ba2ff · outbound

This paper cites Exploring Nearest Neighbor Approaches for Image Captioning.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Exploring Nearest Neighbor Approaches for Image Captioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T11:42:45.333632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:42:45.333632Z digest=sha256:58338c921f02b78e16e8526157140b15195c054d697cff2801be8f814443b562

Observation 4e2beee5-4c90-44a8-9036-bef619024f81 · outbound

This paper cites Long-term recurrent convolutional net- works for visual recognition and description.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Long-term recurrent convolutional net- works for visual recognition and description

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.778463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.337518Z digest=sha256:da5838449c7724eaaf295aeb5722936935fc44864a816973f26b109e5f94e737

Observation 4761cb56-8639-4782-8b4d-282539e9fef0 · outbound

This paper cites From captions to vi- sual concepts and back.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning From captions to vi- sual concepts and back

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.768594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.340695Z digest=sha256:67f4530ffb09b1253cf373eefb0d713d15741805ad93253024eeb21ad0c82f55

Observation 3768b671-a6ad-44ed-8e80-43617fa3a930 · outbound

This paper cites Every picture tells a story: Generating sentences from images.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Every picture tells a story: Generating sentences from images

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.758433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.343685Z digest=sha256:635970191b7a5f5dbaa2984b3f987e60ada8b05a5b5cf3041dfb3586146ba81c

Observation bc86d92f-13ff-48e9-ac3a-6f85ad0d0fb2 · outbound

This paper cites Sequential neural models with stochastic lay- ers.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Sequential neural models with stochastic lay- ers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.748850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.346749Z digest=sha256:b5458c84e95673166dc41c00c7a8ac6a2a962946e49070f0451efa019e30fb65

Observation 57fe8c85-8398-4484-9e77-5f51c75d4c87 · outbound

This paper cites Generative Adversarial Networks.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Generative Adversarial Networks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.739150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.350016Z digest=sha256:e59a994ba1f720cbcefd42357961d2946c643a87473962bd4413424b09003e5b

Observation 40a71266-e5e1-4a98-9171-6fbf80f0824a · outbound

This paper cites Z-forcing: Training stochastic recurrent networks.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Z-forcing: Training stochastic recurrent networks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.729356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.352974Z digest=sha256:7749d729d29d2ade4c8dda684073708bd03e44e611d202802c6143dcf9374724

Observation fd76bcec-3593-4985-a043-1aadc5714142 · outbound

This paper cites Long short-term memory.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Long short-term memory

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.719499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.355909Z digest=sha256:649d8caac866e11d81d779eda720422ca5e9e41cc75a7038b0f199de7e69af2f

Observation c64d429c-2141-4544-866b-0f24b7178907 · outbound

This paper cites an unresolved cited work.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:42:45.709913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.358852Z digest=sha256:c235504e71451061a64245f11d7d8105fd00d77da9973f720cc6322ebc557faa

Observation 7b0a9ca3-ff94-4fa4-bc10-e15c7f1b4892 · outbound

This paper cites an unresolved cited work.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:42:45.699264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.361867Z digest=sha256:ce1c2e616120c29e5f17258bb380c0af6d0f0646128c096b489d005fb481d83f

Observation 208f81c8-242e-4491-9f13-df0e2dfce9ac · outbound

This paper cites Densecap: Fully convolutional localization networks for dense caption- ing.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Densecap: Fully convolutional localization networks for dense caption- ing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.688616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.365006Z digest=sha256:80ba309afd6b6d0b36f0ced0a6d98329f70ef9a76c0b72560ac695a08c60f9a7

Observation 73d04d50-909b-413f-b31c-09bf6e17a5e3 · outbound

This paper cites Deep visual-semantic align- ments for generating image descriptions.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Deep visual-semantic align- ments for generating image descriptions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.678274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.369226Z digest=sha256:8cc4e6f4530e9af57bc1836bffb32ed394c795179028dfe7c23797677bacdfcd

Observation e1e2bd66-e51e-47f9-a294-90ffcab1ba8d · outbound

This paper cites Semi-supervised learning with deep gen- erative models.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Semi-supervised learning with deep gen- erative models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.668560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.372488Z digest=sha256:997786d7dd4a4dc3be7926e6596545e089ba91cdad97818f38cf1c6dde726ad7

Observation 4ee4aba7-4ecd-417b-8683-788a00074d8f · outbound

This paper cites Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T11:42:45.375598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:42:45.375598Z digest=sha256:833c39842f838e3d5099ffae5ee8024c44f83a5bef36e4dea0d80c5e32e32c33

Observation df5142b0-c528-4986-84b6-f42a6e883515 · outbound

This paper cites Babytalk: Understanding and generating simple image descriptions.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Babytalk: Understanding and generating simple image descriptions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.659261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.379146Z digest=sha256:f325bd24eabc1cf83899edd47dc77260940c8ece5a355fefaf9aeb5ade6923f3

Observation 4c43c69b-3cbb-42b5-a172-e45c2c8639f4 · outbound

This paper cites Generating Diverse and Accurate Visual Captions by Comparative Adversarial Learning.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Generating Diverse and Accurate Visual Captions by Comparative Adversarial Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T11:42:45.382255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:42:45.382255Z digest=sha256:90a1dc165beebf95ff4c4edc845b0a7caef118ad41d41f6acba46e6b2d2547c0

Observation a9e20b80-8a94-41fb-b1f9-eb8e1accf9d7 · outbound

This paper cites Microsoft coco: Common objects in context.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Microsoft coco: Common objects in context

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.649498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.385706Z digest=sha256:ab9b5844a8fa51729ca668dcad6634f56f361e0daef650041682a1ac21f7a174

Observation c0333e07-efa2-4da9-91ad-fd96bad3056b · outbound

This paper cites Improved Image Captioning via Policy Gradient optimization of SPIDEr.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Improved Image Captioning via Policy Gradient optimization of SPIDEr

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T11:42:45.389128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:42:45.389128Z digest=sha256:930f1107e1f5468750f1ab6b5913a549aa63300066bad67f95eca329c777a465

Observation 47a82e14-cb06-49f6-9be2-9c9bedde142f · outbound

This paper cites Visualizing data using t-sne.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Visualizing data using t-sne

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.640143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.392597Z digest=sha256:fdd902d1c3fa7bc00158b29144ff00534fd15d652cfbd933746aa116e365d363

Observation 8cc3889c-38b9-425d-8b3f-bc652e4a40dc · outbound

This paper cites Deep Captioning with Multimodal Recur- rent Neural Networks (m-rnn).

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Deep Captioning with Multimodal Recur- rent Neural Networks (m-rnn)

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.630316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.395766Z digest=sha256:37c7a4daac9920f483d8ec268559dd49a9e804878e7f0ce8cfaeacad9d4f5592

Observation b96efe42-5d23-4076-ae91-e140963135b3 · outbound

This paper cites Peters, Mark Neumann, Mohit Iyyer, Matt Gard- ner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Peters, Mark Neumann, Mohit Iyyer, Matt Gard- ner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.620604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.398903Z digest=sha256:c864424c5ec201d87be28d72907ffa3fe9db1505e1c5f3118062d71990df153a

Observation 2d64583b-0804-43c5-bb9c-2660676d9322 · outbound

This paper cites Faster R-CNN: Towards real-time object detection with re- gion proposal networks.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Faster R-CNN: Towards real-time object detection with re- gion proposal networks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.610463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.402194Z digest=sha256:12e30affbcd7807f28344514f4bdb96b5713e815ed0e3f33c285d80346d763a8

Observation 48095380-5270-489c-83d5-7f1d2bf4c63e · outbound

This paper cites Rennie, Etienne Marcheret, Youssef Mroueh, Jarret Ross, and Vaibhava Goel.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Rennie, Etienne Marcheret, Youssef Mroueh, Jarret Ross, and Vaibhava Goel

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.600438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.405390Z digest=sha256:6975f5042015fa1462d23083969be5fabea2dc80aa9eb831342a29051ce787d8

Observation 95cfd86f-7ddc-41da-9446-d14dc90bf09a · outbound

This paper cites Speaking the same language: Matching machine to human captions by adversarial training.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Speaking the same language: Matching machine to human captions by adversarial training

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.590795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.408625Z digest=sha256:6f0fce207cfe6394aff04a7e633c350378a530e6173ada831c136605e80ff5ba

Observation 9b07a541-d1f1-413a-8a2e-9cc62ef1aaa8 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-14T11:42:45.411699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T11:42:45.411699Z digest=sha256:c50a2b453aac5712074f8aa30623f17abb3dc7cad6905ba1b1150255cd44102b

Observation 1b303fbd-b57d-4faa-9d62-e1dc3ac18f15 · outbound

This paper cites Grounded compositional se- mantics for finding and describing images with sentences.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Grounded compositional se- mantics for finding and describing images with sentences

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.581217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.415520Z digest=sha256:4626904db2cd6cc60899373528c374711ae309b79f3efc5a3136ed53e539bfdd

Observation c28f4082-3ab6-40b7-ba27-fc0a244b6db9 · outbound

This paper cites Learning structured output representation using deep conditional gen- erative models.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Learning structured output representation using deep conditional gen- erative models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.570755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.418677Z digest=sha256:e39cc57e667e6f7b2b7cc765d56b8eae08e9304faab85d1782878f2b72a14c92

Observation 7b1ae024-3194-4b9f-aced-105cfa49217c · outbound

This paper cites Vijayakumar, Michael Cogswell, Ramprasaath R.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Vijayakumar, Michael Cogswell, Ramprasaath R

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.559475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.421777Z digest=sha256:50e9d9003e7b63d10e79ddd3deb4509786d9646ed23e972f19aa2a561a4e5019

Observation 26be94f5-f1d4-42dc-a90b-8f950b17142e · outbound

This paper cites Show and tell: A neural image caption gen- erator.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Show and tell: A neural image caption gen- erator

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.549246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.424959Z digest=sha256:fabcac2879330353e4d2426d7935690959e613ced52d02d542351129b673c214

Observation ad615eb0-5e96-4094-b6ea-c2ab1c973442 · outbound

This paper cites Schwing, and Svetlana Lazebnik.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Schwing, and Svetlana Lazebnik

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.538789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.428161Z digest=sha256:d86eecbba48d789fb9a0a3ab75929e8c17193d5b75f971d37ac308a717e700ec

Observation 0df11933-7117-4beb-bf17-095f49a22be9 · outbound

This paper cites Show, attend and tell: Neural image caption gen- eration with visual attention.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Show, attend and tell: Neural image caption gen- eration with visual attention

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.529150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.431465Z digest=sha256:75b8af81671e4598a7b76bf248182b4ecf64fde7389b5519b25ac8a19fad1372

Observation 9dea1fe9-f5c4-426d-b8c7-69f932ab3682 · outbound

This paper cites Boosting image captioning with attributes.

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning Boosting image captioning with attributes

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:42:45.518544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T11:42:45.434477Z digest=sha256:8728cc6d1616f7af9d74afc1155a4f432223f30a6b1b4281952c95d589a5254a

Pith citing papers

Observation f4ac3de0-ec1d-4fd2-9dd6-d203e4c0f0fa · inbound

Image Embedding Sampling Method for Diverse Captioning cites this paper.

Image Embedding Sampling Method for Diverse Captioning Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T19:25:35.475849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T19:25:33.534230Z digest=sha256:fef16c2797bc24b9a243d4a9b1fae1390aac658fd55921a12225caa68854c50c

Observation a7168d6e-c615-4cb9-ab9e-ecd1c3c81e91 · inbound

Variational Prefix Tuning for Diverse and Accurate Code Summarization Using Pre-trained Language Models cites this paper.

Variational Prefix Tuning for Diverse and Accurate Code Summarization Using Pre-trained Language Models Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:46:41.404184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:46:41.404184Z digest=sha256:e80a27111f2029ef7a1a5b56b0ea6cf9089a0ab9a76094b7082fd35dbf303972