Pith. sign in

Paper Citation Record · LEDGER

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network

As of 17 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:1908.10072.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.10072 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T10:57:38.415005Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1873e6e7-db75-48f2-994e-6742a4b95208 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.651622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.041580Z digest=sha256:6e3c11a0a91f2f74eb972b275b9bceec1c2e7c8da96edb6eb0843c0be9379710

Observation 021b16e0-d146-4917-b487-dd4c5234eea9 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Quo vadis, action recognition? a new model and the kinetics dataset

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.629639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.047580Z digest=sha256:2361282584358cf0d616846b3830caef00be564238e017eab8f9807cfd6a2dff

Observation c7d2996c-835a-41d4-ba78-3a57282b3503 · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Collecting highly parallel data for paraphrase evaluation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.612282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.053330Z digest=sha256:8952131f14c6e645a382773128dc9559e2f332970761ec0b1405cf00dd2a2d91

Observation 9df306c4-4e7e-4c5c-a4db-86ac34445603 · outbound

This paper cites Video captioning with guidance of multimodal latent topics.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Video captioning with guidance of multimodal latent topics

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.594185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.059178Z digest=sha256:1c79a72f1cb79acaad150fed0a1106b1321c521bb217c122f89c062856b44c7e

Observation 41befd8a-49ea-4e0f-926d-ae5b944f8b9c · outbound

This paper cites Motion guided spatial attention for video captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Motion guided spatial attention for video captioning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.573543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.064496Z digest=sha256:667fa1d9ffad220bd3ad99f83c3c49bd3b0ba49029aa7df8ab76c0dcd7bd00e7

Observation 7de028ad-7c2e-4fde-a3a9-8c9e009071cb · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T10:57:38.069975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:57:38.069975Z digest=sha256:a10f9775102420779b42ad2d474dd876bee12e970361ad9b859eca834d4aeca2

Observation 481f45ee-4ef8-412c-8145-ec313b212b6f · outbound

This paper cites Regularizing rnns for caption generation by reconstructing the past with the present.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Regularizing rnns for caption generation by reconstructing the past with the present

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.554748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.076492Z digest=sha256:f7a011b62e669333dc36b7ea003baced8bd1a13deafb494ee9062bbb4d103205

Observation d75d3be0-b701-4797-88fd-00040d78ea1e · outbound

This paper cites Less is more: Picking informative frames for video captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Less is more: Picking informative frames for video captioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.538263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.081380Z digest=sha256:845fd8def058e203a38531cebc44e62f2437a3019a7c7327ecfa3b7ef7b60826

Observation 34ce0a48-94d5-47ae-9fe5-34e52be0ad7d · outbound

This paper cites Embodied question answering.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Embodied question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.521481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.086671Z digest=sha256:4aa0057e627377d0d1b1dba45cd6dfe2f8d11dee83c222ab9043c5b2dade85df

Observation cab863b8-345a-49e2-95ac-b02654e6a053 · outbound

This paper cites Diverse and controllable image captioning with part-of-speech guidance.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Diverse and controllable image captioning with part-of-speech guidance

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.504035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.091702Z digest=sha256:622578fe96b4b4309b9e9cd3ba3d6bf52a6f296a179ca555cbce7ae25e66f4d1

Observation 2c82ff3d-679d-4f8d-ba5a-e0381c22deb6 · outbound

This paper cites Long-term recurrent convolutional networks for visual recognition and description.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Long-term recurrent convolutional networks for visual recognition and description

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.487390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.099165Z digest=sha256:75bf5df4f1fbd31ebd6e9955c7a35f5b1526d7f56bf5b2caf637aad983cdb872

Observation f8fd452f-6470-4535-bc1a-c6e806eb8b79 · outbound

This paper cites Unsupervised image captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Unsupervised image captioning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.468755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.104751Z digest=sha256:c6b4dca931eb74529ecd835c8fa1f38c3b580ba40736775fadf44098047fb3cf

Observation 5091e6f7-4a65-40e0-ab93-8ad3849c18b5 · outbound

This paper cites Multimodal compact bilinear pooling for visual question answering and visual grounding.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Multimodal compact bilinear pooling for visual question answering and visual grounding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.450592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.109755Z digest=sha256:385fec93108c0707dcbea748d8af41966e81827ee7bfe065c2be203e54bb7385

Observation 4df725c0-297e-43d6-b786-a5bca031cd73 · outbound

This paper cites Semantic compositional networks for visual captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Semantic compositional networks for visual captioning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.433743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.115264Z digest=sha256:1bd99221f94a21dc55a38abaeb18078bc9e2fd943ac15ce7226b2390ecf81c09

Observation 5f1b3ffe-4ac0-4288-9a15-f6affc13d424 · outbound

This paper cites Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.414528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.121213Z digest=sha256:fd4a1e27d697e68bd0dbd9ed14dbee740bf8d80dbe6b8406bde38ce54fa402bb

Observation 3a93c80b-12b3-45ec-86a7-4ca44ba535a6 · outbound

This paper cites Image caption generation with part of speech guidance.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Image caption generation with part of speech guidance

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.392423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.126470Z digest=sha256:8a24f45d0446ef4eff371e1d34ec84c6c17faf8196e36d6ac8aacbbcdc7fcbfa

Observation d137cd4b-5e2f-46a4-b370-e29aa33a6ecd · outbound

This paper cites Recurrent fusion network for image captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Recurrent fusion network for image captioning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.373534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.131519Z digest=sha256:62e56c194ed134bcf788b9d5084e2e9d54902bd1599c4dab70866edc78353335

Observation 37133a6b-b580-46ff-9cf8-4f59c9d78d39 · outbound

This paper cites Describing videos using multi-modal fusion.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Describing videos using multi-modal fusion

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.354958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.137521Z digest=sha256:8948e924c176f100eff7e1e774147406d55556befa9061c1832ee86b85954c3b

Observation b11ab0e7-a50b-4e99-af94-a08eca01a820 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network The Kinetics Human Action Video Dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T10:57:38.143226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:57:38.143226Z digest=sha256:984b2bf73459e726a9e36fa181cbe4d5b6fd170fb5214d2a5cfea72517397746

Observation c7fcec0e-8678-4bc7-ae8b-d7fee8afe461 · outbound

This paper cites Hadamard product for low-rank bilinear pooling.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Hadamard product for low-rank bilinear pooling

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.337697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.151471Z digest=sha256:b30076a31d687d57cbab8927defdb1bf41e4b8f5d998fae39d1b375d791eadfb

Observation fd5448d2-06ec-4839-94be-8a646803e163 · outbound

This paper cites Natural language description of human activities from video images based on concept hierarchy of actions.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Natural language description of human activities from video images based on concept hierarchy of actions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.319316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.157564Z digest=sha256:3bc2c8e444d8069a7f526f8a8461e62147b89cd17bd03695c82901d020426e5e

Observation 5baa7aad-7032-48e8-aa80-26f5f6fc6183 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Rouge: A package for automatic evaluation of summaries

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T10:57:38.168749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:57:38.168749Z digest=sha256:f33f6c628e8f4e4c5edb9fbda9b444fb4445f7464ee1fb33317b102481c5de44

Observation d6ffe39b-b72a-4915-8561-e5da7eb56eaf · outbound

This paper cites Sibnet: Sibling convolutional encoder for video captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Sibnet: Sibling convolutional encoder for video captioning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.279723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.176170Z digest=sha256:7f3e0a72ebb9df6dddaba7cae5685c8d585ca95c8fdfb9816b4eee4e25745ff6

Observation 1bc14e2a-7463-42a6-a223-3813c20f5cce · outbound

This paper cites Matching image and sentence with multi-faceted representations.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Matching image and sentence with multi-faceted representations

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.260563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.181550Z digest=sha256:1a846fbb8c68b78dbdcbb72e601d9f32215b29584d75c3bf8c98f1e50d83ec64

Observation 700d4d2c-c917-4829-bd71-ea8df1bb4df0 · outbound

This paper cites Learning to answer questions from image using convolutional neural network.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Learning to answer questions from image using convolutional neural network

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.244241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.186885Z digest=sha256:9c53cf24b820d0d542758e0fdc207409d9be1aa74df3feafec0962e0786956ac

Observation 9a572c67-6bea-4cd4-be14-2846aa3b6a65 · outbound

This paper cites Multimodal convolutional neural networks for matching image and sentence.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Multimodal convolutional neural networks for matching image and sentence

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.227761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.192462Z digest=sha256:b38f4d3cceb3075e4c0e4a1c4eaa2dd0873f74cfda233cb18a00ee11b447b766

Observation c9dae1cb-c071-497e-8df2-ade0bfbab609 · outbound

This paper cites Jointly modeling embedding and translation to bridge video and language.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Jointly modeling embedding and translation to bridge video and language

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.211142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.198856Z digest=sha256:48e793bd9a795957d65c3ff9d041d125c47327e7242e4ede4cf1e3330d5702d2

Observation c5c7377c-0eea-44d0-9384-882e23422678 · outbound

This paper cites Video captioning with transferred semantic attributes.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Video captioning with transferred semantic attributes

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.192777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.204428Z digest=sha256:06eb70cb2efa4c9147c27f4b6d0d92d317ad16ca8d37644988d530a6247995cf

Observation 17f6b06a-f575-46f8-9748-75a7407ac00d · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Bleu: a method for automatic evaluation of machine translation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.176958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.209681Z digest=sha256:fed13bb3a74a8e470dc55a8bb1c2c6d39eef99ecf58efe136d68f268c010b412

Observation d16035c1-df84-4eb6-9130-bbee5849f4fc · outbound

This paper cites Tv-l1 optical flow estimation.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Tv-l1 optical flow estimation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.159957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.215643Z digest=sha256:dbf5b48e396bc7739f94b17b4e97eace68957df9bff3e70e3310ecc123244603

Observation 1f5c2d93-c0ba-4599-8f61-dec9b559090e · outbound

This paper cites Multimodal video description.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Multimodal video description

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.143375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.220829Z digest=sha256:d8448f1cb42fd0cd0325d637775ee553f3e3f644e8c6c8c0106e8183a6341b7d

Observation 125438e3-74ba-4c9c-a42c-ed07a30c43a5 · outbound

This paper cites Self-critical sequence training for image captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Self-critical sequence training for image captioning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.126286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.226951Z digest=sha256:9e7d71e06460b49825d8ac93d4d26e08c5bfa9ede703282dd4928547942b13ca

Observation 0a6eba4c-8f05-4e06-99ed-98f2b8e01e39 · outbound

This paper cites Coherent multi-sentence video description with variable level of detail.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Coherent multi-sentence video description with variable level of detail

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.108828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.232533Z digest=sha256:6829494db32c3717066f64e28990cae3ff0948bd9024ad2493cfc36c63900a71

Observation 25f0ed17-dd4c-440b-b6ac-b34f75bacfd9 · outbound

This paper cites Translating video content to natural language descriptions.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Translating video content to natural language descriptions

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.093009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.238054Z digest=sha256:4a40b7a00b7b1c634d6422d24d6ed3e43cb8550bfd64ee6519f9f8df810f6fb1

Observation 1bb049a4-948f-4de3-8b78-667998310f40 · outbound

This paper cites Imagenet large scale visual recognition challenge.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Imagenet large scale visual recognition challenge

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.075893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.246095Z digest=sha256:f78ea62861167437e7d4068afc7394b438de083a2dc9af7d5e0ef562409a4034

Observation 5d73ab8b-b969-4a39-a3b5-52386a26d2d1 · outbound

This paper cites Frame-and segment-level features and candidate pool evaluation for video caption generation.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Frame-and segment-level features and candidate pool evaluation for video caption generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.058047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.252404Z digest=sha256:c0212a93c5574df1b905ce7f8f79411220c5e3f6643f207e61b905320ffd5fd3

Observation c7427a89-25d2-4d56-ac26-5c1b8981e5b1 · outbound

This paper cites Quantization-based hashing: a general framework for scalable image and video retrieval.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Quantization-based hashing: a general framework for scalable image and video retrieval

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.039809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.258742Z digest=sha256:be06c5722aa23163a9e30e2332b214672017735e9f6b738400cecb46413152ef

Observation 328ed9a7-2499-4e16-8773-4450a66adadf · outbound

This paper cites Inception-v4, inception-resnet and the impact of residual connections on learning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Inception-v4, inception-resnet and the impact of residual connections on learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.022436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.265410Z digest=sha256:71df31b172d179487ebd7844f2f913c6fe3f8ef867c2b1080b4e282f079b017f

Observation a8ef05c7-5dff-4789-9594-33bf05158b2c · outbound

This paper cites Feature-rich part-of-speech tagging with a cyclic dependency network.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Feature-rich part-of-speech tagging with a cyclic dependency network

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:39.003880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.272559Z digest=sha256:d12e6b6e84365a1991bdf098f67c83504a0a5c54010f78969f17ebf63427a7f6

Observation f73729ed-960a-4aea-9987-e82694235fea · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Learning spatiotemporal features with 3d convolutional networks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.981064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.279403Z digest=sha256:f166c8ea58c373c40451ce95bc40da28dcd52fd17abb63921e652459074a9a83

Observation 9e318745-2eb7-46a5-bc12-d5c4a9cc3744 · outbound

This paper cites Cider: Consensus-based image description evaluation.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Cider: Consensus-based image description evaluation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.957250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.284378Z digest=sha256:be065fe84ab0332958a5d04da346a09f77a2589f6a9a5c8ab50cb501a888540c

Observation c4fb9dbb-b469-4a6e-a081-6a7517c5fc11 · outbound

This paper cites Sequence to sequence-video to text.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Sequence to sequence-video to text

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.940058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.290185Z digest=sha256:7274a8257187f7b9e4851c3413771b5d0e6792a9ab367d3aadad19da09c00e63

Observation e5e326db-5329-435d-8d2e-2e5f3a6da194 · outbound

This paper cites Translating videos to natural language using deep recurrent neural networks.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Translating videos to natural language using deep recurrent neural networks

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.921555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.297238Z digest=sha256:6c6bb1d39724544c7e8f1374715ff2f03f6fd7c512ae4bb04e588269c26f9a20

Observation 408adeae-2df5-4cde-994f-197dd7e1f8e2 · outbound

This paper cites Hierarchical photo-scene encoder for album storytelling.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Hierarchical photo-scene encoder for album storytelling

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.898224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.303765Z digest=sha256:3ac5daf8023f7af7e6e1eab67a78dc4d39685e426187e53116a0a994789050bd

Observation c28207be-6357-408c-9aee-dd507a4f322e · outbound

This paper cites Reconstruction network for video captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Reconstruction network for video captioning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.879932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.309023Z digest=sha256:41e3dcff4acd90a8639091150b4c2bc0a925ba6de9e4475aaf6251eb84c62d84

Observation 0da1353e-262b-4929-8318-89f3cf4832e4 · outbound

This paper cites Bidirectional attentive fusion with context gating for dense video captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Bidirectional attentive fusion with context gating for dense video captioning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.861169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.314056Z digest=sha256:dd4e4419a79a0ee99eb7823dc4f279a46ef410481d9458fb25245ff9ec64f927

Observation fb9e6b81-abac-4d4c-b220-38e3d1fa02fa · outbound

This paper cites M3: Multimodal memory modelling for video captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network M3: Multimodal memory modelling for video captioning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.843411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.319099Z digest=sha256:599ec8a5fe15560651b9bdcdcfaf96b6235b6205058d760b027299440c37a6fa

Observation f54d4121-39f4-4b6a-8d35-7ff8fae36104 · outbound

This paper cites A survey on learning to hash.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network A survey on learning to hash

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.825218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.324097Z digest=sha256:5ed416934b79d1ca99a5e09ab6207fc5360f064bfc739fc46a121dff8ddabd09

Observation 8a66c135-e4c8-4d55-b55b-1cd98878b26f · outbound

This paper cites Interpretable video captioning via trajectory structured localization.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Interpretable video captioning via trajectory structured localization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.807467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.329406Z digest=sha256:bf595a25a15266e03753d50a9ba7be21c8a9c28c14d5771edf1eb7a5e202e248

Observation 28b260c5-3bdc-4cb1-abaa-56167eece4f1 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Msr-vtt: A large video description dataset for bridging video and language

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.788354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.334432Z digest=sha256:a43ddc5382fb085254fa900d9dba586b06312d363b5307c959d7c7d8d0a0c527

Observation 50162c04-e9e6-4aa3-85f5-32c46c1ffb3d · outbound

This paper cites Learning multimodal attention lstm networks for video captioning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Learning multimodal attention lstm networks for video captioning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.765691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.339586Z digest=sha256:d4823cb0e20d903376f92c565662fcc9163d1442a3ab59b3a1b287dbfd119585

Observation 3f0aaf19-3837-458e-9e34-87f77d321aa7 · outbound

This paper cites Jointly modeling deep video and compositional text to bridge vision and language in a unified framework.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Jointly modeling deep video and compositional text to bridge vision and language in a unified framework

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.746357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.345161Z digest=sha256:1ede98126cd8e0753d7eb05cd0cf8ae3f8943f5454dec9852a839a34357988f5

Observation 658ae573-f81c-4559-bda4-e7e81f1d461b · outbound

This paper cites Describing videos by exploiting temporal structure.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Describing videos by exploiting temporal structure

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.724518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.354088Z digest=sha256:b3685441be8187b72ddba1993bb888039f1eaa7b4e1cba88e641fcfc1d9f76e3

Observation a010624a-10c4-45f7-b188-3f0f6482ef5b · outbound

This paper cites Video paragraph captioning using hierarchical recurrent neural networks.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Video paragraph captioning using hierarchical recurrent neural networks

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.703201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.360033Z digest=sha256:187760ce3d5ad2c3d642e880c75a525fede4b24bac6561acbb886317f0a1ca52

Observation efd89af8-df75-4e7f-b627-91de1c30eefc · outbound

This paper cites ADADELTA: An Adaptive Learning Rate Method.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network ADADELTA: An Adaptive Learning Rate Method

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-14T10:57:38.365230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:57:38.365230Z digest=sha256:97a7563ec9440b8d28d682e2b015991fcb87d602060007ac35f7d7c53db18a35

Observation 4296d482-517e-4e48-964d-212286737a98 · outbound

This paper cites Reconstruct and represent video contents for captioning via reinforcement learning.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Reconstruct and represent video contents for captioning via reinforcement learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.680241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.371492Z digest=sha256:7ea7ca87c48af31d405afd236f3c450b26b4dee2e171f64e2e99634ead4a5386

Observation cab27f6e-071f-41d5-9e55-f7c21fd2b2b2 · outbound

This paper cites an unresolved cited work.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:57:38.656895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.376450Z digest=sha256:d0f4a81a1913eb8fb51e7eec6f2fb10eeb01aee9609953dbce513289a271ab0c

Observation 9575ebe7-20b6-4b41-90cd-6d53b600021d · outbound

This paper cites Reinforced Video Captioning with Entailment Rewards.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Reinforced Video Captioning with Entailment Rewards

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-14T10:57:38.381673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:57:38.381673Z digest=sha256:217b4952a05c2cac2de7be11f6b7723d544eb0959f09207ec7af90d8b7475c0f

Observation 18fdd141-200f-401e-9a10-76be6874bb12 · outbound

This paper cites an unresolved cited work.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:57:38.636429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.387488Z digest=sha256:4171bdfca6ee20b423e058c8fda6c20e1c3cb7025b3760a407ec1566456b9f1d

Observation 71e4be24-20f1-4bf1-aaba-8833c9e222e4 · outbound

This paper cites an unresolved cited work.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:57:38.618714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.393036Z digest=sha256:fa4d8ceb4a261d269eb5486e52001c5f845e9188b32c7d6279db95a1ba155448

Observation 5d62de1a-c435-4887-85d2-63b8ad041f82 · outbound

This paper cites an unresolved cited work.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:57:38.600523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.398564Z digest=sha256:331a4e98e6bf541b8b8d09b6c0623d973011000a16511921da2a190066d1c2a3

Observation c8db4f33-38ba-4b64-afa1-e2ab34ba190f · outbound

This paper cites Vedantam, C.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Vedantam, C

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T10:57:38.582150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.404074Z digest=sha256:81476a784c27a7b9dcb46190ca602e8fd7f853963ab3835db8a27e2cd0e77685

Observation 5a5ab054-33a5-47b3-a10b-26015aab2806 · outbound

This paper cites an unresolved cited work.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-14T10:57:38.561785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-14T10:57:38.409600Z digest=sha256:e8ea267611abf0992bfede32534e9c3403e35fc630c08f1fa9e8047d6bb95609

Observation 37f52123-9752-4d5d-9975-2604d131c8bb · outbound

This paper cites Reinforcement Learning Neural Turing Machines - Revised.

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network Reinforcement Learning Neural Turing Machines - Revised

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-14T10:57:38.415005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:57:38.415005Z digest=sha256:9d0084c4b01206a8026b1bb90a82c650ced75fd6e55670fa5e2e833f7e8a3a6a

Pith citing papers

No inbound Pith citation observations are available.