Pith. sign in

Paper Citation Record · LEDGER

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization

As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2506.08649.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08649 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:11:46.728553Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy43
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d549bf15-a081-4f71-99d9-3c23b26c444d · outbound

This paper cites Beit: Bert pre-training of image transformers, in: International Conference on Learning Representations, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Beit: Bert pre-training of image transformers, in: International Conference on Learning Representations, pp

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.278240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.549152Z digest=sha256:94cb421893969b2b00e8dfd296fe146582bfa5e6d1e749f88f4d01530d02c71c

Observation d4533415-1463-4bb9-ade0-ebc2debd5557 · outbound

This paper cites Emerging properties in self-supervised vision transformers, in: Proceedings of the IEEE/CVF international confer- ence on computer vision, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Emerging properties in self-supervised vision transformers, in: Proceedings of the IEEE/CVF international confer- ence on computer vision, pp

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.269193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.553980Z digest=sha256:0696da9646d9ef3603cdda15faddbe163c467b21eb2813ae534a8b0af000cfde

Observation 5f6d40fb-6dbc-4058-ade9-bc74f41d3314 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Quo vadis, action recognition? a new model and the kinetics dataset, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.258773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.557687Z digest=sha256:eef74fa820f9610e61c3952ae0db2218a22a7c7dad39d4f3ff403d141fdaa21f

Observation 781eb04d-9054-47f5-b515-80bf6f8f6ffd · outbound

This paper cites An empirical study of training self-supervisedvisiontransformers,in:ProceedingsoftheIEEE/CVF International Conference on Computer Vision, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization An empirical study of training self-supervisedvisiontransformers,in:ProceedingsoftheIEEE/CVF International Conference on Computer Vision, pp

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.249109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.561115Z digest=sha256:84bb77fed17d60b622ad14c2dc014125804d24d7816b9b78b953992182c5b34a

Observation f7f2fcf9-4a59-4e6d-84ee-f354477db27b · outbound

This paper cites Videomem:Constructing,analyzing,predictingshort-termandlong- term video memorability, in: Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Videomem:Constructing,analyzing,predictingshort-termandlong- term video memorability, in: Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, pp

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.239926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.564420Z digest=sha256:13189b0d85d0b8cff72e685ad4c4f8bc1cf054ea748b1ab6c0d08e8e61d320db

Observation 5073d918-1ea4-4e6f-9efb-10efdc8bca9f · outbound

This paper cites Annotating, understanding, and predicting long-term video memora- bility, in: Proceedings of the 2018 ACM on International Conference on Multimedia Retrieval, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Annotating, understanding, and predicting long-term video memora- bility, in: Proceedings of the 2018 ACM on International Conference on Multimedia Retrieval, pp

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.230167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.567791Z digest=sha256:051e422f280140fffb86e19c3efe0990a4a45052d0d49cbf9488333ae2923ab1

Observation 6095a6a9-f846-4065-a195-77aa90910f05 · outbound

This paper cites an unresolved cited work.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:11:47.220995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.571702Z digest=sha256:15c63b203b2dd01a95291a33ee5b0f38916664a3fdcc3ed7b114a51bc5b0e453

Observation 3ffd4819-98e8-4c73-b6fe-5d4919e5d333 · outbound

This paper cites Aimultimedialab at mediaeval 2022:Predictingmediamemorabilityusingvideovisiontransformers and augmented memorable moments , 12–16.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Aimultimedialab at mediaeval 2022:Predictingmediamemorabilityusingvideovisiontransformers and augmented memorable moments , 12–16

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.212077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.574728Z digest=sha256:a88df8fc5ec5dcf271535c16de29ee495c6b756455096b7d083966d65d1a5b39

Observation 5ce3ed93-6776-41c0-901a-45c0efaa3776 · outbound

This paper cites an unresolved cited work.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:11:47.201644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.577868Z digest=sha256:709fcd04009c0bad78f6aa866b640e17ec5a62c680e926e1b2c3e9fabad61d0e

Observation 6e36c3a1-ad21-47ba-8486-1d1bdfc0bd25 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale, in: International Conference on Learning Representations, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization An image is worth 16x16 words: Transformers for image recognition at scale, in: International Conference on Learning Representations, pp

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.181724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.585252Z digest=sha256:b5460ed2bba1f049ba9f817ae51ab027eda9a90ad5fae78549183734d6b04e7f

Observation 2a9e927e-a2ce-4c9b-8ee3-b13738bf9dc4 · outbound

This paper cites Modular memorability: Tieredrepresentationsforvideomemorabilityprediction,in:Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Modular memorability: Tieredrepresentationsforvideomemorabilityprediction,in:Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.171965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.588453Z digest=sha256:d4097c1602c13b199e9f211e9d0c7dc25fdbd9f35d6e59438a8363630e67828c

Observation 84ec52ee-4516-4af3-b617-78011458f673 · outbound

This paper cites Memory: A contribution to experimental psychology.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Memory: A contribution to experimental psychology

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.162467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.591673Z digest=sha256:e2c406e45adeb890c1935621a255e5fa58afcb7149fd937c552866ee54d48227

Observation de1c3754-e5bd-4c67-8476-58f973fd7335 · outbound

This paper cites Amnet: Memorability estimation with attention, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Amnet: Memorability estimation with attention, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.153275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.595240Z digest=sha256:00f92e94f3b29ba80cbbe9c5673b97b7052633527ab9e4d2899000a30f7364ee

Observation c0f6c0f8-f23b-4f2e-94c7-236cc1b7868c · outbound

This paper cites Supervised video summarization via multiple feature sets with parallel attention, in: 2021 IEEE International Conference on Multimedia and Expo, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Supervised video summarization via multiple feature sets with parallel attention, in: 2021 IEEE International Conference on Multimedia and Expo, pp

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.143515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.598319Z digest=sha256:984181ba9145fd673762e8f31b2c0300f56f5744f427a7cca5a91bd6eb4c406e

Observation be5b7807-0dae-4325-b6c1-3758efa4560c · outbound

This paper cites Creating summaries from user videos, in: Computer Vision–ECCV 2014: 13th European Conference, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Creating summaries from user videos, in: Computer Vision–ECCV 2014: 13th European Conference, pp

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.133542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.601286Z digest=sha256:b206a3f3c5e3c7a9627aaf611892e6faed04bda3302c9f2cf2f115846d603b67

Observation fd0c0786-15b0-4c1a-921d-bab3f1d12ecb · outbound

This paper cites Learning computational models of video memorability from fmri brain imag- ing.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Learning computational models of video memorability from fmri brain imag- ing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.124114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.604322Z digest=sha256:1701c49aa4ceb2c2b018c78c4b11f1230fbe702175dc71a0f5c5f6a6c7ca40f4

Observation ffe99691-de32-4b92-ba29-32342f3291fb · outbound

This paper cites Self-supervised co-training for video representation learning.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Self-supervised co-training for video representation learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.114758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.607559Z digest=sha256:aa53c0c51d19d7dc1ed39c878471b9d0b4fb14198cb4f134234215444497144a

Observation ede9dd67-556e-44a8-a51b-f425d216a1c5 · outbound

This paper cites Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.105493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.611647Z digest=sha256:3d6a65046feb1a06f9169ca1c9c66ba9b7ed986b5d7d63bfe8bece525aa6e24c

Observation d0174df5-c239-4170-b3ba-5ff567d05f2d · outbound

This paper cites Momentum contrast for unsupervised visual representation learning, in: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Momentum contrast for unsupervised visual representation learning, in: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.095543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.614733Z digest=sha256:bc7118beaeb4b18969595b929b2cf42bcdc9bd1a173d8438d07cd65d524cba39

Observation a059f577-50fa-40ae-9079-d16b7f94146b · outbound

This paper cites Deep residual learning for image recognition, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Deep residual learning for image recognition, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:46.617686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:46.617686Z digest=sha256:0fbb25276e675b23117e0ecf109772bfec853c2446c66f3840eb1f172af6871f

Observation 8dec3552-4bd3-42df-a5d2-3c1405283945 · outbound

This paper cites Densely connected convolutional networks, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Densely connected convolutional networks, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.080297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.620718Z digest=sha256:1ea3b01ce27ee50f169209e092a45dc15ffd8e1c81a32fcefb373f7b8319e86b

Observation ce92ed41-a7c0-464d-836a-ce8fd1d80404 · outbound

This paper cites Whatmakesanimage memorable?, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Whatmakesanimage memorable?, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.070066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.624420Z digest=sha256:e080827dc6f0191b28fd53d064abd1b0570736bc2f36191ec33ba9ded08b2055

Observation 2ea92ec0-1555-4f97-9453-a060091b111f · outbound

This paper cites an unresolved cited work.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:11:47.059531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.627680Z digest=sha256:e83aa3790aea0973ae5ecac559d44c653e1f311c7990bc9c8a6a344b5533e19a

Observation bea60b69-7984-4d34-b94b-86b41c1fe728 · outbound

This paper cites Understanding and predicting image memorability at a large scale, in: Proceedings oftheIEEEInternationalConferenceonComputerVision,pp.2390– 2398.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Understanding and predicting image memorability at a large scale, in: Proceedings oftheIEEEInternationalConferenceonComputerVision,pp.2390– 2398

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.048543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.630767Z digest=sha256:bdeaa5f12c3aa30a3bed63489c52c618c7ecc9f0f6e75824c49303c3bc6409ce

Observation 91762437-a41f-4ea4-89f3-4819249ab71e · outbound

This paper cites Topic-oriented text features can match visual deep models of video memorability.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Topic-oriented text features can match visual deep models of video memorability

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.038975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.634020Z digest=sha256:599cbfd346c50e09d8599c1e021d1a3f2624806f5101c5fa92d0f5b87e934563

Observation 72896d0d-303f-47a7-862a-0fd231a61be8 · outbound

This paper cites an unresolved cited work.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:11:47.029606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.637286Z digest=sha256:a9614498a6ab6bfabfe281c230e8fc813d60d1c420e534727d889b66ff563b3d

Observation 48757189-2542-41bc-aac0-c056b1ca5b4b · outbound

This paper cites Scene memory is more detailed than you think: The role of categories in visual long-term memory.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Scene memory is more detailed than you think: The role of categories in visual long-term memory

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.020132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.640442Z digest=sha256:ecbed2349e498a120c9cd96cd6200aa78830904a2382557aa70d650a6ffb889b

Observation f08918d2-d602-4c82-9f48-86c00bf63e79 · outbound

This paper cites Multimodal deep features fusion for video memorability prediction, in: Working Notes Proceedings of the MediaEval 2019 Workshop (CEUR Workshop Proceedings), pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Multimodal deep features fusion for video memorability prediction, in: Working Notes Proceedings of the MediaEval 2019 Workshop (CEUR Workshop Proceedings), pp

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.009187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.643491Z digest=sha256:1f5225916553a90ade3b068b9d1d422facbcdc7648c68335250074a594855dfb

Observation e034b08a-7b4c-4fec-8f8a-0b6c0e015bf6 · outbound

This paper cites Adaptive multi- modalensemblenetworkforvideomemorabilityprediction.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Adaptive multi- modalensemblenetworkforvideomemorabilityprediction

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.998252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.646654Z digest=sha256:15768327a85a90da4eace1a2a59cfde030f053828b79b6e98d6cf0a5e1a35d22

Observation 71079d0d-67ba-4b7e-b32f-27c3ec9ed57d · outbound

This paper cites Deephierarchicallstmnetworks with attention for video summarization.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Deephierarchicallstmnetworks with attention for video summarization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.988125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.649773Z digest=sha256:5692a43d7dbbedc88e4a6c8a41d0e8c03bb4492419aab6768f9f2925519412ba

Observation ff956c57-9814-42bf-b4dd-ec3d6c424eae · outbound

This paper cites an unresolved cited work.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:11:46.976935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.652766Z digest=sha256:b5817cf1e12e43731c1f89f208b30e66eb1364fd5de7ef9f73d4d99209c98ec9

Observation 23248fe8-ccad-4499-a197-a514ea451e42 · outbound

This paper cites an unresolved cited work.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:46.659487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:46.659487Z digest=sha256:36df124460c069e907db2a36fd81d2a31a7aa07444a1559d9d61db5ae09af240

Observation 91ecfd1f-9f8f-41fa-a58d-052615cc9d23 · outbound

This paper cites Video storytelling based on gated video memorability filtering.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Video storytelling based on gated video memorability filtering

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.943516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.666658Z digest=sha256:e436ef001954e261ac404485a47688db06ae95e5a581ac92eaadc9362f7f3ed3

Observation 1d0b8dae-449e-4cbc-aff0-1aaf2fabcbce · outbound

This paper cites Audio-visual instance discrimination with cross-modal agreement, in: Proceedings of the IEEE/CVFConferenceonComputerVisionandPatternRecognition, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Audio-visual instance discrimination with cross-modal agreement, in: Proceedings of the IEEE/CVFConferenceonComputerVisionandPatternRecognition, pp

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.934791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.669574Z digest=sha256:47cccbe316996f65d16e02cb26484ed86398631ff26ba79c03dcf5ae45b43fda

Observation 97421f1b-278c-400a-9857-ab4a7a232d88 · outbound

This paper cites 10012–10022.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization 10012–10022

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:46.663446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:46.663446Z digest=sha256:32d5e02700fc8ab128303311ecb569864355a28a76305706f0d9d483702dd905

Observation 570c3d7d-ef52-4e09-8b49-7ed65ff3e591 · outbound

This paper cites Multimodal memorability: Modeling effects of semantics anddecayonvideomemorability,in:ComputerVision–ECCV2020: 16th European Conference, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Multimodal memorability: Modeling effects of semantics anddecayonvideomemorability,in:ComputerVision–ECCV2020: 16th European Conference, pp

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.925463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.677044Z digest=sha256:ba36912f40570c3eabf767c7ef1fa5ab3edbf71363dd0a19d6f1827256bc05ac

Observation 8906cfb4-8275-4239-b9c0-d7363c616322 · outbound

This paper cites Spatiotemporal contrastive video representation learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Spatiotemporal contrastive video representation learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.916181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.680222Z digest=sha256:4ba3a2baab06c567bfff57d0e184960edaaee4c724921f01a52c191b648f9fc9

Observation 47a9df1e-e41e-4f9d-b905-a0743f5d9dff · outbound

This paper cites Does Video Summarization Require Videos? Quantifying the Effectiveness of Language in Video Summarization.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Does Video Summarization Require Videos? Quantifying the Effectiveness of Language in Video Summarization

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:11:46.772873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.672648Z digest=sha256:6e2e0cd6e47370cb7950cad2270cb6ca9f94e6d1ca374c32357364f6e037d224

Observation 50758686-3c06-448b-8f17-18b60185de27 · outbound

This paper cites Self-supervised video transformer, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Self-supervised video transformer, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.897162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.686690Z digest=sha256:d035f043e144504f3efa06373221642c5bd380e6bad49ddebcdb6d5d245105d4

Observation 6bba91b0-db5e-4035-b671-d0e3647c7f9b · outbound

This paper cites Ex- ploringmultimodality,perplexityandexplainabilityformemorability prediction, in: Working Notes Proceedings of the MediaEval 2021 Workshop (CEUR Workshop Proceedings), pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Ex- ploringmultimodality,perplexityandexplainabilityformemorability prediction, in: Working Notes Proceedings of the MediaEval 2021 Workshop (CEUR Workshop Proceedings), pp

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.887360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.689785Z digest=sha256:fd674583d0b5970c64755acc7c878ff4265d1ae4a1758d59c3c1813c94e14ec9

Observation 326a40c3-738d-40ba-a555-8067694ee21c · outbound

This paper cites Learning transferable visual models from natural language supervision, in: International Conference on Machine Learning, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Learning transferable visual models from natural language supervision, in: International Conference on Machine Learning, pp

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.906842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.683330Z digest=sha256:889f3c891f2463f9a03c2ce8898ea5e0fd7920e9e7990e7084583a84c6bf2e4c

Observation 5c39147e-8b28-4b8b-9caa-19a1c1050184 · outbound

This paper cites an unresolved cited work.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:11:46.867763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.695940Z digest=sha256:f918bb1554df70db8830a7b4bff89772c37498fe6df4a6956c48542f02b5e54f

Observation e2f732db-0da7-458a-8223-282ff3392ce8 · outbound

This paper cites A network linking scene perception and spatial memory systems in posterior cerebral cortex.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization A network linking scene perception and spatial memory systems in posterior cerebral cortex

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.848255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.702055Z digest=sha256:3c40af6a93657000c9e80eb6bdeba9a8f62918343bcfeacba78ac5729ec81d63

Observation afbe63c6-0be1-47be-af86-05dae2b3d1fe · outbound

This paper cites Tvsum: Summarizing web videos using titles, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Tvsum: Summarizing web videos using titles, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.877589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.692753Z digest=sha256:10b78ceb0f206b121d664944c3c41b731ab65fe448c51d27eca3a988935e3748

Observation d06d11f4-aae0-4123-9feb-a02475e4e491 · outbound

This paper cites Predicting media memorability:Comparingvisual,textualandauditoryfeatures,103– 105.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Predicting media memorability:Comparingvisual,textualandauditoryfeatures,103– 105

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.829024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.708918Z digest=sha256:dc0a069253561b812709562f77495cc7d2b2f7845e09558577f41b8589e584c7

Observation 929fe93a-ef38-4ba4-8e2b-b044a5a33f5e · outbound

This paper cites Diffusing Surrogate Dreams of Video Scenes to Predict Video Memorability.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Diffusing Surrogate Dreams of Video Scenes to Predict Video Memorability

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:46.712366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:46.712366Z digest=sha256:6be64de035c82928def987fa9983b5e192000a60ddf5cd59845f8af8f070e1ab

Observation 6c5972c9-04a2-4f90-b2b6-8b390cfe2712 · outbound

This paper cites Siamese image modeling for self-supervised vision representationlearning,in:ProceedingsoftheIEEE/CVFConference on Computer Vision and Pattern Recognition, pp.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Siamese image modeling for self-supervised vision representationlearning,in:ProceedingsoftheIEEE/CVFConference on Computer Vision and Pattern Recognition, pp

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.819385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.715755Z digest=sha256:70abda4472e9dd655c7e46d89db4eaa0d90d87ce45cc8fb369420c5284470569

Observation 88fcf709-254d-43dc-94db-1cbe309cd1de · outbound

This paper cites Recurrent unit augmented memory network for video summarisation.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Recurrent unit augmented memory network for video summarisation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.838910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.705054Z digest=sha256:5edbf75cd58b537ee847a334d38f8b4d04fba3801fe9de4cd0fe8d5498271398

Observation 3c5dd633-7268-40ca-b90e-a4cbc551ed13 · outbound

This paper cites Modelling of video memorability using ensemble learning and transformers , 7–11.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Modelling of video memorability using ensemble learning and transformers , 7–11

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.803999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.722577Z digest=sha256:c155ce245b252f82d11f8d2537ade9644a7be37c2bf488616b3821db50ade640

Observation 7f3b7039-30bb-47c2-9dd0-42c8088d0041 · outbound

This paper cites Reconstructivesequence-graph network for video summarization.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Reconstructivesequence-graph network for video summarization

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.793870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.725551Z digest=sha256:e35a5ad954847f405c31f4d7311c7b677be4a83e9df03cf23c863e4ed6cdff71

Observation 47d8ed23-a8be-460a-9a4c-4d5c6f4ef6b5 · outbound

This paper cites Learning multiscale hierarchical attention for video summarization.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Learning multiscale hierarchical attention for video summarization

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.784016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.728553Z digest=sha256:ce5af375f6c1e5ee0aa28d4fd30adc4e3db06e560a7109f91b02dfb559976acc

Observation 4f87002f-97f5-4cc9-a326-97089498140c · outbound

This paper cites Training data-efficient image transformers & distillation throughattention,in:InternationalConferenceonMachineLearning, PMLR.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization Training data-efficient image transformers & distillation throughattention,in:InternationalConferenceonMachineLearning, PMLR

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:46.718784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:46.718784Z digest=sha256:1cf6ffcc9cb385f038e79ace49c7c65c187a9971a92c08d19aecb88e9800ecb0

Observation 71f02226-52a6-4ebb-824f-cb86936bc0d2 · outbound

This paper cites 2371–2375.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization 2371–2375

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.857713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.699116Z digest=sha256:6cd530383a89546f92da78d0e8a24fdc1893df968024dea534aba91710b237cf

Observation 7cf91807-10fb-4bf6-88e2-4e356758669b · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 4065–4080.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 4065–4080

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:47.191839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.581336Z digest=sha256:45aadf2a145e3ade1942bea017102fe8a654f15c12f555408753c9a9526ff613

Observation b9de527e-2891-4ccf-a400-2d771e263e9a · outbound

This paper cites IEEE Transactions on Image Processing 31, 1573–1586.

Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization IEEE Transactions on Image Processing 31, 1573–1586

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:46.965734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:46.655922Z digest=sha256:4a925c8b9d4ca19295846acc9d1c24df18055694566c743350fd062ed3879db7

Pith citing papers

No inbound Pith citation observations are available.