Pith. sign in

Paper Citation Record · LEDGER

Unsupervised Transcript-assisted Video Summarization and Highlight Detection

As of 7 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2505.23268.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23268 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:54:19.856159Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 255ca5ff-7fb4-4850-be8d-db87a2bb0e1a · outbound

This paper cites Align and attend: Multimodal summarization with dual contrastiv e losses,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Align and attend: Multimodal summarization with dual contrastiv e losses,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:25.528941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:18.096395Z digest=sha256:a413a5e48a01d22a170c80d3a589513a9f1f822c6f01fec4ac374393fb0917a3

Observation 6e930162-5b9e-46fd-bbd3-55ae43317b2f · outbound

This paper cites Clover: T owards a unified video-language alignment and fusion model,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Clover: T owards a unified video-language alignment and fusion model,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:25.313406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:18.203435Z digest=sha256:ccadc59127c0798c5a8f90e20a0e75a307bd7ebe1bb3c24fd06f71e38f0583eb

Observation 37a123c6-8d8f-45e2-9e3d-cf28bb2b2ff3 · outbound

This paper cites SUSiNet: See, understand and summarize it,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection SUSiNet: See, understand and summarize it,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:25.093946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:18.296863Z digest=sha256:bbf508b0e57d07e776c705e661daf8d0c4ac8e10ab02972f40f2b2703cb0785f

Observation b95352db-86b7-431a-b4c9-d784454a8592 · outbound

This paper cites UMT: Un ified multi-modal transformers for joint video moment retrieval and highlight detection,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection UMT: Un ified multi-modal transformers for joint video moment retrieval and highlight detection,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:24.873048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:18.391846Z digest=sha256:a465a06dd672815a38916a4b3acacb4425f965101c0faf5b0ddeb0a0c08119d4

Observation b18d4533-263d-4070-8ab2-c6c5b36c6e65 · outbound

This paper cites CLIP-it! la nguage-guided video summarization,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection CLIP-it! la nguage-guided video summarization,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:24.575792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:18.485127Z digest=sha256:a6ec1505800cc661c93adb30e37e12b0eb1fe747a2878b087ea37485dccd5bc8

Observation eed8a926-2057-4442-8544-0f509613b852 · outbound

This paper cites Creating summaries from user videos,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Creating summaries from user videos,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:24.373861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:18.562836Z digest=sha256:0fd1d307c826588941940ea5e12b38a448f78d2b8f29cd08dddaa398512b0cc8

Observation 45204a32-4423-49fc-9ad1-098b3bba9675 · outbound

This paper cites TVSum: Summarizing web videos using titles,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection TVSum: Summarizing web videos using titles,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:24.135808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:18.624966Z digest=sha256:64d6679fd27270298ecab70e85f11d52d7d1189941d192570478e60350a1fe53

Observation e7d54eac-11c2-4a59-a3f9-22799e0ee230 · outbound

This paper cites Deep reinforcement learning for uns upervised video summarization with diversity-representativeness reward,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Deep reinforcement learning for uns upervised video summarization with diversity-representativeness reward,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:23.948873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:18.696700Z digest=sha256:882a90e8874cb1a666a342f575e3200827a630b591a7388ab3f1afc5eb9824b2

Observation b50f2ff4-c28b-4b04-9487-0885bc7b84bd · outbound

This paper cites Summarizing videos using concentrated attention and considering the un iqueness and diversity of the video frames,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Summarizing videos using concentrated attention and considering the un iqueness and diversity of the video frames,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:23.704468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:18.750814Z digest=sha256:ec9d940cf8574a8cc12ec3f71edba92247200f7f4edd0776923e3d699db8fbe4

Observation e77ff443-b4d3-477e-aa60-56085cc752ce · outbound

This paper cites TL;DW? Summarizing Instructio nal Videos with Task Relevance and Cross-Modal Saliency,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection TL;DW? Summarizing Instructio nal Videos with Task Relevance and Cross-Modal Saliency,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:23.487700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:18.814953Z digest=sha256:884d85496bcaeba40afb75c2605e9108e12846e766e727dea17658c9dfad9feb

Observation 55f3bec8-60de-408a-881d-0d810d202a93 · outbound

This paper cites Rethinkin g spatiotem- poral feature learning: Speed-accuracy trade-offs in vide o classification,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Rethinkin g spatiotem- poral feature learning: Speed-accuracy trade-offs in vide o classification,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:23.276417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:18.880999Z digest=sha256:51116fd5b742975074f4654bce63949f7ed055261946829c2699d98b715952ef

Observation 9d8e2308-f7ff-452b-bd09-0dece83de2df · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:18.964559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:18.964559Z digest=sha256:008e3e2e1099ad29c6dbd38606062013281e78f8f8f4881c559dc191876542bb

Observation 8f372dae-3bed-4e7b-80a0-5702f1a0bce0 · outbound

This paper cites AREDSUM: Adaptive redundancy-aware iterative sentence ranking for extracti ve document summarization,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection AREDSUM: Adaptive redundancy-aware iterative sentence ranking for extracti ve document summarization,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:23.095873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:19.039626Z digest=sha256:f1dbeb84456d37a563ccd0f19c2ead4eca9de4bef06e013be6fc2082862f280f

Observation cff0ff2b-9299-4c67-88a8-4af7ca625b7f · outbound

This paper cites Attention is all you need,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Attention is all you need,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:22.903832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:19.086216Z digest=sha256:72e44f940504535463be160a68ae70c3964437ea96494fa1d67d6bb37f56f4f3

Observation 701e8788-8a6e-4869-bdc3-bb00788978c6 · outbound

This paper cites VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:19.140326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:19.140326Z digest=sha256:7aa43afd5c2ada4c247b0e7d15c16d9d0f332481d2037f5c50ef651d0205d0f6

Observation 5254dac2-47db-4fad-a325-79a852aafcc6 · outbound

This paper cites Simple statistical gradient-followi ng algorithms for connectionist reinforcement learning,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Simple statistical gradient-followi ng algorithms for connectionist reinforcement learning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:22.639669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:19.191277Z digest=sha256:470994e94e92c3189bdcaa2ba3a9f53c9c65f366f6235c05b6d2dece36a2a2e1

Observation a98664f6-bc7e-442c-ae1b-07415d1cbf90 · outbound

This paper cites Adam: A method for stochastic opt imization,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Adam: A method for stochastic opt imization,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:19.241770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:19.241770Z digest=sha256:263d877bfc1f60e8f142565adffa1486aa5f55fe77c2a010b423fe654f8f8b8c

Observation 58d02cb6-099b-4849-8583-bff9ff7666aa · outbound

This paper cites Abstractive text summarization using sequence-to-seque nce RNNs and beyond,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Abstractive text summarization using sequence-to-seque nce RNNs and beyond,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:22.405014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:19.289591Z digest=sha256:795d15e245d692e60182f60157f0e1f351296d20aeb264dc7dd7aaa17419ba64

Observation 61934291-2a83-4dc6-9eb2-afc055363ec3 · outbound

This paper cites HowTo100M: Learning a text-video embedding by watching hu ndred million narrated video clips,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection HowTo100M: Learning a text-video embedding by watching hu ndred million narrated video clips,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:22.235780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:19.335478Z digest=sha256:f00edb4700bce0860626fb9b1aabfb762021cd9f25cea5ac15b93bb0d8deecfd

Observation a23ffb87-3931-4893-8939-5b73da35ad07 · outbound

This paper cites Mr. HiSum: A large-scale data set for video highlight detection and summarization,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Mr. HiSum: A large-scale data set for video highlight detection and summarization,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:22.019824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:19.400921Z digest=sha256:8adc1e3d4b4246b5160b05a975b453b23df9f174a708a4b1aae3d59504b8cca8

Observation 75c4784f-b26e-4d75-83ae-6f872e9962c4 · outbound

This paper cites Reth inking the evaluation of video summaries,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Reth inking the evaluation of video summaries,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:21.831846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:19.459604Z digest=sha256:7b03fca714e63a1ae4cdc6055f16e26965e7a8d80ebfb77a2fae4d677d203d8f

Observation 72018b24-f092-453f-bdfc-639cb90eaf01 · outbound

This paper cites Zwillinger and S.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Zwillinger and S

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:21.598933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:19.537886Z digest=sha256:ba05fd5b834ee8bc580211c840cf81b25455054666a7e65143d740143c538c93

Observation e035518a-dc10-473d-bac5-0deeccae4290 · outbound

This paper cites The treatment of ties in ranking problem s,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection The treatment of ties in ranking problem s,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:21.180234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:19.632314Z digest=sha256:c8404b063426d1938c3e2c04df39795a9f6f77e6fb85887841111c0f6a7c412d

Observation 245fb392-2849-4c05-a7e9-de4e43be91e6 · outbound

This paper cites Cate gory-specific video summarization,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Cate gory-specific video summarization,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:20.866198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:19.724624Z digest=sha256:917a1dcdaf0f754554c1e3dba7dc6b4bdfdf3e1fa08f8c6c8a82ab1236aea4bc

Observation eb81ee09-0ff3-4f1d-b697-77588af1849d · outbound

This paper cites WikiHow: A Large Scale Text Summarization Dataset.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection WikiHow: A Large Scale Text Summarization Dataset

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:19.775314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:19.775314Z digest=sha256:440067b197f089801a40706f074ab59d7e8643be93fd02afb379d55d939583e6

Observation 1b59550b-3d56-4209-82fb-6d6801571aac · outbound

This paper cites COG N- IMUSE: a multimodal video database annotated with saliency , events, semantics and emotion with application to summarization,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection COG N- IMUSE: a multimodal video database annotated with saliency , events, semantics and emotion with application to summarization,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:20.479050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:54:19.856159Z digest=sha256:be4b05c5c43732d15b9af1773baa74fab263ffc33662447491739e5713499a1f

Pith citing papers

No inbound Pith citation observations are available.