Pith. sign in

Paper Citation Record · LEDGER

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos

As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.12623.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12623 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:48:30.581939Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 34ea7804-2841-406c-a928-502d210f5a08 · outbound

This paper cites BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:29.881132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:29.881132Z digest=sha256:201bdcc68a77ec665998936379f6d8edcb4df8ac1a7da4552fe7665fe6e00c9e

Observation bf0cc22b-2fe0-4413-a07d-e2ec1215f176 · outbound

This paper cites InProceedings of the 2020 Con- ference on Empirical Methods in Natural Language Processing (EMNLP), pages 4707–4716, Online.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos InProceedings of the 2020 Con- ference on Empirical Methods in Natural Language Processing (EMNLP), pages 4707–4716, Online

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:31.404579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:48:29.990078Z digest=sha256:0c5551a32679be3402a49a1d0b49911d0339967a6506e1ac7f5d458198fa90b2

Observation 795be711-6f3f-4c84-a376-8019c03f01d3 · outbound

This paper cites MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of Videos.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of Videos

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:48:30.791494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:48:30.247861Z digest=sha256:844db0d9944ac866ccc05d0f6b513476ae3244bdcc0c9a3eec72b7796af224be

Observation 2de4cfa6-e2a6-4e58-af28-867351195999 · outbound

This paper cites How2: A Large-scale Dataset for Multimodal Language Understanding.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos How2: A Large-scale Dataset for Multimodal Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.364282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.364282Z digest=sha256:e52b8581f579c73e997decc8f4762d2681ca2b0f7542ea89c099618d2b9fa0e0

Observation 5f5296aa-02e4-4b02-892e-3f33da3e84d6 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.467237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.467237Z digest=sha256:2d77f38fa628f92829983cf88552d2449f71e058aa1567d3c830b6f1a574395f

Observation 70c46022-23c8-4b93-8fac-770eac918e80 · outbound

This paper cites InProceed- ings of the IEEE conference on computer vision and pattern recognition, pages 1059–1067.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos InProceed- ings of the IEEE conference on computer vision and pattern recognition, pages 1059–1067

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:31.211184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:48:30.581939Z digest=sha256:5a454f0fe429ec6edd27f0e392fa094f46786c43ccb7c820ad5946429639b155

Observation 7607221a-21b7-4272-9467-33ef23071a16 · outbound

This paper cites InComputer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VII 13, pages 505–520.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos InComputer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VII 13, pages 505–520

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:31.904581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:48:29.588655Z digest=sha256:17b5b001014c1e33f3d92318d4d50238489fb09818ba733f9810527b71ba3f75

Observation 4a6f7f72-e090-462a-9b56-96a4427c9d89 · outbound

This paper cites Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.130052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.130052Z digest=sha256:0de02fee734247c8cf4185c1c63ae6b74cb8713f692db6a6e2969f7fb8687396

Observation 264c2662-35b9-4326-84c6-0d60e4ab4ab3 · outbound

This paper cites Deep Communicating Agents for Abstractive Summarization.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos Deep Communicating Agents for Abstractive Summarization

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:29.217734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:29.217734Z digest=sha256:3e9ef74eda2ec26d7e05592093a2fbed4675f28479b03096d22bf65223980881

Observation 63b842e4-0c12-47d4-98b7-1c8586ae3834 · outbound

This paper cites InPro- ceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pages 1–10.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos InPro- ceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pages 1–10

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:31.740163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:48:29.687369Z digest=sha256:28d984a25f6e06d4dfa07872d1549a120c47e2f8b1b9f7df491dd7f22bd0ff39

Observation 8063f4d8-c1ec-4ac4-89b5-c4169252c7e1 · outbound

This paper cites Multi-modal Summarization for Video-containing Documents.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos Multi-modal Summarization for Video-containing Documents

Reference 2020

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:48:31.016644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:48:29.447866Z digest=sha256:4ab77410b7c4758849375fd5a6918d0e228bfb9b686e1481fea5d8fc0b5f6f21

Observation 76ee0f58-4157-4dd3-b9f3-e554262b00db · outbound

This paper cites bert2BERT: Towards Reusable Pretrained Language Models.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos bert2BERT: Towards Reusable Pretrained Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:29.332019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:29.332019Z digest=sha256:2b6c7094e5f2525b29b025f7dbf0c0a1bc8b3c128c05118725eb7bd4ac4705f7

Observation a5eb272a-6f4d-4fad-a5ed-e63d7755510e · outbound

This paper cites InFindings of the Association for Computa- tional Linguistics: EACL 2023, pages 880–894.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos InFindings of the Association for Computa- tional Linguistics: EACL 2023, pages 880–894

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:31.573166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:48:29.789251Z digest=sha256:41ea4278ee60e9a53a014f066b883ff33f66c3ed07b752f9afe17147298a451f

Pith citing papers

No inbound Pith citation observations are available.