Pith. sign in

Paper Citation Record · LEDGER

Comparing Learning Paradigms for Egocentric Video Summarization

As of 7 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2506.21785.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21785 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:23:43.243274Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact4
  • verified fuzzy1
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c8b10c0-3408-4d37-b3b5-e2fca4dcb349 · outbound

This paper cites UniVTG: Towards Unified Video-Language Temporal Grounding.

Comparing Learning Paradigms for Egocentric Video Summarization UniVTG: Towards Unified Video-Language Temporal Grounding

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:23:44.366891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T22:23:41.171290Z digest=sha256:58665e6acb305a22f5538aca42cf3e5974a3b000d8c4cdc82c1d8ffc023ba203

Observation 90c0f3b1-5916-4edf-8bfe-b087c15b290a · outbound

This paper cites VideoLLM-online: Online Video Large Language Model for Streaming Video.

Comparing Learning Paradigms for Egocentric Video Summarization VideoLLM-online: Online Video Large Language Model for Streaming Video

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:41.274752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:41.274752Z digest=sha256:ac2f99fb4218b3646699f452e09d098dd275944af159e54bd5e8a80a90058ad6

Observation 60fd24f9-7f8e-4d8d-bade-a4d8a55cc9a8 · outbound

This paper cites VideoMamba: State Space Model for Efficient Video Understanding.

Comparing Learning Paradigms for Egocentric Video Summarization VideoMamba: State Space Model for Efficient Video Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:41.396361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:41.396361Z digest=sha256:d6cacef331a35f5ab5ef282a935da5b1b633a81620eb8b43baf57ab97f7084bc

Observation bdf647d9-3f6f-4ac8-bb9d-c1c7908b92a9 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Comparing Learning Paradigms for Egocentric Video Summarization Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:41.495758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:41.495758Z digest=sha256:0d7a6ffa77bde81f06d4fc8a4e5a38c49a7290956853d810013b6bd5d7388af5

Observation bd7e80c6-f1a6-4542-be12-6e0b683c42a7 · outbound

This paper cites Is Space-Time Attention All You Need for Video Understanding?.

Comparing Learning Paradigms for Egocentric Video Summarization Is Space-Time Attention All You Need for Video Understanding?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:41.597659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:41.597659Z digest=sha256:57c7d173baa993c84024cf29538e2c84f4fec30d405b325637efa18e19780c67

Observation 3ae6917d-ea86-4819-bad3-0222acd15ed4 · outbound

This paper cites Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives.

Comparing Learning Paradigms for Egocentric Video Summarization Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:41.740633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:41.740633Z digest=sha256:61267a98c4077fefe0b2a7c412977206f17824cb2a6296ef8dd632b12568cf23

Observation 248d10b3-b2d1-4c0d-b58e-6c6220c682b4 · outbound

This paper cites Shotluck Holmes: A Family of Efficient Small-Scale Large Language Vision Models For Video Captioning and Summarization.

Comparing Learning Paradigms for Egocentric Video Summarization Shotluck Holmes: A Family of Efficient Small-Scale Large Language Vision Models For Video Captioning and Summarization

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:23:44.044371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T22:23:41.864051Z digest=sha256:2565bb0a0af983ef7cbcc29627506f08b561a52e3eb1ee1966e96b311aeed66c

Observation e5aed2b2-69a5-42fe-87bd-b4dce7efe8ad · outbound

This paper cites Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos.

Comparing Learning Paradigms for Egocentric Video Summarization Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:41.969901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:41.969901Z digest=sha256:bbf9c4aeaecd8d7cfac4ef78e0b08d184a5955b405be973a6904738f6b87cd20

Observation 5e87f939-0cdb-4a1c-bef7-4e88cf736c47 · outbound

This paper cites Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions.

Comparing Learning Paradigms for Egocentric Video Summarization Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:23:43.860557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T22:23:42.126947Z digest=sha256:af394cf6b0e9f2b83fffa26f3f2a9d976dc19bcd4c69128b1b8c9913a938a21f

Observation 504e0a5d-bbdc-46a6-9602-16acb08fd1e5 · outbound

This paper cites TransNet V2: An effective deep network architecture for fast shot transition detection.

Comparing Learning Paradigms for Egocentric Video Summarization TransNet V2: An effective deep network architecture for fast shot transition detection

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:42.246470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:42.246470Z digest=sha256:31f1f33289f052e3950245fcd039861fe63a3c735ba505ad92fbe8bad524b105

Observation 0335ffe1-ba69-4b56-90af-e15f5577fc2e · outbound

This paper cites Ego4D: Around the World in 3,000 Hours of Egocentric Video.

Comparing Learning Paradigms for Egocentric Video Summarization Ego4D: Around the World in 3,000 Hours of Egocentric Video

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:42.360607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:42.360607Z digest=sha256:6b5f657a052a345e0dec740d5b253b64a8116feefaf28e88eac42e92da5eec8a

Observation 2f91c3c5-a537-41c3-a653-86e22eebc81f · outbound

This paper cites From Sparse to Dense: GPT-4 Summarization with Chain of Density Prompting.

Comparing Learning Paradigms for Egocentric Video Summarization From Sparse to Dense: GPT-4 Summarization with Chain of Density Prompting

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:23:43.644497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T22:23:42.496883Z digest=sha256:173334a9ef7d4150653ee46e5012f2dc46b0327bf3ddccfe2895d3ee364e9ad7

Observation ece7af2f-5675-4b0a-8fb9-4d57d673dfd8 · outbound

This paper cites GPT-4 Technical Report.

Comparing Learning Paradigms for Egocentric Video Summarization GPT-4 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:42.657182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:42.657182Z digest=sha256:6f8f49e25f15dce564927456edbed028cf3684eba3e89afc38f580a3dfbebdec

Observation 998b9a53-fd36-4a54-a714-cae86b08c33a · outbound

This paper cites Cluster-based Video Summarization with Temporal Context Awareness.

Comparing Learning Paradigms for Egocentric Video Summarization Cluster-based Video Summarization with Temporal Context Awareness

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:42.816049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:42.816049Z digest=sha256:069418cfa185462aa6981185bfe73f5cf7c72d7ebb44cc54e003ea29e49462da

Observation a2d84fed-fae2-4413-bc87-bb203fad4ccf · outbound

This paper cites Creating Summaries from User Videos.

Comparing Learning Paradigms for Egocentric Video Summarization Creating Summaries from User Videos

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:42.917546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:42.917546Z digest=sha256:c443c33a08b965318579cd0f6c4e430d10e9475efea366a89615ac07e6241989

Observation 9d10f3e3-ba32-4d81-88d6-3cf103de00d9 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Comparing Learning Paradigms for Egocentric Video Summarization Sigmoid Loss for Language Image Pre-Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:42.985153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:42.985153Z digest=sha256:3cc86a6f46ed2dac990e55a3fc22bfc534622faa78ec38e1378197868ff18134

Observation e64a774c-1eea-46ea-987e-18ebdce0fd69 · outbound

This paper cites BIRCH: An Efficient Data Clustering Method for Very Large Databases.

Comparing Learning Paradigms for Egocentric Video Summarization BIRCH: An Efficient Data Clustering Method for Very Large Databases

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:43.101087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:43.101087Z digest=sha256:6903da27956b9825a93377acd4ff58f3d18767dd42793da23e932dcde8e71ed7

Observation 46b74914-c675-4384-9cba-1e2f60511c7d · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Comparing Learning Paradigms for Egocentric Video Summarization Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:43.159750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:43.159750Z digest=sha256:c61ee264cf12c998a5a1da1114543219707b8fdb00ceca875ecba84d75d04072

Observation 6b7f6f11-5185-4657-9eca-d3c950b1ec78 · outbound

This paper cites Nov 4, 2024, https://build.nvidia.com/nvidia/video-search-and-summarization 13.

Comparing Learning Paradigms for Egocentric Video Summarization Nov 4, 2024, https://build.nvidia.com/nvidia/video-search-and-summarization 13

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:44.649217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T22:23:43.243274Z digest=sha256:7564366d88c57438d54bae12d55d3b713d3abd3d02d86a890210cd56b37b96b6

Pith citing papers

No inbound Pith citation observations are available.