Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:47:28.818411Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2507.17844.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:47:28.818411Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d3859547-cc78-47d7-a1a8-e3fb88d0406d · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Review on wearable technology in sports: Concepts, challenges and opportunities,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 338bca84-3529-4456-bff5-dc53d909da87 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00d64863-bf51-4eca-aa1e-a18a39129839 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74989adb-e10d-4565-aa3a-e6a220763406 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition A path towards autonomous machine intelligence,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41494e1c-70a4-4022-8f07-fb841d1a0c80 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Self-supervised learning from images with a joint- embedding predictive architecture,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7bd5611-59ef-4565-afb7-7d959e6fd077 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition V -JEPA: Latent video prediction for visual representation learning,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 643eb1ac-91a6-4c1e-bbc1-b16805fde231 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 029376f8-78cd-460f-9024-e915cb0e2d86 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 354d65ae-8848-46b7-af8c-44abcdf96f4f · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Computer vision for sports: Current applications and research topics,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc57ca70-1e9c-48d0-8886-435839cb8a2a · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Soccernet-v2: A dataset and benchmarks for holistic understanding of broadcast soccer videos,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79293dc3-e8d5-4646-81d1-c6af34e48e68 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Soccernet: A scalable dataset for action spotting in soccer videos,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fac1c96e-771e-4fe3-bb4b-d68df025fad1 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Fine-grained action recognition on a novel basketball dataset,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8fd68349-1c4d-498b-8c8e-be5fefeb4b81 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Soccernet caption: Dense video captioning for soccer broadcasts commentaries,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f03400f5-4020-4fe7-bf01-b888895bf08e · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Sports video captioning via attentive motion representation and group relationship modeling,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c99070ce-cc63-4f03-b9a5-68d95982b5fe · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Matchtime: Towards automatic soccer game commentary generation,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1709e277-9f4e-4a28-9110-8115693abf3d · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c6771fa-87fd-482a-9d0d-3a0ade95999c · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Fine-grained video captioning for sports narrative,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21d0f4e1-c2ea-4a6d-8575-36f75976bcd0 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Finegym: A hierarchical video dataset for fine -grained action understanding,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 35700ec6-dfc7-473b-8984-aab3d5f17bf9 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Finediving: A fine - grained dataset for procedure-aware action quality assessment,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d7bb799-2d82-4796-b9fd-f7259c4d6e1c · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Tacticai: An AI assistant for football tactics,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96745f7f-0097-49c8-af54-2e418a36fc36 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition VARS: Video assistant referee system for automated soccer decision making from multiple views,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 18204d28-eecd-429a-89a7-660da1fa7d85 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition X -VARS: Introducing explainability in football refereeing with multimodal large language models,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19ac599a-1c98-46cc-b946-2036d5ea9ff0 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Sports-QA: A large -scale video question answering benchmark for complex and professional sports,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e851dbc-fdf2-424a-a7fb-05e4c106ecd6 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition SportQA: A benchmark for sports understanding in large language models,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c80b69f-ef7f-4e47-a145-9fe7c26de692 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05827de8-aa86-48d4-b09d-4a6f57139e35 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Flamingo: a visual language model for few -shot learning,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6326f665-380b-4582-b3d0-ee35e9d3a584 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27132ff3-fe73-447b-85dc-23a36b03cca7 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition BLIP -2: Bootstrapping language- image pre -training with frozen image encoders and large language models,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e314612-64b6-4c4f-bfae-a239a2f1a818 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Learning transferable visual models from natural language supervision,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3e888dd-8ada-490a-b844-0a4a727806c1 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Sigmoid loss for language image pre -training,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8cd9abe6-d824-4158-9990-45f7127f4079 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition MVBench: A comprehensive multi-modal video understanding benchmark,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 09140373-a5f7-4ad0-adf8-01dfc9870cf4 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Llama -vid: An image is worth 2 tokens in large language models,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8613031e-91c7-49b0-9caf-e3f67a6c1907 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Video-llama: An instruction-tuned audio- visual language model for video understanding,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5588f34-ed5d-4267-aedc-2f891d43403c · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Temporal alignment networks for long-term video,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 94b8b072-21f6-433a-b120-af04d79c59e9 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Multi-sentence grounding for long -term instructional video,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b3295252-d1f3-4c72-8202-08f2a77f17de · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Panda- 70M: Captioning 70M videos with multiple cross -modality teachers,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1e043d5-b327-4f58-97d0-c7c0245eea61 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Vid2seq: Large -scale pretraining of a visual language model for dense video captioning,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 542eae72-54e7-47eb-a817-15cb45c4713e · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Streaming dense video captioning,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c1bb864-2069-46bd-ad3c-82db58b13626 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Autoad: Movie description in context,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2287eebb-0f85-4024-bd33-826f84e35ab4 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Autoad II: The sequel —who, when, and what in movie audio description,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b3b75a8e-7dfc-4aaa-9db7-271e0fde46ba · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition Autoad III: The prequel —back to the pixels,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 225b382c-0702-4d58-aeff-48fbc85917e1 · outbound
SV3.3B: A Sports Video Understanding Model for Action Recognition NSVA Subset: Basketball Video -Text Dataset,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.