Pith. sign in

Paper Citation Record · LEDGER

SV3.3B: A Sports Video Understanding Model for Action Recognition

As of 10 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2507.17844.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17844 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:47:28.818411Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact3
  • verified fuzzy34
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3859547-cc78-47d7-a1a8-e3fb88d0406d · outbound

This paper cites Review on wearable technology in sports: Concepts, challenges and opportunities,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Review on wearable technology in sports: Concepts, challenges and opportunities,

Reference 1

Resolution
verified exact
doi, observed 2026-08-06T14:47:28.838947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.735332Z digest=sha256:ad8f80f3e7ebf4bcfc0d5fb48b59f7259d1f13412c12eee08d76448ae7a1ef26

Observation 338bca84-3529-4456-bff5-dc53d909da87 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

SV3.3B: A Sports Video Understanding Model for Action Recognition Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.737949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.737949Z digest=sha256:58cb4eabe4018904724fb53f35732cb4e5fdfeee8172ce730b2da0cf72241de3

Observation 00d64863-bf51-4eca-aa1e-a18a39129839 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

SV3.3B: A Sports Video Understanding Model for Action Recognition LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.740370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.740370Z digest=sha256:0e2217f834fa633ed986e6121a8a4f21e40e62ec8013d1f63db8f5891da99519

Observation 74989adb-e10d-4565-aa3a-e6a220763406 · outbound

This paper cites A path towards autonomous machine intelligence,.

SV3.3B: A Sports Video Understanding Model for Action Recognition A path towards autonomous machine intelligence,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.152382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.742825Z digest=sha256:aa114473831fc2c946d758ff12fd3e4b10262ef1232ceeb490e41f372ae23f89

Observation 41494e1c-70a4-4022-8f07-fb841d1a0c80 · outbound

This paper cites Self-supervised learning from images with a joint- embedding predictive architecture,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Self-supervised learning from images with a joint- embedding predictive architecture,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.146394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.745119Z digest=sha256:473ed9325d50ca08fc7eb38ed626aa3b5743f359cc9f502b2d991867f013cff6

Observation a7bd5611-59ef-4565-afb7-7d959e6fd077 · outbound

This paper cites V -JEPA: Latent video prediction for visual representation learning,.

SV3.3B: A Sports Video Understanding Model for Action Recognition V -JEPA: Latent video prediction for visual representation learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.140706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.747272Z digest=sha256:708bc71335cc8eafc1697cfd341932d443a73c60a665329000b67d4755ac9109

Observation 643eb1ac-91a6-4c1e-bbc1-b16805fde231 · outbound

This paper cites UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity.

SV3.3B: A Sports Video Understanding Model for Action Recognition UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:47:28.940562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.749711Z digest=sha256:75fa0351ecf81d426cf6f2712972a3c6e859b19c30179a8823c56141d81d8c37

Observation 029376f8-78cd-460f-9024-e915cb0e2d86 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

SV3.3B: A Sports Video Understanding Model for Action Recognition V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.752004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.752004Z digest=sha256:2e38b8e278f90ec98ad0efa43cea76290821428bcad9b5ba5d6ff24588a0999f

Observation 354d65ae-8848-46b7-af8c-44abcdf96f4f · outbound

This paper cites Computer vision for sports: Current applications and research topics,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Computer vision for sports: Current applications and research topics,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.134924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.754962Z digest=sha256:39b8294a62ca736df8966d4cd6c0d16187207f77404287f6291ac5e391409867

Observation dc57ca70-1e9c-48d0-8886-435839cb8a2a · outbound

This paper cites Soccernet-v2: A dataset and benchmarks for holistic understanding of broadcast soccer videos,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Soccernet-v2: A dataset and benchmarks for holistic understanding of broadcast soccer videos,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.129217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.757216Z digest=sha256:0d12a0753e943a2671f7c8ae15fe90a4f440d8ea7e148f8efff9763a6e3e5f8b

Observation 79293dc3-e8d5-4646-81d1-c6af34e48e68 · outbound

This paper cites Soccernet: A scalable dataset for action spotting in soccer videos,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Soccernet: A scalable dataset for action spotting in soccer videos,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.123627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.759138Z digest=sha256:e941db810295406575912189138c8f351d494fbee5d4460b02f18ffc5ee15f57

Observation fac1c96e-771e-4fe3-bb4b-d68df025fad1 · outbound

This paper cites Fine-grained action recognition on a novel basketball dataset,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Fine-grained action recognition on a novel basketball dataset,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.118061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.761422Z digest=sha256:050368ccf9aa1b09d48b97da7b43481510a399948a07138a426d39a9e98cd93a

Observation 8fd68349-1c4d-498b-8c8e-be5fefeb4b81 · outbound

This paper cites Soccernet caption: Dense video captioning for soccer broadcasts commentaries,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Soccernet caption: Dense video captioning for soccer broadcasts commentaries,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.112178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.763299Z digest=sha256:df94dbf21b66ff5eb96165795b6771eed9afb0bf39a797047959b080140995c3

Observation f03400f5-4020-4fe7-bf01-b888895bf08e · outbound

This paper cites Sports video captioning via attentive motion representation and group relationship modeling,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Sports video captioning via attentive motion representation and group relationship modeling,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.106353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.765170Z digest=sha256:1ec2f298cb9bcbf1553d5afe6ccd38940eefd3786f6ab8134bb1d92b312f28e9

Observation c99070ce-cc63-4f03-b9a5-68d95982b5fe · outbound

This paper cites Matchtime: Towards automatic soccer game commentary generation,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Matchtime: Towards automatic soccer game commentary generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.100713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.767048Z digest=sha256:d4d0eacfe79cfab8906ab3ebea545fbc4870b06fd5a3ef085501ebcd19d73003

Observation 1709e277-9f4e-4a28-9110-8115693abf3d · outbound

This paper cites Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark.

SV3.3B: A Sports Video Understanding Model for Action Recognition Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:47:28.925700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.768978Z digest=sha256:f99a3823cf5417a52d6d1edd3a315fb1ddd40f6e68bbd94c6761055707f8cabf

Observation 7c6771fa-87fd-482a-9d0d-3a0ade95999c · outbound

This paper cites Fine-grained video captioning for sports narrative,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Fine-grained video captioning for sports narrative,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.095279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.771064Z digest=sha256:b645a8b496b8048a24ecb230a605ecdde7e169c9ca9e3f8d483941f2696001db

Observation 21d0f4e1-c2ea-4a6d-8575-36f75976bcd0 · outbound

This paper cites Finegym: A hierarchical video dataset for fine -grained action understanding,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Finegym: A hierarchical video dataset for fine -grained action understanding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.089788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.772938Z digest=sha256:a995d83db5ecab5dfb382c8cc732955775cf4a02089830a55e56ebaa51668f48

Observation 35700ec6-dfc7-473b-8984-aab3d5f17bf9 · outbound

This paper cites Finediving: A fine - grained dataset for procedure-aware action quality assessment,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Finediving: A fine - grained dataset for procedure-aware action quality assessment,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.083763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.775051Z digest=sha256:b50e617121765bf61d8ed83bec5e677fb5b6ab89c6a7ba71a791560d8092b971

Observation 3d7bb799-2d82-4796-b9fd-f7259c4d6e1c · outbound

This paper cites Tacticai: An AI assistant for football tactics,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Tacticai: An AI assistant for football tactics,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.077698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.776961Z digest=sha256:59d41bacbfd35e37cb275beea31ada77aacf5eea499ee72aebc0477d87ca55d6

Observation 96745f7f-0097-49c8-af54-2e418a36fc36 · outbound

This paper cites VARS: Video assistant referee system for automated soccer decision making from multiple views,.

SV3.3B: A Sports Video Understanding Model for Action Recognition VARS: Video assistant referee system for automated soccer decision making from multiple views,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.071515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.778846Z digest=sha256:6f0bdb3fdd163840593b4e984a586cd28b173eff5126540eee9aa83d161d9a77

Observation 18204d28-eecd-429a-89a7-660da1fa7d85 · outbound

This paper cites X -VARS: Introducing explainability in football refereeing with multimodal large language models,.

SV3.3B: A Sports Video Understanding Model for Action Recognition X -VARS: Introducing explainability in football refereeing with multimodal large language models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.065929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.780649Z digest=sha256:30b9df5eb70e1be168eeb3db9d954e1e5711a3d227a689f5912ca4077d30dc78

Observation 19ac599a-1c98-46cc-b946-2036d5ea9ff0 · outbound

This paper cites Sports-QA: A large -scale video question answering benchmark for complex and professional sports,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Sports-QA: A large -scale video question answering benchmark for complex and professional sports,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.782472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.782472Z digest=sha256:234763a6a0db209429f1e3c4d60aa4cc94a55034b48f090e4de964c2c4e21761

Observation 1e851dbc-fdf2-424a-a7fb-05e4c106ecd6 · outbound

This paper cites SportQA: A benchmark for sports understanding in large language models,.

SV3.3B: A Sports Video Understanding Model for Action Recognition SportQA: A benchmark for sports understanding in large language models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.060399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.784339Z digest=sha256:d4d2927d0135cbf2f7075d6d37f7345c99c7020a263d81ef6d4cf463e4c50752

Observation 5c80b69f-ef7f-4e47-a145-9fe7c26de692 · outbound

This paper cites SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models.

SV3.3B: A Sports Video Understanding Model for Action Recognition SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.786186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.786186Z digest=sha256:7277cb1c917dd7be15faae147014231ddae7c5972620847b8f08324aaf9a8195

Observation 05827de8-aa86-48d4-b09d-4a6f57139e35 · outbound

This paper cites Flamingo: a visual language model for few -shot learning,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Flamingo: a visual language model for few -shot learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.054842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.788390Z digest=sha256:a1b9b68f43ccf888449a06761ca0cb2d16211ece7333539577f7ae87b2f6c5d5

Observation 6326f665-380b-4582-b3d0-ee35e9d3a584 · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

SV3.3B: A Sports Video Understanding Model for Action Recognition BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.049117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.790361Z digest=sha256:89d8b3a84c73f1612be3e773a58fb782370d53039d7fdaf1a110dbc58d50829c

Observation 27132ff3-fe73-447b-85dc-23a36b03cca7 · outbound

This paper cites BLIP -2: Bootstrapping language- image pre -training with frozen image encoders and large language models,.

SV3.3B: A Sports Video Understanding Model for Action Recognition BLIP -2: Bootstrapping language- image pre -training with frozen image encoders and large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.043418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.792353Z digest=sha256:f0063e058fac1a12cf3d555bdfb47babc6b75ec1fcbd896eb4f5aa6b926b51cc

Observation 1e314612-64b6-4c4f-bfae-a239a2f1a818 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Learning transferable visual models from natural language supervision,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.037574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.794410Z digest=sha256:a3acffd40a5797efbc4ecb04a2b648a10d8b8b0d73e32516f0f9a0468a62ecef

Observation c3e888dd-8ada-490a-b844-0a4a727806c1 · outbound

This paper cites Sigmoid loss for language image pre -training,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Sigmoid loss for language image pre -training,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.031708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.796296Z digest=sha256:088840a5c445315f299e5660413de6bbe3985cc3d1a6d10a8b1599eca1ea41c9

Observation 8cd9abe6-d824-4158-9990-45f7127f4079 · outbound

This paper cites MVBench: A comprehensive multi-modal video understanding benchmark,.

SV3.3B: A Sports Video Understanding Model for Action Recognition MVBench: A comprehensive multi-modal video understanding benchmark,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.026208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.798156Z digest=sha256:ec303fa871178670668afaef8e8a508448ead63e62117da7fdeaaef7d5961efb

Observation 09140373-a5f7-4ad0-adf8-01dfc9870cf4 · outbound

This paper cites Llama -vid: An image is worth 2 tokens in large language models,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Llama -vid: An image is worth 2 tokens in large language models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.020360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.799820Z digest=sha256:6b221cd6a99e0b4710bb86c742d5e41db5634b90afd2f4ff645e7c9426d7d48f

Observation 8613031e-91c7-49b0-9caf-e3f67a6c1907 · outbound

This paper cites Video-llama: An instruction-tuned audio- visual language model for video understanding,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Video-llama: An instruction-tuned audio- visual language model for video understanding,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.014390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.801681Z digest=sha256:e6ccd4b9eb148873d22bcbc712dc0c41593eb321e2a9de325c3140ee4dde9fc8

Observation b5588f34-ed5d-4267-aedc-2f891d43403c · outbound

This paper cites Temporal alignment networks for long-term video,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Temporal alignment networks for long-term video,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.008523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.803498Z digest=sha256:9b57d6095bb2ea2dab92f2847728798fc144f138fde5271089c1cc738b4a2f6d

Observation 94b8b072-21f6-433a-b120-af04d79c59e9 · outbound

This paper cites Multi-sentence grounding for long -term instructional video,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Multi-sentence grounding for long -term instructional video,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.002751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.805361Z digest=sha256:68dcd78e2af61c567e94848d070433f9dd2fe633f4f08a11d0f5351188637fe2

Observation b3295252-d1f3-4c72-8202-08f2a77f17de · outbound

This paper cites Panda- 70M: Captioning 70M videos with multiple cross -modality teachers,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Panda- 70M: Captioning 70M videos with multiple cross -modality teachers,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.996492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.807354Z digest=sha256:f0b9ac3990f3188419e1560eb94e3718a344d461ee2e3d067a5b1a83cda078dc

Observation a1e043d5-b327-4f58-97d0-c7c0245eea61 · outbound

This paper cites Vid2seq: Large -scale pretraining of a visual language model for dense video captioning,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Vid2seq: Large -scale pretraining of a visual language model for dense video captioning,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.990368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.809252Z digest=sha256:7e9de0515f217bf63fcf344a3cc480481899b9f9a5447fdd8957b95bfa725611

Observation 542eae72-54e7-47eb-a817-15cb45c4713e · outbound

This paper cites Streaming dense video captioning,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Streaming dense video captioning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.984046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.811016Z digest=sha256:db256dcf6772afb370efc9e26d1a304211163a413c082ce92d1cac4f8ec260af

Observation 8c1bb864-2069-46bd-ad3c-82db58b13626 · outbound

This paper cites Autoad: Movie description in context,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Autoad: Movie description in context,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.978020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.812820Z digest=sha256:6023d833811c9213ae8169c0e723c91947eb623a7a83a07e717c99ec5dbb240d

Observation 2287eebb-0f85-4024-bd33-826f84e35ab4 · outbound

This paper cites Autoad II: The sequel —who, when, and what in movie audio description,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Autoad II: The sequel —who, when, and what in movie audio description,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.971953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.814682Z digest=sha256:593a877a13c46ef6248d5f5744363d25cef264c5d8f6a9d9322b60782c959a34

Observation b3b75a8e-7dfc-4aaa-9db7-271e0fde46ba · outbound

This paper cites Autoad III: The prequel —back to the pixels,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Autoad III: The prequel —back to the pixels,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.965659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.816566Z digest=sha256:d55217c39d32c25096e4340cc4d926da6fae36d1bdf5c36526915f0718a8dfc9

Observation 225b382c-0702-4d58-aeff-48fbc85917e1 · outbound

This paper cites NSVA Subset: Basketball Video -Text Dataset,.

SV3.3B: A Sports Video Understanding Model for Action Recognition NSVA Subset: Basketball Video -Text Dataset,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.959004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T14:47:28.818411Z digest=sha256:d76d606684fc154b901bed8457fe0927d7a881cf51e22f836dc1c58819bac6ad

Pith citing papers

No inbound Pith citation observations are available.