Pith. sign in

Paper Citation Record · LEDGER

SV3.3B: A Sports Video Understanding Model for Action Recognition

As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2507.17844.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17844 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:47:28.818411Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact3
  • verified fuzzy34
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3859547-cc78-47d7-a1a8-e3fb88d0406d · outbound

This paper cites Review on wearable technology in sports: Concepts, challenges and opportunities,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Review on wearable technology in sports: Concepts, challenges and opportunities,

Reference 1

Resolution
verified exact
doi, observed 2026-08-06T14:47:28.838947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.735332Z digest=sha256:526783ee05253bf1b1851cb86d62897e3368c17cf60b0d586ceb92607659a91b

Observation 338bca84-3529-4456-bff5-dc53d909da87 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

SV3.3B: A Sports Video Understanding Model for Action Recognition Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.737949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.737949Z digest=sha256:4649f1385a1cdc5757cef6d69b0a46c4c38ee36aa69875e83ba84101d94719db

Observation 00d64863-bf51-4eca-aa1e-a18a39129839 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

SV3.3B: A Sports Video Understanding Model for Action Recognition LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.740370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.740370Z digest=sha256:356258d0f6e8e5cbd3e0ad08bfc9b7888772e6b2b85de14a6b2f76b9c7a50312

Observation 74989adb-e10d-4565-aa3a-e6a220763406 · outbound

This paper cites A path towards autonomous machine intelligence,.

SV3.3B: A Sports Video Understanding Model for Action Recognition A path towards autonomous machine intelligence,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.152382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.742825Z digest=sha256:e7b9800289506c8886abcda4fd6cba6af13a170c1a9b645f578803bf123b8d19

Observation 41494e1c-70a4-4022-8f07-fb841d1a0c80 · outbound

This paper cites Self-supervised learning from images with a joint- embedding predictive architecture,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Self-supervised learning from images with a joint- embedding predictive architecture,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.146394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.745119Z digest=sha256:25c48adc19706e83deb9363532b100015c131e58641180c054ddc9b008c747af

Observation a7bd5611-59ef-4565-afb7-7d959e6fd077 · outbound

This paper cites V -JEPA: Latent video prediction for visual representation learning,.

SV3.3B: A Sports Video Understanding Model for Action Recognition V -JEPA: Latent video prediction for visual representation learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.140706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.747272Z digest=sha256:25c48fa0a3f98b490879a6fa4652c7c6d56841b257f5f40357cfbfc5cd0014cb

Observation 643eb1ac-91a6-4c1e-bbc1-b16805fde231 · outbound

This paper cites UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity.

SV3.3B: A Sports Video Understanding Model for Action Recognition UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:47:28.940562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.749711Z digest=sha256:cf309b3634f532004c6e1081165f4840e477ed01f4c3773c2826ca96e9e91e68

Observation 029376f8-78cd-460f-9024-e915cb0e2d86 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

SV3.3B: A Sports Video Understanding Model for Action Recognition V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.752004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.752004Z digest=sha256:dfed6468eef24d5bf007cd754753f2d3fc977c9f60b0e8a8166bddfe30fa1e9d

Observation 354d65ae-8848-46b7-af8c-44abcdf96f4f · outbound

This paper cites Computer vision for sports: Current applications and research topics,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Computer vision for sports: Current applications and research topics,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.134924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.754962Z digest=sha256:565650d77a7829cefcc6240c7b3ce82a297e1340229ad1efd28433e87b1f3055

Observation dc57ca70-1e9c-48d0-8886-435839cb8a2a · outbound

This paper cites Soccernet-v2: A dataset and benchmarks for holistic understanding of broadcast soccer videos,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Soccernet-v2: A dataset and benchmarks for holistic understanding of broadcast soccer videos,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.129217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.757216Z digest=sha256:1a1e13e2d438dc2011176f55a21ab4e94e3cc544f8bb7e289a9b51c6fb3fa377

Observation 79293dc3-e8d5-4646-81d1-c6af34e48e68 · outbound

This paper cites Soccernet: A scalable dataset for action spotting in soccer videos,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Soccernet: A scalable dataset for action spotting in soccer videos,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.123627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.759138Z digest=sha256:77e1c80d7a25d1488fbf8c51c5a65dba4e1c189624aeaae1428363c2d3f728d0

Observation fac1c96e-771e-4fe3-bb4b-d68df025fad1 · outbound

This paper cites Fine-grained action recognition on a novel basketball dataset,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Fine-grained action recognition on a novel basketball dataset,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.118061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.761422Z digest=sha256:d0cebdac07f895b51e5fb104d3f157935f17e6878eae03e6b5364c56bf42942a

Observation 8fd68349-1c4d-498b-8c8e-be5fefeb4b81 · outbound

This paper cites Soccernet caption: Dense video captioning for soccer broadcasts commentaries,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Soccernet caption: Dense video captioning for soccer broadcasts commentaries,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.112178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.763299Z digest=sha256:13dcac47209d46d2787ca582850a4119e51eb2197124a740186b18f3c9fdd3a6

Observation f03400f5-4020-4fe7-bf01-b888895bf08e · outbound

This paper cites Sports video captioning via attentive motion representation and group relationship modeling,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Sports video captioning via attentive motion representation and group relationship modeling,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.106353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.765170Z digest=sha256:489bb2d697822acba2628a1f57e0750f3712b3d9d92c1dc46dffaca3a5c9ddda

Observation c99070ce-cc63-4f03-b9a5-68d95982b5fe · outbound

This paper cites Matchtime: Towards automatic soccer game commentary generation,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Matchtime: Towards automatic soccer game commentary generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.100713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.767048Z digest=sha256:c38412bcb6e73112d9d748ce304f9bbb64c17d7366b37cf389e1ab8b028a3770

Observation 1709e277-9f4e-4a28-9110-8115693abf3d · outbound

This paper cites Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark.

SV3.3B: A Sports Video Understanding Model for Action Recognition Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:47:28.925700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.768978Z digest=sha256:50331339f30e174d937f5ce607ac1753143b9abe974b9bd56844e5b86ddb62ef

Observation 7c6771fa-87fd-482a-9d0d-3a0ade95999c · outbound

This paper cites Fine-grained video captioning for sports narrative,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Fine-grained video captioning for sports narrative,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.095279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.771064Z digest=sha256:239f057de6b237c8bac24e5f20fe19885524b0239250bd51332e84318904e5c1

Observation 21d0f4e1-c2ea-4a6d-8575-36f75976bcd0 · outbound

This paper cites Finegym: A hierarchical video dataset for fine -grained action understanding,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Finegym: A hierarchical video dataset for fine -grained action understanding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.089788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.772938Z digest=sha256:5b38ef1dda9d7036c93b41f188e3443b601b9a188f5bd0445ca1d7348f127dd1

Observation 35700ec6-dfc7-473b-8984-aab3d5f17bf9 · outbound

This paper cites Finediving: A fine - grained dataset for procedure-aware action quality assessment,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Finediving: A fine - grained dataset for procedure-aware action quality assessment,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.083763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.775051Z digest=sha256:d295d46f039ac5df135be5c02a9444ccb0f15d8252088d971a2f2fba07d5261c

Observation 3d7bb799-2d82-4796-b9fd-f7259c4d6e1c · outbound

This paper cites Tacticai: An AI assistant for football tactics,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Tacticai: An AI assistant for football tactics,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.077698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.776961Z digest=sha256:adca5e66a6097fdea5f8cbd5eb0d2f6896637b5bd579306fdd6fad7c857bf1b1

Observation 96745f7f-0097-49c8-af54-2e418a36fc36 · outbound

This paper cites VARS: Video assistant referee system for automated soccer decision making from multiple views,.

SV3.3B: A Sports Video Understanding Model for Action Recognition VARS: Video assistant referee system for automated soccer decision making from multiple views,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.071515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.778846Z digest=sha256:5e65677955f6c4c0a7c717a696e040413d2107118285acadf0cdbde7acf8f2f6

Observation 18204d28-eecd-429a-89a7-660da1fa7d85 · outbound

This paper cites X -VARS: Introducing explainability in football refereeing with multimodal large language models,.

SV3.3B: A Sports Video Understanding Model for Action Recognition X -VARS: Introducing explainability in football refereeing with multimodal large language models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.065929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.780649Z digest=sha256:54a95e6fb7f279caa5ef53115894cf86f530425a4364a79410a67275c4204f73

Observation 19ac599a-1c98-46cc-b946-2036d5ea9ff0 · outbound

This paper cites Sports-QA: A large -scale video question answering benchmark for complex and professional sports,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Sports-QA: A large -scale video question answering benchmark for complex and professional sports,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.782472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.782472Z digest=sha256:9c1003bb6d55dd72c18dd41923145466828e99a5ae2c5ac886078c1d05c07add

Observation 1e851dbc-fdf2-424a-a7fb-05e4c106ecd6 · outbound

This paper cites SportQA: A benchmark for sports understanding in large language models,.

SV3.3B: A Sports Video Understanding Model for Action Recognition SportQA: A benchmark for sports understanding in large language models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.060399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.784339Z digest=sha256:ece1a353d42485f275a0701f1ad143db9d4ca39532cc36a3810db9cc32f05613

Observation 5c80b69f-ef7f-4e47-a145-9fe7c26de692 · outbound

This paper cites SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models.

SV3.3B: A Sports Video Understanding Model for Action Recognition SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:47:28.786186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:47:28.786186Z digest=sha256:bab2255f52d58a06c3b4d025e57699f452db3acd3ef28a9ca22556f75aeaf7a6

Observation 05827de8-aa86-48d4-b09d-4a6f57139e35 · outbound

This paper cites Flamingo: a visual language model for few -shot learning,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Flamingo: a visual language model for few -shot learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.054842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.788390Z digest=sha256:09b0479dc221c12461940378196e783466fd2380703477f8c2877743d00433e1

Observation 6326f665-380b-4582-b3d0-ee35e9d3a584 · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

SV3.3B: A Sports Video Understanding Model for Action Recognition BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.049117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.790361Z digest=sha256:84ac77ec86d8c32199fe627e421b132deeb239c19210b5edc83c9043b06e895c

Observation 27132ff3-fe73-447b-85dc-23a36b03cca7 · outbound

This paper cites BLIP -2: Bootstrapping language- image pre -training with frozen image encoders and large language models,.

SV3.3B: A Sports Video Understanding Model for Action Recognition BLIP -2: Bootstrapping language- image pre -training with frozen image encoders and large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.043418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.792353Z digest=sha256:f66b42d8dec3b568dcb156c81315247efca389b698df65239458245b6fa7b105

Observation 1e314612-64b6-4c4f-bfae-a239a2f1a818 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Learning transferable visual models from natural language supervision,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.037574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.794410Z digest=sha256:6c46d6da41591528f3d6c243aa82c7ffd735aef4aaf01c8267190d3003ef6877

Observation c3e888dd-8ada-490a-b844-0a4a727806c1 · outbound

This paper cites Sigmoid loss for language image pre -training,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Sigmoid loss for language image pre -training,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.031708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.796296Z digest=sha256:f2596741ebd5eb84f91e5fe76cde31306e564ec1b1d80afa46923b1f530dd62f

Observation 8cd9abe6-d824-4158-9990-45f7127f4079 · outbound

This paper cites MVBench: A comprehensive multi-modal video understanding benchmark,.

SV3.3B: A Sports Video Understanding Model for Action Recognition MVBench: A comprehensive multi-modal video understanding benchmark,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.026208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.798156Z digest=sha256:86231f229f52a8a5f76aed2667b9274781e99d3c3d5273b6f717ae2a956a6ab3

Observation 09140373-a5f7-4ad0-adf8-01dfc9870cf4 · outbound

This paper cites Llama -vid: An image is worth 2 tokens in large language models,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Llama -vid: An image is worth 2 tokens in large language models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.020360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.799820Z digest=sha256:71dfd334e3f11b2bb62bf9a275d2e160847c887975ed4b70bce2511bda54c4ac

Observation 8613031e-91c7-49b0-9caf-e3f67a6c1907 · outbound

This paper cites Video-llama: An instruction-tuned audio- visual language model for video understanding,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Video-llama: An instruction-tuned audio- visual language model for video understanding,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.014390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.801681Z digest=sha256:9f6f8288910e0f11d7168d2e1df1ccfd890522136d3af6cd8f374294843c9019

Observation b5588f34-ed5d-4267-aedc-2f891d43403c · outbound

This paper cites Temporal alignment networks for long-term video,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Temporal alignment networks for long-term video,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.008523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.803498Z digest=sha256:3ba656ab103d5f128632615c3acdfedcb378205b3e316afe944807542cfa0b87

Observation 94b8b072-21f6-433a-b120-af04d79c59e9 · outbound

This paper cites Multi-sentence grounding for long -term instructional video,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Multi-sentence grounding for long -term instructional video,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:29.002751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.805361Z digest=sha256:dc9b4f79b9e09ed020833bdab56dcc050074eb5f1e839a37029408faf023f1ba

Observation b3295252-d1f3-4c72-8202-08f2a77f17de · outbound

This paper cites Panda- 70M: Captioning 70M videos with multiple cross -modality teachers,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Panda- 70M: Captioning 70M videos with multiple cross -modality teachers,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.996492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.807354Z digest=sha256:fe98303b82fc90e6e05089ca5ac861e88664565d319af6630c6dd3a7e3da4546

Observation a1e043d5-b327-4f58-97d0-c7c0245eea61 · outbound

This paper cites Vid2seq: Large -scale pretraining of a visual language model for dense video captioning,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Vid2seq: Large -scale pretraining of a visual language model for dense video captioning,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.990368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.809252Z digest=sha256:e6090d0eea27128927f66247c6e31c7f90c9f0da568c20db5333616326a1fb7c

Observation 542eae72-54e7-47eb-a817-15cb45c4713e · outbound

This paper cites Streaming dense video captioning,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Streaming dense video captioning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.984046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.811016Z digest=sha256:025ac586f8ef95dfef785580c363ed5d89cc787f5f8aa3cab883755ad92a805b

Observation 8c1bb864-2069-46bd-ad3c-82db58b13626 · outbound

This paper cites Autoad: Movie description in context,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Autoad: Movie description in context,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.978020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.812820Z digest=sha256:2411d9fd443c1044958a4e0a26ee57afa23a8bdf057a1822f5839874e32183b0

Observation 2287eebb-0f85-4024-bd33-826f84e35ab4 · outbound

This paper cites Autoad II: The sequel —who, when, and what in movie audio description,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Autoad II: The sequel —who, when, and what in movie audio description,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.971953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.814682Z digest=sha256:c50e04b3ada76a728cf33af33a48343e768a329b371710fc427239c9d6aaa05b

Observation b3b75a8e-7dfc-4aaa-9db7-271e0fde46ba · outbound

This paper cites Autoad III: The prequel —back to the pixels,.

SV3.3B: A Sports Video Understanding Model for Action Recognition Autoad III: The prequel —back to the pixels,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.965659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.816566Z digest=sha256:5710cd454c813e5e4fc07b96bbfaf02cf6129fe75182af891049f2b93eeaebb3

Observation 225b382c-0702-4d58-aeff-48fbc85917e1 · outbound

This paper cites NSVA Subset: Basketball Video -Text Dataset,.

SV3.3B: A Sports Video Understanding Model for Action Recognition NSVA Subset: Basketball Video -Text Dataset,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:47:28.959004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:47:28.818411Z digest=sha256:c58357db72dc017ab1da6e6d125cc41d10fe31a06f72642af4339c9fd099e949

Pith citing papers

No inbound Pith citation observations are available.