Pith. sign in

Paper Citation Record · LEDGER

SCBench: A Sports Commentary Benchmark for Video LLMs

As of 13 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2412.17637.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17637 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:24:17.842065Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:10:07.699225Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T23:25:51.674329Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b6535f5-61bb-4049-acd1-2fe75a8907d3 · outbound

This paper cites write newline.

SCBench: A Sports Commentary Benchmark for Video LLMs write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.683409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.683409Z digest=sha256:f71821f45b0736a3df3adebc486630e32b03ad00cccd2b97d6581af5c57d61fb

Observation e55d5ec1-caed-4f9d-bfec-fc3518af17e9 · outbound

This paper cites Spice: Semantic propositional image caption evaluation, 2016.

SCBench: A Sports Commentary Benchmark for Video LLMs Spice: Semantic propositional image caption evaluation, 2016

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.277345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.687975Z digest=sha256:ff45fa14b96f0e81003c342ba1fe4aaec03baf77d8d3b18331fe279908aac234

Observation 92643bb3-53a7-4370-82e1-8eb769cc1569 · outbound

This paper cites METEOR : An automatic metric for MT evaluation with improved correlation with human judgments.

SCBench: A Sports Commentary Benchmark for Video LLMs METEOR : An automatic metric for MT evaluation with improved correlation with human judgments

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.267206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.691725Z digest=sha256:258e78674f4ce81c2406e688f5558fe62ec8d3ee27909635850caac21494d901

Observation 81aa4ebc-e894-4060-853e-1f87748c46ba · outbound

This paper cites P2anet: A dataset and benchmark for dense action detection from table tennis match broadcasting videos, 2024.

SCBench: A Sports Commentary Benchmark for Video LLMs P2anet: A dataset and benchmark for dense action detection from table tennis match broadcasting videos, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.257020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.695503Z digest=sha256:4aeb000bc1abdf9a72bb22c9373402e424fef9e51be598bd8abc3491a1855f48

Observation d38a36f9-cdc7-4668-81bc-d2bff5b5d90e · outbound

This paper cites an unresolved cited work.

SCBench: A Sports Commentary Benchmark for Video LLMs Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.699465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.699465Z digest=sha256:89239a32b29615a6b4b65df7530f6e0bfe8f55f74b0c4fc2e1367f14b8864fee

Observation 01fb1590-e1ab-4220-ae56-24e768fe32bf · outbound

This paper cites Autoeval-video: An automatic benchmark for assessing large vision language models in open-ended video question answering, 2024 a.

SCBench: A Sports Commentary Benchmark for Video LLMs Autoeval-video: An automatic benchmark for assessing large vision language models in open-ended video question answering, 2024 a

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.240253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.703060Z digest=sha256:850e83a05e2eae8e103f845e3f63fe653a7f986a926ddcccfe94b4d6021c436a

Observation 4885a10a-2b21-4003-bcdf-004717ab064e · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks, 2024 b.

SCBench: A Sports Commentary Benchmark for Video LLMs Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks, 2024 b

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.229784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.706887Z digest=sha256:fe255ddc6c4b45f7dbce0f893df506bc27b0bb4f626346e59ccfedc42e5bd1a8

Observation 49590249-4420-459c-86fb-112a787c439a · outbound

This paper cites Sports re-id: Improving re-identification of players in broadcast videos of team sports, 2022.

SCBench: A Sports Commentary Benchmark for Video LLMs Sports re-id: Improving re-identification of players in broadcast videos of team sports, 2022

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.218808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.710525Z digest=sha256:14bf309e6623195cd8e6914961793ebeecdca4b77f5ac833d7ef87b39ff34552

Observation a09997db-526f-4884-9426-88cfca800994 · outbound

This paper cites Seikavandi, Jacob V.

SCBench: A Sports Commentary Benchmark for Video LLMs Seikavandi, Jacob V

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.208594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.713845Z digest=sha256:e1d24b7394dd7c4c2b9acb4af4be023306d91bcbe00e0f11f7fae7fd62661595

Observation bd052060-cac3-4270-9fd0-2fc4fa6f6edc · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding, 2024.

SCBench: A Sports Commentary Benchmark for Video LLMs Mmbench-video: A long-form multi-shot benchmark for holistic video understanding, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.199517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.717112Z digest=sha256:435e610bd04a95f1ed77de3e7a65baa6a13af5dfaec434a89ed8e64e2e1ce649

Observation b6b813a9-e419-4a29-be55-548f911c8fe4 · outbound

This paper cites an unresolved cited work.

SCBench: A Sports Commentary Benchmark for Video LLMs Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:24:18.190434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.720502Z digest=sha256:ea5de173314df4b96f265dca654e4c19ea7cc3fcf182a52598c2ad875a57082d

Observation 201e4a8c-2feb-4498-8750-0a5980363c45 · outbound

This paper cites Chatpose: Chatting about 3d human pose.

SCBench: A Sports Commentary Benchmark for Video LLMs Chatpose: Chatting about 3d human pose

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.180323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.723725Z digest=sha256:541f626942c3b57aac77084b9c737185c29d24a13bddb319ba6c6cb8743596a9

Observation ee94795f-eb39-4741-ba3c-83ad9bbe4e93 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis, 2024.

SCBench: A Sports Commentary Benchmark for Video LLMs Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.731047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.731047Z digest=sha256:3bba017afd1624618355586b50cc8eba3421f1ec7b940cd08c36ee860770fdc1

Observation f753cd98-bff0-460d-ac1f-22c9961d08f5 · outbound

This paper cites Mini-internvl: A flexible-transfer pocket multimodal model with 5.

SCBench: A Sports Commentary Benchmark for Video LLMs Mini-internvl: A flexible-transfer pocket multimodal model with 5

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.163828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.734832Z digest=sha256:f78397cf8f420ad18f0ddd5300be1c3c47525375f2c84d3e7207f31a2ac7ef23

Observation 5dab9f72-2add-4922-81fc-e92616989e7d · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering, 2017.

SCBench: A Sports Commentary Benchmark for Video LLMs Tgif-qa: Toward spatio-temporal reasoning in visual question answering, 2017

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.153707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.738244Z digest=sha256:709a731841445ab3eee84f8c04f1fc9d442505bf3a1b4b04ff7843d1aa327c14

Observation f367ceff-082e-4348-8159-16ea306ededa · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding, 2024.

SCBench: A Sports Commentary Benchmark for Video LLMs Chat-univi: Unified visual representation empowers large language models with image and video understanding, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.143230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.741346Z digest=sha256:54ec65a610b046b1134c410c7c622c1b5edbe3a8f512cef2fa0b66385fef9601

Observation 7fa9b383-63be-4c0d-96d7-e83b3e6a0fb5 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks, 2021.

SCBench: A Sports Commentary Benchmark for Video LLMs Retrieval-augmented generation for knowledge-intensive nlp tasks, 2021

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.744141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.744141Z digest=sha256:9a2dabc52d83552c4a28f6eedf27768a7592beb39254dd0a2f5356ef71881481

Observation d0c9d076-ab13-4cf7-ba77-886c8f7d102d · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark, 2024.

SCBench: A Sports Commentary Benchmark for Video LLMs Mvbench: A comprehensive multi-modal video understanding benchmark, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.127754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.746975Z digest=sha256:c15b36b68e2ea60e286f14b945d4486ae47cdb911070ffb46e5499dd9a76a0d9

Observation 63c2705e-2a9b-4b8f-94ca-53e371c0145a · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection, 2023.

SCBench: A Sports Commentary Benchmark for Video LLMs Video-llava: Learning united visual representation by alignment before projection, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.749606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.749606Z digest=sha256:c2b07a5d4ce673cd5dc6ff118231c82f1fa4f73cf3e31994f39dc3b7e97d8a75

Observation a8cb6f5d-a8c1-4014-abe8-39cf5a8b1c5f · outbound

This paper cites ROUGE : A package for automatic evaluation of summaries.

SCBench: A Sports Commentary Benchmark for Video LLMs ROUGE : A package for automatic evaluation of summaries

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.752192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.752192Z digest=sha256:902fbef8afa90f3e7cae2a0daf7fea50b071c80bbe29727675f5a9e8fd1a5dcb

Observation 2dde1edb-3b92-4512-b7ae-b2c24263de32 · outbound

This paper cites Vila: On pre-training for visual language models, 2024.

SCBench: A Sports Commentary Benchmark for Video LLMs Vila: On pre-training for visual language models, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.105842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.755116Z digest=sha256:e0ccec3e0515a8f715d5a6e4bf48a17e4bda11e53d537c66e5288fd0cadfa012

Observation 272a74c3-c035-45fa-b6d7-4b5bcc9a23e1 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024 a.

SCBench: A Sports Commentary Benchmark for Video LLMs Llava-next: Improved reasoning, ocr, and world knowledge, 2024 a

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.094806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.757693Z digest=sha256:5bda7608dbcef6cedc3ca6043ee2718c9cddb8918dcd1dd355551d5bb1db9ddf

Observation 5f312c6c-c424-4e8b-9f60-e60a9de35532 · outbound

This paper cites Kangaroo: A powerful video-language model supporting long-context video input, 2024 b.

SCBench: A Sports Commentary Benchmark for Video LLMs Kangaroo: A powerful video-language model supporting long-context video input, 2024 b

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.083902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.760420Z digest=sha256:b3d5d3dc270a834b188590ae49012f769ffb94c699065b028f9e25bd694eb73d

Observation d441747c-01e2-428a-a79d-1b96fcec39e9 · outbound

This paper cites Fineaction: A fine-grained video dataset for temporal action localization.

SCBench: A Sports Commentary Benchmark for Video LLMs Fineaction: A fine-grained video dataset for temporal action localization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.073076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.763437Z digest=sha256:5af526ec4843717f7e1b0f582ecc14ca4c5901f65b75f0c57801ecf8ab24b0bb

Observation a5b2fb07-83a1-4be2-8113-58add5c6f016 · outbound

This paper cites Tempcompass: Do video llms really understand videos?, 2024 c.

SCBench: A Sports Commentary Benchmark for Video LLMs Tempcompass: Do video llms really understand videos?, 2024 c

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.063005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.767152Z digest=sha256:962fa9677105b344a11cc3966146437d583b7124d014656b448bbbbf8d82a66f

Observation 495a771b-3799-45de-8bc8-c6f62180a8e1 · outbound

This paper cites Cross-block fine-grained semantic cascade for skeleton-based sports action recognition, 2024 d.

SCBench: A Sports Commentary Benchmark for Video LLMs Cross-block fine-grained semantic cascade for skeleton-based sports action recognition, 2024 d

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.053087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.770554Z digest=sha256:14023478c3bc0f2add6b4f813d21b06687a3930c876c4926cf99210d260ec550

Observation 1bf485d7-49e8-4bf9-87c5-ea611345cbe1 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding, 2023.

SCBench: A Sports Commentary Benchmark for Video LLMs Egoschema: A diagnostic benchmark for very long-form video language understanding, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.773955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.773955Z digest=sha256:4d9bad91b919821e36611e424f7ec98389d397137cea0691b9ca090ee0727789

Observation 914eb97b-14ed-45ae-b0f7-c9c1f81aa64d · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

SCBench: A Sports Commentary Benchmark for Video LLMs Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.777547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.777547Z digest=sha256:9a3d7d8003cba3ed8c7af9a0cfa4d78c97ea82e46566e9e5ca93ce22e9ca9be7

Observation 790bf743-95bc-4de3-be39-e8ff36a81b80 · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

SCBench: A Sports Commentary Benchmark for Video LLMs Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.780895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.780895Z digest=sha256:978052ba39d0701a887206d08536c50d0a4c0396360e0a9b5de67f30dc9e4bf9

Observation 00e92d4e-b813-4d32-9c87-1de79dd35e53 · outbound

This paper cites Video-bench: A comprehensive benchmark and toolkit for evaluating video-based large language models, 2023 b.

SCBench: A Sports Commentary Benchmark for Video LLMs Video-bench: A comprehensive benchmark and toolkit for evaluating video-based large language models, 2023 b

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.031890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.784738Z digest=sha256:9a9fc3925b7a507992d15d70af39d6d147a339aaf9515f2cffd21b532a045835

Observation 318f1244-9d23-4391-96f8-c4276da2d98e · outbound

This paper cites B leu: a method for automatic evaluation of machine translation.

SCBench: A Sports Commentary Benchmark for Video LLMs B leu: a method for automatic evaluation of machine translation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.020253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.787974Z digest=sha256:0ecb144c9db374dccb12098b5aeb2c7d31918e377d461d67580c24720b54ae07

Observation 1e62d750-37c1-48f2-a52e-ff97f2b59ddd · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models, 2023.

SCBench: A Sports Commentary Benchmark for Video LLMs Perception test: A diagnostic benchmark for multimodal video models, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:18.009190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.791466Z digest=sha256:54809b895f3257660195934c4aa1cbc0262f43f404ea3545ebfe2b6d387af6a5

Observation 7ec357c2-f255-4e71-b76b-c16c38beb508 · outbound

This paper cites A survey of video datasets for grounded event understanding.

SCBench: A Sports Commentary Benchmark for Video LLMs A survey of video datasets for grounded event understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:17.996400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.794736Z digest=sha256:f1d5bceb8d475bbb60aa27d629a9c5bddf3da0b780af7f9b2ca104a6bfab8d4e

Observation 194d3318-23d9-447d-b2c2-4ee768d1fea6 · outbound

This paper cites Finegym: A hierarchical video dataset for fine-grained action understanding, 2020.

SCBench: A Sports Commentary Benchmark for Video LLMs Finegym: A hierarchical video dataset for fine-grained action understanding, 2020

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.797956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.797956Z digest=sha256:55a16e7e27ac8500b56ff7e61e7c9f8b3997d8e97058cdde3937c34c7bfd42a1

Observation 1ad71774-beb3-43d1-9c46-8f1b44452bf0 · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning, 2024.

SCBench: A Sports Commentary Benchmark for Video LLMs Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.801306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.801306Z digest=sha256:8744ada247554cadfc96ea4e17b8e31b046c30c402ffa2bd13e92f75247c75b9

Observation 22ceb1ea-6e79-4b8f-a840-4c7cec682770 · outbound

This paper cites Playertv: Advanced player tracking and identification for automatic soccer highlight clips, 2024.

SCBench: A Sports Commentary Benchmark for Video LLMs Playertv: Advanced player tracking and identification for automatic soccer highlight clips, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:17.973119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.804672Z digest=sha256:4a6a107137dc8e5c26798a08852ba35120efe1212197f09cbf9ac1b87bb4e655

Observation 0e4d8d06-c469-4c5d-b539-8f68e0a4bde8 · outbound

This paper cites Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy.

SCBench: A Sports Commentary Benchmark for Video LLMs Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:17.962914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.808227Z digest=sha256:7a3aa279c808789aae6dbf834f5c9e39502b0a4e119677859c71023c2e0acd76

Observation 712f6b79-4947-472c-a454-94d31e375393 · outbound

This paper cites CIDEr: Consensus-based Image Description Evaluation.

SCBench: A Sports Commentary Benchmark for Video LLMs CIDEr: Consensus-based Image Description Evaluation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.811515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.811515Z digest=sha256:13f00b1d2ff46a036d94eeffa061a6e9d6da2217409e3d93c9c140a0c1a240b6

Observation 48281dc9-9464-4065-9549-cac8cc4408b2 · outbound

This paper cites Lvbench: An extreme long video understanding benchmark, 2024.

SCBench: A Sports Commentary Benchmark for Video LLMs Lvbench: An extreme long video understanding benchmark, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.815018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.815018Z digest=sha256:acb0df5e0e3ec74c5a0678283ca6b0252efd64db96ec1058d6a5b4a3319d030d

Observation 9138a27a-d2d9-4035-8d99-52e1913656b4 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models, 2023.

SCBench: A Sports Commentary Benchmark for Video LLMs Chain-of-thought prompting elicits reasoning in large language models, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.818259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.818259Z digest=sha256:55349c218ea078ba3ff85048194ce9db70a2012fb8a8b5668eac817bb799fcc5

Observation 2b9fd42c-5dab-4c68-852f-ececbfabeb6a · outbound

This paper cites Sportshhi: A dataset for human-human interaction detection in sports videos, 2024.

SCBench: A Sports Commentary Benchmark for Video LLMs Sportshhi: A dataset for human-human interaction detection in sports videos, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:17.941986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.821236Z digest=sha256:a3451c930da83913212067add3d6bd0c41c36838c0fe8e8134760f3b17ef33aa

Observation 191afc41-22d3-4cb1-8203-7215a9c1bc65 · outbound

This paper cites Next-qa:next phase of question-answering to explaining temporal actions.

SCBench: A Sports Commentary Benchmark for Video LLMs Next-qa:next phase of question-answering to explaining temporal actions

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:17.932170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.824496Z digest=sha256:c2dfac0bde6e3b061c7982ce836080db846d8ece235a63a3b2c198dba68b7a4a

Observation 5a12612e-03b0-4cd0-92ba-11cf5040ecb1 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

SCBench: A Sports Commentary Benchmark for Video LLMs Video question answering via gradually refined attention over appearance and motion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:17.923406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.827835Z digest=sha256:d5acb2e3704408cff0cd5fe302cd0a1051b180647d31ed77af411c38dd309f51

Observation 6ece9f76-4af6-42f9-8fe9-ec73996239ee · outbound

This paper cites Youku-mplug: A 10 million large-scale chinese video-language dataset for pre-training and benchmarks, 2023.

SCBench: A Sports Commentary Benchmark for Video LLMs Youku-mplug: A 10 million large-scale chinese video-language dataset for pre-training and benchmarks, 2023

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:24:17.913531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T05:24:17.830956Z digest=sha256:0a92bd3b98ea01bf78b8582d5eb280bc6ec2c198c1e1139782e91a7ef55c9db6

Observation 2f72312a-1c0c-4b73-93f4-b98ac8a319bb · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering, 2019.

SCBench: A Sports Commentary Benchmark for Video LLMs Activitynet-qa: A dataset for understanding complex web videos via question answering, 2019

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.834360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.834360Z digest=sha256:8af6368d167b190e2816c1e0183133a5151c233c314c2bbedd4fc233ea767a08

Observation ec67df3c-0e60-44f4-b323-567b1480ea59 · outbound

This paper cites Long context transfer from language to vision, 2024.

SCBench: A Sports Commentary Benchmark for Video LLMs Long context transfer from language to vision, 2024

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.837559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.837559Z digest=sha256:1174e3d77800ab9d134170a186083e38bae4251688f653dd41e649a408dc1d2d

Observation f966328a-0266-435b-8b64-9392d51e1834 · outbound

This paper cites A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming.

SCBench: A Sports Commentary Benchmark for Video LLMs A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:17.842065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:24:17.842065Z digest=sha256:ac272b27082f22629bfc3a6853e0c15cb9c94dcf9abcf601ab60eb20eb0f804e

Pith citing papers

Observation 9412f7ea-5c9d-4aa8-b13c-2020277ff3bd · inbound

BoxComm: Benchmarking Category-Aware Commentary Generation and Narration Rhythm in Boxing cites this paper.

BoxComm: Benchmarking Category-Aware Commentary Generation and Narration Rhythm in Boxing SCBench: A Sports Commentary Benchmark for Video LLMs

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:25:51.682028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T19:10:07.699225Z digest=sha256:67aee3253341eb916b0685cbf951cf190939fc1b9957b5b3399900341f1064be

Observation 5ba8e445-dd51-4bbc-964e-2b2d7c78c243 · inbound

RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees cites this paper.

RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees SCBench: A Sports Commentary Benchmark for Video LLMs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:48:02.282471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T08:38:27.081358Z digest=sha256:20a7accb7b333c58760a21ef43a9c1d9dc777550170ab20bc6769099d30ed600