Pith. sign in

Paper Citation Record · LEDGER

TextVidBench: A Benchmark for Long Video Scene Text Understanding

As of 8 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2506.04983.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04983 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:35:18.850729Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d84e5b36-0132-433a-9ef7-73a3368f5491 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.712205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.712205Z digest=sha256:5c25ef875c60f0224985da08955ceb169cf8e37636d1c678a6e111ea9dee5f79

Observation 6bf83b2e-a48c-4215-a0af-067dc447f7c5 · outbound

This paper cites Qwen Technical Report.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.715556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.715556Z digest=sha256:40a91ea5cb654a8a3d4d5afab0b0a0a1c2125c2e5183f412b6e36e4ad43f4b46

Observation cc093305-7e84-45dc-8f87-67444dbb058b · outbound

This paper cites Qwen2.5-VL Technical Report.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.718987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.718987Z digest=sha256:151cdb41935bf4b3b337fa4a2c117a1e2279aea40342c333e88024a0ee6880c2

Observation fb90f1e1-b9bf-4b35-a329-2e38dbdc1f89 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.722304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.722304Z digest=sha256:4460022f0ae0a06e981ec0ad1a34e2628bf5e00be001b1ed54befc98b7ed36f4

Observation 692e5820-f16e-441f-8b78-987d9d3936dc · outbound

This paper cites VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding.

TextVidBench: A Benchmark for Long Video Scene Text Understanding VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.725324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.725324Z digest=sha256:92d461d90555edeacd3f6529f1b1c87fa4fb85b3e8f2cd7de1fcb94fd7c195e6

Observation f1f8aefa-6e75-4ddb-a7f4-b4f29d55b1bb · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.728423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.728423Z digest=sha256:5db97b0f945ed1b3e2460fbba5b73b4622d7cdcd66a7e0b076503a53006b1efb

Observation 8323916c-5354-4581-80d9-9d0e94cfa7c0 · outbound

This paper cites LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens.

TextVidBench: A Benchmark for Long Video Scene Text Understanding LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.731567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.731567Z digest=sha256:f212ac8f804f66f81d94a255ea66b68334a1a5418a497bad0b79af374583e310

Observation 1f484c36-ab46-4e9d-bd2f-90875d639028 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:35:19.241090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:35:18.734490Z digest=sha256:ecdda2c31e71d299e56efcf316dd6eeded34491681d50788b8e39f9ff3c42fe7

Observation 70e534b0-ede5-4740-bcd7-10d40ea9d9d8 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.737180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.737180Z digest=sha256:a3dfdb51f287c4f13db677006edfe8e7507de860a4c9a2843dbc4073892a9919

Observation 86d84e25-8b1c-4e00-a23d-28b0c49b1487 · outbound

This paper cites The Llama 3 Herd of Models.

TextVidBench: A Benchmark for Long Video Scene Text Understanding The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.740049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.740049Z digest=sha256:481eb45bc628c7bcaef07975a9320e7daa118c58e30f341df23f61124d676318

Observation 44120cab-638a-4f32-8b1d-b6625c4b4954 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:35:19.224862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:35:18.743167Z digest=sha256:b739e8de9b7ecc13b2fb0a04710be2010763d3d33699e0a6603c6e3acea2311d

Observation 4e0ac14e-64e5-454f-a330-5945d7a76fbb · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:35:19.216232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:35:18.745903Z digest=sha256:f5c4d72dcfecee145186cd342f094b75b9afb77c4fff1c66632e05d797f24998

Observation 2edea8af-6358-4836-b67b-8fd7451826b7 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.748572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.748572Z digest=sha256:53da9f4f5273b11ff1cb30a6bf3ea51d6e3c44291da69787a603787bceae9d8b

Observation 6c063a6d-30a1-425b-9ce7-f98e6fc2a168 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:35:19.199858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:35:18.751221Z digest=sha256:f401fe5f5f4c5cf9f0bef26f0c6ad89e75c72481756eddce4d2bf60e40c3c4c6

Observation 27d0473b-a613-48cc-8957-ab1f2a2084d8 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

TextVidBench: A Benchmark for Long Video Scene Text Understanding LoRA: Low-Rank Adaptation of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.754501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.754501Z digest=sha256:b5199cb86f98f7b8c023bec53b74b030549bec685f03e6f067db49066b0fdefc

Observation 5580c4de-fc44-4d8e-afd0-0d20ffad5f23 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:35:19.191036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:35:18.757346Z digest=sha256:6c29765a40f95638d05ed65f4cded94bc5f17af8c1df315d24fec30b5e1d8167

Observation a2f35c1a-7f6e-4776-8b4f-46814d970c23 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:35:19.181696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:35:18.760395Z digest=sha256:234728cf68b1f451d5a5417f15bab171cf6a644e41af7d1d03e5d62d72660e3c

Observation 649a6979-1daf-43ba-a276-18031daef767 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.763204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.763204Z digest=sha256:2e37ae395d7f588bea74ba209fc0f2697b246531ece5fcd01d7425e87d813f76

Observation 89ef1f9b-86be-4bea-a632-c6eb69d17a44 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:35:19.166753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:35:18.766069Z digest=sha256:4a6b9cc00ac28e2541dc5449a48cf3069f59c59da10244c4d1742c777b979753

Observation ddda0886-6765-46c2-957d-95ba1f9af3bb · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

TextVidBench: A Benchmark for Long Video Scene Text Understanding TVQA: Localized, Compositional Video Question Answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.768653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.768653Z digest=sha256:1de3a19f2540354eb2f5e90bd1e6d8b2e35abe2d14020d1e02935b1e5a10b8d5

Observation db414071-d74d-4059-bbd8-d526f99d6b4e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

TextVidBench: A Benchmark for Long Video Scene Text Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.771710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.771710Z digest=sha256:0d8f77e0801f01ce25f99e7175000fd0bd372a0f01b9b556bd1b1a6c70d52b2a

Observation ae6379de-b6d5-48c2-b106-3f10132f757d · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

TextVidBench: A Benchmark for Long Video Scene Text Understanding SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.774968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.774968Z digest=sha256:9b3d8da016120f34cb4db1becc51734bd2786373c567257020fac60cc6d34605

Observation c62b7654-d564-4af8-bb92-a58225cfd4cb · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.778106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.778106Z digest=sha256:e64c96621513c64e9de70779a7092e9d0ea56987a05c3b4b4b4a3941c84e2212

Observation e8216c7b-6030-4433-b699-c252e2a34def · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:35:19.150608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:35:18.780925Z digest=sha256:0ff3019b7b73e092cc2f184c533172f6449c55577473ae47039e4a6764133f2c

Observation 835cf121-144c-4b1d-97ca-2f4a8bb14475 · outbound

This paper cites GPT-4 Technical Report.

TextVidBench: A Benchmark for Long Video Scene Text Understanding GPT-4 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.783906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.783906Z digest=sha256:e0648ff6167d3902cf4a2a0d74a3b4b278067b82d7e59134c4c746ff5b572e11

Observation a0549b04-a468-4022-b661-1bde09cc35bb · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.786787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.786787Z digest=sha256:e39abb60cdb43c5df9dfcd7490e62389072b27609efda6f28114d9434696a0e9

Observation db2feae8-49ac-463b-abbd-36bc4592cc42 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.789673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.789673Z digest=sha256:e3095618267a0e5736c4305c5f5cc9de62b220d0101a6680cb6392e579f20900

Observation 96e8a43c-5a05-4d4e-a631-db2b19611c37 · outbound

This paper cites Reading Between the Lanes: Text VideoQA on the Road.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Reading Between the Lanes: Text VideoQA on the Road

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.792524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.792524Z digest=sha256:1db074b9373ac3bf29d501d0c71eb245a245ebac7cb1325c4bc183c3d9e781ed

Observation 05c341f4-66da-4e22-a021-331cb3888233 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.795639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.795639Z digest=sha256:ef5772a3c75b4acff0942f8188c88b65ee16116a68948446c2af660d8d6dfd7c

Observation 9c75f9be-642d-4630-9ee1-dfe1a9b78f60 · outbound

This paper cites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images.

TextVidBench: A Benchmark for Long Video Scene Text Understanding COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.799068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.799068Z digest=sha256:388e72841585dc929eacf487dcaea0a18703b13ee4d2138c5c700b5ad87a0876

Observation 680c6ada-9cf1-4cf4-91f0-e2489beda7c4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.802966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.802966Z digest=sha256:cb916fb27d990b636c7db58dd18ee7f8060672c654a96ccbe8781674a8ee53fa

Observation 360ab4bf-a420-4bd0-96b6-52d94ba8acea · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:35:19.127122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:35:18.805735Z digest=sha256:8865c080ff94a773ad9f21abb20bc35c287b3425e76703291b6836e0b6a53e4a

Observation db80d668-d646-4277-a99d-17adbb856bde · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.808585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.808585Z digest=sha256:57bc780332e2668821882b3d163b87bee314b49aad48f578193ded89e4930f6f

Observation 3d99a1c4-a845-44b1-b4af-e6847fd32f22 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.812430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.812430Z digest=sha256:2b2fa2ebbdf92ba18c0ed77437acbf279963bcc43e30aafb36c92fb70d6addc8

Observation 8cf51cf3-8f4e-4f4b-a8ee-f94c685ffb68 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:35:19.104244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:35:18.815479Z digest=sha256:14a85f08ff8fc1957bc601630d0911aa55f4de90469d93466bd89aa42dfa9ea1

Observation 90256300-b541-4ad1-a4c1-06973ecb44f1 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

TextVidBench: A Benchmark for Long Video Scene Text Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.818281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.818281Z digest=sha256:9ee4c3047ce401e3c56936e036897277c0ca0c8327edb3a8c7130d7f4bc854e6

Observation d95a24a7-32b9-44fe-bb84-81e1294b2af8 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

TextVidBench: A Benchmark for Long Video Scene Text Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.821165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.821165Z digest=sha256:a1feea303e049c69082edd2fb62bf906d9a4b1f43571ba01e450199bf99dc5d5

Observation 1f12a3a1-cb80-4eb9-9e94-ba6a2310e420 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.824624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.824624Z digest=sha256:2296b85fa0fdbef34b73461a098e49b37f7973110974b1ec14426cea696da29f

Observation a953a613-48c1-4045-903e-2287e95efbb0 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.827589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.827589Z digest=sha256:c46b81561a407c1ed643a7c12c03358713cab5fd9228bfce97bca6a95e0f6c28

Observation 6eae8b80-451a-4cc6-ab7d-9268c2002e87 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:35:19.080995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:35:18.830401Z digest=sha256:24da70e13bf7dd703ea8d14cdde0a32750dc05470685cf5fb975317c663954ac

Observation 414c1749-d341-4e53-997f-cd0c64927358 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Sigmoid Loss for Language Image Pre-Training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.833261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.833261Z digest=sha256:8031dcc8404e7078be94b8555eac4db31e12b2c3efc6b554f9c30f81d45548d4

Observation b10c4db0-dc68-4031-b0c2-9d98e407ecf4 · outbound

This paper cites Long Context Transfer from Language to Vision.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Long Context Transfer from Language to Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.836228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.836228Z digest=sha256:2e641d7e74f4af225e79575f071218bd55aa4dd8c5f1a1d1aa84145c91cda6b7

Observation a58274db-6f1a-4235-bbbd-59ccb7aaae56 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:35:19.071198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:35:18.841905Z digest=sha256:69998e0785777fa0f0e4188b5fe1e0e315bca19826d69a4fe41a6a1d21c0d419

Observation 414cff74-895b-47a4-8848-bdd2979f58b8 · outbound

This paper cites an unresolved cited work.

TextVidBench: A Benchmark for Long Video Scene Text Understanding Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:35:19.061792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:35:18.844911Z digest=sha256:5c2c2ff44785f854c92679f2aaab2743c554c4fdfb22675cd6178c4cb1d1ced1

Observation cb6284ac-2f4d-421d-8026-4ec1b428979a · outbound

This paper cites online" 'onlinestring :=.

TextVidBench: A Benchmark for Long Video Scene Text Understanding online" 'onlinestring :=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.847607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.847607Z digest=sha256:a33982099d8e684fea76ce34d075ae22792ddfa2750a9dec4a20da988e1bc93f

Observation d04e1152-177c-4677-9969-5d520847b655 · outbound

This paper cites write newline.

TextVidBench: A Benchmark for Long Video Scene Text Understanding write newline

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:18.850729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:35:18.850729Z digest=sha256:c8278479de30c5cd16686f077f1ed6b79c749d9ce3e3fccb73249bf77ef1e613

Pith citing papers

No inbound Pith citation observations are available.