Pith. sign in

Paper Citation Record · LEDGER

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts

As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2506.04999.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04999 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:34:03.546459Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T05:51:08.895411Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy50
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19b7fbcd-7a0e-47a9-91ed-980b563b0049 · outbound

This paper cites Word spotting and recognition with embedded attributes.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Word spotting and recognition with embedded attributes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.370790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:54.768761Z digest=sha256:f011ac8681a3daeb1e56b78f9c62e1337b4f08289789ad8d720656e1219a7338

Observation 8d125acd-11c2-4932-a0cd-de2d8ac1427c · outbound

This paper cites Integrating scene text and visual appearance for fine-grained image classification.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Integrating scene text and visual appearance for fine-grained image classification

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.355245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:54.851730Z digest=sha256:4c3aec2b4c06468579aab06a9c678778ffd534ffb4609f07e08719ddbeecf3f7

Observation 725917d1-32b0-4003-9e5d-048aed65f324 · outbound

This paper cites Composed image retrieval using contrastive learning and task-oriented clip-based features.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Composed image retrieval using contrastive learning and task-oriented clip-based features

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.337860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.023632Z digest=sha256:f03f2c44f1e61c6d1f932d7e563aacbee8487af87dfeb14494f8a95c1d204403

Observation 6ab1ddc1-e6d6-446d-bdd9-94d793bc8a99 · outbound

This paper cites The devil is in fine-tuning and long-tailed problems: A new benchmark for scene text detection.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts The devil is in fine-tuning and long-tailed problems: A new benchmark for scene text detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.322419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.221307Z digest=sha256:fabfd4fad1b741deb9746083d0ed03d87b33195a7a059eea3f217a6a1cbfab81

Observation a9ecefec-40c7-452c-aad5-9d1010d67940 · outbound

This paper cites an unresolved cited work.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:34:05.306842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.383782Z digest=sha256:f20e1668e2059456fed5fb8424249ad59c4474c8865b195d14a486a3f3afa882

Observation 67c6b5b4-6b3e-42ad-893f-1766c4ca0c3c · outbound

This paper cites an unresolved cited work.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:34:05.289434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.567461Z digest=sha256:7b80fb0da757e5d3985a9264530ed1371bbae8bc8f72152d55e5eb4a795cfe78

Observation 7326a63d-173c-49bd-aa02-dfb3177080ea · outbound

This paper cites K., Gomez, L., Karatzas, D., and Valveny, E.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts K., Gomez, L., Karatzas, D., and Valveny, E

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.271142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.750412Z digest=sha256:9aa8a307e988649f60109e850e9024fba3f90f38b1707c22e9793ce6d29751ac

Observation d77fb228-da44-412e-94e4-46a9daa8bf59 · outbound

This paper cites Lsde: Levenshtein space deep embedding for query-by-string word spotting.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Lsde: Levenshtein space deep embedding for query-by-string word spotting

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.254671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.944990Z digest=sha256:db39a2eb22052fb744a23dfb590d01c5a43f52f4acbca70d0b7b76a8ef972b3c

Observation 967183e2-ec4b-478b-b0b5-0e00341d3fc8 · outbound

This paper cites Single shot scene text retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Single shot scene text retrieval

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.239639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.079254Z digest=sha256:c763015a0b6b63b0bdac042e240418ded5df6776decea74fe6cd0804c828dd1e

Observation 9c1f59e1-cfab-4ed4-b07b-eac0c396bfaa · outbound

This paper cites Synthetic data for text localisation in natural images.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Synthetic data for text localisation in natural images

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.224142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.206032Z digest=sha256:b3afc4a5bfcc71b1b02aad66d68082f442b0a7422af0a9989cf2863606ae592a

Observation 7c64ac95-9dbd-4e08-8110-2a4d52a8a483 · outbound

This paper cites Bridging the gap between end-to-end and two-step text spotting.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Bridging the gap between end-to-end and two-step text spotting

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.206480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.368840Z digest=sha256:542179677b4b758144717b5184733833c526a603014b87da8bc1ae6bd41e8b6f

Observation 55c12d19-33a6-49b6-895d-5f8a1cf31640 · outbound

This paper cites Reading text in the wild with convolutional neural networks.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Reading text in the wild with convolutional neural networks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.052301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.502105Z digest=sha256:d73df5b39a8060503c3c62dd1a5c98db83297fbb1718486eaff8d1efcc1f754f

Observation b8e8a25f-ff91-4dea-a15b-86c46a14cfcd · outbound

This paper cites an unresolved cited work.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:34:05.037243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.685119Z digest=sha256:24c14f9740029bb3ed222d738ac81c2f3f95dff5609a454c02f1c522b39e5269

Observation 8ae8f700-b529-44ad-89e1-e9cc258ebd46 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.021425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.792242Z digest=sha256:00eff879aa05a87b2ec2db33a96af7b2a43df294a1b44d9a9d83dbb828b483c8

Observation 09dc5171-96f9-4610-aa75-d5a0d8a44054 · outbound

This paper cites Mask textspotter v3: Segmentation proposal network for robust scene text spotting.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Mask textspotter v3: Segmentation proposal network for robust scene text spotting

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.003389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.915356Z digest=sha256:8d4d5d3df49f032403d0d3b35e284feaf24003f38f07f0dda711e9dff5ee1111

Observation 9bd96913-088a-4be8-9437-307444e8fa51 · outbound

This paper cites Parrot Captions Teach CLIP to Spot Text.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Parrot Captions Teach CLIP to Spot Text

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:33:57.108940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:33:57.108940Z digest=sha256:6abb90308cc41ab32d6629a3b2bd1cfb0e01924e5d3e7bbb6609105bef8f912d

Observation a0e84eb5-63f5-4cce-abd7-ad5238686dfc · outbound

This paper cites Abcnet: Real-time scene text spotting with adaptive bezier-curve network.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Abcnet: Real-time scene text spotting with adaptive bezier-curve network

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.986619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:57.275665Z digest=sha256:7c05f894812afaeabfdcbf108136128b06ac0e11fe4a460fed8aaa5ec3721990

Observation 3394719f-ae53-4136-8686-c37337185e42 · outbound

This paper cites Clip4clip: An empirical study of clip for end-to-end video clip retrieval and captioning.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Clip4clip: An empirical study of clip for end-to-end video clip retrieval and captioning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.968964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:57.405128Z digest=sha256:c66688b9f2250f44d3c29aebe452ac811a45c864a0a3d5f3849c26e61657ac44

Observation a6f5e436-a131-431e-9444-31d7188beb4e · outbound

This paper cites Visual and semantic guided scene text retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Visual and semantic guided scene text retrieval

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.952098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:57.566552Z digest=sha256:5e59fc8ac09f5887abca5af133d865d196d1f8450ec29834ade9c6995d52b0a7

Observation 1290ac95-3c8d-498c-8fba-e6136ab610f2 · outbound

This paper cites Arbitrary reading order scene text spotter with local semantics guidance.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Arbitrary reading order scene text spotter with local semantics guidance

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.935312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:57.698449Z digest=sha256:b900f7840567a843e46a30f58f901488df377da89b39e4cb3533867b1b4cc7e7

Observation ca4dca92-f899-4c60-a392-a58de153c75d · outbound

This paper cites TextBlockV2 : Towards precise-detection-free scene text spotting with pre-trained language model.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts TextBlockV2 : Towards precise-detection-free scene text spotting with pre-trained language model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.917434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:57.885055Z digest=sha256:e36cf65858a2cf2f428ce31a57598a95a73f21a3e9f45f866e52074a3905efd4

Observation 67d2aaff-0a84-42d9-97d3-b142b693894c · outbound

This paper cites F., Gomez, L., and Karatzas, D.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts F., Gomez, L., and Karatzas, D

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.896657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.018490Z digest=sha256:fce627c0057d92d093b281f2db9b30b8ea2649b45b47d33f6cbcdafae95d972e

Observation adf89f9b-e367-484c-a52c-d6b0c15db0ad · outbound

This paper cites Real-time lexicon-free scene text retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Real-time lexicon-free scene text retrieval

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.879134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.137410Z digest=sha256:c325d0f0014d05a70c6e7779b53ee7b30b6559ebde2ace3089dcd65c4ba74ef9

Observation 9a27e253-a92b-4800-904f-084ab52987a0 · outbound

This paper cites Disentangling visual and written concepts in clip.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Disentangling visual and written concepts in clip

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.861704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.298068Z digest=sha256:d1c772bdedf12699fd669da043064c152d088a472836648949e3f003b91d0389

Observation 0fd555bc-2131-47b9-91f9-7effe32c3009 · outbound

This paper cites Word spotting and recognition via a joint deep embedding of image and text.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Word spotting and recognition via a joint deep embedding of image and text

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.845036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.501734Z digest=sha256:a5f113c22a513aa010746a20e98456117b18d816165cd1ecb45a71e12aabc6e7

Observation 494d4877-9c19-4f90-8df1-26544445e99c · outbound

This paper cites Image retrieval using textual cues.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Image retrieval using textual cues

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.824192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.633790Z digest=sha256:04772a7ce578adae471f298972d97d87300e3db233826e5d2fd0b0c704b2b072

Observation 43fdc69e-fd6a-441f-96d4-c51a7af8dabd · outbound

This paper cites Text perceptron: Towards end-to-end arbitrary-shaped text spotting.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Text perceptron: Towards end-to-end arbitrary-shaped text spotting

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.800334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.756760Z digest=sha256:4c0b30fe833ad7355f2a6fb42b87b1cf4ac182a71a8b5b73c43cb397a75d8781

Observation 33c759fe-ff90-40e3-899c-695a9e8c2622 · outbound

This paper cites SEED : Semantics enhanced encoder-decoder framework for scene text recognition.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts SEED : Semantics enhanced encoder-decoder framework for scene text recognition

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.777621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.932601Z digest=sha256:fc082afe243f1177267c03c98655aefefbe1dcb37e8ccab625deefa91c11086f

Observation 26669bf7-34ac-4dc0-9a7a-bb1f1fda7196 · outbound

This paper cites PIMNet : A parallel, iterative and mimicking network for scene text recognition.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts PIMNet : A parallel, iterative and mimicking network for scene text recognition

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.757899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.113247Z digest=sha256:feac51a1c7acefbb36a340352a7b08e8f197924a1ffad62c914d71129e88548f

Observation c812f105-6f83-4140-90a6-86a2dcbe04b7 · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.739445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.311002Z digest=sha256:78f3667bd2638d8ff5492a19594f70631d3de5e0d8fa38940d7b7dfed187779f

Observation 29099d5c-5d3b-49fd-8f8e-2455b3566cfd · outbound

This paper cites Pic2word: Mapping pictures to words for zero-shot composed image retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Pic2word: Mapping pictures to words for zero-shot composed image retrieval

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.721235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.490623Z digest=sha256:9c5744af9da93fc2b55d6c3a14ef40bac413470915deb2338a30f790b00005a3

Observation 3193a5a6-dac0-43c3-8005-6a8100dfc4d0 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.700159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.630256Z digest=sha256:9d7ae17404d83133a8a2fa895b287d12ad52062718aca399a0c04c0e66115cc1

Observation d5948cc0-dc9a-4651-a9cd-0024258b43a4 · outbound

This paper cites LDP : Generalizing to multilingual visual information extraction by language decoupled pretraining.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts LDP : Generalizing to multilingual visual information extraction by language decoupled pretraining

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.677264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.737394Z digest=sha256:a450a38b9f32314182a17096484344d3d0eb9ce65c4500e86bccdaa07756df25

Observation 36a0a6c3-4e20-452e-bf5d-7be37736e743 · outbound

This paper cites and Yang, S.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts and Yang, S

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.655824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.865196Z digest=sha256:00d982cbee9225e2a4cb01f34dcbc91caf7b65eabbd35d219e4e5957e93ae3b0

Observation bb09becc-0361-4559-98c8-9d9091963ae1 · outbound

This paper cites What does clip know about a red circle? visual prompt engineering for vlms.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts What does clip know about a red circle? visual prompt engineering for vlms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.633680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.999588Z digest=sha256:c62e2cda8e3ee812cbd2fe4688034523de3799b38aa2fbc2efb0ec0a661e9fee

Observation 3b8b2a52-613f-4fa0-b765-416e8b945def · outbound

This paper cites Perceiving ambiguity and semantics without recognition: An efficient and effective ambiguous scene text detector.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Perceiving ambiguity and semantics without recognition: An efficient and effective ambiguous scene text detector

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.612274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:00.150960Z digest=sha256:a20be978e54c62889cc631526c88802f822af32bff1d2eb9037ee5409b24aa6c

Observation 7582d782-bcc6-4193-a654-25ce888f7cbe · outbound

This paper cites Visual Text Processing: A Comprehensive Review and Unified Evaluation.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Visual Text Processing: A Comprehensive Review and Unified Evaluation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:00.263138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:00.263138Z digest=sha256:1f1b4f01f95be59ba816c379d441948339183cfc2c863641d5f9f3d6c000dcc9

Observation 8afb498f-98dd-4c1a-a222-54ea38e52ad3 · outbound

This paper cites Text siamese network for video textual keyframe detection.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Text siamese network for video textual keyframe detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.592910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:00.423634Z digest=sha256:77f566dca7e38d1c7fbff548d1c6e025e24a2a10b2d76c7d28c7da3afa32f62f

Observation 0b702319-8127-42fb-ab6a-b59835d7f895 · outbound

This paper cites ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:00.550841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:00.550841Z digest=sha256:02dbb966b8351d8cb32b7d36e1c6f6271386325fa302e0c76f5cfa43dd399e87

Observation 7bab5433-11c0-4b08-a4fa-649066da084a · outbound

This paper cites Alpha-clip: A clip model focusing on wherever you want.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Alpha-clip: A clip model focusing on wherever you want

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.567265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:00.739489Z digest=sha256:b1e7ef59748ef9ffe99d3d2f882f847bb75375722637904d318d7a8bf8b90274

Observation 6208d676-89e0-4553-b63e-49921c64e1d4 · outbound

This paper cites Scene text retrieval via joint text detection and similarity learning.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Scene text retrieval via joint text detection and similarity learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.550844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:00.863099Z digest=sha256:d9f629d299eb91e235c7f6367fada36772047a0d5466a318c9311e9c7b064ac2

Observation 98b0e76a-9b63-4201-90b8-5e6fa7dcc40b · outbound

This paper cites and Belongie, S.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts and Belongie, S

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.533936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:00.975641Z digest=sha256:470c36ca63b678bb02941b8ca350e1f93a24b4ef35d427df35f80efc6b6f2e96

Observation 0230ef5f-da59-4ce7-8441-fa8aa021f513 · outbound

This paper cites TPSNet : Reverse thinking of thin plate splines for arbitrary shape scene text representation.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts TPSNet : Reverse thinking of thin plate splines for arbitrary shape scene text representation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.509147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:01.158728Z digest=sha256:14155d67da54227b0cb2c29400a1fd64a3a8afc7699224d94c83dba30a47a356

Observation 840cf871-9a2c-4b19-9ba7-38f7781a32d6 · outbound

This paper cites TextBlock : Towards scene text spotting without fine-grained detection.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts TextBlock : Towards scene text spotting without fine-grained detection

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.492576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:01.308373Z digest=sha256:5ec10d79fc50ab9a647bc7baad7a840ba0dd40533cb68be35752b3901670f286

Observation 05d1d8c1-31d6-442f-aadd-fb08bb350c8a · outbound

This paper cites Visual matching is enough for scene text retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Visual matching is enough for scene text retrieval

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.475436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:01.467766Z digest=sha256:30e6214bb256d7daaee466a0067ae172d9128aca6fdc0278a5f4ca4758d4c812

Observation d868ad56-f596-4837-bcc9-a5ae60c5d046 · outbound

This paper cites and Brun, A.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts and Brun, A

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.452994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:01.645696Z digest=sha256:830ec4c5141ea8990665c882a466dcda15a09fe3e71ffcab2bdb11bf4a1d2fdf

Observation 7787ef5a-3a50-4eac-af92-27cb0176e811 · outbound

This paper cites Bridging semantic gaps for language-supervised semantic segmentation.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Bridging semantic gaps for language-supervised semantic segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.434978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:01.767073Z digest=sha256:fda2c5f109aa3fba489b396b00da4cdf46e42f930544ebf9709e28cdae460de2

Observation f1967491-8899-4f2d-85a9-d16e43bfea0e · outbound

This paper cites Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:01.936414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:01.936414Z digest=sha256:0648fc19037a62f832e5adab64709e3d75f91fcdf3088a0496cab6a18ce76428

Observation d672d0c0-2f1d-46bb-81ac-1a78f2cbf2ed · outbound

This paper cites IPAD : Iterative, parallel, and diffusion-based network for scene text recognition.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts IPAD : Iterative, parallel, and diffusion-based network for scene text recognition

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.415705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.139112Z digest=sha256:f29139994b4e72d8b237881ee215424dd94fc81ea77493c78497deef5636b6cc

Observation 329c6e7e-26df-4a66-af7f-658495b5cf18 · outbound

This paper cites Turning a clip model into a scene text detector.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Turning a clip model into a scene text detector

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.393041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.295822Z digest=sha256:bd1f5b2f06a59ce5198c724a8420ef327425afa89130efe73a24de7efd39514b

Observation ebecad2c-826a-4bac-b244-797a5e3837c7 · outbound

This paper cites Beyond OCR + VQA : Towards end-to-end reading and reasoning for robust and accurate textvqa.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Beyond OCR + VQA : Towards end-to-end reading and reasoning for robust and accurate textvqa

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.372619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.427022Z digest=sha256:36ab09bbbe5a30d002e18bee98addfc0b9939c108bc62c8b7fbef01005dcaa44

Observation 0b96d416-462b-4063-8ac6-9952abdee07c · outbound

This paper cites Focus, distinguish, and prompt: Unleashing CLIP for efficient and flexible scene text retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Focus, distinguish, and prompt: Unleashing CLIP for efficient and flexible scene text retrieval

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.345456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.623053Z digest=sha256:050b9bda46ed1fa0f8e466967e8ea032cff478154395cae567c597dd7a2e0c32

Observation 3cec4b21-caac-42c8-99a8-3e5dc85313e0 · outbound

This paper cites TextCtrl : Diffusion-based scene text editing with prior guidance control.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts TextCtrl : Diffusion-based scene text editing with prior guidance control

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.314377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.764417Z digest=sha256:4a1c8f1396e2701e737d3212dd7219718886c038da67b3c66f574cf82641613c

Observation f22ec2ea-6ebf-4912-a5cb-199b77e2db53 · outbound

This paper cites Icdar 2019 robust reading challenge on reading chinese text on signboard.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Icdar 2019 robust reading challenge on reading chinese text on signboard

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.285972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.932231Z digest=sha256:2e8176fe31d955a456382721b64bf7256eb2928f2f76c62ef62fc0a592d33a22

Observation 108619cb-9a22-4768-9e8b-9aa182cd25cf · outbound

This paper cites Linguistics-aware masked image modeling for self-supervised scene text recognition.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Linguistics-aware masked image modeling for self-supervised scene text recognition

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.257084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:03.135547Z digest=sha256:edde425a1e2a7685856cba90c35c1527c1bf5d6087fc78dcc60c8792f625e9ee

Observation 2457c7fe-0a2f-45d4-bcff-032fe5305315 · outbound

This paper cites C., and Dai, B.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts C., and Dai, B

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.142177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:03.265589Z digest=sha256:8330833b5335bda939056bdd865761c6bf24fba8b745eda9d6ae7d3e993f3551

Observation 7934d71d-86ab-4dcf-8d81-aed4847aa900 · outbound

This paper cites C., and Liu, Z.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts C., and Liu, Z

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:03.908191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:34:03.386877Z digest=sha256:9d0352104215345162cfb543b4aa7f78b5fddbd01c346df80ee939374dfdc99a

Observation ca56f343-47b7-4f98-b6f0-f9d6456a2e1a · outbound

This paper cites write newline.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts write newline

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:03.546459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:03.546459Z digest=sha256:208ed379c2d45bb2da89e694b5a178e760b466a99a7d59e98f575edef91a0427

Pith citing papers

Observation dad30a42-fba3-4c1c-b7a1-f2edd0a12c4e · inbound

MORE: A Multilingual Document Parsing Benchmark and Evaluation cites this paper.

MORE: A Multilingual Document Parsing Benchmark and Evaluation Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T05:51:08.895411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:51:08.895411Z digest=sha256:d9859641743dba9a89c3b6b7481abfae6d1e39d7b6e3ae16418d07260c662317