Pith. sign in

Paper Citation Record · LEDGER

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts

As of 19 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2506.04999.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04999 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:34:03.546459Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T05:51:08.895411Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy50
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19b7fbcd-7a0e-47a9-91ed-980b563b0049 · outbound

This paper cites Word spotting and recognition with embedded attributes.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Word spotting and recognition with embedded attributes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.370790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:54.768761Z digest=sha256:1aeb96411c5a201bc1566736712e59c221d1b7d95c21d3e43edbda959868d9ad

Observation 8d125acd-11c2-4932-a0cd-de2d8ac1427c · outbound

This paper cites Integrating scene text and visual appearance for fine-grained image classification.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Integrating scene text and visual appearance for fine-grained image classification

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.355245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:54.851730Z digest=sha256:963d2d4ba66fd18887af0fd9e3cb3af45e8c73bc882f8d52555f915f685a74af

Observation 725917d1-32b0-4003-9e5d-048aed65f324 · outbound

This paper cites Composed image retrieval using contrastive learning and task-oriented clip-based features.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Composed image retrieval using contrastive learning and task-oriented clip-based features

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.337860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.023632Z digest=sha256:052a3e76c3766a0e6206c38acce9605f6dc33411b843807566fe9e7bf72e84db

Observation 6ab1ddc1-e6d6-446d-bdd9-94d793bc8a99 · outbound

This paper cites The devil is in fine-tuning and long-tailed problems: A new benchmark for scene text detection.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts The devil is in fine-tuning and long-tailed problems: A new benchmark for scene text detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.322419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.221307Z digest=sha256:4c0059ac12aea1da16f7ad396a1a0a229a1b4a42e6fb50881fe5db1177f90bef

Observation a9ecefec-40c7-452c-aad5-9d1010d67940 · outbound

This paper cites an unresolved cited work.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:34:05.306842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.383782Z digest=sha256:ed219cd653b8c05b610419329f39bebd64565eb38586466d90631778e9339397

Observation 67c6b5b4-6b3e-42ad-893f-1766c4ca0c3c · outbound

This paper cites an unresolved cited work.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:34:05.289434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.567461Z digest=sha256:98324bbaf521341b978f2fecbbc6e96400ce5c0832791e20a2712f340513205b

Observation 7326a63d-173c-49bd-aa02-dfb3177080ea · outbound

This paper cites K., Gomez, L., Karatzas, D., and Valveny, E.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts K., Gomez, L., Karatzas, D., and Valveny, E

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.271142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.750412Z digest=sha256:e842841b3195c639ab519c5bf29949ab749a8644b3b16af99b48e71f73a55138

Observation d77fb228-da44-412e-94e4-46a9daa8bf59 · outbound

This paper cites Lsde: Levenshtein space deep embedding for query-by-string word spotting.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Lsde: Levenshtein space deep embedding for query-by-string word spotting

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.254671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:55.944990Z digest=sha256:896411d7346e5fe15ad6563a3d3b1b7d59d613e1c6d8387a103b7264d63af2be

Observation 967183e2-ec4b-478b-b0b5-0e00341d3fc8 · outbound

This paper cites Single shot scene text retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Single shot scene text retrieval

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.239639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.079254Z digest=sha256:cab954af885f2e81a4293d52a384a3fbb015b8f43c6736ab295de3062b750e07

Observation 9c1f59e1-cfab-4ed4-b07b-eac0c396bfaa · outbound

This paper cites Synthetic data for text localisation in natural images.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Synthetic data for text localisation in natural images

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.224142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.206032Z digest=sha256:2dc2e2aeb3c8244426b6d77e6f51f57f17b0561ffc69588e6952140802e2224a

Observation 7c64ac95-9dbd-4e08-8110-2a4d52a8a483 · outbound

This paper cites Bridging the gap between end-to-end and two-step text spotting.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Bridging the gap between end-to-end and two-step text spotting

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.206480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.368840Z digest=sha256:c3f3e0fdecdecbecbc3a75e89e26bcb15b431eb8abecc398eac12bf815838482

Observation 55c12d19-33a6-49b6-895d-5f8a1cf31640 · outbound

This paper cites Reading text in the wild with convolutional neural networks.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Reading text in the wild with convolutional neural networks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.052301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.502105Z digest=sha256:a9b25978ab77b8c2f84a92c8da942d36a0ae7891fe9be871428e6c8c6d7e5560

Observation b8e8a25f-ff91-4dea-a15b-86c46a14cfcd · outbound

This paper cites an unresolved cited work.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:34:05.037243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.685119Z digest=sha256:b1bc82b07b0475b3be284b74fdd8a553fb3f7619c849085d4317e5ae16d4221e

Observation 8ae8f700-b529-44ad-89e1-e9cc258ebd46 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.021425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.792242Z digest=sha256:e7837aa0093a5448ec42a9a8d11dd5157a8abad2f6fd46728f039c2c88554afc

Observation 09dc5171-96f9-4610-aa75-d5a0d8a44054 · outbound

This paper cites Mask textspotter v3: Segmentation proposal network for robust scene text spotting.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Mask textspotter v3: Segmentation proposal network for robust scene text spotting

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:05.003389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:56.915356Z digest=sha256:08ef680d754d30308edd16b211337d687fc732ca9ed24ab7dd97ff44f1a887d9

Observation 9bd96913-088a-4be8-9437-307444e8fa51 · outbound

This paper cites Parrot Captions Teach CLIP to Spot Text.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Parrot Captions Teach CLIP to Spot Text

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:33:57.108940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:33:57.108940Z digest=sha256:9d099aa7ee47885e30d5d725747c0bbd980269c2ef2ef9c29e15438c80da287e

Observation a0e84eb5-63f5-4cce-abd7-ad5238686dfc · outbound

This paper cites Abcnet: Real-time scene text spotting with adaptive bezier-curve network.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Abcnet: Real-time scene text spotting with adaptive bezier-curve network

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.986619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:57.275665Z digest=sha256:32f72b11ced1782ca0b2f4e2cf30927e33409b90bcd875c7dbac0d0825478abb

Observation 3394719f-ae53-4136-8686-c37337185e42 · outbound

This paper cites Clip4clip: An empirical study of clip for end-to-end video clip retrieval and captioning.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Clip4clip: An empirical study of clip for end-to-end video clip retrieval and captioning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.968964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:57.405128Z digest=sha256:36c981894b4607ef9ee7510db9467423b228b28ce79840225455d818157fbfec

Observation a6f5e436-a131-431e-9444-31d7188beb4e · outbound

This paper cites Visual and semantic guided scene text retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Visual and semantic guided scene text retrieval

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.952098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:57.566552Z digest=sha256:9fbc1e7a3b2c3901a1c6e838c05f24ee58b0501bed6640d4f3572cbbb640cb4e

Observation 1290ac95-3c8d-498c-8fba-e6136ab610f2 · outbound

This paper cites Arbitrary reading order scene text spotter with local semantics guidance.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Arbitrary reading order scene text spotter with local semantics guidance

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.935312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:57.698449Z digest=sha256:168bcb625d4386cbc1e560613de3987bb4b2ba7e0e8fe54c9b888d23f67cacac

Observation ca4dca92-f899-4c60-a392-a58de153c75d · outbound

This paper cites TextBlockV2 : Towards precise-detection-free scene text spotting with pre-trained language model.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts TextBlockV2 : Towards precise-detection-free scene text spotting with pre-trained language model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.917434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:57.885055Z digest=sha256:ac6d472a131c26b590113fb6643bfc4a676d683b5e8570d07f37985964e187f9

Observation 67d2aaff-0a84-42d9-97d3-b142b693894c · outbound

This paper cites F., Gomez, L., and Karatzas, D.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts F., Gomez, L., and Karatzas, D

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.896657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.018490Z digest=sha256:ac37b8f7396c34721c610b46e9017baeaac42b06e0bd4173c32f4bebeb9ecfa7

Observation adf89f9b-e367-484c-a52c-d6b0c15db0ad · outbound

This paper cites Real-time lexicon-free scene text retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Real-time lexicon-free scene text retrieval

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.879134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.137410Z digest=sha256:011ffc78743710ce68f4d4139f0d7415ae703b99c1c2cf8f768d9916929f537e

Observation 9a27e253-a92b-4800-904f-084ab52987a0 · outbound

This paper cites Disentangling visual and written concepts in clip.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Disentangling visual and written concepts in clip

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.861704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.298068Z digest=sha256:4c0b21d0c41b4783e78423cb8db30633338f57ef2c4d9d0502e6f3b977ec7934

Observation 0fd555bc-2131-47b9-91f9-7effe32c3009 · outbound

This paper cites Word spotting and recognition via a joint deep embedding of image and text.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Word spotting and recognition via a joint deep embedding of image and text

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.845036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.501734Z digest=sha256:65c0ac643fba2958e923be09bdc44ce71968ab881205ab3e5ce025df3fea9e84

Observation 494d4877-9c19-4f90-8df1-26544445e99c · outbound

This paper cites Image retrieval using textual cues.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Image retrieval using textual cues

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.824192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.633790Z digest=sha256:4864b117d1e33a574f292524a379fdfeee35532dbbe6653eb4f95d507cf190f3

Observation 43fdc69e-fd6a-441f-96d4-c51a7af8dabd · outbound

This paper cites Text perceptron: Towards end-to-end arbitrary-shaped text spotting.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Text perceptron: Towards end-to-end arbitrary-shaped text spotting

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.800334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.756760Z digest=sha256:1407a1e085e402ea1b13f67a0db0382dc5dd1526ebc9d6618c10c6ad5570d01b

Observation 33c759fe-ff90-40e3-899c-695a9e8c2622 · outbound

This paper cites SEED : Semantics enhanced encoder-decoder framework for scene text recognition.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts SEED : Semantics enhanced encoder-decoder framework for scene text recognition

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.777621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:58.932601Z digest=sha256:b3c68f607370debf939fb2ad1e8165746ed9c75c18f5e64ef0b5a0672cba0690

Observation 26669bf7-34ac-4dc0-9a7a-bb1f1fda7196 · outbound

This paper cites PIMNet : A parallel, iterative and mimicking network for scene text recognition.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts PIMNet : A parallel, iterative and mimicking network for scene text recognition

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.757899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.113247Z digest=sha256:229091d56644d8e1059d9f02d932db4394d8e740a175f9c0a4f8e3e1e4ad2e5e

Observation c812f105-6f83-4140-90a6-86a2dcbe04b7 · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.739445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.311002Z digest=sha256:cfe59941aa0c36e948e58b69a6b31890e954be6a357adc6e82e9748864db1e5d

Observation 29099d5c-5d3b-49fd-8f8e-2455b3566cfd · outbound

This paper cites Pic2word: Mapping pictures to words for zero-shot composed image retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Pic2word: Mapping pictures to words for zero-shot composed image retrieval

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.721235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.490623Z digest=sha256:0ffa9e4fbb7f786c1bea1d6de3033e63b0c455aa5e05674034d93f7b85d0f773

Observation 3193a5a6-dac0-43c3-8005-6a8100dfc4d0 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.700159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.630256Z digest=sha256:0c24b284d95330527d24e3948eb22c1ccbc98b4b2f51032226fc8d5769c52291

Observation d5948cc0-dc9a-4651-a9cd-0024258b43a4 · outbound

This paper cites LDP : Generalizing to multilingual visual information extraction by language decoupled pretraining.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts LDP : Generalizing to multilingual visual information extraction by language decoupled pretraining

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.677264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.737394Z digest=sha256:3929e071cc79f59da35754a1aeb07b82dfef6e70eb95e26d09eb69824d8d7e51

Observation 36a0a6c3-4e20-452e-bf5d-7be37736e743 · outbound

This paper cites and Yang, S.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts and Yang, S

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.655824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.865196Z digest=sha256:ec5190b3807333a096222206eddb8b23f2270716933db536f47f94ecc656e253

Observation bb09becc-0361-4559-98c8-9d9091963ae1 · outbound

This paper cites What does clip know about a red circle? visual prompt engineering for vlms.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts What does clip know about a red circle? visual prompt engineering for vlms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.633680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:33:59.999588Z digest=sha256:fbf02bf54f2090f3aabdb0f38c350c540ffa6cc49c8dd59f63e81d2384c4caa9

Observation 3b8b2a52-613f-4fa0-b765-416e8b945def · outbound

This paper cites Perceiving ambiguity and semantics without recognition: An efficient and effective ambiguous scene text detector.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Perceiving ambiguity and semantics without recognition: An efficient and effective ambiguous scene text detector

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.612274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:00.150960Z digest=sha256:8a50a1df41483d8b5ef5f26e9f606785c0398fd5ecc2058102c094dfe2921f4b

Observation 7582d782-bcc6-4193-a654-25ce888f7cbe · outbound

This paper cites Visual Text Processing: A Comprehensive Review and Unified Evaluation.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Visual Text Processing: A Comprehensive Review and Unified Evaluation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:00.263138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:00.263138Z digest=sha256:ce2a4b8c210049496666df806b5a509424036bbf48ba7fc6b64b42187c78736d

Observation 8afb498f-98dd-4c1a-a222-54ea38e52ad3 · outbound

This paper cites Text siamese network for video textual keyframe detection.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Text siamese network for video textual keyframe detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.592910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:00.423634Z digest=sha256:2e03c18d2711e40d0da95034faaf94f94cc46c99aaad52a729fbbbd964090324

Observation 0b702319-8127-42fb-ab6a-b59835d7f895 · outbound

This paper cites ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:00.550841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:00.550841Z digest=sha256:6140a626d89b474994b7eb8023b7d2921895767cb7e7ddbc8930c4374276d187

Observation 7bab5433-11c0-4b08-a4fa-649066da084a · outbound

This paper cites Alpha-clip: A clip model focusing on wherever you want.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Alpha-clip: A clip model focusing on wherever you want

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.567265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:00.739489Z digest=sha256:852b518fbc31d49bc40f4b05cd995cbec00debb56bbe6321f66172a975eba12f

Observation 6208d676-89e0-4553-b63e-49921c64e1d4 · outbound

This paper cites Scene text retrieval via joint text detection and similarity learning.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Scene text retrieval via joint text detection and similarity learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.550844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:00.863099Z digest=sha256:d7056f03a876fc7ef896a01a06d2b76835c88924e8e3ea28ac239f359f801866

Observation 98b0e76a-9b63-4201-90b8-5e6fa7dcc40b · outbound

This paper cites and Belongie, S.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts and Belongie, S

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.533936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:00.975641Z digest=sha256:d9b9698ca2fa864f6925d7651a6593606b116148377f8a1e74feb537e0f495bb

Observation 0230ef5f-da59-4ce7-8441-fa8aa021f513 · outbound

This paper cites TPSNet : Reverse thinking of thin plate splines for arbitrary shape scene text representation.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts TPSNet : Reverse thinking of thin plate splines for arbitrary shape scene text representation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.509147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:01.158728Z digest=sha256:dd998a5a5555bb67b51a6107b11c088d8cb1ec2e58608b55a1f107c218912ebf

Observation 840cf871-9a2c-4b19-9ba7-38f7781a32d6 · outbound

This paper cites TextBlock : Towards scene text spotting without fine-grained detection.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts TextBlock : Towards scene text spotting without fine-grained detection

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.492576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:01.308373Z digest=sha256:049588e2642031c1539b166dc2fb18b7881fa541e57d584d3ba6881336c27fea

Observation 05d1d8c1-31d6-442f-aadd-fb08bb350c8a · outbound

This paper cites Visual matching is enough for scene text retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Visual matching is enough for scene text retrieval

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.475436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:01.467766Z digest=sha256:b5699f3b63f215e90820c4eaeb834a894826308f7b72fad34250985957141471

Observation d868ad56-f596-4837-bcc9-a5ae60c5d046 · outbound

This paper cites and Brun, A.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts and Brun, A

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.452994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:01.645696Z digest=sha256:b6dd174e1f358ae4b853a8a0c238acfbd3b986cb007a79ca789531698f3b5a78

Observation 7787ef5a-3a50-4eac-af92-27cb0176e811 · outbound

This paper cites Bridging semantic gaps for language-supervised semantic segmentation.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Bridging semantic gaps for language-supervised semantic segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.434978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:01.767073Z digest=sha256:32c37b45a664dcb6d8c4674873dbcdbd9c2eef10e10e4618e5176be918992681

Observation f1967491-8899-4f2d-85a9-d16e43bfea0e · outbound

This paper cites Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:01.936414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:01.936414Z digest=sha256:c8becce262acbd253efee53652325f1b7192ec6251c2c4fe8abc6bf9d7cbc87b

Observation d672d0c0-2f1d-46bb-81ac-1a78f2cbf2ed · outbound

This paper cites IPAD : Iterative, parallel, and diffusion-based network for scene text recognition.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts IPAD : Iterative, parallel, and diffusion-based network for scene text recognition

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.415705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.139112Z digest=sha256:0b1a83a4f6673ab7ffb765850adda6761d7b841d43bf6927da3f30e5d3de2c4f

Observation 329c6e7e-26df-4a66-af7f-658495b5cf18 · outbound

This paper cites Turning a clip model into a scene text detector.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Turning a clip model into a scene text detector

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.393041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.295822Z digest=sha256:a2255e17cbeb451750b4ac17ffc031a7766ee10a93bec471c495e292799d8393

Observation ebecad2c-826a-4bac-b244-797a5e3837c7 · outbound

This paper cites Beyond OCR + VQA : Towards end-to-end reading and reasoning for robust and accurate textvqa.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Beyond OCR + VQA : Towards end-to-end reading and reasoning for robust and accurate textvqa

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.372619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.427022Z digest=sha256:75a71a11c9d735cabeea91753599938bae188302142b9bc31909adec9ff8a387

Observation 0b96d416-462b-4063-8ac6-9952abdee07c · outbound

This paper cites Focus, distinguish, and prompt: Unleashing CLIP for efficient and flexible scene text retrieval.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Focus, distinguish, and prompt: Unleashing CLIP for efficient and flexible scene text retrieval

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.345456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.623053Z digest=sha256:fe48c31c334de25e7c3eb6e8c0a35d5b42ae0bc206ef0f71942a42a15467cb06

Observation 3cec4b21-caac-42c8-99a8-3e5dc85313e0 · outbound

This paper cites TextCtrl : Diffusion-based scene text editing with prior guidance control.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts TextCtrl : Diffusion-based scene text editing with prior guidance control

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.314377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.764417Z digest=sha256:02ee2d99c673012da0624fc605a1e1d763da547e1cfa4522d420dc9287ac952a

Observation f22ec2ea-6ebf-4912-a5cb-199b77e2db53 · outbound

This paper cites Icdar 2019 robust reading challenge on reading chinese text on signboard.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Icdar 2019 robust reading challenge on reading chinese text on signboard

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.285972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:02.932231Z digest=sha256:c79e610eb09a7ed173544682d0ed78ec942a42715c169f69a34b0e5ccbaa3700

Observation 108619cb-9a22-4768-9e8b-9aa182cd25cf · outbound

This paper cites Linguistics-aware masked image modeling for self-supervised scene text recognition.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Linguistics-aware masked image modeling for self-supervised scene text recognition

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.257084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:03.135547Z digest=sha256:6066ca5066a2211c7819c2dfca68edeca20ed307f6ff5d05a92862d4e13232c0

Observation 2457c7fe-0a2f-45d4-bcff-032fe5305315 · outbound

This paper cites C., and Dai, B.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts C., and Dai, B

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:04.142177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:03.265589Z digest=sha256:420706a3d5d3e3a5045bfcac2615c8c08daa85e3d88f585b833d60e30a1e42da

Observation 7934d71d-86ab-4dcf-8d81-aed4847aa900 · outbound

This paper cites C., and Liu, Z.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts C., and Liu, Z

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:34:03.908191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:34:03.386877Z digest=sha256:056705c63bcc385b887e91d92fe38edafb39ac3541e42a031cf29872cf1751f3

Observation ca56f343-47b7-4f98-b6f0-f9d6456a2e1a · outbound

This paper cites write newline.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts write newline

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:03.546459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:03.546459Z digest=sha256:6c909e1db6c40e2bd13a1cd58bfd262a72388bb7b1828d57109a7ae93fcd4702

Pith citing papers

Observation dad30a42-fba3-4c1c-b7a1-f2edd0a12c4e · inbound

MORE: A Multilingual Document Parsing Benchmark and Evaluation cites this paper.

MORE: A Multilingual Document Parsing Benchmark and Evaluation Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T05:51:08.895411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:51:08.895411Z digest=sha256:50e126bcd06984a68a47fd79966b5f76cc0181c3e2d173b2bf10bc0298c384ee