Pith. sign in

Paper Citation Record · LEDGER

Hear the Scene: Audio-Enhanced Text Spotting

As of 11 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2412.19504.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.19504 v3

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:21:07.885817Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3e3736c0-0c32-4234-87ed-1374dbbce715 · outbound

This paper cites SPTS: Single-Point Text Spotting.

Hear the Scene: Audio-Enhanced Text Spotting SPTS: Single-Point Text Spotting

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:21:08.243152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:21:07.849502Z digest=sha256:f4df7ebb29f9eec663527305457862a026496e555088c312198d1b918254ead5

Observation bf08b69b-0585-4048-88e9-224ba16c255b · outbound

This paper cites Wenhao Sun, Xue-Mei Dong, Benlei Cui, and Jingqun Tang.

Hear the Scene: Audio-Enhanced Text Spotting Wenhao Sun, Xue-Mei Dong, Benlei Cui, and Jingqun Tang

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:07.858949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:07.858949Z digest=sha256:61372920689d7a0bc9bd79c82c4ef681ca2af8612aaf2f2e396bd416aa617c64

Observation fba9246a-b314-43ea-950d-e636f0264b63 · outbound

This paper cites Optimal boxes: boosting end-to-end scene text recognition by adjusting annotated bounding boxes via reinforcement learning.

Hear the Scene: Audio-Enhanced Text Spotting Optimal boxes: boosting end-to-end scene text recognition by adjusting annotated bounding boxes via reinforcement learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:07.862859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:07.862859Z digest=sha256:413339725a169cca4ff5f9cacdbcef71e8f094ecc7a6f23e26a7dd65b3be1870

Observation 0d154a4d-4040-4e90-9a6d-ddc36b4af98f · outbound

This paper cites MaRI: Material Retrieval Integration across Domains.

Hear the Scene: Audio-Enhanced Text Spotting MaRI: Material Retrieval Integration across Domains

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:07.872147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:07.872147Z digest=sha256:b4a27e199bb49463bce8e94e62e4ee1b950d16083c6d4375230e4c48393da42b

Observation e9488525-4c20-4884-98c2-a7b4aa274469 · outbound

This paper cites Pgnet: Real-time arbitrarily-shaped text spotting with point gathering network.

Hear the Scene: Audio-Enhanced Text Spotting Pgnet: Real-time arbitrarily-shaped text spotting with point gathering network

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:07.876992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:07.876992Z digest=sha256:8ac570297e1799d24e2a765e512da4b87fc225cf1cf3c6421a3a3060b3971a46

Observation 221ac8c3-4498-43ad-bd65-493db31b42b6 · outbound

This paper cites Detecting Curve Text in the Wild: New Dataset and New Solution.

Hear the Scene: Audio-Enhanced Text Spotting Detecting Curve Text in the Wild: New Dataset and New Solution

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:07.881364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:07.881364Z digest=sha256:8057553b4edade9132c668c01d3d6c7c0944932b05307fd772aca15fb690b426

Observation 10a0ad06-6139-4658-a805-234a3e2a94fa · outbound

This paper cites Icdar 2015 competition on robust reading.

Hear the Scene: Audio-Enhanced Text Spotting Icdar 2015 competition on robust reading

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:08.301731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:21:07.830619Z digest=sha256:d1ba276b5f2c8d2dcec0376b1a1db5543659dd1bfff83ea28ff73cde7f41a9d2

Observation 463dba09-af75-409e-a9ab-143d5bb24204 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Hear the Scene: Audio-Enhanced Text Spotting LLaVA-OneVision: Easy Visual Task Transfer

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:07.835018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:07.835018Z digest=sha256:04700d542e60fdf4ecd41cd08d54e9f235ee7cd98ab57fbe20a6ceba4a46f13b

Observation eb421500-3124-493c-90e8-6e3db399d5af · outbound

This paper cites Harmonizing Visual Text Comprehension and Generation.

Hear the Scene: Audio-Enhanced Text Spotting Harmonizing Visual Text Comprehension and Generation

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:07.885817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:07.885817Z digest=sha256:22e28c4bf168c280edf24eeb97b71f60f096955ec6351110ed0ae0fbeff320fa

Observation 50070c02-cf12-4a36-b1ab-9051814e57b0 · outbound

This paper cites an unresolved cited work.

Hear the Scene: Audio-Enhanced Text Spotting Unresolved cited work

Reference 2018

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:21:08.269825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:21:07.845228Z digest=sha256:c6ea95ec4ceb037b0ca2faaf0953a09229ccd7062ac18162a904799eb705afcc

Observation 0a055a81-5913-41f8-ae1d-5c65edf415b0 · outbound

This paper cites Icdar 2013 robust reading competition.

Hear the Scene: Audio-Enhanced Text Spotting Icdar 2013 robust reading competition

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:08.317089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:21:07.825758Z digest=sha256:a9807bb374df8eb9558e9a14d59a0d037f9a3ca91744e39ba406d9675b420ffe

Observation 38eee840-95b0-48bb-84eb-c4aab8a10d0a · outbound

This paper cites ISBN 978-3-030-58621-8.

Hear the Scene: Audio-Enhanced Text Spotting ISBN 978-3-030-58621-8

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:21:08.286019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T00:21:07.840297Z digest=sha256:c65ac352b1fcb889ada4f30139b73dc5528d7ad918b39d1441b3baec4855b6b0

Observation 973347f9-4992-4b2b-9ad0-b61ebbe9ebbd · outbound

This paper cites TextSquare: Scaling up Text-Centric Visual Instruction Tuning.

Hear the Scene: Audio-Enhanced Text Spotting TextSquare: Scaling up Text-Centric Visual Instruction Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:07.867247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:07.867247Z digest=sha256:9844c7e7b5d0fe48e931152f747965e29afc53c3f1bef5dbc9d237d1bcb36e97

Observation 53f70042-0b80-4414-89d9-c00deccd73be · outbound

This paper cites MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark.

Hear the Scene: Audio-Enhanced Text Spotting MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T00:21:07.854835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:21:07.854835Z digest=sha256:643b1fc74dec0dece2a24f3e628eed0ff946c36f866bff0a0f19bddfec261db5

Pith citing papers

No inbound Pith citation observations are available.