Pith. sign in

Paper Citation Record · LEDGER

Towards VQA Models That Can Read

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:1904.08920.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1904.08920 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:41:28.888433Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T12:59:52.380964Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8062b052-dd92-41b9-8a75-13f2b2487229 · inbound

ICDAR 2019 Competition on Scene Text Visual Question Answering cites this paper.

ICDAR 2019 Competition on Scene Text Visual Question Answering Towards VQA Models That Can Read

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-25T12:16:56.505425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T12:16:41.394335Z digest=sha256:9a4127b71b3488e8c0b3c762dd279f6024cd0907ecba979f920f44130e2ec698

Observation 850faad0-0984-4bbb-ba49-1c5f2833e346 · inbound

Preserving Knowledge in Large Language Model with Model-Agnostic Self-Decompression cites this paper.

Preserving Knowledge in Large Language Model with Model-Agnostic Self-Decompression Towards VQA Models That Can Read

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:13:39.632976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-24T00:09:52.093810Z digest=sha256:0712e69e07033e5e6acda09eaf78c0b34c617f4e8ba913fc32c89ad756b6fb74

Observation 49ad3306-3508-4f09-a5e0-f5f8f5b47b18 · inbound

Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models cites this paper.

Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models Towards VQA Models That Can Read

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:40:43.111204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T22:40:39.892802Z digest=sha256:bf0881179fecef0ecb36b57a8c1cb96763af237186475578e63a7b5eb84a1c9a

Observation 79822f8d-424e-4bce-8edd-cbd1201db123 · inbound

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality cites this paper.

VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality Towards VQA Models That Can Read

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T11:11:20.827420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:11:20.827420Z digest=sha256:08a062c5883f1ef50508a4c2006aad359be2a2ff31335bc5876cce81fe481d9e

Observation 655061f3-86cb-4125-892c-1055b50def25 · inbound

DODO: Discrete OCR Diffusion Models cites this paper.

DODO: Discrete OCR Diffusion Models Towards VQA Models That Can Read

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T22:27:43.378257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:27:43.378257Z digest=sha256:84414eea775f790f59679a0344dedc1dab348bec9bead78adb3f0dd614317e9a

Observation a0fbcb54-03fa-495d-8e2f-8a30308d5bcf · inbound

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding cites this paper.

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding Towards VQA Models That Can Read

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:01.559349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T14:50:37.022338Z digest=sha256:e005d2ec64f554dec7ac33f7f19dcee945682ba7f05354637e722f17ad2122ab

Observation 5fb65f1f-13c0-4c5a-838d-6e6a4392de2b · inbound

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics cites this paper.

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics Towards VQA Models That Can Read

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:36:26.554180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:51:38.455604Z digest=sha256:72254fc98a3106457489680edfb2c66ffa74f31e61fb77fb6ec6312a2c32c925

Observation dc92ee85-2095-4ae6-957a-2a24db50dc8b · inbound

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder cites this paper.

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder Towards VQA Models That Can Read

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:37:40.212732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T13:10:14.308216Z digest=sha256:be01c400185d9faecabc52e17a19af625ac32b114006d4df524e9c331acbb89f

Observation d55f6a5c-a662-41f2-86a5-c533cd9b1e3e · inbound

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference cites this paper.

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference Towards VQA Models That Can Read

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-04T12:59:52.382133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T05:41:39.052865Z digest=sha256:e452bfe6e8876b49cb0d5e31767a0691dfd9258b58a49a7a54356c48fce6d042

Observation 5206297c-dea9-48ea-929b-a52ba970d5ba · inbound

TrustCLIP: Learning Private Visual Features via Adversarial Reconstruction cites this paper.

TrustCLIP: Learning Private Visual Features via Adversarial Reconstruction Towards VQA Models That Can Read

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-11T18:48:48.581852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T18:48:48.581852Z digest=sha256:3abc18ec8cf7fc070b80a9eef28ad1f3a3c07aa3b4397f3194aa95a5af5d5060

Observation b019ead3-b359-4a55-8593-47693e5179b3 · inbound

SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models cites this paper.

SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models Towards VQA Models That Can Read

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T16:41:28.888433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:41:28.888433Z digest=sha256:1ce005ad08960bca35faf1afbe692521a64de3b414d5783d8f4049d9d778bf06