Pith. sign in

Paper Citation Record · LEDGER

Can Visual Encoder Learn to See Arrows?

As of 21 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2505.19944.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19944 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:07:20.058994Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 39471ad9-2149-4212-ab77-4c84f03b3515 · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

Can Visual Encoder Learn to See Arrows? Understanding intermediate layers using linear classifier probes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:17.749975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:17.749975Z digest=sha256:50228ac96b466f9c9632cb2988a0893bb63751b0178ad5ca88f131eb3de25803

Observation 15425ba8-7583-45ec-bcb5-af119347e278 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Can Visual Encoder Learn to See Arrows? OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:17.838017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:17.838017Z digest=sha256:cba60b0c366615448b1056a0e83baab563c25c818cbaea436e2a38d2ca758491

Observation 6840dfbf-15ea-4ce4-b255-7e5c5f94f55d · outbound

This paper cites Neural codes for image retrieval.

Can Visual Encoder Learn to See Arrows? Neural codes for image retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:23.472836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:17.881086Z digest=sha256:23785cb88e369cdc2da17a0345b02ff1d00a743a17fc21ec2abc159dfefbd668

Observation 3c66b4bf-0e7e-434e-b2d1-9df26c205382 · outbound

This paper cites GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning.

Can Visual Encoder Learn to See Arrows? GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:17.952748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:17.952748Z digest=sha256:a8c58d17a58be04d597b49497f2e3737a22855dca0e63ce95f53094d5b33920c

Observation a519e5e2-26d5-4070-a650-1918f41a2bd2 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? InAdv.

Can Visual Encoder Learn to See Arrows? Are we on the right way for evaluating large vision-language models? InAdv

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:23.314526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:18.029280Z digest=sha256:3087326c7d3f4ec9c697362f15f735bd8f65b3b4a047170d52ad9c5a9ed285dd

Observation 01660096-ceee-4595-967a-1b7a508dab39 · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

Can Visual Encoder Learn to See Arrows? PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:18.135269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:18.135269Z digest=sha256:0df52ed3fe15485bfcccdddff359cd0ad96c0dd5ba7bc8694b8ef73714458b7a

Observation cbf9bf9f-8a64-4cde-8b7d-f57ec038c2ad · outbound

This paper cites What you can cram into a single $&!#* vector: Probing sentence embeddings for lin- guistic properties.

Can Visual Encoder Learn to See Arrows? What you can cram into a single $&!#* vector: Probing sentence embeddings for lin- guistic properties

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:23.141131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:18.206189Z digest=sha256:7829ee143a2cce042ecb0f05fb6da5a3ddf9298ef615229ba87a8af33ef4cfc7

Observation c5f4f1d8-5e21-41dd-bd76-3b8db9d45745 · outbound

This paper cites an unresolved cited work.

Can Visual Encoder Learn to See Arrows? Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:07:22.993229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:18.275398Z digest=sha256:320d8871e28a8265bc08c62d3a769b20e257998b843204760c17cd6039b063f8

Observation feca93f9-d9bf-46aa-bb58-2b31043ea144 · outbound

This paper cites an unresolved cited work.

Can Visual Encoder Learn to See Arrows? Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:07:22.873637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:18.339017Z digest=sha256:6e104a015a9bc1ebbae6811c4ec955deaf2109e5614cd26cd0b530fcd415ecf8

Observation 10cd801b-d7d7-47e9-a18c-c49f6d9f1bd6 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Can Visual Encoder Learn to See Arrows? BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:18.386462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:18.386462Z digest=sha256:1a9ddb14df23b40e9442d1708981ed1fd1afd3285a70e0ddda32bbe91040f61a

Observation 9256d60c-dabc-4af2-b317-1c2f7b9f2ee2 · outbound

This paper cites Visual Instruction Tuning.

Can Visual Encoder Learn to See Arrows? Visual Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:18.468295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:18.468295Z digest=sha256:b73bf868041ab1dd502f583bdf85c2703eb69c373ea7dccfb1f43b97a1b2b30f

Observation e21d1d39-45dd-4cbe-9757-dd4859b1091e · outbound

This paper cites LLaV A-NeXT: Im- proved reasoning, OCR, and world knowledge.https: / / llava - vl.

Can Visual Encoder Learn to See Arrows? LLaV A-NeXT: Im- proved reasoning, OCR, and world knowledge.https: / / llava - vl

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:22.708865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:18.580748Z digest=sha256:14ba3ddd902ad1acb02288704f8c830f4c0fae90d998c1150731419e93ec50de

Observation 3d4b7091-b4eb-404e-a5c5-cf13591b1663 · outbound

This paper cites MathVista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.

Can Visual Encoder Learn to See Arrows? MathVista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:22.554493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:18.662136Z digest=sha256:8be36abbad81c434cc3c2ec151548de77cc9a96e04be0490c55f585fcb75bb9b

Observation 0ee13ad4-fb7e-4074-bc2f-54c22c545f54 · outbound

This paper cites CLIP ViT-B/32.https://huggingface.co/ openai/clip-vit-base-patch32, 2021.

Can Visual Encoder Learn to See Arrows? CLIP ViT-B/32.https://huggingface.co/ openai/clip-vit-base-patch32, 2021

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:22.368834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:18.768103Z digest=sha256:6f5dd4f6103926a53e89087a6d07dd6c17c51a67fd3d3bfdf857ffc198f78f4c

Observation ef4f94de-b285-4dcf-b4fb-68e9e6189ad8 · outbound

This paper cites CLIP ViT-L/14-336.https://huggingface.

Can Visual Encoder Learn to See Arrows? CLIP ViT-L/14-336.https://huggingface

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:22.193856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:18.860984Z digest=sha256:69b64d95b9d28c2af4afa7484505ec2c21bb3c56814ce60d4063706fd606d1af

Observation f2afb6e4-0f69-402e-803c-94fd551c98a8 · outbound

This paper cites GPT-4 Technical Report.

Can Visual Encoder Learn to See Arrows? GPT-4 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:19.043623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:19.043623Z digest=sha256:25db6ceb53fd642cbfa8e8a46ff0e726f45ff317c432cc4e208cd567c8bcf487

Observation 021ebe13-dda6-4a68-9d94-6be23c62d492 · outbound

This paper cites Hello, GPT-4o, 2023.

Can Visual Encoder Learn to See Arrows? Hello, GPT-4o, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:21.849237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:19.114429Z digest=sha256:906159eb3f8a2f46af8b62317842fffbad13eae3e264c3778c642187c9d07f49

Observation 4a48a3ef-7915-4365-80ee-d887729d4ad4 · outbound

This paper cites Language models are unsu- pervised multitask learners.OpenAI blog, 1(8):9, 2019.

Can Visual Encoder Learn to See Arrows? Language models are unsu- pervised multitask learners.OpenAI blog, 1(8):9, 2019

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:19.211866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:19.211866Z digest=sha256:1d66c822e81994de82f4ff67d1775894106e743c5d6d26f9934131a53642c1ef

Observation bae674e7-28f6-4f70-9d20-ea7ac8cf3d74 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Can Visual Encoder Learn to See Arrows? Learning transferable visual models from natural language supervi- sion

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:19.321307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:19.321307Z digest=sha256:6e8419c63fd337ab02f048f065fe06558167498311d25039f973d5d2190310f8

Observation 583e82d8-3d67-4a2a-b75b-3dc7d83a6f03 · outbound

This paper cites Vision language models are blind.

Can Visual Encoder Learn to See Arrows? Vision language models are blind

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:21.645904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:19.426053Z digest=sha256:4e664a8fe199a9fb41d0d521b7bae31001034e9d860a16eedfbf5135b34b236d

Observation 269978b4-6d20-4519-b803-4644348d04fa · outbound

This paper cites CNN features off-the-shelf: an astound- ing baseline for recognition.

Can Visual Encoder Learn to See Arrows? CNN features off-the-shelf: an astound- ing baseline for recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:21.450556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:19.484458Z digest=sha256:5264495892c7343438c96d3df6ef41962e26f862e3948ea8f43f0c958be686e9

Observation 122c9e54-676f-46a2-b903-7e38d0211383 · outbound

This paper cites FlowVQA: Mapping multimodal logic in visual question an- swering with flowcharts.

Can Visual Encoder Learn to See Arrows? FlowVQA: Mapping multimodal logic in visual question an- swering with flowcharts

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:21.259613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:19.576412Z digest=sha256:d9262a3c95381ee8855a3cf43b99ae43de8d3f6d83657432a41b754f8af98c7b

Observation 6b44838c-7224-42f5-87f6-6b742d9b635c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Can Visual Encoder Learn to See Arrows? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:19.645208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:19.645208Z digest=sha256:e317839a834673902012efd1382fa6ff81b755530708cbd0aed98c3eb74c3128

Observation da8f3efc-55c7-4383-8605-3a9269e26372 · outbound

This paper cites How well do vision models encode diagram attributes? Inthe ACL 2024 Student Research Workshop, 2024.

Can Visual Encoder Learn to See Arrows? How well do vision models encode diagram attributes? Inthe ACL 2024 Student Research Workshop, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:21.086930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:19.770636Z digest=sha256:63762fb7a2a9deda6a446b005ae93c9f0f589ce2c4e42b643510a5ac281e43db

Observation 4d5510b3-200f-4346-8ad0-b15a579ca76b · outbound

This paper cites MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert AGI.

Can Visual Encoder Learn to See Arrows? MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert AGI

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:20.800916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:19.873890Z digest=sha256:74959189f60d7ab15d762152809b20d989477855796dba8db76899af34de2a80

Observation 53cb6495-33d5-4ab9-892d-a7e523cbd428 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

Can Visual Encoder Learn to See Arrows? Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:20.608602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:19.966361Z digest=sha256:2185594910a5037c64651000aa8700c216fff3c3d4c605d017fb377e080783d1

Observation 96ce0fac-082c-43f0-8b31-9f7c8e30d11e · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Can Visual Encoder Learn to See Arrows? MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:20.058994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:20.058994Z digest=sha256:ad621efd0c24398cc805c5a9ae4158856324c01aad32469184b152e66ff78078

Observation c3a0f9ef-7804-437b-8a68-3ca0b5960f85 · outbound

This paper cites Accessed: 2025-04-12.

Can Visual Encoder Learn to See Arrows? Accessed: 2025-04-12

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:22.028089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T14:07:18.935665Z digest=sha256:8f5d0a0426691b69a029dd0cda292ab3f564b16c25fc618d1323bcba836915ef

Pith citing papers

No inbound Pith citation observations are available.