Pith. sign in

Paper Citation Record · LEDGER

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models

As of 11 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2501.15144.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15144 v2

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:38:24.547522Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b9600f08-70cc-46cc-8d16-9e5bd6b977ae · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.015154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.015154Z digest=sha256:b2d6d5f9623b6ad002a726d39a20109ec38f92e44a20105f108f07364ebe3beb

Observation 0a930986-248c-450f-9d8e-60d6af521c45 · outbound

This paper cites Alayrac, J.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Alayrac, J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.890814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.020398Z digest=sha256:c18c5b1760e05cc8daf83181461a7185da6ce807e6f1216b2a04eb2e6dc80912

Observation fe7bcbf2-fa40-4db2-a84a-a1d8bbfed217 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.025536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.025536Z digest=sha256:c3b7fe3d9f0e50890c850b26d6beefa45653e31b7a5e4ea777ea6c7a01867e71

Observation 940a74c5-8ba4-48df-a740-0a32275092f2 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.045964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.045964Z digest=sha256:c4d3de3b4f85411dc4d6285b865df0026f6d9cf9380dd32c77caaa5a1b92026e

Observation 10576d6b-f8c2-49ee-b744-fed08389c296 · outbound

This paper cites Carion, F.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Carion, F

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.875882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.156760Z digest=sha256:bc8972061f8836e8a2b208eb19077d4aaeec47d6b1d690dae068bad0a3710e23

Observation 3c1a8817-28ba-4336-8eb7-c38a2dcd6382 · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.861327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.193224Z digest=sha256:981859f081cc6c671a60b867abf470862632dc914045bc75f62ca05999ed0430

Observation 4dceec40-7eab-41d6-9659-523db1b678d9 · outbound

This paper cites The Llama 3 Herd of Models.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.199332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.199332Z digest=sha256:105e32ec1bc3f14a96a84bbab61c5ae745c498a410f321b81a25ad7436a56343

Observation 238f5e4a-4b42-4556-a9d0-95e355648d00 · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.654140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.204219Z digest=sha256:3e0f951a13c629a15dbc3ac3b6dc1821e70e7e772141bf1087724e5d45b6f463

Observation 6b657000-0ba8-412d-aa29-55ab72f5b5b0 · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.639103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.209906Z digest=sha256:818c1c94753faa76fa47fdc236973cbff191267c562f60ec1cb4e9795f9b1b2e

Observation bec82237-c7d1-4272-81e2-49e398be6a08 · outbound

This paper cites Kv and A.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Kv and A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.623751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.214634Z digest=sha256:371172063a1d2296a2e3de63c5cbe045959fbc0c5491160fcd98486cc3445c6f

Observation 0977a93c-da7a-4db5-8d65-abb39d5a0576 · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.609875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.219493Z digest=sha256:de7fc8a92c5791bd8ba8043ce5c0c0dfcfe634df462f25c3b310bb17dda5e3d1

Observation 5d547e25-110d-46d1-9a00-7089337b122f · outbound

This paper cites Minervini, A.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Minervini, A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.420462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.223550Z digest=sha256:cb1afb5364c2863aac8bdf6766bf6c9f97d517a32d3e37e486fff3de73678d5d

Observation b1114157-1104-4bb7-a156-a3ff464717d8 · outbound

This paper cites Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-10T14:38:24.771826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.228150Z digest=sha256:933cae0f8efba7f2443e085530fdebdb7010d1dcc8512d567006ddc70df4491f

Observation d6e0c9c3-502f-4b1d-8630-292561fcc54d · outbound

This paper cites Radford, J.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Radford, J

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.389601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.233127Z digest=sha256:20673007ef5ca89ace48f5d5636d59558e229d7a4d96e561dcc0f5cf11f28336

Observation 4492f9b0-c03f-4c34-9738-dd928133df44 · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Vision language models are blind: Failing to translate detailed visual features into words

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.238672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.238672Z digest=sha256:c724265e147a10ef5da93bfd83af69d70801ef8c880c6be5fd979c9f4c72411b

Observation 235333c6-971a-48ab-9389-ae5ccdd6eb78 · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.374655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.293516Z digest=sha256:50ad219a9cab45605ed0ad6fa3d1d25a7f796963ff646a2c8f3dd229030c68cb

Observation c63b0f55-e4ae-4d82-a09a-99c4ce743a26 · outbound

This paper cites Salewski, A.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Salewski, A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.358212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.411039Z digest=sha256:f8a0f3ef5c71a0f1b50c0c0751bca0b9f7a7462bba4515f1376a10fbc70867ff

Observation 96a05e5f-1e9b-4574-be11-c733b22f262a · outbound

This paper cites an unresolved cited work.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:38:25.176242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.481439Z digest=sha256:9fc3d6b9c654e86a7f8ec73b02c4fceebf0b560ce8c44f4a7f975ceff4739c9f

Observation 6eef9f09-924d-460d-8b05-be5df45f10d9 · outbound

This paper cites FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.505902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.505902Z digest=sha256:77d0cd70015a6d94d6255419c3fe6813429c34adecbe12494de877c82055c11f

Observation e1cdf967-880c-4933-8791-c87a16da6ac7 · outbound

This paper cites Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.510393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.510393Z digest=sha256:811e3830dc0df12344e6d9d8898c9f71edbca42c35f3dc01413b116962351bb7

Observation d08b5476-9cec-49ff-883a-52c4d94ca8a0 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.514755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.514755Z digest=sha256:da30a90384c6634e029e6d27a3aad41b87e8e76f3efa0d16f372b2afe3c76f39

Observation 0cb8647d-f12e-44d0-bd35-61754c089210 · outbound

This paper cites Qwen2 Technical Report.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Qwen2 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.519030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.519030Z digest=sha256:9a9ddaab6a30378f35696049555c494ce2a187654af935f8ed1cfc8d33140127

Observation a700a717-608f-4fbb-b18d-4bd2b1c7d44b · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:24.523145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:24.523145Z digest=sha256:f9f5b17e918b95b12e63d39f68c8108588a583745ba9b026703348eb0ac192fb

Observation 9eea5803-ea97-434c-ace7-79e3569581ad · outbound

This paper cites Specifically, the shape limit is increased to 5–6 shapes per image, and the occlusion limit is also raised to 5–6 shapes.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Specifically, the shape limit is increased to 5–6 shapes per image, and the occlusion limit is also raised to 5–6 shapes

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.121899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.528507Z digest=sha256:ec22f1cfbe973a5d5739602e6efc7f960ab2a46acf8fa3708ac2b42ac08854d0

Observation 06aa8683-e30e-4912-a802-ed735ad8f655 · outbound

This paper cites This setup tests the model’s ability to detect and attribute more shapes in configurations that were not present in the training data.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models This setup tests the model’s ability to detect and attribute more shapes in configurations that were not present in the training data

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.106796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.533223Z digest=sha256:91801efedee99690417f224ba836641fc1712d46772c8bbab1565ee1a65ddff2

Observation 628f8bf5-7b8d-4947-b61f-b63869b57ebf · outbound

This paper cites This scenario assesses the model’s ability to accurately detect and attribute shapes under previously unseen levels of occlusion.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models This scenario assesses the model’s ability to accurately detect and attribute shapes under previously unseen levels of occlusion

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:25.038081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.538152Z digest=sha256:d4dfca67d9302c5e548c9de7cd9b34bcd31343777a42f53c3666c2459033e536

Observation c0a8e2e7-c228-495c-b296-9ff3fdc85bc8 · outbound

This paper cites The ability to generalize to these out-of-domain rotations Model OD Comp.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models The ability to generalize to these out-of-domain rotations Model OD Comp

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:24.879894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.543002Z digest=sha256:338fea5e9c39183caff75872568d806c5d0d4710737dfc3ff12e9a4d0e99839f

Observation 1387a2fc-f641-4791-98f9-ac2311053ecc · outbound

This paper cites Sentence Format Tuple Format MiniCPM-V2.5 Figure S4. Sentence Vs Tuple Output Comparison for MiniCPM-V2.5 Models for Validation Dataset.

Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models Sentence Format Tuple Format MiniCPM-V2.5 Figure S4. Sentence Vs Tuple Output Comparison for MiniCPM-V2.5 Models for Validation Dataset

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:38:24.864409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:38:24.547522Z digest=sha256:51daf8bd40d30483570b4c0e1bf08cf5217a0c2f80bafc48221e26d9f38b5626

Pith citing papers

No inbound Pith citation observations are available.