Pith. sign in

Paper Citation Record · LEDGER

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models

As of 11 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2501.09672.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09672 v3

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:51:25.887796Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8c55e834-d644-4453-8609-4029da6d3356 · outbound

This paper cites write newline.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.601135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.601135Z digest=sha256:b97d7f2265e9c8d4a1e7c5403623ce593aca437e8078fc0c463481b8a310cc9e

Observation e5e8a8d6-36d6-4fbd-932d-ce30f36737e1 · outbound

This paper cites Scaling laws for generative mixed-modal language models.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Scaling laws for generative mixed-modal language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.609109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.609109Z digest=sha256:25683ec4ffd1ed67a0df5512c23c6f41e00ed8710996fd4d44d39defcb5ec9ed

Observation 965ccf86-cf63-490a-897b-12180214a7c0 · outbound

This paper cites Lawrence Zitnick, Dhruv Batra, and Devi Parikh.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Lawrence Zitnick, Dhruv Batra, and Devi Parikh

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:27.092949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.614692Z digest=sha256:498a5f9613b805545b7c84adca81915f202eef818e9c4dfdb36b9ad33c71edef

Observation 6d0c078c-27ad-407d-a6ed-4e6ece8b1d9c · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Pythia: A suite for analyzing large language models across training and scaling

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:27.070866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.626791Z digest=sha256:df539a4cdd71d46888be11d6e10f150e9fc570e27f56f7fda91c71312a3ac2a9

Observation c1e84e37-393c-4d22-8f32-ba5994008391 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.633642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.633642Z digest=sha256:5fa86398a78027955d985b977a77e6272b53b713ee52b2cd14c0e21f56ff66bc

Observation 4a69c8f5-4d07-41d0-a6df-3c0be1f202f7 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Reproducible scaling laws for contrastive language-image learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:27.024753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.639573Z digest=sha256:dca090edcc8b0d76220a16159611dc9d50a681d92934809c921b8ec635025ab6

Observation 071960fb-a29d-496e-8816-80874bbf38f9 · outbound

This paper cites A coefficient of agreement for nominal scales.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models A coefficient of agreement for nominal scales

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.645192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.645192Z digest=sha256:6074f92c40ff76cc554076a8b4cfb255090eafde89684dff4e33f1e0eefe133a

Observation 988c5c71-494e-4961-830f-8d3f9587a9b7 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.985702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.659470Z digest=sha256:f801c89961f24341d06d185423202ddf569ec47f1aca317a8028af8a299b9c98

Observation c6eb6c26-6232-4cd0-9d38-a5c1ee17599a · outbound

This paper cites Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.963085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.665895Z digest=sha256:27d888f8d2ed21e5ae2d884f5fd6951c5dbb5e477139636d6647e06af19781f9

Observation 6c988c71-d243-4fc3-9eef-0e37e0e2f402 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.924608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.671507Z digest=sha256:d00164b4b26bfe86bbdd80e787ef14ba471e0a5d14ae3d7755c8aaf0f6e356d2

Observation 774765d2-7178-4461-83e5-b0083ccf2e89 · outbound

This paper cites Hudson and Christopher D.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Hudson and Christopher D

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.884643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.679340Z digest=sha256:953ebdc5353c61d4e5fff5f8960a90debec51fb62dd7f681ea266037dac65d4c

Observation 8db6cee1-d9a2-4de2-87ea-42797afbda03 · outbound

This paper cites Scaling laws for downstream task performance of large language models.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Scaling laws for downstream task performance of large language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.695590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.695590Z digest=sha256:f95ceb957296900d5a3735b3a931e9366e287db2775a089889b0b9aae9dd490c

Observation 3b5fe3cc-cb98-4dd6-aad3-c45194caa9b6 · outbound

This paper cites Large Language Models as Automated Aligners for benchmarking Vision-Language Models.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Large Language Models as Automated Aligners for benchmarking Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.705056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.705056Z digest=sha256:6ea499246c95d8507b73e4dad436d90a1033465c991ef179f6fea4e0a68941d7

Observation 8ef415a1-fb34-4b4f-abe4-43da84becfff · outbound

This paper cites Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.857063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.715803Z digest=sha256:d333ddac7efe309577d91aa926d25cb21adfb00967f2202202297676f3bf98e2

Observation 6b7b52ec-d510-4b2a-ae3e-41f8d6522a8a · outbound

This paper cites Benchmarking Cognitive Biases in Large Language Models as Evaluators.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.722030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.722030Z digest=sha256:8929d28e4688da3ef15ccc081c0e8e702f7652a1634cea1b01146135efab8413

Observation e27249bc-83bf-4033-be89-10cf8fb087ee · outbound

This paper cites Shamma, Michael S.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Shamma, Michael S

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.835213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.727861Z digest=sha256:777ee18836fbde3a4720c4c3c8f67a045dad870d52933f3c03f79023eb461f20

Observation 467e27ca-f6fd-446c-abb4-d96ef405a73f · outbound

This paper cites Richard Landis and Gary G.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Richard Landis and Gary G

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.732913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.732913Z digest=sha256:d0574984f45e3b6e532ff47d53640f943e546a1836ed72b5c2a0caa654660aa1

Observation 295d2df2-2f49-4ca0-9b70-0c7fd212d330 · outbound

This paper cites Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.740068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.740068Z digest=sha256:5bb21d3b696035b4e01c834edf9c23ca02dc123b449b596fb6cdf9402f4368a6

Observation 832ef200-8aa7-41d0-b0c2-4e95a6ee8bcc · outbound

This paper cites Hashimoto.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Hashimoto

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.813110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.746643Z digest=sha256:c5a0bea50a1218bbefab181b23292fe7dc8d21757c4ba1b7028ad3141351c071

Observation 42bc1c55-1abe-433d-827b-4849628fe22f · outbound

This paper cites Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.758465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.758465Z digest=sha256:2c491b594df6658fe3d26d0dbb78ef5ebbb1aeddb216e452c7f0d28ed8f22cac

Observation f93b1673-c777-4a6a-804c-9f1346b0fbe6 · outbound

This paper cites Lawrence Zitnick, and Piotr Dollár.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Lawrence Zitnick, and Piotr Dollár

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.785676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.764824Z digest=sha256:de5349cff4a65daa64585d0f67070699509159046ffaeed4942e5c3f7a3249bf

Observation f734faa6-9cdb-458f-896e-ed141fc023fb · outbound

This paper cites Improved Baselines with Visual Instruction Tuning , 2023 a.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Improved Baselines with Visual Instruction Tuning , 2023 a

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.768831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.769256Z digest=sha256:17de16278c3e63631fef378c84aefe17a42cb811b4e4b1d4120fc898f17d5783

Observation 2334d8b9-6bf9-41b2-a240-d44b5099203c · outbound

This paper cites Visual Instruction Tuning , 2023 b.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Visual Instruction Tuning , 2023 b

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.743919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.774490Z digest=sha256:bd6041af7abb6edbfa8d699ba7b510c29d30f4a0c3d425bf432de8c65c9608c3

Observation 02fdbc1d-1582-4b13-afc6-ae23afd06637 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.724928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.780485Z digest=sha256:475a47b2cd5757e8069c3d14f24895e17d77b0f404327c9979fda73f70dba12b

Observation ea1380d3-e1ef-4445-b172-78babf8c9908 · outbound

This paper cites OCR-VQA: Visual Question Answering by Reading Text in Images.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models OCR-VQA: Visual Question Answering by Reading Text in Images

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.788155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.788155Z digest=sha256:aaed206df3a710a27edd1147d3a485c3bb6e5b3af787644ffe9d1cd041c1dab0

Observation ee15cd68-f4a2-422f-953d-733d0c764825 · outbound

This paper cites Gpt-4v(ision) system card.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Gpt-4v(ision) system card

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.699551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.799891Z digest=sha256:49f890a42e72c03adeeb09b57be81264636f94d4c37943091e9d8e3735187069

Observation 4899d78b-0e84-4e0c-bf73-8b28a0560d7b · outbound

This paper cites Openai model documentation.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Openai model documentation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.669145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.807624Z digest=sha256:878a3fa6d22fe3270748f8335d614356bde65bd626fed54919984bec8f61cb0e

Observation 85cee05d-d53c-48d8-804c-823981f7e093 · outbound

This paper cites Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.645480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.820188Z digest=sha256:4369b1560bcdab7f06c551ad30a807cfb9a3b2a3304bdefbae32d5d7ff8c8b84

Observation 5520d006-a753-4346-baf9-c42044b301bb · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.827746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.827746Z digest=sha256:412ea79134b212e9f2905ad557a29141ba3218df40ba003767fa645409b3093d

Observation a6f9c3a3-9ba9-439d-9ba8-d1b2ac43344c · outbound

This paper cites Revisiting the train loss: an efficient performance estimator for neural architecture search, 06 2020.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Revisiting the train loss: an efficient performance estimator for neural architecture search, 06 2020

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.595617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.833267Z digest=sha256:725eb00d1181e325eea2f869ec6ea2461e98b30743b17f3d84c8cdea98022a8e

Observation 2671ea9d-312b-46c0-9bf5-8dc13fd2d1c5 · outbound

This paper cites Towards vqa models that can read.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Towards vqa models that can read

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.841536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.841536Z digest=sha256:e43492c510c5adf43c1d4f585d6a551d5b2c93dcd93d05f2241abbad13617ffc

Observation 20c47a55-3e47-4ebe-9ece-ddbf32f870df · outbound

This paper cites Generative Multimodal Models are In-Context Learners.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Generative Multimodal Models are In-Context Learners

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.852957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.852957Z digest=sha256:1cb167062d1faebb499691288f8a33a8863cf946df6ceaeec26dd375147f302d

Observation 85d87149-19c4-45c4-bece-1d2d06fd2cbb · outbound

This paper cites Székely, Maria L.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Székely, Maria L

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.866465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.866465Z digest=sha256:ed633eedcb8a5a978fde7b541e339939423a4d546e3d482fb79ea881fba1a1e7

Observation 454a078a-a3e9-4704-9f36-cc8df588f1eb · outbound

This paper cites Style Over Substance: Evaluation Biases for Large Language Models.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Style Over Substance: Evaluation Biases for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T19:51:25.872740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:51:25.872740Z digest=sha256:84f2c18fc3ec0e5168f273a29dbc25b0317b82504754dd34b02b8a52a4479a23

Observation 223c606b-c254-4773-86a6-2f50624c24d7 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities , 2023.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities , 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.542116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.880433Z digest=sha256:d3a52ee6a1721885ccaf307898814228da9d928da5fd94a7b3839f1989ca859b

Observation 7f77c188-126e-456f-9203-a7c136f166a3 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models Xing, Hao Zhang, Joseph E

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:51:26.500908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T19:51:25.887796Z digest=sha256:7e64f64339017fd739f8dbeb85648b90f032722fad51042ec3026068ada87de3

Pith citing papers

No inbound Pith citation observations are available.