Pith. sign in

Paper Citation Record · LEDGER

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models

As of 3 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2604.04172.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.04172 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T16:45:17.022062Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T16:12:23.659005Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-30T12:16:12.909640Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact21
  • verified fuzzy2
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3aad4050-c983-46f0-9cdb-2582f8018c3e · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.872301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:8154ad76202af319e40986e7c7722ae97547504ca2e67e4663483a4f47c89b2a

Observation 29923d39-5047-4c56-a629-c4fddb464879 · outbound

This paper cites SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.127435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:59cb37186c21ec1e7659e2f0f0ee70fa7e76ab694b405cd6cc26a900cdfcbb49

Observation 637a6ca9-398d-4514-b9a2-aa414119ae6e · outbound

This paper cites Figure Captioning with Reasoning and Sequence-Level Training.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Figure Captioning with Reasoning and Sequence-Level Training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.153640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:fc57457ef1c54999b6c6e7e528fa64b1143f6184b4d31842d1c4d97fdca51aaa

Observation 3c45bc01-ce71-4cf9-90a9-a5378ebe2afe · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.856423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:0ebc7f75604c9bf1af50289ad634bce13a6d958fba09692f6d165cdba4cb73a9

Observation 4906e079-76c2-4fdb-bc59-428e16635311 · outbound

This paper cites TIAM -- A Metric for Evaluating Alignment in Text-to-Image Generation.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models TIAM -- A Metric for Evaluating Alignment in Text-to-Image Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.134881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:95feacf5863561a8b42ae0ad6421eee77904f85690f03288ee31bb6cd5c3cd47

Observation 9ac896bd-787d-49b9-8aa1-0384152d2871 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.853550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:bd3c328a001b207c166fbc6e648b67e154bd2c51c2ec3dbcfe3780fe9f08a681

Observation aa87a366-a820-4db3-9be1-bdd36d720a5b · outbound

This paper cites SciCap: Generating Captions for Scientific Figures.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models SciCap: Generating Captions for Scientific Figures

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.157019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:ce6e18fad1acdf22f44beedb0fd18d5cfb1addf881ac580717d9165782748c53

Observation 1a8df55f-d682-43a3-9150-19411977c5c1 · outbound

This paper cites mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.131525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:cb73dc2f068f9d48b4c59605420d5c243005696e6d8d8fe0ddec4f02240d3830

Observation 1a663a3f-b1ac-4b2e-aa3d-8ca67ed77482 · outbound

This paper cites GPT-4o System Card.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models GPT-4o System Card

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T16:48:03.142195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:3b744d627142c5fb278e5ef6ebee80aa8d40f3de98af801155effdc57e56e1b1

Observation 6ea16a10-3e8f-4df2-98b2-5816654ae642 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.859100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:6c8b2199bf9f50def26b0ec630ed6681e8f7e2c9e41f5517da9bebf4451f1d4c

Observation f9e1b439-4b0a-440e-8fe7-509dc4e9a0e0 · outbound

This paper cites DVQA: Understanding Data Visualizations via Question Answering.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models DVQA: Understanding Data Visualizations via Question Answering

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.145919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:5890d878aad97ee2b22eb319c881f2364bd0652a2500a9598547fcf5692304c5

Observation 57e0cd0d-85d4-4dc5-8cb7-eb0fa457a6e5 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.120684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:7d09812a4e14c4d2b009549d4cb6f6447f27aa2d283ffa41b38036782400dff8

Observation 36d86ebd-6a76-4eb8-bf23-5d56f63093bb · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.894565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:2bb028b84cd0be8ec311f687c6dfba1924460f3166d8f07bf99b3562caf48366

Observation 5ab9584a-fef9-4d1f-ba7b-4957d8e3abe5 · outbound

This paper cites Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.177074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:a333822a2dd2499f9000dd44a33f979cb4b9cdab3326894a85e048f4d29ef427

Observation 4241a15d-582b-4c18-aafa-a8dd0b191c77 · outbound

This paper cites MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.187477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:d538d10ec14238c2856abbd7f246f0529304f0bfd25067318448237b049da742

Observation 728949d0-95c7-4b12-8945-2e66611f446e · outbound

This paper cites Aligning Large Language Models with Human Preferences through Representation Engineering.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Aligning Large Language Models with Human Preferences through Representation Engineering

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.173387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:83a975e987ffc6aeabb49d5fa04397b08a7b5f18427948184a32fde2a61316bd

Observation a68b9895-21cb-4fc4-bbda-60bd4415c2d0 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.889617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:2fcd8578416bf347fd0be62cc0d2c4c2d7ef7d7a8ca79c6dc36d2e579c1195f5

Observation c1c364f7-8a10-4067-b9ba-964d218cacd4 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.904571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:b6338dee4cdac3273dddb406f12a0dddea250affeaa130939fa40802870a9259

Observation bd49d785-8948-4577-a935-5beb3f527ee1 · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-13T16:48:03.160299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:cac5d2bec7827352f1f7772f80fd9584af92693015610bfba2e5172bff7a11dc

Observation f1cfd909-86c2-4065-8e3c-e448d9621a5c · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-13T16:48:03.169996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:e50adf78ae816f676d9df56f4039eae20d8fdb1f4c6345a1be823f79b64c3dd7

Observation 8d6ccf26-197f-4b07-a06d-602920319e59 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.902116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:011773a0849a607be77383787a1037fbe2a6ce4ecbbad153ba32ab5ba99aedde

Observation 670a2808-f012-4078-851e-45ada37ec5d8 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.897160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:5909e2627bed64f0e3bb37693e227159038ddcf79c79ba5b452fc7b90aa044b2

Observation ffe3a154-5aaf-4027-80b5-847567de1cfd · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.899686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:1cc2768f715284972c678dbf49e991fa4b7d3a1f41b7dd9cf2c5ca105e21b63e

Observation 6a1702af-bec8-460d-b80e-41888bc8672b · outbound

This paper cites SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.163678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:366d6965be7700dfce2d6211b8f562926a006295f1e9b541dfc81cbb2ff341d3

Observation 99457d57-9e9b-488e-9493-2aa71d9cee3b · outbound

This paper cites StarVector: Generating Scalable Vector Graphics Code from Images and Text.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models StarVector: Generating Scalable Vector Graphics Code from Images and Text

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T16:48:03.166902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:ac28c57131f8668abfcb213f9127be45629b3bfa31d0a1de53f93958233fbf00

Observation 87f3003e-6cf7-43a5-9a1f-829d500ca49d · outbound

This paper cites OCR-VQGAN: Taming Text-within-Image Generation.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models OCR-VQGAN: Taming Text-within-Image Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.180740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:9b52d96f2c50bedb34c7c15cf4464961993b9a284a80d4eff08e300d6444cccd

Observation a2666c8e-65e4-4d7e-bc8c-bf595c51b326 · outbound

This paper cites FigGen: Text to Scientific Figure Generation.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models FigGen: Text to Scientific Figure Generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.184022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:fc2cec402dcd1c41779e15d6eda7af9e87521b0796d6a8cf7b85530a6ae7e7b1

Observation 8fcb09b1-f074-4784-a7d7-675820ef6211 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.891982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:b2e42ee81794e576dabd63abf245e37466b26092529eff9c5ebe916241efb11d

Observation 77b4005e-1dd4-46f5-b1b1-8be724f3edf7 · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-13T16:48:03.149583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:9f7eaf6ff070c45ea21a4d306240678f161b343002d0fb152b6e978400ac62c5

Observation a725aeec-e143-4fca-a19f-070a6fa01b0b · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.867052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:ba3be824f5a8b09af39af5f12538b401fe53db56277a88f6fc9368cc2135a62f

Observation 8aeb5649-b9f0-406c-a534-cf6ff8eb1f1b · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.861534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:b5dd5d6b1d4e431b6df08981978d5594c02ceaba0c343f6b82fc6b0676ef9eca

Observation 841567f4-5b03-4a87-a5bf-e8314b9bf0e4 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.884429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:26390a292ff8f9171fbeb8ce6d73fee737ba48909c44775fbed0e487c630ac4e

Observation a2252e7b-776c-41c1-a6b9-0db6ed906358 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.875150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:5195c9dd201ca112ae3e5bf6ff5c742e78e01a663aa35a6cee3c4bcf692307e8

Observation 6d1625b7-d971-470e-97a3-00b0d46a3d21 · outbound

This paper cites CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.138494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:de2f4aaf10c64b78a567e8eaf83d9a403d205b0487dc4ae3341e7d837404bcfc

Observation e977d784-24ad-476a-aef0-f7058ec1b1fd · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.887186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:6cbaa5013e226b97ed2dce2c209039480755e17995879218f3df23086905d324

Observation 2098109e-c795-4eb8-a2f0-e990eeeb7b7c · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.869763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:a118dfdb246890490deaaf93aec5919b5e0ad9ac3764d27a7f3b8c8c51f9a89b

Observation 73f37f0d-40e9-4022-b32b-978c601e84b5 · outbound

This paper cites SciCap+: A Knowledge Augmented Dataset to Study the Challenges of Scientific Figure Captioning.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models SciCap+: A Knowledge Augmented Dataset to Study the Challenges of Scientific Figure Captioning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.123892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:377a8f1797d6b3efebb62c718392eab13e4b6a27fb22058be92ff194fb748399

Observation 0b982cf7-55e5-4ecf-82bd-c1a2c83e607e · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.864256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:8282e149c761ebefb110a746ad44f5bb211be92602c6b99a102112debcbfafa5

Observation d9f239f0-318a-4107-9de3-8ec44bafc90b · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.877915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:11256dfd86bf4553d141bec4d71e731c7ac248b8fd13a2af028e688fd71d0b9c

Observation 83ad4aca-ed07-4ca4-be35-c520170bfa20 · outbound

This paper cites You are an expert on translating academic writing to visual specification.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models You are an expert on translating academic writing to visual specification

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.115963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:be14ee6829cfd96246a8cff52849c685ca6889ae693a26142510dee9b31dc551

Observation 2487533b-bfc4-4ebc-8995-757631700859 · outbound

This paper cites Autofigure: Generating and refining publication-ready scientific illustrations.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Autofigure: Generating and refining publication-ready scientific illustrations

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.111073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:088b3fe04ea459a8b5680da71f6f80a01760a0570d2d5199f820b8faf6864b79

Observation aaddee79-f288-4f75-8b22-6b81650df147 · outbound

This paper cites online" 'onlinestring :=.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models online" 'onlinestring :=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T05:58:14.881515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:9143ce668767769f6ffb35f8c6cbe720dcf7b0d2fe373c430f9691da1a4895d5

Observation f235db4d-e6cf-4d13-9104-83fca1482a0d · outbound

This paper cites write newline.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models write newline

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T05:58:14.850483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:6ff9aef2ee95bcdf2fe7d967f8b8ab67fa692b10cf6c66bae9205a9415e4dfaf

Pith citing papers

Observation 079c2a9e-6acc-4a07-b591-945fb24f61b2 · inbound

SciForma: Structure-Faithful Generation of Scientific Diagrams cites this paper.

SciForma: Structure-Faithful Generation of Scientific Diagrams GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T16:12:23.659005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:12:23.659005Z digest=sha256:271e74e9cfd68af7ab12c9e8980025d36be949d5f8c9bdda7e88310e49a98454

Observation 87fa53f8-e43c-4cdb-ae0d-e5e9fbe6fa81 · inbound

SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence cites this paper.

SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-30T12:14:15.738599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T12:14:15.738599Z digest=sha256:8a05cb8e91307b570fd3a701fc00b6fc881cb8d229c348a660c54b77bd7a0201

Observation 39b48dcb-1e7a-4f9f-9980-a454bbd1f682 · inbound

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context cites this paper.

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-30T11:41:21.736732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-07-30T11:37:27.095506Z digest=sha256:bc67ccf605c84a519e1f9532c0d8c622f2ee71cad8d89fefe71df4dac8a94f8c