Pith. sign in

Paper Citation Record · LEDGER

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models

As of 11 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2501.12418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12418 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:16:25.745948Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3f1bcd6-c70b-4d87-898f-ed34590ff693 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.478505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.478505Z digest=sha256:7c7d3e38670f19aaee3067eca19cd4d0ec0cb8a98a5a2797e94a9ff28480f534

Observation 78ad0596-1cf7-42be-9a48-27e086ca04b9 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.489908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.489908Z digest=sha256:07936c521f54bc12528966e50daf4824a8f3d17232dc09ff53e279ac71d6b2e7

Observation fd44d442-f1f2-4410-9704-411d6df1e24c · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.855782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.500593Z digest=sha256:15cfe054135b947fb95399be37ba85d062055f68f59df697468830e53377177c

Observation 5d661999-f36f-4037-b99f-76c5e0f1400f · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.836411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.506459Z digest=sha256:430c5828f8b48f44233a7967120efdd43a71f73743195c2940259ef963bbfdaf

Observation 6cfd4a4f-804c-41d2-a245-989a647ecb83 · outbound

This paper cites Learning to Plan and Generate Text with Citations.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Learning to Plan and Generate Text with Citations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.512501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.512501Z digest=sha256:1686c19212101dace8a8823f030f05e50ef00502399e51587033806fa1eed613

Observation 3ecb4581-1e0e-4db6-9bf9-3adb572429e4 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.518335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.518335Z digest=sha256:4ec6338a86942ef719fefbe57a1642f23bcdf186a9a0d43ce956cc379df32250

Observation ca1e3466-f52a-470b-b095-c64f46e697cc · outbound

This paper cites Enabling Large Language Models to Generate Text with Citations.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Enabling Large Language Models to Generate Text with Citations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.525389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.525389Z digest=sha256:6614f62a70afa6710e53a4ee95c264b3bc40320a146b4c98a1d407943101dbf6

Observation b3931a67-f4e3-4821-96e8-8b3a7ffa821d · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.531582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.531582Z digest=sha256:f82519bb6bfecbb677df04379a19eb92e997ad07e32999d94029c8cdaa5f529d

Observation cd25219e-df64-40b9-b348-54153b77cf7f · outbound

This paper cites A Comprehensive Survey of Retrieval-Augmented Generation (RAG): Evolution, Current Landscape and Future Directions.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models A Comprehensive Survey of Retrieval-Augmented Generation (RAG): Evolution, Current Landscape and Future Directions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.537683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.537683Z digest=sha256:5d5a104712d1d95c7b5ec67df664b8b54dae71d209a86e7c0fd452d156267566

Observation 2440a93c-3f7d-491a-a1f3-68822f06b681 · outbound

This paper cites Towards Verifiable Text Generation with Symbolic References.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Towards Verifiable Text Generation with Symbolic References

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.543464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.543464Z digest=sha256:76112c14785ac8b10fc305d32c7afce31589a78e36f2cea429f3d9f7f3a43ba8

Observation 61773371-c9bc-4a60-bf0b-9c0035504e97 · outbound

This paper cites Training Language Models to Generate Text with Citations via Fine-grained Rewards.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Training Language Models to Generate Text with Citations via Fine-grained Rewards

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.549812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.549812Z digest=sha256:ba63c140eea000cf1f6fd693ce2960986aac060a94026fda30674fbd60c63ac0

Observation c2001045-7c55-4b96-a412-b82bd13399a4 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.819478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.555646Z digest=sha256:686fb4d609e08b7e0c9287c2e23088903281e959703368a50e7cf2f0e4e9172f

Observation 3942e530-2209-421a-bc6d-ed3cbdb689c9 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.800976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.561384Z digest=sha256:4ec633d6cc2f0dc09bf74fca83aac84d3a81bc50ce57d3fbaef12927ded43213

Observation c9148c76-5f53-48ab-a03c-9e23f89f2945 · outbound

This paper cites GPT-4o System Card.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.567338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.567338Z digest=sha256:0c761abf079f9942056815b0f51bd36f59470034a67679c8f95a105aee0dce79

Observation 1b394101-d2d2-4ddb-bec2-f56c560aded5 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.779772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.573600Z digest=sha256:86d49deb6044e5e8fd3360914c94aaf2e8832ecf314674a86ca9329e978e47ca

Observation 089bacb6-b49e-45d5-9fcb-a5911443e629 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.755031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.579592Z digest=sha256:78057998ca7345ea2ca11e772700e35ff94631024cef41c0add169940882af56

Observation a8484ff9-f8e0-4ab7-a617-943a4ead9f4a · outbound

This paper cites Towards Reliable and Fluent Large Language Models: Incorporating Feedback Learning Loops in QA Systems.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Towards Reliable and Fluent Large Language Models: Incorporating Feedback Learning Loops in QA Systems

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:16:26.120152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.585379Z digest=sha256:b1a6a26a8ab1d03308104672d85d18cb3e36fd7d2c5182f2bf913b1aa2a5cc46

Observation 649546f3-a34b-40fe-ac77-b628926cbb85 · outbound

This paper cites Improving Attributed Text Generation of Large Language Models via Preference Learning.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Improving Attributed Text Generation of Large Language Models via Preference Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:16:26.089678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.591478Z digest=sha256:93be07374afdcbfa592ab7e0f7f58fc9e99c1f28a9802ebff0bd36a671cea29f

Observation 97b966fe-596c-4582-a2d5-ff7f850a019a · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.734247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.598412Z digest=sha256:799e4c793268c3900ad9a92ebebb5f03542bd6e1757b32a1e0ca55dae6d49b1e

Observation d097dc70-7d00-4245-b8a4-ab888f08bf8f · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.714210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.604077Z digest=sha256:371b206962d3c42d4a5964e03f51fdba8314f765b4a97c2a1c6fe60766b0805e

Observation 4799352b-a7d1-4f1b-b085-ad4609652a67 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.690333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.610224Z digest=sha256:74e17886a283aafbb2a36651713fb2a74da8c9f5feec63d66845ccc29358407a

Observation fc911385-58c2-4d29-8960-28a2efed4a5f · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.668979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.615354Z digest=sha256:0b4aac80e2ac4c45f96e2d6227854447ac27cb0bda5a399914b4d317fbb4da14

Observation 7957f7e0-13a0-49d0-bc50-6a40f7e65fa1 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.621128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.621128Z digest=sha256:03b3e0b9f8595d501db60adf7d972b70024f8667601301da6b1d2defb61fcb1b

Observation 0ba157aa-1050-4032-b4de-ad8a13b995d8 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.644069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.627439Z digest=sha256:183a254ebd71c08b127e74dea9798e0097a7d75406e01369b3ea6e9b936e33ff

Observation 5484146a-643a-449c-898a-9a95aadc7e39 · outbound

This paper cites DyKnow: Dynamically Verifying Time-Sensitive Factual Knowledge in LLMs.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models DyKnow: Dynamically Verifying Time-Sensitive Factual Knowledge in LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.633276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.633276Z digest=sha256:7eff74e606e74471538a2fbb201a4308d58a518b75c08f52ee590d69c7181d0d

Observation 39cb3076-c777-471f-b7a5-f00248c93f6f · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.620189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.640027Z digest=sha256:0a54bc89967ec6955968c5019ea52cf94354bc4e7ad6f77944654258862ecf1d

Observation 87af93bd-6f40-4b73-bba1-372d988aa146 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.587531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.646855Z digest=sha256:8a761530840e0b32622c81fd307f471a3b3f0f2a686587dbf9f150b68b0659cf

Observation 7e2fdcbe-a844-4989-97bc-092846acb64e · outbound

This paper cites Synchromesh: Reliable code generation from pre-trained language models.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Synchromesh: Reliable code generation from pre-trained language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.653670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.653670Z digest=sha256:4dd29bdda859dc22565e4761237e3170eb1256b5e5f405471f41a6be06d302dd

Observation 58397e66-d170-454b-926a-d1997761ea6a · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.560725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.659144Z digest=sha256:f300ecb9692889884a704f57e80c536fc2adee0ff1eac2ab0cba186b083aed8d

Observation e253b605-890b-4671-9ba9-84ef2227c51d · outbound

This paper cites Citekit: A Modular Toolkit for Large Language Model Citation Generation.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Citekit: A Modular Toolkit for Large Language Model Citation Generation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:16:25.974877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.665773Z digest=sha256:501bbcf912d57adc7d5a1b18d0816f416e71cf0ac54f45997a8f111cce31f0a1

Observation 6b93e68f-306c-4521-9ca8-fa02621dbc03 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.532493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.671292Z digest=sha256:dff9bead9ae18ed70fc64ebcffff2bb16b0062b23b4f4faacc605ff2efcd9f7b

Observation b53eb951-6531-4e0b-95ed-faa89f2851c2 · outbound

This paper cites A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.676939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.676939Z digest=sha256:d3a5b38e7c9f99cb22d9208752f034b8545c50b8eece70f75c46df155bd2d834

Observation 344f8f6f-a051-4bd0-83d1-85b5b879be5f · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.511687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.682144Z digest=sha256:b42906ffee538a17eb2f18f254a0d59ed6fcfee1792f9223031e29a1b407c34c

Observation 019bafe7-4c6c-4911-b4a1-5ae5d81c959a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.687214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.687214Z digest=sha256:8c42fb765011eb63e5805ec04a83cc4045314361abe6cac3438acf40a097f17c

Observation e402273f-a06f-44c4-bf4d-67f400b16324 · outbound

This paper cites Hallucination is Inevitable: An Innate Limitation of Large Language Models.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Hallucination is Inevitable: An Innate Limitation of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.692816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.692816Z digest=sha256:aa30d3dbd7f5ecd7d8e457b4b2b6ddcf63cb46b3eaec7cfd993a50c7ff345602

Observation 241b834d-f3d9-4120-84ee-13065d37bbd8 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.490873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.698440Z digest=sha256:b1bb30bc16dd0e10dcd2859ea783f4fea18435c06e7388d7ee004b2492f416b5

Observation c812ddfb-7440-4723-a8db-be9867f96b6b · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.704108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.704108Z digest=sha256:4480ee6bae88d139a0d6d1c71555cf2df8574d27693a37bafc8befdca72dbdf1

Observation 4fc75928-77d6-4d36-98e4-d9bab9f44392 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.469151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.710476Z digest=sha256:c2207be44703ceef2bd92c825f519fe39dffd2f9d92de19cf3e9abb5eec4023f

Observation 1bfce028-bca9-48bf-b799-e1a9dbccfb10 · outbound

This paper cites LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.717571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.717571Z digest=sha256:99ce5628a35c5f0765c38d67b5b0924325ef5865d216759d1b679fdd7e99bd10

Observation fe6e9f42-aa12-431a-8146-69411c6c6d09 · outbound

This paper cites Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.725604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.725604Z digest=sha256:f6276436693cee1118e4368fc6f5020af44e2310b3b665cda7245ec812050aa7

Observation 3c686bcb-82a8-49ce-9c05-5c9128ca258b · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.731564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.731564Z digest=sha256:8bb3d034c3f2e19efbab92c4c2a106af9ee5fe8594d293e73d7a7ce303cb6ad0

Observation 2f5fddca-8002-424b-8c90-b1eb12485805 · outbound

This paper cites online" 'onlinestring :=.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models online" 'onlinestring :=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.737702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.737702Z digest=sha256:6bc6c9ffd698f4a1f00c13acd9a5dec4b55d694d3d05d990fce4156ce7f7ee87

Observation 30148be3-d543-49a2-8f81-e7d975819348 · outbound

This paper cites write newline.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models write newline

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.745948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.745948Z digest=sha256:eaf973cab0621e801510e216b26fd678ce0b3b21d93da51c69b59340bef9d526

Pith citing papers

No inbound Pith citation observations are available.