Pith. sign in

Paper Citation Record · LEDGER

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study

As of 7 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2506.06232.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06232 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:01:41.017116Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact3
  • verified fuzzy3
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ca26ed5-5aae-4719-8641-0f3311ba4dd7 · outbound

This paper cites Czempiel, M.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Czempiel, M

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:01:42.109452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:01:40.781723Z digest=sha256:b6c691bf0d041823d5ad447211ce93e2c0113a79a656877019a29de1557ada24

Observation 834c7b66-b21e-4042-8fb0-8244dd16a8ab · outbound

This paper cites Yamlahi, T.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Yamlahi, T

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:01:42.073782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:01:40.789737Z digest=sha256:1d82925f052814578bfb2869a6ba99c54fe60c78aa999ae81d1bc087c09cd276

Observation ec5623ee-2697-4d0a-9f03-2508a90841b1 · outbound

This paper cites Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.799960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.799960Z digest=sha256:ba3fc0a0c38337dd4b0ac7645c113b24c5a5d819e11350a1a5c35010f6c1b47b

Observation 995d6d0c-7fa4-416c-9f9a-960bbb276a1c · outbound

This paper cites Segment Anything.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Segment Anything

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.811423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.811423Z digest=sha256:542b63504d1996001718e31d031fcb8b30d516b2d64812ecbc88a8770a52f0f2

Observation a22fdc4b-c2a9-4fbb-86f2-ac1058f66381 · outbound

This paper cites Surgical SAM 2: Real-time Segment Anything in Surgical Video by Efficient Frame Pruning.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Surgical SAM 2: Real-time Segment Anything in Surgical Video by Efficient Frame Pruning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.821259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.821259Z digest=sha256:0dce323ba3e061f5d4f4f3eb89fffb119ca87117da94d1ffbf37a34ceaedf92e

Observation 7fb5124e-50fd-4c73-ae2d-bcdd1713e5a8 · outbound

This paper cites an unresolved cited work.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Unresolved cited work

Reference 6

Resolution
verified exact
raw_fallback, observed 2026-08-07T06:01:41.693085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:01:40.828320Z digest=sha256:72951417e364ddf77d23435d099b94a357dd4a9014e45f6d476a04e40a61e48f

Observation 3272a767-fbbb-4438-a217-a4cf59934f05 · outbound

This paper cites an unresolved cited work.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:01:42.039685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:01:40.841726Z digest=sha256:989d5802536df81759822f37f48bfda4352e1936387515750be32fcdfe833ccf

Observation fa5dfb38-def3-473e-8931-d4994e655c26 · outbound

This paper cites VidLPRO: A $\underline{Vid}$eo-$\underline{L}$anguage $\underline{P}$re-training Framework for $\underline{Ro}$botic and Laparoscopic Surgery.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study VidLPRO: A $\underline{Vid}$eo-$\underline{L}$anguage $\underline{P}$re-training Framework for $\underline{Ro}$botic and Laparoscopic Surgery

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.854985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.854985Z digest=sha256:ed07f7dea0bcdf09c634d60d99db999fdd1dc9f1c226f7ebae1acf92eb556324

Observation 73ac3da9-42d6-4537-9d6f-860b60e4b9ba · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Learning Transferable Visual Models From Natural Language Supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.874635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.874635Z digest=sha256:ecf6556d5e0c2182e7f7d61a0506d18967c47e273e7d725c05387db72d316013

Observation 73f50197-5c0a-467a-86b8-fcb825cc339b · outbound

This paper cites Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:01:41.495440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:01:40.884579Z digest=sha256:5786739dd511d9cdf62bcdad4501cdac6ccc5f943fe8a58cc68631a7f2bb58a8

Observation e7176091-eeb6-4474-827e-d896fbe6dec6 · outbound

This paper cites Depth Anything V2.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Depth Anything V2

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.893107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.893107Z digest=sha256:2c8d6da20ce2bd45d35b08e0d8ac3aa49c2db757e71b27ccb4d7df14c075dbe4

Observation 2d189f2b-a1f3-4e59-a6ec-335fb4838dc6 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.901934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.901934Z digest=sha256:2ad1a6ad4dbf34ff3330e019b392b6ba05dab105f4ede05c2840b1cbdd90fd07

Observation 07cc04ad-a7a3-43d3-b28e-d2d6df63f747 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.910317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.910317Z digest=sha256:e3bb1b3353c9bc35a078ffa08aac0f14837e9391a0f73d8072452bb1a49cd4ff

Observation 1450493b-8a05-469b-ad41-cb09613d2247 · outbound

This paper cites Maier-Hein, M.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Maier-Hein, M

Reference 14

Resolution
verified exact
doi, observed 2026-08-07T06:01:42.009368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:01:40.918117Z digest=sha256:2399728557e74fcfa9097b3cb402f20dccea0f2f5a9b5abb72bcd7e2c0ce79d0

Observation a3dc95d1-c380-4249-8786-63462564bf3a · outbound

This paper cites an unresolved cited work.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:01:41.970928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:01:40.935612Z digest=sha256:2e96c010767e5bf925a20472803f2ae3ae111eef9d943839e809ed717f6e6d63

Observation c1b49f04-b31e-4710-9c24-e25d882de839 · outbound

This paper cites an unresolved cited work.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:01:41.942958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:01:40.960150Z digest=sha256:2cbf106186c9eb36ae878337dd1dbe98b1719578b794e333827fd04b5dcaa024

Observation 319b16a7-4d70-4833-8b3b-0303ddd696d6 · outbound

This paper cites The Endoscapes Dataset for Surgical Scene Segmentation, Object Detection, and Critical View of Safety Assessment: Official Splits and Benchmark.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study The Endoscapes Dataset for Surgical Scene Segmentation, Object Detection, and Critical View of Safety Assessment: Official Splits and Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.968221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.968221Z digest=sha256:3c4639ec1bb937c42a7880a877a2bf5156e285415b013769206cd053465b21f9

Observation f93fd5fb-b60d-464e-a55f-8de8f0dbe950 · outbound

This paper cites Gautam, A.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Gautam, A

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.982318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.982318Z digest=sha256:9c8f9b97b436fa4c8f6da4500c17f0505c4bf96d530969b0a8fa13454b2fd9d3

Observation 31091282-8b37-4cdf-9ed2-542b709e2519 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.989310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.989310Z digest=sha256:26f803e69b22aa75f466d778722c84c086e3bdb09034a44902429acbffda6b84

Observation fd55f037-e5a8-499c-80cb-7ce337155cc7 · outbound

This paper cites Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:41.004736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:41.004736Z digest=sha256:d0c8e2aa5ab1b9f5b9ffb8b11b414d004703eebbfd13c6002afd8beee64a7393

Observation e262dcf8-ca8d-48c6-b394-1666109cff89 · outbound

This paper cites Godau, L.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study Godau, L

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:01:41.904021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T06:01:41.017116Z digest=sha256:42af9b80edd386269c7b8a4cfce22209344c80b5f1ef0fc8699f699ed65c5c9e

Observation ba7e69c1-7a21-4706-bf1d-7051023c9849 · outbound

This paper cites URL https://doi.org/10.1007/s11548-024-03141-y.

Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study URL https://doi.org/10.1007/s11548-024-03141-y

Reference 1417

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:40.951582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:40.951582Z digest=sha256:c84721a87e507c2b837fdfe04de953dd647c5197c5b1f702f8b371850c25c21e

Pith citing papers

No inbound Pith citation observations are available.