Pith. sign in

Paper Citation Record · LEDGER

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis

As of 11 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2501.09555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09555 v3

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:58:30.055411Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:21:40.759357Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T20:21:44.721684Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 43ea087c-64f4-4113-ad3d-894f4d3fe015 · outbound

This paper cites an unresolved cited work.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:58:31.353065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:29.586616Z digest=sha256:faac16f3fc7549b76ed3bd39d0cad12777e9c7edbc62ebca659ed3432bf00449

Observation b8aaaa07-3a72-45e2-bf3c-230500716579 · outbound

This paper cites an unresolved cited work.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:58:31.199339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:29.593947Z digest=sha256:a92cd8e11f0dd2946d34f2eaa3e3eca5266a696eb7974acd6715256d6386c98a

Observation 6fa50221-dd08-49a9-bb00-1583b5ef32ed · outbound

This paper cites an unresolved cited work.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:58:31.140077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:29.601216Z digest=sha256:fa4fa274ef424732c937662f4a222228ef1000fdc764ea03b875db264b452599

Observation a853ea1a-95ba-4cd3-bd33-906fad56707c · outbound

This paper cites I can’t believe there’s no im- ages! learning visual tasks using only language supervision, in: Pro- ceedings of the IEEE /CVF International Conference on Computer Vi- sion, pp.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis I can’t believe there’s no im- ages! learning visual tasks using only language supervision, in: Pro- ceedings of the IEEE /CVF International Conference on Computer Vi- sion, pp

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:31.126610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:29.608083Z digest=sha256:0bb7f020545ac1c9a54d73f5e86c1960f7ea28f39313bee446602d58b1803b5c

Observation 418a45d2-4161-48c8-8203-beedaefa8e8c · outbound

This paper cites A k-means clustering algorithm.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis A k-means clustering algorithm

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:31.082359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:29.615929Z digest=sha256:62fe5687e4ad224ee5d6a3081251efdd8b1e97b79a3a246df8d4e1bf4ca64a8e

Observation a0b99eda-ff0c-46c0-a332-8f39729378e6 · outbound

This paper cites an unresolved cited work.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:58:31.010264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:29.621377Z digest=sha256:e13fa52f1e1ed67c58de10a3b1f07fe01b5bca0bccc3bc2691c818c1c57530d1

Observation 5c524ee1-4658-4400-a9e4-d589fd2efd31 · outbound

This paper cites Latent graph representations for critical view of safety assessment.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Latent graph representations for critical view of safety assessment

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:30.846622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:29.629723Z digest=sha256:7fd84cacf41638a3694a6d77f61f7d16b1b67f4bc87af16eb4623fe07d97e039

Observation 2e5fd39c-03ab-4f00-9e12-a4e80447c76a · outbound

This paper cites Text-only training for image captioning using noise-injected CLIP, in: Goldberg, Y ., Kozareva, Z., Zhang, Y.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Text-only training for image captioning using noise-injected CLIP, in: Goldberg, Y ., Kozareva, Z., Zhang, Y

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T19:58:29.634023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:58:29.634023Z digest=sha256:5471db46afdabcf6e0ac34f6970cfcb9ea3f9a9c24e56033127b9f816c3b80f8

Observation 6d73ad59-8501-4043-b93b-da954d441220 · outbound

This paper cites Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:30.785614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:29.638762Z digest=sha256:8acb044d2ba88b9155a6ab6b37af958ca8aff1f629e0fbed79b427691492a360

Observation 1965676f-65cc-4e1d-861d-8a14eeab5c6e · outbound

This paper cites Machine and deep learning for workflow recognition during surgery.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Machine and deep learning for workflow recognition during surgery

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:30.771197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:29.687621Z digest=sha256:6e861c8674609b7b800c6e77bac435f3f80b9a9dd0b477f4818dd1a608cf05bc

Observation 392f555d-026f-4d63-ab46-717e82c9a4de · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation, in: Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Bleu: a method for automatic evaluation of machine translation, in: Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:30.729598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:29.744728Z digest=sha256:4424a182f4e2278e9499f8b09eb2230b8a806e575b9fe533b73745efe2134da4

Observation e2cf401d-bf9f-4012-8792-6b0356b881e3 · outbound

This paper cites PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T19:58:29.793971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:58:29.793971Z digest=sha256:459747107668797f2b5387afc37036c0c2d81d1b6794dd8824b1bc935241123b

Observation b32e7f4e-248b-41db-9aaf-452cdd0926e2 · outbound

This paper cites Learning transferable visual models from natural language supervision, in: Inter- national conference on machine learning, PMLR.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Learning transferable visual models from natural language supervision, in: Inter- national conference on machine learning, PMLR

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:30.674429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:29.878254Z digest=sha256:9711ca6a2e2855234e1baca3a66ae289b783d3ef3339c62fc07faf304b91e86b

Observation 74c254b5-b088-4852-b8c6-6c829984cf24 · outbound

This paper cites Language models are unsupervised multitask learners.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Language models are unsupervised multitask learners

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:30.660063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:29.894306Z digest=sha256:61b6e4b5a77728b2df06d3fb51e837539fed4cc6bed2b1a1fc21d799629b0c7a

Observation 4c83c0dd-4217-4d19-a01f-f2ba4154141d · outbound

This paper cites Ima- genet large scale visual recognition challenge.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Ima- genet large scale visual recognition challenge

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:30.578523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:29.964339Z digest=sha256:3321b1117b59fdb858c9916a424abed62386a018826e02fd67fe0622b6f7345e

Observation 02577e59-e60e-4716-a404-478873dcb6e5 · outbound

This paper cites Surgical- vqa: Visual question answering in surgical scenes using transformer, in: MICCAI, pp.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Surgical- vqa: Visual question answering in surgical scenes using transformer, in: MICCAI, pp

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:30.507813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:29.974247Z digest=sha256:351a863ca3107dca77a01ecf412de8f97821c98f46795024164226cdf0e0f7bc

Observation 354a920a-9b7b-4a31-8ada-460a1a272f25 · outbound

This paper cites Endonet: a deep architecture for recognition tasks on laparoscopic videos.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Endonet: a deep architecture for recognition tasks on laparoscopic videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:30.493209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:29.983824Z digest=sha256:f102609bacbbf0e9f193599e4cc30f4166be8975efa52913fb697470cdf530ef

Observation ccb0c833-738d-4f41-be2f-6b9ce11161c8 · outbound

This paper cites Cider: Consensus-based image description evaluation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Cider: Consensus-based image description evaluation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:30.477714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:30.002057Z digest=sha256:693ac5e5ad5703e8d0f2d484e5369f55a87e99f056add7368bf8d409c456bc18

Observation 63691d2c-6679-47e0-9330-9d243bc9096f · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis ActionCLIP: A New Paradigm for Video Action Recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T19:58:30.021304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:58:30.021304Z digest=sha256:9c2a4bf60b612015a96f2d3d091ffd76b336835968f3f3a5470984317ec47997

Observation 17fb5981-f75b-4346-a199-e060373369e2 · outbound

This paper cites Visual-language prompt tuning with knowledge-guided context optimization, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Visual-language prompt tuning with knowledge-guided context optimization, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:30.444666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:30.029323Z digest=sha256:c0782f852b3f26f0b102d22893faa0a441ac3d0ac4f840e83ef91c9578c4c5b6

Observation 988509e9-faa8-45f5-82e9-50d5c03521be · outbound

This paper cites Anticipation for sur- gical workflow through instrument interaction and recognized signals.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Anticipation for sur- gical workflow through instrument interaction and recognized signals

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:30.377144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:30.033568Z digest=sha256:fa0d87f7bdfd96abeb7c35e8d37f690ae3561de2174edc547fd7e631d0518404

Observation 7f9cb348-5232-45bd-b0a1-1272a58a5e48 · outbound

This paper cites Advancing surgical vqa with scene graph knowledge.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Advancing surgical vqa with scene graph knowledge

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:30.302547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:30.037780Z digest=sha256:410d7a81f9c70b0d52c16c4712f7af7c3f38235be111073a8c7a6faf2c20748b

Observation 70e91694-db7f-421d-960c-bd6c8c9f1a1b · outbound

This paper cites an unresolved cited work.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:58:30.229130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:30.042669Z digest=sha256:d95e885c69ff168ddec92e37fdd5cc2a0ae77d12c64434de617d77a19d369355

Observation e5781e70-2e2d-45da-8ec0-44300e107a4b · outbound

This paper cites Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T19:58:30.046785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:58:30.046785Z digest=sha256:7003913ed7897b7c398d12d0756df35991472a12587bb64c620f7a30fee7e854

Observation 4e77adc0-8dd8-4725-a0f8-b9346ea271ca · outbound

This paper cites Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T19:58:30.051097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:58:30.051097Z digest=sha256:ea103126dcd706c91c2d809d0052c16610257588e96a9c3cd5a6fbac8cbd1db4

Observation 94e025d8-1456-4453-8894-67f4ca9e9abc · outbound

This paper cites Conditional prompt learn- ing for vision-language models, in: Proceedings of the IEEE /CVF con- ference on computer vision and pattern recognition, pp.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis Conditional prompt learn- ing for vision-language models, in: Proceedings of the IEEE /CVF con- ference on computer vision and pattern recognition, pp

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:30.164155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:30.055411Z digest=sha256:06d159ebc7516fad88b1a86756cd12e42f491a40167dc7b3f4b2552973768a2a

Observation 53533d37-a7f8-4f09-b947-26b9992524d3 · outbound

This paper cites MedIA 76, 102306.

Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis MedIA 76, 102306

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:58:30.945144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:58:29.625621Z digest=sha256:3016432484384218ea2781ca85468e5397f5d76e1f614e0739627fd9e3737174

Pith citing papers

Observation 26345b51-1f17-4046-8379-ddeed7f4d3a3 · inbound

CPKD: Clinical Prior Knowledge-Constrained Diffusion Models for Surgical Phase Recognition in Endoscopic Submucosal Dissection cites this paper.

CPKD: Clinical Prior Knowledge-Constrained Diffusion Models for Surgical Phase Recognition in Endoscopic Submucosal Dissection Text-driven Adaptation of Foundation Models for Few-shot Surgical Workflow Analysis

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:21:44.820536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T20:21:40.759357Z digest=sha256:f1f972a596fd34473e76a50edc3011ec713d8b01cdc0be76180a486f15eccc5c