Pith. sign in

Paper Citation Record · LEDGER

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training?

As of 7 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2604.18134.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.18134 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T05:50:40.056376Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact4
  • verified fuzzy37
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b7f0d21a-f965-4b5e-b9e3-83b9bf4d7c4c · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.arXiv.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.arXiv

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.186477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:59478eb0228900e2e064df701528cbb72dfa946cb02ce55567645662115d893b

Observation eb9d6db4-31ee-4116-87ee-1b18518ada9f · outbound

This paper cites EndoViT: pretraining vision transformers on a large collection of endoscopic images.International Journal of Computer Assisted Radiology and Surgery, 19(6): 1085–1091.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? EndoViT: pretraining vision transformers on a large collection of endoscopic images.International Journal of Computer Assisted Radiology and Surgery, 19(6): 1085–1091

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.183631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:f7a3364588b90bc4fd96a0248bd8da8b38e01b01bf968c8b01bbe057072a606d

Observation 0eb4b824-f695-4f34-be9c-f24ead5cf0cc · outbound

This paper cites Unsupervised learning of visual features by contrasting cluster assignments.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Unsupervised learning of visual features by contrasting cluster assignments

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.233533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:72ed776aae6a36a43ad900d8209d409731e018e36e95d542692a68d4be4923da

Observation afc4ee06-ada0-4ec5-82f8-5d83b7684252 · outbound

This paper cites Emerg- ing Properties in Self-Supervised Vision Transformers.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Emerg- ing Properties in Self-Supervised Vision Transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.209436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:9c3c56249cd476cfa270f9ae1f76058a5328563ae35a9075bce0cc53772ede8c

Observation cf18a84f-5590-4f52-ad08-2f2e55c73f8f · outbound

This paper cites Garcia-Peraza-Herrera.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Garcia-Peraza-Herrera

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.272388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:4cbd4171d77160dc3434a4f7ac45fa17a64d43a9bc60ee884709b4107a8361bd

Observation 98f648b4-9f88-438a-afb1-1d12c9059e19 · outbound

This paper cites Garcia-Peraza-Herrera.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Garcia-Peraza-Herrera

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.196142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:cb050c159cc1c51446c86520dd677f7d0c5cfb434134d5902b2ae54c9c05ad1e

Observation 1913f95a-34ef-472a-98ed-548408ff9818 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? A simple framework for contrastive learning of visual representations

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.265405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:d3972c3449851fc5d08a10c9027584873c4c2033f6271e7c383718f8e11b37dd

Observation cdea6b76-9097-40cc-9655-f6789d4baf5c · outbound

This paper cites Improved Baselines with Momentum Contrastive Learning.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Improved Baselines with Momentum Contrastive Learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.211512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:b2591764fe50f9c8a1173ed64eb02cad8157fe2780868807da4c2897abcdce12

Observation 7b8ec03c-6238-40ad-9d15-d0d170a67c93 · outbound

This paper cites Unsupervised Hyper- spectral Image Super-Resolution via Self-Supervised Modal- ity Decoupling.International Journal of Computer Vision.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Unsupervised Hyper- spectral Image Super-Resolution via Self-Supervised Modal- ity Decoupling.International Journal of Computer Vision

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.241496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:975eb8665935834e56d61c9d70f2a2c935ad65828d1e2ece57936b1d52c88104

Observation 10e846fa-174d-495c-9980-26a6a734eb98 · outbound

This paper cites Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multi- modal Reasoning.arXiv.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multi- modal Reasoning.arXiv

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.199041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:f9087676569c3ebc96f9c408663de52552cb25db6810d0ea02992b732f0c593c

Observation de9d3590-c5ef-445b-b9d0-d58eb1df448b · outbound

This paper cites Domain-Specific Language Model Pre- training for Biomedical Natural Language Processing.ACM Trans.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Domain-Specific Language Model Pre- training for Biomedical Natural Language Processing.ACM Trans

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.204622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:abc3a64a97713fd18e6be947aa677f27f0555d0a5a8dc1e1d76b4b2eda82e453

Observation ca5d9f0e-510f-411a-91a9-ab1a8c83b80e · outbound

This paper cites Masked Autoencoders Are Scalable Vision Learners.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Masked Autoencoders Are Scalable Vision Learners

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.256536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:f58cbc019977d0f8f886a628f2e8f833d5b8a25e0100090490d047c061c0820a

Observation 5f389e16-7219-4158-9638-4b73abb90d44 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-10T05:51:09.750117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:fc67da116f64e16ba11ccec95543f0d5869f89dc7bd32137fc958182e5aa883d

Observation 75433114-e759-4f05-8e9e-022c259bea0f · outbound

This paper cites Jaspers, Ronald L.P.D.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Jaspers, Ronald L.P.D

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.193850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:64e447ea5e72d7cc3aa4a41e2f154b217ac7ae759524e2cb8a348ff00c8aaa75

Observation 541c6cd0-6380-45a4-a418-d620049f2a6c · outbound

This paper cites Survey of hallucination in natural language generation.ACM Comput.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Survey of hallucination in natural language generation.ACM Comput

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.274783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:76ec35c89bab160a093d8f4f88a73ca81b64084fb5ddf44f9c7406ef6658f697

Observation 1709b0b8-b947-4de1-a619-c03d21ed31df · outbound

This paper cites Scaling Up Visual and Vision-Language Representa- tion Learning With Noisy Text Supervision.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Scaling Up Visual and Vision-Language Representa- tion Learning With Noisy Text Supervision

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.244190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:4e6ac3d01a414e068d5706d1e8885c872203ac5b70f9f84e677c72321bf8a09a

Observation 877cf544-f528-4deb-969a-afb16d7894fe · outbound

This paper cites Align before Fuse: Vision and Language Representation Learning with Momentum Distillation.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.201763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:53d5df34fef0d853f257fd29286ef5c130c4c1003f19d33398e1406a93c9622a

Observation e91a39ef-05a4-4d88-961c-4c261e2c9181 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Uni- fied Vision-Language Understanding and Generation.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? BLIP: Bootstrapping Language-Image Pre-training for Uni- fied Vision-Language Understanding and Generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.219075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:49c2757a06b00e51ba539b6dd64b2edde9e3f1cbb780cd2698f3e627fc95f241

Observation 1ef1d810-821d-4cae-9a1e-cffa32104835 · outbound

This paper cites Meireles, Guy Rosman, Maria S.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Meireles, Guy Rosman, Maria S

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.251834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:1d5144c74b1310890a1b01508a0234eaa566c9f5cf864e8e0b0447cb99a891af

Observation 7d5f0969-218d-4291-aee5-2e24498482fb · outbound

This paper cites A systematic review of annotation for surgical pro- cess model analysis in minimally invasive surgery based on video.Surgical Endoscopy, 37(6):4298–4314.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? A systematic review of annotation for surgical pro- cess model analysis in minimally invasive surgery based on video.Surgical Endoscopy, 37(6):4298–4314

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.228556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:dd552d524ab679b1673d034d113ae33e239cb6cfa22c6eed8d5d0223c7dc9e86

Observation b0186eec-df22-4b63-a242-42dfac215d3b · outbound

This paper cites Sur- gLaVi: Large-scale hierarchical dataset for surgical vision- language representation learning.Medical Image Analysis, 110:103982.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Sur- gLaVi: Large-scale hierarchical dataset for surgical vision- language representation learning.Medical Image Analysis, 110:103982

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.262050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:c72598886397f1e864d72c82e0c0e85682b1cfee139deb3ccb0fcad632445878

Observation 9db65e0d-a197-4d60-aaa2-32acba388688 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Learning Transferable Visual Models From Natural Language Supervision

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.249738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:f2217a3ed50dd331f0a08f60f240238cbf22c1ac27f5be55009913911eec1a2f

Observation 7d43969e-6a10-416c-87aa-771ded0e530d · outbound

This paper cites LAION-5B: an open large-scale dataset for training next generation image-text models.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? LAION-5B: an open large-scale dataset for training next generation image-text models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.247180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:f55e9f53e568e3427e967b5bfbdb042a41c4ebe94ce74088354601584f9a3826

Observation 603281e6-f751-49f3-a645-cea8b99297b8 · outbound

This paper cites Conceptual Captions: A Cleaned, Hypernymed, Im- age Alt-text Dataset For Automatic Image Captioning.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Conceptual Captions: A Cleaned, Hypernymed, Im- age Alt-text Dataset For Automatic Image Captioning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.276964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:9709f0d6c63f35a87bdaa690a0419468314c03ad7100ecc453629f0e50365646

Observation 319c6820-bd52-47ca-83bd-a1f29179d3b3 · outbound

This paper cites Transnet v2: An effective deep network architecture for fast shot transition detection.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Transnet v2: An effective deep network architecture for fast shot transition detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.230809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:e52a471f4d1294ec43590dbc8f059dff3958dc5c9d38b04d16689d935ad5a6a2

Observation 36fa8543-bfbc-4e94-99d1-fa2f4b472c73 · outbound

This paper cites Regionaligner: Bridging ego-exo views for object correspondence via unified text- visual learning.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Regionaligner: Bridging ego-exo views for object correspondence via unified text- visual learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.190558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:a7a34da05c58917401698fc3c07d263170500ef9ac26f0d2282a3ad918247e2d

Observation 4068f84f-b9fc-4e44-88fd-4cf6e9ced7f3 · outbound

This paper cites MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-10T05:51:09.755366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:a0d387803f191b27afca8a90877ff52ddd5bbf0d6bba1b8d5076963c4015d974

Observation 59998f53-caed-4f4c-b7be-3a9d1b319120 · outbound

This paper cites Parwani, and Muhammad Khalid Khan Niazi.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Parwani, and Muhammad Khalid Khan Niazi

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.216558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:3e9e3abe01d0d7c844fb6b1626a4dc4eb9cad885d20544536e99ab096b1f3d10

Observation 325258b3-530d-4aa5-9970-d94ae1cc2704 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Gemini: A Family of Highly Capable Multimodal Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-10T05:51:09.752692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:f98dc42f67a72c7dd0aa6d1e043f3679394a1fcc2f74ccba693078232d05655c

Observation 24c125e9-5c23-4de8-8e5e-81558f2b61e6 · outbound

This paper cites Twinanda, Sherif Shehata, Didier Mutter, Jacques Marescaux, Michel de Mathelin, and Nicolas Padoy.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Twinanda, Sherif Shehata, Didier Mutter, Jacques Marescaux, Michel de Mathelin, and Nicolas Padoy

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.226019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:11792d4a4bd1c937f3dd961e7e676132053a88c8983c1369e9ac8ae713ded7dc

Observation 829e8f0e-8067-4477-a914-25a1a60b455b · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Representation Learning with Contrastive Predictive Coding

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-10T05:51:09.758046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:58fb8522529bf4815f32ee96492dc38c6d298a6d02ccea62dac7e4c0fa66e54d

Observation 660ab2d5-d2f8-4f9e-ba7c-e6735183a1f4 · outbound

This paper cites VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.238699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:3e55d8f471ff59cffbb105cb165dd4d83bfa7d0f9a071e9d27d68517b330afe0

Observation eea168c9-3f95-48e0-97d6-6d33d8b2000a · outbound

This paper cites AutoLaparo: A New Dataset of Integrated Multi-tasks for Image-guided Sur- gical Automation in Laparoscopic Hysterectomy.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? AutoLaparo: A New Dataset of Integrated Multi-tasks for Image-guided Sur- gical Automation in Laparoscopic Hysterectomy

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.267802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:eca9297e486b32f8993ca58ac250c07afca3d0446a1d032ead4c4c99c62b897f

Observation 8939082c-6987-4233-9051-eb36b0592c0f · outbound

This paper cites Foun- dation Model for Endoscopy Video Analysis via Large-Scale Self-supervised Pre-train.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Foun- dation Model for Endoscopy Video Analysis via Large-Scale Self-supervised Pre-train

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.206854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:7281529d8f8b3471aac5b4034abe9d94d24407bdfea02773e0797036b1f47bf8

Observation 0825d2ff-0285-4f72-99d7-c4977d647622 · outbound

This paper cites Challenges in surgical video annotation.Computer Assisted Surgery, 26 (1):58–68.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Challenges in surgical video annotation.Computer Assisted Surgery, 26 (1):58–68

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.254297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:801e4884be6570a53381666e51c0a21619e21875420e0903591848285ff5bef1

Observation 1da81fec-5a34-4430-9f55-129142904ef9 · outbound

This paper cites SpatiaLQA: A Benchmark for Evaluating Spatial Logical Reasoning in Vision-Language Models.arXiv.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? SpatiaLQA: A Benchmark for Evaluating Spatial Logical Reasoning in Vision-Language Models.arXiv

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.221464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:a9e3a181545caceb5cab2ebe941dd4eb10ea688773004afaf5a67adeb24c6cab

Observation 540dac1e-036a-4dff-8805-34781650609b · outbound

This paper cites Lavanchy, Jacques Marescaux, Pietro Mascagni, Nassir Navab, and Nicolas Padoy.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Lavanchy, Jacques Marescaux, Pietro Mascagni, Nassir Navab, and Nicolas Padoy

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.223914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:e892c29f4c0bed5a5f60be3eec15cca3e1e645f2d17a498922ea9df8b2770f2b

Observation 4a04d59e-b710-4640-b6d5-c4f35d4e4ced · outbound

This paper cites Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring.arXiv.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring.arXiv

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.259586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:c01799089a51c891dfc3a279d951b20d35aea48e4f53982d44520f386f1f6949

Observation 56f9caa7-f108-49cd-929f-e279ab4fd2c3 · outbound

This paper cites Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning.arXiv.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning.arXiv

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.236073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:7dd35ec05ee7ea273178b287d72fbda4006d5a962f9f702eb4961877fd592ce8

Observation 5f773138-3c4f-4f06-b07b-8a3d22f8aa1b · outbound

This paper cites iBOT: Image BERT Pre- Training with Online Tokenizer.International Conference on Learning Representations (ICLR).

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? iBOT: Image BERT Pre- Training with Online Tokenizer.International Conference on Learning Representations (ICLR)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.270347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:5efabd07865d0d625387dd7d76db6932592e21ba6ca153db3bd5698103cc4bb5

Observation 496ddb58-0d99-4c9a-8033-29bdae276a34 · outbound

This paper cites Can we trust AI doctors? a survey of medical hallucination in large language and large vision- language models.

Can LLM-Generated Text Empower Surgical Vision-Language Pre-training? Can we trust AI doctors? a survey of medical hallucination in large language and large vision- language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T18:45:29.214178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:50:40.056376Z digest=sha256:5b01c211f086a9805d64ab762334c7e9a44b9234ea8169619fd2d46b98331d1d

Pith citing papers

No inbound Pith citation observations are available.