Pith. sign in

Paper Citation Record · LEDGER

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI

As of 17 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2504.19918.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19918 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:43:43.323375Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:15:40.944411Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T03:16:18.683507Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact1
  • verified fuzzy43
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65b91c13-a6f5-46ea-aba4-1869d9f6ba08 · outbound

This paper cites Gaze-assisted automatic captioning of fetal ultrasound videos using three-way multi-modal deep neural networks.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Gaze-assisted automatic captioning of fetal ultrasound videos using three-way multi-modal deep neural networks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.444080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:41.901887Z digest=sha256:cc0874afbf81aac0fdf58deea54894f5dfb4d341ac03669e7d55eadcdc76bdf7

Observation 280b488c-fd8d-487b-9ffa-85a235bc249e · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Bottom-up and top-down attention for image captioning and visual question answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.045511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.045511Z digest=sha256:7aa7c8b5b296caecb53d56fcd3587cc08b4841fde6ab1b6fcdcb330b8d00022c

Observation 022d5a76-e071-460c-856a-0ba400f6845e · outbound

This paper cites Croma: Cross-modal attention for visual question answering in robotic surgery.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Croma: Cross-modal attention for visual question answering in robotic surgery

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.344698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.094303Z digest=sha256:7b8d2ec0306595a153b018da294a0ce8ed63641ef0b97c52744e27c82b22b0f3

Observation 51231b20-71de-4633-97ae-83c22eb59cdd · outbound

This paper cites Vivit: A video vision transformer.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Vivit: A video vision transformer

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.333637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.098989Z digest=sha256:acc6bb5ffd3af70da84f302c32fdc0377d3d7a152e7354299a8c00e9be278a1c

Observation 039eefb5-dc0a-4d47-8dd1-0f44e18d48ad · outbound

This paper cites Video-based coaching in surgical education: a systematic review and meta-analysis.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Video-based coaching in surgical education: a systematic review and meta-analysis

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.321896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.103523Z digest=sha256:8682c1ad867066500dd7510048219559e4d8edc8b5eb625a9b69e27fe7291798

Observation 603f2386-1211-4df2-b3dd-0cd525f62106 · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.108010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.108010Z digest=sha256:1bf9dc7293c8b80a0019ba913f634e80b85cda8f99ea257911091e4145f47c71

Observation 6914d4e4-bb35-4512-aaaa-0345ebd44a15 · outbound

This paper cites Space-time attention networks for video understanding.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Space-time attention networks for video understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.310249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.112743Z digest=sha256:e819230f5be66335ba5b57e48fa442fbb69d8ad60449c4b68ea3c74a6951fb63

Observation c6600e4a-58ef-4a81-9729-6d828ffecccb · outbound

This paper cites Language models are few-shot learners.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Language models are few-shot learners

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.300076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.117842Z digest=sha256:1614b72a8d3b144f2b0b60fe23f3a90136f75dc2398660aa7b2d11eae10048c0

Observation 87b33889-2b43-46fe-be96-4dbfdc430a75 · outbound

This paper cites Cholect50: A dataset for surgical video understanding, 2020.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Cholect50: A dataset for surgical video understanding, 2020

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.288051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.184774Z digest=sha256:ec44caf7caa052551c5909104104e7f3eaab2c688d6d02e5d05ff3d450a95b45

Observation a32d99c5-1b30-4ab5-b1a5-ff6614355c87 · outbound

This paper cites Unmasking deception: a topic-oriented multimodal approach to uncover false information on social media.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Unmasking deception: a topic-oriented multimodal approach to uncover false information on social media

Reference 10

Resolution
verified exact
doi, observed 2026-08-16T05:43:43.435845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.281716Z digest=sha256:9025574d3e1e46ba87a8935f51ed885e4ff0ec92559adf4144a42ff133e895fc

Observation 06637a61-a7e7-4ab1-b454-4217e529dcd1 · outbound

This paper cites Harnessing prompt- based large language models for disaster monitoring and automated reporting from social media feedback.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Harnessing prompt- based large language models for disaster monitoring and automated reporting from social media feedback

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.145499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.285629Z digest=sha256:20c6314c8f069ee270c24ebfa6c891fc3e632464914cc67afa3143ebf4d5f02d

Observation 39b92217-b8ee-4594-b58d-ed878c367b21 · outbound

This paper cites Surgical video captioning with mutual-modal concept alignment.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Surgical video captioning with mutual-modal concept alignment

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.079828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.290474Z digest=sha256:4522552e7209fb7cd7f74b5907630a1574b839c6339c658ab11928989cd1d58d

Observation ded858a9-602a-4671-877b-82ee7d967e02 · outbound

This paper cites Vision language models in medicine.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Vision language models in medicine

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.068097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.294015Z digest=sha256:c29f2028330fc72cf1f926356f5ca61475775991e523c398be95dbb5d495551b

Observation 93f429a1-a0c4-44f9-bfc0-ba361da86f91 · outbound

This paper cites Meshed-memory transformer for image captioning.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Meshed-memory transformer for image captioning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.056339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.297352Z digest=sha256:542f175f3d0ce9869a09f44a53761ac694be80cbb15fa3723635896e7f94f32e

Observation d8d9dc8d-ba59-44dc-b943-e148c86a717b · outbound

This paper cites Attribution-noncommercial-sharealike 4.0 international (cc by-nc-sa 4.0).

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Attribution-noncommercial-sharealike 4.0 international (cc by-nc-sa 4.0)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.044287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.301598Z digest=sha256:44f4f9b1ce591f5d3449b45a946233efea8f3fe5d10fef364130ca08cc3c58c9

Observation 56de2812-4295-4153-9b2a-c6e09463dc48 · outbound

This paper cites Class-balanced loss based on effective number of samples.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Class-balanced loss based on effective number of samples

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.306279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.306279Z digest=sha256:e4f9a4b82b1c30741e7341ad7e48e995d4331569cf08e6bd47d914c7da503d08

Observation 194f44f0-0ef8-430b-aea0-ab09b47168ac · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.310201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.310201Z digest=sha256:bcab0b477a4972a6adf59d7f55b9781394cf137696c577f4f75e8820281a136a

Observation 621034ec-7d7a-4a42-9e82-aab2b673014d · outbound

This paper cites Surgical video analysis: an emerging tool for improving surgeon performance, 2015.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Surgical video analysis: an emerging tool for improving surgeon performance, 2015

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.025018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.321600Z digest=sha256:b24b6e0c8991e1e1ddae829b07c7142dae64d42bc4c76edab591c8b3477029af

Observation f30f51f4-3509-40ea-85c8-9be32a31300a · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.420366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.420366Z digest=sha256:a374fcaf6774eed93d179c1be19f51fcbf6c02acb7ecab2d150681dc2dee266f

Observation 33f7feff-207a-4e92-9c5b-053e6c498d97 · outbound

This paper cites VideoOrion: Tokenizing Object Dynamics in Videos.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI VideoOrion: Tokenizing Object Dynamics in Videos

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.424641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.424641Z digest=sha256:1b1fcbd0db569221db262e628ae514411ee532112af9bd20df47bfbe10b175b7

Observation 31eb75b2-ec8e-4279-b99c-6a5079cf7248 · outbound

This paper cites Image-text surgery: Efficient concept learning in image captioning by generating pseudopairs.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Image-text surgery: Efficient concept learning in image captioning by generating pseudopairs

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:45.012197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.428992Z digest=sha256:1233241acae79c96ad734a400c2580c77a654500f9992a9f8432e42485b6da76

Observation bc4e316c-c4f0-4553-95c9-6d6039c5bad7 · outbound

This paper cites Using surgical video to improve technique and skill.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Using surgical video to improve technique and skill

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.999052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.432844Z digest=sha256:413c153ff62cbcb0e204914dc1da45ab03700f65e4d827cee72ef620980bb9fd

Observation 76f39dcb-1c49-4e43-83c0-b8b096c61cf7 · outbound

This paper cites Weinberger.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Weinberger

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.987453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.436742Z digest=sha256:479b42ff66519c479f01ac89ed0a7ced977ee6b092253125afd950f4a7bf6108

Observation affd8197-634a-4798-bb4c-b1d19b77990e · outbound

This paper cites Hashimoto, Guy Rosman, Daniela Rus, and Ozanan R.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Hashimoto, Guy Rosman, Daniela Rus, and Ozanan R

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.440064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.440064Z digest=sha256:a149816ab3d06e4e5b0e85225699a79ee7b00787147b4b9e4506397583250074

Observation 24aa3131-0181-4d5f-a397-da3ab2eab46e · outbound

This paper cites Deep residual learning for image recognition.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Deep residual learning for image recognition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.443698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.443698Z digest=sha256:4710a2ffd628f7f132e1d70ab5bbb8d898358f14f05825dfeba9c8e58d9c8e9c

Observation 18f0a98e-45a9-43c5-8002-ec7d69782cfb · outbound

This paper cites What do we need to build explainable AI systems for the medical domain?.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI What do we need to build explainable AI systems for the medical domain?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.447190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.447190Z digest=sha256:30cf38c4c91007e618777d3662bb1a97a264274d03471100eccb10c930016b42

Observation f900bf73-ce3b-4439-a8b4-329ac5725203 · outbound

This paper cites Advancing medical imaging with language models: featuring a spotlight on chatgpt.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Advancing medical imaging with language models: featuring a spotlight on chatgpt

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.480347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.480347Z digest=sha256:5ecc551873cb24e94cafc8193d90a83573a9299d0117fd2c5fdfc58f551a5d1e

Observation 31430eef-cd38-4ddb-84a7-3b5354f949b2 · outbound

This paper cites Exploring video captioning techniques: A comprehensive survey on deep learning methods.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Exploring video captioning techniques: A comprehensive survey on deep learning methods

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.794772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.538829Z digest=sha256:cc97cdcc4c527f8e7159efcccfb0ecbf93a0b91384047c772dad0fe7b0fedea8

Observation 1ab8b5de-f54d-4030-8a4f-742972c7a017 · outbound

This paper cites Multi-task recurrent convolutional network with correlation loss for surgical video analysis.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Multi-task recurrent convolutional network with correlation loss for surgical video analysis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.569371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.569371Z digest=sha256:c1ab09be64ab9eaf0589eccbdacd1ce15ba72967a99f93288033a66b5936c0d6

Observation edeb8745-7979-49a4-9857-c02cfe1e1b2f · outbound

This paper cites an unresolved cited work.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:43:44.740682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.572825Z digest=sha256:840bc020e84477a4086347db1da632aeee61550016052f253b5dd7e4177d7c42

Observation f61e7426-f910-40f8-8c08-9a521495ec05 · outbound

This paper cites Predicting decompression surgery by applying multimodal deep learning to patients’ structured and unstructured health data.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Predicting decompression surgery by applying multimodal deep learning to patients’ structured and unstructured health data

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.728676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.576369Z digest=sha256:04ab1e9cb74d5ff994b6e31f1ce4b35ec7f88295bb7da16a121cafcca87d6ebe

Observation 875b45cd-286d-42fb-94fb-211e21c07b17 · outbound

This paper cites Stvs: Spatio-temporal feature fusion for video summarization.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Stvs: Spatio-temporal feature fusion for video summarization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.717527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.580245Z digest=sha256:f2a715a21ab8217a63d93ff134f0547289ec52973071e8c670a4f6a6334b6a09

Observation 56c5463d-1a87-45e6-92ec-1370e65ebf37 · outbound

This paper cites Intraoperative video analysis and machine learning models will change the future of surgical training.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Intraoperative video analysis and machine learning models will change the future of surgical training

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.705961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.584521Z digest=sha256:93fdd1bfe70fe83c53108a36aeb25dee9e938efec639b5e94788ce54ad0f3a75

Observation 854b319b-648e-4b96-afa4-23398b71d643 · outbound

This paper cites Video question-answering techniques, benchmark datasets and evaluation metrics leveraging video captioning: A comprehensive survey.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Video question-answering techniques, benchmark datasets and evaluation metrics leveraging video captioning: A comprehensive survey

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.694043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.589002Z digest=sha256:ddb296f5ae6be20c290fd285e2cc5e138e80ee704b65f9fcf6cf7bf4ef56cd5a

Observation dd7a2d58-84ae-47f6-81b5-4bbd8eb01835 · outbound

This paper cites Generating automatic surgical captions using a contrastive language-image pre-training model for nephrectomy surgery images.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Generating automatic surgical captions using a contrastive language-image pre-training model for nephrectomy surgery images

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.680800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.593164Z digest=sha256:107796a4ea79e54bab93297488bc27fe9920454725c3250e710edda43378d9b7

Observation d1d62cd2-c2e6-484d-80b5-e3a8ae531713 · outbound

This paper cites Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.580260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.596149Z digest=sha256:9135bf6e90e0ef19de3146eaa2cd3348655b2acb9b6746a4a77e8de433421557

Observation c909f190-b858-4177-bbd7-26a6fe8458c0 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.478307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.599127Z digest=sha256:777669bf9ed1f5a38d7ec36a5700f098915c952f685ec8501352ddb843664f1c

Observation 89765e17-031a-42ed-8228-629600724cc1 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Rouge: A package for automatic evaluation of summaries

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.658846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.658846Z digest=sha256:3f51e9f6fe55e1fee4707d870aa8aee4dc2581cf47d825c8059492915f789a6b

Observation a2fe0918-af0a-4acb-a986-d5de792d3d04 · outbound

This paper cites Focal loss for dense object detection.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Focal loss for dense object detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.723455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.723455Z digest=sha256:4b6c0608fc0b9570dbf9fe7970327b5cda1442849aa86d92e4ad3ce32ed4ec2c

Observation c337ad92-c209-44d3-b340-6dc6b35f3806 · outbound

This paper cites A survey on deep learning in medical image analysis.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI A survey on deep learning in medical image analysis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.781863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.781863Z digest=sha256:f02af3cf7e2579ea03a046c4228e90419ec778e591f163400027b905f197bf11

Observation 43827adc-7873-4fc2-9db6-e29b2b319165 · outbound

This paper cites Video content analysis of surgical procedures.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Video content analysis of surgical procedures

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.447588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.785565Z digest=sha256:8c63c0de05b7205511ecc96eb33ed838ecb9680259daffe56c1f955e608e1cd6

Observation 2f6bb367-4920-4ce2-9860-cea82207a6a0 · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.788616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.788616Z digest=sha256:1752aea140ab8dee9ebeb7bb1b238ab9895b158342d88ccf2fbea469b0e78974

Observation e0c75f67-a4c2-4e39-8646-426c82a5e4cd · outbound

This paper cites BioGPT: generative pre-trained transformer for biomedical text generation and mining.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI BioGPT: generative pre-trained transformer for biomedical text generation and mining

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.793034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.793034Z digest=sha256:ac1aa026404fb46e9b7783124a74c583ff8347a12ce5fd77fcab0e747640f581

Observation 2ccc63c2-d92e-4809-bc7b-05c429c369a7 · outbound

This paper cites Surgical data science for next- generation interventions.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Surgical data science for next- generation interventions

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.435637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.796684Z digest=sha256:d9e6fdfcc00765f3aea50bc7375340f849032de75a15c221b3aef74c3a4bdf28

Observation 60222a66-fa30-4a7b-a46b-4040224b16cd · outbound

This paper cites Is online video-based education an effective method to teach basic surgical skills to students and surgical trainees? a systematic review and meta-analysis.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Is online video-based education an effective method to teach basic surgical skills to students and surgical trainees? a systematic review and meta-analysis

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.424288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.800372Z digest=sha256:e627a05700faa51124c6c646b6bc15417751b6e8c3c2dbcf01c9a400750ee14c

Observation ebb80fc1-898e-4d06-9d57-2c6d1f0d20db · outbound

This paper cites Gpt-4 technical report, 2023.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Gpt-4 technical report, 2023

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.804020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.804020Z digest=sha256:c3eb881e598b60290433a1505c46a84df42de7587802288e9aabf939c982b686

Observation b4b81d0e-c758-4e95-a55d-13f7aa103878 · outbound

This paper cites Training language models to follow instructions with human feedback.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Training language models to follow instructions with human feedback

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.807245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.807245Z digest=sha256:3fdda379c6c0951d8fb0809ce2e1c844b7336897fbab48e4432ebf281459c96f

Observation 5192ce61-c9ab-4f48-aa7c-97b85abee2c4 · outbound

This paper cites Bleu: A method for automatic evaluation of machine translation.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Bleu: A method for automatic evaluation of machine translation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.281583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.811499Z digest=sha256:3f0bdc2f8c0704aebc7f6b2bf8c6f3d7355719ec11a9001d7f9935b9c1dfaca4

Observation 6f62def6-996b-4313-afd1-14255631a386 · outbound

This paper cites Dense video captioning: A survey of techniques, datasets and evaluation protocols.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Dense video captioning: A survey of techniques, datasets and evaluation protocols

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.244441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.814496Z digest=sha256:bf5ad8c4034b0605b2f0f60ae0eb7121ae19fb70ef9a080508cbc0fa5effb832

Observation e9597801-e764-448d-8094-729d3e9a4ed6 · outbound

This paper cites Language models are unsupervised multitask learners.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Language models are unsupervised multitask learners

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.818065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.818065Z digest=sha256:8c0d5faf578ff19c0b35693f8d8183da87e0cea96ae54dd789eaafe9d13c4fea

Observation a2d14e7b-3278-431b-b548-cfa2e625cae2 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.223163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.849230Z digest=sha256:51511a024544474ea68a115a64db956f2b513e6e7a57ab79967815d8f0dbb868

Observation a8db6b68-3d6a-441d-b31f-1f5e81d6c2ea · outbound

This paper cites Dl4burn: Burn surgical candidacy prediction using multimodal deep learning.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Dl4burn: Burn surgical candidacy prediction using multimodal deep learning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.209756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:42.948706Z digest=sha256:d62eedcd71e874951334f4d38a1bff14944c8d335737d0bfafda6c8e5efef48f

Observation 72a4f4ca-b025-429d-ae1f-59c2dc181554 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:43.025376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:43.025376Z digest=sha256:84c63b3d8da1e9ded2c8fa3ce0a462c016e81616a22c257a3cfe55e17c325916

Observation 6b1be94f-f1bf-43dc-9574-4cdb699098fb · outbound

This paper cites Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.168239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.101180Z digest=sha256:2251a81157179d4228a6ce893830180402bfe22b9b4cd2576d6cfb8426d1c98f

Observation 00ee135e-a59f-423f-b686-cb6bf8cc72fe · outbound

This paper cites Evolution of visual data captioning methods, datasets, and evaluation metrics: A comprehensive survey.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Evolution of visual data captioning methods, datasets, and evaluation metrics: A comprehensive survey

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.091053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.105167Z digest=sha256:ee6e6e29c0c699b7c44553eeb755a6ed2e773dafb561fd4e68678e0fca44fd34

Observation 080d5a68-750d-4eed-8abb-1f2fa956df12 · outbound

This paper cites Towards Expert-Level Medical Question Answering with Large Language Models.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Towards Expert-Level Medical Question Answering with Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:43.109468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:43.109468Z digest=sha256:35c11b7c74c3017115e1f7d2ec4c1d66dae9118ac244a4e5acd209160827351d

Observation b93d691c-fb74-4455-9ca1-f88d7ac64bf8 · outbound

This paper cites Automated radiology report generation: A review of recent advances.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Automated radiology report generation: A review of recent advances

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.077498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.113530Z digest=sha256:848cf16055567043215002943f6ffd1e2e8c41f2d39959b724ccbd78bbe1ca0d

Observation b121b26e-e170-4c4e-8181-0ab025da6cfc · outbound

This paper cites The role of large language models in medical image processing: a narrative review.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI The role of large language models in medical image processing: a narrative review

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.064101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.116866Z digest=sha256:8309b85fb106700eb7c6c97d332dfdae25e80e2114b83eea0c7b2271afa1b18a

Observation 5a201b1c-71ab-45af-be40-162d8b6f8ebe · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Training data-efficient image transformers & distillation through attention

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:44.013534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.121206Z digest=sha256:3f93915e04b4a61532b11a3780b697109512987a6559461ce6629d952dcf9857

Observation dd672e10-d9ce-4a40-840c-77e56c8c2c97 · outbound

This paper cites On large visual language models for medical imaging analysis: An empirical study.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI On large visual language models for medical imaging analysis: An empirical study

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:43.125685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:43.125685Z digest=sha256:59a73ca6ada6f2572de6423ab44aa79ce15f227a99fe94a67e3da7fed5a46927

Observation 6d3aafef-a6f4-489c-86ec-f0c9fa1398d6 · outbound

This paper cites Attention is all you need.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Attention is all you need

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:43.954375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.129378Z digest=sha256:399058aecf67cdb26c355f42ca11a7439b04aef8403a9d3204a98bbca7d091f3

Observation 35513820-fcbf-4ea4-a480-08bee0340569 · outbound

This paper cites A novel multimodal deep learning model for preoperative prediction of microvascular invasion and outcome in hepatocellular carcinoma.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI A novel multimodal deep learning model for preoperative prediction of microvascular invasion and outcome in hepatocellular carcinoma

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:43.939261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.133062Z digest=sha256:f6b50a2856e9f2f45ff1962795658b397a19a5accf54c93f94c4624c55211bdb

Observation 86232122-99d9-4ead-bb87-b5f9a2c435f1 · outbound

This paper cites EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:43.205100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:43.205100Z digest=sha256:dd1d77f03b33d0cd177a8fea21fcb04cae3c030a28c91870d23d08c8b660a01b

Observation 0757bc3e-862d-41ab-a634-bdf0b97c9697 · outbound

This paper cites ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:43.297751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:43.297751Z digest=sha256:d3c427847d82c974949e679d60bc7698363cbc0a7cf2390e4ed622c603c54b25

Observation 3794ff53-9cc3-4163-814a-2a6337dfae74 · outbound

This paper cites Learning domain adaptation with model calibration for surgical report generation in robotic surgery.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Learning domain adaptation with model calibration for surgical report generation in robotic surgery

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:43.905987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.301562Z digest=sha256:c047517b108263ffb50081aefd0b9b63828eda8a12c888f1aef2759c04cea66c

Observation 2967e3a6-3463-43d7-a1e2-9fabf322452f · outbound

This paper cites Benchmarking Large Language Models for News Summarization.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Benchmarking Large Language Models for News Summarization

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:43.305442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:43.305442Z digest=sha256:c0f50d2fd6754bcf4c18540b331699d02b304247416769ebb1e4144cc96f5511

Observation c767ffd8-2e02-4c35-abda-429d62379f7f · outbound

This paper cites Weinberger, and Yoav Artzi.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Weinberger, and Yoav Artzi

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:43.784836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.309955Z digest=sha256:261ad2ef385fd2bd1beb20f1c08a4736dd520f287154a23e6b15c73c4599166e

Observation 2e0eda0c-4008-4bed-9363-b4aec77c412d · outbound

This paper cites Dilated temporal relational adversarial network for generic video summarization.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Dilated temporal relational adversarial network for generic video summarization

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:43.717687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.314461Z digest=sha256:da35ce2ccce03bbe2407c7e0f81510bcdd9919b021bfe90cae9f182b17258edb

Observation 7f1238d6-d592-4abf-8762-97061ca11778 · outbound

This paper cites Dense video captioning using graph-based sentence summarization.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Dense video captioning using graph-based sentence summarization

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:43.704809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.317969Z digest=sha256:86c36ee0561101938738401da824c917af6447927f1afd3312b889bdfd82c582

Observation 897c983f-7c11-485f-a915-05066dc267e2 · outbound

This paper cites Surgical activity recognition in robot-assisted radical prostatectomy using deep learning.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI Surgical activity recognition in robot-assisted radical prostatectomy using deep learning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:43:43.692625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:43:43.323375Z digest=sha256:8cebc227c47676bbd69feb41f383ee729cb4d5f6257292470b2373391089df64

Pith citing papers

Observation 5758488c-4032-45a5-a1b7-4bada09099ed · inbound

From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation cites this paper.

From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:18.685120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T03:15:40.944411Z digest=sha256:d5f1d1d44661a4cc6c929cd866ff1e7fa8653d956c9a9d2a79d93d8bedd460b9