Pith. sign in

Paper Citation Record · LEDGER

Sharingan: Extract User Action Sequence from Desktop Recordings

As of 18 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2411.08768.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.08768 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:25:57.458346Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 83ff7106-7c8c-4849-8f84-6d86cb4b9c59 · outbound

This paper cites an unresolved cited work.

Sharingan: Extract User Action Sequence from Desktop Recordings Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:25:57.882023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.329469Z digest=sha256:20c462d6eb43530a2f204fb15d258a4d3fe980a6016608d76421f3b7e475a898

Observation 00a0fe10-4be5-4260-854d-63449dcc48fe · outbound

This paper cites an unresolved cited work.

Sharingan: Extract User Action Sequence from Desktop Recordings Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T21:25:57.334259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:25:57.334259Z digest=sha256:4b5a68726fa8f179521305dce2c14bec3dab1c0a4f70682ac06d24f14afc51a5

Observation 68732f94-8c18-40a7-b17e-6253cb1d3cd9 · outbound

This paper cites Gui-world: A dataset for gui-oriented multimodal llm-based agents, 2024.

Sharingan: Extract User Action Sequence from Desktop Recordings Gui-world: A dataset for gui-oriented multimodal llm-based agents, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.862547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.338531Z digest=sha256:992841bbe67608843191cc85367359734f0e2f116548ffd3154b29ddf3159189

Observation 00c757c2-c5fb-4f95-8913-8350b7f206f0 · outbound

This paper cites Vg4d: Vision-language model goes 4d video recog- nition.

Sharingan: Extract User Action Sequence from Desktop Recordings Vg4d: Vision-language model goes 4d video recog- nition

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.850388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.342821Z digest=sha256:dfa20148a2adbb7a9fc01c8e350020dbd46ffc4377bcb58f09f792ddf8a60561

Observation 5e58c4ea-2cfa-4003-a831-f0dff2666657 · outbound

This paper cites Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd, 2024.

Sharingan: Extract User Action Sequence from Desktop Recordings Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.837336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.347203Z digest=sha256:8087d2126dd2ad8b759e9ab72f7509a3a87cfc940b0333a42917007ef6a3a102

Observation 355b6808-dc24-4c77-982f-c2d9b5d89bac · outbound

This paper cites Videoagent: A memory-augmented multimodal agent for video understanding, 2024.

Sharingan: Extract User Action Sequence from Desktop Recordings Videoagent: A memory-augmented multimodal agent for video understanding, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.824397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.351443Z digest=sha256:b04eb35ecd850e9a31cdbc08e0a185ff7ea4fab476445fc2105b4df02e634edf

Observation e3196b28-7e80-49e6-88a4-23ab1557d86c · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition.

Sharingan: Extract User Action Sequence from Desktop Recordings Video-of-thought: Step-by-step video reasoning from perception to cognition

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.810735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.356101Z digest=sha256:2fc37ecbd3b526b8f27e70e861c3e0c8ccb70437db8094bf44c79badf9ffa74a

Observation 5a33fbcd-994a-41ab-b620-37c3d3653a5c · outbound

This paper cites Gemini models.

Sharingan: Extract User Action Sequence from Desktop Recordings Gemini models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.796532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.359971Z digest=sha256:21c7694f3a1093040c2fc78c6db0a37e6be692d92f9a9927be20fca13145df0d

Observation 6f931a34-e272-45a8-87dd-924404ff4428 · outbound

This paper cites an unresolved cited work.

Sharingan: Extract User Action Sequence from Desktop Recordings Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:25:57.784644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.364174Z digest=sha256:c1f38b42a417161ee4f3edc403eddf086bab25b0951f017ad961a18a9bbef293

Observation b0519d1d-4927-4f95-8f48-a11fbcd5ead7 · outbound

This paper cites Smartflow: Robotic process automation using llms, 2024.

Sharingan: Extract User Action Sequence from Desktop Recordings Smartflow: Robotic process automation using llms, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.773669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.368073Z digest=sha256:737d27725d6f89ab95bc4755917967dc05a9976da551de60b85a8f328b812e77

Observation d7429d13-c4e0-4478-a73a-3e1fc0f3b75e · outbound

This paper cites Dream2real: Zero-shot 3d object rearrangement with vision-language models.

Sharingan: Extract User Action Sequence from Desktop Recordings Dream2real: Zero-shot 3d object rearrangement with vision-language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.760739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.372039Z digest=sha256:dc1813ccad9af19d0c1a8c65ed6b24755f75bd19e6da073c08343a6c163839f6

Observation 0fdc0360-92bd-4fab-8c77-6a1e5bd41df9 · outbound

This paper cites Wolf: Captioning everything with a world summarization framework, 2024.

Sharingan: Extract User Action Sequence from Desktop Recordings Wolf: Captioning everything with a world summarization framework, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.747701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.376285Z digest=sha256:9bc5ec88608777bd9f82a032bdf757c09f7d80549a0a02f8cf91829d8ecf8c10

Observation b4cd499a-0675-4d0a-a5c5-b0cc12bc989e · outbound

This paper cites Videochat: Chat-centric video understanding, 2023.

Sharingan: Extract User Action Sequence from Desktop Recordings Videochat: Chat-centric video understanding, 2023

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T21:25:57.381034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:25:57.381034Z digest=sha256:5daa3457fef459440c40f87e68015ecda0a24799a56ee4aa6bcaeac1a9dc94dc

Observation d95b2274-eb15-4d8d-bd30-8ffcf2d5af75 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

Sharingan: Extract User Action Sequence from Desktop Recordings Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T21:25:57.384739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:25:57.384739Z digest=sha256:53af037aa83bffb5e4c37ce018b38a68162494b01aa7a538b7546d5d7366b923

Observation 5b7f4954-487a-4a59-9acd-d9488ae86eb9 · outbound

This paper cites Visualwebbench: How far have multi- modal llms evolved in web page understanding and grounding?, 2024.

Sharingan: Extract User Action Sequence from Desktop Recordings Visualwebbench: How far have multi- modal llms evolved in web page understanding and grounding?, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.721681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.388998Z digest=sha256:9c8106f9fbc51f43bc913f4f6976bb99065a806bba63db3a787c97366b6da7f2

Observation ff63183a-4361-4743-9227-3b5f7be2dbc9 · outbound

This paper cites Online robot navigation and manipulation with distilled vision-language models, 2024.

Sharingan: Extract User Action Sequence from Desktop Recordings Online robot navigation and manipulation with distilled vision-language models, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.709262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.392679Z digest=sha256:fabd6cf092a95ab3c1f043f9457d0297da09f87731103443ed6b1c934bd5e073

Observation 66795c9f-ab15-4121-aa2a-83e2b4443867 · outbound

This paper cites Robollm: Robotic vision tasks grounded on multimodal large language models.

Sharingan: Extract User Action Sequence from Desktop Recordings Robollm: Robotic vision tasks grounded on multimodal large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.697203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.396215Z digest=sha256:41d39fd3f3e5b8e1c2d99a1e049108688019bfdfe154d21d15f765160b5f65ed

Observation 9eece4f9-5a86-4fa3-8b58-2f8c5b8b2f6e · outbound

This paper cites Omniparser for pure vision based gui agent, 2024.

Sharingan: Extract User Action Sequence from Desktop Recordings Omniparser for pure vision based gui agent, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T21:25:57.399303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:25:57.399303Z digest=sha256:a0f201c04b4486a88624eeaffc079ccdad51a2049467ed3d3e9d793dcc92c771

Observation adc0d112-8c0f-4e39-a218-7e07b6d9ca02 · outbound

This paper cites an unresolved cited work.

Sharingan: Extract User Action Sequence from Desktop Recordings Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:25:57.677338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.402623Z digest=sha256:c6247ffc92eeefbf47efbbb44bf4ccb66249b4c7b26862cae83588494165ecce

Observation b7b29a5f-30c8-4001-97d6-7ae24254edaf · outbound

This paper cites an unresolved cited work.

Sharingan: Extract User Action Sequence from Desktop Recordings Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:25:57.666488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.406409Z digest=sha256:b7d9818048d95dd568d6bdec0cd3ea586c524042b57f285ebb4ce938e7d342c7

Observation e008d938-d9a3-4770-a1f0-49e6243e333f · outbound

This paper cites an unresolved cited work.

Sharingan: Extract User Action Sequence from Desktop Recordings Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:25:57.655923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.410428Z digest=sha256:056f88a7604d9663d7925286fde0c58a48a46f4a29d8da740027c251b0a3a818

Observation 0661f689-ed6b-4684-a9d1-bdd96b945477 · outbound

This paper cites Owl - always-on wearable ai.

Sharingan: Extract User Action Sequence from Desktop Recordings Owl - always-on wearable ai

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.644081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.413698Z digest=sha256:df34acaebe7da023506c5d42f087bec13223d09c89a5bf0316deb67546182407

Observation ca310773-f698-4aa7-8aa2-01c334297d8d · outbound

This paper cites bert-embedding.

Sharingan: Extract User Action Sequence from Desktop Recordings bert-embedding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.630518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.417347Z digest=sha256:aedfdc4b2f252c2b037d9902321d010381d16d9ddcfb92b41ccd5c6364cac3f5

Observation ad818ad9-c9f1-42ca-b44d-6e4f198ff4f4 · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

Sharingan: Extract User Action Sequence from Desktop Recordings Learning transferable visual models from natural language supervision, 2021

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T21:25:57.420574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:25:57.420574Z digest=sha256:6b612b830c5711c7e2a5e8855757dd2102bffaa23d544970ad285dfda45317b8

Observation 8caa44b6-4169-4161-acf9-66c196a9b58f · outbound

This paper cites Beltran-Hernandez, Masashi Hamaya, Atsushi Hashimoto, Shohei Tanaka, Kento Kawaharazuka, Kazutoshi Tanaka, Yoshitaka Ushiku, and Shinsuke Mori.

Sharingan: Extract User Action Sequence from Desktop Recordings Beltran-Hernandez, Masashi Hamaya, Atsushi Hashimoto, Shohei Tanaka, Kento Kawaharazuka, Kazutoshi Tanaka, Yoshitaka Ushiku, and Shinsuke Mori

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.609927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.423733Z digest=sha256:f6b73034826a042e3f06298d5baddd10a0f03b6bda645778329822bf7df190e8

Observation 3c09affd-f4ca-4217-b55c-1514fc754e33 · outbound

This paper cites Karlsson, Bo An, Shuicheng Yan, and Zongqing Lu.

Sharingan: Extract User Action Sequence from Desktop Recordings Karlsson, Bo An, Shuicheng Yan, and Zongqing Lu

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.597318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.426886Z digest=sha256:49750698c1ad8bbf0ea4aacdefd556540a5d04041b1391206c6a286d27ae7c7b

Observation 2bc09ff3-9aa8-4972-8322-d1c51f42eaf2 · outbound

This paper cites Sch ¨onberger, Juan Nunez-Iglesias, Franc ¸ois Boulogne, Joshua D.

Sharingan: Extract User Action Sequence from Desktop Recordings Sch ¨onberger, Juan Nunez-Iglesias, Franc ¸ois Boulogne, Joshua D

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.584298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.430293Z digest=sha256:a791517fd3dcaa8751e6c254102f64d9551d1ca589259438dc6661605326a7d6

Observation 84a6865f-48a3-4b4d-aaf9-caa12ae79997 · outbound

This paper cites Drive anywhere: Generalizable end-to-end autonomous driving with multi- modal foundation models, 2023.

Sharingan: Extract User Action Sequence from Desktop Recordings Drive anywhere: Generalizable end-to-end autonomous driving with multi- modal foundation models, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.572406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.434030Z digest=sha256:409f157d2e69618e2efb827f0ffaa4fae379a4e46110918b4476e99ad55c3ed8

Observation 40d8d6d9-6bff-4f49-a4e7-88fca66f861b · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent, 2024.

Sharingan: Extract User Action Sequence from Desktop Recordings Videoagent: Long-form video understanding with large language model as agent, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.559710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.438156Z digest=sha256:13e82b5ceb13c12afb74e6ab25717816094c434097b8a4b2154b7bba4d66324c

Observation 60abc502-0b87-4dc7-89d2-67dbcb3027f5 · outbound

This paper cites Videotree: Adaptive tree- based video representation for llm reasoning on long videos, 2024.

Sharingan: Extract User Action Sequence from Desktop Recordings Videotree: Adaptive tree- based video representation for llm reasoning on long videos, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.546641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.442072Z digest=sha256:73e439d93905f913461e992073d8368656e3821b9474e7552e587c08dabfab17

Observation fd03c251-ece6-4aee-8370-a81d11984e58 · outbound

This paper cites Incorporating scene graphs into pre- trained vision-language models for multimodal open-vocabulary action recognition.

Sharingan: Extract User Action Sequence from Desktop Recordings Incorporating scene graphs into pre- trained vision-language models for multimodal open-vocabulary action recognition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.532019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.445966Z digest=sha256:ca06bd49111198996db624c88655cd55840f155dd7c55e17f2d2c70364952bb3

Observation cb0ee32d-9b8f-4838-8764-11d6715bbb7c · outbound

This paper cites Vlfm: Vision-language frontier maps for zero- shot semantic navigation.

Sharingan: Extract User Action Sequence from Desktop Recordings Vlfm: Vision-language frontier maps for zero- shot semantic navigation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.520461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.450139Z digest=sha256:d063b7a3d0e683f4c6ff7867602ffb911ac78a63dffce14c145078883e6ba2ed

Observation fa58de06-7a6c-4e10-aeae-c06127417073 · outbound

This paper cites Ufo: A ui-focused agent for windows os interaction, 2024.

Sharingan: Extract User Action Sequence from Desktop Recordings Ufo: A ui-focused agent for windows os interaction, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.508722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.454509Z digest=sha256:d1b6e8ba75b0626732d849f51f8e09d042c4d211f9f57e0a779cb12020216069

Observation a81000a5-0509-47cc-8ec0-22ac18b91bfe · outbound

This paper cites global_description.

Sharingan: Extract User Action Sequence from Desktop Recordings global_description

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:25:57.495755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T21:25:57.458346Z digest=sha256:97bc4123ee4ed9e8bad1d68d79de4a1c002dc661c73ab172d3778aa38c20bbdf

Pith citing papers

No inbound Pith citation observations are available.