Pith. sign in

Paper Citation Record · LEDGER

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction

As of 7 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2507.16718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.16718 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:06:22.975019Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 38c28270-dfbb-43b8-9abc-c2780e9f88a4 · outbound

This paper cites Qwen2.5-VL Technical Report.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.938905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.938905Z digest=sha256:63921c221f268d9d9f034c271bb13532f6700c1dca42eecbade50d2b94b6b1ac

Observation 3f2176c7-d82d-4ff1-a70c-a5d5436599ad · outbound

This paper cites Towards Intelligent Speech Assistants in Operating Rooms: A Multimodal Model for Surgical Workflow Analysis.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Towards Intelligent Speech Assistants in Operating Rooms: A Multimodal Model for Surgical Workflow Analysis

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.942371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.942371Z digest=sha256:d8abe3017e01aa37e9c77ea1e507c30117a7e78d01cebdd1ccee8e472a4df04f

Observation 51299db0-0732-41a4-8e82-c53c0aa15b54 · outbound

This paper cites A spatio- temporal network for video semantic segmentation in surgical videos.International Journal of Computer Assisted Radiology and Surgery , 19(2):375–382, 2024.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction A spatio- temporal network for video semantic segmentation in surgical videos.International Journal of Computer Assisted Radiology and Surgery , 19(2):375–382, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:06:23.119196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:06:22.945334Z digest=sha256:ca0981e9359d2fce31009eb53bcc81f093fe20dbfd665e98a0c51d4df3ad8895

Observation f854057d-338e-4e6c-9d27-a28e435522a2 · outbound

This paper cites Temporal memory relation network for workflow recognition from surgical video.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Temporal memory relation network for workflow recognition from surgical video

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:06:23.109869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:06:22.948068Z digest=sha256:89d929f9ef851f1d86edc581cc7a73708773fb36c8bc74a3e022857e1f64abb5

Observation b557b853-3366-48ef-884e-1855b990b60c · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Lisa: Reasoning segmentation via large language model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:06:23.100640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:06:22.950833Z digest=sha256:9dc764dc20a65d9b70fa9103ac6af6e532cbba3d6e559cefbf911b4e6d11f72b

Observation bcda117c-9cef-428e-9288-5287183f452f · outbound

This paper cites Improved baselines with visual instruction tuning.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Improved baselines with visual instruction tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.953421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.953421Z digest=sha256:deb1f66ddc6477f00f453199e1d47c2fbfe4def29307ec8264ea84bed422fbf8

Observation e9ee0c61-ecca-434e-b589-1fe233297714 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction SAM 2: Segment Anything in Images and Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.956091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.956091Z digest=sha256:536488ab50fc90a83d86a1ca8adf46d470e9c130ef39c2be82efa46f0fc67e74

Observation 0e0199d5-8493-4970-832a-e4eeb65e3214 · outbound

This paper cites Position: Foundation Models Need Digital Twin Representations.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Position: Foundation Models Need Digital Twin Representations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.959059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.959059Z digest=sha256:1ef0ccf8aca4709c0b64a7d572ff2aca7a67f96b887999964a38ab0b70e19618

Observation 9c38a73e-c019-4a1e-a51a-9686a1255dcc · outbound

This paper cites RVTBench: A Benchmark for Visual Reasoning Tasks.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction RVTBench: A Benchmark for Visual Reasoning Tasks

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:06:23.035823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:06:22.961694Z digest=sha256:4118fddf66c3dd5940cfa15234df04af4b2be85026061af458156085431261f6

Observation caf26f12-4e97-48c3-b8e7-52893dc495fb · outbound

This paper cites Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.964559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.964559Z digest=sha256:bd9abb14dac795012466f06dec1e2449e90bdaf91c1cbd82844ffe660c8b55e3

Observation 6507a37d-5882-44da-844d-3acb78c39a8e · outbound

This paper cites Reasoning Segmentation for Images and Videos: A Survey.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Reasoning Segmentation for Images and Videos: A Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.967237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.967237Z digest=sha256:dcd5c9f161263e5c42073140caf4a9512ca363a4c19926c163ccbdedf5c0d42b

Observation 86b35c74-1a23-4b64-8af0-19d51503112d · outbound

This paper cites MVOR: A Multi-view RGB-D Operating Room Dataset for 2D and 3D Human Pose Estimation.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction MVOR: A Multi-view RGB-D Operating Room Dataset for 2D and 3D Human Pose Estimation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.969839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.969839Z digest=sha256:189edd01557ccfc4be7b7d08a2d3dcf218b1179be297864b11c9957512507842

Observation bda4ed03-9f55-4b56-a48d-4da90aa1cd5f · outbound

This paper cites an unresolved cited work.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:06:23.084249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:06:22.972763Z digest=sha256:641a45be8d157c4272821478f0d17830929823b0cfbb073a49b2baa1c2508c68

Observation e6792027-f9b5-4af3-855a-b28b710000d3 · outbound

This paper cites Depth Anything V2.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Depth Anything V2

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.975019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.975019Z digest=sha256:0119798aa1c23e46b10026d4402863e976d95504e5f0484b3a0252b171c216a3

Pith citing papers

No inbound Pith citation observations are available.