Pith. sign in

Paper Citation Record · LEDGER

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing

As of 17 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2501.07554.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07554 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:42:31.879391Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact3
  • verified fuzzy21
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 767c366f-a965-4f64-9df7-96d07d65ab07 · outbound

This paper cites Detectron2 object detection & manipulating images using cartooniza- tion.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Detectron2 object detection & manipulating images using cartooniza- tion

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.975975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.559158Z digest=sha256:a6f01920b2f12c8e6b5aa6c15bde79b545738a8fdc560470ab558c6483b3f6b3

Observation 653c8f41-7a89-48b1-baa2-0b89e24d5a83 · outbound

This paper cites A deep learning framework for quality assessment and restoration in video endoscopy.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing A deep learning framework for quality assessment and restoration in video endoscopy

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.955084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.570712Z digest=sha256:fc04da04f1c7ca0a3c7e3df57f192b00d72f9ab30fcb896280d467d7ad7fb281

Observation 942d282c-5328-4b35-a70c-365d22da9a9a · outbound

This paper cites Text2live: Text-driven layered image and video editing.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Text2live: Text-driven layered image and video editing

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.928173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.578175Z digest=sha256:f3b822db041af0714df836c08391c0411f646b39b8cb78a015ed535de3ee70cf

Observation 3aabbcc7-f49a-4d2d-b84d-8476f907a3d5 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:42:31.590708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:42:31.590708Z digest=sha256:d2f0ea21ee410a80899b818e46df9c3c234ea1cedf85191dfffce3b38f7b4c74

Observation 128ee8d3-6fdb-4d6e-a6e0-f45f95cd55a7 · outbound

This paper cites Streaming Video Diffusion: Online Video Editing with Diffusion Models.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Streaming Video Diffusion: Online Video Editing with Diffusion Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:42:31.599729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:42:31.599729Z digest=sha256:84381c1fc94a2a5a509100a62eaff700654e6c1910f01fe33c5c18becb67700d

Observation 094c37a2-89ea-44f0-bf68-1918eebe2c47 · outbound

This paper cites Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:42:31.609586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:42:31.609586Z digest=sha256:ec2b650be6c9219d2a575fca866f1b7e2842afa8574cf9636df593cc779265f5

Observation 3ea98bb1-999f-4318-9c07-99ef3c32ea2e · outbound

This paper cites EditBoard: Towards a Comprehensive Evaluation Benchmark for Text-Based Video Editing Models.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing EditBoard: Towards a Comprehensive Evaluation Benchmark for Text-Based Video Editing Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:42:32.192878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.618642Z digest=sha256:313659c6b5a069168ca79d5d2f1499ae4818c2e43b49dc50aef9a887fe18e725

Observation 4f839925-9cd3-4f87-a370-b4f8dd1f9097 · outbound

This paper cites Clip-adapter: Better vision-language models with fea- ture adapters.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Clip-adapter: Better vision-language models with fea- ture adapters

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.893419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.629915Z digest=sha256:ae97b1fc83f8b36cd1b7b16d92fb4ff7bd9854c82bb0c1f6963c39ca830a997f

Observation 77dce725-1a90-42c1-8461-636a555e1769 · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:42:31.638700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:42:31.638700Z digest=sha256:07d27f8481593095975c45ef8015a51b022c5f8732c759294ca470d3114a9826

Observation 8e1e698a-bd57-4bc2-8507-4990874be4c2 · outbound

This paper cites Enhancing the video editing capabilities of text-to-video generators using ddpm inversion.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Enhancing the video editing capabilities of text-to-video generators using ddpm inversion

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.850981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.647850Z digest=sha256:e4690648b17d09797d05eec1bb408319017e594d7c9b193c84b0b34429885dc6

Observation c0f74ddc-28c9-4946-aa9e-8a805ae7567c · outbound

This paper cites Text-based Talking Video Editing with Cascaded Conditional Diffusion.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Text-based Talking Video Editing with Cascaded Conditional Diffusion

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:42:32.112207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.655451Z digest=sha256:2b1957d4df618a0984c08d64fe5da6c66d99916bbc2dbb1f665a0a32123a1c0f

Observation 7ee745d7-dfcc-4819-893c-bed2e4b7139e · outbound

This paper cites ultralytics/yolov5: v6.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing ultralytics/yolov5: v6

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.822953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.665705Z digest=sha256:629be19c3d66122795532bb85145acd3ff1a49a9826e56e992ffc018dc07d073

Observation 521c2502-1795-4146-ad13-6ade0265f007 · outbound

This paper cites Vilt: Vision- and-language transformer without convolution or region su- pervision.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Vilt: Vision- and-language transformer without convolution or region su- pervision

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.800719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.674915Z digest=sha256:14ac84e8f9538d54b81b6a7eb1580d8d1a9653c8308a91f8555ab7c6c8ab6cf7

Observation c7a53771-99f1-4548-a29b-6de66bc4ec26 · outbound

This paper cites Segment any- thing.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Segment any- thing

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.774534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.684518Z digest=sha256:faf7bfb48f3ebb8e4aa4bf96151a928952269610f0d13de21997dcd202e31df0

Observation 47f2576a-b26c-4b9b-ac2a-11040c167a98 · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:42:31.692746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:42:31.692746Z digest=sha256:e3c7a147d2e4587fa2d6f5b38fffd1bc1ab65c59d58db7cd300f98eb529965b8

Observation 79454fb2-a119-498d-90c7-57f9f6afec31 · outbound

This paper cites Shape-aware text-driven lay- ered video editing.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Shape-aware text-driven lay- ered video editing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.753052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.701768Z digest=sha256:f8c08d14329aa8739720ee9b39a5fde9e599b81c08c8156a1bf55b627714891f

Observation beca871b-6f7f-4781-9cc8-7c2bc8a4c4a8 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:42:31.708208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:42:31.708208Z digest=sha256:7a43c7be2fe3847aec79c242f5489461fdd6d8c0324c4873c317b97d43d9ca28

Observation 20cc6552-8158-4efe-a95c-7fd4a5cce111 · outbound

This paper cites Vidtome: Video token merging for zero-shot video editing.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Vidtome: Video token merging for zero-shot video editing

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.702555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.719066Z digest=sha256:0f6a7bb9670e71f6ed6564f8bd813ed3bdeb28f5b3c27c93eff25df7c9fd4a23

Observation ae2c48f6-4d0a-431f-8163-7adc9f0db45b · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.676393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.729724Z digest=sha256:18a5c347659d1bcb5c78185afecea8b74ff6b85e9e4a4cee95d6ac46a2b87762

Observation 2aae9f3b-ea1c-4fbd-80a5-d965136920b3 · outbound

This paper cites Video-p2p: Video editing with cross-attention control.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Video-p2p: Video editing with cross-attention control

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.646449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.737843Z digest=sha256:d58a1cc58ade1420177c73c850ae3f2ed7c20993dd175e313214e79a84bb0a70

Observation 7818f995-9564-4e6d-b8b9-1c6be52192c9 · outbound

This paper cites TrailBlazer: Trajectory Control for Diffusion-Based Video Generation.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing TrailBlazer: Trajectory Control for Diffusion-Based Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:42:31.744859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:42:31.744859Z digest=sha256:d383537960947a76ef06a72c40e8ec9ab45d1b3b3e1e670fd42bf770528676c5

Observation ee6f2988-171e-4fe0-9bea-86a3cafb8b7a · outbound

This paper cites Gazed–gaze-guided cinematic editing of wide-angle monocular video recordings.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Gazed–gaze-guided cinematic editing of wide-angle monocular video recordings

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.617147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.752271Z digest=sha256:96b68ba2db966c369432874f35740de7e2be7fc03fe05f68d6d5666183b8338e

Observation ad541a63-3434-4ea7-a4c6-be9aabd57398 · outbound

This paper cites Fatezero: Fus- ing attentions for zero-shot text-based video editing.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Fatezero: Fus- ing attentions for zero-shot text-based video editing

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.588539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.764726Z digest=sha256:c0de27234c63c8e11275d369890ed5176d83baeb9449c5bab194ad607498a1cd

Observation 2540c33d-c6d7-4004-afa9-625c574a117f · outbound

This paper cites Enhanced end-to-end video editing: Adaptive customization of path, object, and motion dynam- ics.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Enhanced end-to-end video editing: Adaptive customization of path, object, and motion dynam- ics

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.561206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.779861Z digest=sha256:c4376a21e08d566c2b01ab24ac5fc641e6a1178478546e83f3704580c5a2a5f2

Observation 1b9d67d4-c46f-4683-805a-662d556dcbc5 · outbound

This paper cites Pro- cedural crowd generation for semantically augmented virtual cities.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Pro- cedural crowd generation for semantically augmented virtual cities

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.533346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.791223Z digest=sha256:400c79d3ae830b751be9a946aa5a08775448a7e31ca891f496a92de5b62c3508

Observation c3c4c8ca-a9d7-4fce-b741-b6e35ebebe80 · outbound

This paper cites Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:42:31.800012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:42:31.800012Z digest=sha256:fc73d98d8de4ebf7e2cee95f3073eb18d680a37b4f792762da58177e731ce012

Observation bd7b1a1b-5fee-4b89-82a7-80d80ec6358f · outbound

This paper cites Actionclip: Adapting language-image pretrained models for video action recognition.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Actionclip: Adapting language-image pretrained models for video action recognition

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.492677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.807770Z digest=sha256:912018f39436159806b98716eb12ff1c34939d09604a42ed4979f6253caecd24

Observation 8ad1278e-5828-4f00-ad50-a55d8667cae8 · outbound

This paper cites MedCLIP: Contrastive Learning from Unpaired Medical Images and Text.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing MedCLIP: Contrastive Learning from Unpaired Medical Images and Text

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:42:31.815860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:42:31.815860Z digest=sha256:0b55f0493a271db204d6aab2c1badc2540dc4c7d66764110f1e14596fd283acf

Observation c9d04de6-7bf2-4d94-9cac-776af417d893 · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:42:31.825184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:42:31.825184Z digest=sha256:acc3eb62ae084e6d758730999c95a8a9cf74fe73c4c6cd1cdca8c9334a386570

Observation 93072393-8846-4e5d-b095-09bb36a68b05 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.447522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.834435Z digest=sha256:56585e72879f26a1daa7c60b6d652009e5c6be53256ff58896174115d834a65a

Observation 0e1a16af-b1a6-41a4-9fec-753fb24a3d02 · outbound

This paper cites Temporally consistent semantic video editing.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Temporally consistent semantic video editing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.421605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.844142Z digest=sha256:46b27ce00e427d7937ba8022f45324912fd5acf4a6b56a0e09cc349ea9231826

Observation 0f59556f-a314-4459-9846-2604fd3f692a · outbound

This paper cites Context-aware talking-head video editing.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Context-aware talking-head video editing

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.391770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.851364Z digest=sha256:05804a0a60e7166f59b8f6466c51dee88cd5139ec388b42a3448052f60e33d82

Observation 72ac9507-7d40-4dc9-aa4a-0c4538349109 · outbound

This paper cites Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability,Reproducibility, and Practicality.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability,Reproducibility, and Practicality

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:42:32.004189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.864250Z digest=sha256:4133220b751073000034989b315840631a0b90df9be52654ea12a06f4896d551

Observation d4388d50-4b4b-4850-94d6-cf47385e8094 · outbound

This paper cites Motiondirector: Motion customization of text-to-video diffusion models.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Motiondirector: Motion customization of text-to-video diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:42:32.353751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T20:42:31.873471Z digest=sha256:3f0ac7630b58299e8ba61b72335612c6af807e0d75bdef16b88c57a9005b6fa1

Observation ae470dc1-653f-435b-b9e0-1552428baf80 · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T20:42:31.879391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:42:31.879391Z digest=sha256:ce315e70526cbe801d8485663d789df8995ea0743fd0dfd1ee066a4182570c44

Pith citing papers

No inbound Pith citation observations are available.