Pith. sign in

Paper Citation Record · LEDGER

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection

As of 21 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2505.06663.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.06663 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:41:45.860765Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9a44a258-fa6a-4988-9f77-f78ce830bde1 · outbound

This paper cites Social fabric: Tubelet composi- tions for video relation detection.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Social fabric: Tubelet composi- tions for video relation detection

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.362871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.721310Z digest=sha256:e5b971566237349b368a4af039cbc93705f9a0377db210f8f7aa3d8e00b7176b

Observation 4865361f-2454-4182-a8ab-0ef5dc6e863b · outbound

This paper cites Stacked hybrid-attention and group collaborative learning for unbiased scene graph generation.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Stacked hybrid-attention and group collaborative learning for unbiased scene graph generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.332745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.732086Z digest=sha256:4d69bd4b9c6dfd46bb47529a12efbae13685dca6af7ee0f1548bfd8caedae28f

Observation f700c645-efb0-4785-bd3e-34714c17aaae · outbound

This paper cites Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T22:41:45.746728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:41:45.746728Z digest=sha256:33fb7f918306a2f0df789677176c4475031195b844f19a9f739cc1c4bf1b2794

Observation e2b5a2b0-fa23-4e20-a875-732070209d66 · outbound

This paper cites Align and prompt: Video-and-language pre-training with entity prompts.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Align and prompt: Video-and-language pre-training with entity prompts

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.259711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.762209Z digest=sha256:6ae5d867cbeb2f94006fc50fac59592699d5dd238f7c3d3646e719a94c7ce637

Observation 25f9e06b-b62b-40b4-a7f6-cf4e9d0b1d7a · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.242947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.766775Z digest=sha256:8fdf9eea27fa28ca956f654b78a1be0e644df4e41ed8d34acb2610eb70e0d76c

Observation 709f1e32-706a-4a16-b199-8d1e8c122a42 · outbound

This paper cites Open-vocabulary se- mantic segmentation with mask-adapted clip.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Open-vocabulary se- mantic segmentation with mask-adapted clip

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.227306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.771129Z digest=sha256:0a14fa06d1fa3172736799d7b37139942ed66bfd8f14b24619784e513180f82b

Observation f489c0d1-f809-4de0-844d-ba474d3a15b2 · outbound

This paper cites Microsoft coco: Com- mon objects in context.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Microsoft coco: Com- mon objects in context

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.211167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.775793Z digest=sha256:dfdd900b2e86497356f3c3943c9b02025fddac2262267bdca75bdcdb9459f429

Observation 7a6c256d-2399-4728-a92c-0a3de49f7383 · outbound

This paper cites Beyond short-term snip- pet: Video relation detection with spatio-temporal global context.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Beyond short-term snip- pet: Video relation detection with spatio-temporal global context

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.177273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.784530Z digest=sha256:3e7f5faba88a536f9d5c8a3ec81ed4049ef6672d466e13bc9d062d1947dc9445

Observation eda83f8f-7f36-401f-87dd-b04ba1bff1f2 · outbound

This paper cites Decoupled weight decay regularization.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Decoupled weight decay regularization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.147348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.793440Z digest=sha256:05c270f1156477c9473f838eaea4b3765b293f0b5bb253891f67e914799d8964

Observation ed689957-151d-41e3-b2ba-2827023e3c6c · outbound

This paper cites Video rela- tion detection with spatio-temporal graph.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Video rela- tion detection with spatio-temporal graph

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.116985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.802118Z digest=sha256:686f896e9f04d27b1fcd6a3bca6167b1d39265cf3506e0e1f7648116aa71b7b6

Observation 724552ac-edcd-44b9-995d-498e94d4fdb2 · outbound

This paper cites Learning transferable visual models from nat- ural language supervision.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Learning transferable visual models from nat- ural language supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:41:45.806431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:41:45.806431Z digest=sha256:7a4fcbf88637b7de2d460428c9b09964f66d7311433531856582b371f4a33304

Observation 848d8766-596f-4860-b266-a2baafb88b47 · outbound

This paper cites Ac- tion scene graphs for long-form understanding of egocen- tric videos.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Ac- tion scene graphs for long-form understanding of egocen- tric videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.091473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.810849Z digest=sha256:deb489fe40e6bf368bf2155584357fda2d7f3019aeac135c01e8a5b5237ce83a

Observation d2de1639-c2ad-4b07-ad46-b216818c5398 · outbound

This paper cites Video visual relation detection.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Video visual relation detection

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.076451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.816130Z digest=sha256:4a87eac81d536c1de0cf8ed689ac20d673a59fe80c3d443d5d55828cd1b25460

Observation 7df242c1-d4c0-451b-84f0-99c6c7313c7f · outbound

This paper cites Video visual relation detection via iterative inference.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Video visual relation detection via iterative inference

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.045060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.825138Z digest=sha256:068aa4f5086f72757a430015272a4f652436292424f69411174ffa93230b1344

Observation c69762b8-6d8c-446f-8880-33a6f93ea173 · outbound

This paper cites Video relationship reasoning using gated spatio- temporal energy graph.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Video relationship reasoning using gated spatio- temporal energy graph

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.029000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.829618Z digest=sha256:8b321490ba9b56b10f55937785d36d8ee58a9160ae6394fcd20fb0433a492380

Observation 1991fd2c-6f0d-4f8f-9897-e1640307b889 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection ActionCLIP: A New Paradigm for Video Action Recognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:41:45.834024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:41:45.834024Z digest=sha256:ebee153afd30f1fa77172d47b9fee51ef1ce3f3f3875339f3d358fc97ac484a1

Observation bf2e9519-5265-4148-b797-a4d83d40582c · outbound

This paper cites End-to-end open-vocabulary video vi- sual relationship detection using multi-modal prompting.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection End-to-end open-vocabulary video vi- sual relationship detection using multi-modal prompting

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.011496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.838837Z digest=sha256:56eb44cb8484f8e7496d9afe315935a1febaca02a3eb8a740c3f3efe03dd1ee9

Observation a3d0b1ad-6a10-4483-a461-3be0f514850e · outbound

This paper cites Meta spatio-temporal debiasing for video scene graph generation.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Meta spatio-temporal debiasing for video scene graph generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:45.994364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.843031Z digest=sha256:df3fb1dba60144954c713f2eec4d86b371c9b24e98df12910d0804f91f9b97c3

Observation ae361d15-f1d6-4979-97f7-c221f6ae364e · outbound

This paper cites Multi-modal prompting for open- vocabulary video visual relationship detection.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Multi-modal prompting for open- vocabulary video visual relationship detection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:45.978792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.847417Z digest=sha256:c73a186d14c01f0923b051c60f3e5ff39a456a8b54f5931235c67b572cab8eb1

Observation 9cb81df1-d2ed-46c6-b49e-e03a493e68ad · outbound

This paper cites End-to-end video scene graph generation with temporal propagation transformer.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection End-to-end video scene graph generation with temporal propagation transformer

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:45.963036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.851935Z digest=sha256:857cc1303d003b7b51f39074e20f14c58d51fb42149879e74b8e8f11fe56b8a5

Observation ebf5f5b9-b3ea-45ab-9e22-fda78c4abd67 · outbound

This paper cites Constructing holistic spatio-temporal scene graph for video semantic role labeling.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Constructing holistic spatio-temporal scene graph for video semantic role labeling

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:45.946261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.856480Z digest=sha256:1f0020764e33db115239161f1e504ea549b0e01a789cd01023571c636b08d694

Observation 8b36e4ea-9269-4f8a-9a35-353f3989b246 · outbound

This paper cites Vrdformer: End-to-end video visual relation detec- tion with transformers.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Vrdformer: End-to-end video visual relation detec- tion with transformers

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:45.929043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.860765Z digest=sha256:f36036e7482188688390909d015c5a0b64cb77ab72036bb52f6e43cb5e548f53

Observation 0aa98295-6574-4cd7-877f-f145672e1a54 · outbound

This paper cites Vrdone: One-stage video visual relation detection.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Vrdone: One-stage video visual relation detection

Reference 2002

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.274458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.757440Z digest=sha256:fe0695f052759620e95b68a6a8f25aec88d1ac403e53dcb27d97729c35255e19

Observation ea660bf8-62a9-490d-a8f6-ee0baef8ef68 · outbound

This paper cites Td2-net: To- ward denoising and debiasing for video scene graph gener- ation.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Td2-net: To- ward denoising and debiasing for video scene graph gener- ation

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.194323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.780232Z digest=sha256:385d105bd791098d82a975cd7f4a8f1e4300012a7a257a444ed8790d96b62703

Observation b00aaeaf-f18d-4761-a688-6366966ec79e · outbound

This paper cites Relation understanding in videos: A grand challenge overview.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Relation understanding in videos: A grand challenge overview

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.061364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.820571Z digest=sha256:fa0d8d6a36f6511d82a00a80e254113fa55c7e420f1c1f8958bce673e07283de

Observation 19bc3818-e824-4a61-875c-fca8a92bedcf · outbound

This paper cites Hig: Hierarchical interlacement graph approach to scene graph generation in video understand- ing.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Hig: Hierarchical interlacement graph approach to scene graph generation in video understand- ing

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.131880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.797527Z digest=sha256:27f9d7c5cc69f0c1e2f895b07b214a4e2e18bf4a503c3667fdb66f8809a5e678

Observation 30c5a471-8bc5-4fd6-9e33-70e234e22a04 · outbound

This paper cites Open-vocabulary segmenta- tion with semantic-assisted calibration.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Open-vocabulary segmenta- tion with semantic-assisted calibration

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.162125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.789103Z digest=sha256:e6d327425a3edcb801029479a502bac85424c602999ba5641b2116b25d9cd41d

Observation 07f8fa9b-dee5-4aff-bd20-a3b3a5f96e0a · outbound

This paper cites Spatial-temporal transformer for dynamic scene graph generation.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Spatial-temporal transformer for dynamic scene graph generation

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.347981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.727385Z digest=sha256:535be7dce241d50e7a786d86932473edf300e80f99220bdba550fe0f7d96ed77

Observation a83ea7af-c81b-4f15-a635-16f5b66dd6b1 · outbound

This paper cites Active open-vocabulary recognition: Let intelligent moving mitigate clip limitations.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Active open-vocabulary recognition: Let intelligent moving mitigate clip limitations

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.318507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.736859Z digest=sha256:43175aa0712d612b23a0793ea16c53e5294efaad373489e64f2171cbbf5a33b2

Observation 1c19a07e-4357-46b9-b8ee-88f0eefa85aa · outbound

This paper cites Stochastic neighbor embedding.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Stochastic neighbor embedding

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.288603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.752680Z digest=sha256:6c834c733e8383b520a498dc56d45c4c0cfa722fab1d2eb0a6dbe9a07ecddff5

Observation d9e257df-f28e-4b23-a638-af1d45369923 · outbound

This paper cites Simple image-level classification improves open-vocabulary object detection.

METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection Simple image-level classification improves open-vocabulary object detection

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:41:46.304045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:41:45.741825Z digest=sha256:236e9b9b7403600def42c211a8bb217e8e5be43719aa063f65036c2b6c515b8b

Pith citing papers

No inbound Pith citation observations are available.