Pith. sign in

Paper Citation Record · LEDGER

HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2005.00200.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2005.00200 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:21:01.327799Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:38:56.173895Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f4092bb5-b83a-4515-add3-d7bca8a114e5 · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.162304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:ecc10cb860057bf213cf772432bf8cd3c89b1709758cb82621be403ae3df9ba7

Observation 1e4841d5-236e-42b0-a3c7-faf4e1082a84 · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.191460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:f1f6c6a2c7ba395d5ec7f5dffb910eb1a6d7419203f35c81d6dabc680db3553a

Observation d4736d93-9ddd-4276-a815-ba364627e2bb · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:32:18.256615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:5795bc3807d344dbfe2071b2910833a7a0e4a9333f8146c10397d1eba9c0faf9

Observation da815215-73cc-48dd-b574-c1ff35b1a32b · inbound

Generalizing vision-language models to novel domains: A comprehensive survey cites this paper.

Generalizing vision-language models to novel domains: A comprehensive survey HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 241

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:01.327799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:21:01.327799Z digest=sha256:7cd535f333fd3288632aad9e0c581e46e5f34cef6b37ebd1faa80f7f5d4b7711

Observation d9e9770b-45bc-4ef1-95c7-74c14caa103e · inbound

THYME: Temporal Hierarchical-Cyclic Interactivity Modeling for Video Scene Graphs in Aerial Footage cites this paper.

THYME: Temporal Hierarchical-Cyclic Interactivity Modeling for Video Scene Graphs in Aerial Footage HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:07:16.832602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:07:16.832602Z digest=sha256:5b9eb421047f303f6d2b317d39e50f528f17b85de93fcf8654fc7db8ca20c70a

Observation 0ca66b65-6b30-4633-b844-1cc8f35be18f · inbound

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition cites this paper.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:54.214011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:54.214011Z digest=sha256:496732cbb0612ddf4c221b588fd9bce1e0312488c1a17173b6392cb897767982

Observation 1186ada7-6db8-4df6-9898-174df4cae153 · inbound

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting cites this paper.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:57.908906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:57.908906Z digest=sha256:22cbaeddc0394f987c9c0b80f3068263b4b286cc292492f5f9fda8e119cd88a7

Observation 3254beaf-9463-4b14-8c2b-92733863bfeb · inbound

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding cites this paper.

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T13:01:58.914260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:01:58.914260Z digest=sha256:76dfba869e26c9b78d2819ff992e91926bea85ded043e869501b1e8d8bc4645c

Observation 12cd106b-3b83-45b6-ab5a-bb008a2f4481 · inbound

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding cites this paper.

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T06:36:33.848152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:36:33.848152Z digest=sha256:4f71c88c4fd8d14ab74a6a0c66b835ed11af1d6c51c393dfa382ee32590ac900

Observation fe6319d5-aa72-41db-8405-77fa52d251b9 · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.175406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:54ec0d7f2ab4398b753b539d626d1b62d014f7c3ec2b596868aed357bcb00fdd

Observation 490c03fc-1224-4d82-beb0-35458f2dc982 · inbound

Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing cites this paper.

Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T08:55:44.783824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T08:55:44.783824Z digest=sha256:1e60cea97d0463dad9e9b41ae4b70e0aee7191a059ca3d6e4370ca34b18c258e