Pith. sign in

Paper Citation Record · LEDGER

HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:1906.03327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1906.03327 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:27:50.886671Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:39:29.697481Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aac3bf2f-26c9-4bde-ba18-30b1539858c3 · inbound

HumanNet: Scaling Human-centric Video Learning to One Million Hours cites this paper.

HumanNet: Scaling Human-centric Video Learning to One Million Hours HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:54.100081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:51:08.414394Z digest=sha256:c813072dee52e79c5de66fa561eb224f4e54b4b973643fd3413bf178054f5d5b

Observation b83efce2-60f6-4720-99e4-c3695c01f03a · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.921649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:a7107d769c27204b17cc2ca63312fce6c00d377073469e2fcccdd6ae75f06ccf

Observation 22c02f2b-bdb6-488b-9561-1ca4a1e8c67f · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:46:39.975072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T05:40:33.752341Z digest=sha256:9f1814bd3c7d421fead91a14fca73e42eb21730c4c1a6360cbf5291655f56f38

Observation 17e3486a-d4a0-40ef-b324-1f48e5ea6594 · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-15T11:06:23.089564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:06:23.089564Z digest=sha256:334493c63789d30ae05eaec3fd593e8d1c64c4c6861ef2081740027418a63626

Observation 6047af69-0c9f-4351-98f7-3c54a5e26d1f · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T13:27:50.886671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:27:50.886671Z digest=sha256:e63cd70ed53d7f279197ff309cd3303a36bbfb36d43a128828532c9ef771b183

Observation 8487369e-e5c6-4ade-a681-4925339743b7 · inbound

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation cites this paper.

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.654352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T23:11:43.712391Z digest=sha256:6bfea07430d264858d7b30ec972058c42c63f052d9c7c661cfa4157faca508ec

Observation d80fbd2a-e196-4a80-aacb-ec0f662604df · inbound

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining cites this paper.

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:39:29.699812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T17:53:24.431287Z digest=sha256:55ad3a19614852efe59639e1814f8779e6769ed72b1af728558005335cce1177

Observation 4d429d72-093f-45c4-b88b-ad0ad2c03cbb · inbound

TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living cites this paper.

TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 182

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:09:30.244239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:22:13.147215Z digest=sha256:0e2b1b59bc6aa1af5af8f10c27e6a5faad2fe1050edc80929a199e6dda6b0b69

Observation 1e0810f7-9469-4c59-894d-39418a7e82a0 · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:14.772882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:14.772882Z digest=sha256:899ad85c4c234daca9e025e1ddd1a9910bd1991f8f94e28d7f5494a0f5b3ae0e