Pith. sign in

Paper Citation Record · LEDGER

HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:1906.03327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1906.03327 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:27:50.886671Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:39:29.697481Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aac3bf2f-26c9-4bde-ba18-30b1539858c3 · inbound

HumanNet: Scaling Human-centric Video Learning to One Million Hours cites this paper.

HumanNet: Scaling Human-centric Video Learning to One Million Hours HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:54.100081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:51:08.414394Z digest=sha256:6cf8a67ebf661c1cc8c3f1983c1793d7094c858c6b166a3dfb494990013bc863

Observation b83efce2-60f6-4720-99e4-c3695c01f03a · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.921649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:5451e93a9ab9294bdd96183ea4a67b5b9c5fdef2221cc4e6b8bdfe0c85048b0e

Observation 22c02f2b-bdb6-488b-9561-1ca4a1e8c67f · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:46:39.975072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T05:40:33.752341Z digest=sha256:061f8bef4d3070de57581a6dbbeeb24af51f77fb983ee9019e25524bc08d6a75

Observation 17e3486a-d4a0-40ef-b324-1f48e5ea6594 · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-15T11:06:23.089564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:06:23.089564Z digest=sha256:1bef88a5f47d14bbf95bfbbb21e3f9ac358d8f966986bd50c1f8f6cfdf9c7db8

Observation 6047af69-0c9f-4351-98f7-3c54a5e26d1f · inbound

The TIME Machine: On The Power of Motion for Efficient Perception cites this paper.

The TIME Machine: On The Power of Motion for Efficient Perception HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T13:27:50.886671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:27:50.886671Z digest=sha256:676993382549a1ab238f6657df08b454cfcf9ce4412e014366dff8ad343355f1

Observation 8487369e-e5c6-4ade-a681-4925339743b7 · inbound

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation cites this paper.

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.654352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T23:11:43.712391Z digest=sha256:0692e5ae187371497c011b4e4d1d1001c0ca52b82450d3f5e884162caf729d2f

Observation d80fbd2a-e196-4a80-aacb-ec0f662604df · inbound

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining cites this paper.

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:39:29.699812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T17:53:24.431287Z digest=sha256:d47b4525036152554dafb77a737fba08e0c2a12d4134ecb5c02a3906932e7cce

Observation 4d429d72-093f-45c4-b88b-ad0ad2c03cbb · inbound

TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living cites this paper.

TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 182

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:09:30.244239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T18:22:13.147215Z digest=sha256:1f039860618d256357e0b08bc722e34f388a25ed97abf2acc6050c0199e4cda4

Observation 1e0810f7-9469-4c59-894d-39418a7e82a0 · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:14.772882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:14.772882Z digest=sha256:c6dd7386cdf0b2d24dc811fe0dea676460c4c9c53b1c78a55110e1e9e3374232