Pith. sign in

Paper Citation Record · LEDGER

The "something something" video database for learning and evaluating visual common sense

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:1706.04261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1706.04261 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T17:16:52.452286Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T17:09:58.596248Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ee1876de-1579-4769-9d4e-a8cb6fecf1d4 · inbound

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents cites this paper.

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents The "something something" video database for learning and evaluating visual common sense

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T14:10:15.077625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T14:10:14.929207Z digest=sha256:c534666491513fa08d5ba963bc14feeb1293e1158c3a40fb6ae38af66bcd6b43

Observation 5fdac18d-6351-4019-a772-6945646dee21 · inbound

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models cites this paper.

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models The "something something" video database for learning and evaluating visual common sense

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T21:52:03.013988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T21:52:02.893886Z digest=sha256:d932fada8fb62c43c03877941662fcc34455e21e28f02376bda59be9bb5c79fd

Observation 7183458b-6492-4297-9937-aaba07f7b4e8 · inbound

Co-Evolving Latent Action World Models cites this paper.

Co-Evolving Latent Action World Models The "something something" video database for learning and evaluating visual common sense

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T02:52:21.741814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T02:51:21.569303Z digest=sha256:8a1565165ff7dc3cccae0a9112666e00fa2df0bbb562db77c946a2e5ed1e5204

Observation 994d8252-2e83-4c78-809d-7b71f84a16e5 · inbound

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders cites this paper.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders The "something something" video database for learning and evaluating visual common sense

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:0523f403394b6f66504cb449b4e21f56eebe4d7238a1d51b7d87b3ae6476f4ee

Observation 09f9666b-78fd-4c09-840e-50e786f135d5 · inbound

From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data cites this paper.

From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data The "something something" video database for learning and evaluating visual common sense

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:03:01.175828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T17:02:18.358675Z digest=sha256:967fa2a90d039dc3c76d86d109fc5cf3111aaa8caeab22e24b36c4c910204e1e

Observation 1ee09952-dde3-45bb-9494-20e2eeb9c0ff · inbound

HumanNet: Scaling Human-centric Video Learning to One Million Hours cites this paper.

HumanNet: Scaling Human-centric Video Learning to One Million Hours The "something something" video database for learning and evaluating visual common sense

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:54.198343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T00:51:08.414394Z digest=sha256:05b440b808143cf54c5c690d1ccf0504bb6a3db3e25f457d8402eaf3f2ef324e

Observation 4a2e9053-f973-4bd5-92a1-d09c21bfe8f3 · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI The "something something" video database for learning and evaluating visual common sense

Reference 176

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.947674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:6580bca03b8376e3324218be5986f7cf7299ad0beb1506ecf4de1c36085c0bba

Observation 934f1c32-356b-44da-9f54-8e717575702e · inbound

DiLA: Disentangled Latent Action World Models cites this paper.

DiLA: Disentangled Latent Action World Models The "something something" video database for learning and evaluating visual common sense

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:38:56.295557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T19:35:37.527479Z digest=sha256:c74f20492b81af12747deb3e9ca6c5f273db1e745663e0092ac43d88807bb237

Observation d794f375-357b-4a20-b4b0-ab59df96ca84 · inbound

Structure Abstraction and Generalization in a Hippocampal-Entorhinal Inspired World Model cites this paper.

Structure Abstraction and Generalization in a Hippocampal-Entorhinal Inspired World Model The "something something" video database for learning and evaluating visual common sense

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T19:37:43.759875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T19:36:17.173279Z digest=sha256:903f36b1e22b787c4cf541657c3193a4ab0c8726ef840ea837189431cb561566

Observation f0b5406e-cecb-4ae6-b460-96405cd53248 · inbound

PEIRA: Learning Predictive Encoders through Inter-View Regressor Alignment cites this paper.

PEIRA: Learning Predictive Encoders through Inter-View Regressor Alignment The "something something" video database for learning and evaluating visual common sense

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:43:19.512905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T13:41:52.182460Z digest=sha256:febb67fc83affa2bb4b5413eb63831acbff66edeaf34506be891039ed233374a

Observation 1173a47c-6d92-4b71-9b1e-f9cfbef8840d · inbound

MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning cites this paper.

MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning The "something something" video database for learning and evaluating visual common sense

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:09:58.598027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T23:57:19.880950Z digest=sha256:07c812d07e99889bea861aaa546b3469c73ba246247e560159ca247845a49378

Observation 02496fb4-be63-4472-adb3-5ee9d78284ca · inbound

iFLYTEK-Embodied-Omni Technical Report cites this paper.

iFLYTEK-Embodied-Omni Technical Report The "something something" video database for learning and evaluating visual common sense

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T12:21:16.963000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:21:16.963000Z digest=sha256:6dad3ccb16c44f6a452d1052f32e9bad39e6c7117a0eda6f709959de7835bebf

Observation 037bd44e-3e4d-4bf9-9449-51853ebdf1f8 · inbound

Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution cites this paper.

Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution The "something something" video database for learning and evaluating visual common sense

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T17:16:52.452286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:16:52.452286Z digest=sha256:7d1c59560eb91db59435a59e8e15df1dae643d0d38c03b5c187d463a9e43cc36