Pith. sign in

Paper Citation Record · LEDGER

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders

As of 4 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2604.03919.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.03919 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T12:01:02.997589Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d280a2e6-894a-4d2c-be0b-c14343a7af07 · outbound

This paper cites an unresolved cited work.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:c7aaf93d506598014963358ad262d43f2a541ee5b76f591eddbefa77d3bec9ab

Observation 0717ce59-2fbc-408d-826d-fa861af72f8e · outbound

This paper cites an unresolved cited work.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:ec850305c56a5d743cfb19275f906b7e622781296eb3feb72f35f95ef3df21de

Observation 7b0861a5-8004-419a-84ed-ea0b201911b9 · outbound

This paper cites BatchTopK Sparse Autoencoders.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders BatchTopK Sparse Autoencoders

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:6ba3c0da97e888c434ee6f439ea822d04200bdac749d1fc754671ae3ca2e71be

Observation a74b171e-2ed0-48bc-81f5-983d5823f2fb · outbound

This paper cites an unresolved cited work.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:2c43f7ee0b5976adf92f583254c6619fb8b5ad6423a14f4321be4250ce04a107

Observation bd90e026-701d-4738-b113-791fbf3d894f · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:fcc413fbb9a2f593158d6ab9422246678f96b1c1771454f089337d5da5f5935d

Observation 763d40df-1cf3-4378-a6ec-20a869df75ca · outbound

This paper cites an unresolved cited work.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Unresolved cited work

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-13T12:09:35.531886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:e2003e1b2a86ea8cd92c42f20cb695f3a3810a1f0aeb13b4e41460175be6c60c

Observation f63a962c-1029-44d8-8605-c30e082eb22c · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Scaling and evaluating sparse autoencoders

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:d3c2be303d0c97123cde43ebbe569e6808d4e02874cd8bf053eb80fb9584d9b0

Observation 994d8252-2e83-4c78-809d-7b71f84a16e5 · outbound

This paper cites The "something something" video database for learning and evaluating visual common sense.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders The "something something" video database for learning and evaluating visual common sense

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:0523f403394b6f66504cb449b4e21f56eebe4d7238a1d51b7d87b3ae6476f4ee

Observation 9a224e0d-5c59-48c7-86dc-70addb5af670 · outbound

This paper cites Mechanistic interpretability for steering vision-language-action models.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Mechanistic interpretability for steering vision-language-action models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:544a978eaf54a85c3d7da0c25d7cad9b62b5d6b230cb878f40fe68dde9d2d9c2

Observation b4355137-defe-4ecc-9bf3-16a0739fbb7e · outbound

This paper cites Steering CLIP's vision transformer with sparse autoencoders.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Steering CLIP's vision transformer with sparse autoencoders

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:957a4a1ceb0e55fffd693dd0e14ef2391ced0b0de6e08256a10a590cb8301bbc

Observation ed65cc81-a023-4163-850d-ab7814252275 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders The Kinetics Human Action Video Dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:31bb87968b1ca8f4db93b932fe0b79621fce05b27de76a4035d094db15cf96e4

Observation c81d3e90-2b98-477a-ba9f-b8f7d4ee5b20 · outbound

This paper cites Anwar, and Manzoor A.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Anwar, and Manzoor A

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:9760f5a60ac0c64934097f2b240295ee3945dc4eb712c675890487a625621770

Observation a48979d6-f51a-4c45-a44c-d1053f86a272 · outbound

This paper cites InMechanistic Interpretability Workshop at NeurIPS 2025.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders InMechanistic Interpretability Workshop at NeurIPS 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:2d16398975848840f46239d3bd54572a9781c7df8cb054e6cbf7bed5900d983d

Observation 416f27a4-a02a-4eb7-85fe-7d38b18ba032 · outbound

This paper cites an unresolved cited work.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:34fef9ce95fbc1dfc1f0095ed5347df0748289cdb489bfaea92642ed398b8fbb

Observation 689e70e7-2ddf-4d43-a15d-8e35fe7a6822 · outbound

This paper cites an unresolved cited work.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:e6ac536d67aa64314dd76e2515059811c757b46a49429dc3b2b71bcd0a8e1326

Observation 847ef516-85e4-4054-8d8f-ce401133de67 · outbound

This paper cites an unresolved cited work.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:c553d4bfea7fbfaf82a4fb4b15259f63e1a3dab6d5e662180aee24d9d1f13c51

Observation 6186501e-3046-409c-964d-a05f1c8a5542 · outbound

This paper cites an unresolved cited work.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:43f7efe0d6bff71bbffa1962ca57ecdf2bb19a9b9bb1c4db248af57984a7b7b3

Observation ee29588d-131e-463f-a010-ebb95aa18f57 · outbound

This paper cites an unresolved cited work.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:1eb55e508e4e123740a952accecb4fb5f239a7cb679902077550aab87f49c207

Observation 2331d670-77a6-47b5-9b75-82dd1a6a3543 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Learning Transferable Visual Models From Natural Language Supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:fac26d67b9ff7d2f1be6c13d513a1adeae0a6c240e771ba0362ed234ef157de5

Observation c3294692-5b20-4f7c-96b4-2622d3a64e90 · outbound

This paper cites Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:b8e855052ed7434b12457d4ed9cb90ccc2211e42f585e82557461be850e5a578

Observation 90a1949a-ff3e-467e-aed1-07eda143c752 · outbound

This paper cites an unresolved cited work.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:61bf496ba2b42f983cf078f5ace144672a9473171117a42a646c98b16c11123b

Observation 079297de-ee13-462d-820d-68c9d5e5139a · outbound

This paper cites RAFT: Recurrent All-Pairs Field Transforms for Optical Flow.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders RAFT: Recurrent All-Pairs Field Transforms for Optical Flow

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:3dbd1f8aea5b2da1241dc3a667187da94aebd1361b65d8b4a0f8150c86f58e39

Observation d5c93816-aaf2-4008-9ab9-957d9d8b530f · outbound

This paper cites an unresolved cited work.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:54a6fcb235a47c581baa7caa34bcd5da813839bc2fa2c718736153a58b39d679

Observation a91e62cd-f7f9-41d5-8f25-666a78105900 · outbound

This paper cites an unresolved cited work.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:cc3a779bb0aa09bb6eff03d750bd5755d93d444291f05fa0d6f526d864baebf6

Observation 097ec171-d4f5-48ea-8624-e42e5edbdbd3 · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:d4cd9e8798ee35b449c8e6fee96604d4e35e004b1cfdf20ca38c6749e35916c4

Observation 36df1305-5592-4a67-919d-2b483ae9f7fe · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Representation Learning with Contrastive Predictive Coding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:a599ff715935fd92d60cd5b615438deb50902221f62f7b6f965333196ba6120a

Observation 3d79df1c-4921-4f42-8470-c46817602c7c · outbound

This paper cites STAA: Spatio-Temporal Attention Attribution for Real-Time Interpreting Transformer-based Video Models.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders STAA: Spatio-Temporal Attention Attribution for Real-Time Interpreting Transformer-based Video Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:74de1e228193c3c4451ea8816faa9f0089173c3db9362aac5ceae7adfb574775

Observation 967da691-e205-487a-bc84-2ca367fa5fe7 · outbound

This paper cites an unresolved cited work.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:d16ba9416e0fd61e45d147e9f7fc7bb98758bfb8e923b26ffd12ba903274420a

Observation c4c4a8f9-705f-4575-9ff3-3876af76bf02 · outbound

This paper cites an unresolved cited work.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:5787574a798d58600879e6270aff8945ab6b5091c018e46a98a247bb951ee083

Pith citing papers

No inbound Pith citation observations are available.