Pith. sign in

Paper Citation Record · LEDGER

Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images

As of 21 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2506.13458.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13458 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:04:05.713834Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4ea3692e-0764-437f-9a79-6498f175c86a · outbound

This paper cites URL: " 'urlintro :=.

Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:04:05.668809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:04:05.668809Z digest=sha256:29f39ad28cba9cb739c31f171c061ebdc69e7a59671f16a93efa2a3e4dab1a33

Observation 55db126f-9e24-4cb6-92e5-306255312b7e · outbound

This paper cites write newline.

Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:04:05.673224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:04:05.673224Z digest=sha256:b23002a505adc39f7bffcdc0cfe131eb43a6ff0591026c548910296e33072408

Observation d199879a-032a-45f2-aafe-80b19f53c96b · outbound

This paper cites LeGrad: An Explainability Method for Vision Transformers via Feature Formation Sensitivity.

Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images LeGrad: An Explainability Method for Vision Transformers via Feature Formation Sensitivity

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:04:05.677202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:04:05.677202Z digest=sha256:5366fc088a34ed03fddfb553d4f59d6b2b2732096991bc94dc50c8b4f9983913

Observation 184e942d-4966-4f05-bfab-98d4a650bf68 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:04:05.681327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:04:05.681327Z digest=sha256:fb6646d486c0d80b4aefccd25374f3ca5df20ed4087c9151cbf5bcd770d485ed

Observation 6b92f081-d958-4863-9a2e-34787bdd3fc5 · outbound

This paper cites Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift.

Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:04:05.685545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:04:05.685545Z digest=sha256:c0bf93c5510900c84f51fb6c9497e470517c3da11486b68fe23e8d663a6969d9

Observation 5cfeb47c-7ee2-4d6f-b074-895e571caa25 · outbound

This paper cites Lecun, L.

Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images Lecun, L

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:04:05.689524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:04:05.689524Z digest=sha256:6beec5729c7e67c5d855394dd7b51d811918587c1434bb9a2213d450d4d8929a

Observation aa9ed83c-c0d8-425c-ad00-a7411753c496 · outbound

This paper cites Bengio, and Geoffrey Hinton.

Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images Bengio, and Geoffrey Hinton

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:04:05.693378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:04:05.693378Z digest=sha256:d3f0536cdccfcaa773c6ab3cca1f020900a89f78632be14e64189bf62a5e6fb9

Observation 970cb47e-d881-45a3-a9ca-e2447d066451 · outbound

This paper cites Microsoft COCO: Common Objects in Context.

Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images Microsoft COCO: Common Objects in Context

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:04:05.697291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:04:05.697291Z digest=sha256:c67187953be2361e23ceb22772f0cd9ffa7fdc81fc8c71318d0574bbb5d8d6b2

Observation ffd638f4-86a0-4274-bf0b-b86480e7b0c2 · outbound

This paper cites An Introduction to Convolutional Neural Networks.

Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images An Introduction to Convolutional Neural Networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:04:05.701605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:04:05.701605Z digest=sha256:d19775420ff080044fa8d5753e76bd59cb1936992286ac53f9c0efa7619a15d9

Observation 193430aa-65fa-4c96-a984-eddc89017e3d · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:04:05.705684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:04:05.705684Z digest=sha256:7727402e291bb63cb231260efabd20fa87c7ef34596ebff711df6f244f1047d5

Observation 6c5f843e-6896-4597-b5bf-073bc34b7724 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images Learning Transferable Visual Models From Natural Language Supervision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:04:05.709879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:04:05.709879Z digest=sha256:fae53c28ccaff2488782de4ec66c5cab295a231c45d507fb0fba4585b47cd813

Observation 25c19370-cf0e-4e32-a69b-6306e4c2f82b · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:04:05.713834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:04:05.713834Z digest=sha256:b4dd2fcbb265e525fca08c2c220b19819e79fefaf9ebcc3f4be6bb3cd58276a7

Pith citing papers

No inbound Pith citation observations are available.