Pith. sign in

Paper Citation Record · LEDGER

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP

As of 15 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 0 inbound Pith citation observations for arXiv:2412.09895.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09895 v2

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:41:35.796320Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

81 of 81 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved62
  • parse uncertain5
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1e4052c-37f7-49c0-8ff1-b206c51fa4a5 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:37.008579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.460295Z digest=sha256:be0ad73db4ccc1ab518505e33c9b2729fd403bfdf26f1949fa53641a0a4ea670

Observation 8d83d00c-8508-4542-be35-1b33c1091153 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.993580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.464893Z digest=sha256:9e4868d72d66861c08277d06783b584094d2608abb2a38200ee5d51e4afc882e

Observation cd1035d4-28eb-48c7-9093-06919a7e5916 · outbound

This paper cites A Short Note about Kinetics-600.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP A Short Note about Kinetics-600

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.425233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.425233Z digest=sha256:6574618db7a6a5baf50880c65cc42929fe9eb59ce52a806870d1f57d1c2381c4

Observation b5628ccf-e913-441f-bc6f-25fd9d9b3ff3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.962123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.474613Z digest=sha256:30e9802d8b264d521f4f51363d3967136e73167489cfa3bf3cfd56dc2a5cd3f7

Observation 0748889f-d01e-4c92-8b6d-569f533d136c · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.932708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.479269Z digest=sha256:ed22620c19b9edc961174d1b465e3b73856df8b5e92539dc8cd0c3539ae2b597

Observation cece37e7-1654-4128-9149-d896670024fe · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.918232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.484238Z digest=sha256:8ee15d29463b774128396d49a712020fad4fd6c9f2796719b7b8e9cb47904110

Observation 6e3c96fc-4ad9-4dd4-895c-14abf033f7c0 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.444683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.444683Z digest=sha256:c3ff7a4a03d62237b97f1b8fa616f9c6f223e646bd5688bba394d74485c5c719

Observation abcbb0da-2040-4bd1-90e5-bde471121f85 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP ActionCLIP: A New Paradigm for Video Action Recognition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.449384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.449384Z digest=sha256:d48d51d2dfbd5c74b4b7f3ddf4225413c707b50934680ef30c30ff471dd96705

Observation 7d050f96-786a-4bde-8973-7ca2b622a978 · outbound

This paper cites AIM: Adapting Image Models for Efficient Video Action Recognition.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP AIM: Adapting Image Models for Efficient Video Action Recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.455213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.455213Z digest=sha256:4a7d77930ea58aeb26feac8fff86e871d05345b4c99d56b1c4d2e8ba224e2d17

Observation 2d5d870c-9dd9-4c77-96f1-7ef6de7e99a9 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.977628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.469970Z digest=sha256:bc508dbe9048fdb46eafcffe515ef4d873c017f17832a5b4439f12b8679f970c

Observation 24df1153-402e-4455-a5ab-8c5887c7c98c · outbound

This paper cites Sub-action list:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action list:

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.903915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.489068Z digest=sha256:9ba8d8c8ad36180a342f8a79cb70f3936cec86cb580077a75e618ac4166b023d

Observation 2c3238b8-7ef8-4f67-b250-f54f529a39ba · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 17

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:41:36.888249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.493873Z digest=sha256:58da3b8badf237f1459bc9c49c9973f8d965dd9e4f53a1bcc346b3398a32d34f

Observation 658f3d45-73e9-426b-a703-3bca23e6cac8 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 18

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:41:36.873477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.498702Z digest=sha256:db011656771246331b8fbed266c1add77b73a9c80966aa1f68e65b973f3d16e6

Observation c29aa97d-2e3b-4258-b881-a5d3b385ac85 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.857743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.504067Z digest=sha256:19cafd3f692c98ecab9e97c33bb277a28e86c180847d18b7a77acbf237cc0cc4

Observation 000d4a80-9583-490e-b2a8-012a0159c461 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.842178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.509306Z digest=sha256:70e2d5b7556ec7551a227c744f4033db62d00d8fa40218e616acabacbd6cc85c

Observation d0408fee-b2f2-463d-b198-de4a8d6559a5 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.827080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.514358Z digest=sha256:c69936f134d7fb8efc4575d8bbcbe560d2b34b2993161ebdc80443a050985134

Observation 4f86c36b-5afa-45ce-a71d-9884aa93d8f2 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.812538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.519169Z digest=sha256:de9a1ead6d37f91de96fd0f07157696957973f278880fa71eddaf65370f6657c

Observation 2e37585c-0929-45ce-8feb-356115eb4bfc · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.796896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.524367Z digest=sha256:6bcf7acd5fc8c7b9946ed0eb9b81efd35e2d5535a9e8b5da18720e3b37e480b5

Observation 472d5c52-89ce-4ddc-8f95-977771c9997c · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.781860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.529457Z digest=sha256:963b8b59f0b60974510f0b7eebfdaf41d3f19409e4b4eb415307bcac6610b91b

Observation 907bf1e0-99c7-4291-a594-569dad3c5d78 · outbound

This paper cites Sub-action relation triples for temporal text prompts [T]:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action relation triples for temporal text prompts [T]:

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.765399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.534387Z digest=sha256:441bdb5a1d8a2204b95949be0c2258b58828fc6ee49f429751c3ddc1f7fe8cf0

Observation d464e9af-b9d3-41a4-bf4c-b103b0615b69 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.750950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.538823Z digest=sha256:e9a796a4fb425643b78c3d0d8bf02d5c940bdcb587b90a2c236a359a23bf61c9

Observation 2905054b-43ef-4078-b1a9-f87f9efb41c3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.737523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.543406Z digest=sha256:abf61e1b34fe765eecf8991a84300c7828bf818c36913333a363297200e69d4d

Observation d43fb80d-a3ac-4d7d-b238-37da53c63b3a · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.723612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.548859Z digest=sha256:76e872c0cc54efb23a764c14f5858862009b2f36100291dd217cad61a80472a4

Observation 520871b4-8713-4c6b-97aa-4e9dd1b395b0 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.708199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.554641Z digest=sha256:da55db4eef8b82218a93eb46a28708af5e1917157dbe55681a8b573d0a151f99

Observation c24c487f-8871-495a-9cf6-939cce21ac34 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.689492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.560349Z digest=sha256:29e20e5f9e492f71ac5e2082bf0683fb7b44792a1f6db1bb440407e6d179960b

Observation 34bc97d2-29ae-42d0-bb6c-6090aaa87c95 · outbound

This paper cites Responses for action category “Surfing”: Object list for spatial text prompts [S]:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Responses for action category “Surfing”: Object list for spatial text prompts [S]:

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.675106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.565565Z digest=sha256:5333e680664a0041e89ce799922cf3a7272ce60749115e86b414e35d54b2cbbb

Observation 3508cd0c-47ef-4aea-890d-89a117075176 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.660917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.570823Z digest=sha256:f227870b98f3e03cc16d069752c01f3bc3dcb3b0f2b1395b87de5b176f0f07f6

Observation 4b833654-c6df-455d-8124-7cb60090ae95 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.646147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.574896Z digest=sha256:ca0c26753751c4e903e8b5bd5f7e6df11f74e1a50bc966627e338464af7295c3

Observation c76308dd-8a8f-4355-9f6b-9d027017bb75 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.631409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.579164Z digest=sha256:8ad1524fb6f2a284ecd0da993fc839651902c3e7d7b1183d7a75cb5f494f06df

Observation ed8cb8bb-fcfa-4798-a8b5-2c31f6b40fc6 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.616920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.583420Z digest=sha256:87b93e5b9fd22708ff93eee02674c3599af6d97ca4904ec3813ae1d634889b5b

Observation b407b063-6e62-4508-bfe9-6f13d09b2999 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.600275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.587957Z digest=sha256:8b5e58a5b16e9b90c3bdb9a44b5e87abb5eb42988baf84208f0f2df5f52366cd

Observation 0ff8ec5c-6c18-4ec7-9a34-f818a585fc4d · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.585007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.592642Z digest=sha256:a8a4d6f3a22c5e3059ac0db324429fbd4b36b56f385e64234153c6aa79a1b070

Observation 3310ed52-22ca-41e7-9fd2-dafb21b4346c · outbound

This paper cites Sub-action list:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action list:

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.568910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.596857Z digest=sha256:8151a8984b6ae00bc08e9e01ac7c855a24a300156caabaa1dbf1237550f213a8

Observation 1e8ce5c8-232c-4fa1-a67f-f0418efaf330 · outbound

This paper cites given action name.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP given action name

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.554491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.602066Z digest=sha256:321f40cafd4aa1fbbc69f5a94ab29bbd1b77d0a9f18a829563a71bfac4163877

Observation 715ba0c4-9293-4c8c-9823-1f53f9e321c3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.539984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.607394Z digest=sha256:25cfcf3ac4f490afbe6155737c67926ebccf59721e06ad923a45ca994e9c2598

Observation cb3a4fe5-4b16-4954-8e1f-3e346d6c43fa · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.525766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.612239Z digest=sha256:b176740d92e222e229d9380203a8642ca5c8f928c0af1bf78315eb91228e6c5a

Observation e6896d1f-d830-4e01-bb3b-e727d0669cdb · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.512304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.617075Z digest=sha256:f93998a634ce501ee7de2abbea71fd7639794a4a05c9f861d710063218532295

Observation 4604b74a-3dc0-4a02-810e-741808f2f607 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.498650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.621966Z digest=sha256:b942ef48ce5aeb296f949d495182b1b7417a438d7115c70bf6a85969aaa80a63

Observation 7eef8db0-8f5c-4155-beaa-c48496b87f3f · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.483352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.627226Z digest=sha256:b65796262a854c02f7081f035bbc467521f13666b9626dc00312c883415bdd4e

Observation c327faa7-29f0-4752-a44f-e59524ef7e60 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.468702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.632248Z digest=sha256:61d5781d8c2c0ca91f8ef0899c634fb42aa90c3a791371ce6b5978972f553c38

Observation e4b7d01f-ae90-45de-903b-3630190e1f80 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.453446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.636966Z digest=sha256:4b25547f56161583092378b11b75bbf8b6eb13ba445035675af5756e3dfea82f

Observation 72cbe206-ccde-4d35-9645-5921daf7cccf · outbound

This paper cites Sub-action relation triples for temporal text prompts [T]:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action relation triples for temporal text prompts [T]:

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.438626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.641426Z digest=sha256:5f4d0f5f779e316f0daeb6141a811a81052d2d77a308f02b3c8fea39954abe5e

Observation c69e605d-6355-4977-a6d3-7ab9f20977a8 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.423754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.645907Z digest=sha256:adf5246aa48ce39f3213bb9798a7be301fc7d478c14f28581a358744bbc633c9

Observation ace0df3f-cfcd-4585-9a75-981c6360c71f · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.408609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.650859Z digest=sha256:605ac815ec53ca90b459354d42635f866a67c6df8a43b1f308b0d0f2540e0583

Observation 63a5e824-3a9a-4b6e-bfef-1b7d990d91a9 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.394323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.655428Z digest=sha256:a688278d1a7df5ade0a33b2bac767dbddc1cd3b791d843273792737ed0bd2a53

Observation 054c740e-499b-4282-a1e6-3e626481ffe5 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.379530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.659914Z digest=sha256:45e9d23e8db46a838c52f7b97919b7c9230c5585f898d6f986ae2c7dc0e0bff9

Observation 1be595b1-c095-4013-8b28-87e8dce80b7b · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.365014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.664511Z digest=sha256:1da56f1abe3ec95f2cdbefdadf236547117d8c020f1fea94f27e72a9cb94fae4

Observation e5002ea4-d56b-4fc6-8726-2f32aae58328 · outbound

This paper cites Clean and jerk.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Clean and jerk

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.351397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.668719Z digest=sha256:680fa3628cc06d35bcaa4ca2f03e7b31fd5b2a1891380a21dceb1ef85f75cf2f

Observation 92631da5-4954-48f6-ac02-b668ce0ea1e9 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.337236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.673142Z digest=sha256:5b96fdf7a79e6f8a0de1f70185c0deffedec4ce2246a3b63060fc30830d11f39

Observation 2b06b85c-74b6-4174-90fb-c06b0342bb72 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.322517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.677555Z digest=sha256:1359d8da87c96178b16a1f21c72620786924040f7773a3a796663edfa281057c

Observation 76a85673-baba-4a35-ae23-0d898b89b1ed · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.307906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.682303Z digest=sha256:d06df076af28e7de04018766db87338b88ff86fc87a2fccc35ac92391bf97265

Observation d26f89f6-a4c9-4616-820a-bb6ed704e2b6 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.292915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.687056Z digest=sha256:975fd386fb0bb182bbf674ee2935227aa08ff14dedb178734ac0a5ac01af6e72

Observation 9bbbd63d-cf7b-4899-92c6-cc284b8693bb · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.278314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.692402Z digest=sha256:1f3f56dd2bad9a849571067aa3001093a290df0d226790f2f17f7e8b9d989bb0

Observation 97282e13-6b49-4ac9-890d-c36f39908512 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.264381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.696907Z digest=sha256:a691511b6862ea5ad0efff36bbdb770b5b28054fe15baac5e9140a4f5cb8233a

Observation aa8e3779-b8cb-4e43-9620-c1c47bef0f37 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.249948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.702217Z digest=sha256:bd7c767939df48f8f3cda4212cc3b9ba2e87f92db6a01ca446443b332068cf45

Observation c19b2ae5-1b73-46c3-b185-870babd944b3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.236238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.706824Z digest=sha256:3797929c2dd015f1fa370fcbd2a017f844dca94a097f1ec7fa9c45ce04e4ff24

Observation 6bccac1b-34ee-4e92-bd23-0208af35b7d6 · outbound

This paper cites Sub-action list:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action list:

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.221966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.711922Z digest=sha256:de7c3a161f448552c1babf6cbe714622569746292da9e0ecb9526b3ab4c17c54

Observation 2ac87c45-565a-470e-bc54-2ec053dfc9bf · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 63

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:41:36.207027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.716236Z digest=sha256:528745ddd01d18a51bb4d240d8309a843a1bb55c0fab2e8865e843321434dc90

Observation 9dbecae1-f80c-41ec-8178-5857411271f3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 64

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:41:36.191558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.720612Z digest=sha256:6a18de54b46863e22f1d7b1930fab0539b3670d21be54c0cfbac047e8169aa54

Observation 88f8cd3e-9d23-44be-b1b4-61b21f20c193 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 65

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:41:36.177192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.724883Z digest=sha256:c2f6f660e1b7f545b20d98382a52d053a52c13c3799071156377b55ec28958e6

Observation 994ac44b-c81e-47e8-8238-0f61e20ab453 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.164074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.729353Z digest=sha256:1d110077cf45a75d4e25cc248aef17a78c842e647bc6effb16157861b912554d

Observation 5383e4d2-5546-4226-bd41-b6efc58cadb6 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.150134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.734000Z digest=sha256:707d6df1021eb9e89256c624ec8f495b8a1e540dc7dd76f6b3b9938956ff1814

Observation 56de4781-f5b4-4ce5-9357-2f6d7cd7ac92 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.136455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.740225Z digest=sha256:791bbd9054ad0a3f0593aeaa06d34773dd096a8a3445cf56a398db752cc841e0

Observation a976d8ea-9993-4140-93ab-85344d7a0622 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.121130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.744329Z digest=sha256:b882988bb335b42669f40dc80160b539a56a73e57bc4420e381a20e559e9fba0

Observation 8ff0ee90-0967-4aa0-a5bb-ad98b5030455 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.106517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.748776Z digest=sha256:1744ad9c9249e0a689407c88a7b98b635b07eb00f4c1aa62778a1823f1d38148

Observation 600e6fb0-18a8-43d1-a4f0-ff57bbd78436 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.090831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.753751Z digest=sha256:58061db9d18d515173caef25d1411ea1f4c1c374254adbb1299b42ed2ececab1

Observation 1947000e-f008-4154-9278-f7db4966acf2 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.074870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.757928Z digest=sha256:848f7538dada48e5381da89e6df7c2efc8f6c6061fabf8936a1e6530241803d8

Observation d4832963-bdfc-4ba9-a2ca-c9703a2967af · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.058938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.761853Z digest=sha256:f82bc3d8d63b71a8b4f41ced430c9a0e5e54255721224d0fd7dff2b620ef01e5

Observation 0b614176-2cc3-4a8a-86d9-09896f602c06 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.043397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.765894Z digest=sha256:7bb11da6d9ad6e284c23d427b5429439a7855719046e7588d8319a0ad9a3b90d

Observation 0f47195c-4b35-4bdf-b4be-2eeb045cc8ea · outbound

This paper cites Sub-action relation triples for temporal text prompts [T]:.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Sub-action relation triples for temporal text prompts [T]:

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:36.027022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.769936Z digest=sha256:431d31de3b89868b09110f0fd98438f7b080e119e6da757a3c9eacda30d62888

Observation 62da1bdf-4bef-463b-9509-9bc11828d69d · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:36.009686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.774046Z digest=sha256:eb7f220c93f935180c062b30f3a32e3e133b80048b830cb8e2a0b01e5332b047

Observation 918abf29-90e5-486d-a3d9-c4483d0ac1c6 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:35.994903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.777983Z digest=sha256:d34b2787dbe9692902fee5b7cd8e652ddb6a53877c03d0e3b43b6e744aa7843e

Observation 16f78237-19b5-48a5-8385-a8228b15cbf3 · outbound

This paper cites an unresolved cited work.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:41:35.980916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.782355Z digest=sha256:e0115ed75d639293a4eba1d1940a7c435d3cc54abf304d04f2afc6b30e231d11

Observation 2927d39a-e0fd-471f-83c9-e617c65621fd · outbound

This paper cites B Details of Datasets and Evaluation Protocols Datasets We conduct the training process on Kinetics-400 (Kay et al.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP B Details of Datasets and Evaluation Protocols Datasets We conduct the training process on Kinetics-400 (Kay et al

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:35.965754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.786825Z digest=sha256:79c0c921b84487138a909f3e5a31d8e8a91e6a5508b8a62f020f987e2561619f

Observation 20acf465-aae6-4172-8886-cee4e5fbbbc9 · outbound

This paper cites jump”, “kiss.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP jump”, “kiss

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:35.950023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.791432Z digest=sha256:19bb027b32474977821b3037250e90f66554ce369b4d01267380c95bfd424a1b

Observation 3af675e4-82c9-4254-8a9b-b2d059eaab1d · outbound

This paper cites This can be explained by the fact that larger temporal scales result in sparser interactions for boundary frames during channel mixing.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP This can be explained by the fact that larger temporal scales result in sparser interactions for boundary frames during channel mixing

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:35.932844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.796320Z digest=sha256:c1f995de01862bcb85f4f8d9d6a8ed6c8eadc2abfb2a788800e19fe0316b1a0a

Observation 8cae5fb0-5c0b-4c5a-8374-82e7c0d27291 · outbound

This paper cites In 2011 International conference on computer vision, 2556–.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP In 2011 International conference on computer vision, 2556–

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:37.042145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.435691Z digest=sha256:cfb5acbae30c5a0903b7790977119759d87919a42456939b683172a953c01bf4

Observation 3649627a-c22f-4a9b-9bf1-6911b23e172c · outbound

This paper cites The Kinetics Human Action Video Dataset.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP The Kinetics Human Action Video Dataset

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.431226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.431226Z digest=sha256:67014f5c10cc83153d054c3d96df45935e6c5d35a279a117959996b915b7bfff

Observation f095a1d2-2ee5-49c7-926c-d9e245a1484c · outbound

This paper cites InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 4613–4623.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 4613–4623

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.420099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.420099Z digest=sha256:2f05444066ff006885cafcce473cd6a88131cea081c8ffd1184635bcb3c0b9f9

Observation 37347fda-c884-4304-88e1-dc77d1cbd936 · outbound

This paper cites GPT-4 Technical Report.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP GPT-4 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T16:41:35.414650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:41:35.414650Z digest=sha256:e5320efa41ac35fc686dec6e56dd2010e00fe254eba03c81c861f477c4e0a42f

Observation 55190b9c-e992-4a7c-9057-f09fae1fba50 · outbound

This paper cites Lee, D.; Lee, J.; and Choi, J.

Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP Lee, D.; Lee, J.; and Choi, J

Reference 2563

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:41:37.025080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:41:35.439885Z digest=sha256:197b90b19a83eb52c82805e11a3a087c060ddd51e61882b713fe20488e48b2c1

Pith citing papers

No inbound Pith citation observations are available.