Pith. sign in

Paper Citation Record · LEDGER

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos

As of 13 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2411.15628.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15628 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:08:42.308404Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy54
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bc9bc2a2-bd41-45fd-a3e5-6572c2935db5 · outbound

This paper cites GPT-4 Technical Report.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:40.821947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:40.821947Z digest=sha256:2ec533f5eac2d436ce351a930043bfe95a7b1006630ce3e612e42e30c74169cc

Observation eb663771-06c0-4920-b6a0-434d580c9460 · outbound

This paper cites Ht-step: Aligning instructional articles with how-to videos.Advances in Neural Information Processing Systems, 36, 2024.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Ht-step: Aligning instructional articles with how-to videos.Advances in Neural Information Processing Systems, 36, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.540893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:40.888612Z digest=sha256:a8a7d36f8c46979715d6db9f66ee46e25671d191054095182388ba8d05d7cc5d

Observation 96065e79-07ec-478b-a35d-a020b140eec8 · outbound

This paper cites Exploring synonyms as context in zero-shot action recognition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Exploring synonyms as context in zero-shot action recognition

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.472400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:40.995780Z digest=sha256:68409aec41c01be5c5ce1cd354276000bc0b0a5601c62ec96b24bff28e154a81

Observation 84fa500b-d883-473c-b707-e4b8e2c44fa2 · outbound

This paper cites Hiervl: Learning hierarchical video- language embeddings.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Hiervl: Learning hierarchical video- language embeddings

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.425059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.016803Z digest=sha256:c1b90f05dddd76c53196bf7c882c9a9b31f294bf6ab223bf420c555e63892033

Observation 732405b6-5ed4-4b7a-9f92-fae80af0857d · outbound

This paper cites The ikea asm dataset: Understanding people assem- bling furniture through actions, objects and pose.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos The ikea asm dataset: Understanding people assem- bling furniture through actions, objects and pose

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.409979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.022207Z digest=sha256:c0a94361415b719951f8fd67130fad19d52efc9544e56a23502ea7a6d80895ae

Observation e7551e14-bab1-4879-b945-d862870784d2 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, volume 2, page 4, 2021.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Is space-time attention all you need for video understanding? In ICML, volume 2, page 4, 2021

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.396127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.027101Z digest=sha256:71b1abbecb1c2a3d53fcaeb8cfeb977c54d09ee4e6bc704fa09720f21c60980b

Observation 80e0cd83-12ec-4d51-a6d1-3e5312a09955 · outbound

This paper cites Rethinking zero-shot video classification: End-to-end training for realistic applications.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Rethinking zero-shot video classification: End-to-end training for realistic applications

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.381534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.083324Z digest=sha256:049fb4729c5447b9c810048c76defb19030851fe5eeec7cb35e8a89561c7fe8f

Observation ff4d3c15-0b5a-44a9-8445-fed95ac1fa51 · outbound

This paper cites Regen: A good generative zero-shot video classifier should be rewarded.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Regen: A good generative zero-shot video classifier should be rewarded

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.299158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.090381Z digest=sha256:249a2116a92f440da4f8c0f5b2b4e29e839f4d4360e791dd7913eac129441a1f

Observation 9a249bd3-7f50-4a03-b916-3db3601cda56 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Quo vadis, action recognition? a new model and the kinetics dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.285381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.094991Z digest=sha256:fadbd352a3d256def22a29a814343973994f03bae111d3bdc93befe0eed39d77

Observation 74109dde-6953-4ac5-bf01-dbd2129eacc2 · outbound

This paper cites Elaborative rehearsal for zero-shot action recognition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Elaborative rehearsal for zero-shot action recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.268691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.101071Z digest=sha256:302ebfddb80c33fc3fe7b017b25c87100b24084f229a5d0ca8c800923ba972cb

Observation b606cf03-a89a-4827-a2cf-a92ae546cf73 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:41.148077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:41.148077Z digest=sha256:273a7b8d43ca89c098b80d8237a7e18a32597991a714cb809222e70d991bb4c0

Observation 11326423-3c91-49f8-92de-fb0fedaa099c · outbound

This paper cites Teaching structured vision & language concepts to vision & language models.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Teaching structured vision & language concepts to vision & language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.157402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.215173Z digest=sha256:66566bd7ac56df1bb8dd5f62849675b3241144aee5bc3e453e4256da9b5ccad2

Observation a75dd326-4fb5-44bc-8847-7ade7db59dda · outbound

This paper cites Step- former: Self-supervised step discovery and localization in instructional videos.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Step- former: Self-supervised step discovery and localization in instructional videos

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.142507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.240189Z digest=sha256:7253f0568ab9b0624f5267b030c4a7acb4a49539306bb2c6326597c31e0996fa

Observation 2fec3f59-f902-4123-84b1-945a3c7abf3e · outbound

This paper cites Zero-shot action recognition in videos: A survey.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Zero-shot action recognition in videos: A survey

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:45.071841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.244645Z digest=sha256:e7b0be1bb69f933216357396eb684b7e5f01890d48ab70b0fe879b1e8a1b84f8

Observation 7d0628ee-9900-45d8-868a-4c71a8f9d486 · outbound

This paper cites Ms-tcn: Multi-stage tem- poral convolutional network for action segmentation.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Ms-tcn: Multi-stage tem- poral convolutional network for action segmentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.968430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.329232Z digest=sha256:7accdc285a3e004672497e4e1fce0bb34da2f90e53787e0a272b3cfbcda61998

Observation bea9a6b6-a870-40a8-bd16-a779afd8df73 · outbound

This paper cites Learning to recognize objects in egocentric activities.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning to recognize objects in egocentric activities

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.953527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.421730Z digest=sha256:27500206e5e50d8f95ea0fd5d1cd7d84d3b67c2b69ebb57467901ada407c80bb

Observation d9e4b6be-e75c-4f34-9bd5-bfb91ebfc51d · outbound

This paper cites Slowfast networks for video recognition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Slowfast networks for video recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.936844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.502547Z digest=sha256:d9e7e638b33f99a1e74e537f117061d55fa76e5093331a0b67ace033c3387d93

Observation 9e8b1f8f-6ab6-4707-8485-c0dbb80803cd · outbound

This paper cites Prego: online mistake detection in procedural ego- centric videos.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Prego: online mistake detection in procedural ego- centric videos

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.921960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.535699Z digest=sha256:4644cee3cf4480dd9aed4069ae7a3b7fb6aee68fc87c840b10b780d501ade67e

Observation 1d271bf2-ada1-4cb9-81a9-db6c660651a9 · outbound

This paper cites Improving zero-shot gen- eralization and robustness of multi-modal models.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Improving zero-shot gen- eralization and robustness of multi-modal models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.821681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.558624Z digest=sha256:0bfe472c44f9f248942d10bdb51cbd168a9e3c20c8e258614045df95b4613f68

Observation 0d3393ef-0247-49d9-9225-eb6b77159975 · outbound

This paper cites Weakly-supervised action segmentation and unseen error detection in anomalous instructional videos.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Weakly-supervised action segmentation and unseen error detection in anomalous instructional videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.769725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.563967Z digest=sha256:471795174a7ceca2c2400270a4473681ea677613d0596bcff4f5d872c64aa343

Observation 7263b4e0-2589-4854-b751-0c54a2799b8f · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Ego4d: Around the world in 3,000 hours of egocentric video

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.754663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.569698Z digest=sha256:875eb07e3e71d8f78dd0fcc3cf19079c26b1283978709aaca496ebf9d774e8e2

Observation 90c21025-a295-4590-91b0-ee49897d6703 · outbound

This paper cites Temporal alignment networks for long-term video.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Temporal alignment networks for long-term video

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:41.620096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:41.620096Z digest=sha256:0295ed032f3e790cfc6ec3062013e00205036232729862ef02bb354fde231a76

Observation fa393e83-200d-46a6-825e-39e6971c9c23 · outbound

This paper cites Probing Image-Language Transformers for Verb Understanding.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Probing Image-Language Transformers for Verb Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:41.680083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:41.680083Z digest=sha256:36e863bd9cb37f26aa5a6292612929cab21faaa6b185f86d6c25e6f78ffae5cf

Observation ac89540c-787a-4938-a1f6-9275e2aa9601 · outbound

This paper cites Clover: Towards a unified video-language alignment and fusion model.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Clover: Towards a unified video-language alignment and fusion model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.722893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.709050Z digest=sha256:347c215287d77b04af833874f0032f6f2fcdccb9f29a7b928b4be80b0cb4060a

Observation be5709eb-168f-4b2c-940d-cbf8a686689b · outbound

This paper cites Fine-grained generalized zero-shot learning via dense attribute-based attention.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Fine-grained generalized zero-shot learning via dense attribute-based attention

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.639994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.714001Z digest=sha256:e5e90f06fecf0fb8c76d032decfc742019ac86b2a069813e6154ec78981602a4

Observation 74b88be9-9871-4bae-b5b4-d24e6670ea2b · outbound

This paper cites Objects2action: Classifying and localiz- ing actions without any video example.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Objects2action: Classifying and localiz- ing actions without any video example

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.624457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.718983Z digest=sha256:96ad057388d5498a33257433e95c3898b9c430e6a8fb57f47dd08c00b880bab5

Observation e9921a13-86a5-40c0-af2e-6a837086a824 · outbound

This paper cites Prompting visual-language models for efficient video understanding.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Prompting visual-language models for efficient video understanding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.606588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.723234Z digest=sha256:a2c04e903db843e947c8ee41933dcd81bd9bf1a81be785be121e570636bd8140

Observation df167b3c-640b-476e-92fe-196eaf923582 · outbound

This paper cites Error detection in egocentric procedural task videos.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Error detection in egocentric procedural task videos

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.535690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.783002Z digest=sha256:2109f7d7459792d990e23adbda30661bca702b606ab701dda84ef60b1150843e

Observation e749117b-00d5-444c-995e-e6de40abf633 · outbound

This paper cites Align and prompt: Video-and-language pre-training with entity prompts.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Align and prompt: Video-and-language pre-training with entity prompts

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.468170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.825070Z digest=sha256:5e2af97d736fd12f3bfa6cb7df09bbd2ded61839e50aecc1f142626e0a3b10ca

Observation 2a32aef5-e3fe-42f0-bfeb-c374ee904f9f · outbound

This paper cites Cross-modal representation learning for zero- shot action recognition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Cross-modal representation learning for zero- shot action recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.453055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.829748Z digest=sha256:eb68bf4f484b0a291e417978096d0a244a4f3b4cdf193ea7e1f811217e87fbc0

Observation 8b00805e-9d19-4f0d-93f4-716827babb06 · outbound

This paper cites Match, expand and im- prove: Unsupervised finetuning for zero-shot action recog- nition with language knowledge.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Match, expand and im- prove: Unsupervised finetuning for zero-shot action recog- nition with language knowledge

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.354680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.834114Z digest=sha256:e8610c2d9eb3341d48c85b58a070edce6421a23dceb2c38c90ca597138a56114

Observation bb8f3bb8-0369-4a79-ba98-800b5391f2a3 · outbound

This paper cites Learning to recognize procedural activities with distant supervision.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning to recognize procedural activities with distant supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:41.838632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:41.838632Z digest=sha256:7201ec44bf6fde75fccbd6379cc6404120ca84450b10df604320ee74e4bd2a77

Observation 56a4e60a-3ba5-486c-b277-0223da873a3d · outbound

This paper cites Out-of-distribution detection for gener- alized zero-shot action recognition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Out-of-distribution detection for gener- alized zero-shot action recognition

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.260818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.917429Z digest=sha256:6ef2b0045a0e3588aa8c2087d999e5074810ea92d34a41f7e92d45b2816c5a54

Observation f768d3f8-f83e-4dc0-8c91-00a35ec1b7e3 · outbound

This paper cites Learning to ground instructional articles in videos through narrations.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning to ground instructional articles in videos through narrations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:41.922633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:41.922633Z digest=sha256:a7e1f53da9ba4185b5ae28f55765c426ce5039a9414afd7d4ee91d062b9ce0be

Observation 392843a1-d2b0-42bc-a566-d817eda44526 · outbound

This paper cites Object priors for classifying and localizing unseen actions.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Object priors for classifying and localizing unseen actions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.233566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.927079Z digest=sha256:b98cab2a5e35884672b1381172b32b4300b665d07562320ea4b039e7f3dcad82

Observation f7caf8e7-e8af-42fc-8c45-92f4ced659d3 · outbound

This paper cites End-to-end learning of visual representations from uncurated instruc- tional videos.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos End-to-end learning of visual representations from uncurated instruc- tional videos

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.205632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.931494Z digest=sha256:06becba04d1d949d31ccd5a3bdc250620443a822e15956a01af4296c5fc01e7f

Observation 5439a55e-bf2d-4057-9040-60008bbf1ecc · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips, 2019.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Howto100m: Learning a text-video embedding by watching hundred million narrated video clips, 2019

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.151592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.936116Z digest=sha256:48801f5f23781f4bcb2e943380877b5050e7d340eee8eb424c35179cb7a2ef93

Observation d6fba638-7398-40f5-8d77-219f0861028d · outbound

This paper cites Efficient Estimation of Word Representations in Vector Space.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Efficient Estimation of Word Representations in Vector Space

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:41.941389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:41.941389Z digest=sha256:f830fd434d58a9ce3809f63c4b9a274c26378d816f2378957de0bbaf0e47a648

Observation f35a470c-b545-4034-a694-4679b5b08965 · outbound

This paper cites Verbs in action: Improving verb understanding in video-language models.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Verbs in action: Improving verb understanding in video-language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.123711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.946220Z digest=sha256:644fe964d9e68d235e7872f04122a370a3bb6cab86622fbb67f9a0732a06e0ea

Observation 366828e0-c3b9-4c9c-a70b-1df0eefb606c · outbound

This paper cites Spoken moments: Learning joint audio-visual representations from video descriptions.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Spoken moments: Learning joint audio-visual representations from video descriptions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.028356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.951307Z digest=sha256:c6e5e5b8c17b6382921b5a96090ca0c473a0dcc73d15c0eb68b649711a6f1536

Observation c20dc994-7087-4b41-9e42-0a02bfd86572 · outbound

This paper cites Zero-shot temporal action detection via vision-language prompting.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Zero-shot temporal action detection via vision-language prompting

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:44.007255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:41.996161Z digest=sha256:2eee82df1faad14fd50a94c6312b3986fc1d4aa3c27ccc0aef76537e541f6455

Observation 15acba4f-6f76-4565-bc7e-44d4020b9676 · outbound

This paper cites Expanding language-image pretrained models for gen- eral video recognition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Expanding language-image pretrained models for gen- eral video recognition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.896043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.021559Z digest=sha256:85c838d0ee9db9d19663152404414de488d6fb0aa0797418376b7cb183cb3be3

Observation f86fbf1e-492b-422a-8aa7-857260a243b9 · outbound

This paper cites Learning multimodal representations for unseen activities.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning multimodal representations for unseen activities

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.867587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.027762Z digest=sha256:28a8bc4da777b3155ee2bbce61b7484df5a1191728276d32810b172498fc2331

Observation 4cc7ed0a-b02c-45a6-9cab-ef19423003a3 · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.761687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.032409Z digest=sha256:59cfdccb03c09c97cc0661fc6285f69a728c467712506dc2fd9c25a9a950134e

Observation 5efc7eac-01e5-4a80-a6f9-02e7aae9ae5e · outbound

This paper cites What does a platypus look like? generating customized prompts for zero-shot image classification.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos What does a platypus look like? generating customized prompts for zero-shot image classification

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:42.043056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:42.043056Z digest=sha256:8359a0eec7f2deb0eec865707d2f72cd4342ad97121b16cad8977d529aa48dab

Observation 68d4deaf-c8e6-4719-bcb3-44f0b725a8b2 · outbound

This paper cites Alignment-uniformity aware representation learning for zero-shot video classifica- tion.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Alignment-uniformity aware representation learning for zero-shot video classifica- tion

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.736529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.090035Z digest=sha256:e3faaaa3f2db87169bff2e9f5ffcf23ff7ccd7907ac31970a956a9dd5526574a

Observation 08ddf227-795d-407f-a990-67c7f1f1ef5e · outbound

This paper cites Rethinking zero-shot action recognition: Learning from latent atomic actions.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Rethinking zero-shot action recognition: Learning from latent atomic actions

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.648754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.101490Z digest=sha256:5a246acaca392647ce0f7e5a23238d3c5c431488925a77c5417fdaac1793f421

Observation 28f404fb-d0db-4909-bc7a-7117967cc0c3 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning transferable visual models from natural language supervi- sion

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:42.105973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:42.105973Z digest=sha256:e9917c6d3a1f32963413e0b519012f7d591c6cba50b785837d61044401f9eb01

Observation d556de65-a245-475d-94c7-ea542c650a66 · outbound

This paper cites Language-based action concept spaces improve video self-supervised learn- ing.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Language-based action concept spaces improve video self-supervised learn- ing

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.568529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.110115Z digest=sha256:ab56b063be0ba11ad921f083a13dc7edcb2bd666aa29935e0fcbcab4851fd804

Observation 60c15b2d-c2ce-4b4c-abcd-b2b204a7c148 · outbound

This paper cites Fine-tuned clip models are efficient video learners.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Fine-tuned clip models are efficient video learners

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.480138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.114311Z digest=sha256:7d47298c314fd6ea20c6ed773300ee744457d0eb1cec441fddc2b664b71f8389

Observation 91ad1310-2332-43cf-933a-e037b138d7c1 · outbound

This paper cites Is a Caption Worth a Thousand Images? A Controlled Study for Representation Learning.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Is a Caption Worth a Thousand Images? A Controlled Study for Representation Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:42.119921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:42.119921Z digest=sha256:e83323c86d98df77520b005c9a28f280ebcb1feb197d356d834916864c0122b4

Observation 2aaf919a-3816-4afe-bbb6-cd3f35a5bd35 · outbound

This paper cites As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.414332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.126043Z digest=sha256:fd350e1b2dfdf9698bf1d1cd8b85c0e284af9dee5d5600f28f56123b33370b6a

Observation 230c78fd-ddb1-42ee-9656-e57fd4294125 · outbound

This paper cites Mpnet: Masked and permuted pre-training for language understanding.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Mpnet: Masked and permuted pre-training for language understanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.398085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.157706Z digest=sha256:0e4740346c61687e100a094abd147650544fab8148c411e21ff54c4aa808548c

Observation 43391558-a256-44fc-9939-dd713ce37841 · outbound

This paper cites Coin: A large-scale dataset for comprehensive instructional video analysis.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Coin: A large-scale dataset for comprehensive instructional video analysis

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.295076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.184990Z digest=sha256:349cd3a93660c9ec5ea8db147bc84c2c6d8509eba5dc75e631d3a9d1e3dd3318

Observation 04979466-8c37-4949-a8df-3dead8f8a0aa · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos ActionCLIP: A New Paradigm for Video Action Recognition

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:42.189536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:42.189536Z digest=sha256:294465d73297ae5e9ece6af871bab6fb9e592d7fd4f2bda060253051c7b5ffda

Observation b9aa2de1-0f43-4a3e-9a2b-0f8408ebdcb6 · outbound

This paper cites Vilta: Enhancing vision-language pre-training through textual augmentation.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Vilta: Enhancing vision-language pre-training through textual augmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.235796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.193948Z digest=sha256:8bd0fb93de9589da42b6bc88ffe0a0b011db48fc4ed8617376e0bb5f23c63a9d

Observation c871377a-b35e-4725-98b1-44ce9e340e4f · outbound

This paper cites Pax- ion: Patching action knowledge in video-language founda- tion models.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Pax- ion: Patching action knowledge in video-language founda- tion models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.152810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.198093Z digest=sha256:ead04860cd77daa386db13916ba47fc535420e6d2293c080eea0682539c32f3d

Observation 002a0166-01aa-4425-8f9b-4ebf5272c2d3 · outbound

This paper cites Zero-shot event detection using multi-modal fusion of weakly supervised concepts.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Zero-shot event detection using multi-modal fusion of weakly supervised concepts

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:43.059961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.202643Z digest=sha256:6c2696904b542399acdb3364b6584e27ab89c7134baae9779676995736a82ab2

Observation fa7edda3-578f-463d-942c-58c3ed6cb094 · outbound

This paper cites Cap4video: What can auxiliary captions do for text-video retrieval? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 10704–10713, 2023.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Cap4video: What can auxiliary captions do for text-video retrieval? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 10704–10713, 2023

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:42.945797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.207694Z digest=sha256:24aa0e575f9002e9713cf67654e248d82c5343184232c8bcc9e15bf6119a24b7

Observation d8a3b65a-cb63-4822-9401-d98af018659b · outbound

This paper cites Revisiting clas- sifier: Transferring vision-language models for video recog- nition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Revisiting clas- sifier: Transferring vision-language models for video recog- nition

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:42.879421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.213370Z digest=sha256:9e5c7805996b08181aa7690cb0c90bce029df64f698635a0918f0c99ecb817c6

Observation d61984bc-1b77-4ae9-894d-e0c97895b061 · outbound

This paper cites Bidirectional cross- modal knowledge exploration for video recognition with pre-trained vision-language models.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Bidirectional cross- modal knowledge exploration for video recognition with pre-trained vision-language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:42.835387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.227643Z digest=sha256:1ba5f3f50b7508cabf70e895120d2d9ee2721483b55488733ec97793c9debf4b

Observation cb7ad621-edbc-4d41-8701-0aad67f7b9c2 · outbound

This paper cites Generative action description prompts for skeleton-based action recognition.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Generative action description prompts for skeleton-based action recognition

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:42.651926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.259861Z digest=sha256:6e8a6ad3e59cd96662bb3d11b9f6c8174f574a1eda450bc1826a3edc81601807

Observation 31beffb7-5a7b-4c4d-b56c-48ca0150bb0e · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T14:08:42.293896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:08:42.293896Z digest=sha256:b251cdf8ea50d504fc283aaa2065bacda55aa1d6edac3016cadcfb4ab5ca71e5

Observation cb277bc7-0b42-4946-9107-b6beed78ed43 · outbound

This paper cites Movie genre classification by language augmentation and shot sampling.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Movie genre classification by language augmentation and shot sampling

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:42.622826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.298672Z digest=sha256:3904dfc967a652001f8a1207134e32601c491fc84304db0412639c1ecfc15d87

Observation 5cbe71b3-35a1-47ab-b07f-abd064bcdac0 · outbound

This paper cites Learning video representations from large lan- guage models.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning video representations from large lan- guage models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:42.566586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.303134Z digest=sha256:504cce493f03b796243b7a668ab5777bf69af9afa52da44551d9c450ad385918

Observation 79a20398-21bc-4633-ad0f-c8d514f8455c · outbound

This paper cites Learning procedure-aware video represen- tation from instructional videos and their narrations.

ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning procedure-aware video represen- tation from instructional videos and their narrations

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:08:42.457418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:08:42.308404Z digest=sha256:b719c168a3de4a32900d5e4a9ebae8801b6fbf69b0d6ec5fd34e12f71ec1086c

Pith citing papers

No inbound Pith citation observations are available.