Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:08:42.308404Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2411.15628.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:08:42.308404Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
66 of 66 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bc9bc2a2-bd41-45fd-a3e5-6572c2935db5 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb663771-06c0-4920-b6a0-434d580c9460 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Ht-step: Aligning instructional articles with how-to videos.Advances in Neural Information Processing Systems, 36, 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 96065e79-07ec-478b-a35d-a020b140eec8 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Exploring synonyms as context in zero-shot action recognition
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 84fa500b-d883-473c-b707-e4b8e2c44fa2 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Hiervl: Learning hierarchical video- language embeddings
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 732405b6-5ed4-4b7a-9f92-fae80af0857d · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos The ikea asm dataset: Understanding people assem- bling furniture through actions, objects and pose
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e7551e14-bab1-4879-b945-d862870784d2 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Is space-time attention all you need for video understanding? In ICML, volume 2, page 4, 2021
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 80e0cd83-12ec-4d51-a6d1-3e5312a09955 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Rethinking zero-shot video classification: End-to-end training for realistic applications
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ff4d3c15-0b5a-44a9-8445-fed95ac1fa51 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Regen: A good generative zero-shot video classifier should be rewarded
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9a249bd3-7f50-4a03-b916-3db3601cda56 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Quo vadis, action recognition? a new model and the kinetics dataset
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 74109dde-6953-4ac5-bf01-dbd2129eacc2 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Elaborative rehearsal for zero-shot action recognition
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b606cf03-a89a-4827-a2cf-a92ae546cf73 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11326423-3c91-49f8-92de-fb0fedaa099c · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Teaching structured vision & language concepts to vision & language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a75dd326-4fb5-44bc-8847-7ade7db59dda · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Step- former: Self-supervised step discovery and localization in instructional videos
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2fec3f59-f902-4123-84b1-945a3c7abf3e · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Zero-shot action recognition in videos: A survey
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7d0628ee-9900-45d8-868a-4c71a8f9d486 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Ms-tcn: Multi-stage tem- poral convolutional network for action segmentation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bea9a6b6-a870-40a8-bd16-a779afd8df73 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning to recognize objects in egocentric activities
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d9e4b6be-e75c-4f34-9bd5-bfb91ebfc51d · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Slowfast networks for video recognition
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9e8b1f8f-6ab6-4707-8485-c0dbb80803cd · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Prego: online mistake detection in procedural ego- centric videos
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1d271bf2-ada1-4cb9-81a9-db6c660651a9 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Improving zero-shot gen- eralization and robustness of multi-modal models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0d3393ef-0247-49d9-9225-eb6b77159975 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Weakly-supervised action segmentation and unseen error detection in anomalous instructional videos
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7263b4e0-2589-4854-b751-0c54a2799b8f · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Ego4d: Around the world in 3,000 hours of egocentric video
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 90c21025-a295-4590-91b0-ee49897d6703 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Temporal alignment networks for long-term video
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa393e83-200d-46a6-825e-39e6971c9c23 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Probing Image-Language Transformers for Verb Understanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac89540c-787a-4938-a1f6-9275e2aa9601 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Clover: Towards a unified video-language alignment and fusion model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation be5709eb-168f-4b2c-940d-cbf8a686689b · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Fine-grained generalized zero-shot learning via dense attribute-based attention
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 74b88be9-9871-4bae-b5b4-d24e6670ea2b · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Objects2action: Classifying and localiz- ing actions without any video example
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e9921a13-86a5-40c0-af2e-6a837086a824 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Prompting visual-language models for efficient video understanding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation df167b3c-640b-476e-92fe-196eaf923582 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Error detection in egocentric procedural task videos
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e749117b-00d5-444c-995e-e6de40abf633 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Align and prompt: Video-and-language pre-training with entity prompts
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2a32aef5-e3fe-42f0-bfeb-c374ee904f9f · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Cross-modal representation learning for zero- shot action recognition
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8b00805e-9d19-4f0d-93f4-716827babb06 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Match, expand and im- prove: Unsupervised finetuning for zero-shot action recog- nition with language knowledge
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bb8f3bb8-0369-4a79-ba98-800b5391f2a3 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning to recognize procedural activities with distant supervision
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56a4e60a-3ba5-486c-b277-0223da873a3d · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Out-of-distribution detection for gener- alized zero-shot action recognition
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f768d3f8-f83e-4dc0-8c91-00a35ec1b7e3 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning to ground instructional articles in videos through narrations
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 392843a1-d2b0-42bc-a566-d817eda44526 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Object priors for classifying and localizing unseen actions
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f7caf8e7-e8af-42fc-8c45-92f4ced659d3 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos End-to-end learning of visual representations from uncurated instruc- tional videos
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5439a55e-bf2d-4057-9040-60008bbf1ecc · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Howto100m: Learning a text-video embedding by watching hundred million narrated video clips, 2019
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d6fba638-7398-40f5-8d77-219f0861028d · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Efficient Estimation of Word Representations in Vector Space
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f35a470c-b545-4034-a694-4679b5b08965 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Verbs in action: Improving verb understanding in video-language models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 366828e0-c3b9-4c9c-a70b-1df0eefb606c · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Spoken moments: Learning joint audio-visual representations from video descriptions
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c20dc994-7087-4b41-9e42-0a02bfd86572 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Zero-shot temporal action detection via vision-language prompting
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 15acba4f-6f76-4565-bc7e-44d4020b9676 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Expanding language-image pretrained models for gen- eral video recognition
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f86fbf1e-492b-422a-8aa7-857260a243b9 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning multimodal representations for unseen activities
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4cc7ed0a-b02c-45a6-9cab-ef19423003a3 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Egovlpv2: Egocentric video-language pre-training with fusion in the backbone
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5efc7eac-01e5-4a80-a6f9-02e7aae9ae5e · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos What does a platypus look like? generating customized prompts for zero-shot image classification
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68d4deaf-c8e6-4719-bcb3-44f0b725a8b2 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Alignment-uniformity aware representation learning for zero-shot video classifica- tion
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 08ddf227-795d-407f-a990-67c7f1f1ef5e · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Rethinking zero-shot action recognition: Learning from latent atomic actions
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 28f404fb-d0db-4909-bc7a-7117967cc0c3 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning transferable visual models from natural language supervi- sion
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d556de65-a245-475d-94c7-ea542c650a66 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Language-based action concept spaces improve video self-supervised learn- ing
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 60c15b2d-c2ce-4b4c-abcd-b2b204a7c148 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Fine-tuned clip models are efficient video learners
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 91ad1310-2332-43cf-933a-e037b138d7c1 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Is a Caption Worth a Thousand Images? A Controlled Study for Representation Learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aaf919a-3816-4afe-bbb6-cd3f35a5bd35 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 230c78fd-ddb1-42ee-9656-e57fd4294125 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Mpnet: Masked and permuted pre-training for language understanding
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 43391558-a256-44fc-9939-dd713ce37841 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Coin: A large-scale dataset for comprehensive instructional video analysis
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 04979466-8c37-4949-a8df-3dead8f8a0aa · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos ActionCLIP: A New Paradigm for Video Action Recognition
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9aa2de1-0f43-4a3e-9a2b-0f8408ebdcb6 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Vilta: Enhancing vision-language pre-training through textual augmentation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c871377a-b35e-4725-98b1-44ce9e340e4f · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Pax- ion: Patching action knowledge in video-language founda- tion models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 002a0166-01aa-4425-8f9b-4ebf5272c2d3 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Zero-shot event detection using multi-modal fusion of weakly supervised concepts
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fa7edda3-578f-463d-942c-58c3ed6cb094 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Cap4video: What can auxiliary captions do for text-video retrieval? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 10704–10713, 2023
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d8a3b65a-cb63-4822-9401-d98af018659b · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Revisiting clas- sifier: Transferring vision-language models for video recog- nition
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d61984bc-1b77-4ae9-894d-e0c97895b061 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Bidirectional cross- modal knowledge exploration for video recognition with pre-trained vision-language models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cb7ad621-edbc-4d41-8701-0aad67f7b9c2 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Generative action description prompts for skeleton-based action recognition
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 31beffb7-5a7b-4c4d-b56c-48ca0150bb0e · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb277bc7-0b42-4946-9107-b6beed78ed43 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Movie genre classification by language augmentation and shot sampling
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5cbe71b3-35a1-47ab-b07f-abd064bcdac0 · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning video representations from large lan- guage models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 79a20398-21bc-4633-ad0f-c8d514f8455c · outbound
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos Learning procedure-aware video represen- tation from instructional videos and their narrations
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
No inbound Pith citation observations are available.