← back to paper
arxiv: 2506.13458 · 2 revisions
Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images