Pith. sign in

Paper Citation Record · LEDGER

HuMoCon: Concept Discovery for Human Motion Understanding

As of 15 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 0 inbound Pith citation observations for arXiv:2505.20920.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20920 v1

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:48:17.043562Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

91 of 91 outbound references displayed

  • verified exact2
  • verified fuzzy47
  • unresolved41
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6a04c1be-3b50-48f9-a3aa-6d6e70c25472 · outbound

This paper cites Teach: Temporal action composition for 3d hu- mans.

HuMoCon: Concept Discovery for Human Motion Understanding Teach: Temporal action composition for 3d hu- mans

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:09.924564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:09.924564Z digest=sha256:b8df941d9ef0e572f22bd1f7fe7cc4da92df819906a298feb348036cc1896e07

Observation 20f56a7f-9e9b-4041-9c62-13c5e3c02fe8 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

HuMoCon: Concept Discovery for Human Motion Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.034634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.034634Z digest=sha256:fd3f629da69a7dcbab66a4976a43095a9d6e5d5ce930da06cae2a35f399b7777

Observation f1855ea7-529f-4ad9-8e66-2ca7a18ed8db · outbound

This paper cites Ac- tion quality assessment with temporal parsing transformer.

HuMoCon: Concept Discovery for Human Motion Understanding Ac- tion quality assessment with temporal parsing transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.114885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.114885Z digest=sha256:bf864d78c5c2200255726277abaa77aeb71608c3ba644ac8f44223b055d37354

Observation b59fd234-9607-451d-8ce8-c0a755cf1132 · outbound

This paper cites Implicit neural representations for variable length human motion generation.

HuMoCon: Concept Discovery for Human Motion Understanding Implicit neural representations for variable length human motion generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.177615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.177615Z digest=sha256:7fff1114c9626395a23ba19841ee4fdd5573b40d56666496c44cea1ce9e8edad

Observation 778a610e-bda4-495d-a79b-d180f7739efe · outbound

This paper cites MotionLLM: Understanding Human Behaviors from Human Motions and Videos.

HuMoCon: Concept Discovery for Human Motion Understanding MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.244869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.244869Z digest=sha256:bb1f3b201e82e845cc7ed4679e3a458c6f5ec686cbe577eee3880a5d3d201576

Observation ca57b96e-3861-4524-8fc5-a5addd528d5c · outbound

This paper cites Pose Trainer: Correcting Exercise Posture using Pose Estimation.

HuMoCon: Concept Discovery for Human Motion Understanding Pose Trainer: Correcting Exercise Posture using Pose Estimation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.314111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.314111Z digest=sha256:0842491c649bb2c6dda2f098edfde6023cabfa4dad8c8510026a2e10a57156e4

Observation 8ce6c9d6-75e6-4b1e-a17a-a4a5be0c0cc4 · outbound

This paper cites VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset.

HuMoCon: Concept Discovery for Human Motion Understanding VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.413560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.413560Z digest=sha256:00620e13de0a97e2882c73b8da3ddc56b5488296cdf7767f3398957a514b4059

Observation 34b2f955-4cd2-4802-847c-9557cc812b52 · outbound

This paper cites Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.Advances in Neural Information Processing Sys- tems, 36:72842–72866, 2023.

HuMoCon: Concept Discovery for Human Motion Understanding Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.Advances in Neural Information Processing Sys- tems, 36:72842–72866, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.501301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.501301Z digest=sha256:7f7a8e8d7d3d7b460fcd381edd2b4cd9e71013b60e591dd5c50b8ec0d8743b8d

Observation 43231859-4654-4fd6-bf88-baf8537fb9ec · outbound

This paper cites Posefix: correcting 3d hu- man poses with natural language.

HuMoCon: Concept Discovery for Human Motion Understanding Posefix: correcting 3d hu- man poses with natural language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.565338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.565338Z digest=sha256:5df2000475635d826102012a04015661ed751b8ae53f12d1073a620d72e75132

Observation 2db973a8-70d5-4453-a49a-047820280c13 · outbound

This paper cites Behavior recognition via sparse spatio-temporal features.

HuMoCon: Concept Discovery for Human Motion Understanding Behavior recognition via sparse spatio-temporal features

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.633738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.633738Z digest=sha256:2b570710861d7056f38b76cbcbd5638c67eb2b86d4f59a9005d789c1e27fab4b

Observation 4e4dc90b-7635-42b4-b240-1c6fc46fbd75 · outbound

This paper cites Hierarchical recur- rent neural network for skeleton based action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Hierarchical recur- rent neural network for skeleton based action recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.760959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.760959Z digest=sha256:7e8ee42ffacc2918bbca429a9091159fbb4f75dfd5dce52891178a103d22c72d

Observation 1489a40f-1f6a-4fa6-8b8b-bb7f10d95df4 · outbound

This paper cites Clap learning audio concepts from nat- ural language supervision.

HuMoCon: Concept Discovery for Human Motion Understanding Clap learning audio concepts from nat- ural language supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.850580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.850580Z digest=sha256:7d295d0a994191165a7f18c49ce2d8a8833a8cfa884b0c483946aeb4530849b4

Observation c782aa6f-b0b6-472c-be5e-b22c87b436cd · outbound

This paper cites Motion question answering via modular motion programs.

HuMoCon: Concept Discovery for Human Motion Understanding Motion question answering via modular motion programs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.909322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.909322Z digest=sha256:194097e99c016a0bd69d2e3d9696f34e2bf3ff7bae28da9c83313898e23ca0dc

Observation 417681e7-4df5-4451-9114-711ab491d4b7 · outbound

This paper cites Towards accurate active camera localization.

HuMoCon: Concept Discovery for Human Motion Understanding Towards accurate active camera localization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.992617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.992617Z digest=sha256:14e9e573e623dea60ee49abd27c4ed371229265765bde2cbec098c641cf727a6

Observation a96723f5-5384-4022-9933-3d07570c428b · outbound

This paper cites Cigtime: Corrective instruction generation through inverse motion editing.

HuMoCon: Concept Discovery for Human Motion Understanding Cigtime: Corrective instruction generation through inverse motion editing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:11.057582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:11.057582Z digest=sha256:e7977ec78e5dc02e8b228ed134a5007dcf3713b335656c724bf46cbdfbeee49a

Observation ff860867-19ff-4e5c-bc38-fb68429e0236 · outbound

This paper cites Slowfast networks for video recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Slowfast networks for video recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.903825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:11.142433Z digest=sha256:5679afa5bae738fae8c9450fcfa347dcdb5693c375e67cf8346f79c67e3b1221

Observation 3eac355a-cfcc-4cf5-911f-9df17de7310a · outbound

This paper cites Chatpose: Chatting about 3d human pose.

HuMoCon: Concept Discovery for Human Motion Understanding Chatpose: Chatting about 3d human pose

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.889083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:11.225495Z digest=sha256:b82c59f7db2026aec5318f5ce5a4c512b4bdcb6459ae2a885863a996e58548ae

Observation ea2b70de-3e15-47dc-bbb1-6c4ae6750b3a · outbound

This paper cites Aifit: Automatic 3d human-interpretable feedback models for fitness training.

HuMoCon: Concept Discovery for Human Motion Understanding Aifit: Automatic 3d human-interpretable feedback models for fitness training

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.871720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:11.320638Z digest=sha256:55478ab446a9866d52129ff248fe3d57dea0461bffb18e3d2e43b60e833be15a

Observation 20667788-0537-40e8-a2b5-447e05ac8a1c · outbound

This paper cites Imagebind: One embedding space to bind them all.

HuMoCon: Concept Discovery for Human Motion Understanding Imagebind: One embedding space to bind them all

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.742066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:11.391368Z digest=sha256:81d6fc6c0f45ea1eba160ffecdfcd0e4778e02ce0253dfa10a112cf4e75aba3c

Observation 51dfab67-7197-434c-aa1f-0087d77241be · outbound

This paper cites Ac- tion2motion: Conditioned generation of 3d human motions.

HuMoCon: Concept Discovery for Human Motion Understanding Ac- tion2motion: Conditioned generation of 3d human motions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.407672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:11.475504Z digest=sha256:36d695403715c14061e992ba68768c6af5cdd655311dd71389c43e15a66ef806

Observation 4997ba37-e887-4c0a-b5a8-33be0db7c361 · outbound

This paper cites Generating diverse and natural 3d human motions from text.

HuMoCon: Concept Discovery for Human Motion Understanding Generating diverse and natural 3d human motions from text

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.123330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:11.547308Z digest=sha256:4736e15f5369d095dd4cfe34cf383027926715c3d54c4571e05cbb89a8ffee70

Observation e775fcd9-ebfd-439c-8d60-a8bcd4cb95e9 · outbound

This paper cites Generating diverse and natural 3d human motions from text.

HuMoCon: Concept Discovery for Human Motion Understanding Generating diverse and natural 3d human motions from text

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:25.893693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:11.636934Z digest=sha256:a8cc4f38b3f3e7d0cd05887061c657740632741deb743496f86a9661c7c66a88

Observation e15425d3-6e93-4e05-81ba-f26ababa7d1c · outbound

This paper cites Tm2t: Stochastic and tokenized modeling for the reciprocal genera- tion of 3d human motions and texts.

HuMoCon: Concept Discovery for Human Motion Understanding Tm2t: Stochastic and tokenized modeling for the reciprocal genera- tion of 3d human motions and texts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:11.730244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:11.730244Z digest=sha256:c10e05f70927e87f7c8572b1fa893d573308e0b0d2693cfa9038f12fbcf474df

Observation b78b7213-7196-48e1-8d36-82aa6765d731 · outbound

This paper cites Contrastive learning from ex- tremely augmented skeleton sequences for self-supervised action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Contrastive learning from ex- tremely augmented skeleton sequences for self-supervised action recognition

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:25.653995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:11.815961Z digest=sha256:35f42d5c667b75035bcbe06d3802050e960e759fcb3d53984cd2c1a7a298709a

Observation 29a51d22-1927-4a53-9eef-ab08a9b70ff6 · outbound

This paper cites Autoad ii: The sequel-who, when, and what in movie audio description.

HuMoCon: Concept Discovery for Human Motion Understanding Autoad ii: The sequel-who, when, and what in movie audio description

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:25.434143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:11.935438Z digest=sha256:c3cbd4b6945f1ff9c297179fad147315891100458095215dd884124cd1b215e4

Observation ba9114fe-6119-4eb3-9b85-ac1b458f5310 · outbound

This paper cites Autoad iii: The prequel-back to the pixels.

HuMoCon: Concept Discovery for Human Motion Understanding Autoad iii: The prequel-back to the pixels

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:25.144137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:12.033488Z digest=sha256:bee1c1ae684f51919d5a913fdca8c13f8c70884d3f801b75c674d510601fd808

Observation b07742b7-366d-4efd-96cc-966dcf1311c4 · outbound

This paper cites Masked autoencoders are scalable vision learners.

HuMoCon: Concept Discovery for Human Motion Understanding Masked autoencoders are scalable vision learners

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.973255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:12.126802Z digest=sha256:d9d9f452c8106cec0d47e7819e231ea96f79f7054669364d55b63c57afaa65c9

Observation b482f4d7-6ad4-4402-ae8c-d7f83a1888bb · outbound

This paper cites Phase- functioned neural networks for character control.ACM Transactions on Graphics (TOG), 36(4):1–13, 2017.

HuMoCon: Concept Discovery for Human Motion Understanding Phase- functioned neural networks for character control.ACM Transactions on Graphics (TOG), 36(4):1–13, 2017

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.817282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:12.239768Z digest=sha256:cae0d7d49a703b7269d72f8bb75245b2ca73137468f2291b6173be76275fc463

Observation 03b1b41f-5f04-46c5-9cc2-4d0979884ee3 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

HuMoCon: Concept Discovery for Human Motion Understanding LoRA: Low-Rank Adaptation of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:12.389567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:12.389567Z digest=sha256:9091118977a73ddecc5bd8c17dd030cba922aeca59035bc83ceef47b4be3a6a9

Observation 03f14e86-f0ea-4dca-9e3a-63bced0f1254 · outbound

This paper cites Motiongpt: Human motion as a foreign lan- guage.Advances in Neural Information Processing Systems, 36:20067–20079, 2023.

HuMoCon: Concept Discovery for Human Motion Understanding Motiongpt: Human motion as a foreign lan- guage.Advances in Neural Information Processing Systems, 36:20067–20079, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:12.495813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:12.495813Z digest=sha256:fca1462f9ae755c16456e60dae3479189796fb678f7ab953816f1a3e6610fcfe

Observation e528ae26-35a2-4910-b0f0-dbd5a8e40684 · outbound

This paper cites Hand-object contact consistency reasoning for human grasps generation.

HuMoCon: Concept Discovery for Human Motion Understanding Hand-object contact consistency reasoning for human grasps generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:12.639694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:12.639694Z digest=sha256:0bdb43a8ae90146d2fb5b9689f9ada2cee379042e40d995ad9f213d876519419

Observation 64e433a8-cb03-4d13-a276-bf17fe350259 · outbound

This paper cites SMPLX-Lite: A Realistic and Drivable Avatar Benchmark with Rich Geometry and Texture Annotations.

HuMoCon: Concept Discovery for Human Motion Understanding SMPLX-Lite: A Realistic and Drivable Avatar Benchmark with Rich Geometry and Texture Annotations

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:48:17.565477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:12.726047Z digest=sha256:3eca2f84f6005bafdc8a219732054297c278708514d7a7a50b67c1e7c73be942

Observation 9104fc61-8821-44f7-a4be-9cdafa246fb6 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

HuMoCon: Concept Discovery for Human Motion Understanding Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:12.856895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:12.856895Z digest=sha256:c6876933c1aa0bf146fccc085668ae682c98d6c509b6cc7ea64784421b8e0bb6

Observation 208464c2-a052-4d9d-92b3-afff0104b3fe · outbound

This paper cites A new representation of skele- ton sequences for 3d action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding A new representation of skele- ton sequences for 3d action recognition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.707784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:12.907945Z digest=sha256:ec85a61b09428aed567aa1dca3d373517b0376b69af80d67350fdc3478500df6

Observation 87faa928-de0b-4a98-b43c-6fe7b6f03427 · outbound

This paper cites Flame: Free- form language-based motion synthesis & editing.

HuMoCon: Concept Discovery for Human Motion Understanding Flame: Free- form language-based motion synthesis & editing

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:12.961097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:12.961097Z digest=sha256:e9c3650f85c951be0c3a7a1475c935a61328ad452d939707f57b761472905fc6

Observation ea5edf5d-c586-4251-ab3f-d4051cdb4253 · outbound

This paper cites Danceconv: Dance motion genera- tion with convolutional networks.IEEE Access, 10:44982– 45000, 2022.

HuMoCon: Concept Discovery for Human Motion Understanding Danceconv: Dance motion genera- tion with convolutional networks.IEEE Access, 10:44982– 45000, 2022

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.598734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:13.005451Z digest=sha256:75658f0786094fb5abcbfccdbef6194c03e56a01bb07c3cbe65892039fce23c2

Observation 1721af8e-1829-47c5-9651-76f8e441ebc4 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

HuMoCon: Concept Discovery for Human Motion Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:13.085093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:13.085093Z digest=sha256:f681189444aa89c940e564aeef1a6e3e2626c05f608ced16c20df0e208c8c712

Observation 57916962-adad-4b7f-9fd7-09b260c88146 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

HuMoCon: Concept Discovery for Human Motion Understanding VideoChat: Chat-Centric Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:13.144440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:13.144440Z digest=sha256:a46330391711562f39dbc06f1f7b529f778d73f5d93e60d0204fcee353f304c0

Observation cd0b5d30-6ae9-4a28-bdd1-cbb6461c1914 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

HuMoCon: Concept Discovery for Human Motion Understanding Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.472147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:13.202566Z digest=sha256:48a107cf8bdf3d67d29ffc79bc61da6dc08983a61dc8cef02fe710ce3c0c977c

Observation 16f9336b-0b54-4c04-8488-460a6540235f · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

HuMoCon: Concept Discovery for Human Motion Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:13.245692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:13.245692Z digest=sha256:424432f0194bd7f6ba5c7ec67424c8098815062eb3f11c74677268110dfa776f

Observation 0b529021-9209-4e40-942b-db8c8cd4b16c · outbound

This paper cites Motion-x: A large- scale 3d expressive whole-body human motion dataset.Ad- vances in Neural Information Processing Systems, 36, 2024.

HuMoCon: Concept Discovery for Human Motion Understanding Motion-x: A large- scale 3d expressive whole-body human motion dataset.Ad- vances in Neural Information Processing Systems, 36, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.321591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:13.341764Z digest=sha256:35ba8388b17f34abd3af8fb976e0b7a448de97386baa3d55b1d20fd6775a756d

Observation e6a15de1-018e-4550-b7a3-1cafb400d427 · outbound

This paper cites Actionlet- dependent contrastive learning for unsupervised skeleton- based action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Actionlet- dependent contrastive learning for unsupervised skeleton- based action recognition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.182622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:13.427954Z digest=sha256:2b412dd4e8c749e9f00edcbd23eba20d52dd35247798f019f4d54958a2d67d98

Observation e4c8e21f-edb1-4d66-b8df-79f53e57b561 · outbound

This paper cites Towards unified sur- gical skill assessment.

HuMoCon: Concept Discovery for Human Motion Understanding Towards unified sur- gical skill assessment

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.994813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:13.478624Z digest=sha256:cd1c5d95c4accaf6928914d3bd08448c225e168a6a16aaba9e8d0c855cbacfed

Observation c5dd6ce4-9680-4d59-b44d-fd72a6387429 · outbound

This paper cites Spatio-temporal lstm with trust gates for 3d human action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Spatio-temporal lstm with trust gates for 3d human action recognition

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.822773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:13.547011Z digest=sha256:676f056f74e4cdfa8f53562022e825cdc9a1f2f668f37be47ad40261fe415376

Observation 187ab89a-4c1d-4011-bfd7-2c4d3ab69860 · outbound

This paper cites Enhanced skele- ton visualization for view invariant human action recogni- tion.Pattern Recognition, 68:346–362, 2017.

HuMoCon: Concept Discovery for Human Motion Understanding Enhanced skele- ton visualization for view invariant human action recogni- tion.Pattern Recognition, 68:346–362, 2017

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.683090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:13.607043Z digest=sha256:e560b7cc5bdcf142e03438f4d3d9dd05a80b502da6e031e3da0873d2f6522907

Observation d90d895f-741d-4b2e-8605-da08a2e85ae3 · outbound

This paper cites InfoCon: Concept Discovery with Generative and Discriminative Informativeness.

HuMoCon: Concept Discovery for Human Motion Understanding InfoCon: Concept Discovery with Generative and Discriminative Informativeness

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:48:17.297953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:13.659969Z digest=sha256:86471aef949c416dcb0b2e84b3dd0309aff5b407006b6fad4b823551bb2b8e28

Observation 00d2d0fb-ce91-421d-9212-bbf7b32180b8 · outbound

This paper cites Posegpt: Quantization-based 3d human mo- tion generation and forecasting.

HuMoCon: Concept Discovery for Human Motion Understanding Posegpt: Quantization-based 3d human mo- tion generation and forecasting

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.435492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:13.714895Z digest=sha256:b951b77cad39624d86904b12a618fa3dc46ee6af7fdf15cdbc465a829264c846

Observation 8f552295-2322-4b0f-8068-6be3fc8388ce · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neu- rocomputing, 508:293–304, 2022.

HuMoCon: Concept Discovery for Human Motion Understanding Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neu- rocomputing, 508:293–304, 2022

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.253372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:13.805257Z digest=sha256:95f154c93c5540a42dbb921a24acb886ad2e5317509eb692bf196a0c95da9c1a

Observation 28359c28-23f9-4524-89dc-45a73fa2755f · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

HuMoCon: Concept Discovery for Human Motion Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:13.881665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:13.881665Z digest=sha256:cfd4e09acaaecfb43b3b7f23bf8574998f5360b5e87f88a51e91e73c78d8bd39

Observation d5e98ac8-6eb9-4368-ae87-f5f683df8823 · outbound

This paper cites PG-Video-LLaVA: Pixel Grounding Large Video-Language Models.

HuMoCon: Concept Discovery for Human Motion Understanding PG-Video-LLaVA: Pixel Grounding Large Video-Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:13.947245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:13.947245Z digest=sha256:23768ee40753e9e772fc5f3a58c06a5f805d4db9666397e2de0d4d6a13febfd7

Observation 9b0fb8e3-4431-46bb-a5b7-380fcf050f4d · outbound

This paper cites What and how well you performed? a multitask learning approach to action quality assessment.

HuMoCon: Concept Discovery for Human Motion Understanding What and how well you performed? a multitask learning approach to action quality assessment

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.136730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:13.998154Z digest=sha256:2c3aa85a522bca9f100fd8dd45adbd95447f6d18424271ab6819378943de0189

Observation c3457ee8-ab87-472d-853d-dca2a07d5498 · outbound

This paper cites Action- conditioned 3d human motion synthesis with transformer vae.

HuMoCon: Concept Discovery for Human Motion Understanding Action- conditioned 3d human motion synthesis with transformer vae

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:22.873624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:14.050854Z digest=sha256:4e5ffe797a0d539e8990f1ed98d2bc0dc4c492e8a51e4df416582e93828fd08d

Observation da2de4c7-d835-4237-985d-8b695a69468b · outbound

This paper cites Temos: Generating diverse human motions from textual descriptions.

HuMoCon: Concept Discovery for Human Motion Understanding Temos: Generating diverse human motions from textual descriptions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:14.094528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:14.094528Z digest=sha256:32bed57ea30b7d80ef069e9af3917d71128a3fbb975bf995145e9336bbbedbae

Observation 7a7255e9-6b35-4a8c-b1d6-d10d5156d21b · outbound

This paper cites Babel: Bodies, action and behavior with english la- bels.

HuMoCon: Concept Discovery for Human Motion Understanding Babel: Bodies, action and behavior with english la- bels

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:22.607481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:14.138546Z digest=sha256:db0b59f213c5a34eed49ad02e3500dc56066d074fee16b696f17dffa7381ca38

Observation 65f19c98-0b95-4647-bcf4-e2983ec9d777 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

HuMoCon: Concept Discovery for Human Motion Understanding Learning transferable visual models from natural language supervi- sion

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:22.425341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:14.215738Z digest=sha256:31d7f948432d121fcd63a3a9150cb841725a7e50fba4a6fc66e7a8b476c3ea9d

Observation 831fc922-753d-41b5-aa46-34fb4338c55c · outbound

This paper cites On the benefits of 3d pose and tracking for human action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding On the benefits of 3d pose and tracking for human action recognition

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:22.196705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:14.279679Z digest=sha256:f092e6332132d288428eb67a4db4ea9bbff8714ef632c8625a6ec153a33daebd

Observation 7fe7f90a-a16d-4072-87f5-c4acfd15452d · outbound

This paper cites Zero-shot audio captioning with audio-language model guidance and audio context keywords.

HuMoCon: Concept Discovery for Human Motion Understanding Zero-shot audio captioning with audio-language model guidance and audio context keywords

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:14.346070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:14.346070Z digest=sha256:4aeaf4a726d88f34015f31e51b305c8f83df2b1eed86989514045c679d580df6

Observation 8c342166-b30e-4688-bdf9-2c447fdf3826 · outbound

This paper cites Two- stream adaptive graph convolutional networks for skeleton- based action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Two- stream adaptive graph convolutional networks for skeleton- based action recognition

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:21.862099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:14.405018Z digest=sha256:1b74c9248e67e243c5b79ec12346b4493a0024ffa55e28f06ff386e188335f23

Observation 2b58ca6e-3251-406b-bc50-9d004b6fdca0 · outbound

This paper cites Neural state machine for character-scene interactions.ACM Transactions on Graphics, 38(6):178, 2019.

HuMoCon: Concept Discovery for Human Motion Understanding Neural state machine for character-scene interactions.ACM Transactions on Graphics, 38(6):178, 2019

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:21.647262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:14.464679Z digest=sha256:4e3d3ba156ecbf38c14fedcc24c909dca090ae4c125a6056cf4fc8030d35a895

Observation 3af641c8-f034-457d-b777-cafcd8bcf05a · outbound

This paper cites Local motion phases for learning multi-contact charac- ter movements.ACM Transactions on Graphics (TOG), 39 (4):54–1, 2020.

HuMoCon: Concept Discovery for Human Motion Understanding Local motion phases for learning multi-contact charac- ter movements.ACM Transactions on Graphics (TOG), 39 (4):54–1, 2020

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:14.523931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:14.523931Z digest=sha256:31a4e07f87f0c3d247cd3504e959bcbd21cff6661e07f2210371226c7655dc9e

Observation 6663db44-167f-4f5f-923c-045a26aecb86 · outbound

This paper cites Deepphase: Periodic autoencoders for learning motion phase manifolds.

HuMoCon: Concept Discovery for Human Motion Understanding Deepphase: Periodic autoencoders for learning motion phase manifolds

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:21.436539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:14.601668Z digest=sha256:e08a75632923d61b8e55a97f21066b584aa9803cab74b02ac22ebaf1bf679ddc

Observation d06102dc-abc2-48cb-9248-6985586b7706 · outbound

This paper cites Convolutional learning of spatio-temporal features.

HuMoCon: Concept Discovery for Human Motion Understanding Convolutional learning of spatio-temporal features

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:21.236498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:14.704401Z digest=sha256:dc82c7d6c5bf1893ee8d50a3b30a11c0e34863aef8b8bdb1bf17ba6d984f764e

Observation d1fe6ca9-1bbc-405f-873d-1e16e0da08c3 · outbound

This paper cites Motionclip: Exposing human motion generation to clip space.

HuMoCon: Concept Discovery for Human Motion Understanding Motionclip: Exposing human motion generation to clip space

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:21.032251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:14.788254Z digest=sha256:5f7e5e0cc723b31898acf4c353b0eb190bbf9e75c0c3257acd3cad9fb49363e9

Observation 5e1ffcd6-416e-4ea7-bf77-1091569b3774 · outbound

This paper cites Human Motion Diffusion Model.

HuMoCon: Concept Discovery for Human Motion Understanding Human Motion Diffusion Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:14.844576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:14.844576Z digest=sha256:95c2d3d383549b5aadcd4e3bb00d45805b8e349becd20cbf76b603d32ebeb987

Observation 3c5815f4-3f2d-4da3-bc78-a8f56306f38c · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

HuMoCon: Concept Discovery for Human Motion Understanding Learning spatiotemporal features with 3d convolutional networks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:14.909199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:14.909199Z digest=sha256:f9c547e32c1f100158474d34aadd0f4f9133e159bfb1fab67fa4d51640355c2a

Observation 56c8bae8-3a4c-453b-a1be-7e88c2a6ed36 · outbound

This paper cites Human action recognition by representing 3d skeletons as points in a lie group.

HuMoCon: Concept Discovery for Human Motion Understanding Human action recognition by representing 3d skeletons as points in a lie group

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:20.873807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:14.987206Z digest=sha256:5deb820ee9e6a556723495d2915b529275195c8cdd13e7a32d4801dc9ffb946e

Observation 06f1a0dd-c61b-4c52-9b0c-a9bd360d9b94 · outbound

This paper cites Action recognition with improved trajectories.

HuMoCon: Concept Discovery for Human Motion Understanding Action recognition with improved trajectories

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:20.555895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:15.061942Z digest=sha256:1f01fe1116c476011fba062dcf35454ff11a4de8276c862a7a9b459ed492a293

Observation a3db810f-40c1-4d18-ab05-227898ea01bf · outbound

This paper cites Learn- ing human dynamics in autonomous driving scenarios.

HuMoCon: Concept Discovery for Human Motion Understanding Learn- ing human dynamics in autonomous driving scenarios

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:20.420267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:15.133096Z digest=sha256:d95a283f197e84f4e52d9771c588dffc36ae7744c6bece8b21509c1bb2a7b3f2

Observation 59f471f3-70f9-4451-b87b-921b62699906 · outbound

This paper cites Vatex: A large-scale, high- quality multilingual dataset for video-and-language research.

HuMoCon: Concept Discovery for Human Motion Understanding Vatex: A large-scale, high- quality multilingual dataset for video-and-language research

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:20.259990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:15.244074Z digest=sha256:67e185dfc49eafc6a0adfd211830ef04e97fd852de5a60103ad8efb3c0edac81

Observation 8401b309-0b16-4b9d-ac1b-30f5586f3016 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

HuMoCon: Concept Discovery for Human Motion Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.328435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.328435Z digest=sha256:719afc5f9559e8ea679359a83e95206eb6737bf12ed5c16567cef1c5f649e735

Observation c51851bd-3b04-4f8f-9f48-14366c20f7cd · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

HuMoCon: Concept Discovery for Human Motion Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.401308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.401308Z digest=sha256:e4a97fbe53ecbcb68240f87be571d7e82801bb346cff997d2a4acfe7b51ad6a8

Observation b6deda1a-9d71-425a-98b1-b0e2b25f293a · outbound

This paper cites Unified Human-Scene Interaction via Prompted Chain-of-Contacts.

HuMoCon: Concept Discovery for Human Motion Understanding Unified Human-Scene Interaction via Prompted Chain-of-Contacts

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.499056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.499056Z digest=sha256:030a84fa4faadb4aa35ab7de069959799e01978fa4a6dea90c177ab2a780f51e

Observation be953bc4-6c03-4eb8-8487-f0ea8c17fd8c · outbound

This paper cites AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description.

HuMoCon: Concept Discovery for Human Motion Understanding AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.573298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.573298Z digest=sha256:c38079751e0a63222c0f78f3395f1c431aa8f72ebc8e7ccebaf3fac69cd0cdbb

Observation d007fada-1dd1-47bc-89fb-6af1778e96e1 · outbound

This paper cites Dexterous grasp transformer.

HuMoCon: Concept Discovery for Human Motion Understanding Dexterous grasp transformer

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:20.011797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:15.647947Z digest=sha256:292b33d46fd3b802a20e6e15880340dbe8488b27a1f7640146095e49cc288bce

Observation f76756c4-f32c-49d6-9d50-f76b9c548aa5 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

HuMoCon: Concept Discovery for Human Motion Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.720080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.720080Z digest=sha256:7ce33ee5f930cfa5d0c71ef5dc63619378d8d2bf675c78020248d2ad587f147f

Observation 10c3c04d-cff4-4210-8c7f-e449cfe56b35 · outbound

This paper cites Spatial tempo- ral graph convolutional networks for skeleton-based action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Spatial tempo- ral graph convolutional networks for skeleton-based action recognition

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.777468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.777468Z digest=sha256:70fab2bf82916442b8ce16866ba413dd8a8740ff90d4d40f537ed518a4571a7d

Observation 37fd70ca-82b3-4a2e-8099-524b47f3f24d · outbound

This paper cites Learning to use chopsticks in diverse gripping styles.

HuMoCon: Concept Discovery for Human Motion Understanding Learning to use chopsticks in diverse gripping styles

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:19.789211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:15.830957Z digest=sha256:168e6eb5797e94310688f3659cfec019ff4ed46c89b7520c1f82614f792e4cc7

Observation d7b03f79-b05b-4883-8c10-c6785a696709 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

HuMoCon: Concept Discovery for Human Motion Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:19.523041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:15.910767Z digest=sha256:453e435c0d1b8773f5d1019fdb3eb85c5bdc657c13be6a9146bc77210d6bda05

Observation d21bf28c-63a4-4cd6-ac0b-94068ef65ac2 · outbound

This paper cites Sigmoid loss for language image pre-training.

HuMoCon: Concept Discovery for Human Motion Understanding Sigmoid loss for language image pre-training

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:19.278946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:15.969101Z digest=sha256:34e2184022d77d85c9dbf9c15c4f1d8ed133aeac3dc0d9e41e47a07694b02830

Observation 5c7172aa-8f9f-46c2-bc77-e4df44dd84ca · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

HuMoCon: Concept Discovery for Human Motion Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:16.029474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:16.029474Z digest=sha256:6c02b3168180ca7775be344ea1ebf1a17fe4cd85486db7e3d170f0bffbfea0b7

Observation 47b9ecd6-67b5-43ac-b2b8-7def70600694 · outbound

This paper cites T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations.

HuMoCon: Concept Discovery for Human Motion Understanding T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:16.105127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:16.105127Z digest=sha256:641196d6879fdf534604feb38cf8c51fea71148025f4a1acd1f166fa99d3ff1f

Observation 9d0d6c33-8e81-444f-b42b-1d09f4fc4fd7 · outbound

This paper cites Motiondif- fuse: Text-driven human motion generation with diffusion model.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

HuMoCon: Concept Discovery for Human Motion Understanding Motiondif- fuse: Text-driven human motion generation with diffusion model.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:19.045164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:16.159016Z digest=sha256:c0ed095718a8cc83f1a9d95f45cc62e6a2be2ca627caffde2977ba3a0f8d4978

Observation 54025af5-d1e1-441b-94d3-973e27b5fe10 · outbound

This paper cites Pointclip: Point cloud understanding by clip.

HuMoCon: Concept Discovery for Human Motion Understanding Pointclip: Point cloud understanding by clip

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:18.857074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:16.224869Z digest=sha256:5c4dbd3939b6f438e6984944d434e45e23842fafedd1f7e9bc92b09df4f328b0

Observation 3d7d98d9-8c2b-4cdc-943d-70a929f975b8 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

HuMoCon: Concept Discovery for Human Motion Understanding LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:16.292122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:16.292122Z digest=sha256:2cd637d5d0504982cf20a5310750487a5ae9c1c1cb329655f33b53f3bcca5ab9

Observation 7d80e39c-03c8-4383-9fbd-44a6bb6f78f5 · outbound

This paper cites Streaming dense video captioning.

HuMoCon: Concept Discovery for Human Motion Understanding Streaming dense video captioning

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:18.710966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:16.441949Z digest=sha256:db8fd4c8241f1c0d2e55190179ebb5f5e0b0c95cf5fb45b65fc25a10594db3e2

Observation a48bdeba-b3a4-4799-b1b4-30ab37674bb7 · outbound

This paper cites Avatargpt: All- in-one framework for motion understanding planning gener- ation and beyond.

HuMoCon: Concept Discovery for Human Motion Understanding Avatargpt: All- in-one framework for motion understanding planning gener- ation and beyond

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:18.526788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:16.539175Z digest=sha256:ee53fdfafb0f50f66b0084c661f0e29dc1e599904311cbdccb8c8f48a4ebe3fb

Observation e2d842ec-0c37-4a27-a756-081035a8de30 · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

HuMoCon: Concept Discovery for Human Motion Understanding LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:16.643392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:16.643392Z digest=sha256:184daadc156880d1122860042b487a695201f5be47b6af12992019eab4e938b1

Observation 15700c2e-c4d9-429e-96b6-bb2ca64ecb9e · outbound

This paper cites Limited by GPU memory, we only take 8 key frames for each video following MotionLLM [5].

HuMoCon: Concept Discovery for Human Motion Understanding Limited by GPU memory, we only take 8 key frames for each video following MotionLLM [5]

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:18.383910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:16.773797Z digest=sha256:0dc2627a10cb56b3b47724b4687e9f24b9c12dd32b6c0f3e65c86e2a3fba1200

Observation d23a2469-1fb5-4239-8f6e-29726d222219 · outbound

This paper cites Visualizations for Motion Understanding Additional visualization results for motion understanding are provided in Fig.

HuMoCon: Concept Discovery for Human Motion Understanding Visualizations for Motion Understanding Additional visualization results for motion understanding are provided in Fig

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:18.128667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:16.873059Z digest=sha256:50a0e31768a75f035d0bfd6a6427cf0d34015f21a084834b754e46bc64610504

Observation 20a2d8f4-4502-4b51-8466-9ff35e6c1ca7 · outbound

This paper cites Evaluation on BABEL-QA Benchmark For the BABEL-QA benchmark, we extend the evalua- tion protocol used in previous multi-modality LLMs evalu- ations [80].

HuMoCon: Concept Discovery for Human Motion Understanding Evaluation on BABEL-QA Benchmark For the BABEL-QA benchmark, we extend the evalua- tion protocol used in previous multi-modality LLMs evalu- ations [80]

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:17.934458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:16.966088Z digest=sha256:6ffbfddc1531b72ec26f579ce3a2720700643a06c21ce6ef31b1c137c4b1ed3a

Observation e97ab098-55cd-4ba2-9b0f-dcf258105aed · outbound

This paper cites Structure of Hyper-Networks We utilize hyper-networksH u andH m, for velocity recon- struction, following InfoCon [46].

HuMoCon: Concept Discovery for Human Motion Understanding Structure of Hyper-Networks We utilize hyper-networksH u andH m, for velocity recon- struction, following InfoCon [46]

Reference 91

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:48:17.751549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T13:48:17.043562Z digest=sha256:d0ddd10858961d4243f16e955efcd5c14b0f31469305875a28a3fc89fb2aba5c

Pith citing papers

No inbound Pith citation observations are available.