Pith. sign in

Paper Citation Record · LEDGER

HuMoCon: Concept Discovery for Human Motion Understanding

As of 8 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 0 inbound Pith citation observations for arXiv:2505.20920.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20920 v1

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:48:17.043562Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

91 of 91 outbound references displayed

  • verified exact2
  • verified fuzzy47
  • unresolved41
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6a04c1be-3b50-48f9-a3aa-6d6e70c25472 · outbound

This paper cites Teach: Temporal action composition for 3d hu- mans.

HuMoCon: Concept Discovery for Human Motion Understanding Teach: Temporal action composition for 3d hu- mans

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:09.924564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:09.924564Z digest=sha256:d53473971bc082f7d1dffd0635c059cc32303dfc9ef9a384ba76c2b2686ed625

Observation 20f56a7f-9e9b-4041-9c62-13c5e3c02fe8 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

HuMoCon: Concept Discovery for Human Motion Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.034634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.034634Z digest=sha256:eab50c6ebf679bfe75b52669f622a8d641255e622f43628dadde10d4620e894b

Observation f1855ea7-529f-4ad9-8e66-2ca7a18ed8db · outbound

This paper cites Ac- tion quality assessment with temporal parsing transformer.

HuMoCon: Concept Discovery for Human Motion Understanding Ac- tion quality assessment with temporal parsing transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.114885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.114885Z digest=sha256:15af7c155ec6a86c6b7bd88b2f7d68a92a897e6726118c02437fdbf2bb706b83

Observation b59fd234-9607-451d-8ce8-c0a755cf1132 · outbound

This paper cites Implicit neural representations for variable length human motion generation.

HuMoCon: Concept Discovery for Human Motion Understanding Implicit neural representations for variable length human motion generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.177615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.177615Z digest=sha256:4b9f69f3c6bbc13d3634714e153b84be738ef80155c65e69fbd36cf8ec01081c

Observation 778a610e-bda4-495d-a79b-d180f7739efe · outbound

This paper cites MotionLLM: Understanding Human Behaviors from Human Motions and Videos.

HuMoCon: Concept Discovery for Human Motion Understanding MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.244869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.244869Z digest=sha256:dad5d45aea367ad28c6952a8854b907ca8c0b1bb7bc6eb761da3133318122d0f

Observation ca57b96e-3861-4524-8fc5-a5addd528d5c · outbound

This paper cites Pose Trainer: Correcting Exercise Posture using Pose Estimation.

HuMoCon: Concept Discovery for Human Motion Understanding Pose Trainer: Correcting Exercise Posture using Pose Estimation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.314111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.314111Z digest=sha256:3f9c81c54023da5c8a530640d970e098d4ce138662100aa09c0768b31a75d08d

Observation 8ce6c9d6-75e6-4b1e-a17a-a4a5be0c0cc4 · outbound

This paper cites VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset.

HuMoCon: Concept Discovery for Human Motion Understanding VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.413560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.413560Z digest=sha256:f8b5ecd220e1165e51ba11e8e8b3c2962e2a39ddae43d630eeca778f91c0baed

Observation 34b2f955-4cd2-4802-847c-9557cc812b52 · outbound

This paper cites Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.Advances in Neural Information Processing Sys- tems, 36:72842–72866, 2023.

HuMoCon: Concept Discovery for Human Motion Understanding Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.Advances in Neural Information Processing Sys- tems, 36:72842–72866, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.501301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.501301Z digest=sha256:1ea0c198c90be58b9b9305e812fdacbe178186210ea08880fbbaedd1da7cf9df

Observation 43231859-4654-4fd6-bf88-baf8537fb9ec · outbound

This paper cites Posefix: correcting 3d hu- man poses with natural language.

HuMoCon: Concept Discovery for Human Motion Understanding Posefix: correcting 3d hu- man poses with natural language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.565338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.565338Z digest=sha256:ac3efd821db52bb388ba34eb1d66cbd948434a506a7c78bbc12d099705710bad

Observation 2db973a8-70d5-4453-a49a-047820280c13 · outbound

This paper cites Behavior recognition via sparse spatio-temporal features.

HuMoCon: Concept Discovery for Human Motion Understanding Behavior recognition via sparse spatio-temporal features

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.633738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.633738Z digest=sha256:d92037785901f7323b44f61dc0d6a31a91ae929ff98fb98807edee5c35f78be4

Observation 4e4dc90b-7635-42b4-b240-1c6fc46fbd75 · outbound

This paper cites Hierarchical recur- rent neural network for skeleton based action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Hierarchical recur- rent neural network for skeleton based action recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.760959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.760959Z digest=sha256:51cca8b21ecedc833f4d49d6f849a10e88fea6c5dc245e3f83e87cc9683a3778

Observation 1489a40f-1f6a-4fa6-8b8b-bb7f10d95df4 · outbound

This paper cites Clap learning audio concepts from nat- ural language supervision.

HuMoCon: Concept Discovery for Human Motion Understanding Clap learning audio concepts from nat- ural language supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.850580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.850580Z digest=sha256:572e25c4ed8601cf4c91207c3f77bcbc98b9bdcadce1a3e57ff1063537b78130

Observation c782aa6f-b0b6-472c-be5e-b22c87b436cd · outbound

This paper cites Motion question answering via modular motion programs.

HuMoCon: Concept Discovery for Human Motion Understanding Motion question answering via modular motion programs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.909322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.909322Z digest=sha256:98f57e93bac4ea3a0cb31793cbfe77e3abf7ccddd52599e8540292a818bc1c51

Observation 417681e7-4df5-4451-9114-711ab491d4b7 · outbound

This paper cites Towards accurate active camera localization.

HuMoCon: Concept Discovery for Human Motion Understanding Towards accurate active camera localization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:10.992617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:10.992617Z digest=sha256:d423a2ba87cc6efe5ead250175a2353badf4c663684af212f55c2fb20b3be2e3

Observation a96723f5-5384-4022-9933-3d07570c428b · outbound

This paper cites Cigtime: Corrective instruction generation through inverse motion editing.

HuMoCon: Concept Discovery for Human Motion Understanding Cigtime: Corrective instruction generation through inverse motion editing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:11.057582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:11.057582Z digest=sha256:374232b3fe9af4d11c88537f98c423f33ae1e666f9785d9f2bcc54a4b681bd9f

Observation ff860867-19ff-4e5c-bc38-fb68429e0236 · outbound

This paper cites Slowfast networks for video recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Slowfast networks for video recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.903825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:11.142433Z digest=sha256:148bb66ea76c94eb38511aaa4ffa74af8080d6fb56eb1b74193fe102d34c289c

Observation 3eac355a-cfcc-4cf5-911f-9df17de7310a · outbound

This paper cites Chatpose: Chatting about 3d human pose.

HuMoCon: Concept Discovery for Human Motion Understanding Chatpose: Chatting about 3d human pose

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.889083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:11.225495Z digest=sha256:1eb9e6b07530161720086dabb0e62d2e4f8bf21af3e5e6845369ac3c506f194e

Observation ea2b70de-3e15-47dc-bbb1-6c4ae6750b3a · outbound

This paper cites Aifit: Automatic 3d human-interpretable feedback models for fitness training.

HuMoCon: Concept Discovery for Human Motion Understanding Aifit: Automatic 3d human-interpretable feedback models for fitness training

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.871720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:11.320638Z digest=sha256:41eabee164be22422575268e98096094722ebd8c5f4c5f476b770e3e33e52176

Observation 20667788-0537-40e8-a2b5-447e05ac8a1c · outbound

This paper cites Imagebind: One embedding space to bind them all.

HuMoCon: Concept Discovery for Human Motion Understanding Imagebind: One embedding space to bind them all

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.742066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:11.391368Z digest=sha256:f743434b590f295dffb251f5a21a1e6ceede70fe5225ccab82704621c15643d4

Observation 51dfab67-7197-434c-aa1f-0087d77241be · outbound

This paper cites Ac- tion2motion: Conditioned generation of 3d human motions.

HuMoCon: Concept Discovery for Human Motion Understanding Ac- tion2motion: Conditioned generation of 3d human motions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.407672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:11.475504Z digest=sha256:44ee0d10b1d9fe2b1588a350315d0827f1c73311be11a487d22ab6c25864af6e

Observation 4997ba37-e887-4c0a-b5a8-33be0db7c361 · outbound

This paper cites Generating diverse and natural 3d human motions from text.

HuMoCon: Concept Discovery for Human Motion Understanding Generating diverse and natural 3d human motions from text

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:26.123330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:11.547308Z digest=sha256:4de28c4e21440a6c53ccbf52e3780cbfbabaa901063a3d496d98d9175b93c298

Observation e775fcd9-ebfd-439c-8d60-a8bcd4cb95e9 · outbound

This paper cites Generating diverse and natural 3d human motions from text.

HuMoCon: Concept Discovery for Human Motion Understanding Generating diverse and natural 3d human motions from text

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:25.893693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:11.636934Z digest=sha256:e2a652e03a0662e5fc553903e605fad7b70b0438aab1ba47184be3bb86946ab1

Observation e15425d3-6e93-4e05-81ba-f26ababa7d1c · outbound

This paper cites Tm2t: Stochastic and tokenized modeling for the reciprocal genera- tion of 3d human motions and texts.

HuMoCon: Concept Discovery for Human Motion Understanding Tm2t: Stochastic and tokenized modeling for the reciprocal genera- tion of 3d human motions and texts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:11.730244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:11.730244Z digest=sha256:cdd8c93fd60ee53d272d5f411fb61ca205a3d79bf4cf7dea89027c4a85196dc2

Observation b78b7213-7196-48e1-8d36-82aa6765d731 · outbound

This paper cites Contrastive learning from ex- tremely augmented skeleton sequences for self-supervised action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Contrastive learning from ex- tremely augmented skeleton sequences for self-supervised action recognition

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:25.653995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:11.815961Z digest=sha256:80d0ad11f4090b778148c7505a35ad48a06cd45fbbab2d26f976320cded5c9ab

Observation 29a51d22-1927-4a53-9eef-ab08a9b70ff6 · outbound

This paper cites Autoad ii: The sequel-who, when, and what in movie audio description.

HuMoCon: Concept Discovery for Human Motion Understanding Autoad ii: The sequel-who, when, and what in movie audio description

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:25.434143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:11.935438Z digest=sha256:3a6c2b0000ce9fbb25e1c8c90d04b2e91dca0b7341995bbffa6b238af71b7f40

Observation ba9114fe-6119-4eb3-9b85-ac1b458f5310 · outbound

This paper cites Autoad iii: The prequel-back to the pixels.

HuMoCon: Concept Discovery for Human Motion Understanding Autoad iii: The prequel-back to the pixels

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:25.144137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:12.033488Z digest=sha256:34707918234ae37c094c9fad67e7eb5a0e50261b82b0200a3e336a3a29c48c76

Observation b07742b7-366d-4efd-96cc-966dcf1311c4 · outbound

This paper cites Masked autoencoders are scalable vision learners.

HuMoCon: Concept Discovery for Human Motion Understanding Masked autoencoders are scalable vision learners

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.973255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:12.126802Z digest=sha256:ca434347ee55fd5b4cb2b4a3452ec83450b8c89aea17c17e82f3ecc3f0c82bc4

Observation b482f4d7-6ad4-4402-ae8c-d7f83a1888bb · outbound

This paper cites Phase- functioned neural networks for character control.ACM Transactions on Graphics (TOG), 36(4):1–13, 2017.

HuMoCon: Concept Discovery for Human Motion Understanding Phase- functioned neural networks for character control.ACM Transactions on Graphics (TOG), 36(4):1–13, 2017

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.817282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:12.239768Z digest=sha256:b10b7aa96b1b28f4abdf94965815ab0603525adfa3ab581f3da168888520782d

Observation 03b1b41f-5f04-46c5-9cc2-4d0979884ee3 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

HuMoCon: Concept Discovery for Human Motion Understanding LoRA: Low-Rank Adaptation of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:12.389567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:12.389567Z digest=sha256:372003f095a77919fa9c05f9ced439aeaeab0525c2e9cbc48c10129aa0cf1391

Observation 03f14e86-f0ea-4dca-9e3a-63bced0f1254 · outbound

This paper cites Motiongpt: Human motion as a foreign lan- guage.Advances in Neural Information Processing Systems, 36:20067–20079, 2023.

HuMoCon: Concept Discovery for Human Motion Understanding Motiongpt: Human motion as a foreign lan- guage.Advances in Neural Information Processing Systems, 36:20067–20079, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:12.495813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:12.495813Z digest=sha256:404e5c3b2bb4a1723bded6df805310dee1b4a80755b83c49b3f8b7969c91554d

Observation e528ae26-35a2-4910-b0f0-dbd5a8e40684 · outbound

This paper cites Hand-object contact consistency reasoning for human grasps generation.

HuMoCon: Concept Discovery for Human Motion Understanding Hand-object contact consistency reasoning for human grasps generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:12.639694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:12.639694Z digest=sha256:68280c8f35e14cddd42051cb34647812cb6610195aac776299d4a4914954f4a6

Observation 64e433a8-cb03-4d13-a276-bf17fe350259 · outbound

This paper cites SMPLX-Lite: A Realistic and Drivable Avatar Benchmark with Rich Geometry and Texture Annotations.

HuMoCon: Concept Discovery for Human Motion Understanding SMPLX-Lite: A Realistic and Drivable Avatar Benchmark with Rich Geometry and Texture Annotations

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:48:17.565477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:12.726047Z digest=sha256:8642343a015c4c2b45295cdcba7304e4f0aefde9d84f17fd9b294dbe5aaadbbb

Observation 9104fc61-8821-44f7-a4be-9cdafa246fb6 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

HuMoCon: Concept Discovery for Human Motion Understanding Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:12.856895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:12.856895Z digest=sha256:dcae7f8c08a5dd73a35b9c8186daa0e1de2b49cc6fd78ea548b382727160daab

Observation 208464c2-a052-4d9d-92b3-afff0104b3fe · outbound

This paper cites A new representation of skele- ton sequences for 3d action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding A new representation of skele- ton sequences for 3d action recognition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.707784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:12.907945Z digest=sha256:6fac1fa5b8c15dae7ade8172c31789f32c74b17460ce9823104503d61b119515

Observation 87faa928-de0b-4a98-b43c-6fe7b6f03427 · outbound

This paper cites Flame: Free- form language-based motion synthesis & editing.

HuMoCon: Concept Discovery for Human Motion Understanding Flame: Free- form language-based motion synthesis & editing

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:12.961097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:12.961097Z digest=sha256:b15be1a4e6135baf5e2f44348fe073ef7192c019b4f62cc55ee129bf4c221853

Observation ea5edf5d-c586-4251-ab3f-d4051cdb4253 · outbound

This paper cites Danceconv: Dance motion genera- tion with convolutional networks.IEEE Access, 10:44982– 45000, 2022.

HuMoCon: Concept Discovery for Human Motion Understanding Danceconv: Dance motion genera- tion with convolutional networks.IEEE Access, 10:44982– 45000, 2022

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.598734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:13.005451Z digest=sha256:6f1fdd8e5fc054cadb126b813669df013178582786e308096215382cea92b9c3

Observation 1721af8e-1829-47c5-9651-76f8e441ebc4 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

HuMoCon: Concept Discovery for Human Motion Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:13.085093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:13.085093Z digest=sha256:69c6795eebcbc879a9fdb3f0efdb0774de27402cbec926bed2f18a1a5eab9b67

Observation 57916962-adad-4b7f-9fd7-09b260c88146 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

HuMoCon: Concept Discovery for Human Motion Understanding VideoChat: Chat-Centric Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:13.144440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:13.144440Z digest=sha256:1e5b492048aeba30a7b857589f36fcdb063985f472deb09146f1cff125d445a6

Observation cd0b5d30-6ae9-4a28-bdd1-cbb6461c1914 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

HuMoCon: Concept Discovery for Human Motion Understanding Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.472147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:13.202566Z digest=sha256:8100222a8c4a20df9d3a6b955815f055ff4438f2279255653cd5803e52e5b2d7

Observation 16f9336b-0b54-4c04-8488-460a6540235f · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

HuMoCon: Concept Discovery for Human Motion Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:13.245692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:13.245692Z digest=sha256:830694cad90b4bcb4e2824e897fc551c0506847538de398db5a13cbc133d76a2

Observation 0b529021-9209-4e40-942b-db8c8cd4b16c · outbound

This paper cites Motion-x: A large- scale 3d expressive whole-body human motion dataset.Ad- vances in Neural Information Processing Systems, 36, 2024.

HuMoCon: Concept Discovery for Human Motion Understanding Motion-x: A large- scale 3d expressive whole-body human motion dataset.Ad- vances in Neural Information Processing Systems, 36, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.321591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:13.341764Z digest=sha256:a1ab5d473a362d957d2829bc0f87c693f861a4526bd486a8b8b42ac7760fa114

Observation e6a15de1-018e-4550-b7a3-1cafb400d427 · outbound

This paper cites Actionlet- dependent contrastive learning for unsupervised skeleton- based action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Actionlet- dependent contrastive learning for unsupervised skeleton- based action recognition

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:24.182622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:13.427954Z digest=sha256:6ff9689a1a99b30795d62ec27f87dba4a5ec3f1c0ec8053c8073f394b839c549

Observation e4c8e21f-edb1-4d66-b8df-79f53e57b561 · outbound

This paper cites Towards unified sur- gical skill assessment.

HuMoCon: Concept Discovery for Human Motion Understanding Towards unified sur- gical skill assessment

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.994813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:13.478624Z digest=sha256:13fe9c8d31a8c0fcfad22ba3d8d449fcb508047e26dbe7f315407a5ee622a43a

Observation c5dd6ce4-9680-4d59-b44d-fd72a6387429 · outbound

This paper cites Spatio-temporal lstm with trust gates for 3d human action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Spatio-temporal lstm with trust gates for 3d human action recognition

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.822773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:13.547011Z digest=sha256:34f7d95378c07f0043a36514dc8021c09bf2f37a12ce7846e56f243d807a5e6e

Observation 187ab89a-4c1d-4011-bfd7-2c4d3ab69860 · outbound

This paper cites Enhanced skele- ton visualization for view invariant human action recogni- tion.Pattern Recognition, 68:346–362, 2017.

HuMoCon: Concept Discovery for Human Motion Understanding Enhanced skele- ton visualization for view invariant human action recogni- tion.Pattern Recognition, 68:346–362, 2017

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.683090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:13.607043Z digest=sha256:30ea5bcc4310d5939e3418a761ef6a12ac83c6b2096e1d94d398ef6cd459e6de

Observation d90d895f-741d-4b2e-8605-da08a2e85ae3 · outbound

This paper cites InfoCon: Concept Discovery with Generative and Discriminative Informativeness.

HuMoCon: Concept Discovery for Human Motion Understanding InfoCon: Concept Discovery with Generative and Discriminative Informativeness

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:48:17.297953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:13.659969Z digest=sha256:6627e3c10a4e4dee6e659f7e219248c805384523ede602559d9215f451a08a5e

Observation 00d2d0fb-ce91-421d-9212-bbf7b32180b8 · outbound

This paper cites Posegpt: Quantization-based 3d human mo- tion generation and forecasting.

HuMoCon: Concept Discovery for Human Motion Understanding Posegpt: Quantization-based 3d human mo- tion generation and forecasting

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.435492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:13.714895Z digest=sha256:9c1e8fb8d48b4d9a56ba5360c25ef5ae2901c0408ad9578b08f255d60ad21f90

Observation 8f552295-2322-4b0f-8068-6be3fc8388ce · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neu- rocomputing, 508:293–304, 2022.

HuMoCon: Concept Discovery for Human Motion Understanding Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neu- rocomputing, 508:293–304, 2022

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.253372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:13.805257Z digest=sha256:d3c61d6f2af58c86fac97f26f935fbefdc147fcf4fb766fe83376639e4dd3e19

Observation 28359c28-23f9-4524-89dc-45a73fa2755f · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

HuMoCon: Concept Discovery for Human Motion Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:13.881665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:13.881665Z digest=sha256:8f6ab7fad067d33298573e9b0487f24c363c8cc7d4f78c0ecec720c85a0bb075

Observation d5e98ac8-6eb9-4368-ae87-f5f683df8823 · outbound

This paper cites PG-Video-LLaVA: Pixel Grounding Large Video-Language Models.

HuMoCon: Concept Discovery for Human Motion Understanding PG-Video-LLaVA: Pixel Grounding Large Video-Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:13.947245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:13.947245Z digest=sha256:96e354b1c3ab914317cd06b0784a7a6185566816b5c06d5361b50ef241acc93e

Observation 9b0fb8e3-4431-46bb-a5b7-380fcf050f4d · outbound

This paper cites What and how well you performed? a multitask learning approach to action quality assessment.

HuMoCon: Concept Discovery for Human Motion Understanding What and how well you performed? a multitask learning approach to action quality assessment

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:23.136730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:13.998154Z digest=sha256:dd3ad9a71061fb7fe567c6a71074d2353d2a9b609cd53d0812260af89d550969

Observation c3457ee8-ab87-472d-853d-dca2a07d5498 · outbound

This paper cites Action- conditioned 3d human motion synthesis with transformer vae.

HuMoCon: Concept Discovery for Human Motion Understanding Action- conditioned 3d human motion synthesis with transformer vae

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:22.873624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:14.050854Z digest=sha256:2f2ef9fd32e499aca30ccee0401f611cb4514aa6fbbec6268142aedf3bc73396

Observation da2de4c7-d835-4237-985d-8b695a69468b · outbound

This paper cites Temos: Generating diverse human motions from textual descriptions.

HuMoCon: Concept Discovery for Human Motion Understanding Temos: Generating diverse human motions from textual descriptions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:14.094528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:14.094528Z digest=sha256:3c11d1e05f7afedafdd111a1a179a924e77e4774828880f8cccea1e72b60113a

Observation 7a7255e9-6b35-4a8c-b1d6-d10d5156d21b · outbound

This paper cites Babel: Bodies, action and behavior with english la- bels.

HuMoCon: Concept Discovery for Human Motion Understanding Babel: Bodies, action and behavior with english la- bels

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:22.607481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:14.138546Z digest=sha256:b869ff580a33841fa8be57059a5d49f0817a4adc477ebe521063dd2f5093d80e

Observation 65f19c98-0b95-4647-bcf4-e2983ec9d777 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

HuMoCon: Concept Discovery for Human Motion Understanding Learning transferable visual models from natural language supervi- sion

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:22.425341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:14.215738Z digest=sha256:d9f8993bc9f09222d76bd13f9d4073f916e2c1368dce29f5e0c876ee0eaec931

Observation 831fc922-753d-41b5-aa46-34fb4338c55c · outbound

This paper cites On the benefits of 3d pose and tracking for human action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding On the benefits of 3d pose and tracking for human action recognition

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:22.196705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:14.279679Z digest=sha256:ad7d0a7b8947c72f6f61a28db6350bdb6e887d46dec6051c44e38a058cd02c82

Observation 7fe7f90a-a16d-4072-87f5-c4acfd15452d · outbound

This paper cites Zero-shot audio captioning with audio-language model guidance and audio context keywords.

HuMoCon: Concept Discovery for Human Motion Understanding Zero-shot audio captioning with audio-language model guidance and audio context keywords

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:14.346070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:14.346070Z digest=sha256:d28318f1a2b1515e39f39749341b38ac20592763acb8ad9c68a96afe3b0c27a4

Observation 8c342166-b30e-4688-bdf9-2c447fdf3826 · outbound

This paper cites Two- stream adaptive graph convolutional networks for skeleton- based action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Two- stream adaptive graph convolutional networks for skeleton- based action recognition

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:21.862099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:14.405018Z digest=sha256:de3e484b9bb69ec89e384aab93728482945299b4c941ae24a833d9d9b7ec9ae8

Observation 2b58ca6e-3251-406b-bc50-9d004b6fdca0 · outbound

This paper cites Neural state machine for character-scene interactions.ACM Transactions on Graphics, 38(6):178, 2019.

HuMoCon: Concept Discovery for Human Motion Understanding Neural state machine for character-scene interactions.ACM Transactions on Graphics, 38(6):178, 2019

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:21.647262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:14.464679Z digest=sha256:c17aee0461ebac0475077b68a364045783e658e30f060bd087252c9f01220d4e

Observation 3af641c8-f034-457d-b777-cafcd8bcf05a · outbound

This paper cites Local motion phases for learning multi-contact charac- ter movements.ACM Transactions on Graphics (TOG), 39 (4):54–1, 2020.

HuMoCon: Concept Discovery for Human Motion Understanding Local motion phases for learning multi-contact charac- ter movements.ACM Transactions on Graphics (TOG), 39 (4):54–1, 2020

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:14.523931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:14.523931Z digest=sha256:31e1aca738e09a0b544d9efdcf514ef131fe1aed696dbd81b89cdbf8c67078cb

Observation 6663db44-167f-4f5f-923c-045a26aecb86 · outbound

This paper cites Deepphase: Periodic autoencoders for learning motion phase manifolds.

HuMoCon: Concept Discovery for Human Motion Understanding Deepphase: Periodic autoencoders for learning motion phase manifolds

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:21.436539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:14.601668Z digest=sha256:40fa3f69bf988264ab05c7cd78e2ee5063d6f55d4f800562d8d69cd72521b390

Observation d06102dc-abc2-48cb-9248-6985586b7706 · outbound

This paper cites Convolutional learning of spatio-temporal features.

HuMoCon: Concept Discovery for Human Motion Understanding Convolutional learning of spatio-temporal features

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:21.236498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:14.704401Z digest=sha256:2fbf09877428d50052f2589e85b4472c91d01ca1c4b71cfbe6be91fffb22af94

Observation d1fe6ca9-1bbc-405f-873d-1e16e0da08c3 · outbound

This paper cites Motionclip: Exposing human motion generation to clip space.

HuMoCon: Concept Discovery for Human Motion Understanding Motionclip: Exposing human motion generation to clip space

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:21.032251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:14.788254Z digest=sha256:48cf95906a3223e6531cfe22e90f1c8f0715369de61a0db86e9348308ac3dbd9

Observation 5e1ffcd6-416e-4ea7-bf77-1091569b3774 · outbound

This paper cites Human Motion Diffusion Model.

HuMoCon: Concept Discovery for Human Motion Understanding Human Motion Diffusion Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:14.844576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:14.844576Z digest=sha256:7f49035b862f2cd52251c8f809fda573f8e1d4277f37235631533d7f9915ffa1

Observation 3c5815f4-3f2d-4da3-bc78-a8f56306f38c · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

HuMoCon: Concept Discovery for Human Motion Understanding Learning spatiotemporal features with 3d convolutional networks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:14.909199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:14.909199Z digest=sha256:d4b3579b23ce54494882b54dd61213e0331a1abbdf4b3f59a9f10ab105154fef

Observation 56c8bae8-3a4c-453b-a1be-7e88c2a6ed36 · outbound

This paper cites Human action recognition by representing 3d skeletons as points in a lie group.

HuMoCon: Concept Discovery for Human Motion Understanding Human action recognition by representing 3d skeletons as points in a lie group

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:20.873807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:14.987206Z digest=sha256:78aba270216258eb94650c1995d6d545ece46d812e9f7a88ed870d0011bad6d3

Observation 06f1a0dd-c61b-4c52-9b0c-a9bd360d9b94 · outbound

This paper cites Action recognition with improved trajectories.

HuMoCon: Concept Discovery for Human Motion Understanding Action recognition with improved trajectories

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:20.555895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:15.061942Z digest=sha256:3696aae95ffb692b580552e644d9d542e96d7c6ccf24e16f19bc82ac6aad45be

Observation a3db810f-40c1-4d18-ab05-227898ea01bf · outbound

This paper cites Learn- ing human dynamics in autonomous driving scenarios.

HuMoCon: Concept Discovery for Human Motion Understanding Learn- ing human dynamics in autonomous driving scenarios

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:20.420267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:15.133096Z digest=sha256:ddc20607c7db124f03251d7ba4e325ee8e6ac865a326c841004202288aa9725e

Observation 59f471f3-70f9-4451-b87b-921b62699906 · outbound

This paper cites Vatex: A large-scale, high- quality multilingual dataset for video-and-language research.

HuMoCon: Concept Discovery for Human Motion Understanding Vatex: A large-scale, high- quality multilingual dataset for video-and-language research

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:20.259990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:15.244074Z digest=sha256:def82c0dfeced149ab4e98a731c410fdf347deb49a74ca212a7a2e6129e7d2a7

Observation 8401b309-0b16-4b9d-ac1b-30f5586f3016 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

HuMoCon: Concept Discovery for Human Motion Understanding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.328435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.328435Z digest=sha256:b92090b5ce7dc582613e43edf691f38711b1f8c9a59e5188f0e80c8d0a592ee7

Observation c51851bd-3b04-4f8f-9f48-14366c20f7cd · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

HuMoCon: Concept Discovery for Human Motion Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.401308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.401308Z digest=sha256:4a287db5af7607e79a3e807e65c079bb85e6452a1975842582c74cf2f8ec71cb

Observation b6deda1a-9d71-425a-98b1-b0e2b25f293a · outbound

This paper cites Unified Human-Scene Interaction via Prompted Chain-of-Contacts.

HuMoCon: Concept Discovery for Human Motion Understanding Unified Human-Scene Interaction via Prompted Chain-of-Contacts

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.499056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.499056Z digest=sha256:accefe1d95bc4ad0597794f3cb0ecaf8b845764f4c49239b7ab484d59f26aeab

Observation be953bc4-6c03-4eb8-8487-f0ea8c17fd8c · outbound

This paper cites AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description.

HuMoCon: Concept Discovery for Human Motion Understanding AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.573298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.573298Z digest=sha256:6b24b8ff43e96d6f60523237711304dabcc8c60a83cde13e71ecbe2989537d9d

Observation d007fada-1dd1-47bc-89fb-6af1778e96e1 · outbound

This paper cites Dexterous grasp transformer.

HuMoCon: Concept Discovery for Human Motion Understanding Dexterous grasp transformer

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:20.011797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:15.647947Z digest=sha256:3da70025600e86043b6385da24d0c590b80020b636faabdf4b0567e1473467fa

Observation f76756c4-f32c-49d6-9d50-f76b9c548aa5 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

HuMoCon: Concept Discovery for Human Motion Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.720080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.720080Z digest=sha256:9e5b406bf8694eebe47a368f6f6e6435e0e747a3e65e45b9578e240081c0ae6f

Observation 10c3c04d-cff4-4210-8c7f-e449cfe56b35 · outbound

This paper cites Spatial tempo- ral graph convolutional networks for skeleton-based action recognition.

HuMoCon: Concept Discovery for Human Motion Understanding Spatial tempo- ral graph convolutional networks for skeleton-based action recognition

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:15.777468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:15.777468Z digest=sha256:4920677e974c2c3415d88ab76df2d940e9c39b9c072818ce9e345f92272c9b76

Observation 37fd70ca-82b3-4a2e-8099-524b47f3f24d · outbound

This paper cites Learning to use chopsticks in diverse gripping styles.

HuMoCon: Concept Discovery for Human Motion Understanding Learning to use chopsticks in diverse gripping styles

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:19.789211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:15.830957Z digest=sha256:17aecd4ac49f1608aac3f42c81b0f19ed685da0738530469f4e2918208cd0045

Observation d7b03f79-b05b-4883-8c10-c6785a696709 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

HuMoCon: Concept Discovery for Human Motion Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:19.523041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:15.910767Z digest=sha256:bfcc241ba37f1884d8e38d0207277b1b01496cae1f84ef7c38e56623460c9e31

Observation d21bf28c-63a4-4cd6-ac0b-94068ef65ac2 · outbound

This paper cites Sigmoid loss for language image pre-training.

HuMoCon: Concept Discovery for Human Motion Understanding Sigmoid loss for language image pre-training

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:19.278946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:15.969101Z digest=sha256:e7f7fa85db6f48c1c733e36df6973ff0339657c993a87003903c9cfd108af744

Observation 5c7172aa-8f9f-46c2-bc77-e4df44dd84ca · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

HuMoCon: Concept Discovery for Human Motion Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:16.029474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:16.029474Z digest=sha256:f1ff03bafa114dfc40342a9a59aa6718c519ae2252d18c870a976ee55ececc81

Observation 47b9ecd6-67b5-43ac-b2b8-7def70600694 · outbound

This paper cites T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations.

HuMoCon: Concept Discovery for Human Motion Understanding T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:16.105127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:16.105127Z digest=sha256:d15fae47294e1c7dc4e961d8bb69e43dc0694cbc66dfd0f7ba1d6ed82b6345d4

Observation 9d0d6c33-8e81-444f-b42b-1d09f4fc4fd7 · outbound

This paper cites Motiondif- fuse: Text-driven human motion generation with diffusion model.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.

HuMoCon: Concept Discovery for Human Motion Understanding Motiondif- fuse: Text-driven human motion generation with diffusion model.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:19.045164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:16.159016Z digest=sha256:897f07de69775b7024deedbdd1bf13a73a66ee109337ed312c1b425dc4adfaa3

Observation 54025af5-d1e1-441b-94d3-973e27b5fe10 · outbound

This paper cites Pointclip: Point cloud understanding by clip.

HuMoCon: Concept Discovery for Human Motion Understanding Pointclip: Point cloud understanding by clip

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:18.857074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:16.224869Z digest=sha256:0c00ddfbe3e76a257148bb08277503d3c4df99619608e1204e953c6544cab7b8

Observation 3d7d98d9-8c2b-4cdc-943d-70a929f975b8 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

HuMoCon: Concept Discovery for Human Motion Understanding LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:16.292122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:16.292122Z digest=sha256:76e6dc95c0b93661d566ed8b88a0ed87e4c287e8cc586ad5ef92a935a01dff34

Observation 7d80e39c-03c8-4383-9fbd-44a6bb6f78f5 · outbound

This paper cites Streaming dense video captioning.

HuMoCon: Concept Discovery for Human Motion Understanding Streaming dense video captioning

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:18.710966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:16.441949Z digest=sha256:b082e28cb1a227e9a845f201cd9f88e5eedcd95f0837bfebae31cb2ba14d5c8e

Observation a48bdeba-b3a4-4799-b1b4-30ab37674bb7 · outbound

This paper cites Avatargpt: All- in-one framework for motion understanding planning gener- ation and beyond.

HuMoCon: Concept Discovery for Human Motion Understanding Avatargpt: All- in-one framework for motion understanding planning gener- ation and beyond

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:18.526788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:16.539175Z digest=sha256:d478ba86b9850565ab4ab07590a0429193e9e9917dcdd61ad0b48dc052a16345

Observation e2d842ec-0c37-4a27-a756-081035a8de30 · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

HuMoCon: Concept Discovery for Human Motion Understanding LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:16.643392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:48:16.643392Z digest=sha256:003274f7cb49fe21c7a7b99b8a7d1814111d698c6c8fce16a0c677fc7aefcbed

Observation 15700c2e-c4d9-429e-96b6-bb2ca64ecb9e · outbound

This paper cites Limited by GPU memory, we only take 8 key frames for each video following MotionLLM [5].

HuMoCon: Concept Discovery for Human Motion Understanding Limited by GPU memory, we only take 8 key frames for each video following MotionLLM [5]

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:18.383910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:16.773797Z digest=sha256:98e5a79f55c23a6f91f10efb58b0b52e0a1783dfad28fff58324c63030740233

Observation d23a2469-1fb5-4239-8f6e-29726d222219 · outbound

This paper cites Visualizations for Motion Understanding Additional visualization results for motion understanding are provided in Fig.

HuMoCon: Concept Discovery for Human Motion Understanding Visualizations for Motion Understanding Additional visualization results for motion understanding are provided in Fig

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:18.128667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:16.873059Z digest=sha256:c57dcf893356cfdd243f02d09020e9b6dd088774356764f5c0d47aadfff70767

Observation 20a2d8f4-4502-4b51-8466-9ff35e6c1ca7 · outbound

This paper cites Evaluation on BABEL-QA Benchmark For the BABEL-QA benchmark, we extend the evalua- tion protocol used in previous multi-modality LLMs evalu- ations [80].

HuMoCon: Concept Discovery for Human Motion Understanding Evaluation on BABEL-QA Benchmark For the BABEL-QA benchmark, we extend the evalua- tion protocol used in previous multi-modality LLMs evalu- ations [80]

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:48:17.934458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:16.966088Z digest=sha256:905566adb286a1a608ff5569b1da3f78ea071f0eaccf8a626541e444f0d44360

Observation e97ab098-55cd-4ba2-9b0f-dcf258105aed · outbound

This paper cites Structure of Hyper-Networks We utilize hyper-networksH u andH m, for velocity recon- struction, following InfoCon [46].

HuMoCon: Concept Discovery for Human Motion Understanding Structure of Hyper-Networks We utilize hyper-networksH u andH m, for velocity recon- struction, following InfoCon [46]

Reference 91

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:48:17.751549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:48:17.043562Z digest=sha256:c789c5181c1e0a5d789ba77c9a00f1f8712cea61d14765f6f4accbd48e632cea

Pith citing papers

No inbound Pith citation observations are available.