Pith. sign in

Paper Citation Record · LEDGER

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision

As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2608.01392.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01392 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:18:12.550250Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1d8eb615-bb3a-4df3-ba8c-b07aaedb640b · outbound

This paper cites Vivit: A video vision transformer.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Vivit: A video vision transformer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:09.624279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:09.624279Z digest=sha256:ce8d952840632767207aa53d1b7d4e8a464606b6d1eb815fae49765fb67ebef0

Observation 1e3ab1af-bc8b-4f09-ae93-eefb8c2ba092 · outbound

This paper cites Seamless human motion composition with blended posi- tional encodings.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Seamless human motion composition with blended posi- tional encodings

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:18.524741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:09.694111Z digest=sha256:6e8a0dd6ef004cb10142f7c4d686d7365bbb6fc0658156d59489d60f0b060dc6

Observation 729720f7-9dcb-4e37-bb4f-45cc441c4bbc · outbound

This paper cites A cross- dataset study for text-based 3d human motion retrieval.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision A cross- dataset study for text-based 3d human motion retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:18.350693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:09.761688Z digest=sha256:49db1d07e053c65e40dcbd439320bb948ea8153f14a086a85ba6e549650bd3fc

Observation 3f7840b3-31ac-44ba-a1e3-a313f6891a9d · outbound

This paper cites Is space-time attention all you need for video understanding? InIcml, page 4, 2021.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Is space-time attention all you need for video understanding? InIcml, page 4, 2021

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:09.806437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:09.806437Z digest=sha256:cc3a435a77413df18734276c642d096926b9b1d2a8de60d72990ef87a6b4a262

Observation 110b8fd9-c88c-4d7b-a86f-4d4abafa473c · outbound

This paper cites Sinkhorn distances: Lightspeed computation of optimal transport.Advances in neural information pro- cessing systems, 26, 2013.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Sinkhorn distances: Lightspeed computation of optimal transport.Advances in neural information pro- cessing systems, 26, 2013

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:18.086343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:09.916222Z digest=sha256:53291e07743560e8ef4aa867d0d503c7c8294f28e1d897f7eaf7554c93a9b93a

Observation e1df4d46-a53f-4f95-8e51-83ea6a6c04b6 · outbound

This paper cites Segmo: Segment-aligned text to 3d human motion generation.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Segmo: Segment-aligned text to 3d human motion generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:17.845673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:10.011811Z digest=sha256:8495a1c26dc9e4882cb76ebb74ac357ee6310f1779a8885475467ad9ed47cf53

Observation 2d8e775e-2337-449d-bacd-e94edbb17081 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.Advances in neural information processing systems, 35:35946–35958,.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Masked autoencoders as spatiotemporal learners.Advances in neural information processing systems, 35:35946–35958,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:10.101270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:10.101270Z digest=sha256:56146fb7e27151e9fe62ed76bd99fb364165dac26a29026ea128c4deef5eea2b

Observation 9396af1c-2063-4a0b-af53-c00773747281 · outbound

This paper cites Generating diverse and natural 3d human motions from text.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Generating diverse and natural 3d human motions from text

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:17.689155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:10.184506Z digest=sha256:21af8676d371a229111bc53cd0a294989174e4d5b999818620283ad6cae23ae6

Observation a814d231-cf1b-4a36-a09b-b2ee7c0b31b5 · outbound

This paper cites Momask: Generative masked model- ing of 3d human motions.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Momask: Generative masked model- ing of 3d human motions

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:17.448771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:10.252169Z digest=sha256:7e6c593cad8062bd06fe499e917e604a3d0f5e2889fc9d6b0641df3d64ea55d1

Observation 92145b61-a539-4f5f-bc8d-c4f3fe4662e8 · outbound

This paper cites Snapmogen: Human motion generation from expressive texts.arXiv preprint arXiv:2507.09122, 2025.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Snapmogen: Human motion generation from expressive texts.arXiv preprint arXiv:2507.09122, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:10.316066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:10.316066Z digest=sha256:34616e435f53e41895da21601bfb69d3033d43cc288068b432cc5c79ce02809f

Observation dd972f82-f165-447b-93fd-bb7817a4e9fc · outbound

This paper cites Amd: Autoregressive motion diffusion.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Amd: Autoregressive motion diffusion

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:17.223219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:10.394699Z digest=sha256:6bbf9e7168cdcc5b3b5ed523a3643a808a1490a053da908893c6dd555d84311e

Observation da7653cb-97cf-45b6-9b9b-a0f7481e01ed · outbound

This paper cites Como: Controllable motion generation through language guided pose code edit- ing.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Como: Controllable motion generation through language guided pose code edit- ing

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:17.019348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:10.457611Z digest=sha256:408431a27e84d4e2b117621e58e9db9e4ea6529258a0a37ee9db72cbfb664eb3

Observation 4fbfc23a-42ea-4035-9321-2186955508ae · outbound

This paper cites Motiongpt: Human motion as a foreign lan- guage.Advances in Neural Information Processing Systems, 36:20067–20079, 2023.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Motiongpt: Human motion as a foreign lan- guage.Advances in Neural Information Processing Systems, 36:20067–20079, 2023

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:16.826747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:10.511210Z digest=sha256:3af697b5f8e30b84727c86bdabce57d78edf5ac5eb2f3ba93c8530e613d7dbcf

Observation e9676649-5f43-4969-8a10-99ff725ef267 · outbound

This paper cites Unimotion: Unifying 3d human motion synthesis and understanding.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Unimotion: Unifying 3d human motion synthesis and understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:16.656878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:10.609319Z digest=sha256:1d5d3fb2fb4cba80560a388a99d1c18d0fec182759d2eae123d3524ec840ff3f

Observation b6e88d32-1981-45cc-928b-82dd243fa9dd · outbound

This paper cites Frankenmotion: Part-level hu- man motion generation and composition.arXiv preprint arXiv:2601.10909, 2026.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Frankenmotion: Part-level hu- man motion generation and composition.arXiv preprint arXiv:2601.10909, 2026

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:10.686875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:10.686875Z digest=sha256:8bce94c8b2b761d46837ef29bf08957de2af76b39224fc2a9f026eb7af6515d4

Observation 0baaaaff-ba9b-4d34-9537-72e342d8d4a2 · outbound

This paper cites Motion-x: A large-scale 3d expressive whole-body human motion dataset.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Motion-x: A large-scale 3d expressive whole-body human motion dataset

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:16.419404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:10.779221Z digest=sha256:9aea776763a5d1595133f58874eefd9c5054bc684b72a61bcf8d7db270584925

Observation cabcb29f-76bb-468c-98c7-20dadfc5b32b · outbound

This paper cites Multi-granularity Correspondence Learning from Long-term Noisy Videos.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Multi-granularity Correspondence Learning from Long-term Noisy Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:10.857069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:10.857069Z digest=sha256:744f9c27884d3eb1e617efa97dd2b5c1733768738ef18334a60bc9e7ab3b50c5

Observation fc20a501-8fd9-45ba-8940-7ee9aff243e5 · outbound

This paper cites Rethinking diffusion for text-driven human motion generation: Redundant representations, evaluation, and masked autoregression.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Rethinking diffusion for text-driven human motion generation: Redundant representations, evaluation, and masked autoregression

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:16.238937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:10.885666Z digest=sha256:f16884992fa624c0aee7e7f27f238b07dc0dc0324b0155e61c03e68b202dea0f

Observation 424881ac-e4b6-4324-a00e-6e48c5cc6a0d · outbound

This paper cites Proposal-free temporal action detection via global segmen- tation mask learning.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Proposal-free temporal action detection via global segmen- tation mask learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:16.000195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:10.925323Z digest=sha256:e961ae5821f37187adcf8eb39ca77897af719c76e3353e52f8a707d3c54b766c

Observation ecee7a8d-abd9-4d16-b5e1-d65c7a3ea633 · outbound

This paper cites Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:15.705738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:10.986515Z digest=sha256:43b9137edf87731e64f09574106bad8e843ad9f48905b14a9dc2a6414dd048a1

Observation 7fbdb2c5-6e48-471e-97f2-46202e197982 · outbound

This paper cites Now Foun- dations and Trends, 2019.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Now Foun- dations and Trends, 2019

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:15.475588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:11.034615Z digest=sha256:42ae71fb6609fc8f7e7e0baf90fd9b3aa5855e778e27a384acbc73efd9a243fe

Observation 9a584c9d-c0dd-48e8-a5ec-be2e508a0d10 · outbound

This paper cites Mmm: Generative masked motion model.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Mmm: Generative masked motion model

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:15.162114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:11.074653Z digest=sha256:38bf5d1113811f6ab0ac4e30675aa9467e1f26181116a5f5baf5ebfd1e416901

Observation b977c382-ccef-4a8b-8949-22bc2cc3fab5 · outbound

This paper cites The kit motion-language dataset.Big data, 4(4):236–252,.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision The kit motion-language dataset.Big data, 4(4):236–252,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.165073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.165073Z digest=sha256:571424d5b1a48f36f000e465a89bc5646f4160169f336758c2bdd1d14bd08ca4

Observation 5d8d8fc0-93e7-4e91-b1d5-9a10e7a9f51a · outbound

This paper cites Babel: Bodies, action and behavior with english la- bels.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Babel: Bodies, action and behavior with english la- bels

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.986967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:11.246845Z digest=sha256:81157953dc95fc12c9c7ab487883c553ee2e61ba3807ac018e5a4b85c0b9fd16

Observation 89c6b694-6741-4344-9dcb-dafa476b1c7c · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Learning transferable visual models from natural language supervi- sion

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.279566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.279566Z digest=sha256:c3fcad57c9b968d429e87c3dcdaa46015909560b7bfe14f7165135a42d852e39

Observation b11f5059-ada5-4929-8cb5-585a11dee210 · outbound

This paper cites an unresolved cited work.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.365326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.365326Z digest=sha256:7260f95d9c95e69c776a98526cc06c85f6e83afd6579a01b53349839ba65af2b

Observation 232ce928-4530-4ff9-bb09-b7e511c4dfed · outbound

This paper cites Ot-clip: Un- derstanding and generalizing clip via optimal transport.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Ot-clip: Un- derstanding and generalizing clip via optimal transport

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.742352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:11.436412Z digest=sha256:2a8110104fc0c027f477a2ead38c460bb4f1f4f36aed8b175dfbf1d9a7c2e06a

Observation 6b0e7d07-a54a-420e-9339-8cf55db0588d · outbound

This paper cites OpenAI GPT-5 System Card.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision OpenAI GPT-5 System Card

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.477503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.477503Z digest=sha256:c9fa5e6bee63f57b25619703d115b5807e4765a02dabb1db1605b7b1858c6d74

Observation ee65441b-8813-4b88-8777-29ae06394ff9 · outbound

This paper cites Optimal transport on discrete domains.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Optimal transport on discrete domains

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.580551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:11.529984Z digest=sha256:7cca1ff052d615124ee2dfb1e3db7e155f538b49e70b5e543589576b2500d2fa

Observation c6481434-0f9e-4afa-9cc2-cdf53114e0fb · outbound

This paper cites CoMA: Compositional Human Motion Generation with Multi-modal Agents.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision CoMA: Compositional Human Motion Generation with Multi-modal Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.588885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.588885Z digest=sha256:6a2ba078ba83c3e8b6c335559a0f71f3577d7825d41dd2cc77b51f915f8909fd

Observation 3ca679e8-3bff-47bc-a24a-7f0a4d873767 · outbound

This paper cites Motionclip: Exposing human motion generation to clip space.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Motionclip: Exposing human motion generation to clip space

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.464800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:11.647315Z digest=sha256:fdb0cdabd8c82d4f6ec5ea35f76cb286d106004c2eecdc3cd804bb4cd4987a40

Observation fcaeb8bd-9c66-4074-ae86-5bc2543f89f5 · outbound

This paper cites Human Motion Diffusion Model.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Human Motion Diffusion Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.728862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.728862Z digest=sha256:d7b47313629a7e39f0a2989b83c1d4a3e5532826075a8d370ecb09092fcf7f60

Observation 3f702020-fa09-4a3d-a137-91d7442beead · outbound

This paper cites Scaling Large Motion Models with Million-Level Human Motions.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Scaling Large Motion Models with Million-Level Human Motions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.792396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.792396Z digest=sha256:d9f215b4f27dd6cdc437b151d9fdfccad6d454b6aff0cf5ed6692f31046fa84f

Observation f40afea6-d756-4ab6-bf7a-4babacb90ddc · outbound

This paper cites Mg-motionllm: A unified framework for motion comprehension and gener- ation across multiple granularities.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Mg-motionllm: A unified framework for motion comprehension and gener- ation across multiple granularities

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.206099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:11.853616Z digest=sha256:f50e95c2a58875276f8ef5cf3b64b2f03c2c88cbe719b91649cdb4b206748416

Observation 65b39cf5-8667-4a57-afbd-b6db884a8c0c · outbound

This paper cites Dense motion captioning.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Dense motion captioning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.014695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:11.903996Z digest=sha256:b8abb26f878494be1211482fda12e20083f9725d96ad240c684ecb51c9f61a99

Observation f8392932-71c2-4c4c-acea-841098fe600c · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.969697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.969697Z digest=sha256:13148476af2de19bdfd1b52807c0c5473a0586509dba68d516179058177a779f

Observation c779391c-f27a-4d81-80ab-4fbcf11f13e8 · outbound

This paper cites Generating human motion from textual descrip- tions with discrete representations.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Generating human motion from textual descrip- tions with discrete representations

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:13.891286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:12.034560Z digest=sha256:714e28062ef1cce19f62c47be8c586b04b0e6d5c133470190a6cfb1edfbebd72

Observation f3756519-1f9f-4f69-b1dc-2a1634116f49 · outbound

This paper cites Re- modiffuse: Retrieval-augmented motion diffusion model.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Re- modiffuse: Retrieval-augmented motion diffusion model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:12.118757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:12.118757Z digest=sha256:e192f3c75a46551bd4ce95a70e135034f84a39f33db69c9e2158b27bde9b6db3

Observation b402e152-3973-455d-932c-1835a5056a43 · outbound

This paper cites Finemogen: Fine-grained spatio- temporal motion generation and editing.Advances in Neural Information Processing Systems, 36:13981–13992, 2023.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Finemogen: Fine-grained spatio- temporal motion generation and editing.Advances in Neural Information Processing Systems, 36:13981–13992, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:13.726895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:12.173471Z digest=sha256:b6d7949787e126d90c07bdfa17f728d364bdc577f30f9ac16a448b98f32f0a0e

Observation 8c455f42-98d8-4d38-b96b-1abb68d58607 · outbound

This paper cites Pre- training clip against data poisoning with optimal transport- based matching and alignment.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Pre- training clip against data poisoning with optimal transport- based matching and alignment

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:13.640874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:12.264675Z digest=sha256:9ac5fd5c1eeca17202267317ab08e67d8244cd03a2e88138f6176472b116d498

Observation 0973253c-3618-4620-86b4-6fcc9b32012d · outbound

This paper cites DartControl: A diffusion-based autoregressive motion model for real-time text-driven motion control.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision DartControl: A diffusion-based autoregressive motion model for real-time text-driven motion control

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:13.488133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:12.328454Z digest=sha256:ad52e56ca5cf1a1b78cebbf934c088898836b005b0cdf553d26a385e36b76fab

Observation 79f73070-d566-43e8-923e-878b1e21a7e2 · outbound

This paper cites an unresolved cited work.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:18:13.339249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:12.375527Z digest=sha256:233dc4249d1fc99a2a6099b4eee649044b92715fe78579cb1cc304888914b62d

Observation a6f6d641-ed17-44c8-8ab9-3f0f10634fdb · outbound

This paper cites an unresolved cited work.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:18:13.215087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:12.439527Z digest=sha256:7ab2ae86d1eba233e7caa6e83617a77d3472f2aebd74ad298d5de2a1e940560b

Observation 1845dedb-1e0b-4bd6-b0ef-58b78ec6c43d · outbound

This paper cites The annotation process involves man- ual segment-level boundary identification: for each action segmentj, human annotators specify the temporal interval [tstart, tend]in seconds.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision The annotation process involves man- ual segment-level boundary identification: for each action segmentj, human annotators specify the temporal interval [tstart, tend]in seconds

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:13.060514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:12.510219Z digest=sha256:3f3705d16818ccb563973ff33cba4dae270cb7276a6f9d4df4d929d65d2a51bf

Observation db1aac17-a015-43c4-a123-5f27eda79b80 · outbound

This paper cites specialized human motion analyst.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision specialized human motion analyst

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:12.924197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:18:12.550250Z digest=sha256:fb2dae17e95134ecf93d201647c3f407f63c0886f1722a4cd8aebaf0f8d6f492

Pith citing papers

No inbound Pith citation observations are available.