Pith. sign in

Paper Citation Record · LEDGER

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision

As of 23 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2608.01392.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01392 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:18:12.550250Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1d8eb615-bb3a-4df3-ba8c-b07aaedb640b · outbound

This paper cites Vivit: A video vision transformer.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Vivit: A video vision transformer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:09.624279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:09.624279Z digest=sha256:bad862947e73c64a1cd564fed61cd310bcf2347d63986f4be469057a6a078617

Observation 1e3ab1af-bc8b-4f09-ae93-eefb8c2ba092 · outbound

This paper cites Seamless human motion composition with blended posi- tional encodings.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Seamless human motion composition with blended posi- tional encodings

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:18.524741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:09.694111Z digest=sha256:eeb9ce7251bd1f8908798bbeb05ef225126530bfdbe9e5816915d2b9e6b2b2e5

Observation 729720f7-9dcb-4e37-bb4f-45cc441c4bbc · outbound

This paper cites A cross- dataset study for text-based 3d human motion retrieval.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision A cross- dataset study for text-based 3d human motion retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:18.350693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:09.761688Z digest=sha256:1b495348f32d6b886eee06a387c467fc7e57761bbc2c501f280dea95da122715

Observation 3f7840b3-31ac-44ba-a1e3-a313f6891a9d · outbound

This paper cites Is space-time attention all you need for video understanding? InIcml, page 4, 2021.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Is space-time attention all you need for video understanding? InIcml, page 4, 2021

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:09.806437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:09.806437Z digest=sha256:e1cc2fdc8b4fc9deb02a051cc5dd2455599f3ae04e2118b81c35d7853e1252ec

Observation 110b8fd9-c88c-4d7b-a86f-4d4abafa473c · outbound

This paper cites Sinkhorn distances: Lightspeed computation of optimal transport.Advances in neural information pro- cessing systems, 26, 2013.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Sinkhorn distances: Lightspeed computation of optimal transport.Advances in neural information pro- cessing systems, 26, 2013

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:18.086343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:09.916222Z digest=sha256:dbd87b8a003da77c053351959e9fde76aeff5536bdce2d0f520aa89e582e564c

Observation e1df4d46-a53f-4f95-8e51-83ea6a6c04b6 · outbound

This paper cites Segmo: Segment-aligned text to 3d human motion generation.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Segmo: Segment-aligned text to 3d human motion generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:17.845673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:10.011811Z digest=sha256:89851c955bd87124c11b78c9b44102e8f9b430d07285ba46c116e1131ee900a3

Observation 2d8e775e-2337-449d-bacd-e94edbb17081 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.Advances in neural information processing systems, 35:35946–35958,.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Masked autoencoders as spatiotemporal learners.Advances in neural information processing systems, 35:35946–35958,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:10.101270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:10.101270Z digest=sha256:11e26106996a14504ad85c1f47b5b11278b3d198ac7571ebecc81a0bcaf67287

Observation 9396af1c-2063-4a0b-af53-c00773747281 · outbound

This paper cites Generating diverse and natural 3d human motions from text.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Generating diverse and natural 3d human motions from text

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:17.689155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:10.184506Z digest=sha256:e1b0fe3b7047cf5c697be7bcf03f21c27cee390042cafacbfef780b974032718

Observation a814d231-cf1b-4a36-a09b-b2ee7c0b31b5 · outbound

This paper cites Momask: Generative masked model- ing of 3d human motions.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Momask: Generative masked model- ing of 3d human motions

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:17.448771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:10.252169Z digest=sha256:f06cdd32cfc9d5ff0b6e8156eb33b9dfa9653b3acbfa06a89914ac7e4ecb3adc

Observation 92145b61-a539-4f5f-bc8d-c4f3fe4662e8 · outbound

This paper cites Snapmogen: Human motion generation from expressive texts.arXiv preprint arXiv:2507.09122, 2025.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Snapmogen: Human motion generation from expressive texts.arXiv preprint arXiv:2507.09122, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:10.316066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:10.316066Z digest=sha256:1c0bc8ddf75dd5b2a69802546e20351c7abfdb98ec1a082c3b9118c28dc662fe

Observation dd972f82-f165-447b-93fd-bb7817a4e9fc · outbound

This paper cites Amd: Autoregressive motion diffusion.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Amd: Autoregressive motion diffusion

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:17.223219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:10.394699Z digest=sha256:6497a46798c4fe4d521634ce6523ce71c86a44aafafb045fc1a5b4d9402bf535

Observation da7653cb-97cf-45b6-9b9b-a0f7481e01ed · outbound

This paper cites Como: Controllable motion generation through language guided pose code edit- ing.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Como: Controllable motion generation through language guided pose code edit- ing

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:17.019348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:10.457611Z digest=sha256:4f32f9dad2f36de13e668285301485d76f1b48e7b949aeb3fb446ca444c22e27

Observation 4fbfc23a-42ea-4035-9321-2186955508ae · outbound

This paper cites Motiongpt: Human motion as a foreign lan- guage.Advances in Neural Information Processing Systems, 36:20067–20079, 2023.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Motiongpt: Human motion as a foreign lan- guage.Advances in Neural Information Processing Systems, 36:20067–20079, 2023

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:16.826747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:10.511210Z digest=sha256:0f824abdf447a929405815f32e24490574092843e4cba6ba3d51d426fe31cc4b

Observation e9676649-5f43-4969-8a10-99ff725ef267 · outbound

This paper cites Unimotion: Unifying 3d human motion synthesis and understanding.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Unimotion: Unifying 3d human motion synthesis and understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:16.656878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:10.609319Z digest=sha256:c32e82a7660e57f5e0d89f6477f6152fc67a6ef77e716dcf11e7f1627519b353

Observation b6e88d32-1981-45cc-928b-82dd243fa9dd · outbound

This paper cites Frankenmotion: Part-level hu- man motion generation and composition.arXiv preprint arXiv:2601.10909, 2026.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Frankenmotion: Part-level hu- man motion generation and composition.arXiv preprint arXiv:2601.10909, 2026

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:10.686875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:10.686875Z digest=sha256:0180b4631b87e27b314738c88d85d997329118d985d4aeae75e8891cd5cef7f9

Observation 0baaaaff-ba9b-4d34-9537-72e342d8d4a2 · outbound

This paper cites Motion-x: A large-scale 3d expressive whole-body human motion dataset.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Motion-x: A large-scale 3d expressive whole-body human motion dataset

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:16.419404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:10.779221Z digest=sha256:9deeeddc7124c27293bb31991388abf2c5f5249c41c1efdfa041c927dc19b374

Observation cabcb29f-76bb-468c-98c7-20dadfc5b32b · outbound

This paper cites Multi-granularity Correspondence Learning from Long-term Noisy Videos.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Multi-granularity Correspondence Learning from Long-term Noisy Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:10.857069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:10.857069Z digest=sha256:9f47e3e5cc0952e8b20e88feafc850d5a342179617c291154bc36f9698b42cd5

Observation fc20a501-8fd9-45ba-8940-7ee9aff243e5 · outbound

This paper cites Rethinking diffusion for text-driven human motion generation: Redundant representations, evaluation, and masked autoregression.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Rethinking diffusion for text-driven human motion generation: Redundant representations, evaluation, and masked autoregression

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:16.238937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:10.885666Z digest=sha256:737d1474235225935ee3c58a1108c64041517c6eb5bd15f7d264952ccaba49d1

Observation 424881ac-e4b6-4324-a00e-6e48c5cc6a0d · outbound

This paper cites Proposal-free temporal action detection via global segmen- tation mask learning.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Proposal-free temporal action detection via global segmen- tation mask learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:16.000195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:10.925323Z digest=sha256:0024cbef43d441179e92ec4bdf9acb1591c9b4c3648530ffdf194ab38578121e

Observation ecee7a8d-abd9-4d16-b5e1-d65c7a3ea633 · outbound

This paper cites Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:15.705738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:10.986515Z digest=sha256:ac63608a414647e894d08edb5e20505dce2f30215850ee8a848b1d14bbad1971

Observation 7fbdb2c5-6e48-471e-97f2-46202e197982 · outbound

This paper cites Now Foun- dations and Trends, 2019.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Now Foun- dations and Trends, 2019

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:15.475588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:11.034615Z digest=sha256:7a16ad85395018e5be1a982327a369ff28256433bbb49de9af9f3464766a56c0

Observation 9a584c9d-c0dd-48e8-a5ec-be2e508a0d10 · outbound

This paper cites Mmm: Generative masked motion model.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Mmm: Generative masked motion model

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:15.162114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:11.074653Z digest=sha256:2409e2caa51c086413d1e25b7ad9328fa5bb3c1f9f19425e241433632da30612

Observation b977c382-ccef-4a8b-8949-22bc2cc3fab5 · outbound

This paper cites The kit motion-language dataset.Big data, 4(4):236–252,.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision The kit motion-language dataset.Big data, 4(4):236–252,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.165073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.165073Z digest=sha256:5bbfe344f2b833be4baae3b245c637f718be8dddb733994a647a66746863bd94

Observation 5d8d8fc0-93e7-4e91-b1d5-9a10e7a9f51a · outbound

This paper cites Babel: Bodies, action and behavior with english la- bels.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Babel: Bodies, action and behavior with english la- bels

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.986967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:11.246845Z digest=sha256:c2b4042c4aa7c4461e4e4aeeddf6cc65b8c60b430be54d4a22f3b479fb71032e

Observation 89c6b694-6741-4344-9dcb-dafa476b1c7c · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Learning transferable visual models from natural language supervi- sion

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.279566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.279566Z digest=sha256:e4013b4773425e3cf14d3c892437b85099e33cdadb207cbd043c0441bf688356

Observation b11f5059-ada5-4929-8cb5-585a11dee210 · outbound

This paper cites an unresolved cited work.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.365326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.365326Z digest=sha256:c74c54bd00e60d0416adf17c7b76fe8534237f9d464e9d25ad6ae395ab2c7595

Observation 232ce928-4530-4ff9-bb09-b7e511c4dfed · outbound

This paper cites Ot-clip: Un- derstanding and generalizing clip via optimal transport.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Ot-clip: Un- derstanding and generalizing clip via optimal transport

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.742352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:11.436412Z digest=sha256:b5c9dbe29dabcc4124c30cf464378a710133f4a8ee19ebbd6276a76f2a886b40

Observation 6b0e7d07-a54a-420e-9339-8cf55db0588d · outbound

This paper cites OpenAI GPT-5 System Card.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision OpenAI GPT-5 System Card

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.477503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.477503Z digest=sha256:d429f764857569c3f5ca6a7fa56da79b5a60312f00e47f828211742bd4bd210f

Observation ee65441b-8813-4b88-8777-29ae06394ff9 · outbound

This paper cites Optimal transport on discrete domains.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Optimal transport on discrete domains

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.580551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:11.529984Z digest=sha256:be9acddcc76eb27e929ebae97a438c203d08c5661a8d9107ac4f37663510c42d

Observation c6481434-0f9e-4afa-9cc2-cdf53114e0fb · outbound

This paper cites CoMA: Compositional Human Motion Generation with Multi-modal Agents.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision CoMA: Compositional Human Motion Generation with Multi-modal Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.588885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.588885Z digest=sha256:ab1c61146f7173ebc67010245e0c73b1568abea33dbf59545f35ccdda7c809b0

Observation 3ca679e8-3bff-47bc-a24a-7f0a4d873767 · outbound

This paper cites Motionclip: Exposing human motion generation to clip space.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Motionclip: Exposing human motion generation to clip space

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.464800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:11.647315Z digest=sha256:6d9e5e607272a9446a5c3ae3fce650c8c5a575a996f4885af972a808404f512d

Observation fcaeb8bd-9c66-4074-ae86-5bc2543f89f5 · outbound

This paper cites Human Motion Diffusion Model.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Human Motion Diffusion Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.728862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.728862Z digest=sha256:3df45d021553d4f4b10a5e0bf44f487ecf98b45bc09ed9ef5bc71e27b1e33e02

Observation 3f702020-fa09-4a3d-a137-91d7442beead · outbound

This paper cites Scaling Large Motion Models with Million-Level Human Motions.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Scaling Large Motion Models with Million-Level Human Motions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.792396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.792396Z digest=sha256:c75579ca9f11ff120140feeba98bf0b0ebd5c98516ffb5c64328a7243c026e9a

Observation f40afea6-d756-4ab6-bf7a-4babacb90ddc · outbound

This paper cites Mg-motionllm: A unified framework for motion comprehension and gener- ation across multiple granularities.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Mg-motionllm: A unified framework for motion comprehension and gener- ation across multiple granularities

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.206099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:11.853616Z digest=sha256:dc15ffd601eb41967b7245fc687f0b45f6678bfc0695670f4a0aa326ec324981

Observation 65b39cf5-8667-4a57-afbd-b6db884a8c0c · outbound

This paper cites Dense motion captioning.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Dense motion captioning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.014695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:11.903996Z digest=sha256:e483025a8837daae670c007bdba2325f66d03831e8e0f07852b6713ec770588b

Observation f8392932-71c2-4c4c-acea-841098fe600c · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.969697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.969697Z digest=sha256:029a197cc443b779bbeccc96a78128fe40dd45c1dc4747acdcd2c921646350aa

Observation c779391c-f27a-4d81-80ab-4fbcf11f13e8 · outbound

This paper cites Generating human motion from textual descrip- tions with discrete representations.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Generating human motion from textual descrip- tions with discrete representations

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:13.891286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:12.034560Z digest=sha256:0618df730647b445d9bee1fe3961d02c67893216ef086eec335789297198a32d

Observation f3756519-1f9f-4f69-b1dc-2a1634116f49 · outbound

This paper cites Re- modiffuse: Retrieval-augmented motion diffusion model.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Re- modiffuse: Retrieval-augmented motion diffusion model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:12.118757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:12.118757Z digest=sha256:5e7be4443fa301832a9b3c5003511e11f1d98d71396e6e34e799d7ce8d225859

Observation b402e152-3973-455d-932c-1835a5056a43 · outbound

This paper cites Finemogen: Fine-grained spatio- temporal motion generation and editing.Advances in Neural Information Processing Systems, 36:13981–13992, 2023.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Finemogen: Fine-grained spatio- temporal motion generation and editing.Advances in Neural Information Processing Systems, 36:13981–13992, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:13.726895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:12.173471Z digest=sha256:ecf0c14344cc4b0304034438aef508c10412a38e3e89fce02c9160d6d48c5ed2

Observation 8c455f42-98d8-4d38-b96b-1abb68d58607 · outbound

This paper cites Pre- training clip against data poisoning with optimal transport- based matching and alignment.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Pre- training clip against data poisoning with optimal transport- based matching and alignment

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:13.640874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:12.264675Z digest=sha256:6caf4ec8aa225fa053098aa5d5893d6fd3301441ded8cbf99c126fbb88c20f47

Observation 0973253c-3618-4620-86b4-6fcc9b32012d · outbound

This paper cites DartControl: A diffusion-based autoregressive motion model for real-time text-driven motion control.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision DartControl: A diffusion-based autoregressive motion model for real-time text-driven motion control

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:13.488133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:12.328454Z digest=sha256:3179d9cf3857e798388faa35f297c91e7a98053d18419c038a184af695308912

Observation 79f73070-d566-43e8-923e-878b1e21a7e2 · outbound

This paper cites an unresolved cited work.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:18:13.339249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:12.375527Z digest=sha256:4d1066cb29936695ced5afa7eb34472e6d5ff9dd31eb90bde38a2fd8cd05cd0a

Observation a6f6d641-ed17-44c8-8ab9-3f0f10634fdb · outbound

This paper cites an unresolved cited work.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:18:13.215087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:12.439527Z digest=sha256:e28c2baa6bcb1b2fd82e770154f739a1679ae23011c26816509911feb76a72fd

Observation 1845dedb-1e0b-4bd6-b0ef-58b78ec6c43d · outbound

This paper cites The annotation process involves man- ual segment-level boundary identification: for each action segmentj, human annotators specify the temporal interval [tstart, tend]in seconds.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision The annotation process involves man- ual segment-level boundary identification: for each action segmentj, human annotators specify the temporal interval [tstart, tend]in seconds

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:13.060514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:12.510219Z digest=sha256:0753f7195a130eed2893bede6223f77c216f1f567dade5db57d76050c8f4e07f

Observation db1aac17-a015-43c4-a123-5f27eda79b80 · outbound

This paper cites specialized human motion analyst.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision specialized human motion analyst

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:12.924197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T00:18:12.550250Z digest=sha256:bf5b5cf29af3790bf5eed2fbed1136abf2ae7ad17388ca2a475d00e453b0e386

Pith citing papers

No inbound Pith citation observations are available.