Pith. sign in

Paper Citation Record · LEDGER

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space

As of 7 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2507.23188.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.23188 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:07:20.190660Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy48
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c3d0d4c0-abe6-40c1-825e-40047a55204f · outbound

This paper cites Dual stream relation learning network for image-text retrieval,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Dual stream relation learning network for image-text retrieval,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:21.068747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.917710Z digest=sha256:47ace4542ff50ba8e83504e58606e4ddb85835ed37e4eb31dac6b30ed94fb4f0

Observation 6d2d2a13-d0a9-4b2e-b8d1-c2da8f8b414c · outbound

This paper cites One-shot human motion transfer via occlusion-robust flow prediction and neural texturing,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space One-shot human motion transfer via occlusion-robust flow prediction and neural texturing,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:21.053234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.922972Z digest=sha256:d2594f94cea00f5a424360bf2bf7009bb5c8c522e504f8aed9e155c069266df2

Observation 7ceb4654-f82d-4be2-9590-087d0ae274dc · outbound

This paper cites Ta2v: Text-audio guided video generation,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Ta2v: Text-audio guided video generation,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:21.037886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.927797Z digest=sha256:f87d188ec964b785233e827be7779aa6d9bed3060bc2d07f275e57c73c72e8ae

Observation ac0208f8-bb09-4692-a9a2-d32d9ba81178 · outbound

This paper cites Cross-modal quantization for co-speech gesture generation,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Cross-modal quantization for co-speech gesture generation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:21.022470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.932734Z digest=sha256:8859b233b66605353acbd79f803a3e8b0ec1aed5fbe77fea20f3ae80a32cae1f

Observation d7739b76-5e27-4f66-bc15-a47e715def06 · outbound

This paper cites Generative adversarial graph convolutional networks for human action synthesis,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Generative adversarial graph convolutional networks for human action synthesis,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:21.007627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.937380Z digest=sha256:7f1613342fdd940f3356767a589090638a182fa3ee088c4d703e4e6d6d900608

Observation 522a19a5-af03-4a68-9383-3fd6bc0db884 · outbound

This paper cites Action-conditioned 3d human motion synthesis with transformer vae,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Action-conditioned 3d human motion synthesis with transformer vae,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.992982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.942055Z digest=sha256:7060f85d15681958693268e4a07acbace10f75407979b4f2476dcdee0b0736a3

Observation a865e48a-ab43-4969-8c7d-b0be2125ced4 · outbound

This paper cites Multiact: Long-term 3d human motion generation from multiple action labels,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Multiact: Long-term 3d human motion generation from multiple action labels,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.977904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.947230Z digest=sha256:a46e1f9818e2616e0029e7cdc1ae0436a104904339b48bbe10404c6a0f32a223

Observation 49b08ef6-91b7-48c6-ad88-c94983e8c4f1 · outbound

This paper cites Executing your commands via motion diffusion in latent space,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Executing your commands via motion diffusion in latent space,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.963097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.951851Z digest=sha256:91dfb7f17b216a754d86f9e0493fd6b0fe2f9f5ce3f6bd3cc0c451ed06521330

Observation 01d81e7a-032c-4844-8915-e6e7ae299068 · outbound

This paper cites The kit motion-language dataset,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space The kit motion-language dataset,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.947896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.956741Z digest=sha256:d2aeebcb2fa24eff6f378a7e6284d3a537c119ad40abbe86957cd7164e430955

Observation 095dfe23-4e48-47f8-9e90-2b6f8acdda8e · outbound

This paper cites Generating diverse and natural 3d human motions from text,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Generating diverse and natural 3d human motions from text,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.932758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.961155Z digest=sha256:12b76b39d92b2bbd3899539923d09236794ec3429859e63d72d18d997f4fa02d

Observation 227eaee2-97e9-449c-a7c2-4ab8269e7e2f · outbound

This paper cites Human motion diffusion model,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Human motion diffusion model,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.918269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.965631Z digest=sha256:4cdae6c9cadf454cf1217ba73d8a57059bb6c64d5b01c4377d7474d4d9171f47

Observation e0e44298-732d-4cc5-b0ba-bed362a7e0a1 · outbound

This paper cites Generating human motion from textual descriptions with discrete representations,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Generating human motion from textual descriptions with discrete representations,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.902943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.970336Z digest=sha256:943a8179fbc47f75df8ad5778b72586192f0a1d833aa2be0c4783068f611aeca

Observation 36c58672-c52c-481e-a62c-dc7cb2ad78ac · outbound

This paper cites Groupdancer: Music to multi-people dance synthesis with style collaboration,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Groupdancer: Music to multi-people dance synthesis with style collaboration,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.887444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.974905Z digest=sha256:633253e2a8fc590c74a7dcc383ef9bb9d8b72e33df0f44d6ae6c85f400ec8c83

Observation 743646df-9965-40b7-9484-63b14329913c · outbound

This paper cites Music- driven group choreography,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Music- driven group choreography,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.872221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.979443Z digest=sha256:c4c63713a844b2b38b92723055faf7e17ed5e2327087c1fe7710a338ee20a3bf

Observation c150045f-27fa-4875-a74a-f66c9598bc86 · outbound

This paper cites Edge: Editable dance generation from music,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Edge: Editable dance generation from music,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.856766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.984111Z digest=sha256:0dddd07869ff28cbffa4d71bfd7891931c18825aa4a09b129b9c374b73e042ae

Observation d56ec5b9-7cc4-4c18-b93d-4e9ab369e531 · outbound

This paper cites Pc-dance: Posture- controllable music-driven dance synthesis,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Pc-dance: Posture- controllable music-driven dance synthesis,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.841247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.988790Z digest=sha256:420d6db5b9b99be3e95ce760be56ac0c14c4865b4a1c48c4e82152f624d2952c

Observation 0abead2e-3c61-4f25-aff1-009e29565489 · outbound

This paper cites Couch: Towards controllable human-chair interactions,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Couch: Towards controllable human-chair interactions,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.825215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.993269Z digest=sha256:3864afab20b82b298640cb9d51881cb9ee7908475bafe3ea496bfaa074725cc0

Observation a509508f-d436-49c5-983f-9de104d75d41 · outbound

This paper cites Goal: Generating 4d whole-body motion for hand-object grasping,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Goal: Generating 4d whole-body motion for hand-object grasping,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.809687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:19.997850Z digest=sha256:8ab7a1a3380a765dc910086e2774d917ab019376092adfe92b52e4fca2c115f1

Observation 86592244-b176-4903-be80-4943ac117cde · outbound

This paper cites Human motion generation: A survey,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Human motion generation: A survey,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.795219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.003279Z digest=sha256:a9d5f931310bce146e652f38ce20887476dc396be853a596d140a835eccf15c0

Observation 6e0bc659-6b8e-48e9-878f-c6b549566402 · outbound

This paper cites Phase-functioned neural networks for character control,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Phase-functioned neural networks for character control,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.780450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.009756Z digest=sha256:e8de29c8ae8e7253a27810e233d651b47357cc6d517c1a8f18585db1465a81ba

Observation 0b5a84df-0c83-4c98-bc86-5774b7ba02bb · outbound

This paper cites Learned motion matching,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Learned motion matching,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.765128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.014500Z digest=sha256:06673a733e7e4d0f34020552f7b5f492c1f60d14001e37f839a15b4eb25e4fec

Observation d0429251-a731-407a-bc5d-e9b12eb55dce · outbound

This paper cites Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.750358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.023014Z digest=sha256:6ccbee2fa791275f033599d93d158593189b4bc52f2d6e2b751ae6c717f2efab

Observation 46e64a69-2467-4b76-8040-8d643eca10fd · outbound

This paper cites Tri-modal motion retrieval by learning a joint embedding space,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Tri-modal motion retrieval by learning a joint embedding space,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.734968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.028116Z digest=sha256:35b7cd36fe647944bd76d810e13cdc249867d5493397c6b8645c99d55575a2e7

Observation af304b00-2e5b-41c4-a3b6-7ff024a44696 · outbound

This paper cites GPT-4 Technical Report.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space GPT-4 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.032813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.032813Z digest=sha256:27fe7a67c9f70783acd52c279515242e0603a081b17e8ba473d0177241cef265

Observation cbaf0f20-03b5-457b-ac6b-7415c8d21989 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.716915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.039109Z digest=sha256:27b45e70664c4571fae7cf6a92e1b16e2922cd845152366fe9ff930446f4bc9a

Observation 980f1468-0f6e-48d7-bd95-59f812a4ea20 · outbound

This paper cites Better speech synthesis through scaling.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Better speech synthesis through scaling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.044506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.044506Z digest=sha256:1de2c8a11bf00299d2f63603bc46a4f0344b390f82ed815832fea768dcee9b66

Observation 0e4cc8ae-f314-40d2-8d38-0f7fa9a0b957 · outbound

This paper cites Motionclip: Exposing human motion generation to clip space,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Motionclip: Exposing human motion generation to clip space,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.699757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.051712Z digest=sha256:e0a75c8aeb76ef388fd96e80b6cdcb63f913a9b0222a742cff4a9791e7096750

Observation 5d148ccc-0ace-447d-a638-ecab485f4080 · outbound

This paper cites Auto-encoding variational bayes,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Auto-encoding variational bayes,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.684212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.059703Z digest=sha256:944e61373de8382c41836245d1aaaf5dbab9bc73a71c92eab3fb712d2464c5f5

Observation 9f74b422-15a8-4075-b93f-212138c3eea6 · outbound

This paper cites Temos: Generating diverse human motions from textual descriptions,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Temos: Generating diverse human motions from textual descriptions,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.669620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.064485Z digest=sha256:375215ddcd75b5710af44a8e58d453f71e635ff4f94988573f22e6b2aa152836

Observation a1d28c1c-9ff6-4dae-8c36-96221fdfd196 · outbound

This paper cites Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.655016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.069062Z digest=sha256:ffb88fd6bb414aab28b3db38e54875ee63bf4c171a3c1a000d8b9290cfbad732

Observation cc001b1f-5078-4bf3-8b27-61411de93679 · outbound

This paper cites Neural discrete representation learning,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Neural discrete representation learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.639950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.073904Z digest=sha256:8669dda733a98b5d84ee1fb1feb31de7ea5bcd56bf8e1aa98b4a7b1f6b53c156

Observation 7d1a1a47-e705-484a-aa53-034314937976 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.079034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.079034Z digest=sha256:694ec345783dea7cfaf0eccbbd83982d446a2ced41b5c9d7678681d39d1f03c4

Observation 6f22fbf2-7921-43ad-98df-5f01da09ceb6 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Coca: Contrastive captioners are image-text foundation models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.615160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.084405Z digest=sha256:6babf4c63857b485fc23552261af8a23f945ff8cd918c775cec3510dd46de6ac

Observation 801e400f-3c78-4cc2-8d11-48671b98f329 · outbound

This paper cites Global meets local: Dual activation hashing network for large-scale fine-grained image retrieval,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Global meets local: Dual activation hashing network for large-scale fine-grained image retrieval,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.600313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.089655Z digest=sha256:20d492cb74ea7ac002946933d10fb6ef617644490a105a008a2eb1bf18fae5b9

Observation 085ce84c-cda7-4bf3-a71f-3ed7e673b97b · outbound

This paper cites Dvf: Advancing robust and accurate fine-grained image retrieval with retrieval guidelines,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Dvf: Advancing robust and accurate fine-grained image retrieval with retrieval guidelines,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.585697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.094537Z digest=sha256:039cfd7407fd6888efb86196815dd8c9807815418154fad95cbf291d5c8c2e5f

Observation 3ecb2ba1-991b-4672-ae3d-88e04a36b54b · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.570573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.099676Z digest=sha256:057e1ce9ec5915ce275f9020a529852b78b0d3dd32d24b9c63f5370f32c1bd97

Observation 6d3d0b5d-d03c-4756-a1dc-c42c77076b7a · outbound

This paper cites Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.554167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.103973Z digest=sha256:93969831b6d4f1cfa4ed9e5283bb5bf024c1336096125621b64603296959fc8c

Observation 982ee88c-1a82-46ff-b814-f8c92ab31047 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.539447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.108193Z digest=sha256:eb92deff2a37b75e23385af23b713a9515842060fc07621b09e73dd538e8f0d3

Observation c0871aeb-5f63-4cab-af47-1f51f2579950 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.113044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.113044Z digest=sha256:4efdc0d208e51f56e9f74b49a7b8d9dfdc8a5dc4d587fa349a842099bf9324c9

Observation b59dbc53-86c7-4293-8276-ac490fdaf30c · outbound

This paper cites Audiolm: a language modeling approach to audio generation,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Audiolm: a language modeling approach to audio generation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.524041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.117704Z digest=sha256:591988475c5298ec65c3e3506a76ccf3f30e85ba72e666696eca93160b3523bb

Observation 4f80dfef-9939-46f6-9a24-a32d0ec6e2e1 · outbound

This paper cites Robust speech recognition via large-scale weak super- vision,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Robust speech recognition via large-scale weak super- vision,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.508969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.122266Z digest=sha256:8b2adfd81524db43219501350fcbe0945a66216460202020d94f0be728fed4ad

Observation 4660fb3d-fdf7-40e6-89ae-fe74d8854c43 · outbound

This paper cites SUPERB: Speech processing Universal PERformance Benchmark.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space SUPERB: Speech processing Universal PERformance Benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.126818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.126818Z digest=sha256:c74ba9a571859d6c5c08ea8dceb121ed2d83c76df34565b5676389db8cec7ea9

Observation dc30a259-434d-4211-a537-8b4b5270b804 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.494265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.131594Z digest=sha256:08be8eb9268e15e293e6eb45d2ecfe1468330e8810ee5335c4f481a46bff0997

Observation d449bd73-c834-4a0d-80f1-de423635cec1 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Learning transferable visual models from natural language supervision,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.136841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.136841Z digest=sha256:ca24ea219eeb9ce365681e74fdcf973dbad8481442e174fb27728b53313735a9

Observation ae742115-187d-4977-aa0a-de9c99ed3517 · outbound

This paper cites Videopoet: A large language model for zero-shot video generation,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Videopoet: A large language model for zero-shot video generation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.468393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.141428Z digest=sha256:d7b433b4f636927c1ec31c87a9f9736bba70c95df6c0957a48b64ccb899c66c1

Observation 1b317a51-829d-4de2-b82a-963fb5ab2e10 · outbound

This paper cites Label independent memory for semi-supervised few-shot video classification,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Label independent memory for semi-supervised few-shot video classification,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.453186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.146030Z digest=sha256:969ba5ce4b85d7c5b0c44a2d63de3c7c7dcd6801556f3814f51d7757b710d09b

Observation 801c818e-be16-4958-8cb8-7cc96a45afc4 · outbound

This paper cites Memory-enhanced transformer for representation learning on temporal heterogeneous graphs,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Memory-enhanced transformer for representation learning on temporal heterogeneous graphs,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.438474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.150446Z digest=sha256:b8829753f9cf39f22a01c54349c072d8d5c183d4001f4b548e122548064f262f

Observation 392e1a5d-4b2b-4ed8-b666-15d909101a2c · outbound

This paper cites An efficient memory module for graph few-shot class-incremental learning,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space An efficient memory module for graph few-shot class-incremental learning,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.424127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.155284Z digest=sha256:a56706b11ae1477e88288f0143d155f6e1ba950410519ec72400e0b96da2713d

Observation 5b0982a6-d2e9-414b-bb65-99388e515b22 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Imagebind: One embedding space to bind them all,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.408819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.159521Z digest=sha256:c6e88c522337a25ca598f26ce117d7048af161ee603b0ab2a35e2a8dfc0102fe

Observation a4b8a03c-d931-481a-88e3-c217fb40e66d · outbound

This paper cites Grounded language-image pre- training,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Grounded language-image pre- training,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.392196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.163885Z digest=sha256:5f2f9a934715eea6752b5be820cfce9bf5eeb5e586e817f177c4961007b0debb

Observation 0e899e0e-25bf-4e3e-ab1b-d6b312477551 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space ActionCLIP: A New Paradigm for Video Action Recognition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.168158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.168158Z digest=sha256:0ddc3155cf4c88b226cb3eb3f8d47c208fa21a3cf3afed8f2c72dc2a35876d94

Observation 026a9884-f216-416a-b701-5d015e57fc77 · outbound

This paper cites Delving into multimodal prompting for fine-grained visual classifica- tion,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Delving into multimodal prompting for fine-grained visual classifica- tion,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.377514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.173210Z digest=sha256:c5f7714f9ac8d8ae55d4e54d90d6c4a79e002f4ebfaa80f1ccfe4e66985c78d6

Observation 426b2e26-f8d3-4a76-b4a3-3f56eee89156 · outbound

This paper cites Amass: Archive of motion capture as surface shapes,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Amass: Archive of motion capture as surface shapes,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.361998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.177499Z digest=sha256:94555bca8f4f7bf07889273bc38e03c3c7f5749a54c70b9a6ddefe017abe90a0

Observation 7ef95020-8e13-4f9c-9011-b7434957e4bd · outbound

This paper cites Action2motion: Conditioned generation of 3d human motions,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Action2motion: Conditioned generation of 3d human motions,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.347117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.181731Z digest=sha256:5e22568e5040996e7b60179da99eb7b9f296d12fab79432eb8b5ad0828bf87b7

Observation 0f038371-bae1-4e37-9541-80b9b3c4ff3d · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Representation Learning with Contrastive Predictive Coding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.186032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.186032Z digest=sha256:0a707706e21a7009ef07b544c873cc3cd2f47a71687842b40e6fb95ed346c724

Observation b1a739eb-9df3-4685-bf1f-f5100fa96c4b · outbound

This paper cites Randaugment: Practical automated data augmentation with a reduced search space,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Randaugment: Practical automated data augmentation with a reduced search space,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.331762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:07:20.190660Z digest=sha256:73c345d63c1d0ad6761dfd800035dbd12e9328c58bf65ac80ae3f6868d0ef336

Pith citing papers

No inbound Pith citation observations are available.