Pith. sign in

Paper Citation Record · LEDGER

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing

As of 21 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 3 inbound Pith citation observations for arXiv:2411.16781.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16781 v2

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:32:21.433343Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:37:28.236281Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T21:55:03.803275Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact0
  • verified fuzzy65
  • unresolved19
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4a9d95a-cc54-4242-94f7-8010c3538f76 · outbound

This paper cites Gpt-4 technical report.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Gpt-4 technical report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:20.912733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:20.912733Z digest=sha256:15e4b2172b0f2b62bcd4a01f6b365f4008556a36d6e552ed9fa421ab0e64e8e8

Observation 7876657c-123a-4c00-bd8e-f9b94dd63012 · outbound

This paper cites Star-transformer: a spatio-temporal cross at- tention transformer for human action recognition.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Star-transformer: a spatio-temporal cross at- tention transformer for human action recognition

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:20.917749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:20.917749Z digest=sha256:a3b2850c667d4515d5bd3fa587b0ba46da1cc7ee8e456a40491346b7529e7f20

Observation 64257746-f93a-4247-ae51-840318a12901 · outbound

This paper cites 2d human pose estimation: New benchmark and state of the art analysis.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing 2d human pose estimation: New benchmark and state of the art analysis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:20.922135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:20.922135Z digest=sha256:dc4add5ba87b1bb9e2c5a220ff95a30e07a915e43bfebf742cbbf3d17c0df999

Observation 0710852b-185f-46a2-bc17-5339ce4e5dac · outbound

This paper cites Qwen-vl: A frontier large vision-language model with versatile abilities.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Qwen-vl: A frontier large vision-language model with versatile abilities

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:20.926753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:20.926753Z digest=sha256:cc19ff775e8d3e9f7a3879f108dca87fec244087c0b09c01979bb5719e264bb6

Observation 618c2956-bb17-42a0-9955-61a8404d2a43 · outbound

This paper cites One token to seg them all: Language instructed rea- soning segmentation in videos.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing One token to seg them all: Language instructed rea- soning segmentation in videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:20.931493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:20.931493Z digest=sha256:e8441e703bbc146b5b37700fdd5e3918400fcad565f51f6410a8e7167d222b5d

Observation 84386f5b-e36d-4c59-87fd-b121bf6686b3 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.681723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:20.935657Z digest=sha256:19db0ffc77c600a4cd073ee97b374b72a3f34a6740db4b2b81458dda1c9da645

Observation 18720050-19d7-4b1c-b92d-ea0120947749 · outbound

This paper cites Keep it smpl: Automatic estimation of 3d human pose and shape from a single image.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Keep it smpl: Automatic estimation of 3d human pose and shape from a single image

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.666850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:20.940623Z digest=sha256:b282a92dd6a4ec1fe0a03e8fd36744811ec73c8d70256c63c6353dedcfb324ae

Observation f12619a8-08b8-4fed-89fb-1fc90b4d2111 · outbound

This paper cites Towards bet- ter adversarial synthesis of human images from text.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Towards bet- ter adversarial synthesis of human images from text

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.650345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:20.945599Z digest=sha256:2dfde682b1396964fc11fd70b9a8e8d06f987e452a8bb90ee652846190381de7

Observation eaba6592-26e3-45ca-a011-a1c59fa6334b · outbound

This paper cites Smpler-x: Scaling up expressive human pose and shape estimation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Smpler-x: Scaling up expressive human pose and shape estimation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.632935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:20.949475Z digest=sha256:50c0d474c7e6c700db7ef3e7514fed85f51a323089ae18cb5cf6865d48329a0a

Observation 85af158f-2e60-4438-960e-df3da4e3a9a9 · outbound

This paper cites Motionllm: Understanding human behaviors from human motions and videos.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Motionllm: Understanding human behaviors from human motions and videos

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.614519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:20.953587Z digest=sha256:0eff5e7e0de67361ece2ac466fd9121f3dfcb75df99a0bacfd031f413eefe90b

Observation e130c720-0993-445a-8870-3d8ddb3088ec · outbound

This paper cites Channel-wise topology refinement graph convolution for skeleton-based action recognition.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Channel-wise topology refinement graph convolution for skeleton-based action recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.590230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:20.957532Z digest=sha256:97ae72672214a755389156eb9f600b12efd91c4a25e31b8e82d6ee8d81578683

Observation 489fe822-fbc8-43e0-a267-2e8c147fb47c · outbound

This paper cites Learning phrase representations using rnn encoder-decoder for statistical machine translation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Learning phrase representations using rnn encoder-decoder for statistical machine translation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.570041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:20.961555Z digest=sha256:1aefdc19da3bf1839b554f72cd0c9116a61ce5b1c2a261bacd1f1b4daabef325

Observation b9136a30-9000-4a99-98bd-995a3bab6bac · outbound

This paper cites Posefix: correcting 3d human poses with natural language.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Posefix: correcting 3d human poses with natural language

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.551363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.086359Z digest=sha256:6932bad6930be9afacfb4603b0b8f379217dd3c03dc9f0433df415446d381105

Observation a22aa2d8-6dde-493e-a01c-b73232649c71 · outbound

This paper cites Posescript: Linking 3d human poses and natural language.TPAMI, 2024.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Posescript: Linking 3d human poses and natural language.TPAMI, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.528919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.095375Z digest=sha256:476e6a7ff1b290b706b6634bf89149cf12f359eea11b61632a4940c229e6d58e

Observation 583e4259-9014-49d1-b768-3e52f335a8cd · outbound

This paper cites Poseembroider: Towards a 3d, visual, semantic-aware human pose representation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Poseembroider: Towards a 3d, visual, semantic-aware human pose representation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.514957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.099673Z digest=sha256:4f105814d92389f8f0ecef368f83cbdf707eb2b3f36699b20b210ee7b4645436

Observation 4309ce34-fc14-42a0-ace3-a3515811f555 · outbound

This paper cites Tokenhmr: Advancing human mesh recov- ery with a tokenized pose representation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Tokenhmr: Advancing human mesh recov- ery with a tokenized pose representation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.493319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.104976Z digest=sha256:9d4e20f2810a159c6d1417017c9190d379c2239d85f851ad8632607591fb4467

Observation 6f6d2052-187e-4cef-b672-5be79cf6357f · outbound

This paper cites Re- vitalizing optimization for 3d human pose and shape esti- mation: A sparse constrained formulation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Re- vitalizing optimization for 3d human pose and shape esti- mation: A sparse constrained formulation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.481068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.108877Z digest=sha256:743ae1f65af53f0b0bca8bfd1650e1bfd50aca6fb1e63f2e870853a9e12676c1

Observation 390a6615-27ba-466e-a2e6-5bcc94a7c95d · outbound

This paper cites Learning analytical posterior probabil- ity for human mesh recovery.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Learning analytical posterior probabil- ity for human mesh recovery

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.468678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.112708Z digest=sha256:705c8affcec74be4b6cd8391f4f26063c5c612310d70973950535af0936a7275

Observation ff1c9a1b-ec4d-4141-aec4-8e32ecf33157 · outbound

This paper cites Chatpose: Chatting about 3d human pose.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Chatpose: Chatting about 3d human pose

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.455328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.119077Z digest=sha256:f0fd97071edfef702c35d0f96ed8f9d07c7e669df311a7296aeff289d6fe2a18

Observation 5bb87772-2c92-4eb2-9984-6751464dfd4e · outbound

This paper cites Vq-hps: Hu- man pose and shape estimation in a vector-quantized latent space.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Vq-hps: Hu- man pose and shape estimation in a vector-quantized latent space

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.440633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.123544Z digest=sha256:caae693209519aa2a9bf8578d5b8cd88163a54cef5fd1e7db3227b8aec41a0a7

Observation 41ac489c-5776-447f-b04d-1bacdd5c2c17 · outbound

This paper cites Mega: Masked generative autoencoder for human mesh recovery.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Mega: Masked generative autoencoder for human mesh recovery

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.427046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.130396Z digest=sha256:e6090f60b0a57efaa3220fe0f37f067c99c48a62566127bda568cd24ae400535

Observation e4769516-f9ef-41d0-af9c-9eeb653e25e7 · outbound

This paper cites Aifit: Automatic 3d human-interpretable feedback models for fitness training.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Aifit: Automatic 3d human-interpretable feedback models for fitness training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.413208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.134963Z digest=sha256:b5f59cd4b9955e9a91c656776322876a3db4262e94c4d784d8b0ccde629b0b07

Observation 9466948d-567c-46a4-975c-c3917080cd9e · outbound

This paper cites Unified pose sequence modeling.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Unified pose sequence modeling

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.399690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.139847Z digest=sha256:721beb06fd49c7fc20082dae5cbd917391a90c9f406794dfeaebe468d3a4e060

Observation 68089932-b864-4d82-8b19-118df3e87b86 · outbound

This paper cites Humans in 4d: Re- constructing and tracking humans with transformers.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Humans in 4d: Re- constructing and tracking humans with transformers

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.385678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.143836Z digest=sha256:8959f2c11edc6af3f4fd6c4b62c96715336003f1188e1b9aa44d30ccb5095adb

Observation 33041c90-8215-474f-ac1d-bf82b637c59f · outbound

This paper cites Semantify: Simplifying the control of 3d morphable models using clip.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Semantify: Simplifying the control of 3d morphable models using clip

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.365237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.147953Z digest=sha256:88f4561eb5ff809880a4dcafbb557b008c0cc8e8e73df3c190bd4752aed482ff

Observation 00e4dfc9-7096-410a-8484-5145a9942887 · outbound

This paper cites Textbooks are all you need.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Textbooks are all you need

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.348978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.154379Z digest=sha256:9a6750f84f2517f99225966467c07d1826790c7e1912a44972ebda14d22fab91

Observation 1252aec1-2562-4dd8-b177-1f11c6be110c · outbound

This paper cites Denoising dif- fusion probabilistic models.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Denoising dif- fusion probabilistic models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.159567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.159567Z digest=sha256:237109d3d3c7455c564d4c0615295201fb1cc01080edd1accc62080ef2d70e10

Observation 206c4be5-20dc-492f-9f47-ddb63b54324c · outbound

This paper cites Avatarclip: Zero-shot text- driven generation and animation of 3d avatars.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Avatarclip: Zero-shot text- driven generation and animation of 3d avatars

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.327371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.164266Z digest=sha256:ab9ed96aaac04605b4c8cc7d641f695b4b6054600b20e7eb77ab9954f609fb03

Observation 0e1b1487-f2b4-494d-bbe7-cd694f061bef · outbound

This paper cites Lora: Low-rank adaptation of large language models.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Lora: Low-rank adaptation of large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.312089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.170376Z digest=sha256:e3a285fb49492c20c8f0414dd1a81bf3b93d3ec603c8f7ed50084aa5ca43c0a6

Observation 2a406416-bd38-4005-b7ff-c9325f6a24cf · outbound

This paper cites an unresolved cited work.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:32:22.294645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.175172Z digest=sha256:b974f6ab1d6dddcaaaa244d6f3c581e204c5c7dc0336e0dacca9302c5b5df15a

Observation bfaa98e0-f5ae-4bc7-9858-20ee6442d96f · outbound

This paper cites Mistral 7b (2023).

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Mistral 7b (2023)

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.277048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.179729Z digest=sha256:27c0439b05d271828331d3551e456dd7b50f10e90296a0267f84ec8d39209945

Observation 125c29cc-af2e-4a43-8d5b-34660f115841 · outbound

This paper cites Motiongpt: Human motion as a foreign language.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Motiongpt: Human motion as a foreign language

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.183957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.183957Z digest=sha256:cfbdbe76fbdf1573389c7df79c86eb91034c1e3022dffea3a610a87bddd40de7

Observation daf88416-e854-4dfa-be74-387ffa80b8aa · outbound

This paper cites Learning effective hu- man pose estimation from inaccurate annotation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Learning effective hu- man pose estimation from inaccurate annotation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.252169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.188573Z digest=sha256:e16aabc4cd120e59ae69b64a3ef63ac351d9730a1304f0d1a684117d569890d4

Observation 0d8d3e66-0a16-4507-865d-6e560fbd521d · outbound

This paper cites End-to-end recovery of human shape and pose.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing End-to-end recovery of human shape and pose

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.236153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.192875Z digest=sha256:13f3da0e3b936ea3b19a60019c6e6dbcdfa88c6142dea958c827a7864f90ce9c

Observation e6904c24-4f28-4b1d-9023-75f697861773 · outbound

This paper cites Fixmypose: Pose correctional captioning and re- trieval.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Fixmypose: Pose correctional captioning and re- trieval

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.221965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.197389Z digest=sha256:3c2abe0e8649f6c98a20a41ec11dc4897eb03f8f4cecfbeed787e0bb2ec6171e

Observation 89d5e5ee-99e4-42b4-9bfd-8e4ab556e358 · outbound

This paper cites Segment any- thing.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Segment any- thing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.202172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.202172Z digest=sha256:c1580cf216060ae441cfd9b90ccdc7477e4bdcdd53e669d0576a8ebf70d85125

Observation 3a5ef608-1c65-4fb3-b782-4b28855b6ad0 · outbound

This paper cites Lisa: Reasoning segmenta- tion via large language model.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Lisa: Reasoning segmenta- tion via large language model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.207670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.207670Z digest=sha256:fa75c4849f07827523700b1569b76b2480bcfb35dced8dfeb30a0e9c196062c5

Observation 658cc7c4-5e84-4895-9a1f-623ed5b87032 · outbound

This paper cites Llava-onevision: Easy visual task transfer.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Llava-onevision: Easy visual task transfer

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.189486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.211655Z digest=sha256:665996ab8a85d579aa83b4aa7f1a6d6ef27a59305908cb1912de19464ab5397f

Observation 146ca529-ee33-4ac7-9f00-a43f1cc05f4f · outbound

This paper cites Video-llava: Learning united visual repre- sentation by alignment before projection.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Video-llava: Learning united visual repre- sentation by alignment before projection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.176369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.215654Z digest=sha256:5286e07d2203da008e3afed923b461cd1756aee9708f4ea5c98f522ff0160f62

Observation 9b343677-6b4c-4fce-a9c8-42aadf1ed9fd · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Rouge: A package for automatic evaluation of summaries

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.219603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.219603Z digest=sha256:673f679ac53dd29892ac2aa7d761c9d4ea9a131e31c71afae93d1f3f89604c2b

Observation 39aff61a-637b-4c94-911f-692e15a205b6 · outbound

This paper cites Being comes from not-being: Open-vocabulary text-to-motion generation with wordless training.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Being comes from not-being: Open-vocabulary text-to-motion generation with wordless training

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.155392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.223517Z digest=sha256:1b48ca882f48303a7584a92733a1fc975a8f4c0988489bc9468f7bf2a3301f48

Observation 4d813e64-4013-42dd-b893-2af49af8e60f · outbound

This paper cites Chathuman: Language-driven 3d human understanding with retrieval-augmented tool reasoning.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Chathuman: Language-driven 3d human understanding with retrieval-augmented tool reasoning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.139661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.227355Z digest=sha256:4d8e6c4c956c8c34e52e343287adc521728bcdf0f2959f7d45309d219626a60f

Observation 1755b1c1-0985-4422-9f5b-0235735f5843 · outbound

This paper cites Microsoft coco: Common objects in context.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Microsoft coco: Common objects in context

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.123783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.231695Z digest=sha256:2589eca5e74ddfe9c0b7dfa42099ca2b7b44ccd327507f3cca56bdbd5a153c10

Observation 36489d87-307c-409c-898d-1a5b1d44fe09 · outbound

This paper cites Visual instruction tuning.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Visual instruction tuning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.108677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.235564Z digest=sha256:b072342b1fd336f51b8a1d6183a66c3c0e32f6be57935ab1fa203c5966d53f17

Observation 53c3bed6-6d39-4957-bc45-d127255eae69 · outbound

This paper cites Improved baselines with visual instruction tuning.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Improved baselines with visual instruction tuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.239647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.239647Z digest=sha256:911d9f80a739b2169fc086cd0936e7c112a644c13aa55d723ae0ee436d4a5f64

Observation 6cd696fa-3315-4832-b608-1e46fc3064ba · outbound

This paper cites an unresolved cited work.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:32:22.085004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.243783Z digest=sha256:a46aee068fe06e97f14c0bc462430d785825b13cbf46d042cb5509d8a4182321

Observation 59e728d2-b0e5-45fc-83fb-cb7238f82bfe · outbound

This paper cites M 3gpt: An advanced multimodal, multitask framework for motion comprehension and generation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing M 3gpt: An advanced multimodal, multitask framework for motion comprehension and generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.067985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.248502Z digest=sha256:11e102cb4da4ca69ef19baf7dbc8ea75c5dcfe3b900eb61aae197b6fd48d9fad

Observation ef6605ec-c997-43d6-ba72-277bed693146 · outbound

This paper cites Amass: Archive of mo- tion capture as surface shapes.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Amass: Archive of mo- tion capture as surface shapes

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.252722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.252722Z digest=sha256:62794ff5e18ef877fbec01cbb97b0d1d8536e6ad2da271808c18ca85d5aa3409

Observation d1f61c97-5b83-4c01-9735-f83f1cd7693b · outbound

This paper cites Monocular 3d human pose estimation in the wild using improved cnn supervision.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Monocular 3d human pose estimation in the wild using improved cnn supervision

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.040688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.257226Z digest=sha256:693abe9dddd7a81c60c6e3f8641a39bdf7329b0b6da06e63370680785eeb95a6

Observation 37c397a7-3b3d-4432-8676-b6451b05ad85 · outbound

This paper cites Hummuss: Human motion understanding using state space models.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Hummuss: Human motion understanding using state space models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.026305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.262874Z digest=sha256:04710c45b250a5056a4558d371abe6d628268ebe819968f9a723a016756673e8

Observation 30ec4ef5-5d5b-4636-ba7b-6b76d87d6c93 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Bleu: a method for automatic evaluation of machine translation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.007572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.267017Z digest=sha256:34b7f6b8ee368ce9a2593276a6ee085c66094630a886d221939f1310c083b7a7

Observation c13ea94f-15d4-4a2b-ada4-ba3ebbed6392 · outbound

This paper cites an unresolved cited work.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:32:21.992983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.272160Z digest=sha256:b8240634e6eaf54420bc5cc102af20e1d4256ef0e1e350d70fdb8eb9f992184e

Observation f7996632-5593-4697-9f46-c6da324c6749 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Learn- ing transferable visual models from natural language super- vision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.280170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.280170Z digest=sha256:542d402a1050cdb13a383b0dc2306e2ce75e85b7c6f21de188996764193ebfa9

Observation 47101158-c16a-44a2-a8e3-3f223a612d7f · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.967156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.285405Z digest=sha256:2a59525bec38541c2e51fd11e18b6f0ce7151383360ed0dc783518e1720b1797

Observation 242cf45e-6aec-4f84-b690-1fb76b55d11c · outbound

This paper cites Humor: 3d human motion model for robust pose estimation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Humor: 3d human motion model for robust pose estimation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.951624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.291818Z digest=sha256:b0dc208586b079f9460047d036aa8ab92ddc8a9703c0c853ac5d9bf2df5e9c65

Observation d5e2eb03-b7e9-48f5-93e4-5b09f8689cd2 · outbound

This paper cites Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.937045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.298317Z digest=sha256:70e079bd2d2219def31cc61cde3594cd9db301152b8ae173fc74a0acc7a5166d

Observation 555d9789-8f59-4a3c-9ca5-2b8d3b4e0166 · outbound

This paper cites Body talk: Crowdshaping realistic 3d avatars with words.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Body talk: Crowdshaping realistic 3d avatars with words

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.922478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.304225Z digest=sha256:e7da37511f4755a66b1109fcf6e1affc7e8985c9d3524e2c50bc1ccf73189490

Observation dbdbca8d-4e47-481b-832b-277842654197 · outbound

This paper cites Aios: All-in-one-stage expressive human pose and shape estimation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Aios: All-in-one-stage expressive human pose and shape estimation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.907398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.310717Z digest=sha256:91b86e7e8bb977a51eb3ded199fcd3f7ad6088f3a38cdbe9e625ae6c835b67cd

Observation cd67ffb1-5e74-4d9a-9f17-a9363efe1e01 · outbound

This paper cites Llama: Open and efficient foundation language models.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Llama: Open and efficient foundation language models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.893392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.315582Z digest=sha256:133b91b935d9455ff6bc2a0b8fd37c2beb1c0fae05db5435ceb33ce662625164

Observation 80ceb2ec-20b9-4709-b8dc-1b993f1baf9c · outbound

This paper cites 3d human pose estimation via intuitive physics.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing 3d human pose estimation via intuitive physics

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.878143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.320093Z digest=sha256:fa3fa1528a2de569959e4c864caa8fb4326e3f9db56e4012315649df7b456d67

Observation bbe40c2d-a3bf-4269-bda7-413dee71c8f5 · outbound

This paper cites Neural discrete representation learning.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Neural discrete representation learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.324642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.324642Z digest=sha256:e720529d44bed45081a647668fc4d3304b217e321d9beb131784d9a5ca83ea8a

Observation ca28bed9-137a-4303-8f1b-2a5c094f7e46 · outbound

This paper cites Recovering ac- curate 3d human pose in the wild using imus and a moving camera.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Recovering ac- curate 3d human pose in the wild using imus and a moving camera

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.854282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.329306Z digest=sha256:ab44b950bfa9e1509102fdb384bbf254cf213683de7fc70a8043e80bbb24719d

Observation 0fe78897-2607-4d76-b83f-a916817a0ad6 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Videomae v2: Scaling video masked autoencoders with dual masking

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.839383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.333362Z digest=sha256:6141e013ab8c37539bb047b34903820853eab20a26b3dc920ce82c0fafce5133

Observation 5f8fef19-1172-49b2-8240-48eb37437015 · outbound

This paper cites Masked video distillation: Rethinking masked feature mod- eling for self-supervised video representation learning.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Masked video distillation: Rethinking masked feature mod- eling for self-supervised video representation learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.823214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.337724Z digest=sha256:fe2b8c7d11226ddb4fc4f69c4b9c1b9dae56e639ff4d44c37d4ee0fa522fdff9

Observation c3fde15c-4bdc-42a4-aee5-ab04ef616675 · outbound

This paper cites Zolly: Zoom focal length correctly for perspective- distorted human mesh reconstruction.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Zolly: Zoom focal length correctly for perspective- distorted human mesh reconstruction

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.804549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.342598Z digest=sha256:4a7df4b9bfbf3fb2ba372d0cc72a781d787d27de9a7d07b8ea8227934bde6c36

Observation bf1c9eab-8d85-4415-92b1-89ab2d78b0d3 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Cogvlm: Visual expert for pretrained language models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.790484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.348023Z digest=sha256:370690566694b31cfb42c7b3305d0f2f2b86c18be005d5399abe1e0e9257685b

Observation d9ddb6f0-be64-45a7-8e6e-63d335d8da32 · outbound

This paper cites Internvideo2: Scaling video foundation models for multimodal video understanding.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Internvideo2: Scaling video foundation models for multimodal video understanding

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.771542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.352958Z digest=sha256:eb1f3f76b360cb43cf56ecfdda5284af9183bdaf77fbfd902a4942eaba7ccbfb

Observation 95df5f52-b8bc-43b2-8ed4-e786d8a3cf95 · outbound

This paper cites Occllama: An occupancy-language-action generative world model for au- tonomous driving.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Occllama: An occupancy-language-action generative world model for au- tonomous driving

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.754573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.357567Z digest=sha256:9e85d0182b8445dcab71d1c3cf5e7db7f56e321119d862c444a9ae1bb0090ff4

Observation a3791995-a88d-48ba-a198-cd649437d5a4 · outbound

This paper cites Motionllm: Multimodal motion-language learning with large language models.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Motionllm: Multimodal motion-language learning with large language models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.740719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.361735Z digest=sha256:cc798dfe6c234e7f767365af9cf4bc00c980200cd7473b451c1a6b26e19efbbf

Observation 58fcb0a1-17b4-4742-ba50-e2491ff90c62 · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Show-o: One single transformer to unify multimodal understanding and generation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.726148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.365581Z digest=sha256:01cf30204969ce2bb4f05c8a85c0de83cffdd15da16b91bb1ba492b752cf88ca

Observation dee4e429-7000-48bd-a48d-6304b05284c6 · outbound

This paper cites Smpler: Tam- ing transformers for monocular 3d human shape and pose estimation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Smpler: Tam- ing transformers for monocular 3d human shape and pose estimation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.710679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.369509Z digest=sha256:5df8358d20fa125b45df1fb86202d1b17bbdbbde0842d6102adebffc07df0430

Observation fc1f439a-454b-4263-9c70-d6f06b30ebaf · outbound

This paper cites Qwen2 technical report.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Qwen2 technical report

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.690567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.373384Z digest=sha256:80e4978ed089818653b8b03536e5e681275271aba95f89db044e305ccc99ee23

Observation 6c73b9cf-f7c0-45a3-8d48-94795c21b4dc · outbound

This paper cites Uni- audio 1.5: Large language model-driven audio codec is a few-shot audio task learner.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Uni- audio 1.5: Large language model-driven audio codec is a few-shot audio task learner

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.666633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.377256Z digest=sha256:66cedfca5618ad4c3d1d95cdaadbfe3e56e23445445aaa32bd0697253394df35

Observation 9ab6d6af-98f5-412f-a463-789852b3f402 · outbound

This paper cites mplug- owl3: Towards long image-sequence understanding in multi- modal large language models, 2024.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing mplug- owl3: Towards long image-sequence understanding in multi- modal large language models, 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.651711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.382032Z digest=sha256:67d16aa584752236c8b71d899cb58e195ee9cb6f2f64e4a2372af052c943ae4b

Observation 070da214-fd14-4611-a9e6-035bbb3c628b · outbound

This paper cites Speechgpt: Empow- ering large language models with intrinsic cross-modal con- versational abilities.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Speechgpt: Empow- ering large language models with intrinsic cross-modal con- versational abilities

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.635603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.388321Z digest=sha256:c0733998d3ce434e7c0fcee531802c8f84273c0f55da0431a3a7863c7c5f2b44

Observation 80b6d2f3-338a-4e27-b3d2-389936a221b6 · outbound

This paper cites Pymaf: 3d human pose and shape regression with pyramidal mesh alignment feedback loop.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Pymaf: 3d human pose and shape regression with pyramidal mesh alignment feedback loop

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.608653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.393026Z digest=sha256:b79017dec2e260c0fde6284c632a44ae492e74fa329a7c31869bd85b59499a50

Observation 0923ea77-9991-4f7d-b6ed-0e7059ff899e · outbound

This paper cites Mo- tiongpt: Finetuned llms are general-purpose motion genera- tors.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Mo- tiongpt: Finetuned llms are general-purpose motion genera- tors

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.397086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.397086Z digest=sha256:b5bd103babbfe868caf9b927d67f8d1ba0cbc5cbd8ab4656561bd19b9f9339b3

Observation 10617ccd-78e5-4739-b49e-bbac5808ac5b · outbound

This paper cites Single im- age action recognition using semantic body part actions.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Single im- age action recognition using semantic body part actions

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.576250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.401677Z digest=sha256:6063db4599cfd33645a39d20597155766b6ab41af57c59aa3b20e072a1507991

Observation bd19217a-22b0-4553-a834-422408928f10 · outbound

This paper cites Transfusion: Pre- dict the next token and diffuse images with one multi-modal model.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Transfusion: Pre- dict the next token and diffuse images with one multi-modal model

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.557106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.406441Z digest=sha256:e509e2db71a926a3d2793c19cdb427e02982fb6098f1951e97853e90e39d2e6f

Observation 2747de09-6f07-4641-8364-a8e97054a666 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.541729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.411596Z digest=sha256:28f8558a1f3ea8f1054b811d0ce8220abe0ec2e17a64f0b11d35c79b24de2ada

Observation 2fbf5d57-adee-4b25-9015-171dd9c2f7f7 · outbound

This paper cites Therightkneeandthigharestraight,andtheleftlegisalsostraight.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Therightkneeandthigharestraight,andtheleftlegisalsostraight

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.526223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.416075Z digest=sha256:26bccde8dc5c3029150d680234a6c1ce41a82e8c03c7f5587ae43ba3b05e101f

Observation a8cc1aff-ce1c-4d1c-8c34-94e38ef00ca9 · outbound

This paper cites This extensive data effectively facilitates the alignment of pose and text modalities.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing This extensive data effectively facilitates the alignment of pose and text modalities

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.510767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.421879Z digest=sha256:e79deb42ceff87c222a6c97153056c88ba10ba057cb75ec7a35bd488f92425a5

Observation d480161a-2b97-4133-a588-a36d37ded064 · outbound

This paper cites an unresolved cited work.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:32:21.495015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.427240Z digest=sha256:a2448d0976499928eb26ddcb21b82ab8d9bd967c7cde7d7021219426f9f3a14c

Observation 8ae9533b-a6b3-4252-b69c-ed44c7bf8e5f · outbound

This paper cites The results show that our approach more accurately estimates human poses, even in scenarios with complex limb articulations.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing The results show that our approach more accurately estimates human poses, even in scenarios with complex limb articulations

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.476877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T13:32:21.433343Z digest=sha256:119befed64912509716e18f9d4a6f34c38bc200c364faadef494009a61be2912

Observation ef29c048-4adc-4374-ad36-65e4d83caaf2 · outbound

This paper cites an unresolved cited work.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Unresolved cited work

Reference 2023

Resolution
parse uncertain
no resolver link, observed 2026-08-12T13:32:21.090809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.090809Z digest=sha256:e7928bc0a45bc971569920fff6c96239e839f7620fbe41c48dc283a7ac82bc08

Pith citing papers

Observation deb65451-07b8-4f89-b9a5-20ee598e4a3e · inbound

Human-Centric Foundation Models: Perception, Generation and Agentic Modeling cites this paper.

Human-Centric Foundation Models: Perception, Generation and Agentic Modeling UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T04:41:08.459981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:41:08.459981Z digest=sha256:09757ddf479f0957a5aa7ff2207fdff23406387ffb39f59e821a91647398cf42

Observation f031606a-9685-4d35-967b-9da7383f2d0e · inbound

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model cites this paper.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:03.873453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T21:55:02.001344Z digest=sha256:8e21f4f8a77b6be699f91aa1db4e5e8c796da9d80a8e09f9b0897cd0ddc39a93

Observation 494ed840-153c-4c07-91a5-87a25528ded1 · inbound

Extension of generalized KYP lemma: from LTI systems to LPV systems cites this paper.

Extension of generalized KYP lemma: from LTI systems to LPV systems UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:37:28.236281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:37:28.236281Z digest=sha256:b633cd1a8fa6ed908f70671145f065961b1379b665e0b8e5df5b5e6539e6aeeb