Pith. sign in

Paper Citation Record · LEDGER

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model

As of 14 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 3 inbound Pith citation observations for arXiv:2508.07863.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.07863 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:55:03.686435Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:39:25.303896Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T13:47:57.512626Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact5
  • verified fuzzy28
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dece0443-f9d0-4aed-8a4b-fb9ee5b6a50c · outbound

This paper cites Momask: Generative masked modeling of 3d human motions.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Momask: Generative masked modeling of 3d human motions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:57.061147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:57.061147Z digest=sha256:be8767ec8cb9fe58e6d1ea3abf1330d1f97afb81997471291d88a7da43475526

Observation e7f6a5fe-b310-4021-88d5-e9b0c9f65b11 · outbound

This paper cites Motiongpt: Human motion as a foreign language.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motiongpt: Human motion as a foreign language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:57.194279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:57.194279Z digest=sha256:c8bbc3992a93407d8a02015594c8ebf8386cf65895735f146e34ad597342d8ca

Observation dff48fa4-f852-4f2c-843c-a6a25f92051e · outbound

This paper cites Generating diverse and natural 3d human motions from text.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Generating diverse and natural 3d human motions from text

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:57.356942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:57.356942Z digest=sha256:6440256f6be43e2f34edaf1a7a83390099c87b1fee340c64fca34902ae511ee5

Observation fa0611e9-d169-49ac-8e34-ebb202c2a7a7 · outbound

This paper cites Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advancesin Neural Information Processing Systems, 36, 2024.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advancesin Neural Information Processing Systems, 36, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:11.071203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:54:57.480981Z digest=sha256:0b6c8b5217cf54308e87a313ab68455736d9eb247d80ce989d25323f2042ab22

Observation 8b3db17c-7c61-421d-81f4-80040385add2 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:57.627224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:57.627224Z digest=sha256:21db331bf46fc8142b98ef662d30d3eafd0a560be148d490fae113a57375d9c0

Observation 1c21bda0-ecc1-410b-b15d-9684fde545d4 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Improved baselines with visual instruction tuning, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:57.739675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:57.739675Z digest=sha256:8d8804ef2b8b982c3b2397a721491587e5bf0dc79569c5ba81f0e063de8aafd4

Observation be59e44c-9ed1-4551-8228-2bee0d92fc49 · outbound

This paper cites A large-scale rgb-d database for arbitrary-view human action recognition.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model A large-scale rgb-d database for arbitrary-view human action recognition

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:10.900308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:54:57.899801Z digest=sha256:a11729b5dbdc86e4d99ca945afd8d68b36d6a6c9b5784c30c26b93951490a8e9

Observation 5113d07b-6e8d-48e6-b842-1a78fd74b251 · outbound

This paper cites M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:04.531949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:54:58.024126Z digest=sha256:de0af844ad2002ea00cc8498d13d545e134fdffb779c7b790afa151795a20848

Observation 74225f25-b5f1-48d9-a88a-24eecdcaf43f · outbound

This paper cites Scaling large motion models with million-level human motions.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Scaling large motion models with million-level human motions

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:10.589979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:54:58.170476Z digest=sha256:98e6c1f04509629c042c09dbff8e281d4c067e6365b13afb4e96ef9fae372e27

Observation 2fe7aa93-f7bd-4160-a6be-71415b12f3f5 · outbound

This paper cites Autoregressive image generation using residual quantization.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Autoregressive image generation using residual quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:58.286611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:58.286611Z digest=sha256:62c66d615a93c36bf0aa4642c52d8448a4df67d098da6da8a0ce0d9748015bd7

Observation 01277ddf-1995-4ea3-b713-52a20ec98685 · outbound

This paper cites Temos: Generating diverse human motions from textual descriptions.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Temos: Generating diverse human motions from textual descriptions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:10.171990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:54:58.404718Z digest=sha256:7ee228399a78def4c10b6b09b73b92df305985d270aa2cdb4d4835d0fc779fb5

Observation c7e44568-127d-4782-b635-1021a5a3175c · outbound

This paper cites Language2pose: Natural language grounded pose forecasting.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Language2pose: Natural language grounded pose forecasting

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:09.830066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:54:58.488840Z digest=sha256:3d5953876e354129c18d5fa4321806f99e900243d145e0cc81c011af0247a5c1

Observation 41af36a7-045a-4bf5-b95d-ddffd17f0d17 · outbound

This paper cites Motiongpt: Finetuned llms are general-purpose motion generators.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motiongpt: Finetuned llms are general-purpose motion generators

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:09.518557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:54:58.572722Z digest=sha256:f304b8eb1556df7e8b55e7fbe65a0f3f257a332ba7e13a679b0bc55b266060c3

Observation 6f82544c-813a-4c9c-b787-9e4c43a67498 · outbound

This paper cites MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:58.674215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:58.674215Z digest=sha256:83ebdca2e45c4250e4a6ff9cdb5960e82958d39d1affa3b0e4a5b8fff482da14

Observation b5109d80-5552-4ce4-8338-eef799e9e219 · outbound

This paper cites Recurrent network models for human dynamics.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Recurrent network models for human dynamics

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:09.209540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:54:58.788216Z digest=sha256:16b1309478eef2a43dd341558adf926468f406d15251443f8bbf2f860940eaa4

Observation 34812253-a8e0-43bf-ad8c-09a39ec1b6b2 · outbound

This paper cites A neural temporal model for human motion prediction.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model A neural temporal model for human motion prediction

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:08.899001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:54:58.901849Z digest=sha256:364504e9b750630dca0378cf4b504ca02a76f13cec9239ca1a03a166f745eed0

Observation add6134e-aa7e-4f2e-996d-651df9ac534f · outbound

This paper cites A stochastic conditioning scheme for diverse human motion prediction.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model A stochastic conditioning scheme for diverse human motion prediction

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:08.524160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:54:58.988702Z digest=sha256:a51faf070edbd8f8e834d73ccdafdbe318eda2b06af8b3701013ddfe258f2ca8

Observation f2a32cee-4d38-4827-9e1b-3a6acd8c1de7 · outbound

This paper cites Learning diverse stochastic human-action generators by learning smooth latent transitions.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Learning diverse stochastic human-action generators by learning smooth latent transitions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:08.190878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:54:59.102481Z digest=sha256:42e7a8ea673d1aa332721adad61cc2c13028aa6d19501933db6c7cc93540d869

Observation 3a35875b-133f-408f-856e-523ea6a1330b · outbound

This paper cites MotionChain: Conversational Motion Controllers via Multimodal Prompts.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionChain: Conversational Motion Controllers via Multimodal Prompts

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:04.395699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:54:59.216321Z digest=sha256:4cda5e0146d47c5008adcfc1b03e64ce253aa4e0100ce7d51deb0a93547d7d31

Observation c0d45bf0-6888-4ddc-82c2-125238f37983 · outbound

This paper cites MotionLLM: Understanding Human Behaviors from Human Motions and Videos.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.349922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.349922Z digest=sha256:d746aed15a0b7a5d0c4d3c3f71a3fd371cb429c3e8aed172291c43ebafeb6bfb

Observation e8293c10-4cb8-43e5-b917-088f5d8a3fb0 · outbound

This paper cites Large motion model for unified multi-modal motion generation.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Large motion model for unified multi-modal motion generation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:07.920463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:54:59.472583Z digest=sha256:9035f1915d3dd831643558f57b83ffe9ccf8b7f5c462d59da832d3d6514f2c58

Observation 1dda8c6c-d1f5-4128-9e8f-1cc6bfa07b58 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Flamingo: a visual language model for few-shot learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.555702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.555702Z digest=sha256:225302b500bf8283ed44c706fdef237e994877450b6b333f23e8e9313d79ad60

Observation 341ddcb9-d858-4868-aabf-0b7c45401c19 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.654282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.654282Z digest=sha256:067cac0bfa7d1b590654ae0d988b580b3ac56b37a59db0bc6999f932ff1c348f

Observation 1b49001d-75f1-408a-997f-aed91bd5b1ed · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.795393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.795393Z digest=sha256:b42b9af7c74ff7ec83e29bbd89c4d486d64ac671fadb4c555f7ecc1fb513f160

Observation 20663abe-3cae-4f82-a4be-36d4943a4bcd · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.894391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.894391Z digest=sha256:75a51d61ae360db58505379fae2e6a4156c662934234bfe11b3cc9767c22fa10

Observation e480fb9c-08b8-4a63-ac25-0bbcf0bddad0 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Lisa: Reasoning segmentation via large language model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T21:54:59.990725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:54:59.990725Z digest=sha256:9560b2023954b8fe8be3e511a13a708ffb77cbad5b856224d0e783bc0ce2f334

Observation efa3ce64-3ad0-41e3-b8c0-f154ae49a476 · outbound

This paper cites Neural discrete representation learning.Advancesin neural information processing systems, 30, 2017.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Neural discrete representation learning.Advancesin neural information processing systems, 30, 2017

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.076186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.076186Z digest=sha256:14776ee786be204ad35ea3646c2ccab061cdc417e57f58d36296e709b4623d28

Observation ce96e7aa-98de-4a25-b0fe-a707dd9a63cf · outbound

This paper cites Locally hierarchical auto-regressive modeling for image generation.Advances in Neural Information Processing Systems, 35:16360–16372, 2022.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Locally hierarchical auto-regressive modeling for image generation.Advances in Neural Information Processing Systems, 35:16360–16372, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:07.655546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:00.164563Z digest=sha256:ce1190cbdb53ba51325df0ffd32ad61507a42752492baed46fd41cc594d60e30

Observation 085c8fec-7879-4d83-9ba4-3c5a4fa9332b · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.299651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.299651Z digest=sha256:6502781fe278cdff6fbb43e4ca705d89eeba324fb9a6fbb3c074594393a3a470

Observation a7c21b9c-1ed6-46bd-a882-91d9cc60f152 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Finite Scalar Quantization: VQ-VAE Made Simple

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.438317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.438317Z digest=sha256:69900c143651ccea6157ad1f5ca389558491ea7d9ad10e401c095f8d601eb63c

Observation 19116fe6-42eb-4ba8-9981-4744899a9c44 · outbound

This paper cites The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:04.212186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:00.545792Z digest=sha256:afeb2841df962405af3d29e49e0dad9effa3d926acb85edf0d401b3afc91362b

Observation 37024144-556a-4f56-b275-8c5f52378f5e · outbound

This paper cites HumanTOMATO: Text-aligned Whole-body Motion Generation.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model HumanTOMATO: Text-aligned Whole-body Motion Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.655357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.655357Z digest=sha256:ae039266670221a5d6432509dcf3cbe1ac68d4eb44cfdeeab7db6545bb779efa

Observation 3ed91865-da65-4d6e-a1db-a804263ec9bb · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.759463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.759463Z digest=sha256:2d2b3dfdcbddd1c398fa7f0a5b8fc0d23cf67fbd284287f9120db52c004c7d42

Observation d1bf1a21-a541-42d9-9561-b1a77cfae242 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.851447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.851447Z digest=sha256:76d08f27719975d0b8e613020f0f9f0de4bd2bc9dea4feb3effc6306fc6d13fb

Observation 0dd5e36c-bab4-44e7-aca4-8ca3ae7de1e4 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:00.957517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:00.957517Z digest=sha256:f5a7b9511d97b2ae7a9b9a992daf0508ab185b1db24cbfda00126f078546cf49

Observation 07474f84-7df4-418b-832d-84842521fee3 · outbound

This paper cites Sigmoid loss for language image pre-training.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Sigmoid loss for language image pre-training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.067511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.067511Z digest=sha256:12e277b7b59208dafbee1a56b367e6644bb12a5fd76cbbbff3d427c95e61750a

Observation d0cc928d-8f53-4146-9b96-1e97f4d9a744 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.153815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.153815Z digest=sha256:965156665f80b8d1f6af9711adb8a1261fc9ae800e33965d3aba1ad4929b9e79

Observation 831fad93-aa76-4d93-82c8-7cf4f39ea056 · outbound

This paper cites Smpl: A skinned multi-person linear model.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Smpl: A skinned multi-person linear model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:07.365209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:01.246734Z digest=sha256:23e892bf2993a729bda5fa4cd997de090a2f84d6c74a5f9f9382c2e6adaf3828

Observation 4ae44264-f948-4cca-937d-02e8bf0a804b · outbound

This paper cites ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:04.023236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:01.311777Z digest=sha256:39c8c63d1897adc66e2ce8f3153a03e098eb923b0716c056bde5a1711dc86bfe

Observation d9f424e3-a602-4eb6-bb9f-bff04676e6e3 · outbound

This paper cites You only look once: Unified, real-time object detection.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model You only look once: Unified, real-time object detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.439383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.439383Z digest=sha256:3cdac57b80273771da0b16d8b53b08f9d03c837748922bbc868990b6259a93cb

Observation 16abd518-9938-4017-989e-04af108538c1 · outbound

This paper cites Wham: Reconstructing world-grounded humans with accurate 3d motion.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Wham: Reconstructing world-grounded humans with accurate 3d motion

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.545093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.545093Z digest=sha256:1df79b1ae94ead8b45cd4aca530b9a9df7d17ba87cb65400b822e2a29d2c4881

Observation 1ef57ef7-c289-471c-9781-0d3376cbafff · outbound

This paper cites Perpetual humanoid control for real-time simulated avatars.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Perpetual humanoid control for real-time simulated avatars

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.629310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.629310Z digest=sha256:51320b36a841eb6bd9302e455813b338f1a01c7d9b1d2074956400a9a67785f6

Observation a4e663f2-0d61-4a5d-acaf-797a31be13c3 · outbound

This paper cites MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.735535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.735535Z digest=sha256:1f376494e5480291fb49caf3b7506e54b7bdb2044e91a0f065184b4fff4b9db5

Observation 519ee99e-1e62-4771-9393-ceb389115477 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:01.847632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:01.847632Z digest=sha256:d9cc0ec0a3f2bf0c489727864ccb1d6879c7ae85953712ae693030c737307404

Observation 8e52c3ca-4481-48e3-b569-1d8350ff632b · outbound

This paper cites Posescript: Linking 3d human poses and natural language.IEEE transactions on pattern analysis and machine intelligence, 2024.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Posescript: Linking 3d human poses and natural language.IEEE transactions on pattern analysis and machine intelligence, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:07.069823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:01.922185Z digest=sha256:1d9992e3ac76d4aae5ac210438695ce28b9b1e11f2aeb6dec843651dd1b8d222

Observation f031606a-9685-4d35-967b-9da7383f2d0e · outbound

This paper cites UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:03.873453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:02.001344Z digest=sha256:4ec98a9624926aec665e5cb05fc401407faf10674d47fa34da94a2b58998b695

Observation 26d89571-6974-453b-96e8-826bd48f83a2 · outbound

This paper cites The kit motion-language dataset.Big data, 4(4):236– 252, 2016.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model The kit motion-language dataset.Big data, 4(4):236– 252, 2016

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.837633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:02.083256Z digest=sha256:915dee4023cdfec3bd878afd5e74c99b6c23168784b49371aaa757f552f8017a

Observation c88c3d4d-616c-42ec-b458-2c6cfc015bd6 · outbound

This paper cites Amass: Archive of motion capture as surface shapes.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Amass: Archive of motion capture as surface shapes

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:02.186443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:02.186443Z digest=sha256:425faaac739241fc3aab7f0d1f16d4ad395b97d8543adb6895bd5b9c94df370a

Observation e600edf6-3a56-4126-bdd1-9dcb55995bc4 · outbound

This paper cites Recovering accurate 3d human pose in the wild using imus and a moving camera.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Recovering accurate 3d human pose in the wild using imus and a moving camera

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.713662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:02.270194Z digest=sha256:af8de43d7875dec81c68b45b9a9280111361a24278d318f5067a50c4d85317b6

Observation 95a54321-6915-4f40-b11e-4f536679466a · outbound

This paper cites Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advancesin Neural Information Processing Systems, 36:25268–25280, 2023.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motion-x: A large-scale 3d expressive whole-body human motion dataset.Advancesin Neural Information Processing Systems, 36:25268–25280, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.590197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:02.345136Z digest=sha256:3538e91936b97f10243b4b0ae4d8df13cc8a83ae6a0565e96dfe8824def6a884

Observation ac2b0286-e91c-4598-8878-12aa803647ed · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:02.419755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:02.419755Z digest=sha256:3a1eaf6447bba5e940370d5901d642aba3a7195ff547e7e624b80d069da19c01

Observation 4b1eb614-d23a-4f42-9002-67d3e549c225 · outbound

This paper cites Executing your commands via motion diffusion in latent space.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Executing your commands via motion diffusion in latent space

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.463606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:02.501633Z digest=sha256:a698796f0674ea9f782575a34a98fa428b74970e04da543a699e9f30fd6ada4e

Observation 2938cc62-ac40-4c9e-b2d1-53d03875db61 · outbound

This paper cites MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:02.580215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:02.580215Z digest=sha256:2fbb9431f418966cc21682b1f7ffb673ebadbbf5b6861e00bfa6d3076deb86c4

Observation 159b2500-41a7-49cc-ac83-ec77fdb8669a · outbound

This paper cites Generating human motion from textual descriptions with discrete representations.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Generating human motion from textual descriptions with discrete representations

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.313231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:02.747962Z digest=sha256:277ce4295aeecba098a279431582b7e9e51140adeac56329a162a8f3e3477abd

Observation 63b8103f-e63b-4970-bbd4-3bcc41f92c0e · outbound

This paper cites Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:02.826806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:02.826806Z digest=sha256:444c31b9069110ea97bc47a9c113c9f6f197ac7f713b2ff1a59f011236216ae3

Observation 5c5c879b-afdd-482a-875b-d6ae4e2c782f · outbound

This paper cites Avatargpt: All-in-one framework for motion understanding planning generation and beyond.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Avatargpt: All-in-one framework for motion understanding planning generation and beyond

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:06.135640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:02.910283Z digest=sha256:3215bcdb90d70f94f436b6492902d107411fc3dca6bde001816a73b97506831f

Observation 63da06a9-a3db-4858-89f5-3acdbde51454 · outbound

This paper cites llamacpp.https://github.com/ggml-org/llama.cpp, 2024.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model llamacpp.https://github.com/ggml-org/llama.cpp, 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.927027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:02.982817Z digest=sha256:ada453fd65ae9987a56754b41b5325907abb67704aff164ba5c435b265a61d22

Observation 60cd66d7-ba14-42f0-8102-68c5edc2b3df · outbound

This paper cites Parco: Part- coordinating text-to-motion synthesis.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Parco: Part- coordinating text-to-motion synthesis

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.745015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:03.075835Z digest=sha256:760f619cacd002f44acbb1dc804c307735a3a74c429eccfb63dd48265d6fa131

Observation b5ff2380-7174-4532-9f71-1442afa78375 · outbound

This paper cites Fg-t2m++: Llms-augmented fine-grained text driven human motion generation.International Journal of Computer Vision, pages 1–17, 2025.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Fg-t2m++: Llms-augmented fine-grained text driven human motion generation.International Journal of Computer Vision, pages 1–17, 2025

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.563158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:03.191406Z digest=sha256:04a1ba047ef8121fab8d28a9349e06ef2ff16a5ee38cb34cd51c66d8826b5cf1

Observation f4fe6324-23d1-49cd-8c34-8ab627fa4359 · outbound

This paper cites Motiondiffuse: Text-driven human motion generation with diffusion model.IEEE transactions on pattern analysis and machine intelligence, 46(6):4115–4128, 2024.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Motiondiffuse: Text-driven human motion generation with diffusion model.IEEE transactions on pattern analysis and machine intelligence, 46(6):4115–4128, 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.397712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:03.264774Z digest=sha256:ee546bd520ad8efc3baa86b7b7227475725f861340e0d211d08d866f33b3d197

Observation e8dd1cb8-d25d-4708-83cc-06817a31f57f · outbound

This paper cites Microsoft coco: Common objects in context.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Microsoft coco: Common objects in context

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:03.339648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:55:03.339648Z digest=sha256:de329e77cffce7256d47fefcec9653d955922ed332f9259261c9ed2772924456

Observation 359e48f1-c9e6-47d1-8b3e-316920fdb2b0 · outbound

This paper cites Posetrack: A benchmark for human pose estimation and tracking.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Posetrack: A benchmark for human pose estimation and tracking

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.229691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:03.411500Z digest=sha256:7b07aea1bb22f8c1e3aa8482b57a790baa2bfcdd54ba1fa6a0e92f852205b5e1

Observation 4d4680ae-0b05-493f-b552-7d70fa057295 · outbound

This paper cites Resolving 3d human pose ambiguities with 3d scene constraints.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Resolving 3d human pose ambiguities with 3d scene constraints

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:05.051157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:03.476511Z digest=sha256:17ad1d79c61de77cfec158ea1482d3ec4a5d50ea43b675b75519c5653550bb6a

Observation a18f99fb-b686-4ff0-9033-a97193bf97dd · outbound

This paper cites Behave: Dataset and method for tracking human object interactions.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model Behave: Dataset and method for tracking human object interactions

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:04.881522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:03.555056Z digest=sha256:c69ea72b707b46ca5534e1aa107addeb1543178186a2bd6dc84b213b08389a0c

Observation a0088ee7-8b45-4f9a-9ff4-ac15d6d8754c · outbound

This paper cites the left hand is positioned below the right hand.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model the left hand is positioned below the right hand

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:55:04.677201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:55:03.686435Z digest=sha256:7b127c68cf6cfb493130061c82cd0192dd52f6ab3e5a3d9c523d2e96cfbf8dfe

Pith citing papers

Observation f3f0f2b1-ba20-4233-91d7-cb8d54d2cfd2 · inbound

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation cites this paper.

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T17:06:54.531502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:06:54.531502Z digest=sha256:778acdb808863611e33ea5ad1fdd1d882b01949dd156e5e446bbab036fa8ce1e

Observation 76dac94f-3c6c-4fa7-b748-124078c28352 · inbound

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos cites this paper.

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:47:57.514056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T13:43:26.460480Z digest=sha256:7925f417c659f0a61ab32523a5630f11f936933fe95037982e844649b501e5ba

Observation 5fa7384a-329a-43d7-9e3c-f6db91017093 · inbound

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval cites this paper.

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T00:39:25.303896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:39:25.303896Z digest=sha256:b562fbc21b3e039111941e31ff414a8c89dc7aa0840356c647d220e851f340db