Pith. sign in

Paper Citation Record · LEDGER

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios

As of 23 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2412.14643.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14643 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:05:52.798589Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:41:08.431780Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-08T04:41:08.649124Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy48
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2257bc8d-735d-449e-accf-c92b15c0c00b · outbound

This paper cites Augmented-reality-assisted bearing fault diagnosis in in- telligent manufacturing workshop using deep transfer learning,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Augmented-reality-assisted bearing fault diagnosis in in- telligent manufacturing workshop using deep transfer learning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.802033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.359678Z digest=sha256:4d1e4244f3e0d91e910214c277a736bbde016ebca61b201aad107da98ed693ae

Observation faa7a18f-70f0-48ac-8b48-ee334cf69073 · outbound

This paper cites Understanding user behavior in volumetric video watching: Dataset, analysis and prediction,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Understanding user behavior in volumetric video watching: Dataset, analysis and prediction,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.778852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.366782Z digest=sha256:e9e0c6dc4fa435ce40247ff252577db9b2083b621301f930078e8c4777cc2793

Observation 6cf1f5d2-d2e1-4fb1-8295-b464e2b5c47c · outbound

This paper cites Actions speak louder than goals: Valuing player actions in soccer,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Actions speak louder than goals: Valuing player actions in soccer,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.754560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.372344Z digest=sha256:e9637e28bc0049722d6208c1147ac56cf7a7481159a7361efbb01c0176f71d92

Observation e2cdc847-558f-4a8f-b410-3e98d0ddaf13 · outbound

This paper cites Pass receiver prediction in soccer using video and players’ trajectories,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Pass receiver prediction in soccer using video and players’ trajectories,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.729739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.378338Z digest=sha256:8853a68e68800df1164a245600fd69d72a1f865cabe415e30a951fabdd0f12b9

Observation b62ba296-cb4d-4bf0-9bf7-5eb0abbdc179 · outbound

This paper cites Shape-aware text- driven layered video editing,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Shape-aware text- driven layered video editing,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.678068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.389028Z digest=sha256:b2b7a167fc1f54948338369f94b2b9f2c3d728e297056960cdaa97f0ba29190b

Observation 0f6ce562-0c5d-4ad0-a797-acec2c5537fc · outbound

This paper cites Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.396132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.396132Z digest=sha256:55d507ac273721c2a8907c2066a52e523d5dd027452bde2a36526acc90e7734f

Observation f7338412-47e0-46d9-876e-f4286286a4ad · outbound

This paper cites Lcr-net++: Multi-person 2d and 3d pose detection in natural images,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Lcr-net++: Multi-person 2d and 3d pose detection in natural images,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.655494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.402933Z digest=sha256:9219f982aa8942de2c09ee2fa571084174bc408b83bbddc3564572c761a80ab6

Observation 7c978822-c8ef-4687-945a-727af9d6d5cd · outbound

This paper cites Doublefusion: Real-time capture of human performances with inner body shapes from a single depth sensor,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Doublefusion: Real-time capture of human performances with inner body shapes from a single depth sensor,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.634701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.408311Z digest=sha256:6187a120b1117727b53010156e2e86e3663148a336a188c9b7b91a3fe3deadce

Observation 24e0e391-0fbb-4df1-8280-0765fff71264 · outbound

This paper cites Articulated human detection with flexible mixtures of parts,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Articulated human detection with flexible mixtures of parts,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.617813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.414961Z digest=sha256:cdc5328b2d0ff5e2ebad264b4a0f50bfe4b8bb26d0edcd68823d07ef3b09f477

Observation 58e3f427-e4b5-411d-b176-38d456d3e05f · outbound

This paper cites Alphapose: Whole-body regional multi-person pose estimation and tracking in real-time,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Alphapose: Whole-body regional multi-person pose estimation and tracking in real-time,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.597831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.420881Z digest=sha256:ce16c670450538fb6c30b1392a0d224fb1d798a5810ff01179612fd73488506c

Observation 9b1aa88b-a69c-4077-89ae-51b84f008b26 · outbound

This paper cites Spatial and semantic consistency regularizations for pedestrian attribute recognition,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Spatial and semantic consistency regularizations for pedestrian attribute recognition,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.426058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.426058Z digest=sha256:3285f1f4f2e5ad0c0aad0fbba5c81dbe0dc589bcd334c0d422c9d95452bee461

Observation 1b27c801-059f-4776-ba74-6342af80cb6a · outbound

This paper cites Learning disentangled attribute representations for robust pedestrian attribute recognition,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Learning disentangled attribute representations for robust pedestrian attribute recognition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.581051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.442225Z digest=sha256:7638afe19819b65b287bfa1c86464d825a8ea120dbcfa6bf9a6bd51b38f86ead

Observation 47878e59-510a-424d-8b3e-96f3514d12f3 · outbound

This paper cites Label2label: A language modeling framework for multi-attribute learning,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Label2label: A language modeling framework for multi-attribute learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.449494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.449494Z digest=sha256:b406a2fead1d71900035996723c08569161ca808258c3611f2e6b1964811199b

Observation 8c8e2567-f508-4b2d-80c5-aeff5d0e7ab2 · outbound

This paper cites Ferret: Refer and ground anything anywhere at any granularity,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Ferret: Refer and ground anything anywhere at any granularity,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.563009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.454981Z digest=sha256:759c2b4f012f73ed74750179f5233772f1f50f11a4c7bb06b0e12c80077bb233

Observation ca06b9a7-6d0c-443d-a7d6-8a31b9da209c · outbound

This paper cites Aiparsing: Anchor-free instance-level human parsing,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Aiparsing: Anchor-free instance-level human parsing,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.540282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.461320Z digest=sha256:24f0a0ce160a76117dc39650b24cd93f3522a3c68bf5604f45ea0f41e12dff10

Observation 7573a1d9-8f73-47de-8625-a9a81def0ee1 · outbound

This paper cites On the correlation among edge, pose and parsing,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios On the correlation among edge, pose and parsing,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.517641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.466475Z digest=sha256:f133c1767ffda0d31ba7f15fb29d5316cfe48400eb69856195645b5e915ffdbb

Observation f186663a-9a79-4af6-8909-3212801f4a10 · outbound

This paper cites Clothes co- parsing via joint image segmentation and labeling with application to clothing retrieval,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Clothes co- parsing via joint image segmentation and labeling with application to clothing retrieval,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.490545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.471554Z digest=sha256:ce2b55c820319504ea5f7cdeb1f4b019029fddf359cafeb67f663115179faecf

Observation fcc4c40a-5a39-4453-ae3b-98691296d25a · outbound

This paper cites Beyond appearance: A semantic controllable self-supervised learning framework for human-centric visual tasks,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Beyond appearance: A semantic controllable self-supervised learning framework for human-centric visual tasks,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.477493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.477493Z digest=sha256:dc25682266b8f0c1c98e975c00d90383a27de5ac2187952d22627fddba81290c

Observation 23164368-e684-4286-99e4-906a90cc5386 · outbound

This paper cites Motionbert: A unified perspective on learning human motion representations,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Motionbert: A unified perspective on learning human motion representations,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.455864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.483137Z digest=sha256:a9fe9ed5748039e1d9ac6de64d31038fa43f75d20f313a6a674bc494874af534

Observation d6d1f53d-d693-4cfc-8a96-8dab3290cc7c · outbound

This paper cites Unihcp: A unified model for human-centric perceptions,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Unihcp: A unified model for human-centric perceptions,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.430058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.488914Z digest=sha256:02aab40e0d49cde075eaa331e6e3cf9d8e25827c2aa1ccd4202fcf0b0e88906c

Observation b2b84719-cbb8-43f1-b190-5bfb13919f8c · outbound

This paper cites Hulk: A Universal Knowledge Translator for Human-Centric Tasks.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Hulk: A Universal Knowledge Translator for Human-Centric Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.496298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.496298Z digest=sha256:630eba8ae9e4b5a912fc75a3347fc9ed7d35da3bec2ae644b604e867d1c90d9e

Observation 33504922-fad3-42df-bb8f-33bfe3ea5163 · outbound

This paper cites Modeling Context in Referring Expressions.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Modeling Context in Referring Expressions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.502617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.502617Z digest=sha256:ec3309fdfc41fe051b02d87a35a8c2bdbec12841aa38fc84e57162d382e95f2e

Observation c8bbbb50-72d9-46b9-b1a2-6101871cd288 · outbound

This paper cites Hrformer: High-resolution vision transformer for dense predict,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Hrformer: High-resolution vision transformer for dense predict,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.389799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.509388Z digest=sha256:8cdd3cca1b710e00a46fde11a18cbf156babb33237526c190d20f5f9dd8017d8

Observation 46933ea5-0cd6-46f2-8ded-8a1838817a64 · outbound

This paper cites Devil in the details: Towards accurate single and multiple human parsing,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Devil in the details: Towards accurate single and multiple human parsing,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.369191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.514612Z digest=sha256:eb347ff32bce2f65ba579eb6776331f3895b9f4c16a27b17db207294418a9eb3

Observation e13fd90a-c8c7-4913-ba82-7389a50aff88 · outbound

This paper cites Higherhrnet: Scale-aware representation learning for bottom-up human pose estimation,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Higherhrnet: Scale-aware representation learning for bottom-up human pose estimation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.347384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.520128Z digest=sha256:37ed296819af4ca989309dea98df75032e773c11a7e7a2c96deca00d65508dbf

Observation 776510bd-6807-4668-807b-0c179e2c412c · outbound

This paper cites Bottom-up human pose estimation via disentangled keypoint regression,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Bottom-up human pose estimation via disentangled keypoint regression,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.285914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.532912Z digest=sha256:3e61e41b09940583b7b60f356f12904616ec3b95c5c345ab7eb6432de29f3c12

Observation 8b0ad8a7-1f14-4186-ab7b-4cdb48bebf36 · outbound

This paper cites The devil is in the details: Delving into unbiased data processing for human pose estimation,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios The devil is in the details: Delving into unbiased data processing for human pose estimation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.263450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.540586Z digest=sha256:b9526eb346191d25159bd69540bd89252fdc03f76dd88ad35e034bb794244243

Observation 618fbdc6-8af2-40ce-8bca-4b8210cccec5 · outbound

This paper cites Graphon- omy: Universal human parsing via graph transfer learning,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Graphon- omy: Universal human parsing via graph transfer learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.229669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.546570Z digest=sha256:7ae64d6394595f69706e7a087794f4a644558670a0825a0c429c0cecb6bd0f19

Observation d55f2227-48c0-4bec-8815-3377b0fda5cd · outbound

This paper cites HAP: structure-aware masked image modeling for human-centric perception,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios HAP: structure-aware masked image modeling for human-centric perception,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.208539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.552589Z digest=sha256:680bee42575b7d90d49f8d20fc9d0ec860a20eb78a62d61084f629475e9378aa

Observation 1f02bd3b-4587-4e93-8f97-892f84078b06 · outbound

This paper cites End-to-end object detection with transformers,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios End-to-end object detection with transformers,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.180566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.559017Z digest=sha256:5c7faddb49163eddfd22761bb69972f6f96647c0471c5b489827e7190677a838

Observation 75e4671a-b182-4198-9690-345b3d556abb · outbound

This paper cites Referring human pose and mask estimation in the wild,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Referring human pose and mask estimation in the wild,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.157833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.564919Z digest=sha256:fbefe12edf19ac9de4fa738fba69ab25e764d799cc11ecb463244796eed81324

Observation b877c568-38d6-4e8f-b5af-2066328f8fd3 · outbound

This paper cites Deformable DETR: deformable transformers for end-to-end object detection,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Deformable DETR: deformable transformers for end-to-end object detection,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.126325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.572486Z digest=sha256:93f4edb5f22190cb4334f0eab6267bbde61bcdc307942a5a65ef392c58849037

Observation aa1043d6-2a6e-40ac-be15-0ee506093a0a · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.095453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.577839Z digest=sha256:1ec50b00dafbd5bb75404657ed8f0d8b1011c8dc9e2851b8a2ba565f999bb7ad

Observation 1030eebd-1c87-466f-9a70-f36b73b15408 · outbound

This paper cites BART: denoising sequence-to- sequence pre-training for natural language generation, translation, and comprehension,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios BART: denoising sequence-to- sequence pre-training for natural language generation, translation, and comprehension,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.074980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.583161Z digest=sha256:51043023b5decef6d6a41ce9ba6a67d2dd8b7d193cb4f90e1ab1d08f955ed906

Observation ebe5cec9-cc6d-4797-9f1d-f6fa53f267e3 · outbound

This paper cites Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.588882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.588882Z digest=sha256:6360250e2d1daa0ef596825f16efb090d180bb6d09e28a1c60054a74fb8e6484

Observation 263bee27-1c90-416e-86a9-6355871844e9 · outbound

This paper cites OFA: unifying architectures, tasks, and modal- ities through a simple sequence-to-sequence learning framework,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios OFA: unifying architectures, tasks, and modal- ities through a simple sequence-to-sequence learning framework,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.051174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.594682Z digest=sha256:75a5a945d6695483cc99e73b9cf00b1b5ed91edc3255a2c74fa0d49d28aa2dc2

Observation b2c39058-2a71-471a-aa48-d883a7c4f6a0 · outbound

This paper cites Taming transformers for high- resolution image synthesis,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Taming transformers for high- resolution image synthesis,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.027829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.602414Z digest=sha256:532521bb8f2c377d38eb243084de61c3fffbd9000b8b257ec8f78a92de422011

Observation 46111511-ad05-4352-b9c8-4ec89b24682c · outbound

This paper cites Deep residual learning for image recognition,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Deep residual learning for image recognition,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.611704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.611704Z digest=sha256:5fb0ed7264473e5715e3ec6788ceee8db25bc28cdb0542e6378a777c36939f9b

Observation 7495704d-c6e6-463d-93d1-3f53beb00581 · outbound

This paper cites GPT-4 Technical Report.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios GPT-4 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.618424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.618424Z digest=sha256:383f525e9c1cfb7af17bf8234e514ee4bbd0d671deb3a198cc3ec4fa3a699bb8

Observation f25d076a-bce7-45c5-a6e4-c3dbcadb0271 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Chain-of-thought prompting elicits reasoning in large language models,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.001199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.625651Z digest=sha256:99d5c12b30551f0fcd4f71931af2faa297c26a41e07083ff913e1d8cbcef7999

Observation c4ee94f7-8062-4e2b-b599-796d33be67d3 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.634361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.634361Z digest=sha256:61d378cb6d1e923c715c1d3bb18a1cae9806b33ea33700f94ab2d68b2c504f8a

Observation 065e035e-cf9c-446b-81e0-e9b415dcd9ac · outbound

This paper cites Coco - common objects in context,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Coco - common objects in context,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:53.981374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.640719Z digest=sha256:865a938e738f85672d49c4a9ca855a907afc0467e6910659bd8d2d5b91a897bf

Observation 6c2a599e-39c2-4537-9cc0-352bf55a1810 · outbound

This paper cites Instance-level human parsing via part grouping network,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Instance-level human parsing via part grouping network,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:53.960313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.651442Z digest=sha256:bb296271b5e5c7aa920e5b6c712e509b41761d24d09740dd1a4f2cbffef3e688

Observation 1057e573-e365-4dba-bdba-c087d4d19ed4 · outbound

This paper cites Conceptualcaptions,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Conceptualcaptions,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:53.935031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.658253Z digest=sha256:b3c4b4de47dd8dd1750b23342f2a2dece7dde279f0bfd8b606a0b2bab7ba72ea

Observation 1d891e20-9d55-4b78-993e-60e956042a38 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Visual genome: Connecting language and vision using crowdsourced dense image annotations,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:53.911066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.663422Z digest=sha256:7ab0928bfec7b67ae7c69ed28e15571dd5dcd4910b34659a30a093db98400b4b

Observation 983fdd3d-ef07-4082-80f5-2a84c0c67777 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Referitgame: Referring to objects in photographs of natural scenes,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:53.888033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.670162Z digest=sha256:65885540907f47e1451f77ef741d5f441ba68ef25c38f2a23b44b744b9e6d85d

Observation 8c78b1da-8ef8-4f84-a4a7-2bdce5f8d810 · outbound

This paper cites Large-scale adversarial training for vision-and-language representation learning,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Large-scale adversarial training for vision-and-language representation learning,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:53.865242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.677838Z digest=sha256:d4c7821b37ebfe6748de0328b98862bce998dc3b03dfbe2c892cd55a362638cf

Observation 1c49f622-24cc-4591-8c1e-f431c265cc22 · outbound

This paper cites An empirical study of training end-to-end vision-and-language transformers,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios An empirical study of training end-to-end vision-and-language transformers,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:53.843231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.683423Z digest=sha256:2ea84fd0adba6304d47e2a6e6c73b323bb99533f4b5d41d5b5ec57a58c52c261

Observation 0dbd212c-6c3a-4e26-b0c8-309620fe49cc · outbound

This paper cites UNITER: universal image-text representation learning,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios UNITER: universal image-text representation learning,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:53.820009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.688296Z digest=sha256:58986c9fba19ebb8b3aef0af3f057a441c20cb0eb62719f7535d37726566a7fa

Observation 055fc02d-1ab8-4b48-9563-0610d3724eac · outbound

This paper cites Uninext: Exploring A unified architecture for vision recognition,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Uninext: Exploring A unified architecture for vision recognition,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:53.798298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.693090Z digest=sha256:ca1bbc12575ee88f06c4683ad02f44121e4b8e4db222fdf82358379cb665f23f

Observation 25e2b1e8-13b0-468c-9abf-6d74a846ec49 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Generation and comprehension of unambiguous object descriptions,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:53.779359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.699056Z digest=sha256:5ff7ffb96615b60e58f6569d243e82e9647bdad1da1fa9765139c44346e01e94

Observation 213cf844-e41a-46b9-bd36-84fc644b5ba9 · outbound

This paper cites Ferret: Refer and ground anything anywhere at any granularity,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Ferret: Refer and ground anything anywhere at any granularity,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:53.762199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.707679Z digest=sha256:05739ee77262162392f8d05567254ef1535d67a230fbb13b88f794249f457081

Observation 3c1dfc0d-32df-4ac7-a9bd-ef816faf07ec · outbound

This paper cites Cider: Consensus-based image description evaluation,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Cider: Consensus-based image description evaluation,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.713838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.713838Z digest=sha256:ce05c4c89487f00ac6b3cffe1de32b0549507ad290fae3bcd8c7a1a4293a856b

Observation 368d72de-ae0c-4dbd-a686-948f6aa9fcac · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:53.698851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.730138Z digest=sha256:5294bb8b351371022ec1f4b12b82e8983d60e368b56a6dcc368ecadf7851362b

Observation b3fbf443-8d2f-46cf-a5c6-2f4c810c4184 · outbound

This paper cites ChatPose: Chatting about 3D Human Pose.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios ChatPose: Chatting about 3D Human Pose

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.737573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.737573Z digest=sha256:9a0853f580a11485cb76f90c99d05299b8c6583823d2757a7229a85a55857b31

Observation 7a5ddd79-b5df-41a0-b503-9bcacbf19bd9 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Improved Baselines with Visual Instruction Tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.757388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.757388Z digest=sha256:f309ea0183eb796679200edade15bc3a6f919906a83a5d30efde49e687fefd06

Observation 8bdecde8-5bd6-48cd-99c3-30c96ce04cf7 · outbound

This paper cites When Do We Not Need Larger Vision Models?.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios When Do We Not Need Larger Vision Models?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.763273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.763273Z digest=sha256:adbf059649ed32d046fe6774d79c8e16d9690134d17a80961472759d13d5aba6

Observation e85f55e5-eb72-4326-aacc-9ecd7b6bf8d8 · outbound

This paper cites Vector-quantized image modeling with im- proved VQGAN,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Vector-quantized image modeling with im- proved VQGAN,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:53.674017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.774168Z digest=sha256:0616c7250a687642dcbc7374eceed2808cd4f0216dc479eb86ddba89239c3f72

Observation 4d324536-7a3a-4ad0-9850-3e1a931d37e0 · outbound

This paper cites Autoregressive image generation using residual quantization,.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Autoregressive image generation using residual quantization,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:53.655735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.782429Z digest=sha256:49c2ca221658868d1159bd2054d03345f5c2b1056476507bdd47239e9ba280a1

Observation 0a5a4665-584e-4175-b1e2-4760f916c20e · outbound

This paper cites An Image is Worth 32 Tokens for Reconstruction and Generation.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.789320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.789320Z digest=sha256:bea18e6a86ec21dd01755b155b7d76116df950c0d210e05ae5e50cff827c190c

Observation e5b390fb-a50a-40d1-97a0-3d354b075975 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.798589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.798589Z digest=sha256:16e1d4c4fb4053dfcd1b78b824c35cc2cab7c4c9d3cf8197b0d1a1de28479e57

Observation ca33bf07-7a70-463b-9646-de2f020ef802 · outbound

This paper cites 4566–4575.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios 4566–4575

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:53.723654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.724017Z digest=sha256:25a1a5d9dff04d794c41812047e2bfd9d8983d7be1e09ebb8b255d6d216e0bde

Observation 00eedd1a-34c8-4a77-b3a0-e217b66bd905 · outbound

This paper cites 5385–5394.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios 5385–5394

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.322436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.526379Z digest=sha256:e14781544272e85245325acc3eaa6cf54cb81cd9a4dce653ccbb5c902d0c587a

Observation f583c1ed-d594-4dc4-94e0-6268b88b7a2e · outbound

This paper cites 3502–3511.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios 3502–3511

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:05:54.701014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.384076Z digest=sha256:e381f03ce82a8b11101b6a82c745c6ad43ca2b3cb7086f84cf5811d313b4589b

Observation 4b15e097-7edc-4c7f-87d1-7ecd4dbbac12 · outbound

This paper cites ChatPose: Chatting about 3D Human Pose.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios ChatPose: Chatting about 3D Human Pose

Reference 2023

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T12:05:52.877419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T12:05:52.746940Z digest=sha256:057bc8ece83b90e5b6ff946711aaa150850fea7b5955aecb6fe30c998e8fcc3e

Pith citing papers

Observation bcf3f852-e480-42f0-b09a-17a317120d32 · inbound

Human-Centric Foundation Models: Perception, Generation and Agentic Modeling cites this paper.

Human-Centric Foundation Models: Perception, Generation and Agentic Modeling RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-08T04:41:08.652133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T04:41:08.431780Z digest=sha256:b3755f04f0b918acd9b9f16c4dfa1cda153b655dcd7f69a3aa6488e2bb603ea7