Pith. sign in

Paper Citation Record · LEDGER

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images

As of 21 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 0 inbound Pith citation observations for arXiv:2411.10334.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10334 v1

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:48:32.707991Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

76 of 76 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf177382-30de-4a6a-b86d-b36b39267cb5 · outbound

This paper cites The History Began from AlexNet: A Comprehensive Survey on Deep Learning Approaches.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images The History Began from AlexNet: A Comprehensive Survey on Deep Learning Approaches

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.546778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.546778Z digest=sha256:969d0a98a422e539f83d631c1e8f5f34f14bf10040e95119d0dd5f81c339229e

Observation 0da74ba3-62e7-46ba-9df8-ec345c8790da · outbound

This paper cites an unresolved cited work.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:48:34.766853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:31.563958Z digest=sha256:109310c28605b5012de0c27f828e224529cb0ec1a645d2ac4fa96be932d85c5f

Observation d5ecf148-3001-4698-b3d5-f835d95c5085 · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.567333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.567333Z digest=sha256:0f50d23ae36dbf1f96732760a853c48194c62b5a14bf79b96f9d303298646530

Observation a22820ec-8854-4e2f-b6d1-ff6ecbaabe75 · outbound

This paper cites Multimodal fu- sion via teacher-student network for indoor action recogni- tion.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Multimodal fu- sion via teacher-student network for indoor action recogni- tion

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.676976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:31.571113Z digest=sha256:9877d4372280dbdee8f39c3215924cef6656e9132293d435ab2c6bdebca5afe7

Observation a5a36342-85d7-4dff-8a7e-ed3991173067 · outbound

This paper cites Realtime multi-person 2d pose estimation using part affinity fields.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Realtime multi-person 2d pose estimation using part affinity fields

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.641110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:31.574730Z digest=sha256:9dcea0175b85b876aaed4bea19d06a6d70f5b714a677de0b7953e56e5a2466de

Observation 7b84e4ff-1d62-4062-ad08-39fdd4a17794 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Emerg- ing properties in self-supervised vision transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.578084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.578084Z digest=sha256:d305d1743fd99e8bca1cad191561ed734bc3d4484bc5c64e27e12f7a12b4f8c9

Observation 80cec3df-a0ad-418c-8c13-2ce3ad53992e · outbound

This paper cites Higherhrnet: Scale- aware representation learning for bottom-up human pose es- timation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Higherhrnet: Scale- aware representation learning for bottom-up human pose es- timation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.653336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.653336Z digest=sha256:eb60ea7a086b9750891ef0bfa972e203d108706c59cd9b25e989212f36d0c202

Observation f1a1f4ff-922d-4415-9b13-3faab0d30d20 · outbound

This paper cites Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.717798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.717798Z digest=sha256:5fe765f7cc2dd60dcdddd8b77fdb517b5535e31abf561bc07bc98947c31ab186

Observation a9c51e9b-3c38-4b90-879e-d988bce00bdd · outbound

This paper cites Keras documentation - model checkpoint.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Keras documentation - model checkpoint

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.622847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:31.778044Z digest=sha256:38066db5b8d383aec21c5cbb8ea121d27ef4997373230e6158b53554a80fb095

Observation 74b7c40f-cff7-47c0-a08a-efac74d5d06c · outbound

This paper cites Keras documentation - early stopping.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Keras documentation - early stopping

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.613391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:31.781155Z digest=sha256:d2955710b3f45af78e8156f0bdc30ff48be7c891a286756b52178a30a6911589

Observation dc5e7b47-77b0-4f3f-9d43-dea9600589e9 · outbound

This paper cites Unihcp: A unified model for human-centric perceptions.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Unihcp: A unified model for human-centric perceptions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.604168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:31.786558Z digest=sha256:9928925a1ed62a44a1e8a5deefa02cf664c27fcae2622adf24f22153bf8513cd

Observation a48e97ad-a07a-401c-ae1d-8c995dccc199 · outbound

This paper cites Adam: A method for stochastic opti- mization.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Adam: A method for stochastic opti- mization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.791114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.791114Z digest=sha256:cb262ba34578de5ab64bcfcf165771523b27b56e36ca8000dcc57dcc6651e410

Observation 3c3e908e-6545-46d1-9c17-0fabbd72a5c1 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.795196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.795196Z digest=sha256:d113dac10ab665ea824520876e1534bce140f7e0e5708d763f3e95aeee8f43ba

Observation 73770d07-986b-4b96-a4c5-1cc61453b9f4 · outbound

This paper cites an unresolved cited work.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:48:34.557720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:31.800000Z digest=sha256:bbde07fb62b89fb57c76453318c636768d4ed1328247373a1e4863a076400278

Observation fb4c3631-e627-43b2-b332-dbd6c8cb03a2 · outbound

This paper cites an unresolved cited work.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:48:34.509548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:31.803133Z digest=sha256:13b0db086664d15727cda2b2e5f8e37a999c4c8b710d77645f4c0329e1137cc4

Observation ccd16e5d-00f6-4ca1-b584-6d7b6967377c · outbound

This paper cites A stop list for general text.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images A stop list for general text

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.367547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:31.806265Z digest=sha256:8caef551709342422c5c707ebd66347d399d95028628155d94be14b53c80e476

Observation aecde1c8-a446-4357-9786-b221ac1ae827 · outbound

This paper cites Visiongpt2.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Visiongpt2

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.357075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:31.849735Z digest=sha256:47fe8bb737586b235ff91d3e8a8c500cbea98fa499efef5f6c7319651a042007

Observation 4f07817a-9c7c-4b12-b8dd-dc52422b2760 · outbound

This paper cites Understanding the diffi- culty of training deep feedforward neural networks.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Understanding the diffi- culty of training deep feedforward neural networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.899241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.899241Z digest=sha256:b860d990bfa52a9d4160ac63023d16e6a37b1e47443a0bd580dcc93c2e9c49fc

Observation 12e2bc46-d8fc-4feb-8cd3-37f419ce8229 · outbound

This paper cites Densepose: Dense human pose estimation in the wild.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Densepose: Dense human pose estimation in the wild

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.267321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:31.902904Z digest=sha256:5d9888bc245474d329400774bde75ba496070671983f40270c009080ccc5f6ea

Observation 638306de-22f7-430d-956d-c956bf1ab407 · outbound

This paper cites Deep residual learning for image recognition.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Deep residual learning for image recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.906285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.906285Z digest=sha256:021904633e480be67011f8046edeab877d947bb14df6f5b850f2960bfaaafb4d

Observation 3111ff37-4590-4d2b-a203-69a9a6282961 · outbound

This paper cites Mask r-cnn.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Mask r-cnn

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.909520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.909520Z digest=sha256:a3e9a450067e43dfbc5e69de5672141057dcbf331fc3c5d520973b7f57809b0b

Observation 3a88dbba-1b20-4c14-96a2-a4f869f95bdc · outbound

This paper cites Long short-term memory.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Long short-term memory

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.163039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:31.912656Z digest=sha256:82b86d572614f61b17cde9666e7469f287ab52faaf72f087537cd10d40687543

Observation d17947a0-38b3-464a-9602-b35d860122c3 · outbound

This paper cites A multi-instance multi-label dual learning approach for video captioning.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images A multi-instance multi-label dual learning approach for video captioning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.102615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:31.915800Z digest=sha256:ca3502e5e3cec72d58c9e9d93774dc2145a72602c8727b48a6d3d576b6d1fdd9

Observation 6f1234c7-bf84-4598-9456-bbe1214a6cdc · outbound

This paper cites Densecap: Fully convolutional localization networks for dense caption- ing.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Densecap: Fully convolutional localization networks for dense caption- ing

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.017340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:31.918945Z digest=sha256:205c8ccff03530caed77590c14750d8fa527725c81a90ee488994419206ede25

Observation d47208de-a2dd-4b7b-8d62-14b5c389e0ac · outbound

This paper cites Panoptic studio: A massively multiview system for social motion capture.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Panoptic studio: A massively multiview system for social motion capture

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.922299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.922299Z digest=sha256:dad9ddba16b7dc5f3d0aeeed1aa9f2b92d8243fc8120c15649be7356a8fd1fe4

Observation 40ae1662-109f-44e9-8dfd-0657fd420ee1 · outbound

This paper cites Bag of Tricks for Efficient Text Classification.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Bag of Tricks for Efficient Text Classification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.925830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.925830Z digest=sha256:58402999074ab869f7d4a1dc2bec00b00e3c4c9705f6683dc44adead567bf60f

Observation 76af935e-dc95-477c-bfae-97f1c819d2e4 · outbound

This paper cites Repurpos- ing diffusion-based image generators for monocular depth estimation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Repurpos- ing diffusion-based image generators for monocular depth estimation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.929993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.929993Z digest=sha256:66fc6b54b198faf1018ace99a4822aebae10f14e668622046b804cb5ddca1ae0

Observation 249e9ee9-a0b9-4c23-aa82-3acd43c39d26 · outbound

This paper cites Sapiens: Foundation for Human Vision Models.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Sapiens: Foundation for Human Vision Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.934189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.934189Z digest=sha256:7988a3c855c4b1094441940d9e09a47e0990666cdc6a425b206b13558f59ab40

Observation 546e132e-750a-44e6-8050-4b81d1e8c71f · outbound

This paper cites Sapiens discrete models repository.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Sapiens discrete models repository

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.995550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:31.938095Z digest=sha256:b33ee09a5a742fc5d7385cf59c55ba1751bebb0336ca4711924d8cba09dc313c

Observation 48382ebc-26ae-4339-8b4f-ec86d9f60d72 · outbound

This paper cites Sapiens: Foundation for human vision mod- els.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Sapiens: Foundation for human vision mod- els

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.950686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:31.941568Z digest=sha256:509e04ffad019b815edf3a5c8ac0ae4fe92e9ce043af0aca142e8f436bcaf92c

Observation 817d3746-b839-4b86-a153-18c47e87bd58 · outbound

This paper cites Segment any- thing.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Segment any- thing

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.002094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.002094Z digest=sha256:427945825c8db820bcf29d40790e5965246c8f881021bc552236724b3df3d76d

Observation b2f05a6b-b761-45fe-bae2-a4fd173184fa · outbound

This paper cites Visual genome: 9 Connecting language and vision using crowdsourced dense image annotations.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Visual genome: 9 Connecting language and vision using crowdsourced dense image annotations

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.844342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.083097Z digest=sha256:096ae37d08b2b35f80f528350cdd04c8881e2cc0e719d710aaba217fc1d99b80

Observation 084f3fdd-208b-40d6-bef4-a1f57c08dd64 · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Imagenet classification with deep convolutional neural net- works

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.116834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.116834Z digest=sha256:05660410733582bf9d2a2506f11fd71803eac1d367e22b3e094f0002c97c47cf

Observation 22ba0c89-ac3c-4045-92e5-1622eda7ad7d · outbound

This paper cites Motion-x: A large- scale 3d expressive whole-body human motion dataset.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Motion-x: A large- scale 3d expressive whole-body human motion dataset

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.828702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.120634Z digest=sha256:b21479746bc8bc60e6b26810ea0d44c1ee340be34d2cac5a0a2ca53a5b4a6f70

Observation 526af1bf-e9a0-4af5-9fb1-faa39caf0cbf · outbound

This paper cites Microsoft coco: Common objects in context.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Microsoft coco: Common objects in context

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.123328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.123328Z digest=sha256:cb5b4ea857b42e52a989f83ac591eb79af3b7ad05fd7abefcae94ed757cf4424

Observation a083d145-9c55-4830-b50a-1792b621ce47 · outbound

This paper cites Learning features combination for human action recognition from skeleton sequences.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Learning features combination for human action recognition from skeleton sequences

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.750260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.126488Z digest=sha256:5f13f169e179e5fd0a779a266e3d34265f467ec5f026fc1c12b2eca568fdf91d

Observation 9f99ddc7-2d21-4889-9c4d-6192ed8f157b · outbound

This paper cites Exploring the limits of weakly supervised pretraining.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Exploring the limits of weakly supervised pretraining

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.621617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.129923Z digest=sha256:5ccf6f5dbee57f5154388cc98f4e1d5692b7da3ef3eadfd705c7eeca7506ae5a

Observation 0e543e94-e524-41f3-99a0-726ee0ba036a · outbound

This paper cites Amass: Archive of motion capture as surface shapes.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Amass: Archive of motion capture as surface shapes

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.581219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.132979Z digest=sha256:d3f76173a9254f456b0a541cd8d6b05b6733a4c396c96d11b781caa68c459a80

Observation 3589cf7e-e5e0-435d-b69e-4e18f3fdc356 · outbound

This paper cites Y-net: joint segmen- tation and classification for diagnosis of breast biopsy im- ages.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Y-net: joint segmen- tation and classification for diagnosis of breast biopsy im- ages

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.572281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.135412Z digest=sha256:c556d9610b5cbcf767ed598c1aac43b5998427208e4fcadee7c0e0c5bcab197c

Observation 63c33264-156e-456e-9dcf-ca9c2577765f · outbound

This paper cites Efficient Estimation of Word Representations in Vector Space.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Efficient Estimation of Word Representations in Vector Space

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.138469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.138469Z digest=sha256:1efbca317b4515a30ebec5f204bdd36fefe95de487dd0b5574118cb2bfa04f84

Observation 809cf054-9d5e-4d2b-a071-92e6e90facbb · outbound

This paper cites Distributed representations of words and phrases and their compositionality.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Distributed representations of words and phrases and their compositionality

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.141714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.141714Z digest=sha256:a1385eef0bfa45d6d2ad04f5fdf7eaebd76a2bc26c3653c92ff2fc01b9818761

Observation a6125752-1b9b-4dcc-9a44-3001cd5cf3bf · outbound

This paper cites Stacked hour- glass networks for human pose estimation, 2016.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Stacked hour- glass networks for human pose estimation, 2016

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.557585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.144662Z digest=sha256:d28a263eadcb7a80fa5505584633ba277410a3c1e5d111f83c573c9d20866c5a

Observation 61c1ff5e-8eea-419b-a526-bf2877e9dd4f · outbound

This paper cites Derpa- nis, and Kostas Daniilidis.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Derpa- nis, and Kostas Daniilidis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.549075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.147584Z digest=sha256:813bf4c36a4bf1bbe1ef39b0663d6b65d2f457aa02e68bb82575e8d17fc0e3a1

Observation 6ef337c8-2bc3-48b6-b97d-836f9ae66121 · outbound

This paper cites an unresolved cited work.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:48:33.540330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.150508Z digest=sha256:8f86ff244fa04bf8c89c5aea3ffea2ea0f6af61fb3f4d195d98559be9bbcb8b5

Observation be9da4e2-8603-4bf8-a7f3-db0af0dfea04 · outbound

This paper cites Unidepth: Universal monocular metric depth estimation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Unidepth: Universal monocular metric depth estimation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.154476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.154476Z digest=sha256:17a49003deb3d6dfb5626be3048bf1058b492456cd710de6e63cf28391c9c475

Observation 2086889f-b333-4a04-9184-1f0cfdfa70de · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Learning transferable visual models from natural language supervi- sion

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.157855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.157855Z digest=sha256:ca1e4231562db4946ea6ae1cec8dcef9f2c6649b185ee48c65122e9f070cf10f

Observation 0c00a0a4-438d-4c2d-9180-935209b8ad35 · outbound

This paper cites Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.520879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.263913Z digest=sha256:48cfb36814a81bd1829daf79d29e183ea57c4b83a1671c7e4e0403278f676124

Observation 78ec2cd0-4c57-4c13-a2e5-bcedd7bdea3e · outbound

This paper cites You only look once: Unified, real-time object de- tection.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images You only look once: Unified, real-time object de- tection

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.345416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.345416Z digest=sha256:10cd96d627ae504066ab2d56060c875951cb83ff8a262ec63d81a7e9144c294e

Observation 3980dc14-b5a7-4a81-924a-02abb9a15e68 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.349694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.349694Z digest=sha256:4e1637b41f9f6507ca9da6f2508980afe02a7f0a875f0522930b9f75238e0593

Observation 1aadc054-d7ff-48af-b159-4cdd678cb4db · outbound

This paper cites Lcr-net: Localization-classification-regression for human pose.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Lcr-net: Localization-classification-regression for human pose

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.441396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.353152Z digest=sha256:5da228a7ecfbe30481b47d3dd218bcecc76f920c777bb7c8170cc9dedcc660e3

Observation 78cece2a-b6ed-4ffe-a3d7-cfddf65b5900 · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images U- net: Convolutional networks for biomedical image segmen- tation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.356258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.356258Z digest=sha256:8ef9a8646aef9df0d146e3d6b9d1376666a520107972e23237ede10e628982b4

Observation 0d9cfcae-7323-4eb8-a2ec-739f650676d3 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.359234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.359234Z digest=sha256:d279021f0a8a130365af552b273ae31fbbe22ddb141d548492edeed4ce779979

Observation fb2359f6-b880-46b6-8b69-6fc8656090fe · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.362357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.362357Z digest=sha256:48b15a0876e8d2663e2bdebafdf453283bca8dc7d257a84d133cfde1b7e64d11

Observation aa671b21-aa1b-437b-9dde-05c9a0c9a442 · outbound

This paper cites Revisiting unreasonable effectiveness of data in deep learning era.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Revisiting unreasonable effectiveness of data in deep learning era

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.365988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.365988Z digest=sha256:6d3494d7f883889d68aaa018f8401b93a92b09643db0e58a40be433fbf394654

Observation 48d3e48a-0dd9-46b8-8b38-87ed0ecbc87d · outbound

This paper cites Mask-yolo.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Mask-yolo

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.346046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.369181Z digest=sha256:a0bb5b4caed7a1544ce1b545bbeec5e7abc9890999a91612d5ad0fe5467a860d

Observation bfe92ecd-e62f-48ec-bd44-92f29778b4d2 · outbound

This paper cites Deep high-resolution representation learning for human pose es- timation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Deep high-resolution representation learning for human pose es- timation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.372784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.372784Z digest=sha256:487decf22f2d67c9ebe0f7f9233f30458c5cce3a232cc254939c17c4aa0c4c62

Observation d1cd060e-49eb-4ff1-bee3-d65e31d2f793 · outbound

This paper cites Going deeper with convolutions.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Going deeper with convolutions

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.426648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.426648Z digest=sha256:e73c58537e1bbda994d703f9ccb8c503ca69487b03083584541d803a748c7359

Observation 59ea1e12-48d8-47aa-acf9-f7a9673e7583 · outbound

This paper cites Joint training of a convolutional network and a graphical model for human pose estimation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Joint training of a convolutional network and a graphical model for human pose estimation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.327248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.490731Z digest=sha256:f9368d7cb53948119c6ca19c7512c446f7352175bb88cb9178e9a67da7ba27b9

Observation 341623ba-bfa1-47d8-9676-cff011fdb21b · outbound

This paper cites Deeppose: Human pose estimation via deep neural networks.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Deeppose: Human pose estimation via deep neural networks

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.295813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.494573Z digest=sha256:4283877f7fc57b693447ffea1331fe3a243220d1b323d4bf71ec3e1f494e48c4

Observation 2026cb34-3bec-4e5f-bf61-a4792a93ede3 · outbound

This paper cites Attention is all you need.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Attention is all you need

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.498323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.498323Z digest=sha256:67958188420a7df7582554a7723027626b9a04a7e4176c7712353a21073702a6

Observation 3c14e94f-9f80-4137-b377-7a40cc9bf495 · outbound

This paper cites Deep high-resolution repre- sentation learning for visual recognition.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Deep high-resolution repre- sentation learning for visual recognition

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.501928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.501928Z digest=sha256:2e82d43b140cfb290faf5d2cbcc0186c5b14341bc790c236c10576f0b6f05b1f

Observation 7fa5486b-09f6-49d6-9255-b8ef1b6333d2 · outbound

This paper cites Y-net: a one-to-two deep learning framework for digital holographic reconstruction.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Y-net: a one-to-two deep learning framework for digital holographic reconstruction

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.170351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.504928Z digest=sha256:02e7cb6436f1c71b878ca25981d947acd1129a269f5e510b3380cdd208742494

Observation 2293a741-20c4-470b-a5ae-451939fc08fb · outbound

This paper cites Dust3r: Geometric 3d vi- sion made easy.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Dust3r: Geometric 3d vi- sion made easy

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.508802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.508802Z digest=sha256:c03fabf3a32cccd30d2686db9c31dea1ceb3d3aa0c0e8ba473a2ef10033ee63f

Observation 1d1cda28-5ac5-4b04-8792-b9b17cead7c6 · outbound

This paper cites Detectron2.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Detectron2

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.155649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.511448Z digest=sha256:62026c73f89b521644b5fbb13f6d77f3847296a61c84bb1875f3711ae4d9dc06

Observation 8950cc69-942a-4279-803d-5428e95bd62f · outbound

This paper cites Empirical Evaluation of Rectified Activations in Convolutional Network.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Empirical Evaluation of Rectified Activations in Convolutional Network

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.515471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.515471Z digest=sha256:36aa9aaa4e7c1913701c026b557fb3327aca8ddfda7cc1896095e887baf984b4

Observation 16e4fbf8-a0e3-4195-a75b-374508f50223 · outbound

This paper cites Unifying flow, stereo and depth estimation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Unifying flow, stereo and depth estimation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.049507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.519175Z digest=sha256:a0303491ae402a925dce299b30a96c970e5b3050e2a39aeb7b8fe3b82449fe42

Observation 753987b0-d2fa-4ab8-aaa3-28fd3c6f4e0e · outbound

This paper cites Zoomnas: search- ing for whole-body human pose estimation in the wild.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(4):5296–5313, 2022.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Zoomnas: search- ing for whole-body human pose estimation in the wild.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(4):5296–5313, 2022

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.039696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.522539Z digest=sha256:f98f3d137bdebd00edeb85a7a8f9a898e5a311abfcf32e618ce18272ca849649

Observation a241ed22-71e7-4d35-b08b-1f6dde22ba20 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Depth anything: Unleashing the power of large-scale unlabeled data

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.029033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.525915Z digest=sha256:d8c376159c177b3a309857c4ef5144c485b2eee9d2c080df799be24efd66b9fb

Observation c1354a33-3ee3-4967-b85e-0b0ca0fc6370 · outbound

This paper cites Depth Anything V2.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Depth Anything V2

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.530018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.530018Z digest=sha256:8eaac1c8236e334b31e08facef4bd01c42ae27de6c2f503246d4d9134dfbfd12

Observation c34cc9b5-2ad8-4723-9b2c-823938e1acc8 · outbound

This paper cites End-to-end learning of deformable mixture of parts and deep convolutional neural networks for human pose esti- mation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images End-to-end learning of deformable mixture of parts and deep convolutional neural networks for human pose esti- mation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.016587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.533597Z digest=sha256:a24a0c9eb53ed008b0d8a88c132b818c84ff9581a803b9d898db118f9cc714bf

Observation 34546536-b324-400e-a137-5f4da12cf965 · outbound

This paper cites Dptext-detr: Towards better scene text detection with dynamic points in transformer.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Dptext-detr: Towards better scene text detection with dynamic points in transformer

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.003423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.577494Z digest=sha256:8c83d11dc1391f08e8d16c63cbdbf96f05a67c5750a4a9189afb15bc12f414fc

Observation fc42eb17-c79f-4c14-8074-0ec034090818 · outbound

This paper cites Differential transformer, 2024.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Differential transformer, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:32.990798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.659966Z digest=sha256:eda2c23578b6b1ef9fdd836445ca564a99de728dad26be3008ec832f21e0980c

Observation daf28c49-891e-4fc3-955e-76127723f4e4 · outbound

This paper cites Metric3d: Towards zero-shot metric 3d prediction from a single image.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Metric3d: Towards zero-shot metric 3d prediction from a single image

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.699287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.699287Z digest=sha256:41d43beec41587daa0ba63b98a38f42c8ecabf749621bc677f125b939b53b588

Observation bdc3a0ef-dd0d-4440-bd1c-9e029e715f33 · outbound

This paper cites Learn- ing from multiple teacher networks.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Learn- ing from multiple teacher networks

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:32.971313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.702710Z digest=sha256:532dd206273607a60e4006d9c5013b52d1ab6d17b4f4c5ff872b79943b1c67c3

Observation ce5ed699-a2cf-45ac-b305-b824d190de42 · outbound

This paper cites Lite-hrnet: A lightweight high-resolution network.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Lite-hrnet: A lightweight high-resolution network

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:32.879333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T19:48:32.705719Z digest=sha256:a258b4338dc57afbddb622c8226d1afb09ee783c7625d968a61fe452fce32af8

Observation bd8d7780-775a-429b-ad41-918618667f47 · outbound

This paper cites Attention Heads of Large Language Models: A Survey.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Attention Heads of Large Language Models: A Survey

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.707991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.707991Z digest=sha256:76aeee190c8e758673bb60a035e8c24b35e2f452e78ebb852bc4971f13747c20

Pith citing papers

No inbound Pith citation observations are available.