Pith. sign in

Paper Citation Record · LEDGER

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images

As of 14 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 0 inbound Pith citation observations for arXiv:2411.10334.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10334 v1

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:48:32.707991Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

76 of 76 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf177382-30de-4a6a-b86d-b36b39267cb5 · outbound

This paper cites The History Began from AlexNet: A Comprehensive Survey on Deep Learning Approaches.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images The History Began from AlexNet: A Comprehensive Survey on Deep Learning Approaches

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.546778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.546778Z digest=sha256:240c7feb69ef88ec7d50b274e405c6db7dabe6e7ac9c98109b32d24684ce0b24

Observation 0da74ba3-62e7-46ba-9df8-ec345c8790da · outbound

This paper cites an unresolved cited work.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:48:34.766853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:31.563958Z digest=sha256:ff7b77fcb88aa21f90e31bec1e3a9a3ddbdb2985c0c9ce13a4ca7dd19427cd16

Observation d5ecf148-3001-4698-b3d5-f835d95c5085 · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.567333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.567333Z digest=sha256:768f7ea5fc84935f434cfe8e157b952f4b5981e3ddc961e78e3e2d3e5528a506

Observation a22820ec-8854-4e2f-b6d1-ff6ecbaabe75 · outbound

This paper cites Multimodal fu- sion via teacher-student network for indoor action recogni- tion.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Multimodal fu- sion via teacher-student network for indoor action recogni- tion

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.676976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:31.571113Z digest=sha256:a575a30464424730495acb6745dd6f5b774042b69113f0096b99b25f12e0a28a

Observation a5a36342-85d7-4dff-8a7e-ed3991173067 · outbound

This paper cites Realtime multi-person 2d pose estimation using part affinity fields.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Realtime multi-person 2d pose estimation using part affinity fields

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.641110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:31.574730Z digest=sha256:47d7c7456197ee07dfd44e9289180741195f0cab60830972751583124aa28b2f

Observation 7b84e4ff-1d62-4062-ad08-39fdd4a17794 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Emerg- ing properties in self-supervised vision transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.578084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.578084Z digest=sha256:21804ae78b76a3d1de60d9fef7c1fbf551a938436f1e3c81750d0d936f48ee3a

Observation 80cec3df-a0ad-418c-8c13-2ce3ad53992e · outbound

This paper cites Higherhrnet: Scale- aware representation learning for bottom-up human pose es- timation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Higherhrnet: Scale- aware representation learning for bottom-up human pose es- timation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.653336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.653336Z digest=sha256:b9bfd44362de95d01b7e2440efca0a0a33255adcb0336d2d620c00697a0579b9

Observation f1a1f4ff-922d-4415-9b13-3faab0d30d20 · outbound

This paper cites Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.717798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.717798Z digest=sha256:0c0593212f708d130c7ed919af0780fe5b121b568e0ff64745dedecdc378f278

Observation a9c51e9b-3c38-4b90-879e-d988bce00bdd · outbound

This paper cites Keras documentation - model checkpoint.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Keras documentation - model checkpoint

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.622847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:31.778044Z digest=sha256:521b4470e7d855eeb239fa17762f6faa48430a3aa25a719e3a110eba22c1e73e

Observation 74b7c40f-cff7-47c0-a08a-efac74d5d06c · outbound

This paper cites Keras documentation - early stopping.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Keras documentation - early stopping

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.613391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:31.781155Z digest=sha256:78530a3be7e367b7e8cbff95222d3540a42eff0adfe92da5a3ab1bbfa3200e52

Observation dc5e7b47-77b0-4f3f-9d43-dea9600589e9 · outbound

This paper cites Unihcp: A unified model for human-centric perceptions.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Unihcp: A unified model for human-centric perceptions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.604168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:31.786558Z digest=sha256:2fafcb10a899a54ad6223be78f09833c99edc7cd807c8b9f7ca0527ff51b0b2b

Observation a48e97ad-a07a-401c-ae1d-8c995dccc199 · outbound

This paper cites Adam: A method for stochastic opti- mization.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Adam: A method for stochastic opti- mization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.791114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.791114Z digest=sha256:0ac648020edca92371a8a1895f1852ae6b1476dc682804845d3bf8afcbe5489a

Observation 3c3e908e-6545-46d1-9c17-0fabbd72a5c1 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.795196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.795196Z digest=sha256:5d4ac63541664a6693d45e3af685e30940b2bc73f1b2408a5f519e442258d5bd

Observation 73770d07-986b-4b96-a4c5-1cc61453b9f4 · outbound

This paper cites an unresolved cited work.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:48:34.557720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:31.800000Z digest=sha256:e3452e9a0e828e3466ff926708c1c6e6649eeee63837cc12dd845114c7379e8e

Observation fb4c3631-e627-43b2-b332-dbd6c8cb03a2 · outbound

This paper cites an unresolved cited work.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:48:34.509548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:31.803133Z digest=sha256:0748f500ed9cb9e7df12266e040a103f6ecb7f52a9ca534b0b26095d411700b4

Observation ccd16e5d-00f6-4ca1-b584-6d7b6967377c · outbound

This paper cites A stop list for general text.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images A stop list for general text

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.367547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:31.806265Z digest=sha256:2e7b018d49bdcf3ecec20342a6ab7f0bcced575fcb3b7a777f41450acb3bfbd6

Observation aecde1c8-a446-4357-9786-b221ac1ae827 · outbound

This paper cites Visiongpt2.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Visiongpt2

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.357075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:31.849735Z digest=sha256:cb8ae83e0f031919e02d9992f803c59aad154552fd28d98e0f9a0c642a603c7a

Observation 4f07817a-9c7c-4b12-b8dd-dc52422b2760 · outbound

This paper cites Understanding the diffi- culty of training deep feedforward neural networks.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Understanding the diffi- culty of training deep feedforward neural networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.899241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.899241Z digest=sha256:9cb65ab95ce6d10fa739edf85c6b9af28c50b4ad58e9195d7a9cb92ef5ac86af

Observation 12e2bc46-d8fc-4feb-8cd3-37f419ce8229 · outbound

This paper cites Densepose: Dense human pose estimation in the wild.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Densepose: Dense human pose estimation in the wild

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.267321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:31.902904Z digest=sha256:c976feee0c9256b65970649ec63c6a9a01be7ac259e87c5e44756f15b08f99b7

Observation 638306de-22f7-430d-956d-c956bf1ab407 · outbound

This paper cites Deep residual learning for image recognition.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Deep residual learning for image recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.906285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.906285Z digest=sha256:6a8ba409ac17f8ede81d424560f0755794187a0a0ee9b47e9e6d88cef7c8a8f4

Observation 3111ff37-4590-4d2b-a203-69a9a6282961 · outbound

This paper cites Mask r-cnn.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Mask r-cnn

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.909520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.909520Z digest=sha256:a7506e569f9a57338824969f670415c447983fadaf8425e9ad18061fec9100ef

Observation 3a88dbba-1b20-4c14-96a2-a4f869f95bdc · outbound

This paper cites Long short-term memory.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Long short-term memory

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.163039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:31.912656Z digest=sha256:2ef349bbb9a372415170ecf86738d868ab34af8050ee71a8cb6e522dab881e1a

Observation d17947a0-38b3-464a-9602-b35d860122c3 · outbound

This paper cites A multi-instance multi-label dual learning approach for video captioning.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images A multi-instance multi-label dual learning approach for video captioning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.102615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:31.915800Z digest=sha256:7607737922c0db23b3273ae975a2bbe90a598919afc92070fc8e6c44309c7736

Observation 6f1234c7-bf84-4598-9456-bbe1214a6cdc · outbound

This paper cites Densecap: Fully convolutional localization networks for dense caption- ing.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Densecap: Fully convolutional localization networks for dense caption- ing

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:34.017340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:31.918945Z digest=sha256:d8a89df0cc3614f6e3d376f09b97084c6cbb208ff5e1695505e2b7b1e79de38c

Observation d47208de-a2dd-4b7b-8d62-14b5c389e0ac · outbound

This paper cites Panoptic studio: A massively multiview system for social motion capture.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Panoptic studio: A massively multiview system for social motion capture

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.922299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.922299Z digest=sha256:52804ae5ae05c14be25b528071b9f6e152149f07812901c778080d8e73cc6070

Observation 40ae1662-109f-44e9-8dfd-0657fd420ee1 · outbound

This paper cites Bag of Tricks for Efficient Text Classification.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Bag of Tricks for Efficient Text Classification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.925830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.925830Z digest=sha256:2d460a8b0672ace849e8b8bfc1ee6dc2296e41da28d91065658fd05209f09005

Observation 76af935e-dc95-477c-bfae-97f1c819d2e4 · outbound

This paper cites Repurpos- ing diffusion-based image generators for monocular depth estimation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Repurpos- ing diffusion-based image generators for monocular depth estimation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.929993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.929993Z digest=sha256:7090ca723042ea4386e670b693bf20ee5fac8528984d1d207e55ff1d0ac0d657

Observation 249e9ee9-a0b9-4c23-aa82-3acd43c39d26 · outbound

This paper cites Sapiens: Foundation for Human Vision Models.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Sapiens: Foundation for Human Vision Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:31.934189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:31.934189Z digest=sha256:2752e7a527e0f94b89109aedd26f48e445247691d130dcc12abfdf196bff49d4

Observation 546e132e-750a-44e6-8050-4b81d1e8c71f · outbound

This paper cites Sapiens discrete models repository.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Sapiens discrete models repository

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.995550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:31.938095Z digest=sha256:f0f470ba0dd649a2bb06b307ad6974b20cdbc098a74a962e4844b51fc5864ab1

Observation 48382ebc-26ae-4339-8b4f-ec86d9f60d72 · outbound

This paper cites Sapiens: Foundation for human vision mod- els.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Sapiens: Foundation for human vision mod- els

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.950686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:31.941568Z digest=sha256:2446639e503e98d8b0d3a5342d74513a3f1203f966845a94de8d6705acd14d08

Observation 817d3746-b839-4b86-a153-18c47e87bd58 · outbound

This paper cites Segment any- thing.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Segment any- thing

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.002094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.002094Z digest=sha256:b0018cffd7b94e18225ffe40ea05e4c30004568d0e5be1f18807a4b493dd6929

Observation b2f05a6b-b761-45fe-bae2-a4fd173184fa · outbound

This paper cites Visual genome: 9 Connecting language and vision using crowdsourced dense image annotations.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Visual genome: 9 Connecting language and vision using crowdsourced dense image annotations

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.844342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.083097Z digest=sha256:f2064543dcb9875a0cf8ee076c005f8d81a01f299e42bc56646b60c6278a385a

Observation 084f3fdd-208b-40d6-bef4-a1f57c08dd64 · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Imagenet classification with deep convolutional neural net- works

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.116834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.116834Z digest=sha256:4ec171ab2ad6f5d803376f0766d6db5b89d609b890ef3bb0907748ea7c0dd88a

Observation 22ba0c89-ac3c-4045-92e5-1622eda7ad7d · outbound

This paper cites Motion-x: A large- scale 3d expressive whole-body human motion dataset.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Motion-x: A large- scale 3d expressive whole-body human motion dataset

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.828702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.120634Z digest=sha256:805273399380f577962feb502d40bcd0f095fae26d334a221b65f9339aebabd5

Observation 526af1bf-e9a0-4af5-9fb1-faa39caf0cbf · outbound

This paper cites Microsoft coco: Common objects in context.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Microsoft coco: Common objects in context

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.123328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.123328Z digest=sha256:232b9c00290a21978869c828360e8c790a2c91e86d2cb3b9eea51bcd730af5d3

Observation a083d145-9c55-4830-b50a-1792b621ce47 · outbound

This paper cites Learning features combination for human action recognition from skeleton sequences.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Learning features combination for human action recognition from skeleton sequences

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.750260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.126488Z digest=sha256:01bfe17c3e5c9641426251a1a74464b685debd47dc082487cec19cb4dac126c4

Observation 9f99ddc7-2d21-4889-9c4d-6192ed8f157b · outbound

This paper cites Exploring the limits of weakly supervised pretraining.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Exploring the limits of weakly supervised pretraining

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.621617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.129923Z digest=sha256:f5ea291277afac535995b1a8ad1bdfa0641cc72522d4766eda3056ef442e3ba9

Observation 0e543e94-e524-41f3-99a0-726ee0ba036a · outbound

This paper cites Amass: Archive of motion capture as surface shapes.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Amass: Archive of motion capture as surface shapes

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.581219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.132979Z digest=sha256:04366e07545a16342df321550215637823ddca30ab25cce86350c9b5177fd4c2

Observation 3589cf7e-e5e0-435d-b69e-4e18f3fdc356 · outbound

This paper cites Y-net: joint segmen- tation and classification for diagnosis of breast biopsy im- ages.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Y-net: joint segmen- tation and classification for diagnosis of breast biopsy im- ages

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.572281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.135412Z digest=sha256:02d4380433d7ba1507ab21320fc60a862e8d4d9678404f67ccc27fca0a6f0520

Observation 63c33264-156e-456e-9dcf-ca9c2577765f · outbound

This paper cites Efficient Estimation of Word Representations in Vector Space.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Efficient Estimation of Word Representations in Vector Space

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.138469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.138469Z digest=sha256:9bd6a28a71a7cf3a229a2ca7d6c2530c3ecf164d8abc601c1f3e0234a75ba172

Observation 809cf054-9d5e-4d2b-a071-92e6e90facbb · outbound

This paper cites Distributed representations of words and phrases and their compositionality.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Distributed representations of words and phrases and their compositionality

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.141714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.141714Z digest=sha256:c6e89d6460061d71bb93f709d97610c3e61083a5ba32bf2062b9f6cc1f159514

Observation a6125752-1b9b-4dcc-9a44-3001cd5cf3bf · outbound

This paper cites Stacked hour- glass networks for human pose estimation, 2016.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Stacked hour- glass networks for human pose estimation, 2016

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.557585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.144662Z digest=sha256:e3e74228b687df277af04b39e6d562e38ce64d2da6d3f38712226b8348a95520

Observation 61c1ff5e-8eea-419b-a526-bf2877e9dd4f · outbound

This paper cites Derpa- nis, and Kostas Daniilidis.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Derpa- nis, and Kostas Daniilidis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.549075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.147584Z digest=sha256:50edac044bd859f0944e5a44d4ac08adef238b56cbc8ce43b03bbbf0fa2b76bf

Observation 6ef337c8-2bc3-48b6-b97d-836f9ae66121 · outbound

This paper cites an unresolved cited work.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-12T19:48:33.540330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.150508Z digest=sha256:fd1dcecb677a79bffd9cfce0ef3e8f5aaeadb477e3b9c7446105753e30c049de

Observation be9da4e2-8603-4bf8-a7f3-db0af0dfea04 · outbound

This paper cites Unidepth: Universal monocular metric depth estimation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Unidepth: Universal monocular metric depth estimation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.154476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.154476Z digest=sha256:021528f49a0b4a54712e0e9d2b3b3c9ba6418ad11a2bdcb1855110dcb1eddf02

Observation 2086889f-b333-4a04-9184-1f0cfdfa70de · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Learning transferable visual models from natural language supervi- sion

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.157855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.157855Z digest=sha256:6e9cf85a33bf694607bcb837ccb67d83b4b6db56962bf946c73a7f59eefc4395

Observation 0c00a0a4-438d-4c2d-9180-935209b8ad35 · outbound

This paper cites Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.520879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.263913Z digest=sha256:4f2d23ea6835de6ec348abeb88346b1159321c7d99d944a6a05fee7733c48645

Observation 78ec2cd0-4c57-4c13-a2e5-bcedd7bdea3e · outbound

This paper cites You only look once: Unified, real-time object de- tection.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images You only look once: Unified, real-time object de- tection

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.345416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.345416Z digest=sha256:71d86f4767c0d44d35180558dbd699efa084ad9bbfc190e1877a6aef9497ce0b

Observation 3980dc14-b5a7-4a81-924a-02abb9a15e68 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.349694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.349694Z digest=sha256:bb9bdd321dc9f75bc269dfae695c6f471b5662db171313eb86fa04a5c61995ef

Observation 1aadc054-d7ff-48af-b159-4cdd678cb4db · outbound

This paper cites Lcr-net: Localization-classification-regression for human pose.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Lcr-net: Localization-classification-regression for human pose

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.441396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.353152Z digest=sha256:d8997f316f0e65c311793918b4e0c2a5b789431e47305561d82098927e1808a3

Observation 78cece2a-b6ed-4ffe-a3d7-cfddf65b5900 · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images U- net: Convolutional networks for biomedical image segmen- tation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.356258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.356258Z digest=sha256:65f76251fb1aa0e993ce294f008a1f1f6084f8d236ed479262f600d136caa14e

Observation 0d9cfcae-7323-4eb8-a2ec-739f650676d3 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.359234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.359234Z digest=sha256:0dec0fd6858f7a2436a6806f06772ca3e0f90d587eb2cc82df856c4cea717423

Observation fb2359f6-b880-46b6-8b69-6fc8656090fe · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.362357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.362357Z digest=sha256:b3a4d7ae7f1e6bb9eeaa21bf3fc0536794c14b75d82008d759f6b68eef12d0ed

Observation aa671b21-aa1b-437b-9dde-05c9a0c9a442 · outbound

This paper cites Revisiting unreasonable effectiveness of data in deep learning era.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Revisiting unreasonable effectiveness of data in deep learning era

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.365988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.365988Z digest=sha256:48ae65f08f251e08b8ab92b8978a42ca9c7dbac388765a670a34834472ffee09

Observation 48d3e48a-0dd9-46b8-8b38-87ed0ecbc87d · outbound

This paper cites Mask-yolo.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Mask-yolo

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.346046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.369181Z digest=sha256:74539b8dfde4a2584e118c5dcfd9d15799847725f27a70346bfc678123265009

Observation bfe92ecd-e62f-48ec-bd44-92f29778b4d2 · outbound

This paper cites Deep high-resolution representation learning for human pose es- timation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Deep high-resolution representation learning for human pose es- timation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.372784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.372784Z digest=sha256:26d83c7e8148d4ce069a226ad920994b5016e79a4059392bc4676df363864d95

Observation d1cd060e-49eb-4ff1-bee3-d65e31d2f793 · outbound

This paper cites Going deeper with convolutions.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Going deeper with convolutions

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.426648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.426648Z digest=sha256:b2744ec3dabeeb6f8e7b26a9df5626b92bef6a70046a530e65824f3ddee7bc28

Observation 59ea1e12-48d8-47aa-acf9-f7a9673e7583 · outbound

This paper cites Joint training of a convolutional network and a graphical model for human pose estimation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Joint training of a convolutional network and a graphical model for human pose estimation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.327248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.490731Z digest=sha256:44e3bcf0167b026df50ff8856223417ef954a30201ce2d1a4d6040f0e3bb1cd8

Observation 341623ba-bfa1-47d8-9676-cff011fdb21b · outbound

This paper cites Deeppose: Human pose estimation via deep neural networks.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Deeppose: Human pose estimation via deep neural networks

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.295813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.494573Z digest=sha256:2c36ad38c4cd258366e79923efe87324eb101e5c8a060e3a4c79d4e90d9c1e33

Observation 2026cb34-3bec-4e5f-bf61-a4792a93ede3 · outbound

This paper cites Attention is all you need.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Attention is all you need

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.498323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.498323Z digest=sha256:b54354b087b957b1119a76c1ee449fcb58055c25d1506012dcd2370261ce641a

Observation 3c14e94f-9f80-4137-b377-7a40cc9bf495 · outbound

This paper cites Deep high-resolution repre- sentation learning for visual recognition.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Deep high-resolution repre- sentation learning for visual recognition

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.501928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.501928Z digest=sha256:0fcf172d5d22e0a606a1f3382db254b606398840883b4d96e8a1347f9398011b

Observation 7fa5486b-09f6-49d6-9255-b8ef1b6333d2 · outbound

This paper cites Y-net: a one-to-two deep learning framework for digital holographic reconstruction.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Y-net: a one-to-two deep learning framework for digital holographic reconstruction

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.170351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.504928Z digest=sha256:d4b6887b8e69cd0d55894cbbfa656fcc6165774ac8abd7eaeffd81ebb4877395

Observation 2293a741-20c4-470b-a5ae-451939fc08fb · outbound

This paper cites Dust3r: Geometric 3d vi- sion made easy.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Dust3r: Geometric 3d vi- sion made easy

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.508802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.508802Z digest=sha256:547817176cb19a55c882076a9ad0f6e0185262e748fbabf0bb8f10df9a323779

Observation 1d1cda28-5ac5-4b04-8792-b9b17cead7c6 · outbound

This paper cites Detectron2.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Detectron2

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.155649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.511448Z digest=sha256:7867e63a475bfd43f0b2cf84c6020eeea0d1a55e9f640aa3f2eb6dcef4615b77

Observation 8950cc69-942a-4279-803d-5428e95bd62f · outbound

This paper cites Empirical Evaluation of Rectified Activations in Convolutional Network.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Empirical Evaluation of Rectified Activations in Convolutional Network

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.515471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.515471Z digest=sha256:79fa461a90e96b5775e3b210eae40c09dcf70ea3a11900e7e9dd7f2be900d461

Observation 16e4fbf8-a0e3-4195-a75b-374508f50223 · outbound

This paper cites Unifying flow, stereo and depth estimation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Unifying flow, stereo and depth estimation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.049507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.519175Z digest=sha256:4cc00f5317c59cc60904b3ccf7384d0f0603622aaec000db70e4f1c30adb0280

Observation 753987b0-d2fa-4ab8-aaa3-28fd3c6f4e0e · outbound

This paper cites Zoomnas: search- ing for whole-body human pose estimation in the wild.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(4):5296–5313, 2022.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Zoomnas: search- ing for whole-body human pose estimation in the wild.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(4):5296–5313, 2022

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.039696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.522539Z digest=sha256:4f50ff348832ac73df3f695b4dddaf81e908e628d82daa4dd15389774ebe57c8

Observation a241ed22-71e7-4d35-b08b-1f6dde22ba20 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Depth anything: Unleashing the power of large-scale unlabeled data

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.029033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.525915Z digest=sha256:4af10f6a89cd4f5e57fa64bf0a77f91bfec563bee6f7b84b1a206597b67e6192

Observation c1354a33-3ee3-4967-b85e-0b0ca0fc6370 · outbound

This paper cites Depth Anything V2.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Depth Anything V2

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.530018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.530018Z digest=sha256:824ca42a0b0fee1fd0b0caff4e359d4b0bb70de1f8de84e4dab3940735cfd405

Observation c34cc9b5-2ad8-4723-9b2c-823938e1acc8 · outbound

This paper cites End-to-end learning of deformable mixture of parts and deep convolutional neural networks for human pose esti- mation.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images End-to-end learning of deformable mixture of parts and deep convolutional neural networks for human pose esti- mation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.016587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.533597Z digest=sha256:3f0f668a9cc0e37db4535d621bbc697326879d62ff485fb22b611190a0b7859e

Observation 34546536-b324-400e-a137-5f4da12cf965 · outbound

This paper cites Dptext-detr: Towards better scene text detection with dynamic points in transformer.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Dptext-detr: Towards better scene text detection with dynamic points in transformer

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:33.003423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.577494Z digest=sha256:9c2163b44e7f90b5748a90f3bf0ca49dba00a2d9d31737f77d60223f7bb75cc5

Observation fc42eb17-c79f-4c14-8074-0ec034090818 · outbound

This paper cites Differential transformer, 2024.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Differential transformer, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:32.990798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.659966Z digest=sha256:d800434597eeacf9d987025ccbfe8cba7d818675a4ac650a9c446ae9d40a6b93

Observation daf28c49-891e-4fc3-955e-76127723f4e4 · outbound

This paper cites Metric3d: Towards zero-shot metric 3d prediction from a single image.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Metric3d: Towards zero-shot metric 3d prediction from a single image

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.699287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.699287Z digest=sha256:ec88a83628ebcd0c3f4197c3dac73dc5df7b6b29a0ed9a8a0d7889d52b60e60c

Observation bdc3a0ef-dd0d-4440-bd1c-9e029e715f33 · outbound

This paper cites Learn- ing from multiple teacher networks.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Learn- ing from multiple teacher networks

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:32.971313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.702710Z digest=sha256:35f363d24f5248364be56c6bef521d223d1a67664b9ce14dbb52d15079381448

Observation ce5ed699-a2cf-45ac-b305-b824d190de42 · outbound

This paper cites Lite-hrnet: A lightweight high-resolution network.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Lite-hrnet: A lightweight high-resolution network

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:48:32.879333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T19:48:32.705719Z digest=sha256:da1a28efd4c08feffe6987b82440fa0aa2117702ea61a156723c5a52f7854daf

Observation bd8d7780-775a-429b-ad41-918618667f47 · outbound

This paper cites Attention Heads of Large Language Models: A Survey.

Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images Attention Heads of Large Language Models: A Survey

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T19:48:32.707991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:48:32.707991Z digest=sha256:07dc17cd9cc762f52ea6349df7056878e17e59abe36553b6590e272b14132dd0

Pith citing papers

No inbound Pith citation observations are available.