Pith. sign in

Paper Citation Record · LEDGER

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization

As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2606.11805.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.11805 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T09:59:43.874631Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact9
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e74a529a-0816-4df2-ab39-c63dca8eea03 · outbound

This paper cites Reconstructing hand-object interactions in the wild.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Reconstructing hand-object interactions in the wild

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:bca64579b9da28b912ff985e3252d5cccb88dd412da6649b3f76f02d0db0e3a2

Observation 5bcaacf8-9861-422e-93b8-aa8e73cfcaf7 · outbound

This paper cites Text2hoi: Text-guided 3d motion generation for hand-object interaction.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Text2hoi: Text-guided 3d motion generation for hand-object interaction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:0bca470c829b4fade614a8ddcbfb691cdf8f219a3cdb60e95b4f28caff8fcb5e

Observation 75454c51-d163-43a7-ba1a-9310d0fd6eaa · outbound

This paper cites Generative pretraining from pixels.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Generative pretraining from pixels

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:7ae69c7102159e6460f3c5094375d79632b28f2a5b845a7e59d290d0dc77be8b

Observation 0aef0014-7ad4-4c98-b7ad-fbebc5a56183 · outbound

This paper cites Alignsdf: Pose-aligned signed distance fields for hand-object reconstruction.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Alignsdf: Pose-aligned signed distance fields for hand-object reconstruction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:30d0cce2d2b415e730ceac5884de78534781967c6b51af97418d2287d7361e26

Observation afcd23be-6a81-4667-988a-5a621e790523 · outbound

This paper cites gsdf: Geometry-driven signed distance functions for 3d hand-object reconstruction.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization gsdf: Geometry-driven signed distance functions for 3d hand-object reconstruction

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:61455371edd40bd4a907d49aa80d6fa8dd64d5c77e37a7d1a33f9e5352559ee4

Observation 3a973239-dea6-43c9-962a-9db66658ceb5 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Taming transformers for high-resolution image synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:9fa3ce13b3d1d42295af6396207f6d0d0729cfa80e678a353af9cc5afd90c146

Observation 6633bb25-b669-4d83-b0a2-b952bc4646d4 · outbound

This paper cites Honnotate: A method for 3d annotation of hand and object poses.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Honnotate: A method for 3d annotation of hand and object poses

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:8e0e048fb01b6f0eb893436e57ca5498e122beedc003ddf1e519693ab096c8cf

Observation f3adef3a-cafc-47fa-b417-a7c9ecf0726d · outbound

This paper cites Learning joint reconstruction of hands and manipulated objects.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Learning joint reconstruction of hands and manipulated objects

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:874daf13d4678cd0a81de5105a841155b90b91d3ad5e57762d41dc40dc4d66e8

Observation 47514864-d4a8-4745-970d-489b0c7f065b · outbound

This paper cites Leveraging photometric consistency over time for sparsely supervised hand-object reconstruction.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Leveraging photometric consistency over time for sparsely supervised hand-object reconstruction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:b17b5e08cca6254f3737d59e32357e4c5e8308ec7c8eba840c4070cd9331406e

Observation 9f5c9d72-17ba-47d0-9e17-8ae72eb7e814 · outbound

This paper cites Classifier-Free Diffusion Guidance.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Classifier-Free Diffusion Guidance

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.855035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:b2395805032ba7184bd5b68d02b06d0207b1365a19a4efd7768fd0337b841f18

Observation 39f722dd-d2e9-4b4e-bb3a-7162f60a8462 · outbound

This paper cites The Curious Case of Neural Text Degeneration.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization The Curious Case of Neural Text Degeneration

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.857582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:92b32c39151e267e537a34a028df2ad49575acc0ed0bbd61cc531087596fec35

Observation 3dba5774-d574-4c82-9c7a-783b8395959e · outbound

This paper cites LRM: Large Reconstruction Model for Single Image to 3D.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization LRM: Large Reconstruction Model for Single Image to 3D

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.840287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:678dfeaebbef7b604fa48fabd0ae6b686b7d075205dcd017977a32057ba44373

Observation 4e8b580a-8a3f-430e-a8b1-dc741115c386 · outbound

This paper cites Zero-1-to-3: Zero-shot one image to 3d object.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Zero-1-to-3: Zero-shot one image to 3d object

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:6a8703af74ee9964f93e7daf8a5231f15e6cf22a884423f5947c802bba3954a6

Observation 4862b1c9-c01a-4ca7-9ef9-df1d0c8e38e2 · outbound

This paper cites Decoupled Weight Decay Regularization.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Decoupled Weight Decay Regularization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.846020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:b16338a271050b6e32c0f7103184344f03d9687833eb2c9453ad6f00d4303e63

Observation e761997e-92a6-4dd9-adef-a56f5ac62f10 · outbound

This paper cites Reconstructing hands in 3d with transformers.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Reconstructing hands in 3d with transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:5041dbfb86d4925c8b2d1d88a4abbede4306c0ee3c26fab3a72a2f1a196c9148

Observation 21740990-6cc8-4f06-8e97-1bd63a19891f · outbound

This paper cites DreamFusion: Text-to-3D using 2D Diffusion.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization DreamFusion: Text-to-3D using 2D Diffusion

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.849412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:3e21ee8d1b37b15a7a39dc2bb1d0154436f7a071b7e598205a9054ad254fafce

Observation a1b906c0-8d1d-48b8-abe8-cd4cff78cad8 · outbound

This paper cites Learning transferable visual models from natural language supervision.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Learning transferable visual models from natural language supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:52ac39aa5ec4d5a3238944fe5c4def2b8d7e079d55e428aaea04262b91253b9a

Observation c9793f5c-5e6b-4fb5-8bec-574f23a04e8f · outbound

This paper cites Accelerating 3D Deep Learning with PyTorch3D.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Accelerating 3D Deep Learning with PyTorch3D

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.843276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:abaeee16fd27439c7386fbaba121510057ddbac0093ca6ac5a18301cfd466b06

Observation 944895e8-a486-498b-98f9-64143b1dea53 · outbound

This paper cites Generating diverse high-fidelity images with vq-vae-2.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Generating diverse high-fidelity images with vq-vae-2

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:aab75ee7ca45f243c0247889d929d020cb31f72bd79b98ff17ca69ef0ba60b83

Observation 45e4b46a-af01-4d77-af92-b2e33f728b67 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization High-resolution image synthesis with latent diffusion models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:51469397ea934086e141b8b07273e40c03af3149df2cebe4171c9dcf5f847b83

Observation d70d405b-88e8-4ff6-b92b-ef2a7f8dc4f9 · outbound

This paper cites an unresolved cited work.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:93381377c3fa5e360f7c1aaa0cb98c7417a7fa37059385045fc35c4f4d0402f3

Observation d15bd7d1-bd28-426e-8c9d-1ee2c80a21af · outbound

This paper cites MVDream: Multi-view Diffusion for 3D Generation.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization MVDream: Multi-view Diffusion for 3D Generation

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.837575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:6f4ee385724088d11c187938acc9b38bfbdd15c95d4dd1bec73050abc175561a

Observation fa8219fd-72b9-4af5-90bd-d952167abea5 · outbound

This paper cites MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware Diffusion.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware Diffusion

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.852339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:1fe7082b534bbaf9456ff69f5cd6fac24121e7284e6e39399665903e26f6e8a1

Observation 08085502-c884-4ad7-922f-7df00d222544 · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:879fc1d747e70f6545573e4cc1c2849d7693360532f89466148f8ca1092b3814

Observation 38f560e2-aba6-4985-8b1b-cd86488c0722 · outbound

This paper cites Pixel recurrent neural networks.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Pixel recurrent neural networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:16f3b424b2d6e6618a159082052167fce792ade1033eae848000232130ea13b8

Observation 43b01606-f003-463b-b93a-3b6624e04fb3 · outbound

This paper cites Neural discrete representation learning.Advances in neural information processing systems, 30, 2017.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Neural discrete representation learning.Advances in neural information processing systems, 30, 2017

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:2e56cbe996da86776f02ab69807b5e02ca20d004c648cc95585b0a8ed083f433

Observation c72e84f7-61f3-4142-bda9-7fd213c20923 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612, 2004.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612, 2004

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:956c90e1d43dbef5f925d51a1687ac6cfaf45be3386b8e979dce4197032ad19b

Observation c7dfc73a-ad96-469e-97bd-aa1a5ed5fc19 · outbound

This paper cites InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.860266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:f1243534cfe48a996c3e8909842b49763ba5893aa21dfbd34b862a115d8ba829

Observation 6a08f6c2-9900-4a80-bd15-3acfe315393c · outbound

This paper cites What’s in your hands? 3d reconstruction of generic objects in hands.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization What’s in your hands? 3d reconstruction of generic objects in hands

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:94f67e972a5543e617746cb90fc823828126d42258f8698a6aa249b925c6a3e9

Observation 70b19a03-c4de-4cb5-9edc-75d74104b24b · outbound

This paper cites Moho: Learning single-view hand-held object reconstruction with multi-view occlusion-aware supervision.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Moho: Learning single-view hand-held object reconstruction with multi-view occlusion-aware supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:9de8a2cce0eeafd2b1dd768342e015749a72b1f27463b9e07ce4fde0d8eec72f

Observation 3298a7e5-138d-4a78-a72d-ad8119e15c72 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization The unreasonable effectiveness of deep features as a perceptual metric

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:547042c95c57a738519b2339fddce509e55a30a9a401e3d52032d1ca13a73f85

Pith citing papers

No inbound Pith citation observations are available.