Pith. sign in

Paper Citation Record · LEDGER

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents

As of 10 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2506.04606.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04606 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:41:39.824941Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee0b2cfa-35cf-4910-93c5-1836c154d872 · outbound

This paper cites GPT-4 Technical Report.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.714323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.714323Z digest=sha256:142d2c58f8527bdbba1cc031fec1bb74d3c67a3fac2e962afd25e9b76a5c5ee7

Observation e68f483f-94a4-47ae-9c1d-88bd9c6501d6 · outbound

This paper cites AvatarCraft: Transforming Text into Neural Human Avatars with Parameterized Shape and Pose Control.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents AvatarCraft: Transforming Text into Neural Human Avatars with Parameterized Shape and Pose Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.736092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.736092Z digest=sha256:2ae4aa4fe9f9efd09d0b7226228596b1bfee1845689e2629e5c342ce9fa25b81

Observation c37a48f8-18fc-446c-a3df-fe125ff6eb97 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.740007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.740007Z digest=sha256:791dfe683b75abc62c9c008a2652b44e23b7f5dacab271b002493d46d82b14ca

Observation ac184a21-3bb7-450a-a9d8-33504e110371 · outbound

This paper cites TADA! Text to Animatable Digital Avatars.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents TADA! Text to Animatable Digital Avatars

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.744076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.744076Z digest=sha256:849d61a856a50f5555c2e284ce021fde8d8447163cada20f80ff148b28000689

Observation 795b7840-cf2a-4cf9-9c5b-0025bdffbbe3 · outbound

This paper cites Visual Instruction Tuning.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Visual Instruction Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.747720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.747720Z digest=sha256:6b2f9f5bcd8b70e6837153c5acddcd9c04737a26de7e5ffe76ef898314199e92

Observation 464db041-1c9c-472b-9cc4-221915909147 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents AgentBench: Evaluating LLMs as Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.751445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.751445Z digest=sha256:246fcac314ad2827a484b316d6aed1296d02854b70b28ae9678ae56a9a447bec

Observation 9d43d8c9-e1b8-403a-ad63-96d2acb49759 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Self-Refine: Iterative Refinement with Self-Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.759301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.759301Z digest=sha256:b29c435ac2714816a910b4fdf60079eadf45dc38513b6862f80dafcd701f5300

Observation 1353627b-702d-4fee-9e35-aedf2b9feafa · outbound

This paper cites URL https://arxiv.org/abs/ 2304.03442.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents URL https://arxiv.org/abs/ 2304.03442

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.762937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.762937Z digest=sha256:b02fc685cfca8211a632daa71bb6be78cdbf3da883dfed4da198f3ed35f72671

Observation f6aec695-2dc4-42fe-9694-2ba047b30b1c · outbound

This paper cites CharacterGen: Efficient 3D Character Generation from Single Images with Multi-View Pose Canonicalization.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents CharacterGen: Efficient 3D Character Generation from Single Images with Multi-View Pose Canonicalization

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:41:40.013907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:41:39.771357Z digest=sha256:76b07f829fb8110d69490f72d4cb533f5d95673aa86e7294d3b021cd10399106

Observation 0ae27a19-bb7d-45c8-a90a-163c9f79c48d · outbound

This paper cites Iteration of Thought: Leveraging Inner Dialogue for Autonomous Large Language Model Reasoning.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Iteration of Thought: Leveraging Inner Dialogue for Autonomous Large Language Model Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.775381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.775381Z digest=sha256:1099b3d2df5e0b61530990b55eae3c713d33f05e0d49c2e9e7a9f4177297fe14

Observation 2440ee96-9bca-442c-bc12-f83a7e3550d0 · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.784487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.784487Z digest=sha256:9c3432aff7d4c06c7660b361d156d7bc523a051ed45924302f783dc5c415f8c9

Observation 109e1a66-33b1-4ffd-80fb-7bf8336653d3 · outbound

This paper cites DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.788222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.788222Z digest=sha256:690127b0cc11df9d9b10c2f678888895f0ff633fbbf2ce2e8547771068229e87

Observation 672510bf-4af3-49a9-8d90-f75f67a417b1 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Gemini: A Family of Highly Capable Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.792243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.792243Z digest=sha256:7040f567d00b68967e1150493212700f587892765c3d843f4bf9508e324dbe71

Observation 179e2aa8-ddb8-47f0-8e3e-2f727b916229 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.799981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.799981Z digest=sha256:f4ab9f46981172d19b016cc3e1d86823ae48a7823015070904f6385372e0bf32

Observation ab2b0cbd-497b-4bef-aee8-4226b831beeb · outbound

This paper cites Structured 3D Latents for Scalable and Versatile 3D Generation.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Structured 3D Latents for Scalable and Versatile 3D Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.804402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.804402Z digest=sha256:0c5557f211e115abb2fabee397ae6e1631905e7de95fb52496321f56573e10a7

Observation 71778dda-07f4-4aaa-a986-80a7c5997c52 · outbound

This paper cites WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:41:39.915452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:41:39.808081Z digest=sha256:42fce938c249f91a47bc56596a5c444a3dbe60b988f1919f0d1cbcbf8c6291ac

Observation 46d7bc0c-c24a-429d-b1c0-e129e2f8af17 · outbound

This paper cites AvatarGen: a 3D Generative Model for Animatable Human Avatars.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents AvatarGen: a 3D Generative Model for Animatable Human Avatars

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:41:39.897848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:41:39.811748Z digest=sha256:69eb7cfda7e5a349a2feead70e56c9dcbe5aa6fc8a7897ff82383ee5ed86a09c

Observation 4c3688f1-8754-4aa2-a93a-1081b07e0e2b · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Adding Conditional Control to Text-to-Image Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.816391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.816391Z digest=sha256:68ab9207a762269152e91423cacc51e2911a143047e2679767f8e4267a4f717e

Observation dfa004f7-bd2c-4913-a1c4-b38a39225106 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.824941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.824941Z digest=sha256:a3d708aedf4afb4565c55571082d96979fed743fe9ad36182f398b62e74acea8

Observation 9d6d6703-8fe7-407a-9bd1-6bbbc0a90cac · outbound

This paper cites doi: 10.1145/2816795.2818013.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents doi: 10.1145/2816795.2818013

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.755671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.755671Z digest=sha256:9670e5a1c40a5a98f1d5b12a7539e6345e33058a4f727f5b179dcfb52297ff85

Observation 3ed96b5b-db96-46d3-ae2e-966c36cebbd1 · outbound

This paper cites Expressive Body Capture: 3D Hands, Face, and Body from a Single Image.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Expressive Body Capture: 3D Hands, Face, and Body from a Single Image

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.766934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.766934Z digest=sha256:4b82a396c6da18cc94baa444411247875e41c1472f48abae0b01c85035bc3592

Observation 96cc0822-947f-497b-bf70-8c2d52e32e5b · outbound

This paper cites Zero-Shot Text-to-Image Generation.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Zero-Shot Text-to-Image Generation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.779709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.779709Z digest=sha256:9c9b527013fc06934921fe9b26779cf5e4d9015a4fa2b96dc1207a333996c7a0

Observation ba95ddda-170e-419e-b4f0-e00ec34ef932 · outbound

This paper cites EVA3D: Compositional 3D Human Generation from 2D Image Collections.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents EVA3D: Compositional 3D Human Generation from 2D Image Collections

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.727380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.727380Z digest=sha256:7a80f5224a1b514d7fa597a8d9d781baef8d4b5274e025327631e8a3d9e3c57d

Observation 5a9a7a47-5443-4d02-b457-2dcc451d77ac · outbound

This paper cites AG3D: Learning to Generate 3D Avatars from 2D Image Collections.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents AG3D: Learning to Generate 3D Avatars from 2D Image Collections

Reference 2023

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:41:40.269231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:41:39.723413Z digest=sha256:7a2c6c81aab7affe97866785340dbd6b1ae0fefa877d86714b30334c42c2979b

Observation 0326ba3a-d26e-4e6e-90e5-1cc9b3fe4256 · outbound

This paper cites StructLDM: Structured Latent Diffusion for 3D Human Generation.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents StructLDM: Structured Latent Diffusion for 3D Human Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.732131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.732131Z digest=sha256:2712da4d377dd47fdc2f864d603eb48de662181e7028042f9152cfbf5848e189

Observation 07f8e3c3-0c4b-49da-9c4b-4fc59bc741a6 · outbound

This paper cites LLMs can see and hear without any training.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents LLMs can see and hear without any training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.719320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.719320Z digest=sha256:421ff1c27aa57216dcf92f81b78713364972fb926d8551efb2b7d5355d99b819

Pith citing papers

No inbound Pith citation observations are available.