Pith. sign in

Paper Citation Record · LEDGER

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception

As of 15 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2606.20764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.20764 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T18:22:41.189820Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact14
  • verified fuzzy0
  • unresolved8
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 971ab894-c77f-4a63-8a20-be5ddc178ee0 · outbound

This paper cites One-shot gan: Learning to gen- erate samples from single images and videos,.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception One-shot gan: Learning to gen- erate samples from single images and videos,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T18:22:41.189820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:26b45722e2d47967018db299b2575875160a5b23dd65d02ff0406120e284e0d1

Observation 2c24b681-a9ea-45e8-92f6-67e6818d6003 · outbound

This paper cites Roadwork: A dataset and benchmark for learning to recognize, observe, analyze and drive through work zones,.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception Roadwork: A dataset and benchmark for learning to recognize, observe, analyze and drive through work zones,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-26T18:22:41.189820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:5733ba8fafb458d81c2edbaeb6297b8e7b6f63cd380080f7ea3791029a1b8a04

Observation 80e42bb9-c25b-44ff-97d5-f2ed659e8308 · outbound

This paper cites LaRS: A Diverse Panoptic Maritime Obstacle Detection Dataset and Benchmark.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception LaRS: A Diverse Panoptic Maritime Obstacle Detection Dataset and Benchmark

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:09:30.046484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:b84384c816edb9d257e4b7b7e9a1aff860f8b06b089436055e8f328b60618e05

Observation 0be012c0-84e1-4719-a1b6-69953c5f6c82 · outbound

This paper cites Srmf: A data augmentation and multimodal fusion approach for long-tail uhr satellite image segmentation,.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception Srmf: A data augmentation and multimodal fusion approach for long-tail uhr satellite image segmentation,

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-26T18:29:41.375748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:071c6056b06d6603706e1fa0b07a7a89c00afa875be252397a21e93a7d292562

Observation aae15f54-5bdc-4ba9-beef-511a734d930b · outbound

This paper cites Long-tailed object detection for multimodal remote sensing images,.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception Long-tailed object detection for multimodal remote sensing images,

Reference 5

Resolution
verified exact
doi, observed 2026-06-26T18:29:41.372959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:3101e3d463d5f50e5d075c34b9c70e56e38430c2ed67467986b3c4455017d438

Observation ac6ac4ee-bfd1-4af6-8e22-bf42a745fd00 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer 25 Vision and Pattern Recognition, pp.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception In: Proceedings of the IEEE/CVF Conference on Computer 25 Vision and Pattern Recognition, pp

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T18:29:41.378555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:bb4d1967329561844dfce55ed96d8eadb3fd0fbabfc2d0bfeb1568a77f04bc69

Observation 76b56d75-c241-4a17-b882-03e28022bc7b · outbound

This paper cites Hyun Cho and P.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception Hyun Cho and P

Reference 7

Resolution
verified exact
doi, observed 2026-06-26T18:29:41.370630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:4d08351257eee9a2a25d9e834b3df07f697b9ef9a230429fc7dc2e7d0addba05

Observation bfc76c50-b774-4465-a9d3-7bb0e248f0be · outbound

This paper cites Probabilistic contrastive learning for long-tailed visual recognition,.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception Probabilistic contrastive learning for long-tailed visual recognition,

Reference 8

Resolution
malformed identifier
arxiv_id, observed 2026-07-04T03:09:30.076435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:ff870f03e41a15cceebe8c4be2cd937cd5a3a905d349829984f0a6975fa8c6a5

Observation a804ec5b-f13f-46db-a662-380ae5a06aa0 · outbound

This paper cites Decoupled optimisation for long-tailed visual recognition,.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception Decoupled optimisation for long-tailed visual recognition,

Reference 9

Resolution
verified exact
doi, observed 2026-06-26T18:29:41.380626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:4b0e8db4c5d9cd0a95498a50b97d22b628c26f3f7a3e4b730ef798016c23fee5

Observation be8c30c2-22c7-46a6-8cb4-8341c53b74eb · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:09:30.079876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:3888a1de195838602a9899fe6acf9c5eed39e201744ab87b72ec5eb1b6e5e8ec

Observation 20a27811-3f08-4254-a2c8-0816551f5c93 · outbound

This paper cites Wuerstchen: An Efficient Architecture for Large-Scale Text-to-Image Diffusion Models.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception Wuerstchen: An Efficient Architecture for Large-Scale Text-to-Image Diffusion Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:09:30.089338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:09d4c7aa4303221b7c1c73f2447f64dafa0a5f98ffae19791b9823719ef0b877

Observation b7670fc0-d199-432b-bb85-f706a002418a · outbound

This paper cites Geodiffusion: Text-prompted geometric control for object detection data generation,.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception Geodiffusion: Text-prompted geometric control for object detection data generation,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T18:22:41.189820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:c54ebb0ed3ed9fb89ad005e53197fb518e098f8c8c5187fe039c630796fd7e3f

Observation 81b76007-af42-4a58-95ba-56a193961444 · outbound

This paper cites Data augmentation for object detection via controllable diffusion models,.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception Data augmentation for object detection via controllable diffusion models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T18:22:41.189820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:7211de8acaa00b1c4e459e974fcc4cd9d3cf67427dbc1d1858a9856435e341c1

Observation 3c8af3d5-b50c-4a0e-b497-e7d6a5a84b4c · outbound

This paper cites Mini-gemini: Mining the potential of multi-modality vision language models,.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception Mini-gemini: Mining the potential of multi-modality vision language models,

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-26T18:29:41.368543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:4cc6ceac638fbb0c6ffedb6e49ddd1feed477dcc25ca8fcc253d5bf9ff72a031

Observation a7bf10db-3965-469f-9212-8b0376aa8f1e · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:09:30.076334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:9633163afc3e4c501d381e8300ec996379485136449d2a1f3ccd10f73c94c836

Observation 3b0c1a80-c527-431f-bf57-ae15519b9452 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:09:30.080597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:ea328be59463370b650de5db6569babebecf41dbc5ee628552e66d3e61a996e2

Observation 0a55ebc3-dee9-46b8-b055-2a130ade9151 · outbound

This paper cites LLMGA: Multimodal Large Language Model based Generation Assistant.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception LLMGA: Multimodal Large Language Model based Generation Assistant

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:09:30.056467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:ffa0d6a70cec756e720d8fb277c042c03f267cc341bcb34bf5794d9665406137

Observation 0c9ab60a-bee1-4891-ab69-5a7cc2753189 · outbound

This paper cites Spatial chain-of-thought: Bridging understanding and generation models for spatial reasoning generation,.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception Spatial chain-of-thought: Bridging understanding and generation models for spatial reasoning generation,

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:09:30.072558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:79a75818e9d793e68b644fb066181531dd5996ca5815c7652ba7e6353eabc738

Observation 98e8b463-f5f5-4bff-a80e-9ffcfca3db71 · outbound

This paper cites Available: https://weichens.github.io/spatial chain of thought/.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception Available: https://weichens.github.io/spatial chain of thought/

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T18:22:41.189820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:f9e5b649faff60f6bafec782dcc3189422909fdf2911beeb80448d65e32cffbc

Observation c3b08798-36fd-49f8-9dec-25f9d7daf527 · outbound

This paper cites From yolo v1 to yolo v11: comparative analysis of yolo algorithm,.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception From yolo v1 to yolo v11: comparative analysis of yolo algorithm,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T18:22:41.189820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:c2822fd8007d843dbc10054b5f5bd751cea8a47b2cbeee5dd82dde36b8606ec7

Observation b2013627-a87f-4443-8243-1bb784ad43f6 · outbound

This paper cites Faster r-cnn: Towards real- time object detection with region proposal networks,.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception Faster r-cnn: Towards real- time object detection with region proposal networks,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-26T18:22:41.189820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:8cc6645b65640a763859c55a96faaef9a9a7e4018184b4ab900c620ea271e599

Observation 31688ad0-6c4d-4802-a706-3ddbe872580b · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:09:30.063735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:996399690d0383766ff8709cc90a08a07b8da27068c5f31872eec0f7bdaa4d7e

Observation e30dc27f-1a24-48a1-9ee8-bbe94d19a5df · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:09:30.085409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:2946d5c26b13393e9d5b47287bd23e5d44ca0bf97165f216a327d079f9e710f4

Observation 96b9b556-bfb7-41fc-b2c3-6bcba51f705e · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection,.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception Grounding dino: Marrying dino with grounded pre-training for open-set object detection,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T18:22:41.189820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:91acfdebc343132f5600a7291994fb32a53b4f0f523f6d991f22976b9c3a7524

Pith citing papers

No inbound Pith citation observations are available.