Pith. sign in

Paper Citation Record · LEDGER

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation

As of 7 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2506.11380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11380 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:14:29.482855Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8bf112e4-dc4c-453b-8352-2d8b5b85dced · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:30.500766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:14:29.011158Z digest=sha256:d4e7a4ebc3dc8ee0f1b09cb37f65f01ccaee659ff91dfe80537c884d90c3c0a2

Observation e609a620-eb69-4781-b6fc-ba4c1f745fb6 · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:14:28.867443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:14:28.867443Z digest=sha256:eed9d81a2f308883095046e8e122cba9cfe40506ba6831a50710b17347f53539

Observation 66bc283c-1e62-4a2d-920b-c5d976a44efc · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:30.332972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:14:29.106652Z digest=sha256:78cbd052b4885fcb345e5facfacd6d1f78fae831974d0fd47732187d2ffe93fe

Observation 321b4cd2-42c1-439b-a103-54ddd6c14e09 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:14:28.955153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:14:28.955153Z digest=sha256:818b3b77b4f4e651c6545a1e771cfa6612cc0d1a5dfae230c90a531f30cf21be

Observation 06fdeab4-66da-4dd4-b4dd-9b766fa1cee8 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:30.408760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:14:29.058811Z digest=sha256:8fd59633686c2a372b0e05d58b4fcf8d968b493e603916be042e12b663163f89

Observation 712feb65-3b2e-4312-9fd6-b0966e665e53 · outbound

This paper cites Let’s improve a step toGwith visual information.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Let’s improve a step toGwith visual information

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:14:30.250877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:14:29.150192Z digest=sha256:3177bce9c092b72b236a33e39774e17dc1d136395c9490fa0b902c7aba3c47d0

Observation a05a9ca8-2151-4fd8-9381-5f379e8da042 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:30.153594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:14:29.200602Z digest=sha256:c6f476d5f9ffdf51c253541bfab96c6c294b0853c70cc1719178c4808c86c87d

Observation 77a8981f-afc0-449a-9d31-6e5ea9f32dbd · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:30.057830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:14:29.243794Z digest=sha256:db3cb7289403f6bbe05bc012edb08d5aaeb530bb77bf95f8f53d47d5d2e51f4c

Observation de4054d2-4b13-4090-a2bb-7f310d10c171 · outbound

This paper cites Tie” for “Textual Quality.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Tie” for “Textual Quality

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:14:29.957433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:14:29.290892Z digest=sha256:cb3943831012f580796f6b4ba737a0de1d143329fe9ea295f4738ffefaa7bcb2

Observation 97f39116-2264-4adb-9b17-92084d95a3c1 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:29.879480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:14:29.344193Z digest=sha256:e1d894a8e67aa42ce9b68ac5a8cbce46f663335035affe87de6028fcb1df743e

Observation 52783f23-e09a-4689-be77-670abb294039 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:29.798378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:14:29.396630Z digest=sha256:d071cd071f0a99ecaff19ad6d5718057577fb39a2b13539ba98b2baeddb07eea

Observation 06bc4e8b-cb5f-4fe8-a427-3d13067a0b17 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:29.711568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:14:29.437039Z digest=sha256:a87292c16d6b6d74b98e341de9f304a7bac14b1f9f7e1703d30ff191f0502eb8

Observation 6f0d2a26-9721-4c60-a1a0-4ab35859dfc3 · outbound

This paper cites an unresolved cited work.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:14:29.630085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:14:29.482855Z digest=sha256:7375babaeb3cf68d5dd81006c14a9f820afb6fb3db95478c6eccdb7b302254de

Observation baed8f9a-6e54-423d-8192-56728cce2a6f · outbound

This paper cites InstructPix2Pix: Learning to Follow Image Editing Instructions.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T04:14:28.819047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:14:28.819047Z digest=sha256:62da7f5d3cbf2003a124598bde1428b09069cd60a4379c402aec59666e8ef5aa

Observation b11bbddd-215b-449a-951e-b61d403315af · outbound

This paper cites LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models.

Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:14:28.908967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:14:28.908967Z digest=sha256:7843b035998ce2f8fa0a60f21e0d99d5d2bdba29140b3486583b4ec9914baa00

Pith citing papers

No inbound Pith citation observations are available.