Pith. sign in

Paper Citation Record · LEDGER

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning

As of 3 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2604.04746.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.04746 v3

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:16:58.323955Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:07:28.254518Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c48b0112-00ed-4818-9745-dc8d6737318e · outbound

This paper cites Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:10:53.823287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:459a59a672c7147e73093ef61f88c3957eed2fa9bf3de3c669c3dc8a62df8bf2

Observation 122e0cea-d14a-401c-9c97-0662210c0192 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:10:53.745000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:f295a0a4c49a0eac7c9ff50d45b35c74e8fafb38e7824db0ede9491e255d94f1

Observation f87889a0-6e2c-4247-a13e-21de85e6ef2e · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:27:54.173931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:15140f6a2a4ebabe808189ac5de44c5dbb88573c4d38b32e3a4f8bd5db7d7383

Observation fe7b2ad6-500a-4de1-bdef-44db87f0268f · outbound

This paper cites GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:10:53.793691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:d3cfd95b4cadb606f28285f41ae3316c78599e4ce29849945808fdf1fa452158

Observation 9c2a2f93-6c03-460c-b878-df31db549271 · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:10:53.770009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:7d93a601a1b8a21634c92d85c9104d55f0944c8335c6d6440b50924744e12ff5

Observation 390e76f5-04cc-4123-aafa-e0252ab906b1 · outbound

This paper cites Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:10:53.787696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:c6bb44bd85df2d4df58ba3f2b30c36e589f342db8b8d178a6a8937c74265816e

Observation cf9c15bd-dda4-4d12-bb6b-89ed2ba72ab1 · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:10:53.779102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:d9b5a6674c4b0689e10bc61c2ae70532797ca2a6de7af1cef40b25601728224d

Observation b922f013-e444-4eac-bea0-709163224b4c · outbound

This paper cites EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:10:53.736351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:5e4574d38ebc69b680602635e19db75f117f4da87a14a344700f6b719b21b5d6

Observation 4243c24a-2608-44c6-bb20-c55d35749a40 · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:24:27.819650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:c50008a637bb53f1fa07ed508260ae3a3ef60992148afebae2edce135c70d992

Observation 211c994b-f3f4-45f2-a9ff-72e6f011aac3 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:10:53.714152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:b5406045f941700a196d6e29026e4ccb8da02137685650dd752c28faf2c049ee

Observation c9bb5c41-7094-4727-b2ce-cf082f638fbd · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:51:13.882527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:8557ffeb95bff4bbd9452b7cf70254b9409a765e7f0858d425bb9e7d0a430081

Observation 1f819ba7-250d-49e9-8d67-6f4ce8e40496 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:27.005672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:ce88facb5fd47ce26dfa9d51dfbe7115fc3e63a15dee6e466945ef24038442ae

Pith citing papers

Observation 7281173a-0a6c-4164-a079-ad1a89d19ed4 · inbound

Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence cites this paper.

Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:28.254518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:28.254518Z digest=sha256:a8da36c950173f6eb1773fe8a20336750576eb3b01bdcca9a317c923898c0bc1