Pith. sign in

Paper Citation Record · LEDGER

SEED-Story: Multimodal Long Story Generation with Large Language Model

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2407.08683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.08683 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:46:46.835013Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:58.344114Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a1f203de-ad29-49b0-afba-206648d31de8 · inbound

When Attention Sink Emerges in Language Models: An Empirical View cites this paper.

When Attention Sink Emerges in Language Models: An Empirical View SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:41:03.772015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T17:41:03.674759Z digest=sha256:bca8c504c5a88ea69ca527ce8cac32096c5cbd2058f8150ed6813ba60ed7b60b

Observation e1af66a8-8689-4657-acb1-efc641385222 · inbound

Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models cites this paper.

Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T18:32:12.134723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:32:12.134723Z digest=sha256:9b155fca62df889b06b740c547899883a294877197dbb43657f11c0e9dc00784

Observation b20491cf-043c-4fd9-8ce2-7f055c0f8b0a · inbound

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection cites this paper.

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:56:52.702184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:56:52.702184Z digest=sha256:d8cf280bd04f73a30a7311b44f8e738d763c4b4ed337a08eae8020b2e6cfe9ac

Observation aee6e05b-faa5-45b5-b5ff-64917286e43d · inbound

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation cites this paper.

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T11:10:54.868302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:10:54.868302Z digest=sha256:be0d6a8ce5ed24ea48ff447c9c42f61b16ed34ccb0c6a343bca75ec82b888615

Observation d6d15063-c71c-42ca-afbd-4471e4f63b73 · inbound

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation cites this paper.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.593471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.593471Z digest=sha256:ef1d0118b1e762656446cf18485b71e67c6633317833a03927ddf66e6d0a30ed

Observation 5c8e3fd8-4ab9-4867-a02e-6da472943178 · inbound

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation cites this paper.

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T18:45:41.323022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:45:41.323022Z digest=sha256:0f6d0b2da82964b80900a6095a934cba58ceb53fec89e550412dc29f05594630

Observation 1c9b2f8d-ec45-4230-8951-73f294e3a390 · inbound

Olympus: A Universal Task Router for Computer Vision Tasks cites this paper.

Olympus: A Universal Task Router for Computer Vision Tasks SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.647447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.647447Z digest=sha256:a832ee26a062a4eaed5047f5737c49ad16000d2792d2cd8b9821d4e30163b28b

Observation 894f27ee-f34b-4f90-a4cb-36c856540de2 · inbound

SpearBot: Leveraging Large Language Models in a Generative-Critique Framework for Spear-Phishing Email Generation cites this paper.

SpearBot: Leveraging Large Language Models in a Generative-Critique Framework for Spear-Phishing Email Generation SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T15:21:48.443272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:21:48.443272Z digest=sha256:96cba5b92fb05fda41e26cc91658f5e4e3cf45cf3f804a6a35a69c2b54ecd305

Observation 8a22b6f7-d150-430c-a479-7813c71dccbe · inbound

IDEA-Bench: How Far are Generative Models from Professional Designing? cites this paper.

IDEA-Bench: How Far are Generative Models from Professional Designing? SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:39:28.407921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:39:28.407921Z digest=sha256:ad7c5ae30585cb051c66ce5390cdb53c01e4d4e5e0ea162706a9dcd1d9b45e2e

Observation 18d34d83-48a0-498f-a4b9-08d055b977c7 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.485683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:a3c8ec6b54c76bd7a151db81e34e73d4313e99abc1772f55c61cde97f0f90273

Observation 14abd264-1d5d-4849-bf56-5785657f131a · inbound

VideoAuteur: Towards Long Narrative Video Generation cites this paper.

VideoAuteur: Towards Long Narrative Video Generation SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T21:08:51.263917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:08:51.263917Z digest=sha256:15c0bf4456f1fc71d0b9664458095c4264390e5d680dc3e0793929853e0270e7

Observation 2f8f7139-3aad-4d1e-916c-170fa4030d46 · inbound

Curiosity-Driven Reinforcement Learning from Human Feedback cites this paper.

Curiosity-Driven Reinforcement Learning from Human Feedback SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T18:21:37.455009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:21:37.455009Z digest=sha256:b87e45e07ace139e97ebbe38043de93200d4eccc39584b4686c0a635ae74fe6d

Observation a0e16d60-5f0d-4563-b1fa-a0d19c91b1f6 · inbound

One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt cites this paper.

One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T15:53:51.274934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:53:51.274934Z digest=sha256:2ef8f1eb6e242593fcb1bba78dc44c7fa9905607b863feac6c27c2cee981048f

Observation b003c0d9-950a-40cd-b724-cc0b7cff4dba · inbound

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths cites this paper.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.456245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.456245Z digest=sha256:9e297c35f6c69ca62261d06c95735634299b53a3175593c7b7f2dc5acecbc6ba

Observation 1b6bf300-0782-4bd2-a95d-d118452250fe · inbound

CineVerse: Consistent Keyframe Synthesis for Cinematic Scene Composition cites this paper.

CineVerse: Consistent Keyframe Synthesis for Cinematic Scene Composition SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:46:46.835013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:46:46.835013Z digest=sha256:dc80fd2fa0145a54d569075fbb003f64b83712eb98f3ea3f954dcc17be8e46c0

Observation 0a291368-bccf-4c33-95d8-25e55937bc4e · inbound

Character-Centered Dialogue Generation from Scene-Level Prompts cites this paper.

Character-Centered Dialogue Generation from Scene-Level Prompts SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:34:53.671351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T13:31:43.083678Z digest=sha256:aa26280eaf1735610155ef4cd287d29457e601ac6ffd3aef58c66868d2982743

Observation 70b007af-864d-4a5a-9857-76c8a9245cb5 · inbound

Rethinking Causal Mask Attention for Vision-Language Inference cites this paper.

Rethinking Causal Mask Attention for Vision-Language Inference SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:32.799948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:32.799948Z digest=sha256:53bdf64980289f28b7d30a981e72f49da91b38349ce9986e8f6ef7692bd00074

Observation bc59aa92-dab7-4113-ae87-2e9112b87892 · inbound

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation cites this paper.

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:11.474850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:11.474850Z digest=sha256:0a98736fd141b561f72382986b49dc2e3d7206022c5f27d5930bf1e8e48a39f0

Observation 32a27c34-22c9-4719-b64f-79592ae3d8ca · inbound

Audit & Repair: An Agentic Framework for Consistent Story Visualization in Text-to-Image Diffusion Models cites this paper.

Audit & Repair: An Agentic Framework for Consistent Story Visualization in Text-to-Image Diffusion Models SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T18:47:33.059856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:47:33.059856Z digest=sha256:85f849166eeed9d3e60a88a2774544f01f57c1244806fc5975c5e5ef5c63205c

Observation cc1ed02f-1786-4f0c-b0e8-d83496e3c337 · inbound

FairyGen: Storied Cartoon Video from a Single Child-Drawn Character cites this paper.

FairyGen: Storied Cartoon Video from a Single Child-Drawn Character SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:35:00.968329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:35:00.968329Z digest=sha256:bfe0df889ce7b01c3b149b6e25457f16d038f9613fddcb7ba87dbbbc6f8e1a56

Observation e0ff973f-9df2-4bfb-b8b7-588ae49f19f2 · inbound

Captain Cinema: Towards Short Movie Generation cites this paper.

Captain Cinema: Towards Short Movie Generation SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T14:37:19.553823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:37:19.553823Z digest=sha256:f2bb3156933e47c4cd95ea4d7aaf4fb683cd14357280f78f68212441b2aa2dc1

Observation 7c844ef5-0779-4b9f-8dfc-83be9bb08477 · inbound

Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs cites this paper.

Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:19:01.972919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:19:01.972919Z digest=sha256:bd9918dccef8c53c9ccc14cbff7374295f032aaf84506ec6e8a737c9489e7f0c

Observation f698b1de-2d72-4ff7-af37-a1589c6f4d32 · inbound

StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization cites this paper.

StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T10:47:00.754074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:47:00.754074Z digest=sha256:b0a2b0883d49834d5453ca7a3d26f8416b961148463619cae1e9152a499ad3f2

Observation 3a6adcfe-357c-4ea9-b545-c1d4cb8846ca · inbound

Numerical Study of Oblique Detonation Initiation Assisted by Local Energy Deposition cites this paper.

Numerical Study of Oblique Detonation Initiation Assisted by Local Energy Deposition SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T17:32:40.194630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:32:40.194630Z digest=sha256:25a5f9366c7c63b847f9baa1e3717662cc755b62b86f1f0081920e5c6cf6e343

Observation 026d94fe-7633-4f3e-b69e-6c4944e7ad2c · inbound

Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation cites this paper.

Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T17:33:24.931590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:33:24.931590Z digest=sha256:3f08fdfccee7493dc39f18641d69de16703d362718c86095acd8592a079e5901

Observation 385f0895-f073-4f56-95f5-e73eaf925da7 · inbound

Story2Board: A Training-Free Approach for Expressive Storyboard Generation cites this paper.

Story2Board: A Training-Free Approach for Expressive Storyboard Generation SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:09.401643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:44:09.401643Z digest=sha256:e1a36a640de5cbf9e0a409f87132ca29d649fd37935c1e6e2df4c1e59a4a494b

Observation d6ca206d-e27c-4a3d-9ecc-e3e62a9f4729 · inbound

Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models cites this paper.

Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:15.429255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:15.429255Z digest=sha256:c858df27601bc88783ca20a45461cd5de874477f9903f08326329774f181c675

Observation f37382bd-6040-4149-a3d0-81b46c1fd6be · inbound

LongLive: Real-time Interactive Long Video Generation cites this paper.

LongLive: Real-time Interactive Long Video Generation SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 107

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T03:52:59.507936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T03:52:59.287555Z digest=sha256:8b9387491982fc8377700741946342717e24cf2c9ede3ee3e962fbb09f2d4aac

Observation d7fca0e8-a08a-4c49-b5a8-96b303c0cdd7 · inbound

Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization cites this paper.

Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:52:06.073559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T20:50:53.673037Z digest=sha256:9de2109e9f3e6ad1d39b7474ab4a534a74c0f95610ceb8ff6659c62864be3ace

Observation 0ba59551-5d98-4bd8-84b4-a83972c1a104 · inbound

FreeStory: Training-Free Character Consistency for Free-Form Visual Storytelling cites this paper.

FreeStory: Training-Free Character Consistency for Free-Form Visual Storytelling SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:58.345540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T00:10:35.009818Z digest=sha256:532384796543a531d0cb2dade71cb7a58af61f692289e21263cd912216b9986a

Observation 540310c3-fd74-4837-aed5-84e2f6369bff · inbound

FreeStory: Training-Free Character Consistency for Free-Form Visual Storytelling cites this paper.

FreeStory: Training-Free Character Consistency for Free-Form Visual Storytelling SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T12:23:07.999687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:23:07.999687Z digest=sha256:4aaed14192039b419ff7d8b8c5275f424fd2ea29c9a8ac5f76709e0e34c360b4