REVIEW 2 cited by
Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Creating engaging narratives from visual data is crucial for automated digital media consumption, assistive technologies, and interactive entertainment. This survey covers methodologies used in the generation of these narratives, focusing on their principles, strengths, and limitations. The survey also covers tasks related to automatic story generation, such as image and video captioning, and visual question answering, as well as story generation without visual inputs. These tasks share common challenges with visual story generation and have served as inspiration for the techniques used in the field. We analyze the main datasets and evaluation metrics, providing a critical perspective on their limitations.
Forward citations
Cited by 2 Pith papers
-
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
A new multi-frame visual storytelling dataset with explicit entity grounding, plus a fine-tuned Qwen2.5-VL baseline that reduces measured hallucinations by 12.3%.
-
Entity Re-identification in Visual Storytelling via Contrastive Reinforcement Learning
A contrastive reinforcement learning method with synthetic negative stories improves cross-frame entity grounding and re-identification for a 7B visual storyteller, evaluated only on the authors' own dataset.
Discussion (0). Continue with ORCID to comment.