{"id":"5d2bf57a-9617-4316-b228-2775d412b8e6","arxiv_id":"2411.14213","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Applying M3DDM and MOTIA outpainting to short videos shifts AI-predicted memorability scores, generally helping low-memorability videos and hurting high-memorability ones, but the effects are small and not statistically validated.","lead":"This study tests whether expanding the borders of short videos with generative outpainting changes how memorable they are, using AI-predicted memorability scores on 100 videos. It finds small, inconsistent effects: low-scoring videos tend to gain, high-scoring ones tend to lose, and neither outpainting model clearly wins.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim is supported only by an automated memorability predictor never validated on outpainted content, and the reported baseline-dependent pattern is also consistent with regression to the mean.","rationale":"The reader's verdict is REJECT, and my read reaches the same conclusion. The decisive weakness is that every quantitative conclusion rests on score changes from an automated predictor that is never validated on outpainted videos. This is the same load-bearing concern the reader identified. I would sharpen it further: the specific result that low-baseline videos improve and high-baseline videos decline is a textbook regression-to-the-mean pattern, so even the direction of the claimed effect could be an artifact of the measurement procedure rather than of outpainting. The paper is an honest exploratory study and uses a standard benchmark, but the measurement chain is not sufficient for the central claim. The absence of human validation, significance tests, and a control condition are not minor omissions; they are directly load-bearing. The saliency section in the Conclusions also overstates a result that Section 4 describes as non-significant. These concerns are correctable through additional experiments, so a revised version with human memorability judgments and proper controls could become viable, but as presented the claim that generative outpainting enhances memorability is unsupported. The reader's REJECT verdict should stand unchanged.","tokens_in":7701,"tokens_out":7751,"duration_ms":75872,"concrete_test":"Run a human recognition-memory experiment on a stratified sample of, say, 30 of the 100 videos, including both videos the predictor says improved and videos it says declined. For each video, show one group the original and another group the outpainted version, then test recognition with a Memento10k-style protocol using enough participants per video to estimate memorability reliably. Compare the direction and magnitude of human memorability changes with the predictor's changes. If the human data do not reproduce the predictor's baseline-dependent pattern, the central claim fails; if they do, this settles the concern about predictor bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3 states that \"the primary metric used for evaluation of the effect of outpainting is the change in those memorability scores\" produced by a fine-tuned memorability predictor. Figure 1 validates that predictor only against ground-truth scores for original Memento10k videos, not for outpainted videos. Outpainting changes resolution, aspect ratio, and adds synthetically generated borders: M3DDM downsamples to 256x256 and MOTIA expands the frame with generated content. A predictor sensitive to these artifacts could report memorability shifts that do not reflect human memorability. The paper's main finding—low-memorability videos improve, high-memorability videos decline—is exactly the signature of regression to the mean when a noisy predictor is used to compute change from baseline, and no no-op control (e.g., re-encoding or letterboxing originals without generative outpainting) is provided. The paper also contains an internal inconsistency: Section 4 says the saliency-based outpainting \"did not yield any significant difference,\" while Section 5 calls it \"a significant increase.\" Together, these issues mean the central claim about enhancing memorability is not yet empirically supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether generative video outpainting can improve the memorability of short-form videos. Using 100 videos sampled from Memento10k, the authors generate outpainted versions with two models (M3DDM and MOTIA), compute memorability scores with a fine-tuned vision transformer memorability predictor, and compare score changes between original and outpainted videos. They report that outpainting tends to improve the memorability of low-memorability videos, tends to diminish the memorability of high-memorability videos, that MOTIA marginally outperforms M3DDM, and that saliency-guided outpainting has an effect that is described inconsistently in the results and conclusions sections.","tokens_in":7817,"tokens_out":3628,"duration_ms":31230,"significance":"If the reported effects corresponded to genuine human memorability, the paper would be a useful empirical contribution to a relatively underexplored area, since it demonstrates a content-agnostic post-processing operation that shifts memorability in a predictable direction and compares two recent outpainting models. The use of a standard benchmark subset (Memento10k) and the inclusion of two modern generative outpainting methods are strengths, as is the attempt to incorporate saliency into the outpainting prompt. However, the claimed findings rest entirely on an automated memorability predictor that is never validated on outpainted content, so the significance of the results cannot be assessed from the current evidence.","major_comments":[{"comment":"The primary metric is the change in memorability scores output by a fine-tuned vision transformer memorability predictor, as stated in Section 3.3, but Figure 1 validates that predictor only against ground-truth scores for original Memento10k videos, not for outpainted videos. Because outpainting changes resolution (M3DDM downsamples to 256x256), aspect ratio, and adds synthetically generated borders, a predictor that responds to these artifacts rather than to human memorability would invalidate every reported improvement. The paper needs either a human memorability study on a subset of outpainted videos or a no-op control condition (e.g., letterboxing or resizing original videos without generative outpainting) to establish that the predictor is not biased by outpainting artifacts.","section":"Section 3.3"},{"comment":"The central finding that outpainting improves low-memorability videos and diminishes high-memorability videos is exactly the signature of regression to the mean when a noisy predictor is used to compute change from baseline. No significance tests, confidence intervals, or a control condition are reported, so the observed pattern may be an artifact of measurement noise rather than a genuine effect of outpainting. The authors should provide statistical tests and demonstrate that the pattern is not explained by regression to the mean, for example by including a comparison against videos that are re-encoded or resized without generative outpainting.","section":"Section 4, Figures 6-7 and Section 5"},{"comment":"The results section states that saliency-based outpainting \"did not yield any significant difference in the memorability scores as compared to using MOTIA without saliency,\" while the conclusions state that \"we found a significant increase in memorability when salient parts of the image were brought into the center of the frame.\" These statements are directly contradictory, and the stronger claim in Section 5 is not supported by any statistical evidence presented in Section 4. The contradiction must be resolved, and if the effect is not significant, the stronger claim should be removed.","section":"Section 4 vs. Section 5"}],"minor_comments":[{"comment":"There are several typos and inconsistent model names: \"basline\" in Section 2.2, \"divideded\" in Section 3.3, \"ide\" in Section 5, \"M3DMM\" for M3DDM in the conclusions, \"M2DDM\" in the Figure 8 caption, and \"inpainting models\" in the conclusions where \"outpainting models\" is meant.","section":"Throughout"},{"comment":"The axis labels are not defined; the caption should explicitly state which axis corresponds to the prediction model and which to ground truth, and clarify the sample of 100 videos.","section":"Figure 1"},{"comment":"The table contains thumbnail images that are not visible in the manuscript text; the authors should describe the selection criteria or provide the thumbnails in an appendix or supplementary material.","section":"Table 1"},{"comment":"The text says \"very recent work reported in [16] compares their performance\" but reference [16] is the M3DDM paper; it is unclear whether [16] or another source contains the comparison of three outpainting approaches, and this should be clarified.","section":"Section 2.4"},{"comment":"Reference [16] has the malformed arXiv number \"22309.02119\" (likely 2309.02119), and reference [19] has \"22009.01835\" (likely 2009.01835); these should be corrected.","section":"References"}],"recommendation":"reject","confidential_remarks":"This submission resembles a workshop paper more than a full journal article. The evaluation is built entirely on the authors' own memorability predictor, which is never validated on outpainted content, and the reported baseline-dependent pattern is consistent with regression to the mean. The internal contradiction between the results and conclusions regarding saliency would alone require major revision, but the more fundamental issue is that no human validation or control condition supports the central claim. Even with added statistical tests, the paper would need a human memorability experiment to be convincing, which is beyond the scope of a minor revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing you should know: this is the first paper to apply generative video outpainting to memorability, and that is a legitimate new application. The authors compare two recent outpainting models (M3DDM and MOTIA) on 100 Memento10k videos and add a saliency-guided prompting variant. Credit where due: the task framing is clear, the related work is fine, and the exploratory tone is honest. They openly note the results are inconsistent and that model performance is close. That said, the central claim is not supported by the evidence as presented. All memorability measurements come from a fine-tuned predictor developed by the authors' own group (Section 3.3), and they never validate that predictor on outpainted videos. Outpainting changes resolution, aspect ratio, and adds synthetic content; a predictor sensitive to those artifacts could report shifts that have nothing to do with human memorability. The paper's main finding - low-memorability videos improve, high-memorability videos decline - is exactly the signature of regression to the mean when you compute change from a noisy baseline. No no-op control (e.g., letterboxing without generative outpainting) is provided. There is also an internal inconsistency: Section 4 says the saliency-based outpainting did not yield any significant difference, while the Conclusions call it a significant increase. Minor but telling. The effect sizes are small and no significance tests are reported. The numbers in Table 2 (e.g., 0.668 to 0.675 for MOTIA frame 1) are within what you would expect from predictor noise. I do not think the authors are hiding anything - they explicitly say the significance is unclear - but the headline 'enhance memorability' overstates what the data show. Who is this for? Researchers working on memorability manipulation or video outpainting might find it a useful starting point, and the failure mode is instructive for evaluation methodology. But as a result, it is not reliable enough to cite for the claimed effect. My recommendation: this deserves a serious referee rather than a desk reject, because the idea is new and the paper could be salvaged with human-subject evaluation, significance tests, and a control condition. As is, I would not accept it. In peer review, I would lean reject with a clear path to revision: validate the predictor on outpainted videos, add human memorability judgments, report confidence intervals, and reconcile the saliency contradiction. If the authors do that, the claim might hold up - but right now it does not.","headline":"First application of generative outpainting to video memorability, but the central claim rests entirely on an unvalidated self-predictor; the paper is honest and exploratory but the evidence is not yet there.","tokens_in":779,"tokens_out":1092,"would_cite":false,"duration_ms":25998,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Generative outpainting can raise the predicted memorability of low-memorability short-form videos, while the same operation tends to lower memorability for videos that already score high.","keywords":["video memorability","generative outpainting","diffusion models","short-form video","Memento10k","saliency","MOTIA","M3DDM"],"falsifier":"Run a human recognition-memory test on the same 100 Memento10k videos and their outpainted versions: if outpainting improves human recall for low-scoring videos and hurts recall for high-scoring ones, the predictor-based claim is supported; if human recall does not track the predictor's deltas, the central claim fails.","tokens_in":7432,"feed_emoji":"🎬","tokens_out":9419,"duration_ms":79464,"temperature":0.7,"pith_summary":"Short-form video creators need content that viewers will remember, and memorability scores measure how likely a stranger is to recognise a clip on a later viewing. This paper asks whether a purely generative post-processing step—outpainting, or extending a video frame beyond its original borders with synthesised content—can move those scores. Testing two diffusion-based outpainting models on 100 videos from the Memento10k dataset, the authors report that outpainting improves predicted memorability for videos that start low and tends to reduce it for videos that start high. The better of the two models, MOTIA, showed only a marginal average advantage over M3DDM, which the paper itself describes as inconclusive. The authors conclude that outpainting can strengthen forgettable short-form videos, but it should not be applied blindly to already memorable ones.","feed_headline":"Extending video frames boosts forgettable clips, hurts memorable ones","feed_subtitle":"In 100 short videos, AI-generated border content shifted predicted memorability up for weak clips and down for strong ones.","key_machinery":"The machinery is a pair of generative outpainting models applied frame-by-frame to short videos: M3DDM, a masked 3D diffusion model that generates frames together for temporal consistency, and MOTIA, a two-phase model that learns input-specific patterns on the source video before outpainting with text guidance. The outcome variable is the change in memorability score computed by a fine-tuned Vision Transformer predictor, which is applied to original and outpainted versions of the same 100 Memento10k videos. To test saliency-aware outpainting, the authors compute per-frame saliency maps, divide them into four quadrants, identify the most salient quadrant, and prompt MOTIA with keywords describing the salient region so the added border content is drawn around it. The load-bearing operation is the delta between original and outpainted scores: a positive delta means outpainting helped, a negative delta means it hurt.","core_discovery":"The central finding is that generative outpainting generally improved the predicted memorability of videos with low initial memorability, while for videos with high original memorability outpainting tended to diminish it, and this pattern appeared with both M3DDM and MOTIA. The authors attribute the drop for memorable videos to added borders weakening saliency or spreading attention across a larger screen, and attribute MOTIA's marginal overall advantage to its input-specific adaptation and text conditioning. Saliency-guided outpainting produced a significant memorability increase when it moved the most salient content toward the center of the frame, but the gain was not uniform across videos. The work is presented as a new application of outpainting, evaluated not by human memory tests but by changes in an automated memorability predictor's scores.","pith_inferences":["A direct check the paper does not run is binning the 100 videos by initial memorability score and regressing the memorability delta on it; the reported pattern predicts a clear negative slope that a reader could verify from the published dumbbell charts.","Because the memorability predictor was trained on original videos, its scores on outpainted frames could carry a systematic bias from the new aspect ratio, resolution, and synthetic borders; a human recognition-memory study on the same 100 outpainted clips would show whether the predicted deltas reflect real memorability changes.","The effect plausibly extends to image outpainting and to vertical 9:16 outpainting for mobile-first platforms, neither tested here, which would make memorability enhancement a routine step in automated ad production.","Combining saliency-guided outpainting with previously demonstrated saliency-based cropping could both remove distracting borders and add context, potentially yielding larger memorability gains than either operation alone."],"forward_implications":["Creators of forgettable short-form ads or social clips can use outpainting as a post-production step to raise predicted memorability without reshooting or editing the original footage.","Already memorable videos should be left unmodified, because expanding their borders tends to dilute the focus that made them score high.","Between the two models tested, MOTIA is the preferable outpainting choice, though its advantage over M3DDM is marginal and inconsistent across individual videos.","Saliency-guided outpainting is a viable refinement only when it succeeds in recentering the salient content; applying it uniformly is not supported by the results.","The low-go-up, high-go-down pattern indicates that outpainting shifts visual attention rather than adding memorability in absolute terms."],"supporting_citations":[{"why":"Supplies the Memento10k dataset of 10,000 annotated short videos, from which the 100 test videos and their ground-truth memorability scores are drawn.","marker":"[12]"},{"why":"Provides the fine-tuned Vision Transformer approach used to compute memorability scores for each original and outpainted video.","marker":"[8]"},{"why":"Provides the memorability prediction model based on CLIP features and Bayesian ridge regression that the paper uses as its quantitative evaluation baseline.","marker":"[13]"},{"why":"Introduces M3DDM, the masked 3D diffusion model used as one of the two outpainting methods, including its 256-by-256 square output format.","marker":"[16]"},{"why":"Introduces MOTIA, the input-specific-adaptation outpainting model used as the second method, with text-conditioned pattern-aware outpainting.","marker":"[21]"},{"why":"Defines the benchmark task and dataset protocols that motivate the evaluation setup and the interpretation of memorability scores.","marker":"[2]"},{"why":"Shows that saliency-based cropping can improve video memorability, the prior result this paper extends by steering outpainting with saliency.","marker":"[15]"},{"why":"Supplies the saliency-detection technique used to compute saliency maps and quadrant analysis for saliency-guided outpainting.","marker":"[23]"}],"fun_headline_variants":["Outpainting boosts weak clips, dents strong ones","AI border extension shifts video memorability both ways","Generative outpainting lifts forgettable videos, drops memorable","Border expansion helps forgettable clips, hurts memorable ones","Saliency-guided outpainting nudges video memorability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes that the automated memorability predictor, trained on original videos, gives trustworthy scores for outpainted videos whose resolution, aspect ratio, and synthetic borders differ from anything it saw in training.","fun_headline_variants_meta":{"raw":{"variants":["Outpainting boosts weak clips, dents strong ones","AI border extension shifts video memorability both ways","Generative outpainting lifts forgettable videos, drops memorable","Border expansion helps forgettable clips, hurts memorable ones","Saliency-guided outpainting nudges video memorability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000161,"raw_usage":{"total_tokens":1177,"prompt_tokens":827,"completion_tokens":350,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":443,"completion_tokens_details":{"reasoning_tokens":271}},"tokens_in":443,"tokens_out":350,"duration_ms":3431,"temperature":1.0,"reasoning_tokens":271,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:24:33.145306+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a human recognition-memory test on the same 100 Memento10k videos and their outpainted versions: if outpainting improves human recall for low-scoring videos and hurts recall for high-scoring ones, the predictor-based claim is supported; if human recall does not track the predictor's deltas, the central claim fails.","supporting_citations":[{"cited_title":"Newman, C","cited_arxiv_id":null,"evidence_quote":"Supplies the Memento10k dataset of 10,000 annotated short videos, from which the 100 test videos and their ground-truth memorability scores are drawn."},{"cited_title":"Cummins, L","cited_arxiv_id":null,"evidence_quote":"Provides the fine-tuned Vision Transformer approach used to compute memorability scores for each original and outpainted video."},{"cited_title":"Predicting Media Memorability: Comparing Visual, Textual and Auditory Features","cited_arxiv_id":"2112.07969","evidence_quote":"Provides the memorability prediction model based on CLIP features and Bayesian ridge regression that the paper uses as its quantitative evaluation baseline."},{"cited_title":"Hierarchical Masked 3D Diffusion Model for Video Outpainting","cited_arxiv_id":"2309.02119","evidence_quote":"Introduces M3DDM, the masked 3D diffusion model used as one of the two outpainting methods, including its 256-by-256 square output format."},{"cited_title":"Be-Your-Outpainter: Mastering Video Outpainting through Input-Specific Adaptation","cited_arxiv_id":"2403.13745","evidence_quote":"Introduces MOTIA, the input-specific-adaptation outpainting model used as the second method, with text-conditioned pattern-aware outpainting."},{"cited_title":"Overview of The MediaEval 2022 Predicting Video Memorability Task","cited_arxiv_id":"2212.06516","evidence_quote":"Defines the benchmark task and dataset protocols that motivate the evaluation setup and the interpretation of memorability scores."},{"cited_title":"Mudgal, Q","cited_arxiv_id":null,"evidence_quote":"Shows that saliency-based cropping can improve video memorability, the prior result this paper extends by steering outpainting with saliency."},{"cited_title":"Ullah, M","cited_arxiv_id":null,"evidence_quote":"Supplies the saliency-detection technique used to compute saliency maps and quadrant analysis for saliency-guided outpainting."}],"review_version":1}