{"id":"f7aec484-2e03-4951-b510-b464069b68f2","arxiv_id":"2501.09707","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A new multimodal dataset from Goofus and Gallant comics, with normativity labels and LLM-generated, human-checked social principle labels, along with baseline classifiers.","lead":"Goofus and Gallant comic strips from 1995 to 2017 are collected into a multimodal dataset with normative and social-principle labels, plus baseline classifiers showing the labels can be learned. The paper argues this curated children's material is a better value alignment training source than larger, noisier crowdsourced story corpora.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Normative labels are assigned by character identity with no human validation, and the mixed-action example in §VII shows the binary scheme can conflate contradictory behaviors; the 'ideal dataset' claim hinges on unmeasured label noise.","rationale":"The reader's weakest assumption—that Section III.A's automatic labeling is unvalidated by humans—is exactly the load-bearing concern identified here. The paper's value rests on the accuracy of these labels, because the central claim is that the dataset is 'ideal' for training normative agents. The paper honestly mentions the camera/apology example but does not quantify how often such mixed-action cases occur, nor does it provide any human evaluation of the normative labels themselves. The human verification step in the principles dataset (91%/93% agreement on at least one principle, 61% on both) addresses a different set of labels and does not validate the normative labels. The proposed concrete test would directly measure label noise and ambiguity, and the outcome would determine whether the claim needs qualification or can stand as is. This is not a reason to reject the paper outright; the dataset may still be useful, but the 'ideal' claim requires stronger evidence. The verdict should remain CONDITIONAL pending the validation study. The paper deserves credit for releasing baselines and for its clear framing, but the missing validation is the weakest link in the argument.","tokens_in":11362,"tokens_out":4496,"duration_ms":51081,"concrete_test":"Select a random sample of 200 text descriptions from the GnG Normative dataset. Have three annotators, blinded to character identity, label each as 'normative', 'non-normative', or 'mixed/ambiguous' (with instructions to use the last category when the text contains both normative and non-normative elements or is unclear). Compute the majority-vote agreement with the automatic labels and report the proportion of 'mixed' cases. If majority agreement is below 90% or mixed cases exceed 5%, the automatic labeling is not a reliable gold standard and the paper should qualify its central claim or re-annotate the dataset.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—that the GnG corpus is an ideal dataset for training socially normative agents—depends on the accuracy of the normative/non-normative labels. Section III.A generates these labels by convention: every Gallant action is normative, every Goofus action is non-normative. No human verification of these labels is reported. The paper itself, in Section VII, identifies at least one class of examples that break this assumption: \"I think I broke your camera, Dad. I'm sorry\" is a mixed action—breaking the camera is non-normative while apologizing is normative—yet it receives a single binary label. This is not a rare edge case; it reveals that a comic panel can contain multiple action components with different normative statuses, and the character-based rule collapses them into one label. If a nontrivial fraction of the 1,387 texts and 819 images have such ambiguity, then (a) the dataset's clean signal is diluted, and (b) baseline accuracies (e.g., 94% text-only BERT) may partly reflect learning Goofus/Gallant character identity rather than socially generalizable norms, even after name removal. The 'ideal dataset' argument—that the comic is deliberately designed to teach norms—gives reason to expect mostly correct labels, but it does not establish that every panel is unambiguously one or the other. Without quantifying label noise, the central claim is unverified. This is the most load-bearing assumption because the entire dataset's value as a training resource depends on it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces the Goofus & Gallant Story Corpus, a multimodal dataset built from 1995-2017 Highlights magazine comic strips. The corpus has two components: the GnG Normative dataset (1,387 panel texts and 819 images) in which actions are labeled normative or non-normative based on whether the actor is Gallant or Goofus; and the GnG Principles dataset (819 examples) in which each action is assigned two social principles from a 27-value taxonomy derived from Kiesel et al. (2022), generated by GPT-4o with a limited human verification step. The paper reports baselines for two tasks: normativity classification (binary) and principles classification (multi-label), with best accuracies of 76.77% for the multimodal normativity model, 94.05% for text-only BERT normativity, and 74.84% relaxed accuracy for principles classification. The central claim is that, because the comics were deliberately designed to teach young children social principles, this corpus is an ideal dataset for training socially normative agents.","tokens_in":11659,"tokens_out":6127,"duration_ms":60097,"significance":"If the labels are valid, the corpus provides a compact, carefully designed, multimodal alternative to large crowdsourced moral datasets such as Moral Stories and Social Chemistry 101. The paper's release of the data and its transparent baselines are useful, and the multimodal pairing is an asset for future value alignment research. However, the significance is conditional: the strong performance of the baselines is unsurprising if the models are learning the comic's internal character convention, and the dataset's value depends on label validity and generalizability, both of which remain unverified. The paper also offers a useful contrast to sentiment analysis, showing that normativity is not reducible to sentiment, though this is a negative result.","major_comments":[{"comment":"The normative labels are assigned by character identity (Gallant = normative, Goofus = non-normative) and no human validation of these labels is reported. The paper itself identifies a counterexample in Section VII—'I think I broke your camera, Dad. I'm sorry'—where the breaking and the apology have different normative valences, so a single binary label is not well-defined for such composed actions. Because the proportion of ambiguous panels is not quantified, the 'ideal dataset' claim rests on an unverified premise. The authors should validate a sample of the normative labels with independent annotators, quantify ambiguous or mixed panels, and rerun the classifiers on a cleaned subset to measure the impact of label noise.","section":"Section III.A and Section VII"},{"comment":"The GnG Principles labels are generated by GPT-4o and the human verification is constrained: reviewers select among the model's proposed principles or choose 'None,' with 93% agreement on at least one principle and 61% on both in a 100-sample check. This design cannot identify systematic omissions in the candidate set, and the principles classifier is trained and evaluated on these machine-generated labels. The relaxed evaluation rule (correct if the prediction matches either of the two machine-generated true principles) further inflates apparent accuracy. To support the principles dataset, the authors should create a fully human-annotated test set in which annotators choose from the full 27-value taxonomy without LLM suggestions, report inter-annotator agreement, and evaluate the model against that independent gold standard.","section":"Section III.B.2.c and Section V.B"},{"comment":"The 94.05% text-only BERT accuracy on the GnG test set does not demonstrate that the model learned generalizable social norms; it may reflect dataset-specific stylistic differences between Goofus and Gallant panels even after name removal. The sentiment-analysis comparison only shows that normativity is not sentiment; it does not test generalization. The authors should evaluate the normativity classifier on external normative corpora (e.g., Moral Stories or ETHICS) or on a newly collected set of Goofus & Gallant panels with human-provided normative labels, and report error patterns on the ambiguous and mixed examples discussed in Section VII.","section":"Table III and Section VI.A"},{"comment":"The claims that the corpus is 'an ideal dataset' and contains 'rich knowledge about socio-cultural norms and values' go beyond the evidence presented. The corpus reflects the values of a single children's magazine over 1995-2017, and the paper does not analyze the coverage, cultural specificity, or potential biases of those values. The authors should either moderate the claims to describe the corpus as a useful, deliberately designed resource, or add an analysis comparing the value distribution in the corpus with other normative datasets and discussing its limitations as a general value-alignment resource.","section":"Abstract and Section IV"}],"minor_comments":[{"comment":"The word 'datasat' in the contributions summary should be 'dataset'.","section":"Section I"},{"comment":"The word 'assgined' should be 'assigned'.","section":"Section III.B.2.b"},{"comment":"Clarify whether the GnG Principles 'Text only' row refers to action descriptions, scene descriptions, or both; the row label is ambiguous.","section":"Table I"},{"comment":"Report how the majority-vote final label was chosen when no single principle received a majority, and discuss the relationship between the Fleiss kappa of 0.54 and the reliability of the final labels.","section":"Section III.B.1"},{"comment":"The compliance information fed to the principles classifier is derived from the character-identity labels; state explicitly that this makes the principles task depend on the unvalidated normativity labels.","section":"Section VI.B"},{"comment":"Add a data availability statement with the URL or access conditions for the dataset.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is a dataset contribution with moderate novelty. The main risk is overclaiming in the abstract and Section IV; the authors should be encouraged to add human validation and cross-dataset evaluation. The release of the multimodal data is the strongest part. The manuscript fits a journal on AI/ML ethics and datasets, but the technical depth of the baseline experiments is thin."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the GnG corpus is a real addition to the value-alignment data zoo, and the paper is honest about its size and coarseness. The main thing to know is that the normative labels are assigned by character identity, not by human judgment, and the paper never validates them; the 'ideal dataset' pitch is therefore stronger than the evidence.\n\nWhat's new: no one has used Goofus & Gallant for value alignment before. The combination of child-oriented comics, paired text and image, binary normativity labels, and a 27-principle taxonomy is genuinely novel. The authors also do useful baseline work, showing that text-only BERT hits 94% accuracy on normativity, that image-text fusion helps a bit, and that sentiment analysis does not capture normativity. Those baselines are a reasonable demonstration that the data carries a learnable signal.\n\nWhat it does well: the paper is transparent about the dataset's limits—small size, coarse labels, possible mixed actions. It reports inter-annotator agreement for the crowdsourced principle labeling (Fleiss' kappa 0.54) and provides a human check on GPT-4o's principle labels (93% at least one, 61% both). That is more honesty than many dataset papers show.\n\nSoft spots, in proportion. The biggest one is the normative labels. Defining Gallant as normative and Goofus as non-normative is a clever automatic labeling trick, but it inherits every ambiguity in the comic. The apology-after-breaking-the-camera example in Section VII is a real counterexample to a clean binary split, and the paper does not quantify how often such mixed cases occur. The high baseline accuracy could partly reflect stylistic differences between the two characters rather than generalizable norm learning, even with names replaced. A quick human validation study on a few hundred labels would settle this. Second, the principles dataset is generated by GPT-4o and verified by humans only in a select-one-of-the-suggested-principles manner, so the verification is weaker than it looks. The 61% both-principles agreement is a more honest number than the 93%. Third, no external benchmark shows that training on GnG transfers to held-out moral scenarios from ETHICS or Social Chemistry; without that, the 'ideal dataset' claim is an assertion, not a result.\n\nWho this is for: researchers building or benchmarking value-alignment datasets, especially multimodal ones. It is a small resource, but it is deliberately designed and may be useful as a complement to large noisy corpora. I would send it to review, with a request that the authors add a label-validation study and a transfer experiment.\n\nRecommendation: deserves a serious referee. With those additions it could be a solid dataset contribution; without them, it is still a reasonable short paper, but the central claim needs tempering.","headline":"A small, honest multimodal dataset paper that is a plausible new resource, but the 'ideal dataset' claim outruns the evidence.","tokens_in":12183,"tokens_out":2175,"would_cite":false,"duration_ms":23355,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that a curated children's comic-strip corpus, where one character always acts well and the other always acts badly, gives machine learning models a grounded, low-noise source of social norms for value alignment.","keywords":["value alignment","social norms","normative behavior","multimodal dataset","children's comic strip","machine ethics","text classification","principle taxonomy"],"falsifier":"Take a sample of strips and ask a diverse group of adult readers to label each depicted action normative or non-normative; if a substantial share disagree with the Goofus/Gallant assignment, or if a model trained only on this corpus predicts human labels near chance on everyday scenarios never shown in the comic, the corpus's value-alignment claim is undercut.","tokens_in":11190,"feed_emoji":"⚖️","tokens_out":9228,"duration_ms":94731,"temperature":0.7,"pith_summary":"The paper introduces the Goofus & Gallant Story Corpus, a multimodal dataset built from decades of a children's comic strip designed to teach social lessons. Its core idea is that deliberately pedagogical stories are a better source for value alignment than large crowdsourced or scraped corpora, because every strip already encodes a clear normative message. In the corpus, actions performed by Gallant are labeled normative and actions performed by Goofus are labeled non-normative, giving automatic ground truth; a second layer labels each action with up to two principles from a 27-value taxonomy. The paper reports that transformer models trained on this corpus learn to classify normative behavior and to predict social principles. If valid, this would make value-alignment training more tractable and less reliant on noisy internet data.","feed_headline":"Comic-strip corpus trains AI to tell good from bad behavior","feed_subtitle":"Curated Goofus & Gallant strips give machine-learning models paired examples of normative and non-normative actions.","key_machinery":"The load-bearing object is the Goofus & Gallant comic strip itself, a long-running children's feature whose two recurring characters perform opposite behaviors in the same situation. That structure turns into an automatic labeling machine: every action by the 'good' character becomes a normative example and every action by the 'bad' character becomes a non-normative example, with no manual moral judgment needed. On top of this, a 27-value taxonomy of social principles, derived by running a pretrained value model over the corpus texts, gives each example up to two principle labels, generated by an LLM and verified by human annotators. The baselines use transformer encoders for text and images, including a dual encoder that fuses both modalities, to demonstrate that the labels are learnable.","core_discovery":"The central claim is that the corpus—not any single model—is the contribution: it provides a compact, deliberately designed collection of normative and non-normative behaviors, and this design makes it an appropriate training set for socially normative agents. The paper supports the claim in two ways. First, normativity labels are constructed from the comic's own structure, so no crowd worker has to decide what is good or bad. Second, the principle labels are anchored to a fixed 27-value taxonomy derived from a broader social-science value framework, with the final labels produced by repeated language-model prompting and checked by human reviewers. The baseline experiments are offered as evidence that the value information is actually learnable: models classify text-only normativity with high accuracy, and multimodal input helps; principle prediction works under a relaxed evaluation that accepts either of the two true principles.","pith_inferences":["The comic's paired-panel structure is a natural fit for contrastive learning—explicitly training a model to push normative and non-normative descriptions apart—but the paper's baselines treat examples independently; exploiting the pairs could improve sample efficiency.","The normative ground truth is culturally specific: a children's magazine's etiquette from 1995 to 2017. A deployer should test transfer across countries, socioeconomic groups, and time before treating these labels as universal values.","Because the 27-value taxonomy was selected by running a pretrained value model on the corpus, model bias could enter before human annotation; human verification checks agreement with predicted labels, not the adequacy of the taxonomy itself.","A natural next step is to use the trained normativity classifier as a reward signal for reinforcement learning agents, turning the dataset from a benchmark into an actual alignment mechanism."],"forward_implications":["A modest, purpose-built corpus can substitute for much larger crowdsourced or scraped datasets when the goal is learning everyday normative judgments.","The 27-principle taxonomy makes value learning a bounded classification problem, which is more tractable for downstream agents than open-ended moral reasoning.","Since both text and image versions are released, multimodal value-aligned systems can be trained and evaluated against a common benchmark.","The corpus offers a testbed for whether models trained on children's stories generalize to adult everyday situations, and the paper invites further work on augmenting existing value-alignment systems with this data."],"supporting_citations":[{"why":"Supplies the social-science value taxonomy from which the corpus's 27 principle labels are drawn.","marker":"[12]"},{"why":"Gives the crowdsourced story-corpus approach the paper contrasts with curated children's stories.","marker":"[9]"},{"why":"Large crowdsourced ethics corpus used as a comparison for scale and labeling approach.","marker":"[10]"},{"why":"Community-sourced normative anecdotes that illustrate the limitations of forum-derived data.","marker":"[19]"},{"why":"Large social-norm corpus that the paper positions its smaller curated dataset against.","marker":"[20]"},{"why":"Unified norm bank built from internet corpora, cited as potentially biased relative to curated stories.","marker":"[21]"},{"why":"Provides the inter-annotator agreement measure whose moderate result motivates the switch to LLM annotation.","marker":"[22]"},{"why":"Language model used to generate the principle labels that are then human-verified.","marker":"[23]"},{"why":"Vision transformer used for image-only and multimodal normativity baselines.","marker":"[24]"},{"why":"Language model used for text-only normativity and principle classification baselines.","marker":"[25]"}],"fun_headline_variants":["Kids' comic strips teach AI right from wrong","Goofus & Gallant dataset helps AI learn social norms","New corpus uses classic comics to align AI values","Comic-strip pairs train AI on good vs bad behavior","Classic comic strips become AI training for ethics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole dataset rests on the convention that every action by Gallant is normative and every action by Goofus is non-normative, with no independent check that the comic's intended lessons match the norms a deployed AI should actually follow.","fun_headline_variants_meta":{"raw":{"variants":["Kids' comic strips teach AI right from wrong","Goofus & Gallant dataset helps AI learn social norms","New corpus uses classic comics to align AI values","Comic-strip pairs train AI on good vs bad behavior","Classic comic strips become AI training for ethics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000153,"raw_usage":{"total_tokens":1160,"prompt_tokens":848,"completion_tokens":312,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":464,"completion_tokens_details":{"reasoning_tokens":235}},"tokens_in":464,"tokens_out":312,"duration_ms":4213,"temperature":1.0,"reasoning_tokens":235,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:43:51.751535+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a sample of strips and ask a diverse group of adult readers to label each depicted action normative or non-normative; if a substantial share disagree with the Goofus/Gallant assignment, or if a model trained only on this corpus predicts human labels near chance on everyday scenarios never shown in the comic, the corpus's value-alignment claim is undercut.","supporting_citations":[{"cited_title":"Identifying the human values behind arguments,","cited_arxiv_id":null,"evidence_quote":"Supplies the social-science value taxonomy from which the corpus's 27 principle labels are drawn."},{"cited_title":"Moral stories: Situated reasoning about norms, intents, actions, and their consequences,","cited_arxiv_id":null,"evidence_quote":"Gives the crowdsourced story-corpus approach the paper contrasts with curated children's stories."},{"cited_title":"Scruples: A Corpus of Community Ethical Judgments on 32,000 Real-Life Anecdotes","cited_arxiv_id":"2008.09094","evidence_quote":"Community-sourced normative anecdotes that illustrate the limitations of forum-derived data."},{"cited_title":"Delphi: Towards machine ethics and norms,","cited_arxiv_id":null,"evidence_quote":"Unified norm bank built from internet corpora, cited as potentially biased relative to curated stories."},{"cited_title":"Gpt-4: Technical report,","cited_arxiv_id":null,"evidence_quote":"Language model used to generate the principle labels that are then human-verified."}],"review_version":1}