{"id":"6bda2ecd-6c18-44f2-b5a6-ae23b905eac5","arxiv_id":"2505.22084","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"MObyGaze provides 6072 expert-annotated segments across 43 hours of film, with a multimodal thesaurus, and shows that current vision and text models can detect objectification above chance, while audio models cannot.","lead":"This paper introduces MObyGaze, a dataset of 20 films densely annotated by two experts for on-screen objectification, with segment boundaries, objectification levels, and multimodal concept labels. It benchmarks vision, text, and audio models on classification and localization tasks, using the dataset to test whether these models can learn to detect objectification.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground truth validity is the load-bearing point: the two annotators designed and refined the thesaurus before annotating all films, so benchmark feasibility may only reflect their shared perspective unless independent annotators confirm the labels.","rationale":"I read the paper as a dataset contribution whose central contributions are (1) the first audiovisual objectification dataset and (2) evidence that classification and localization are feasible. The first is plausible from the related-work survey. The second depends on the labels being a trustworthy operationalization of the construct. The weakest point is not the modeling choices, which are honestly reported and compared to trivial baselines, but the chain from thesaurus to annotation to IAA. The two expert annotators are also the people who designed and refined the thesaurus; the refinement step in App. A.2.3 explicitly modifies concept definitions after seeing their own disagreements. This makes high agreement partly a product of mutual calibration. It also makes the low per-concept KS values hard to interpret: the authors say disagreements are mostly omissions, but no independent evidence is given. If the labels reflect a shared but idiosyncratic reading, the benchmark numbers do not transfer to other annotators or audiences, which weakens the feasibility claim and the broader social-purpose argument. The proposed test distinguishes these cases concretely. If independent judges agree with the MObyGaze labels, the concern is resolved and the conditional can be lifted; if not, the dataset remains useful as a resource but the central claim should be re-scoped. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":23145,"tokens_out":4523,"duration_ms":53436,"concrete_test":"Select a stratified sample of about 300 segments (15 per movie) spanning EN, HN, S and all concepts, or alternatively 10 one-minute clips per movie, and have 3-5 independent annotators with film studies or psychology training, who did not participate in thesaurus design, annotate them with the published thesaurus and tool. Compute per-concept KS and level-wise agreement (e.g., Cohen's kappa for S vs not-S) between each independent annotator and the MObyGaze labels. If mean S-vs-not-S kappa is below 0.4, or per-concept KS for Sound and Expression of emotion remains below 0.5, the ground truth is annotator-specific; the feasibility conclusions in Tables 2-5 and the audio negative result should then be re-scoped as conditional on the original annotators' perspective.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, that MObyGaze is a valid dataset of multimodal objectification and that current models can detect it, rests on the annotations being a measurement of the construct rather than of the annotators' private calibration. This is the least secure point. Per App. A.2.3, the same two experts who designed the thesaurus annotated two movies, then compared their disagreements and changed the definitions (expanding Activities, trimming Appearance) before annotating the remaining 18 movies. The ground-truth definition and the measured agreement are therefore co-produced by the same pair; the reported IAA of sigma = 0.74 validates consistency within that pair but not the construct's generalizability. Per-concept IAA in Table 9 is low precisely for concepts that matter to the audio and emotion channels: Sound KS = 0.22, Expression of emotion 0.37, Look and Voice 0.41. The paper attributes most concept disagreements to one annotator overlooking concepts, but that attribution is made after the fact by the same annotators and is not independently checked. Consequently the audio negative result (Table 10) is ambiguous: it may show that audio is not discriminative, or merely that audio labels are too noisy. The benchmark claims in Tables 2-5 are conditional on this unvalidated ground truth.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MObyGaze, a dataset of 20 feature-length films selected from MovieGraphs, densely annotated by two expert annotators for the high-level construct of objectification. The annotation protocol uses a custom thesaurus of 5 sub-constructs and 11 concepts spanning vision, text, and sound, with freely delimited segments, yielding 6072 segments over 43 hours. The authors formulate classification and localization tasks, propose several label-aggregation strategies for learning from a small number of annotators, and benchmark X-CLIP-based video models, Actionformer, DistilRoBERTa, LLaMA-2, and wav2vec2. They report that vision and text models outperform trivial baselines, while audio-only models do not, and claim this is the first audiovisual dataset annotated for objectification as a multimodal construct.","tokens_in":23383,"tokens_out":7882,"duration_ms":87697,"significance":"If the annotations are accepted as valid measurements of objectification, MObyGaze is a valuable new resource: it is the first dense expert annotation of a high-level multimodal interpretive construct over full-length films, with reproducible splits, released code, datasheet documentation, and Croissant metadata. The benchmark study is honest in reporting modest margins over baselines and a negative audio result, and the label-diversity analysis addresses a real problem. The significance is conditional on ground-truth validity: because the two annotators also designed and refined the thesaurus, the reported inter-annotator agreement validates consistency within one pair rather than generalizability of the construct, and several concept-level agreements are low (Sound KS = 0.22, Expression of emotion KS = 0.37). If independent validation is added, the dataset could support a new line of computational media analysis.","major_comments":[{"comment":"The load-bearing assumption is that the MObyGaze labels measure objectification rather than the private calibration of the two annotators. Per App. A.2.3, the same two experts designed the thesaurus, annotated two movies, compared their disagreements, and changed the definitions (expanding Activities, trimming Appearance) before annotating the remaining 18 films. The reported IAA sigma = 0.74 is a consistency measure for this pair, and the per-concept KS values in Table 9 are low for exactly the audio and emotion channels (Sound 0.22, Expression of emotion 0.37, Look and Voice 0.41). The explanation that most concept disagreements are due to one annotator overlooking concepts is made post hoc by the same annotators and is not independently checked. The benchmark claims in Tables 2-5 and the audio null result in Table 10 are therefore conditional on an unvalidated ground truth. The manuscript needs an external validation study, for example independent annotators applying the final thesaurus to a subset of films, or a comparison against an established objectification scale, or the claims must be substantially restricted.","section":"App. A.2.3 and Table 9"},{"comment":"The conclusion that audio alone is not discriminative for objectification is not supported, because the audio labels themselves have the lowest reliability. Sound has KS = 0.22 and Voice 0.41 in Table 9. The wav2vec2 result in Table 10 (AUC 0.589, std 0.089) is within one standard deviation of chance, but that could reflect label noise rather than the intrinsic non-discriminability of the audio channel. The 6% standalone-occurrence argument in Sec. 4.4 is not a substitute for a clean-label analysis. The authors should either re-run audio experiments on segments where both annotators agree on an audio concept, or explicitly state that the audio null result is confounded by label noise.","section":"Table 10 and Sec. 4.4"},{"comment":"The evaluation of label-diversity strategies is circular for Ragg2labv and favorable to Ragg2labm. The inverse variety loss trains on the closer of the two labels, namely L = min_l l(f(vid), l^l), and the Evar evaluation uses a winner-takes-all metric that compares the model output to its closest label. Thus Ragg2labv, and to a lesser extent Ragg2labm, are scored under the same closest-label principle used in training. The Evar columns in Table 4 therefore do not demonstrate that label separation is preferable to aggregation; they only show that a model trained and evaluated with closest-label matching achieves high closest-label agreement. The authors should report a single fixed evaluation metric, such as hard-label agreement on aggregated labels or agreement with both annotators, for all training strategies, and treat Evar only as a diagnostic of output diversity rather than as a performance metric.","section":"Sec. 4.2 and Table 4"},{"comment":"In Table 2, the WSL rows report standard deviations of 0.004 for AUC-ROC and F1 over five leave-4-movies-out folds (e.g., EN vs S: AUC 0.719 +/- 0.004, F1 0.517 +/- 0.004). Because the test sets are disjoint movies, such near-zero variance is implausible unless the metrics are computed on pooled predictions or there is leakage between folds. Please clarify how per-fold metrics and standard deviations are computed. This is important because the feasibility claim rests in part on the margin over trivial baselines in this table.","section":"Table 2"}],"minor_comments":[{"comment":"The header 'MoviGraphs' in Table 8 is a typo for 'MovieGraphs', and App. A.2.3 contains 'Actitivies' for 'Activities'.","section":"Table 8"},{"comment":"The cell for Llama-2-7B fine-tuned Recall uses a comma decimal separator ('0,431 (0,102)') unlike the rest of the table; please standardize the number format.","section":"Table 5"},{"comment":"Section 5 describes the annotation granularity as 'scene-level', but the dataset uses freely delimited segments that do not necessarily coincide with scenes; please clarify the relationship between segments and scenes.","section":"Sec. 5"},{"comment":"The datasheet's answer to 'Are there any errors, sources of noise, or redundancies?' is 'N/A', which is inconsistent with the paper's discussion of annotator overlook in App. A.2.5; please document known noise sources in the datasheet.","section":"Datasheet, A.1"},{"comment":"In App. A.4.1, 'HNUS' should be 'HN∪S' in the descriptions of the classification tasks.","section":"App. A.4.1"}],"recommendation":"major_revision","confidential_remarks":"The contribution is potentially strong, but the measurement-validity issue is central and will require either new data (independent annotations or an external validation study) or a substantial narrowing of the claims. The Evar evaluation issue is also a correctness concern for the label-diversity section, though it is secondary to the dataset-validity question. I would not reject on novelty grounds, but I would not accept without addressing these points."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [colleague],\n\nThe thing to know: MObyGaze is the first audiovisual dataset annotated for a high-level construct (objectification) with a structured thesaurus and free temporal boundaries. That claim holds up. The paper also ships code, Croissant metadata, datasheet, and benchmarks across vision, text, audio with fold splits. If you work on media bias or multimodal interpretation, this is a resource worth having on your radar.\n\nWhat it does well: The thesaurus is grounded in film studies and psychology, not invented ad hoc. The annotation procedure is described in detail, including the remediation step. They report IAA on both unitization and categorization, and they are transparent about per-concept agreement. The benchmark choices are sensible: leave-4-movies-out, three binary class definitions, trivial baselines, standard deviations. The label-diversity experiments are a real contribution, even if the results are modest. The vision and text results do support the modest claim of feasibility; audio does not, and they say so.\n\nWhere it gets soft: The load-bearing assumption is that two expert annotators, who designed and refined the thesaurus before annotating the full set, produce a measurement of objectification that transfers outside their shared perspective. The reported σ=0.74 validates consistency within that pair, not construct generalizability. Per-concept agreement is low precisely for the audio and emotion channels (Sound 0.22, Expression of emotion 0.37), which makes the audio negative result ambiguous. The paper attributes most concept disagreements to 'overlook' by one annotator, but that attribution comes from the same annotators and is not independently checked.\n\nA second soft spot is evaluation: the Evar metric (winner-takes-all, closest label) can flatter results. The Ehard numbers in Table 4 are the honest ones, and they are still decent. The Ragg2labv training and Evar evaluation share the 'closest label' logic, which nudges the comparison in its favor. None of this sinks the paper; it just means the benchmark numbers are upper-bound-ish.\n\nOverall: This is a carefully built dataset, honestly presented, with a plausible novelty claim. The main unresolved question is construct validity, and that is an empirical matter: independent annotators following the same thesaurus, or a transfer test to new annotators, would go a long way. I would send it to review, with a request for that validation before acceptance.\n\nRead it if you care about how to build datasets for interpretive constructs; skip if you need decisive benchmark numbers.","headline":"A genuinely new multimodal dataset for a high-level interpretive construct, with honest benchmarks—but the label validity rests on two co-designing annotators, so treat the feasibility numbers as conditional on construct validation.","tokens_in":23966,"tokens_out":2447,"would_cite":true,"duration_ms":29097,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces MObyGaze, presented as the first audiovisual dataset labeled for the high-level construct of objectification, and shows that current vision and language models can classify and localize objectification in films.","keywords":["objectification","movie dataset","multimodal annotation","film studies","video classification","temporal localization","inter-annotator agreement","gender representation bias"],"falsifier":"Have annotators who did not design the thesaurus independently re-annotate a sample of the films; if their agreement with the original labels on Sound and Expression of emotion is as low as the reported values of 0.22 and 0.37, then the measured construct is annotator-specific rather than a stable property of the films.","tokens_in":22949,"feed_emoji":"🎬","tokens_out":10797,"duration_ms":96506,"temperature":0.7,"pith_summary":"The paper builds a computational bridge to a concept from film studies: objectification, the ways a character is composed on screen to be perceived more as an object of desire than a subject of action. It claims that this high-level, multimodal construct can be pinned down in a structured thesaurus of eleven concepts spanning vision, speech, and sound, and that two expert annotators can densely label twenty films with free temporal boundaries and graded levels of objectification. The result is the MObyGaze dataset, 6,072 annotated segments over 43 hours of film, and the paper further claims that classifying and localizing objectification is accessible to current vision and text models, though audio alone is not. The payoff is a shared, quantified instrument for studying subtle gendered representation in cinema, usable by media scholars and machine learning researchers alike.","feed_headline":"Machines can detect objectification in films","feed_subtitle":"A new 6,072-segment expert dataset across 20 films makes subtle on-screen bias measurable.","key_machinery":"The load-bearing object is the thesaurus of objectification: five sub-constructs, grounded in film studies and social psychology, operationalized as eleven annotatable concepts spanning vision, speech, and sound. Because it decomposes the high-level construct into concrete, codeable dimensions, it makes the labeling reproducible enough that two experts reach an inter-annotator agreement of $\\sigma = 0.74$ on objectification levels, and it lets the paper formulate classification and localization as well-defined machine learning tasks. The complementary mechanism is the four-level rating scale, in which the hard-negative level captures segments that contain objectifying elements but are not perceived as objectifying; the paper shows these hard negatives are the strongest confusers for models.","core_discovery":"The central claim is that objectification in audiovisual storytelling is a measurable, multimodal construct that can be turned into a learning task. The authors define objectification through a thesaurus of five sub-constructs grounded in film studies and psychology, manifested through eleven concepts—type of shot, look, body, posture, clothing, appearance, activities, expression of emotion, voice, speech, and sound—spanning vision, text, and audio. Two expert annotators watched each of the twenty films in full, freely delimiting every timespan where at least one objectifying concept was present, and rated each segment on a four-level scale: easy negative, hard negative, sure objectification, and not sure. The resulting dataset contains 6,072 segments totaling 43 hours, with fine-grained concept tags and hard negatives that expose ambiguity. The paper then benchmarks vision, text, and audio models on classification and temporal localization, reporting that the visual and textual tasks are accessible to existing models while the audio-only task is not, and that a label-aggregation strategy merging the two annotators' labels into one per segment yields the best performance on the raw data.","pith_inferences":["Re-annotating a subset of films with fresh experts from a different cultural background would test whether the thesaurus transfers; the low agreement on Sound (KS 0.22) and Expression of emotion (KS 0.37) are the most likely fault lines.","The dataset could double as a fairness probe for vision models, since objectifying shots often frame body parts without heads and may expose systematic detection failures in pose estimators and person detectors.","The same thesaurus-based recipe could be extended to other interpretive constructs such as agency or vulnerability, turning qualitative film-studies notions into trainable datasets.","Because annotation stops at scene level, a natural follow-up is to test whether segment-level models can be composed to capture whole-film narrative tropes, the open direction the paper itself flags."],"forward_implications":["A shared, quantified definition of on-screen objectification becomes available to computational media studies, allowing subtle gender-representation patterns to be measured across large film corpora rather than only through close reading.","Classification of objectification from visual and textual modalities is already feasible with current models, so the bottleneck shifts to multimodal integration and better concept-level representations rather than dataset existence.","Hard negatives prove to be a decisive design choice: they are the main source of model confusion, so datasets for interpretive constructs should include them to separate presence of elements from perception of the construct.","Label aggregation into a single merged label per segment outperforms per-annotator separation on raw-data evaluation, supporting the use of expert-consensus labels even when the number of annotators is very small.","Audio-only objectification classification fails to beat trivial baselines, indicating that future multimodal models must weight audio contributions carefully."],"supporting_citations":[{"why":"Introduces the male gaze, the film-studies notion of objectification that the paper operationalizes.","marker":"Mulvey (1975)"},{"why":"Qualitative analysis of over 120 scenes that supplies the temporal and multimodal patterns the thesaurus systematizes.","marker":"Brey (2020)"},{"why":"Validated psychology questionnaire that grounds the gaze and appearance sub-constructs.","marker":"Calogero et al. (2011)"},{"why":"Shows clothing alone does not produce objectification, supporting the thesaurus's multi-concept design.","marker":"Bernard et al. (2019)"},{"why":"Precedent for defining a high-level construct with a codebook and for the importance of hard negatives.","marker":"Samory et al. (2021)"},{"why":"Supplies the inter-annotator agreement metric used to validate the level annotations.","marker":"Braylan et al. (2022)"},{"why":"Source dataset of 51 movies from which the 20 films are selected to preserve genre distribution.","marker":"Vicol et al. (2018)"},{"why":"The label-separation result that motivates the label-diversity training and evaluation strategies.","marker":"Wei et al. (2023)"},{"why":"ActionFormer, the model adapted as baseline for temporal localization and classification.","marker":"Zhang et al. (2022)"},{"why":"X-CLIP supplies the frozen video features used in the vision benchmarks.","marker":"Ni et al. (2022)"}],"fun_headline_variants":["MObyGaze: 6,072 expert-tagged segments for objectification AI","20 films, 6,072 segments: a new benchmark for objectification","Expert-annotated film dataset enables objectification detection","New AI task: measuring objectification in films from 20 movies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset is valid only if the thesaurus and the two expert annotators who designed it capture objectification as other audiences would perceive it, rather than encoding a private reading by its own authors.","fun_headline_variants_meta":{"raw":{"variants":["MObyGaze: 6,072 expert-tagged segments for objectification AI","20 films, 6,072 segments: a new benchmark for objectification","Expert-annotated film dataset enables objectification detection","New AI task: measuring objectification in films from 20 movies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000778,"raw_usage":{"total_tokens":3460,"prompt_tokens":984,"completion_tokens":2476,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":600,"completion_tokens_details":{"reasoning_tokens":2398}},"tokens_in":600,"tokens_out":2476,"duration_ms":17114,"temperature":1.0,"reasoning_tokens":2398,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:15:12.830546+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have annotators who did not design the thesaurus independently re-annotate a sample of the films; if their agreement with the original labels on Sound and Expression of emotion is as low as the reported values of 0.22 and 0.37, then the measured construct is annotator-specific rather than a stable property of the films.","supporting_citations":[{"cited_title":"Visual Pleasure and Narrative Cinema // Screen","cited_arxiv_id":null,"evidence_quote":"Introduces the male gaze, the film-studies notion of objectification that the paper operationalizes."},{"cited_title":"Le regard f \\'e minin-Une r \\'e volution \\`a l' \\'e cran","cited_arxiv_id":null,"evidence_quote":"Qualitative analysis of over 120 scenes that supplies the temporal and multimodal patterns the thesaurus systematizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Validated psychology questionnaire that grounds the gaze and appearance sub-constructs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows clothing alone does not produce objectification, supporting the thesaurus's multi-concept design."},{"cited_title":"Call me sexist, but","cited_arxiv_id":null,"evidence_quote":"Precedent for defining a high-level construct with a codebook and for the importance of hard negatives."},{"cited_title":"Measuring Annotator Agreement Generally across Complex Structured, Multi-object, and Free-text Annotation Tasks","cited_arxiv_id":"2212.09503","evidence_quote":"Supplies the inter-annotator agreement metric used to validate the level annotations."},{"cited_title":"MovieGraphs : Towards Understanding Human - Centric Situations from Videos // IEEE Conference on Computer Vision and Pattern Recognition ( CVPR )","cited_arxiv_id":null,"evidence_quote":"Source dataset of 51 movies from which the 20 films are selected to preserve genre distribution."},{"cited_title":"To Aggregate or Not ? Learning with Separate Noisy Labels // Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","cited_arxiv_id":null,"evidence_quote":"The label-separation result that motivates the label-diversity training and evaluation strategies."}],"review_version":1}