{"id":"81a9a047-eed7-4c2a-a56d-cc5d490b6b91","arxiv_id":"2606.25980","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"FoleySet is a Creative Commons dataset of 10,000 audio clips with two-level human annotations for Foley sound classification, retrieval, and generation.","lead":"The paper releases FoleySet, a new public dataset of 10,000 human-annotated Foley sound clips organized under a two-level taxonomy. A smart generalist might read it to understand available resources for training AI systems that generate or classify sound effects for film and video.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Annotation process, taxonomy definition, and quality metrics are not described, leaving 'high-quality' and 'standardized' claims unsupported","rationale":"The reader's weakest_assumption directly identifies the usefulness of the dataset as a standardized resource; that usefulness is load-bearing on annotation quality, which is not evidenced in the supplied abstract. This matches the reader's UNVERDICTED stance and would shift the verdict to CONDITIONAL once the missing methodological details are examined. No other internal inconsistency is visible from the given text.","tokens_in":1617,"tokens_out":342,"duration_ms":21140,"concrete_test":"Locate the sections on taxonomy design and human annotation; verify whether they report (a) the explicit two-level hierarchy with example labels, (b) annotator count and instructions, and (c) any agreement metric (e.g., Fleiss' kappa > 0.6). If these elements are absent or agreement is unreported, the quality claim remains unverified.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that FoleySet supplies a high-quality, standardized, CC-licensed resource of 10k clips with a two-level taxonomy that addresses the scarcity of training data for classification/retrieval/generation. This requires that the annotations are reliable and the taxonomy is well-specified. The abstract (and any full text that omits these details) provides zero information on taxonomy construction, annotation guidelines, number of annotators, collection protocol, clip durations/diversity, or reliability statistics such as inter-annotator agreement. Without these, it is impossible to evaluate whether the dataset actually meets the quality threshold asserted.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper presents FoleySet, a publicly available Creative Commons-licensed dataset of 10,000 audio clips annotated with a two-level Foley taxonomy, intended to address the scarcity of high-quality annotated resources for data-driven Foley classification, retrieval, and generation tasks.","tokens_in":1733,"tokens_out":278,"duration_ms":13141,"significance":"If the annotations prove reliable and the taxonomy well-specified, the dataset could provide a useful standardized resource for Foley-related machine learning tasks. The contribution is a direct data release rather than a methodological advance, so its impact hinges entirely on the documented quality and reproducibility of the annotations.","major_comments":[{"comment":"Abstract: the central claims that the dataset is 'high-quality' and 'standardized' are unsupported because the manuscript provides no description of the annotation process, taxonomy construction, number of annotators, collection protocol, clip selection criteria, or any reliability statistics such as inter-annotator agreement.","section":"Abstract"},{"comment":"No section: without details on how the two-level taxonomy was defined or validated, it is impossible to assess whether the annotations meet the quality threshold asserted for training classification/retrieval/generation models.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. The comments correctly identify that the current manuscript lacks the necessary documentation on the annotation process and taxonomy to substantiate claims of high quality and standardization. We will address these points through revisions.","responses":[{"response":"We agree that the manuscript as submitted does not include these details, which are required to support the stated claims. In the revised version we will add a dedicated Methods section describing the two-level taxonomy construction, annotation protocol, number of annotators and their qualifications, clip sourcing and selection criteria, and any computed reliability statistics including inter-annotator agreement.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central claims that the dataset is 'high-quality' and 'standardized' are unsupported because the manuscript provides no description of the annotation process, taxonomy construction, number of annotators, collection protocol, clip selection criteria, or any reliability statistics such as inter-annotator agreement."},{"response":"We accept this assessment. The revised manuscript will contain an explicit section on taxonomy definition, validation procedures, and supporting evidence so that readers can evaluate suitability for downstream tasks.","revision_made":"yes","referee_comment":"[—] No section: without details on how the two-level taxonomy was defined or validated, it is impossible to assess whether the annotations meet the quality threshold asserted for training classification/retrieval/generation models."}],"tokens_in":1164,"tokens_out":318,"duration_ms":13780,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main contribution is the release of FoleySet: 10,000 audio clips tagged with a two-level Foley taxonomy and put out under Creative Commons. That specific combination and the public licensing are new relative to the stated motivation around data scarcity for Foley tasks.\n\nIt does the basic job of identifying a practical gap in audiovisual post-production data and offering a resource aimed at classification, retrieval, and generation work. The CC license is a clear positive for anyone who might want to use it.\n\nThe soft spot is straightforward and central. The abstract asserts a high-quality, standardized dataset but contains no information on taxonomy construction, annotation guidelines, number of annotators, collection protocol, clip diversity, or any reliability measure such as inter-annotator agreement. Without those, the quality claims cannot be evaluated. Dataset papers stand or fall on exactly this documentation.\n\nIf the full manuscript adds those details and shows they were done carefully, the release could be useful. As presented, the evidence does not yet support the description.\n\nThis is for people working on audio ML for media sound effects who need labeled Foley data. A reader already in that niche might get value once the annotation process is documented. It deserves a serious referee to examine the collection and labeling methodology, because that is where the actual contribution lives or dies.","headline":"FoleySet releases 10k Foley clips with a two-level taxonomy, but the abstract supplies zero details on annotation or validation.","tokens_in":2175,"tokens_out":339,"would_cite":false,"duration_ms":18630,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"FoleySet supplies 10,000 human-annotated clips under a two-level taxonomy to support Foley classification, retrieval, and generation.","keywords":["Foley sound","audio dataset","sound effects","taxonomy","classification","retrieval","generation"],"falsifier":"A controlled experiment in which classifiers or generators trained on FoleySet show no accuracy or quality gain over models trained only on prior smaller collections would indicate the dataset does not deliver the expected benefit.","tokens_in":2512,"feed_emoji":"🔊","tokens_out":537,"duration_ms":18210,"temperature":0.7,"pith_summary":"The paper identifies the scarcity of high-quality annotated data for Foley sounds, which are recreated effects for on-screen actions in audiovisual production. It releases FoleySet, a collection of 10,000 clips labeled with a two-level taxonomy and offered under a Creative Commons license. The resource targets data-driven work on classifying existing sounds, retrieving matches for video, and generating new ones. A reader would care because professional Foley recording is labor-intensive, so better training data could enable scalable alternatives.","feed_headline":"Dataset releases 10,000 Foley clips with two-level labels","feed_subtitle":"Standardized resource targets classification, retrieval, and generation of action-linked sound effects","key_machinery":"The two-level Foley taxonomy applied to annotate the 10,000 audio clips of human actions and prop interactions.","core_discovery":"We present FoleySet, a publicly available Foley dataset of 10,000 audio clips annotated with a two-level Foley taxonomy. This dataset provides a standardized, Creative Commons-licensed resource for data-driven Foley classification, retrieval, and generation.","pith_inferences":["The two-level structure could support hierarchical training where coarse labels regularize fine-grained predictions.","Integration with video datasets might enable joint audio-visual models for automatic Foley creation.","Widespread adoption could create shared evaluation protocols similar to those used in speech recognition."],"forward_implications":["Models for Foley classification can be trained directly on the labeled clips.","Retrieval systems can match sounds to specific on-screen actions using the taxonomy.","Generative models can learn to produce new Foley audio from the annotated examples.","Research groups gain a common, licensed benchmark for comparing methods."],"fun_headline_variants":["FoleySet: 10,000 clips with two-level taxonomy","10k Foley clips with two-level annotations","10,000 Foley clips in two-level taxonomy dataset","FoleySet dataset features 10k two-level labeled clips"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"High-quality annotated Foley datasets for training remain scarce and the new dataset with its two-level taxonomy will serve as a useful standardized resource.","fun_headline_variants_meta":{"raw":{"variants":["FoleySet: 10,000 clips with two-level taxonomy","10k Foley clips with two-level annotations","10,000 Foley clips in two-level taxonomy dataset","FoleySet dataset features 10k two-level labeled clips"]},"model":"grok-4.3","cost_usd":0.006093,"raw_usage":{"total_tokens":2730,"prompt_tokens":532,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":60928000,"prompt_tokens_details":{"text_tokens":532,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2134,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":532,"tokens_out":64,"duration_ms":13242,"temperature":1.0,"reasoning_tokens":2134,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-25T19:54:34.946891+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled experiment in which classifiers or generators trained on FoleySet show no accuracy or quality gain over models trained only on prior smaller collections would indicate the dataset does not deliver the expected benefit.","supporting_citations":[],"review_version":1}