{"id":"d3580c83-49b7-426e-80a7-f2e54db73f27","arxiv_id":"2507.10016","paper_version":2,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A multi-agent audio-language model framework can automatically profile private attributes, such as age, health, and income, directly from general audio recordings.","lead":"This paper shows that multimodal AI models can infer sensitive personal details like age, health, and income just from audio clips. The authors built a new benchmark dataset, an AI agent system, and tested defenses, arguing that this is a real and understudied privacy risk.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"AP2-Com labels leak into the audio: clips were curated to be indicative of each attribute value before being assigned to profiles, so the 86.7% accuracy may reflect audio-event classification, not real-world covert profiling.","rationale":"I read the paper as an empirical feasibility study: the central claim is that a specific multi-agent pipeline can infer twelve sensitive attributes from audio with high accuracy, and that this constitutes a realistic privacy threat. The most load-bearing evidence for that claim is the AP2-Com benchmark, and the benchmark's construction creates a direct circularity: audio samples were retrieved because they were indicative of a target attribute value, then assigned to synthetic profiles with that value as ground truth. Asking a model to infer the attribute from such audio is closer to audio event classification with a commonsense label mapping than to covert attribute profiling of a real person. This is an internal design issue, not a disagreement with external consensus, and it affects the headline number, the human comparison, and the defense evaluations. I do not see this concern in the reader's weakest_assumption, which focuses on AP2-TV annotation subjectivity and the human-evaluation asymmetry; hence partial agreement. The paper does have independent support in the form of reproducible prompt templates, ablation studies, and a coherent pipeline, and I am not questioning the authors' honesty. However, the current manuscript does not justify the real-world attack claim. The concern is testable: a control experiment with unselected audio would settle whether the reported accuracy survives when the label no longer determines the audio content. Because the evidence could be repaired with such a control or with a naturalistic held-out dataset, I recommend CONDITIONAL rather than outright rejection: the manuscript should not be accepted in its current form as establishing the feasibility of audio private attribute profiling.","tokens_in":34239,"tokens_out":4225,"duration_ms":52918,"concrete_test":"Construct a control version of AP2-Com in which non-demographic attributes (HAB, OCC, SOP, PER, HEA, INC, etc.) are inferred from audio clips that were not selected to exemplify those attributes: randomly sample an equal number of clips from the same source corpora (WildDESED, AudioSet, Freesound, Sound Bible, disease-voice datasets) without expert curation, or take continuous ambient segments from the same recording sessions with no relevance filter. Run the full Gifts pipeline with the exact prompts of Appendix G and the metrics of Appendix E.1 on this control set. If average accuracy drops substantially, e.g., HAB/OCC/SOP falling from the reported 76-83 range toward the chance levels of Table 1, then the reported 86.7% is an artifact of label-selective curation and the real-world feasibility claim is unsupported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that off-the-shelf MLLMs can be assembled into an effective covert audio attribute profiling attack rests on the AP2-Com results, but AP2-Com is constructed in a way that leaks the target labels into the test audio. Section 5.1 states that for attributes beyond CommonVoice demographics, experts performed 'targeted retrieval of audio data from existing public resources ... identifying samples indicative of typical behaviors, dialogues, and activities associated with each entry,' and then 'randomly assigned the validated attribute values and their corresponding audio samples to speakers.' Appendix A adds that each AP2-Com individual is a stochastic composite and 'does not include any real humans.' Consequently, the model is not asked to infer an attribute from naturalistic audio of a person; it is asked to classify audio that was explicitly selected because it exemplified the target label, e.g., vacuum-cleaner sounds for HAB and occupation-relevant Sound Bible clips for OCC. High accuracy on this task largely reflects audio-event recognition plus the LLM's world knowledge about the curated scene, rather than profiling a real person from covertly captured audio. The same leakage inflates the human comparison in Table 4, which uses three AP2-Com individuals, and it also inflates the defense evaluations, which are measured on the same benchmark. AP2-TV does not fully remedy this because its clips are extracted from narrative contexts that often dramatize the annotated attribute. The paper therefore does not currently establish the headline threat.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces AP2, a two-part audio benchmark (AP2-Com, composed from curated public audio resources, and AP2-TV, extracted from recent TV dramas) annotated with twelve sensitive attributes, and proposes Gifts, a multi-agent framework that combines an audio-language model (Gemini1.5-Pro) with a large language model (Claude3.5-Sonnet) through guidance, inference, forensics, scrutinization, and consolidation phases. The authors report that Gifts infers all twelve private attributes on AP2-Com with an average accuracy of 86.7%, outperforming all baselines and surpassing a human participant pool by 22.5% in average accuracy while using about one quarter of the time. They also evaluate two defenses, in-context unlearning and phoneme-based noise jamming, and report that both reduce profiling accuracy. The paper claims to be the first study of MLLM-based audio private attribute profiling and positions AP2 and Gifts as resources for future research and defense development.","tokens_in":34471,"tokens_out":6007,"duration_ms":67700,"significance":"If the reported measurements were valid, the paper would demonstrate a serious and easily assembled privacy risk: off-the-shelf ALMs and LLMs can be combined into an agentic pipeline that infers sensitive attributes from audio more accurately and faster than human listeners. The work also makes concrete contributions: a substantial annotation effort across twelve attributes, a multi-phase agent design with ablations showing each phase's contribution, and an initial exploration of model-level and data-level defenses. However, the empirical foundation currently does not support the headline claims. The construction of AP2-Com leaks the target labels into the audio itself, the human evaluation is small and procedurally inconsistent, and the AP2-TV labels are not per-clip ground truth. These issues are load-bearing because the paper's central contribution is an empirical measurement of attack feasibility rather than a theoretical derivation. The dataset and framework are potentially useful, but the reported accuracy numbers and the human-competition claim need to be re-established on a non-leaky, properly validated benchmark.","major_comments":[{"comment":"AP2-Com is constructed so that the attribute label is a property of the audio clip rather than of the individual. The text states that experts performed 'targeted retrieval of audio data from existing public resources ... identifying samples indicative of typical behaviors, dialogues, and activities associated with each entry' and then 'randomly assigned the validated attribute values and their corresponding audio samples to speakers.' This means, for example, HAB clips are domestic event sounds (e.g., vacuum cleaner) and OCC clips are occupation-related Sound Bible sounds; the audio itself transparently signals the label. The 86.7% average accuracy on Table 2 therefore largely measures audio-event recognition plus commonsense mapping from the curated scene to the label, not covert profiling of a person's private attributes from naturalistic audio. Because Table 4's human comparison and Tables 5-6's defense evaluations use the same benchmark, the headline risk claim is inflated. A valid evaluation would require audio clips that were not selected because they exemplify the target label, or per-clip labels that are independent of the audio content that the model is asked to interpret.","section":"Section 5.1"},{"comment":"The human-vs-machine comparison is not a fair or statistically meaningful head-to-head. It uses only three individuals from AP2-Com, all subject to the label leakage described above. The appendix states that participants were forbidden from using LLMs ('the use of large language models ... to generate attribute descriptions was strictly prohibited'), while the main text says participants were 'permitted to use search engines or LLMs to retrieve relevant information.' In addition, the model pipeline receives automatically generated event descriptions and spoken-word transcriptions, whereas human participants listen only to raw audio. The 22.5% accuracy gap and the quarter-time claim in Table 4 are therefore not a valid basis for the conclusion that Gifts outperforms humans at covert audio profiling.","section":"Section 7.3 and Appendix E.2"},{"comment":"AP2-TV does not provide per-clip ground truth for the audio segments used in evaluation. Character-level labels are derived from the full series plus external discourse ('promotional materials, media coverage, and online forum discussions'), and Appendix D.2 explicitly instructs annotators to consider character development across the entire narrative arc. An individual audio clip extracted from one episode may therefore not match the aggregate character label, especially when the character changes over time. Moreover, using fictional characters annotated from external discourse as a proxy for real human sensitive attributes is an unvalidated assumption. Figures 3 and 4 report no numeric tables or error bars, so the actual per-attribute differences between Gifts and the baselines on AP2-TV cannot be assessed from the paper.","section":"Section 5.2 and Appendix D.2"},{"comment":"For the fuzzy attributes (ACC, PER, SOP, OCC, HAB), the paper uses Claude3.7-Sonnet as the automatic judge, which is from the same model family as the Claude3.5-Sonnet LLM agent used inside Gifts. No evidence is provided that this judge agrees with human raters or that it is neutral across the compared model families. This creates an unmeasured risk of systematic bias in the main quantitative claims. The paper should validate the judge (e.g., human-judge agreement on a sample) and report results with at least one alternative judge or a human-rated subset, especially because these fuzzy attributes contribute directly to the reported averages in Tables 2-4.","section":"Section 7.1, Evaluation Metrics"}],"minor_comments":[{"comment":"The statement that 'AP2-Com does not include any real humans' is contradicted by the use of CommonVoice recordings and other real-human speech datasets as base profiles; the stochastic composition of attributes does not remove the fact that the audio originates from real people. The claim should be rephrased to say that no individual profile maps to a single real person.","section":"Appendix A"},{"comment":"The three repeated runs yield very small reported variances, but no significance tests are provided; the paper should report per-attribute statistical tests (or at least confidence intervals) for the Gifts-versus-best-baseline differences.","section":"Section 7.1 and Table 2"},{"comment":"The AP2-TV results are presented only as figures without numeric values or error bars; adding a table with means and standard deviations would allow the claimed consistent superiority of Gifts to be verified.","section":"Figures 3 and 4"},{"comment":"The 'Time Spent' comparison mixes API latency for MLLMs with human self-reported total time; these quantities are not directly comparable, and no variance or per-participant breakdown is reported.","section":"Table 4"},{"comment":"The human participant pool is highly skewed (92% aged 21-30, mean 24.2 years), which limits the generalizability of the human baseline; the paper should acknowledge this and, ideally, recruit a more diverse sample.","section":"Appendix E.2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things before you read. Gifts, the ALM+LLM agentic framework, is genuinely well-built: LLM-guided inference, forensic question-asking, scrutinization, and consolidation are clearly described, with full prompts and ablations. The paper is transparent about design. However, the headline 86.7% average accuracy on AP2-Com is not measuring what you'd think. For most attributes beyond CommonVoice demographics, audio clips were deliberately retrieved because they exemplified the attribute value—vacuum-cleaner sounds for habit, occupation-relevant Sound Bible clips for occupation—then randomly assigned to synthetic speaker profiles. So the model is often doing audio-event classification plus a world-knowledge lookup, not covert profiling of a real person. The authors state this construction openly (Section 5.1, Appendix A), but don't draw the conclusion: the benchmark gives the model a shortcut.\n\nWhat's new and useful: AP2-TV is a clever idea—recent dramas premiering after Sept 2024 to avoid memorization, expert annotations from external discourse. Gifts is a solid contribution to the agentic-MLLM literature, and the defense analysis (in-context unlearning, phoneme-based jamming) is a reasonable first pass, inheriting the same benchmark caveats.\n\nSoft spots: (1) label leakage in AP2-Com, as above; (2) AP2-TV annotations are subjective and clips may leak via explicit dialogue—the claim that attributes aren't explicitly disclosed needs a check against actual transcripts; (3) human evaluation uses only three individuals from AP2-Com, and humans were barred from using LLMs while the model freely uses them, so the \"beats humans by 22.5%\" is not a clean comparison; (4) fuzzy scoring is done by Claude3.7-Sonnet, the same model family as the LLM agent, which could share biases.\n\nVerdict: the central threat claim is not established as stated, but the paper is honest, substantial, and worth a serious referee. The review should demand a decomposition of results by attribute type and a frank discussion of what the benchmark does and doesn't show. I'd bring it to reading group; the label-leakage issue is a useful methodology case study.","headline":"A well-engineered and transparent framework, but the main benchmark's label-curated audio inflates the headline accuracy and the central covert-profiling claim needs reframing.","tokens_in":35103,"tokens_out":5484,"would_cite":true,"duration_ms":59164,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Off-the-shelf audio and text models can be combined into an automated attack that infers a person's private attributes from sound alone, with average accuracy of 86.7 percent on the paper's benchmark.","keywords":["audio private attribute profiling","multimodal large language models","audio language models","attribute inference attack","privacy leakage","multi-agent framework","benchmark dataset","in-context unlearning"],"falsifier":"Run the Gifts pipeline on a fresh sample of consenting real speakers whose age, income, health, occupation, and other attributes are self-reported, using short audio clips with no speech content and no identifying context; if per-attribute accuracy falls toward the captioning-only baselines rather than the reported 86.7 percent average, the benchmark result was carried by its constructed labels rather than by acoustic inference.","tokens_in":33982,"feed_emoji":"🎧","tokens_out":7682,"duration_ms":75345,"temperature":0.7,"pith_summary":"The paper claims that current off-the-shelf multimodal models can already be assembled into a practical covert attack that infers a person's sensitive attributes from audio alone, a capability it calls audio private attribute profiling. To demonstrate this, it introduces AP2, a benchmark of eighty assembled real-world audio profiles and forty characters from recent TV dramas annotated with twelve sensitive attributes, and proposes Gifts, a two-agent framework in which a text LLM guides, interrogates, reviews, and consolidates the outputs of an audio-language model. On AP2-Com, Gifts reports an average accuracy of 86.7 across all twelve attributes, beating every baseline by 4.4 to 40.7 percentage points and beating human participants by 22.5 points while taking about one quarter of the time. The paper also tests two defenses, showing that model-level in-context unlearning and data-level phoneme-based noise jamming both substantially reduce the attack's accuracy, and argues the result is a current risk that warrants safety alignment of audio models.","feed_headline":"Two-model audio agent infers private traits at 86.7%","feed_subtitle":"An audio model paired with a text model beats humans on age, income, health, and more from sound alone.","key_machinery":"Gifts is the central mechanism: a multi-agent pipeline pairing an audio-language model (Gemini 1.5 Pro) with a general-purpose text LLM (Claude 3.5 Sonnet). It runs in five phases: the LLM generates attribute-specific guidance from audio captions and transcripts; the ALM makes a short inference; the LLM poses concise true/false/uncertain clue-validation questions that the ALM answers from the acoustic signal, avoiding the long-form reasoning that makes ALMs hallucinate; the LLM scrutinizes whether the inference is supported and, if not, orders a single re-inference with the previous answer negated; and the LLM consolidates evidence across multiple audio clips into the final profile. The AP2 dataset is the supporting object, built so that no individual maps to a real person in AP2-Com and so that AP2-TV's post-September-2024 dramas prevent memorization-based inference.","core_discovery":"The central claim is that audio private attribute profiling is feasible today with models available through public APIs. The paper shows that neither an audio-language model alone nor a text-only large language model fed with transcripts can profile reliably, but a hybrid agent can: Gifts combines an ALM with an LLM in a guidance–inference–forensics–scrutinization–consolidation loop, and reaches an average accuracy of 86.7 on the twelve attributes of AP2-Com, with perfect gender accuracy and near-perfect scores on health condition and marital status. On AP2-TV the same pattern holds, and the gains are largest for acoustic-driven attributes such as age, gender, accent, and health, indicating that the framework lets an LLM exploit acoustic features it cannot hear directly. The paper's stated conclusion is that the privacy risk from audio is not hypothetical, that current ALMs lack adequate safety alignment, and that defenses at both model and data level are needed.","pith_inferences":["If the reported accuracy transfers beyond constructed benchmarks, the practical exposure is larger than the paper states, because microphones in phones, smart speakers, and meeting software make audio capture continuous, passive, and hard to detect.","A natural follow-up experiment would strip the textual channel from the final consolidation phase to test whether the LLM is reasoning from acoustics or primarily from textual stereotypes in transcripts and captions; the paper leaves this decomposition untested.","The AP2-TV temporal-independence design is directly extensible: applying Gifts to later seasons or to sequels would reveal how much of the attack depends on memorized media knowledge versus genuinely transferable acoustic cues.","The defense results likely upper-bound the protection, since an adaptive adversary aware of noise jamming or in-context unlearning could adjust prompts or retrain, so the residual risk after deployed defenses is probably higher than the paper's post-defense numbers."],"forward_implications":["Anyone with API access to current audio and text models can reconstruct a twelve-attribute profile of a victim from passively captured or scraped audio, without the victim saying anything sensitive.","The reported 22.5-point accuracy advantage over humans, at roughly a quarter of the time, removes the cost barrier that previously confined audio profiling to trained analysts.","Because ALMs already leak attributes under naive prompts and do not exhibit refusal behavior, model providers cannot rely on current safety alignment to stop the attack.","The two defenses provide distinct mitigation paths: in-context unlearning, which providers or users can apply at inference time, and phoneme-based noise jamming, which individuals can deploy at the source.","The AP2 dataset and Gifts implementation enable future measurement of audio privacy leakage, giving defenders a benchmark to test alignment and countermeasures against."],"supporting_citations":[{"why":"Supplies the ALM evaluation baseline and documents the long-context hallucination that motivates the forensics phase.","marker":"[59]"},{"why":"Provides Qwen2-Audio, used for event captioning and as an ALM option inside Gifts.","marker":"[12]"},{"why":"Provides Gemini 1.5 Pro, the ALM agent that performs transcription, captioning, and inference inside Gifts.","marker":"[65]"},{"why":"Provides Claude 3.5 Sonnet, the LLM agent that does guidance, scrutiny, and consolidation.","marker":"[4]"},{"why":"Common Voice supplies AP2-Com's core speech clips with age, gender, and accent annotations.","marker":"[5]"},{"why":"WildDESED supplies domestic environment sounds used for habit annotations in AP2-Com.","marker":"[73]"},{"why":"AudioSet is a source for retrieved audio samples in AP2-Com.","marker":"[24]"},{"why":"Provides in-context unlearning, the model-level defense evaluated against Gifts.","marker":"[52]"},{"why":"Provides phoneme-based noise jamming, the data-level defense evaluated against Gifts.","marker":"[32]"},{"why":"Establishes the prior result that LLMs infer private attributes from text, positioning audio profiling as the unexplored extension.","marker":"[62]"}],"fun_headline_variants":["Two-model agent infers private traits from voice alone","Audio profiling: hybrid agent beats humans on sensitive traits","Voice-only privacy attack: MLLM agent reaches 86.7% accuracy","Agent combines audio and text models to predict age, health, more","Sound reveals secrets: new framework profiles private attributes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the benchmark labels measure real sensitive attributes: AP2-Com's people are randomly assembled composites of audio and expert-chosen attribute values with no real-world referent, and AP2-TV's labels are three annotators' judgments about fictional characters, informed by forums, promotional material, and media coverage, so the headline accuracy is only as valid as those labels as a proxy for genuine human attributes.","fun_headline_variants_meta":{"raw":{"variants":["Two-model agent infers private traits from voice alone","Audio profiling: hybrid agent beats humans on sensitive traits","Voice-only privacy attack: MLLM agent reaches 86.7% accuracy","Agent combines audio and text models to predict age, health, more","Sound reveals secrets: new framework profiles private attributes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000206,"raw_usage":{"total_tokens":1446,"prompt_tokens":1047,"completion_tokens":399,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":663,"completion_tokens_details":{"reasoning_tokens":315}},"tokens_in":663,"tokens_out":399,"duration_ms":5073,"temperature":1.0,"reasoning_tokens":315,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:40:45.506019+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the Gifts pipeline on a fresh sample of consenting real speakers whose age, income, health, occupation, and other attributes are self-reported, using short audio clips with no speech content and no identifying context; if per-attribute accuracy falls toward the captioning-only baselines rather than the reported 86.7 percent average, the benchmark result was carried by its constructed labels rather than by acoustic inference.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the ALM evaluation baseline and documents the long-context hallucination that motivates the forensics phase."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"WildDESED supplies domestic environment sounds used for habit annotations in AP2-Com."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides in-context unlearning, the model-level defense evaluated against Gifts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides phoneme-based noise jamming, the data-level defense evaluated against Gifts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the prior result that LLMs infer private attributes from text, positioning audio profiling as the unexplored extension."}],"review_version":1}