Pith. sign in

REVIEW 4 major objections 4 minor 73 references

CallScreenBench: Benchmarking On-Device Models as Phone Secretaries

T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The most fluent phone-secretary AIs are the most likely to relay a scam caller's callback number, and apparent triage skill is largely a suspicion artifact.

desk verdict A transparent first benchmark for on-device phone secretaries with real methodological ideas, but its headline safety claim about the cloud model rests on a single unvalidated judge. read the letter →

arxiv 2608.01033 v1 pith:FG7VOXM2 submitted 2026-08-02 cs.CR cs.AI

classification cs.CRcs.AI
keywords callsecretarydelegatedagenton-deviceLLMscreeningbenchmarkguardednessLLM-as-judgeevaluationfloors
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

CallScreenBench evaluates small on-device language models (0.6–4B parameters, 4-bit) in a setting no prior benchmark scores: answering an unknown inbound call on the owner's behalf, with no task to complete, no cooperative caller, and no oracle. The paper's claim is that the decisive question—would the owner endorse this handling?—cannot be compressed into one number, so it reports five quality dimensions each billed by an opposing counter-metric, plus three guardedness counts. Its two load-bearing findings are that quality rises with capability while triage does not, and that the apparent triage gradient is a suspicion-rate artifact: after subtracting the always-suspicious floor via Youden's J, none of the 15 model pairs separate at the pre-registered k=1 operating point. The most striking result is a decoupling at the top of the range: the most capable, most fluent answerer in the matrix is the most willing to write the scam caller's supplied callback number into the note for the owner. The paper matters because if these claims hold, phone-secretary AIs must be evaluated with paired counter-metrics and empirical floors, never by a single averaged score.

What carries the argument

The load-bearing instrument is the paired score-and-bill design. Each of the five quality scores (triage, message fidelity, representation, caller experience, interaction) is printed beside a counter-metric from the opposing family—recall beside fabrication, elicitation beside verbosity, scam true-positive rate beside legitimate false-positive rate at the same k—and no score is averaged into a composite. The triage argument runs through Youden's index $J=\mathrm{TPR}-\mathrm{FPR}$, computed with an empirically run 'always-suspicious' scripted agent as the floor; all between-model separation claims are made on paired-difference bootstraps over shared scenario resamples. A secondary guardednes

What would settle it

Run the paper's preregistered stratified human audit on the 48-scenario core: have human raters label each transcript for owner endorsement and the three guardedness harms. If human labels agree with the automated judge below the paper's $\kappa<0.6$ threshold, or if humans do not reproduce the finding that the most fluent answerers relay attacker-supplied callback numbers more often, the headline decoupling is judge artifact rather than model property.

Watch

Extended reading notes

Core claim

The central discovery is that capability, triage, and guardedness are three separate axes in a delegated phone-answering agent, and conflating them produces wrong conclusions. On four of the five quality dimensions, six on-device models form a quality gradient with capability, corroborated by a judge-free external anchor. Triage does not follow that gradient: on raw scam true-positive rate 11 of 15 model pairs appear to separate, but once the always-suspicious agent's floor is subtracted and the same false-positive rate at the same k is included, only 2 of 15 separate at k=2 and none at the pre-registered k=1—so the 'triage' ordering is mostly an ordering of suspicion. Separately, guardednes

Load-bearing premise

The judge model's ratings are trustworthy measures of the qualities they name, even though the same model wrote the call scenarios and played the caller; if that judge is grading its own preferred style rather than the agents, the specific scores and guardedness counts collapse.

Editorial extensions

If this is right

  • Any evaluation of an on-device phone secretary that reports one quality average is uninterpretable; scores must ship with their counter-metric, an empirical floor, and a bootstrap interval.
  • Reporting scam true-positive rate without the false-positive rate at the same operating point will mislead; the paper's panel shows the apparent triage gradient is mostly a suspicion gradient.
  • Safety and capability are not a single trade-off axis: within the same models, guardedness on disclosure and obedience improves with capability while guardedness against relaying attacker-supplied callback numbers worsens.
  • Text-mode benchmark scores are an upper bound on deployed behavior; audio-pipeline failures such as a 45% turn-take rate are invisible to transcript rubrics.
  • No pass/fail bar is currently defensible for call-secretary quality, because owner-endorsement and human ratings are future work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If this decoupling generalizes, the same score-and-bill design could be applied to other delegated-adversarial tasks—email triage, package-delivery confirmation, message screening—where an absent principal's interests oppose the interlocutor's goals.
  • The G3 pattern suggests an explicit mechanism worth testing: the same conversational fluency that makes an agent endorse legitimate callers may be what carries an attacker's framing into the trusted channel; comparing reasoning traces on keep-versus-relay decisions would test it.
  • Because the paper shows its judge's guardedness gates fall below its own pre-registered agreement threshold, benchmark users should treat any single-model judge as a paired-comparison instrument and demand cross-family agreement before quoting absolute safety counts.
  • The degenerate-floor method is a reusable lower-bound test for agentic benchmarks: any metric on which a scripted hangup-and-echo or always-suspicious agent beats the best real model cannot be read as measuring the capability it names.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces CallScreenBench, a text-mode benchmark for evaluating on-device LLMs acting as phone secretaries. It defines five quality dimensions (triage, message fidelity, representation, caller experience, interaction), each paired with a counter-metric, and three guardedness counts (G1–G3) for a toolless proxy. The authors evaluate six 4-bit on-device models (0.6–4B) plus one cloud row on 48 scenarios with a MiniMax-M2.5 simulator and judge. Main findings: quality scales with capability, triage separations collapse from 11/15 to 0/15 at the pre-registered k=1 after correcting for the always-suspicious floor, delegated guardedness (G3 callback relay) comes apart from quality at the top of the device range, and a cloud model shows the highest relay rate. Extensive limitations are disclosed, including judge circularity, cross-family disagreement, and judge non-determinism.

Significance. The benchmark addresses an important and under-served setting: delegated, adversarial call handling by small on-device models. The design choices—counter-metrics, degenerate scripted floors, no averaging, pre-registered operating points, and unusually transparent limitations—are methodologically valuable and could influence future agent benchmarks. The paper ships released artifacts and a measured cross-family judge analysis rather than asserting independence. If the central claims survive independent judging, the 'guardedness decoupling' finding would be a notable security result. However, the paper's own cross-family agreement data (best κ=0.500) and the single-judge dependence of the cloud-row G3 claim mean the strong headline is currently not established to the paper's own standard. The contribution remains a solid benchmark plus a cautionary measurement study.

major comments (4)
  1. [§4.2, Fig. 7, Appendix E.2] The sentence 'The most capable, most fluent answerer in the matrix is the one most willing to launder an attacker’s number through the trusted channel: the paper’s central claim, one tier above the device range' rests on a single MiniMax-M2.5 judge-scored count (callback_accepted = 17/29 for gpt-4.1-mini). The paper’s own cross-family re-scoring of the 288 device transcripts (E.2) shows that per-model G3 counts invert between judges for the top two device rows (Llama-3.2-3B 8/29 vs Gemma-3-4B 5/29 under MiniMax; 6/29 vs 9/29 under gpt-4.1-mini). The cross-family judge is the same model as the cloud candidate, so it cannot validate the cloud row, and no human labels cover it. At minimum, the claim should be explicitly demoted to 'a single-judge observation' or supported by an independent judge/human audit of the cloud row. As written, this load-bearing claim is not supported to the paper’
  2. [§4.2, Table 11, Limitations (Judge non-determinism)] The triage collapse (11/15 pairs separate on TPR(2) but 0/15 on J at pre-registered k=1) is a negative result established by paired-difference bootstrap intervals that do not exclude zero. The paper states that the bootstrap intervals quantify scenario sampling, not judge variance (Limitations). The conditional judge instability for suspicion_d is 31.4% (E.2), roughly an order of magnitude above the marginal rate, and Q1 is a judge-produced suspicion sweep. Under this noise, 'none of the 15 separate' and the abstract’s 'appearance that it does is an artifact of measurement' may be an artifact of measurement noise rather than a true null. A sensitivity analysis that re-runs the paired intervals under judge-verdict resampling (or reports the maximum plausible J difference after perturbation) is needed before the strong 'triage does not rise with capability' claim can be maintained.
  3. [§3.3, Appendix E.2] The paper’s own pre-registered demotion rule—κ<0.6 demotes a guardedness gate to a score—is failed by both gates (comply κ=0.052, leak κ=0.268), and no re-scored item reaches substantial agreement (best κ=0.500). The paper continues to headline G1–G3 as 'guardedness counts' in the abstract and Figure 7. Since the demotion rule was fixed in advance, the benchmark’s secondary contribution should be presented as judge-specific measurements by a named judge, not as model properties. This does not require new experiments if the framing is changed, but it is a load-bearing consistency issue between the pre-registered criterion and the claims.
  4. [§4, Appendix C] The dataset validation section reports that the leakage judge—the generator model itself—returned 0/89 on the leakage item while an independent regex found 18 simulator-visible hits in 6 scenarios, and that three specified controls (regeneration gates, name caps, disfluency mix) did not run as designed. The text correctly says 'the claim that zero simulator-visible text carries a flagged term is false' and that the released QUALITY_TRIAGE.md repeats the error. Because CallScreenBench is released as a reusable artifact, the repository should be corrected and the 31 retained flagged scenarios (or the leakage claim) should be fixed or explicitly excluded from the released evaluation half before external use. As published, this is a correctness defect in the shipped dataset.
minor comments (4)
  1. [Abstract and §4.2] The abstract and results sections refer to G1–G3 as 'counts' without consistently noting that they are judge-specific measurements. Given Appendix E.2, consider using 'judge-scored counts' in the abstract and at each first mention.
  2. [Table 7 / Figure 5] Gemma-3-1B’s responsiveness is not a model score (1 of 48 cells scored) and is correctly marked 'n/a' in Table 7, but Figure 5 plots the interaction line without this annotation. Adding a visible note or marker would prevent a misleading visual ranking.
  3. [Table 10 / §4.1] The VoiceBench anchor uses a different letter extractor, 4-bit quantization, and a regeneration policy; the paper notes these are not drop-in comparable. A one-sentence reminder in the main text that this is a text-mode, non-official protocol would be helpful.
  4. [Appendix A] The A1 win rate is computed over the 30 pairs on which the judge was self-consistent; this is stated in the caveats but might be repeated where the 26.7% is first introduced in Table 2.

Circularity Check

2 steps flagged · score 6.0 of 10

G3/cloud-ceiling claim rests on the same model as judge, generator, and caller; cross-family check inverts per-model counts and cannot validate the cloud row.

  1. other [§3.3 Judge protocol and reporting rules; Figure 2 caption; Limitations E.2]
    "Transcripts are scored by MiniMax-M2.5 at temperature 0 — the same model, not merely the same family, that generated the scenario content and its per-scenario judge_criteria, a circularity disclosed here and discussed in the Limitations section."

    The judge that produces every judge-scored count (including G3 callback_accepted) is the same model that authored the scenarios and the per-scenario judge_criteria. The paper's own cross-family re-scoring (E.2) shows per-model G3 counts invert: MiniMax gives Llama-3.2-3B 8/29 and Gemma-3-4B 5/29, while gpt-4.1-mini gives 6/29 and 9/29. Thus the G3 result is a property of the scoring judge, not an independent measurement. The cloud row (17/29) was not re-scored by any independent judge; the only alternate judge used in the paper is the cloud candidate itself, so the headline 'one tier above the device range' rests on the same-model judge.

  2. other [§4.2 Results; §D.3; Limitations E.2]
    "The pattern continues above the on-device range. As a ceiling opposite the degenerate floors we ran a cloud model (gpt-4.1-mini-2025-04-14) ... and by a wide margin its worst row on the relay harm — 17 of 29 scam calls end with the attacker's number in the note to the owner — against a maximum of 8/29 anywhere in the device panel ... The most capable, most fluent answerer in the matrix is the one most willing to launder an attacker's number through the trusted channel: the paper's central claim, one tier above the device range."

    This flagship claim is produced by the MiniMax judge, which wrote the scenarios and criteria. The paper's only 'cross-family' judge is gpt-4.1-mini — the same model as the cloud candidate — so it cannot independently validate the 17/29 count for that row. E.2 reports that per-model G3 ordering inverts between judges, and the pre-registered human audit has not been run. The specific number '17/29' is therefore a single-instrument artifact, not an externally grounded model property.

full rationale

The benchmark is largely self-validating against external or judge-free anchors: the capability split is corroborated by VoiceBench OpenBookQA (Figure 6), triage separation collapses to zero only after subtracting the always-suspicious floor (analytically TPR=FPR=1.000), and the paper publishes degenerate-agent floors that defeat its own metrics. No load-bearing self-citation chain is present; the cited forger-as-judge result (Wu et al. 2026) only motivates the cross-family check, which is itself reported. The circularity is real but localized to the judge-scored guardedness profile, specifically G3 and the cloud ceiling row: MiniMax-M2.5 wrote the scenarios and criteria and scores the transcripts, the only cross-family judge is the cloud candidate itself, and the paper's own E.2 shows per-model G3 counts invert between judges. Because the paper's central 'most capable, most fluent ... most willing to launder' claim is this single-judge count, the derivation chain for that claim reduces to the judge's own output; hence partial circularity (6), not total (the coarse capability ordering and triage-floor collapse stand independently).

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The dominant input the central claim rests on is the LLM judge, which is also the author of the test content; that single instrument is the main 'axiom' of the measurement. The remaining axioms are the standard domain assumptions of a benchmark: LLM-judge validity as a proxy for owner endorsement, the text-mode upper-bound interpretation borrowed from tau-Voice, the simulated caller as an adequate adversary model, and US-English/single-profile scope. The protocol constants (k=1, turn cap 8, temperatures, anti-gaming thresholds) are hand-chosen rather than fitted, and several were found insufficient to stop degenerate agents.

free parameters (4)
  • k (triage operating point) = 1 (pre-registered; 2 as secondary sweep)
    Hand-chosen operating point that changes the headline result: at k=1, zero of 15 model pairs separate on Youden's J; at k=2, 2 of 15 separate.
  • judge/simulator temperatures = 0 for judge, 0.7 for simulator
    Design choice; judge non-determinism at temperature 0 still changes 2.9-31.4% of verdicts on re-judged cells.
  • turn cap = 8 round trips
    Fixed cap; a 6-turn cap sweep is released unrun.
  • coherence threshold theta_p and echo overlap = theta_p sweep 0.6-1.0 released unrun; echo overlap 0.60
    Anti-gaming thresholds that failed to catch the paraphrase staller and the parrot.
assumptions (5)
  • domain assumption An LLM judge with a binary per-turn checklist approximates whether an owner would endorse the handling of a call.
    The paper declares human ratings future work and states every Q-dimension is an LLM-judge measurement, so construct validity is assumed, not measured.
  • domain assumption Text-mode performance is an upper bound on deployed audio behavior.
    Used to justify text-mode evaluation as interpretable; relies on tau-Voice's measured 30-45% retention statistic.
  • domain assumption The scripted caller simulator following an authored escalation ladder is an adequate model of an adversarial caller.
    The paper notes it is not an adaptive adversary and makes no evasion-resistance claim.
  • domain assumption The single shared owner profile and US-English scenarios support the stated conclusions.
    The paper acknowledges households, business lines, and Mandarin calls are unmeasured.
  • standard math Standard resampling statistics (paired-difference bootstrap, Holm-Bonferroni, rule of three).
    Used for intervals and multiplicity control.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CallScreenBench: Benchmarking On-Device Models as Phone Secretaries." pith.science (2026). https://pith.science/paper/FG7VOXM2

@misc{pith2026260801033,
  author       = {Pith},
  title        = {Pith review of: CallScreenBench: Benchmarking On-Device Models as Phone Secretaries},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FG7VOXM2}},
  note         = {Machine review of arXiv:2608.01033}
}
read the original abstract

Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting for their user, making on-device task automation newly plausible. One such task is answering the phone. A phone secretary takes an unknown inbound call on its owner's behalf. Unlike the agents evaluated by most benchmarks, it has no task to complete and no cooperative user: the caller holds the goal, may be an adversary, and must be judged from the opening turn with no oracle. What matters is not task success, but whether the owner would endorse how their proxy handled the call. We present CallScreenBench, which scores this setting on five quality dimensions. Each dimension is printed beside the counter-metric that bills it and is never averaged into a single number. We also report a guardedness profile for a toolless proxy that holds no credentials and calls no tools. Across six on-device models (0.6-4B parameters, 4-bit quantization), quality scales with capability, but triage does not. The appearance that it does is an artifact of measurement. Scripted degenerate agents supply the missing floors: after correcting for them, the number of model pairs whose triage performance separates falls from 11 of 15 to zero at the preregistered operating point. An agent that simply hangs up and echoes the caller also scores perfect message fidelity. We report which of our own metrics these floors defeat and declare no pass/fail threshold.

Figures

Figures reproduced from arXiv: 2608.01033 by the authors.

Figure 1
Figure 1. Delegation and goal inversion. Cooperative benchmarks (left) score an agent serving a user’s goal; CallScreenBench (right) scores a delegated secretary answering on the owner’s behalf, where the caller holds the goal — decided from the same opening turn with no oracle. on one family and a product failure on the other — so quality cannot be a single average: every CallScreenBench score ships a counter-metric from the… view at source ↗
Figure 2
Figure 2. The CallScreenBench text-mode evaluation loop, drawn by role rather than by model so the protocol is not [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. CallScreenBench measurement, organized around one question. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Composition of the CallScreenBench-89 eval [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: The quality gradient (higher is better on every [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: External anchor. Our six on-device checkpoints (solid, 95% Wilson) against published VoiceBench systems (hatched) on OpenBookQA, text￾instruction condition. Dotted line: chance (25.0). Dashed: the open 8B cascade baseline (66.2). ends are unambiguous, a judge-independe…
Figure 7
Figure 7. Figure 7: Secondary guardedness (counts in Table [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

73 extracted references · 55 canonical work pages

  1. [1]

    Ray, Soham and Dhandhania, Keshav and Barres, Victor and Narasimhan, Karthik , journal =

  2. [2]

    arXiv preprint arXiv:2604.04847 , year =

    Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency , author =. arXiv preprint arXiv:2604.04847 , year =

  3. [3]

    Yan, Ruiqi and Li, Xiquan and Chen, Wenxi and Niu, Zhikang and Yang, Chen and Ma, Ziyang and Yu, Kai and Chen, Xie , journal =

  4. [4]

    Liu, Hongcheng and Hou, Yixuan and Liu, Heyang and Wang, Yuhao and Wang, Yanfeng and Wang, Yu , journal =

  5. [5]

    Hou, Yixuan and Liu, Heyang and Wang, Yuhao and Cheng, Ziyang and Wu, Ronghua and Gu, Qunshan and Wang, Yanfeng and Wang, Yu , journal =

  6. [6]

    Deng, Yayue and Hu, Guoqiang and Sun, Haiyang and Zhang, Xiangyu and Zhang, Haoyang and Tian, Fei and Yang, Xuerui and Yu, Gang and Chng, Eng Siong , journal =

  7. [7]

    Xu, Pengyu and Li, Shijia and Sun, Ao and Zhang, Feng and Li, Yahan and Wu, Bo and Ma, Zhanyu and Li, Jiguo and Xu, Jun and Gao, Jiuchong and Hao, Jinghua and He, Renqing and Wang, Rui and Liu, Yang and Hu, Xiaobo and Yang, Fan and Zheng, Jia and Yao, Guanghua , journal =

  8. [8]

    and Li, Haizhou , journal =

    Chen, Yiming and Yue, Xianghu and Zhang, Chen and Gao, Xiaoxue and Tan, Robby T. and Li, Haizhou , journal =

Show all 73 references
  1. [9]

    Barres, Victor and Dong, Honghua and Ray, Soham and Si, Xujie and Narasimhan, Karthik , journal =

  2. [10]

    arXiv preprint arXiv:2607.14846 , year =

    Ayll. arXiv preprint arXiv:2607.14846 , year =

  3. [11]

    International Conference on Learning Representations (ICLR) , year =

    Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics , author =. International Conference on Learning Representations (ICLR) , year =

  4. [12]

    IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , year =

    Full-Duplex-Bench: A Benchmark to Evaluate Full-Duplex Spoken Dialogue Models on Turn-Taking Capabilities , author =. IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , year =

  5. [13]

    Ekstedt, Erik and Skantze, Gabriel , booktitle =

  6. [14]

    Proceedings of Interspeech 2022 , year =

    Voice Activity Projection: Self-Supervised Learning of Turn-Taking Events , author =. Proceedings of Interspeech 2022 , year =

  7. [15]

    Proceedings of the National Academy of Sciences , volume =

    Universals and Cultural Variation in Turn-Taking in Conversation , author =. Proceedings of the National Academy of Sciences , volume =. 2009 , doi =

  8. [16]

    arXiv preprint arXiv:2203.16502 , year =

    Generative Spoken Dialogue Language Modeling , author =. arXiv preprint arXiv:2203.16502 , year =

  9. [17]

    arXiv preprint arXiv:2410.00037 , year =

    Moshi: A Speech-Text Foundation Model for Real-Time Dialogue , author =. arXiv preprint arXiv:2410.00037 , year =

  10. [18]

    Semantic-

    Roy, Somnath , journal =. Semantic-

  11. [19]

    , journal =

    Kim, Suyoun and Arora, Abhinav and Le, Duc and Yeh, Ching-Feng and Fuegen, Christian and Kalinli, Ozlem and Seltzer, Michael L. , journal =. Semantic Distance: A New Metric for

  12. [20]

    arXiv preprint arXiv:2605.29430 , year =

    Towards Human-Like Interactive Speech Recognition with Agentic Correction and Semantic Evaluation , author =. arXiv preprint arXiv:2605.29430 , year =

  13. [21]

    arXiv preprint arXiv:2212.04356 , year =

    Robust Speech Recognition via Large-Scale Weak Supervision , author =. arXiv preprint arXiv:2212.04356 , year =

  14. [22]

    Saeki, Takaaki and Xin, Detai and Nakata, Wataru and Koriyama, Tomoki and Takamichi, Shinnosuke and Saruwatari, Hiroshi , journal =

  15. [23]

    arXiv preprint arXiv:2104.09494 , year =

    Mittag, Gabriel and Naderi, Babak and Chehadi, Assmaa and M. arXiv preprint arXiv:2104.09494 , year =

  16. [24]

    arXiv preprint arXiv:2407.12707 , year =

    Minixhofer, Christoph and Klejch, Ond. arXiv preprint arXiv:2407.12707 , year =

  17. [25]

    arXiv preprint arXiv:2603.01467 , year =

    Conversational Speech Naturalness Predictor , author =. arXiv preprint arXiv:2603.01467 , year =

  18. [26]

    and Khan, Ali Sartaz and Sirichotedumrong, Warit and Pipatanakul, Kunat and Held, William Barr and Yang, Diyi , booktitle =

    Manakul, Potsawee and Gan, Woody Haosheng and Ryan, Michael J. and Khan, Ali Sartaz and Sirichotedumrong, Warit and Pipatanakul, Kunat and Held, William Barr and Yang, Diyi , booktitle =. 2026 , address =

  19. [27]

    Wang, Hui and Zhao, Jinghua and Yang, Yifan and Liu, Shujie and Chen, Junyang and Zhang, Yanzhe and Zhao, Shiwan and Li, Jinyu and Zhou, Jiaming and Sun, Haoqin and Lu, Yan and Qin, Yong , journal =

  20. [28]

    and Emmons, J

    Sayyad, A. and Emmons, J. and Jones, S. and Lin, T. and Krishnan, H. , journal =. A Reliability Assessment of. 2026 , note =

  21. [29]

    Luo, Xuan and Yao, Lewei and Zhao, Libo and Hong, Lanqing and Chen, Kai and Tao, Dehua and Tan, Daxin and Xu, Ruifeng and Li, Jing , journal =

  22. [30]

    Proceedings of the 29th USENIX Security Symposium (USENIX Security '20) , year =

    Who's Calling? Characterizing Robocalls Through Audio and Metadata Analysis , author =. Proceedings of the 29th USENIX Security Symposium (USENIX Security '20) , year =

  23. [31]

    Proceedings of the 32nd USENIX Security Symposium (USENIX Security '23) , year =

    Combating Robocalls with Phone Virtual Assistant Mediated Interaction , author =. Proceedings of the 32nd USENIX Security Symposium (USENIX Security '23) , year =

  24. [32]

    Diving into Robocall Content with

    Prasad, Sathvik and Dunlap, Trevor and Ross, Alexander and Reaves, Bradley , booktitle =. Diving into Robocall Content with

  25. [33]

    Robocall Audio from the

    Prasad, Sathvik and Reaves, Bradley , institution =. Robocall Audio from the

  26. [34]

    IEEE Symposium on Security and Privacy (S&P) , year =

    Characterizing Robocalls with Multiple Vantage Points , author =. IEEE Symposium on Security and Privacy (S&P) , year =

  27. [35]

    Using Chatbots Against Voice Spam: Analyzing

    Sahin, Merve and Relieu, Marc and Francillon, Aur. Using Chatbots Against Voice Spam: Analyzing. Proceedings of the Thirteenth Symposium on Usable Privacy and Security (SOUPS '17) , pages =

  28. [36]

    Send to Which Account? Evaluation of an

    Siadati, Hossein and Jafarian, Haadi and Jafarikhah, Sima , journal =. Send to Which Account? Evaluation of an

  29. [37]

    , journal =

    Shen, Zitong and Wang, Kangzhong and Zhang, Youqian and Ngai, Grace and Fu, Eugene Y. , journal =. Combating Phone Scams with

  30. [38]

    and Zhang, Hao and Gonzalez, Joseph E

    Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , booktitle =. Judging

  31. [39]

    arXiv preprint arXiv:2305.17926 , year =

    Large Language Models are not Fair Evaluators , author =. arXiv preprint arXiv:2305.17926 , year =

  32. [40]

    arXiv preprint arXiv:2505.09388 , year =

    Qwen3 Technical Report , author =. arXiv preprint arXiv:2505.09388 , year =

  33. [41]

    arXiv preprint arXiv:2503.19786 , year =

    Gemma 3 Technical Report , author =. arXiv preprint arXiv:2503.19786 , year =

  34. [42]

    arXiv preprint arXiv:2502.02737 , year =

    Allal, Loubna Ben and Lozhkov, Anton and Bakouch, Elie and Bl. arXiv preprint arXiv:2502.02737 , year =

  35. [43]

    arXiv preprint arXiv:2407.21783 , year =

    The Llama 3 Herd of Models , author =. arXiv preprint arXiv:2407.21783 , year =

  36. [44]

    arXiv preprint arXiv:2404.14219 , year =

    Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone , author =. arXiv preprint arXiv:2404.14219 , year =

  37. [45]

    Murthy, Rithesh and Yang, Liangwei and Tan, Juntao and Awalgaonkar, Tulika Manoj and Zhou, Yilun and Heinecke, Shelby and Desai, Sachin and Wu, Jason and Xu, Ran and Tan, Sarah and Zhang, Jianguo and Liu, Zhiwei and Kokane, Shirley and Liu, Zuxin and Zhu, Ming and Wang, Huan a...

  38. [46]

    2024 , note =

    Screen your calls before you answer them , author =. 2024 , note =

  39. [47]

    2025 , note =

    Screen and block calls on iPhone , author =. 2025 , note =

  40. [48]

    2026 , note =

    Speech-to-Speech Models and Providers Analysis (Big Bench Audio) , author =. 2026 , note =

  41. [49]

    2026 , note =

    Voice agent evaluation framework: 6 pillars explained , author =. 2026 , note =

  42. [50]

    2025 , note =

    How to evaluate. 2025 , note =

  43. [51]

    2026 , note =

    Voice. 2026 , note =

  44. [52]

    2026 , note =

    How to Evaluate Voice Agents: A Complete Framework for Testing , author =. 2026 , note =

  45. [53]

    Findings of the Association for Computational Linguistics: ACL 2023 , year =

    Discovering Language Model Behaviors with Model-Written Evaluations , author =. Findings of the Association for Computational Linguistics: ACL 2023 , year =

  46. [54]

    International Conference on Learning Representations (ICLR) , year =

    Towards Understanding Sycophancy in Language Models , author =. International Conference on Learning Representations (ICLR) , year =

  47. [55]

    Transactions on Machine Learning Research , year =

    Inverse Scaling: When Bigger Isn't Better , author =. Transactions on Machine Learning Research , year =

  48. [56]

    How Johnny Can Persuade

    Zeng, Yi and Lin, Hongpeng and Zhang, Jingwen and Yang, Diyi and Jia, Ruoxi and Shi, Weiyan , booktitle =. How Johnny Can Persuade. 2024 , note =

  49. [57]

    Not What You've Signed Up For: Compromising Real-World

    Greshake, Kai and Abdelnabi, Sahar and Mishra, Shailesh and Endres, Christoph and Holz, Thorsten and Fritz, Mario , journal =. Not What You've Signed Up For: Compromising Real-World. 2023 , note =

  50. [58]

    Mireshghallah, Niloofar and Kim, Hyunwoo and Zhou, Xuhui and Tsvetkov, Yulia and Sap, Maarten and Shokri, Reza and Choi, Yejin , booktitle =. Can. 2024 , note =

  51. [59]

    2024 , note =

    Shao, Yijia and Li, Tianshi and Shi, Weiyan and Liu, Yanchen and Yang, Diyi , booktitle =. 2024 , note =

  52. [60]

    Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =

    Can Large Language Models Be an Alternative to Human Evaluations? , author =. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =

  53. [61]

    and Feng, Shi , journal =

    Panickssery, Arjun and Bowman, Samuel R. and Feng, Shi , journal =

  54. [62]

    Transactions on Machine Learning Research , year =

    Holistic Evaluation of Language Models , author =. Transactions on Machine Learning Research , year =

  55. [63]

    Beyond Accuracy: Behavioral Testing of

    Ribeiro, Marco Tulio and Wu, Tongshuang and Guestrin, Carlos and Singh, Sameer , booktitle =. Beyond Accuracy: Behavioral Testing of. 2020 , note =

  56. [64]

    and Fitz, Stephen and Hendrycks, Dan , booktitle =

    Ren, Richard and Basart, Steven and Khoja, Adam and Gatti, Alice and Phan, Long and Yin, Xuwang and Mazeika, Mantas and Pan, Alexander and Mukobi, Gabriel and Kim, Ryan H. and Fitz, Stephen and Hendrycks, Dan , booktitle =. Safetywashing: Do. 2024 , note =

  57. [65]

    Washington Law Review , volume =

    Privacy as Contextual Integrity , author =. Washington Law Review , volume =

  58. [66]

    Li, Margaret and Weston, Jason and Roller, Stephen , journal =

  59. [67]

    Laban, Philippe and Hayashi, Hiroaki and Zhou, Yingbo and Neville, Jennifer , journal =

  60. [68]

    The Instruction Hierarchy: Training

    Wallace, Eric and Xiao, Kai and Leike, Reimar and Weng, Lilian and Heidecke, Johannes and Beutel, Alex , journal =. The Instruction Hierarchy: Training

  61. [69]

    and Wang, Di , booktitle =

    Yang, Shu and Zhu, Shenzhe and Wu, Zeyu and Wang, Keyu and Yao, Junchi and Wu, Junchao and Hu, Lijie and Li, Mengdi and Wong, Derek F. and Wang, Di , booktitle =. 2025 , note =

  62. [70]

    2025 , note =

    Apate.ai:. 2025 , note =

  63. [71]

    and Zhou, Y

    Wu, J. and Zhou, Y. and Ng, D. T. and Shen, X. and Zewde, K. and Raj, A. and Duong, T. and Ren, Simiao , year =. When the Forger Is the Judge:. 2604.25213 , archivePrefix =

  64. [72]

    and Ren, Simiao and Raj, A

    Zhang, Y. and Ren, Simiao and Raj, A. and Wei, E. and Ng, D. and Shen, A. and Xue, J. and Zhang, Y. and Marotta, E. , year =. 2603.11442 , archivePrefix =

  65. [73]

    and Shen, X

    Ren, Simiao and Zhou, Y. and Shen, X. and Zewde, K. and Duong, T. and Huang, G. and Wei, E. and Xue, J. , year =. How Well Are Open-Sourced. 2602.07814 , archivePrefix =

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.