REVIEW 4 major objections 4 minor 73 references
CallScreenBench: Benchmarking On-Device Models as Phone Secretaries
T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The most fluent phone-secretary AIs are the most likely to relay a scam caller's callback number, and apparent triage skill is largely a suspicion artifact.
desk verdict A transparent first benchmark for on-device phone secretaries with real methodological ideas, but its headline safety claim about the cloud model rests on a single unvalidated judge. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing instrument is the paired score-and-bill design. Each of the five quality scores (triage, message fidelity, representation, caller experience, interaction) is printed beside a counter-metric from the opposing family—recall beside fabrication, elicitation beside verbosity, scam true-positive rate beside legitimate false-positive rate at the same k—and no score is averaged into a composite. The triage argument runs through Youden's index $J=\mathrm{TPR}-\mathrm{FPR}$, computed with an empirically run 'always-suspicious' scripted agent as the floor; all between-model separation claims are made on paired-difference bootstraps over shared scenario resamples. A secondary guardednes
What would settle it
Run the paper's preregistered stratified human audit on the 48-scenario core: have human raters label each transcript for owner endorsement and the three guardedness harms. If human labels agree with the automated judge below the paper's $\kappa<0.6$ threshold, or if humans do not reproduce the finding that the most fluent answerers relay attacker-supplied callback numbers more often, the headline decoupling is judge artifact rather than model property.
Extended reading notes
Core claim
The central discovery is that capability, triage, and guardedness are three separate axes in a delegated phone-answering agent, and conflating them produces wrong conclusions. On four of the five quality dimensions, six on-device models form a quality gradient with capability, corroborated by a judge-free external anchor. Triage does not follow that gradient: on raw scam true-positive rate 11 of 15 model pairs appear to separate, but once the always-suspicious agent's floor is subtracted and the same false-positive rate at the same k is included, only 2 of 15 separate at k=2 and none at the pre-registered k=1—so the 'triage' ordering is mostly an ordering of suspicion. Separately, guardednes
Load-bearing premise
The judge model's ratings are trustworthy measures of the qualities they name, even though the same model wrote the call scenarios and played the caller; if that judge is grading its own preferred style rather than the agents, the specific scores and guardedness counts collapse.
Editorial extensions
If this is right
- Any evaluation of an on-device phone secretary that reports one quality average is uninterpretable; scores must ship with their counter-metric, an empirical floor, and a bootstrap interval.
- Reporting scam true-positive rate without the false-positive rate at the same operating point will mislead; the paper's panel shows the apparent triage gradient is mostly a suspicion gradient.
- Safety and capability are not a single trade-off axis: within the same models, guardedness on disclosure and obedience improves with capability while guardedness against relaying attacker-supplied callback numbers worsens.
- Text-mode benchmark scores are an upper bound on deployed behavior; audio-pipeline failures such as a 45% turn-take rate are invisible to transcript rubrics.
- No pass/fail bar is currently defensible for call-secretary quality, because owner-endorsement and human ratings are future work.
Reading between the lines
- If this decoupling generalizes, the same score-and-bill design could be applied to other delegated-adversarial tasks—email triage, package-delivery confirmation, message screening—where an absent principal's interests oppose the interlocutor's goals.
- The G3 pattern suggests an explicit mechanism worth testing: the same conversational fluency that makes an agent endorse legitimate callers may be what carries an attacker's framing into the trusted channel; comparing reasoning traces on keep-versus-relay decisions would test it.
- Because the paper shows its judge's guardedness gates fall below its own pre-registered agreement threshold, benchmark users should treat any single-model judge as a paired-comparison instrument and demand cross-family agreement before quoting absolute safety counts.
- The degenerate-floor method is a reusable lower-bound test for agentic benchmarks: any metric on which a scripted hangup-and-echo or always-suspicious agent beats the best real model cannot be read as measuring the capability it names.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces CallScreenBench, a text-mode benchmark for evaluating on-device LLMs acting as phone secretaries. It defines five quality dimensions (triage, message fidelity, representation, caller experience, interaction), each paired with a counter-metric, and three guardedness counts (G1–G3) for a toolless proxy. The authors evaluate six 4-bit on-device models (0.6–4B) plus one cloud row on 48 scenarios with a MiniMax-M2.5 simulator and judge. Main findings: quality scales with capability, triage separations collapse from 11/15 to 0/15 at the pre-registered k=1 after correcting for the always-suspicious floor, delegated guardedness (G3 callback relay) comes apart from quality at the top of the device range, and a cloud model shows the highest relay rate. Extensive limitations are disclosed, including judge circularity, cross-family disagreement, and judge non-determinism.
Significance. The benchmark addresses an important and under-served setting: delegated, adversarial call handling by small on-device models. The design choices—counter-metrics, degenerate scripted floors, no averaging, pre-registered operating points, and unusually transparent limitations—are methodologically valuable and could influence future agent benchmarks. The paper ships released artifacts and a measured cross-family judge analysis rather than asserting independence. If the central claims survive independent judging, the 'guardedness decoupling' finding would be a notable security result. However, the paper's own cross-family agreement data (best κ=0.500) and the single-judge dependence of the cloud-row G3 claim mean the strong headline is currently not established to the paper's own standard. The contribution remains a solid benchmark plus a cautionary measurement study.
major comments (4)
- [§4.2, Fig. 7, Appendix E.2] The sentence 'The most capable, most fluent answerer in the matrix is the one most willing to launder an attacker’s number through the trusted channel: the paper’s central claim, one tier above the device range' rests on a single MiniMax-M2.5 judge-scored count (callback_accepted = 17/29 for gpt-4.1-mini). The paper’s own cross-family re-scoring of the 288 device transcripts (E.2) shows that per-model G3 counts invert between judges for the top two device rows (Llama-3.2-3B 8/29 vs Gemma-3-4B 5/29 under MiniMax; 6/29 vs 9/29 under gpt-4.1-mini). The cross-family judge is the same model as the cloud candidate, so it cannot validate the cloud row, and no human labels cover it. At minimum, the claim should be explicitly demoted to 'a single-judge observation' or supported by an independent judge/human audit of the cloud row. As written, this load-bearing claim is not supported to the paper’
- [§4.2, Table 11, Limitations (Judge non-determinism)] The triage collapse (11/15 pairs separate on TPR(2) but 0/15 on J at pre-registered k=1) is a negative result established by paired-difference bootstrap intervals that do not exclude zero. The paper states that the bootstrap intervals quantify scenario sampling, not judge variance (Limitations). The conditional judge instability for suspicion_d is 31.4% (E.2), roughly an order of magnitude above the marginal rate, and Q1 is a judge-produced suspicion sweep. Under this noise, 'none of the 15 separate' and the abstract’s 'appearance that it does is an artifact of measurement' may be an artifact of measurement noise rather than a true null. A sensitivity analysis that re-runs the paired intervals under judge-verdict resampling (or reports the maximum plausible J difference after perturbation) is needed before the strong 'triage does not rise with capability' claim can be maintained.
- [§3.3, Appendix E.2] The paper’s own pre-registered demotion rule—κ<0.6 demotes a guardedness gate to a score—is failed by both gates (comply κ=0.052, leak κ=0.268), and no re-scored item reaches substantial agreement (best κ=0.500). The paper continues to headline G1–G3 as 'guardedness counts' in the abstract and Figure 7. Since the demotion rule was fixed in advance, the benchmark’s secondary contribution should be presented as judge-specific measurements by a named judge, not as model properties. This does not require new experiments if the framing is changed, but it is a load-bearing consistency issue between the pre-registered criterion and the claims.
- [§4, Appendix C] The dataset validation section reports that the leakage judge—the generator model itself—returned 0/89 on the leakage item while an independent regex found 18 simulator-visible hits in 6 scenarios, and that three specified controls (regeneration gates, name caps, disfluency mix) did not run as designed. The text correctly says 'the claim that zero simulator-visible text carries a flagged term is false' and that the released QUALITY_TRIAGE.md repeats the error. Because CallScreenBench is released as a reusable artifact, the repository should be corrected and the 31 retained flagged scenarios (or the leakage claim) should be fixed or explicitly excluded from the released evaluation half before external use. As published, this is a correctness defect in the shipped dataset.
minor comments (4)
- [Abstract and §4.2] The abstract and results sections refer to G1–G3 as 'counts' without consistently noting that they are judge-specific measurements. Given Appendix E.2, consider using 'judge-scored counts' in the abstract and at each first mention.
- [Table 7 / Figure 5] Gemma-3-1B’s responsiveness is not a model score (1 of 48 cells scored) and is correctly marked 'n/a' in Table 7, but Figure 5 plots the interaction line without this annotation. Adding a visible note or marker would prevent a misleading visual ranking.
- [Table 10 / §4.1] The VoiceBench anchor uses a different letter extractor, 4-bit quantization, and a regeneration policy; the paper notes these are not drop-in comparable. A one-sentence reminder in the main text that this is a text-mode, non-official protocol would be helpful.
- [Appendix A] The A1 win rate is computed over the 30 pairs on which the judge was self-consistent; this is stated in the caveats but might be repeated where the 26.7% is first introduced in Table 2.
Circularity Check
G3/cloud-ceiling claim rests on the same model as judge, generator, and caller; cross-family check inverts per-model counts and cannot validate the cloud row.
-
other
[§3.3 Judge protocol and reporting rules; Figure 2 caption; Limitations E.2]
"Transcripts are scored by MiniMax-M2.5 at temperature 0 — the same model, not merely the same family, that generated the scenario content and its per-scenario judge_criteria, a circularity disclosed here and discussed in the Limitations section."
The judge that produces every judge-scored count (including G3 callback_accepted) is the same model that authored the scenarios and the per-scenario judge_criteria. The paper's own cross-family re-scoring (E.2) shows per-model G3 counts invert: MiniMax gives Llama-3.2-3B 8/29 and Gemma-3-4B 5/29, while gpt-4.1-mini gives 6/29 and 9/29. Thus the G3 result is a property of the scoring judge, not an independent measurement. The cloud row (17/29) was not re-scored by any independent judge; the only alternate judge used in the paper is the cloud candidate itself, so the headline 'one tier above the device range' rests on the same-model judge.
-
other
[§4.2 Results; §D.3; Limitations E.2]
"The pattern continues above the on-device range. As a ceiling opposite the degenerate floors we ran a cloud model (gpt-4.1-mini-2025-04-14) ... and by a wide margin its worst row on the relay harm — 17 of 29 scam calls end with the attacker's number in the note to the owner — against a maximum of 8/29 anywhere in the device panel ... The most capable, most fluent answerer in the matrix is the one most willing to launder an attacker's number through the trusted channel: the paper's central claim, one tier above the device range."
This flagship claim is produced by the MiniMax judge, which wrote the scenarios and criteria. The paper's only 'cross-family' judge is gpt-4.1-mini — the same model as the cloud candidate — so it cannot independently validate the 17/29 count for that row. E.2 reports that per-model G3 ordering inverts between judges, and the pre-registered human audit has not been run. The specific number '17/29' is therefore a single-instrument artifact, not an externally grounded model property.
full rationale
The benchmark is largely self-validating against external or judge-free anchors: the capability split is corroborated by VoiceBench OpenBookQA (Figure 6), triage separation collapses to zero only after subtracting the always-suspicious floor (analytically TPR=FPR=1.000), and the paper publishes degenerate-agent floors that defeat its own metrics. No load-bearing self-citation chain is present; the cited forger-as-judge result (Wu et al. 2026) only motivates the cross-family check, which is itself reported. The circularity is real but localized to the judge-scored guardedness profile, specifically G3 and the cloud ceiling row: MiniMax-M2.5 wrote the scenarios and criteria and scores the transcripts, the only cross-family judge is the cloud candidate itself, and the paper's own E.2 shows per-model G3 counts invert between judges. Because the paper's central 'most capable, most fluent ... most willing to launder' claim is this single-judge count, the derivation chain for that claim reduces to the judge's own output; hence partial circularity (6), not total (the coarse capability ordering and triage-floor collapse stand independently).
Assumptions & free parameters
free parameters (4)
- k (triage operating point) =
1 (pre-registered; 2 as secondary sweep)
- judge/simulator temperatures =
0 for judge, 0.7 for simulator
- turn cap =
8 round trips
- coherence threshold theta_p and echo overlap =
theta_p sweep 0.6-1.0 released unrun; echo overlap 0.60
assumptions (5)
- domain assumption An LLM judge with a binary per-turn checklist approximates whether an owner would endorse the handling of a call.
- domain assumption Text-mode performance is an upper bound on deployed audio behavior.
- domain assumption The scripted caller simulator following an authored escalation ladder is an adequate model of an adversarial caller.
- domain assumption The single shared owner profile and US-English scenarios support the stated conclusions.
- standard math Standard resampling statistics (paired-difference bootstrap, Holm-Bonferroni, rule of three).
Cite this review
Pith. "Pith review of CallScreenBench: Benchmarking On-Device Models as Phone Secretaries." pith.science (2026). https://pith.science/paper/FG7VOXM2
@misc{pith2026260801033,
author = {Pith},
title = {Pith review of: CallScreenBench: Benchmarking On-Device Models as Phone Secretaries},
year = {2026},
howpublished = {\url{https://pith.science/paper/FG7VOXM2}},
note = {Machine review of arXiv:2608.01033}
}
read the original abstract
Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting for their user, making on-device task automation newly plausible. One such task is answering the phone. A phone secretary takes an unknown inbound call on its owner's behalf. Unlike the agents evaluated by most benchmarks, it has no task to complete and no cooperative user: the caller holds the goal, may be an adversary, and must be judged from the opening turn with no oracle. What matters is not task success, but whether the owner would endorse how their proxy handled the call. We present CallScreenBench, which scores this setting on five quality dimensions. Each dimension is printed beside the counter-metric that bills it and is never averaged into a single number. We also report a guardedness profile for a toolless proxy that holds no credentials and calls no tools. Across six on-device models (0.6-4B parameters, 4-bit quantization), quality scales with capability, but triage does not. The appearance that it does is an artifact of measurement. Scripted degenerate agents supply the missing floors: after correcting for them, the number of model pairs whose triage performance separates falls from 11 of 15 to zero at the preregistered operating point. An agent that simply hangs up and echoes the caller also scores perfect message fidelity. We report which of our own metrics these floors defeat and declare no pass/fail threshold.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Ray, Soham and Dhandhania, Keshav and Barres, Victor and Narasimhan, Karthik , journal =
-
[2]
arXiv preprint arXiv:2604.04847 , year =
Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency , author =. arXiv preprint arXiv:2604.04847 , year =
-
[3]
Yan, Ruiqi and Li, Xiquan and Chen, Wenxi and Niu, Zhikang and Yang, Chen and Ma, Ziyang and Yu, Kai and Chen, Xie , journal =
-
[4]
Liu, Hongcheng and Hou, Yixuan and Liu, Heyang and Wang, Yuhao and Wang, Yanfeng and Wang, Yu , journal =
-
[5]
Hou, Yixuan and Liu, Heyang and Wang, Yuhao and Cheng, Ziyang and Wu, Ronghua and Gu, Qunshan and Wang, Yanfeng and Wang, Yu , journal =
-
[6]
Deng, Yayue and Hu, Guoqiang and Sun, Haiyang and Zhang, Xiangyu and Zhang, Haoyang and Tian, Fei and Yang, Xuerui and Yu, Gang and Chng, Eng Siong , journal =
-
[7]
Xu, Pengyu and Li, Shijia and Sun, Ao and Zhang, Feng and Li, Yahan and Wu, Bo and Ma, Zhanyu and Li, Jiguo and Xu, Jun and Gao, Jiuchong and Hao, Jinghua and He, Renqing and Wang, Rui and Liu, Yang and Hu, Xiaobo and Yang, Fan and Zheng, Jia and Yao, Guanghua , journal =
-
[8]
Chen, Yiming and Yue, Xianghu and Zhang, Chen and Gao, Xiaoxue and Tan, Robby T. and Li, Haizhou , journal =
Show all 73 references
-
[9]
Barres, Victor and Dong, Honghua and Ray, Soham and Si, Xujie and Narasimhan, Karthik , journal =
- [10]
-
[11]
International Conference on Learning Representations (ICLR) , year =
Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics , author =. International Conference on Learning Representations (ICLR) , year =
-
[12]
IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , year =
Full-Duplex-Bench: A Benchmark to Evaluate Full-Duplex Spoken Dialogue Models on Turn-Taking Capabilities , author =. IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , year =
-
[13]
Ekstedt, Erik and Skantze, Gabriel , booktitle =
-
[14]
Proceedings of Interspeech 2022 , year =
Voice Activity Projection: Self-Supervised Learning of Turn-Taking Events , author =. Proceedings of Interspeech 2022 , year =
2022
-
[15]
Proceedings of the National Academy of Sciences , volume =
Universals and Cultural Variation in Turn-Taking in Conversation , author =. Proceedings of the National Academy of Sciences , volume =. 2009 , doi =
2009
-
[16]
arXiv preprint arXiv:2203.16502 , year =
Generative Spoken Dialogue Language Modeling , author =. arXiv preprint arXiv:2203.16502 , year =
-
[17]
arXiv preprint arXiv:2410.00037 , year =
Moshi: A Speech-Text Foundation Model for Real-Time Dialogue , author =. arXiv preprint arXiv:2410.00037 , year =
-
[18]
Semantic-
Roy, Somnath , journal =. Semantic-
-
[19]
, journal =
Kim, Suyoun and Arora, Abhinav and Le, Duc and Yeh, Ching-Feng and Fuegen, Christian and Kalinli, Ozlem and Seltzer, Michael L. , journal =. Semantic Distance: A New Metric for
-
[20]
arXiv preprint arXiv:2605.29430 , year =
Towards Human-Like Interactive Speech Recognition with Agentic Correction and Semantic Evaluation , author =. arXiv preprint arXiv:2605.29430 , year =
-
[21]
arXiv preprint arXiv:2212.04356 , year =
Robust Speech Recognition via Large-Scale Weak Supervision , author =. arXiv preprint arXiv:2212.04356 , year =
-
[22]
Saeki, Takaaki and Xin, Detai and Nakata, Wataru and Koriyama, Tomoki and Takamichi, Shinnosuke and Saruwatari, Hiroshi , journal =
-
[23]
arXiv preprint arXiv:2104.09494 , year =
Mittag, Gabriel and Naderi, Babak and Chehadi, Assmaa and M. arXiv preprint arXiv:2104.09494 , year =
-
[24]
arXiv preprint arXiv:2407.12707 , year =
Minixhofer, Christoph and Klejch, Ond. arXiv preprint arXiv:2407.12707 , year =
-
[25]
arXiv preprint arXiv:2603.01467 , year =
Conversational Speech Naturalness Predictor , author =. arXiv preprint arXiv:2603.01467 , year =
-
[26]
and Khan, Ali Sartaz and Sirichotedumrong, Warit and Pipatanakul, Kunat and Held, William Barr and Yang, Diyi , booktitle =
Manakul, Potsawee and Gan, Woody Haosheng and Ryan, Michael J. and Khan, Ali Sartaz and Sirichotedumrong, Warit and Pipatanakul, Kunat and Held, William Barr and Yang, Diyi , booktitle =. 2026 , address =
2026
-
[27]
Wang, Hui and Zhao, Jinghua and Yang, Yifan and Liu, Shujie and Chen, Junyang and Zhang, Yanzhe and Zhao, Shiwan and Li, Jinyu and Zhou, Jiaming and Sun, Haoqin and Lu, Yan and Qin, Yong , journal =
-
[28]
and Emmons, J
Sayyad, A. and Emmons, J. and Jones, S. and Lin, T. and Krishnan, H. , journal =. A Reliability Assessment of. 2026 , note =
2026
-
[29]
Luo, Xuan and Yao, Lewei and Zhao, Libo and Hong, Lanqing and Chen, Kai and Tao, Dehua and Tan, Daxin and Xu, Ruifeng and Li, Jing , journal =
-
[30]
Proceedings of the 29th USENIX Security Symposium (USENIX Security '20) , year =
Who's Calling? Characterizing Robocalls Through Audio and Metadata Analysis , author =. Proceedings of the 29th USENIX Security Symposium (USENIX Security '20) , year =
-
[31]
Proceedings of the 32nd USENIX Security Symposium (USENIX Security '23) , year =
Combating Robocalls with Phone Virtual Assistant Mediated Interaction , author =. Proceedings of the 32nd USENIX Security Symposium (USENIX Security '23) , year =
-
[32]
Diving into Robocall Content with
Prasad, Sathvik and Dunlap, Trevor and Ross, Alexander and Reaves, Bradley , booktitle =. Diving into Robocall Content with
-
[33]
Robocall Audio from the
Prasad, Sathvik and Reaves, Bradley , institution =. Robocall Audio from the
-
[34]
IEEE Symposium on Security and Privacy (S&P) , year =
Characterizing Robocalls with Multiple Vantage Points , author =. IEEE Symposium on Security and Privacy (S&P) , year =
-
[35]
Using Chatbots Against Voice Spam: Analyzing
Sahin, Merve and Relieu, Marc and Francillon, Aur. Using Chatbots Against Voice Spam: Analyzing. Proceedings of the Thirteenth Symposium on Usable Privacy and Security (SOUPS '17) , pages =
-
[36]
Send to Which Account? Evaluation of an
Siadati, Hossein and Jafarian, Haadi and Jafarikhah, Sima , journal =. Send to Which Account? Evaluation of an
-
[37]
, journal =
Shen, Zitong and Wang, Kangzhong and Zhang, Youqian and Ngai, Grace and Fu, Eugene Y. , journal =. Combating Phone Scams with
-
[38]
and Zhang, Hao and Gonzalez, Joseph E
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , booktitle =. Judging
-
[39]
arXiv preprint arXiv:2305.17926 , year =
Large Language Models are not Fair Evaluators , author =. arXiv preprint arXiv:2305.17926 , year =
-
[40]
arXiv preprint arXiv:2505.09388 , year =
Qwen3 Technical Report , author =. arXiv preprint arXiv:2505.09388 , year =
-
[41]
arXiv preprint arXiv:2503.19786 , year =
Gemma 3 Technical Report , author =. arXiv preprint arXiv:2503.19786 , year =
-
[42]
arXiv preprint arXiv:2502.02737 , year =
Allal, Loubna Ben and Lozhkov, Anton and Bakouch, Elie and Bl. arXiv preprint arXiv:2502.02737 , year =
-
[43]
arXiv preprint arXiv:2407.21783 , year =
The Llama 3 Herd of Models , author =. arXiv preprint arXiv:2407.21783 , year =
-
[44]
arXiv preprint arXiv:2404.14219 , year =
Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone , author =. arXiv preprint arXiv:2404.14219 , year =
-
[45]
Murthy, Rithesh and Yang, Liangwei and Tan, Juntao and Awalgaonkar, Tulika Manoj and Zhou, Yilun and Heinecke, Shelby and Desai, Sachin and Wu, Jason and Xu, Ran and Tan, Sarah and Zhang, Jianguo and Liu, Zhiwei and Kokane, Shirley and Liu, Zuxin and Zhu, Ming and Wang, Huan a...
-
[46]
2024 , note =
Screen your calls before you answer them , author =. 2024 , note =
2024
-
[47]
2025 , note =
Screen and block calls on iPhone , author =. 2025 , note =
2025
-
[48]
2026 , note =
Speech-to-Speech Models and Providers Analysis (Big Bench Audio) , author =. 2026 , note =
2026
-
[49]
2026 , note =
Voice agent evaluation framework: 6 pillars explained , author =. 2026 , note =
2026
-
[50]
2025 , note =
How to evaluate. 2025 , note =
2025
-
[51]
2026 , note =
Voice. 2026 , note =
2026
-
[52]
2026 , note =
How to Evaluate Voice Agents: A Complete Framework for Testing , author =. 2026 , note =
2026
-
[53]
Findings of the Association for Computational Linguistics: ACL 2023 , year =
Discovering Language Model Behaviors with Model-Written Evaluations , author =. Findings of the Association for Computational Linguistics: ACL 2023 , year =
2023
-
[54]
International Conference on Learning Representations (ICLR) , year =
Towards Understanding Sycophancy in Language Models , author =. International Conference on Learning Representations (ICLR) , year =
-
[55]
Transactions on Machine Learning Research , year =
Inverse Scaling: When Bigger Isn't Better , author =. Transactions on Machine Learning Research , year =
-
[56]
How Johnny Can Persuade
Zeng, Yi and Lin, Hongpeng and Zhang, Jingwen and Yang, Diyi and Jia, Ruoxi and Shi, Weiyan , booktitle =. How Johnny Can Persuade. 2024 , note =
2024
-
[57]
Not What You've Signed Up For: Compromising Real-World
Greshake, Kai and Abdelnabi, Sahar and Mishra, Shailesh and Endres, Christoph and Holz, Thorsten and Fritz, Mario , journal =. Not What You've Signed Up For: Compromising Real-World. 2023 , note =
2023
-
[58]
Mireshghallah, Niloofar and Kim, Hyunwoo and Zhou, Xuhui and Tsvetkov, Yulia and Sap, Maarten and Shokri, Reza and Choi, Yejin , booktitle =. Can. 2024 , note =
2024
-
[59]
2024 , note =
Shao, Yijia and Li, Tianshi and Shi, Weiyan and Liu, Yanchen and Yang, Diyi , booktitle =. 2024 , note =
2024
-
[60]
Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =
Can Large Language Models Be an Alternative to Human Evaluations? , author =. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =
-
[61]
and Feng, Shi , journal =
Panickssery, Arjun and Bowman, Samuel R. and Feng, Shi , journal =
-
[62]
Transactions on Machine Learning Research , year =
Holistic Evaluation of Language Models , author =. Transactions on Machine Learning Research , year =
-
[63]
Beyond Accuracy: Behavioral Testing of
Ribeiro, Marco Tulio and Wu, Tongshuang and Guestrin, Carlos and Singh, Sameer , booktitle =. Beyond Accuracy: Behavioral Testing of. 2020 , note =
2020
-
[64]
and Fitz, Stephen and Hendrycks, Dan , booktitle =
Ren, Richard and Basart, Steven and Khoja, Adam and Gatti, Alice and Phan, Long and Yin, Xuwang and Mazeika, Mantas and Pan, Alexander and Mukobi, Gabriel and Kim, Ryan H. and Fitz, Stephen and Hendrycks, Dan , booktitle =. Safetywashing: Do. 2024 , note =
2024
-
[65]
Washington Law Review , volume =
Privacy as Contextual Integrity , author =. Washington Law Review , volume =
-
[66]
Li, Margaret and Weston, Jason and Roller, Stephen , journal =
-
[67]
Laban, Philippe and Hayashi, Hiroaki and Zhou, Yingbo and Neville, Jennifer , journal =
-
[68]
The Instruction Hierarchy: Training
Wallace, Eric and Xiao, Kai and Leike, Reimar and Weng, Lilian and Heidecke, Johannes and Beutel, Alex , journal =. The Instruction Hierarchy: Training
-
[69]
and Wang, Di , booktitle =
Yang, Shu and Zhu, Shenzhe and Wu, Zeyu and Wang, Keyu and Yao, Junchi and Wu, Junchao and Hu, Lijie and Li, Mengdi and Wong, Derek F. and Wang, Di , booktitle =. 2025 , note =
2025
-
[70]
2025 , note =
Apate.ai:. 2025 , note =
2025
-
[71]
and Zhou, Y
Wu, J. and Zhou, Y. and Ng, D. T. and Shen, X. and Zewde, K. and Raj, A. and Duong, T. and Ren, Simiao , year =. When the Forger Is the Judge:. 2604.25213 , archivePrefix =
-
[72]
and Ren, Simiao and Raj, A
Zhang, Y. and Ren, Simiao and Raj, A. and Wei, E. and Ng, D. and Shen, A. and Xue, J. and Zhang, Y. and Marotta, E. , year =. 2603.11442 , archivePrefix =
-
[73]
and Shen, X
Ren, Simiao and Zhou, Y. and Shen, X. and Zewde, K. and Duong, T. and Huang, G. and Wei, E. and Xue, J. , year =. How Well Are Open-Sourced. 2602.07814 , archivePrefix =
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.