REVIEW 3 major objections 4 minor 69 references
This paper argues that the standard KL-regularized RL objective is the equilibrium of a sequential detection game, so the regularization coefficient is fixed by the reward and reference policy rather than hand-tuned.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 17:27 UTC pith:RZV247E7
load-bearing objection A clean game-theoretic derivation of KL-regularized RL with a real interpretation for β*, but the calibration margin ρ quietly reintroduces a free hyperparameter and the experiments outrun the theory. the 3 major comments →
Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The discovery is that the policy pair consisting of an agent playing πβ*(y|x) ∝ πref(y|x) exp(r(x,y)/β*) and a monitor running a sequential probability ratio test against exactly that tilted policy is a Nash equilibrium of the sequential detection game, interpreted at leading order as the false-positive level α tends to 0. The equilibrium coefficient β* is the unique maximizer of expected reward per unit of KL divergence to the reference policy, equivalently the unique root of the function M(β) = E[r] − β E[KL]. The standard KL-regularized RL objective is therefore recovered from first principles, and the KL penalty gains an operational meaning: it is the price an agent pays to remain statis
What carries the argument
The central object is the fractional program that maximizes expected reward divided by expected KL divergence to the reference, which emerges after replacing the monitor's expected detection time with its information-theoretic lower bound. The argument is carried by the scalar function M(β) = E_{πβ}[r] − β·E[KL(πβ∥πref)]: β* is its unique root, and classical fractional-programming theory shows the policy maximizing that ratio is the KL-regularized tilt at the root. The monitor's side is carried by the sequential probability ratio test, which is the leading-order optimal test against the tilted alternative.
Load-bearing premise
The load-bearing premise is that the agent's utility is well approximated by reward divided by KL through the leading-order lower bound on sequential detection time, and that each KL-regularized subproblem can be solved as if by an exact oracle; both assumptions weaken at practical confidence levels and with approximate LLM optimizers.
What would settle it
Construct a small environment with a known policy class and reward, fix α at a practical level such as 0.05, and compute the agent's exact best response in the sequential detection game by simulation or dynamic programming, then compare it with πβ*. If the finite-α best response differs materially from the tilt—or its realized reward per unit detection time exceeds the equilibrium value—the leading-order equilibrium claim is falsified at that confidence level.
If this is right
- If the equilibrium characterization is right, the KL penalty in RL fine-tuning is not a convenience term: it is the statistical price of evading sequential detection, so β is a shadow price rather than a tuning knob.
- Because β* is determined by the reward, prompt distribution, and reference policy, practitioners can estimate it by stochastic bisection using O(log(1/ε)) RL subproblems, replacing grid search with an adaptive procedure.
- The monitor's equilibrium strategy—the SPRT against πβ*—controls type-I error and has power one, so auditors can use it to detect hidden fine-tuning while keeping false positives at the advertised level.
- In the reported experiments, bisection-selected policies land near the elbow of the reward–retention frontier in continual learning, achieving competitive trade-offs without a post-hoc model-selection step.
- If the agent and monitor each run the bisection algorithm separately, the resulting profile is an approximate equilibrium, so the same framework can be used by both sides without coordination.
Where Pith is reading between the lines
- Editorial inference: The root-finding view suggests a natural online variant in which β is updated continuously toward the sign of M(β), turning fine-tuning into a self-balancing procedure that tracks the reward–retention frontier as the reward model changes.
- Editorial inference: The auditing story assumes the auditor knows the reward being optimized; coupling this framework with reward learning or inverse RL could let auditors detect strategic fine-tuning when the served model's objective is only partially known—a next step the paper itself flags.
- Editorial inference: At practical α values such as 0.05, the leading-order theory may need finite-sample corrections; a testable extension would compare the exact equilibrium at fixed α against πβ* in a small control environment and, if the gap is material, replace expected stopping times by boundary-corrected versions.
- Editorial inference: The same tilt machinery could be applied to adjacent uses the authors did not test, such as tuning the coefficient that balances a model's utility against its resistance to distillation, replacing manual coefficient choice in that setting.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a game-theoretic interpretation of KL-regularized RL fine-tuning. It models a sequential detection game in which an agent chooses a policy to maximize cumulative reward before a monitor can detect a deviation from a reference policy, while the monitor controls type-I error and minimizes expected detection time. The authors argue that, in the high-confidence regime (α→0), the agent's equilibrium policy is the KL-regularized tilt π_β* ∝ π_ref exp(r/β*), where β* maximizes reward per unit of KL divergence to the reference policy; the monitor's best response is an SPRT. They use Dinkelbach fractional programming to characterize β* as the unique root of M(β)=E_πβ[r]-β·KL(πβ||π_ref), and present a stochastic bisection algorithm (Algorithm 1) that estimates β* by solving O(log(1/ε)) RL subproblems. Experiments on Qwen3-8B and Llama-3.2-1B in continual-learning and model-auditing settings are reported as preliminary evidence. The paper also presents robustness results for tie-breaking, approximate equilibria, and confidence-sequence instantiations of the bisection procedure.
Significance. If the central derivation holds, the paper gives an operational meaning to the KL penalty in RL fine-tuning: β is the shadow price of statistical detectability under sequential monitoring, rather than a purely heuristic hyperparameter. The reduction to Dinkelbach fractional programming is elegant, and the oracle complexity of O(log(1/ε)) RL calls is attractive. The auditing application is concrete and falsifiable, and the authors are careful to state several important caveats, notably the high-confidence interpretation and the exact-RL-oracle assumption. However, the paper's main selling point—that β* is determined endogenously rather than by hyperparameter search—is currently undermined by an unmodeled calibration margin ρ that shifts β* in the experiments, and the practical algorithm used in the experiments is not covered by Theorem 4.1. These issues are load-bearing and need to be addressed before the claims can be accepted.
major comments (3)
- [Section 5; Eq. (4) and Lemma A.1] The reward calibration step shifts the reward by c = -μ_ref - ρ before bisection. The statement that 'adding a constant to the reward does not change the optimal policy at any fixed β' is correct for the tilt πβ, but the equilibrium coefficient β* is not invariant under this shift. The fractional program (4) becomes max_π (E_π[r]+c)/KL(π||π_ref), i.e., M_c(β)=M(β)+c, so the unique root characterized by Lemma A.1 moves. The paper's own results confirm this: for Llama-3.2-1B, β* ≈ 0.0838 for ρ=0.1 and β* ≈ 0.0603 for ρ=0.2 (Appendix C.1). Thus ρ is a free hyperparameter that determines β*, contradicting the abstract and introduction's claim that β* is determined solely by the reward function, prompt distribution, and reference policy. This should be modeled as part of the game (e.g., as the training-cost constant c in Assumption 3.1) or explicitly acknowledged and analyzed; at minimum, sen
- [Section 4, Algorithm 1 / Theorem 4.1] Theorem 4.1 requires a confidence sequence rad(β,n,δ) that is valid simultaneously for all n, and it assumes an exact RL oracle. The experiments explicitly 'forego the confidence sequence' and use a fixed n=4096, and the RL oracle is GRPO, an approximate solver. Therefore the main theoretical result does not cover the empirical bisection runs, and the claim that Algorithm 1 is a principled replacement for β-search in LLM fine-tuning is not established by the paper's analysis. The authors flag this as a 'practical relaxation,' but the scope of the claim should be narrowed, or a finite-sample/empirical validation of the sign test should be provided.
- [Section 3, Eq. (4) and Section 5, Table 1] The equilibrium characterization is leading-order in α↓0; at the practical level α=0.05 used in the auditing experiments, SPRT overshoot and discrete-time boundary effects can make the exact optimal policy differ from πβ*. The caveat is stated in Section 3, but the abstract and introduction present the result as an unqualified Nash equilibrium. Since the experiments run at α=0.05, this gap is directly relevant to the empirical claims. A small-scale finite-state simulation comparing πβ* to the exact solution of the sequential game at α=0.05 would quantify the discrepancy; otherwise all equilibrium claims should be qualified consistently.
minor comments (4)
- [Algorithm 1] The pseudocode initializes n=100, while the text says to collect 'enough samples' n such that the confidence interval is bounded away from zero. The relationship between n and rad(·,·,·) should be made explicit in the pseudocode to avoid confusion.
- [Section 5.1, Figures 1-3] The labels 'Agent' and 'Monitor' in the continual-learning plots are inherited from the auditing experiment; in this experiment they denote two independent runs of the same bisection procedure, which is confusing. A neutral label such as 'bisection run 1/2' would be clearer.
- [Section 2, Assumption 3.1] Assumption 3.1 introduces c as the minimum improvement in expected reward required for the agent to update the reference policy, but c is not used again in the paper. Its relation to the calibration margin ρ in Section 5 should be clarified, since ρ appears to play the same role.
- [Section 4, proof of Lemma B.3] The proof of Lemma B.3 uses the Donsker-Varadhan formula with f = λ(r - μ_ref) and then minimizes over λ. The step is standard, but the notation 'P' and 'Q' is reused from the sequential-testing section without explicit restatement; a one-line reminder would improve readability.
Circularity Check
No circularity: the KL-regularized equilibrium is derived from the sequential-detection lower bound, and the ρ calibration shift is a hyperparameter-dependence concern, not a circular reduction.
full rationale
The central derivation is not circular. In Section 3, the agent's objective (4) is obtained from the sequential detection game by replacing the expected stopping time with the Robbins–Siegmund lower bound log(1/α)/KL(π∥πref), so the KL denominator enters through the monitor's detection power, not through an assumed KL penalty in the agent's utility. The subsequent identification of the maximizer of (4) with the KL-regularized tilt πβ* is a classical concave–convex fractional programming equivalence (Dinkelbach 1967; Schaible 1976), explicitly stated in Theorem 3.3 and used via Lemma A.1; this is a mathematical reduction, not a hidden assumption or a fitted prediction. The paper's self-citations (Harris et al. 2021; Gauthier et al. 2026; Capitaine et al. 2026) appear in related-work discussion and are not load-bearing for the equilibrium characterization or the bisection algorithm. The most substantial caveat is empirical, not circular: in Section 5, the reward is shifted by an arbitrary calibration margin ρ, and the paper's own numbers show β* ≈ 0.0838 for ρ = 0.1 versus β* ≈ 0.0603 for ρ = 0.2, so the claim that β* is 'determined solely by the reward function, prompt distribution, and reference policy' is overstated. But this is a dependence on an unmodeled input/hyperparameter, not a reduction of the prediction to its own input; for any fixed calibrated reward, the derivation is self-contained. The paper also explicitly flags its high-confidence α↓0 idealization and the RL-oracle assumption, which are acknowledged limitations rather than circular steps.
Axiom & Free-Parameter Ledger
free parameters (2)
- Calibration margin ρ =
0.1 (Qwen3-8B); 0.01, 0.1, 0.2 (Llama-3.2-1B)
- Fixed bisection sample count n =
4096
axioms (8)
- standard math Robbins–Siegmund lower bound E_Q[τ] ≥ log(1/α)/KL(Q,P), tight at leading order; SPRT is asymptotically optimal.
- domain assumption Assumption 3.1: some policy yields positive expected reward and μ_ref≤0.
- domain assumption Assumption 3.2: for every π'≠πref there is a test with finite expected stopping time.
- standard math Wald's equation: E[Σ_{t=1}^τ r_t] = E[τ] E[r].
- standard math Dinkelbach/Schaible fractional programming equivalence and convexity of M(β).
- domain assumption High-confidence regime α↓0: discrete-time overshoot terms vanish at leading log(1/α) order.
- domain assumption Exact RL oracle returns the tilt πβ for any β (Algorithm 1/Theorem 4.1).
- domain assumption Rewards are σ-sub-Gaussian and reference token probabilities are bounded below by γ>0 (Assumptions B.1, B.4/B.6).
read the original abstract
Reinforcement learning (RL) fine-tuning is widely used in language model training to improve model performance on a target task while limiting drift from a reference policy. A standard way to balance this trade-off is via a KL-regularized RL objective, although this formulation does not by itself provide a principled way to set the regularization coefficient. In practice, the coefficient is typically chosen heuristically or via hyperparameter search, which can lead to unnecessary overhead in training cost or undesirable reward-retention trade-offs. We instead propose a game-theoretic framework that gives this trade-off an explicit statistical interpretation. Specifically, we study a sequential game in which an agent chooses a policy to maximize cumulative reward while a monitor observes policy outputs over time and tests for deviations from the reference policy. Although not originating from the same perspective, we show that the resulting equilibrium policy can nonetheless be expressed as the solution to a KL-regularized RL problem for an optimal regularization parameter that can be viewed as maximizing reward per unit of statistical distinguishability. Drawing on classical results from concave-convex fractional programming, we provide a principled method for learning this equilibrium coefficient via reduction to the KL-regularized RL objective, thus allowing for flexible integration into standard fine-tuning pipelines. In experiments with Qwen3-8B and Llama-3.2-1B, we demonstrate that our methods result in competitive reward-retention trade-offs in a continual learning setting, and illustrate how our framework may be used to audit API providers serving open-source models.
Figures
Reference graph
Works this paper leans on
-
[1]
Agrawal, Shubhada and Ramdas, Aaditya , journal=
-
[2]
1967 , publisher=
Dinkelbach, Werner , journal=. 1967 , publisher=
1967
-
[3]
1976 , publisher=
Schaible, Siegfried , journal=. 1976 , publisher=
1976
-
[4]
1985 , publisher=
Crouzeix, JP and Ferland, JA and Schaible, S , journal=. 1985 , publisher=
1985
-
[5]
2021 , publisher=
Howard, Steven R and Ramdas, Aaditya and McAuliffe, Jon and Sekhon, Jasjeet , journal=. 2021 , publisher=
2021
-
[6]
2017 , publisher=
Kirkpatrick, James and Pascanu, Razvan and Rabinowitz, Neil and Veness, Joel and Desjardins, Guillaume and Rusu, Andrei A and Milan, Kieran and Quan, John and Ramalho, Tiago and Grabska-Barwinska, Agnieszka and others , journal=. 2017 , publisher=
2017
-
[7]
Zhang, Han and Lei, Yu and Gui, Lin and Yang, Min and He, Yulan and Wang, Hui and Xu, Ruifeng , booktitle=
-
[8]
Zhang, Han and Gui, Lin and Zhai, Yuanzhao and Wang, Hui and Lei, Yu and Xu, Ruifeng , journal=
-
[9]
Mohammadzadeh, Shahrad and Chmura, Jacob and Anokhin, Ivan and Tian, Jacob-Junqi and Samiei, Mandana and Scott-Talib, Taz and Rish, Irina and Precup, Doina and Rabbany, Reihaneh and Anand, Nishanth , booktitle=
-
[10]
Gauthier, Etienne and Bach, Francis and Jordan, Michael I , journal=
-
[11]
Capitaine, Aymeric and Scheid, Antoine and Boursier, Etienne and Durmus, Alain and Jordan, Michael I , journal=
-
[12]
2024 , publisher=
Hu, Yinan and Chen, Juntao and Zhu, Quanyan , journal=. 2024 , publisher=
2024
-
[13]
Ziegler, Daniel M and Stiennon, Nisan and Wu, Jeffrey and Brown, Tom B and Radford, Alec and Amodei, Dario and Christiano, Paul and Irving, Geoffrey , journal=
-
[14]
Rafailov, Rafael and Sharma, Archit and Mitchell, Eric and Manning, Christopher D and Ermon, Stefano and Finn, Chelsea , journal=
-
[15]
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric and others , journal=
-
[16]
Donsker, Monroe D and Varadhan, SR Srinivasa , journal=
-
[17]
2019 , organization=
Geist, Matthieu and Scherrer, Bruno and Pietquin, Olivier , booktitle=. 2019 , organization=
2019
-
[18]
arXiv preprint arXiv:2405.08448 , year=
Tang, Yunhao and Guo, Daniel Zhaohan and Zheng, Zeyu and Calandriello, Daniele and Cao, Yuan and Tarassov, Eugene and Munos, R. arXiv preprint arXiv:2405.08448 , year=
-
[19]
arXiv preprint arXiv:1705.07798 , year=
Neu, Gergely and Jonsson, Anders and G. arXiv preprint arXiv:1705.07798 , year=
-
[20]
Singh, Aaditya and Fry, Adam and Perelman, Adam and Tart, Adam and Ganesh, Adi and El-Kishky, Ahmed and McLaughlin, Aidan and Low, Aiden and Ostrow, AJ and Ananthram, Akhila and others , journal=
-
[21]
Stiennon, Nisan and Ouyang, Long and Wu, Jeffrey and Ziegler, Daniel and Lowe, Ryan and Voss, Chelsea and Radford, Alec and Amodei, Dario and Christiano, Paul F , journal=
-
[22]
Ouyang, Long and Wu, Jeffrey and Jiang, Xu and Almeida, Diogo and Wainwright, Carroll and Mishkin, Pamela and Zhang, Chong and Agarwal, Sandhini and Slama, Katarina and Ray, Alex and others , journal=
-
[23]
Touvron, Hugo and Martin, Louis and Stone, Kevin and Albert, Peter and Almahairi, Amjad and Babaei, Yasmine and Bashlykov, Nikolay and Batra, Soumya and Bhargava, Prajjwal and Bhosale, Shruti and others , journal=
-
[24]
Lin, Yong and Lin, Hangyu and Xiong, Wei and Diao, Shizhe and Liu, Jianmeng and Zhang, Jipeng and Pan, Rui and Wang, Haoxiang and Hu, Wenbin and Zhang, Hanning and others , booktitle=
-
[25]
Noukhovitch, Michael and Lavoie, Samuel and Strub, Florian and Courville, Aaron C , journal=
-
[26]
Yang, An and Li, Anfeng and Yang, Baosong and Zhang, Beichen and Hui, Binyuan and Zheng, Bo and Yu, Bowen and Gao, Chang and Huang, Chengen and Lv, Chenxu and others , journal=
-
[27]
2024 , howpublished =
2024
-
[28]
International Conference on Machine Learning , pages=
Jaques, Natasha and Gu, Shixiang and Bahdanau, Dzmitry and Hern. International Conference on Machine Learning , pages=. 2017 , organization=
2017
-
[29]
Korbak, Tomasz and Perez, Ethan and Buckley, Christopher , booktitle=
-
[30]
Bai, Yuntao and Jones, Andy and Ndousse, Kamal and Askell, Amanda and Chen, Anna and DasSarma, Nova and Drain, Dawn and Fort, Stanislav and Ganguli, Deep and Henighan, Tom and others , journal=
-
[31]
2023 , organization=
Gao, Leo and Schulman, John and Hilton, Jacob , booktitle=. 2023 , organization=
2023
-
[32]
1992 , publisher=
Wald, Abraham , booktitle=. 1992 , publisher=
1992
-
[33]
1948 , publisher=
Wald, Abraham and Wolfowitz, Jacob , journal=. 1948 , publisher=
1948
-
[34]
1951 , publisher=
Robbins, Herbert and Monro, Sutton , journal=. 1951 , publisher=
1951
-
[35]
2013 , publisher=
Waeber, Rolf and Frazier, Peter I and Henderson, Shane G , journal=. 2013 , publisher=
2013
-
[36]
2017 , publisher=
Li, Zhizhong and Hoiem, Derek , journal=. 2017 , publisher=
2017
-
[37]
2017 , organization=
Zenke, Friedemann and Poole, Ben and Ganguli, Surya , booktitle=. 2017 , organization=
2017
-
[38]
Lopez-Paz, David and Ranzato, Marc'Aurelio , journal=
-
[39]
Sun, Fan-Keng and Ho, Cheng-Hao and Lee, Hung-Yi , journal=
-
[40]
Razdaibiedina, Anastasia and Mao, Yuning and Hou, Rui and Khabsa, Madian and Lewis, Mike and Almahairi, Amjad , journal=
-
[41]
Wu, Tongtong and Luo, Linhao and Li, Yuan-Fang and Pan, Shirui and Vu, Thuy-Trang and Haffari, Gholamreza , journal=
-
[42]
2025 , publisher=
Shi, Haizhou and Xu, Zihao and Wang, Hengyi and Qin, Weiyi and Wang, Wenyuan and Wang, Yibin and Wang, Zifeng and Ebrahimi, Sayna and Wang, Hao , journal=. 2025 , publisher=
2025
-
[43]
Gao, Irena and Liang, Percy and Guestrin, Carlos , booktitle=
-
[44]
Richter, Leo and He, Xuanli and Minervini, Pasquale and Kusner, Matt , booktitle=
-
[45]
Raji, Inioluwa Deborah and Smart, Andrew and White, Rebecca N and Mitchell, Margaret and Gebru, Timnit and Hutchinson, Ben and Smith-Loud, Jamila and Theron, Daniel and Barnes, Parker , booktitle=
-
[46]
AI and Ethics , volume=
M. AI and Ethics , volume=. 2024 , publisher=
2024
-
[47]
Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=
Casper, Stephen and Ezell, Carson and Siegmann, Charlotte and Kolt, Noam and Curtis, Taylor Lynn and Bucknall, Benjamin and Haupt, Andreas and Wei, Kevin and Scheurer, J. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=
2024
-
[48]
Hardt, Moritz and Megiddo, Nimrod and Papadimitriou, Christos and Wootters, Mary , booktitle=
-
[49]
Dong, Jinshuo and Roth, Aaron and Schutzman, Zachary and Waggoner, Bo and Wu, Zhiwei Steven , booktitle=
-
[50]
2020 , publisher=
Kleinberg, Jon and Raghavan, Manish , journal=. 2020 , publisher=
2020
-
[51]
International Conference on Machine Learning , pages=
Perdomo, Juan and Zrnic, Tijana and Mendler-D. International Conference on Machine Learning , pages=. 2020 , organization=
2020
-
[52]
2022 , organization=
Brown, Gavin and Hod, Shlomi and Kalemaj, Iden , booktitle=. 2022 , organization=
2022
-
[53]
Harris, Keegan and Heidari, Hoda and Wu, Steven Z , journal=
-
[54]
Schulman, John and Wolski, Filip and Dhariwal, Prafulla and Radford, Alec and Klimov, Oleg , journal=
-
[55]
Shao, Zhihong and Wang, Peiyi and Zhu, Qihao and Xu, Runxin and Song, Junxiao and Bi, Xiao and Zhang, Haowei and Zhang, Mingchuan and Li, YK and Wu, Yang and others , journal=
-
[56]
Savani, Yash and Trockman, Asher and Feng, Zhili and Xu, Yixuan and Schwarzschild, Avi and Robey, Alexander and Finzi, Marc and Kolter, Zico , journal=
-
[57]
Allouah, Youssef and Haghifam, Mahdi and Koyejo, Sanmi and Shokri, Reza , journal=
-
[58]
Hu, Edward J and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Liang and Chen, Weizhu and others , journal=
-
[59]
Kingma, Diederik P and Ba, Jimmy , journal=
-
[60]
2025 , howpublished =
2025
-
[61]
Velasco, Ander Artola and Rontogiannis, Dimitrios and Tsirtsis, Stratis and Gomez-Rodriguez, Manuel , journal=
-
[62]
Cao, Yuhan and Wang, Yu and Liu, Sitong and Li, Miao and Tao, Yixin and He, Tianxing , booktitle=
-
[63]
1945 , publisher=
Wald, Abraham , journal=. 1945 , publisher=
1945
-
[64]
Abraham Wald , year=
-
[65]
1974 , publisher=
Robbins, Herbert and Siegmund, David , journal=. 1974 , publisher=
1974
-
[66]
2023 , organization=
Zhang, Tianjun and Liu, Fangchen and Wong, Justin and Abbeel, Pieter and Gonzalez, Joseph E , booktitle=. 2023 , organization=
2023
-
[67]
1985 , isbn =
David Siegmund , title =. 1985 , isbn =
1985
-
[68]
Jaques, Natasha and Ghandeharioun, Asma and Shen, Judy Hanwen and Ferguson, Craig and Lapedriza, Agata and Jones, Noah and Gu, Shixiang and Picard, Rosalind , journal=
-
[69]
Jang, Youngsoo and Kim, Geon-Hyeong and Kim, Byoungjip and Kim, Yu Jin and Lee, Honglak and Lee, Moontae , booktitle=
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.