REVIEW 3 major objections 2 minor 1 cited by
The largest generalised eigenvalue of the Jacobian pencil between proxy and target embeddings governs worst-case local displacement under semantic paraphrase perturbations.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-26 18:48 UTC pith:RY5RYK5J
load-bearing objection The attackability index from the largest generalized eigenvalue of the Jacobian pencil is the actual new piece, but it rests on an unverified continuous proxy approximation for semantic neighborhoods. the 3 major comments →
Generalised Eigenvalue Geometry of Semantic Adversarial Attacks
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
We develop a continuous local model of semantic paraphrase perturbations that captures this two-model structure. We show that the worst-case local displacement of the target representation, subject to a proxy-model budget, is governed by the largest generalised eigenvalue of a matrix pencil (A,B) constructed from the Jacobians of the two embedding maps. The resulting attackability index λ*(x) is intrinsic to the local paraphrase geometry and the chosen embedders, yields a closed-form prediction-flip condition for affine readouts, and supports conservative population and finite-sample attackability certificates.
What carries the argument
the largest generalised eigenvalue of the matrix pencil (A,B) assembled from the Jacobians of the proxy and target embedding maps, which defines the attackability index λ*(x)
Load-bearing premise
Semantic paraphrase perturbations admit a continuous local approximation whose allowed displacements are adequately captured by a budget constraint on a separate proxy embedding map.
What would settle it
A dataset of generated paraphrases in which the observed maximum displacement of the target representation systematically exceeds the value predicted by the largest generalised eigenvalue at many input points would falsify the governing role of that eigenvalue.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a continuous local model of semantic paraphrase perturbations in a two-embedding setting (proxy and target). It claims that the worst-case local displacement of the target representation under a proxy budget is given by the largest generalized eigenvalue of the Jacobian pencil (A,B), yielding an intrinsic attackability index λ*(x), a closed-form prediction-flip condition for affine readouts, conservative population and finite-sample certificates, a distribution-free VC bound on binary attackability indicators, a scale-sensitive margin bound, and a covering condition linking the continuous and discrete paraphrase-search regimes. An empirical verification framework using soft-token relaxations is also proposed.
Significance. If the modeling assumptions and derivations hold, the work supplies a geometric, largely parameter-free characterization of semantic attackability together with explicit certificates and bounds; this would constitute a substantive theoretical contribution to multi-model robustness analysis in NLP, particularly for applications such as financial sentiment classification.
major comments (3)
- [Abstract / continuous local model] Abstract and continuous-local-model section: the central claim that worst-case target displacement is governed by the largest generalized eigenvalue of the pencil constructed from the two Jacobians is asserted without any derivation steps, intermediate lemmas, or error analysis; all subsequent results (attackability index, closed-form flip condition, VC bounds) rest on this unshown step.
- [Abstract] Abstract: the modeling premise that semantic paraphrases admit a first-order continuous approximation whose feasible set is exactly the ellipsoid defined by the proxy embedding budget is invoked to obtain λ*(x) and the certificates, yet no justification, tightness argument, or analysis of higher-order remainder terms is supplied; if the proxy map fails to tightly contain the true semantic neighborhood, the eigenvalue no longer upper-bounds actual attackability.
- [VC and margin bounds] VC-bound and margin-bound paragraphs: the distribution-free VC bound and the attackability-adjusted margin bound are stated as consequences of the eigenvalue geometry, but without the explicit derivation of λ*(x) it is impossible to verify that the geometric penalty term is correctly subtracted or that the finite-sample certificate remains conservative.
minor comments (2)
- [Abstract] Notation for the matrix pencil (A,B) and the attackability index λ*(x) should be introduced with an explicit equation number at first use.
- [Empirical verification framework] The empirical verification framework is described at a high level; concrete details on the soft-token relaxation and the generated paraphrase sets would improve reproducibility.
Simulated Author's Rebuttal
We thank the referee for the thorough review and valuable comments. We address each of the major comments below, agreeing that additional explicit derivations and justifications are needed to strengthen the presentation. We will make the requested revisions to the manuscript.
read point-by-point responses
-
Referee: [Abstract / continuous local model] Abstract and continuous-local-model section: the central claim that worst-case target displacement is governed by the largest generalized eigenvalue of the pencil constructed from the two Jacobians is asserted without any derivation steps, intermediate lemmas, or error analysis; all subsequent results (attackability index, closed-form flip condition, VC bounds) rest on this unshown step.
Authors: We acknowledge that the derivation steps for the generalized eigenvalue characterization were not presented with sufficient intermediate lemmas in the continuous local model section. The manuscript derives the result from the constrained optimization problem max ||J_t δ|| subject to ||J_p δ|| ≤ ε, which reduces to the generalized eigenvalue problem via the Lagrangian, but we agree this needs to be expanded. In the revision, we will insert explicit lemmas detailing the steps, including the equivalence to the largest generalized eigenvalue λ* and a first-order error bound. This will support all subsequent claims. revision: yes
-
Referee: [Abstract] Abstract: the modeling premise that semantic paraphrases admit a first-order continuous approximation whose feasible set is exactly the ellipsoid defined by the proxy embedding budget is invoked to obtain λ*(x) and the certificates, yet no justification, tightness argument, or analysis of higher-order remainder terms is supplied; if the proxy map fails to tightly contain the true semantic neighborhood, the eigenvalue no longer upper-bounds actual attackability.
Authors: The first-order approximation is motivated in the introduction and related work sections by the local linearity of embedding maps around x. However, we agree that a dedicated justification and analysis of remainder terms is missing. We will add a new paragraph in the continuous local model section providing a Taylor expansion analysis of the remainder, discussing conditions for the ellipsoid to contain the semantic neighborhood, and citing relevant literature on local approximations in NLP embeddings. This will clarify when λ*(x) serves as a valid upper bound. revision: yes
-
Referee: [VC and margin bounds] VC-bound and margin-bound paragraphs: the distribution-free VC bound and the attackability-adjusted margin bound are stated as consequences of the eigenvalue geometry, but without the explicit derivation of λ*(x) it is impossible to verify that the geometric penalty term is correctly subtracted or that the finite-sample certificate remains conservative.
Authors: The VC bound and margin bound follow from applying standard results to the attackability indicator defined via λ*(x). Once the derivation of λ*(x) is expanded as planned, we will also elaborate on how the geometric penalty is incorporated into the margin and why the certificate is conservative (by using the upper bound property). We will include the explicit steps in the revision to make this verifiable. revision: yes
Circularity Check
No circularity: eigenvalue result follows from standard linear algebra on explicit modeling assumptions
full rationale
The paper derives the attackability index λ*(x) as the largest generalized eigenvalue of the pencil (A,B) built from the Jacobians of the two embedding maps. This is the direct mathematical solution to the worst-case displacement problem under the posited local linear approximation and proxy budget ellipsoid; it is not obtained by fitting to data or by self-referential definition. The continuous local model of semantic paraphrases is stated as an explicit premise enabling the Jacobian construction, but the subsequent eigenvalue claim and certificates are independent consequences of that premise rather than reductions to it. No self-citations are invoked as load-bearing for the core geometry, and no fitted input is relabeled as a prediction. The derivation is therefore self-contained against external linear-algebra benchmarks.
Axiom & Free-Parameter Ledger
axioms (2)
- domain assumption Embedding maps are differentiable so that Jacobians exist and the local linear approximation holds.
- domain assumption Semantic perturbations can be modeled as continuous displacements subject to a proxy-model budget constraint.
invented entities (1)
-
attackability index λ*(x)
no independent evidence
read the original abstract
Recent empirical work shows that semantically equivalent paraphrases can fool financial sentiment classifiers: although a paraphrase remains close to the original under a strong reference embedding, it may shift the target model's representation enough to change the predicted class. Existing robustness theory either assumes a single-model threat model or focuses mainly on empirical attack algorithms. We develop a continuous local model of semantic paraphrase perturbations that captures this two-model structure. We show that the worst-case local displacement of the target representation, subject to a proxy-model budget, is governed by the largest generalised eigenvalue of a matrix pencil $(A,B)$ constructed from the Jacobians of the two embedding maps. The resulting attackability index $\lambda^*(x)$ is intrinsic to the local paraphrase geometry and the chosen embedders, yields a closed-form prediction-flip condition for affine readouts, and supports conservative population and finite-sample attackability certificates. For uniform control over classes of affine readouts, we derive a distribution-free VC bound for binary attackability indicators and a scale-sensitive margin bound based on an attackability-adjusted margin that subtracts a local geometric penalty from the standard classifier margin. We also connect the continuous theory to discrete paraphrase search, identify an asymmetry between successful and unsuccessful finite searches, and give a covering condition under which the discrete and continuous settings agree. Finally, we propose an empirical verification framework using soft-token relaxations and generated paraphrase sets to assess the local eigenvalue geometry, prediction-flip condition, and finite-search approximation on a deployed financial-text classifier.
Figures
Forward citations
Cited by 1 Pith paper
-
Finite-Sample Coverage Audits for High-Recall Candidate Generation: Certification and Learning-Theoretic Design
Excluded-pool auditing is minimax rate-optimal for certifying missed relevant mass: any valid zero-miss certificate needs Ω(N0/m) excluded labels, and exact binomial/hypergeometric bounds achieve this rate.
Reference graph
Works this paper leans on
-
[1]
International Conference on Artificial Intelligence and Statistics , pages=
A manifold view of adversarial risk , author=. International Conference on Artificial Intelligence and Statistics , pages=. 2022 , organization=
2022
-
[2]
Advances in neural information processing systems , volume=
Pytorch: An imperative style, high-performance deep learning library , author=. Advances in neural information processing systems , volume=
-
[3]
2017 IEEE symposium on security and privacy (SP) , pages=
Towards evaluating the robustness of neural networks , author=. 2017 IEEE symposium on security and privacy (SP) , pages=. 2017 , organization=
2017
-
[4]
Journal of the Association for Information Science and Technology , volume=
Good debt or bad debt: Detecting semantic orientations in economic texts , author=. Journal of the Association for Information Science and Technology , volume=. 2014 , publisher=
2014
-
[5]
Sentence-bert: Sentence embeddings using siamese bert-networks , author=. Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP) , pages=
2019
-
[6]
FinBERT: Financial Sentiment Analysis with Pre-trained Language Models
Finbert: Financial sentiment analysis with pre-trained language models , author=. arXiv preprint arXiv:1908.10063 , year=
work page internal anchor Pith review Pith/arXiv arXiv 1908
-
[7]
Journal of Machine Learning Research , volume=
Automatic differentiation in machine learning: a survey , author=. Journal of Machine Learning Research , volume=
-
[8]
International Conference on Learning Representations , year =
Szegedy, Christian and Zaremba, Wojciech and Sutskever, Ilya and Bruna, Joan and Erhan, Dumitru and Goodfellow, Ian and Fergus, Rob , title =. International Conference on Learning Representations , year =
-
[9]
and Shlens, Jonathon and Szegedy, Christian , title =
Goodfellow, Ian J. and Shlens, Jonathon and Szegedy, Christian , title =. International Conference on Learning Representations , year =
-
[10]
International Conference on Learning Representations , year =
Madry, Aleksander and Makelov, Aleksandar and Schmidt, Ludwig and Tsipras, Dimitris and Vladu, Adrian , title =. International Conference on Learning Representations , year =
-
[11]
International Conference on Learning Representations , year =
Tsipras, Dimitris and Santurkar, Shibani and Engstrom, Logan and Turner, Alexander and Madry, Aleksander , title =. International Conference on Learning Representations , year =
-
[12]
Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages =
Alzantot, Moustafa and Sharma, Yash and Elgohary, Ahmed and Ho, Bo-Jhang and Srivastava, Mani and Chang, Kai-Wei , title =. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages =. 2018 , address =
2018
-
[13]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Jin, Di and Jin, Zhijing and Zhou, Joey Tianyi and Szolovits, Peter , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2020 , doi =
2020
-
[14]
Battle of Transformers: Adversarial Attacks on Financial Sentiment Models , journal =
Can T. Battle of Transformers: Adversarial Attacks on Financial Sentiment Models , journal =. 2026 , doi =
2026
-
[15]
Gradient-based Adversarial Attacks against Text Transformers , booktitle =
Guo, Chuan and Sablayrolles, Alexandre and J. Gradient-based Adversarial Attacks against Text Transformers , booktitle =. 2021 , address =
2021
-
[16]
International Conference on Learning Representations , year =
Jang, Eric and Gu, Shixiang and Poole, Ben , title =. International Conference on Learning Representations , year =
-
[17]
Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples
Papernot, Nicolas and McDaniel, Patrick and Goodfellow, Ian , title =. arXiv preprint arXiv:1605.07277 , year =
work page internal anchor Pith review Pith/arXiv arXiv
-
[18]
28th USENIX Security Symposium (USENIX Security 19) , pages =
Demontis, Ambra and Melis, Marco and Pintor, Maura and Jagielski, Matthew and Biggio, Battista and Oprea, Alina and Nita-Rotaru, Cristina and Roli, Fabio , title =. 28th USENIX Security Symposium (USENIX Security 19) , pages =
-
[19]
Ensemble Adversarial Training: Attacks and Defenses , booktitle =
Tram. Ensemble Adversarial Training: Attacks and Defenses , booktitle =. 2018 , note =
2018
-
[20]
and Chervonenkis, Alexey Ya
Vapnik, Vladimir N. and Chervonenkis, Alexey Ya. , title =. Theory of Probability and Its Applications , volume =. 1971 , doi =
1971
-
[21]
, title =
Valiant, Leslie G. , title =. Communications of the ACM , volume =. 1984 , doi =
1984
-
[22]
, title =
Anthony, Martin and Bartlett, Peter L. , title =. 1999 , note =
1999
-
[23]
Discrete Applied Mathematics , volume =
Anthony, Martin , title =. Discrete Applied Mathematics , volume =. 1995 , doi =
1995
-
[24]
Shalev-Shwartz, Shai and Ben-David, Shai , title =
-
[25]
, title =
Lee, John M. , title =
-
[26]
and Van Loan, Charles F
Golub, Gene H. and Van Loan, Charles F. , title =
-
[27]
The Annals of Mathematical Statistics , volume =
Dvoretzky, Aryeh and Kiefer, Jack and Wolfowitz, Jacob , title =. The Annals of Mathematical Statistics , volume =. 1956 , doi =
1956
-
[28]
The Annals of Probability , volume =
Massart, Pascal , title =. The Annals of Probability , volume =. 1990 , doi =
1990
-
[29]
The Annals of Statistics , volume =
Koltchinskii, Vladimir and Panchenko, Dmitry , title =. The Annals of Statistics , volume =. 2002 , doi =
2002
-
[30]
2018 , howpublished =
Khim, Justin and Loh, Po-Ling , title =. 2018 , howpublished =
2018
-
[31]
Proceedings of the 36th International Conference on Machine Learning , series =
Yin, Dong and Ramchandran, Kannan and Bartlett, Peter , title =. Proceedings of the 36th International Conference on Machine Learning , series =. 2019 , publisher =
2019
-
[32]
Proceedings of the 37th International Conference on Machine Learning , series =
Awasthi, Pranjal and Frank, Natalie and Mohri, Mehryar , title =. Proceedings of the 37th International Conference on Machine Learning , series =. 2020 , publisher =
2020
-
[33]
and Mendelson, Shahar , title =
Bartlett, Peter L. and Mendelson, Shahar , title =. Journal of Machine Learning Research , volume =
-
[34]
Algorithmic Learning Theory (ALT 2016) , publisher =
Maurer, Andreas , title =. Algorithmic Learning Theory (ALT 2016) , publisher =
2016
-
[35]
and Jerrum, Mark R
Goldberg, Paul W. and Jerrum, Mark R. , title =. Machine Learning , volume =
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.