{"id":"1c65403d-bf70-49b2-84e2-0f96a65a15d4","arxiv_id":"2505.21173","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Speech recognition filters are built from SO(3) rotations of two contrast-maximizing 3x3 templates, with claimed accuracy gains in low-noise phoneme tasks.","lead":"This paper proposes topology-inspired convolutional filters for speech recognition, generated by applying random 3D rotations to two base kernels. It reports better phoneme accuracy in low-noise conditions, but only as plots, without error bars, and the geometric claims contain apparent errors.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'OF' evaluated in Section 4.3 is not the Orthogonal Feature layer defined in Section 4.1; the paper replaces it with non-orthogonal [v,0,±v] kernels while keeping the name, so the headline performance claim is not supported.","rationale":"The reader's formal objection is valid and checkable: at the v1+v3=0 stratum the kernel has the form [v,0,-v], and the stabilizer of this point under Q·A is exactly the SO(2) of rotations about v. Hence the orbit is SO(3)/SO(2) ≅ S^2; the claimed SO(3)/(SO(2)⋊Z2) ≅ RP^2 fiber is false. I nevertheless think the more load-bearing problem for the paper's advertised central result is empirical: the architecture called OF in the success figures is not the OF layer of Definition 4.1. Section 4.3 explicitly relaxes orthogonality, switches to [v1,0,±v1] kernels, and announces it will keep the name OF; the figures showing ~70% and superiority are labelled Non-Orthogonal and are said to outperform the 'previous orthogonal components'. Therefore the abstract's claim 'our proposed OF layer achieves superior performance' is not supported by the experiments as described, even before considering the lack of error bars, the post-hoc balanced sampling, or the Section 5.1 normalization admission. A rerun with the exact Definition 4.1 layer would decide this cleanly. Since the reader already recommends REJECT and this analysis strengthens rather than redirects that rejection, the verdict should remain unchanged.","tokens_in":10844,"tokens_out":9371,"duration_ms":97930,"concrete_test":"Re-run the SpeechBox clean and SNR=20 phoneme classification using exactly Definition 4.1: initialize every kernel as α R M with M∈{±A1,±A2} and R=exp(θ_x L_x+θ_y L_y+θ_z L_z), θ_i∼N(0,σ^2), using the same architecture and training protocol as Section 4. Report mean±std over at least five seeds. Compare the resulting accuracy with Figure 5's ~70% and with KF, CF, and plain CNN. If the exact OF layer does not reproduce the claimed superiority, or matches only after adding the Section 4.3 non-orthogonal kernels, the abstract's attribution to the OF layer fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing problem is that the paper's central empirical claim is attributed to the wrong object. Definition 4.1 defines the OF layer as W_k = α R_k M_k with R_k∈SO(3) randomly sampled via the Lie-algebra exponential and M_k∈{±A1,±A2}. But Section 4.3 says: 'if we relax the orthogonality condition... we can obtain a more canonical set of convolution kernels' (4), namely Q[1 0 -1;1 0 -1;1 0 -1]/√6 and Q[1 0 1;1 0 1;1 0 1]/√6, i.e. vertical-stripe detectors [v1,0,±v1], and 'for simplicity' names these OF. The accuracy gain (~70%) shown in Figure 5 that carries the abstract's claimed superiority is for this non-orthogonal replacement, and Figures 6–9 are explicitly labelled 'Non-Orthogonal'. The paper also describes this family as outperforming the 'previous orthogonal components', i.e. the actual Definition 4.1 layer. Section 5.1 compounds the problem by admitting an omitted audio-normalization step is a plausible root cause of earlier contradictory results. Independently, the reader's stabilizer objection is correct: for [v,0,-v] the stabilizer under the left SO(3) action is SO(2), so the preimage is S^2, not RP^2, contradicting Section 3.2's bullet and Remark 3.12. But even if that were repaired, the practical claim would still not land, because the evaluated network is not the proposed OF layer.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a framework for constructing convolutional kernels for speech spectrograms from an SO(3) group action on a constrained matrix space M = {A in M3x3(R) : ||A||=1, v1+v2+v3=0}. It claims M is S^5, the quotient B=M/SO(3) is D^2 with a stratified fiber bundle whose fibers are RP^2, S^2, L(4,1), or SO(3) depending on the stratum. It then defines an 'Orthogonal Feature' (OF) layer by sampling rotations of two base kernels, and reports phoneme-recognition experiments on SpeechBox, TIMIT, and LJSpeech, plus word and image experiments, claiming superior accuracy in low-noise conditions.","tokens_in":11244,"tokens_out":13570,"duration_ms":147781,"significance":"If the theoretical and empirical claims were correct, the paper would offer a novel bridge between topological/geometric ideas and CNN kernel design, with concrete gains on speech tasks. The paper also makes a commendable attempt to test across three speech datasets and two additional domains, and it explicitly uses a balanced phoneme subset to address class imbalance. However, the theoretical core is not presently correct, and the headline empirical result is attributed to a different layer than the one proposed; the manuscript therefore cannot be accepted in its current form.","major_comments":[{"comment":"The claimed fiber for the v1+v3=0 stratum is incorrect. For the matrix A = (1/sqrt(6))[[1,0,-1],[1,0,-1],[1,0,-1]], the stabilizer under left multiplication by SO(3) is the subgroup of rotations about the diagonal axis, which is SO(2), not SO(2)⋊Z2. The preimage set displayed in Remark 3.12 is therefore the orbit S^2, not RP^2; indeed the displayed set itself is a 2-sphere. Hence the stratified fiber-bundle decomposition stated in Section 3.2 is not established, and the RP^2 claim is false.","section":"Section 3.2, Remark 3.12 and following bullet"},{"comment":"The asserted fiber SO(3)/Z2 is also unsupported. For a rank-2 matrix with x=y but v1 not parallel to v3, the only rotation Q with Qv1=v1 and Qv3=v3 is the identity, because v1 and v3 are independent and span a plane; the 180-degree rotation that swaps v1 and v3 does not fix the matrix since it exchanges the first and third columns. Thus the stabilizer is trivial and the orbit is SO(3), not L(4,1). The proposed stratification by x=y and z^2=xy does not match the actual orbit types, which are SO(3)/SO(2) for rank-1 matrices and SO(3) for rank-2 matrices.","section":"Section 3.2, bullet for the region x=y"},{"comment":"The claimed projection onto M is not well-defined because step (2) does not preserve the unit-norm constraint. For a general input A, the matrix Ã = QA - (lambda/3)11^T satisfies the zero-sum condition but its Frobenius norm is not necessarily 1, so Ã need not lie in M as defined in Definition 3.6. The theorem also leaves the normalization step unspecified, and the 'except when the three column vectors are identical' caveat is not linked to any obstruction in the rotation step. A corrected statement would need to either include normalization or weaken the claim.","section":"Theorem 3.9"},{"comment":"The network that carries the abstract's accuracy claim is not the OF layer defined in Section 4.1. Definition 4.1 defines kernels W_k = alpha * R_k * M_k with R_k sampled from SO(3) via the exponential map and M_k in {±A1, ±A2}. Section 4.3 then replaces this construction by non-orthogonal vertical-stripe kernels [v1,0,±v1], explicitly calls them 'Non-Orthogonal', and states that the name OF is retained 'for simplicity'. Since Figures 5-9 and the roughly 70% accuracy on SpeechBox refer to this replacement, the experimental results do not support the claim that the proposed Orthogonal Feature layer achieves superior phoneme recognition. The actual Definition 4.1 layer is only compared in Figure 4, where the reported advantage is marginal.","section":"Section 4.3, Eq. (4)"},{"comment":"The paper states that audio normalization was omitted in earlier implementations and identifies this omission as a 'plausible root cause' for earlier contradictory results. This admission applies to the preprocessing used in the Section 4 experiments, so the comparisons in Figures 4-9 cannot be taken as valid controlled evaluations of the proposed kernels. No corrected or re-run results are supplied for the earlier experimental groups.","section":"Section 5.1"},{"comment":"All accuracy comparisons are reported as single training/validation curves without error bars, confidence intervals, or repeated-seed statistics, and the experimental section does not report the values of alpha and sigma or the network training hyperparameters. The central empirical claim is a modest accuracy margin, so the absence of uncertainty quantification makes it impossible to assess whether the observed differences are significant.","section":"Figures 4-15"}],"minor_comments":[{"comment":"The equation for the normalization quotient is written as 'theta1(q) = q2' immediately after theta1 has already been defined; this appears to be a typographical error, and the informal quotient construction leading to the Klein bottle would benefit from a more explicit statement.","section":"Section 2.2"},{"comment":"The introductory paragraph refers to 'the convolutional neural network architecture discussed in this chapter is identical to the one in Chapter 6', but the manuscript has only five chapters; this suggests an unrevised draft.","section":"Chapter 5"},{"comment":"The paper uses 'OF' both for the orthogonal layer of Definition 4.1 and for the non-orthogonal replacement of Section 4.3; the notation should be disambiguated everywhere, including in figure captions, since the two are different kernel families.","section":"Notation throughout"},{"comment":"There are numerous typos (e.g., 'conlutional', 'Wikipidea', and 'da t a' in the title) and inconsistent use of 'dissertation' versus 'paper', which should be corrected in a revision.","section":"General presentation"}],"recommendation":"reject","confidential_remarks":"To the editor: the main theoretical claim is contradicted by a direct stabilizer computation, and the principal experimental claim is obtained with a different layer from the one the paper defines. The manuscript is not ready for review in its current form; I would not recommend a major revision unless the authors are willing to redo both the theory and the experiments. The paper also shows signs of an unrevised draft (e.g., references to 'Chapter 6'), which may indicate that the scope of the submission needs clarification."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead arXiv:2505.21173. My take: the core idea—build spectrogram convolution kernels by rotating two contrast-maximizing seed matrices with random SO(3) elements—is coherent and is not in the cited literature. The quotient-disk description of M/SO(3) with the x,y,z coordinates is a reasonable parametrization, and the decision to test on phoneme recognition with cross-checks on Speech Commands and CIFAR-10 is sensible. Credit where due: the author identifies a real gap in applying topological kernel ideas to speech, releases a GitHub link, and the Section 5.1 admission about omitted audio normalization is unusually honest.\n\nThe problems are serious. The stress-test note is correct: Definition 4.1 defines OF as W_k = α R_k M_k with M_k = ±A1, ±A2, but Section 4.3 replaces this with non-orthogonal [v,0,±v] kernels and, \"for simplicity\", keeps the name OF. Every headline accuracy figure—the ~70% SpeechBox result, the SNR experiments—evaluates the non-orthogonal replacement. The abstract's \"OF layer achieves superior performance\" is therefore attributed to the wrong object. This is not a minor naming slip; it makes the central empirical claim unsupported.\n\nThe theory has a load-bearing error too. The reader's stabilizer computation is right: for (v1,0,-v1), the stabilizer under SO(3) is SO(2), so the fiber is S^2, not RP^2. Remark 3.12 even notes the preimage is two-dimensional and that M→B is not a fiber bundle, then the paper asserts the stratified bundle anyway. Theorems 3.7, 3.9, and 3.11 are sketches, and Theorem 3.9's projection does not preserve unit norm as stated, so the theoretical scaffolding is not established.\n\nOther soft spots: no error bars or summary tables, only plots; the balanced 500-sample-per-phoneme selection is post hoc; and the paper does not engage the large equivariant deep learning literature on group-rotated filters, which would contextualise the novelty more honestly.\n\nWho is this for? Researchers interested in topological priors for speech kernels. The paper deserves a serious referee—there is a real idea and honest reporting of failures—but it needs major revision: either prove the stabilizer strata correctly or drop the bundle claim, and rerun the experiments with Definition 4.1's actual OF layer (or explicitly frame the paper as studying the relaxed non-orthogonal family). As is, I would not cite it.","headline":"A coherent kernel-rotation idea for speech CNNs, undermined by a false fiber-bundle claim and by experiments that swap the named OF layer for a different non-orthogonal kernel.","tokens_in":11691,"tokens_out":2342,"would_cite":false,"duration_ms":26737,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["55R10","57S15","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims an Orthogonal Feature layer, whose 3×3 kernels are SO(3)-rotations of two contrast-maximizing templates, beats kernel-filter and circular-filter baselines on phoneme recognition in clean and low-noise audio.","keywords":["topological data analysis","convolutional neural networks","phoneme recognition","SO(3) group action","fiber-bundle decomposition","spectrogram kernels","speech recognition","kernel design"],"falsifier":"Two concrete checks settle the claim. Theoretically: take $A = (1/\\sqrt{6})\\, ([1,1,1]^{\\top},\\, 0,\\, [-1,-1,-1]^{\\top})$, the paper's own boundary point, compute its stabilizer in $\\mathrm{SO}(3)$, and observe that it is the circle of rotations about the all-ones axis, so the orbit is $S^2$ and the asserted $\\mathbb{R}P^2$ fiber for the $\\mathbf{v}_1+\\mathbf{v}_3 = 0$ stratum does not hold. Empirically: re-run the SpeechBox, TIMIT, and LJSpeech phoneme comparisons with audio normalization and all preprocessing held identical across OF, KF, CF, and plain CNN, and check whether OF's clean-audio margin over KF and CF persists; the paper's Section 5.1 concedes that an omitted normalization step may have altered earlier relative rankings, so this is a live test rather than a formality.","tokens_in":10644,"feed_emoji":"🎙️","tokens_out":34507,"duration_ms":289829,"temperature":0.7,"pith_summary":"This paper tries to establish that topological structure, not just spectral shape, can be engineered into the convolutional kernels of a speech recognizer. Its practical claim is that an Orthogonal Feature (OF) layer — whose $3\\times 3$ kernels are generated as $\\mathrm{SO}(3)$ rotations of two high-contrast seed templates — reaches the highest phoneme-classification accuracy among the compared kernel families under clean and lightly noised ($\\mathrm{SNR}=20$ dB) conditions, while kernel filters win when noise is heavy ($\\mathrm{SNR}=0$ dB). Its theoretical claim is that the space of unit-norm, zero-sum $3\\times 3$ kernels fibers over a disk under the rotation group, with four distinct orbit types, and that this decomposition justifies generating every filter from a pair of seeds. A sympathetic reader would care because, if the claims hold, kernel design for audio gains a principled geometric prior and a smaller search space, and topological data analysis gains a concrete speech application beyond the natural-image patches where it began.","feed_headline":"Rotated seed kernels beat phoneme baselines in low-noise audio","feed_subtitle":"Rotating just two seed filters yields the clean-audio edge, and the construction carries to words and images.","key_machinery":"The load-bearing object is the constrained kernel space $M = \\{A = [\\mathbf{v}_1,\\mathbf{v}_2,\\mathbf{v}_3] \\in M_{3\\times 3}(\\mathbb{R}) : \\lVert A\\rVert = 1,\\ \\mathbf{v}_1+\\mathbf{v}_2+\\mathbf{v}_3 = 0\\}$, equipped with the contrast measure $\\mathrm{con}(A) = \\sqrt{\\lVert\\mathbf{v}_1-\\mathbf{v}_2\\rVert^2 + \\lVert\\mathbf{v}_2-\\mathbf{v}_3\\rVert^2}$. The machinery that carries the argument is the claimed stratified fiber bundle of the quotient $B = M/\\mathrm{SO}(3) \\cong D^2$ with the four fiber types listed above. That decomposition is what licenses the OF layer's generation rule: instead of learning nine free weights, every kernel is a rotated seed, $W_k = \\alpha R_k M_k$ with $R_k \\in \\mathrm{SO}(3)$ obtained by exponentiating Gaussian-sampled Lie-algebra elements $\\theta = \\sum_i \\theta_i L_i$, and $M_k$ one of the four augmented seeds. The stratum structure of $B$ is also offered as a coordinate system on kernel space, so that filter families can in principle be indexed by their position in the disk.","core_discovery":"The paper's central contribution is the combination of a geometric claim about kernel space with a concrete layer built from it. The geometric claim: the space $M$ of $3\\times 3$ real kernels with unit Frobenius norm and column sum $\\mathbf{v}_1+\\mathbf{v}_2+\\mathbf{v}_3 = 0$ is homeomorphic to $S^5$, and the left action of $\\mathrm{SO}(3)$ on $M$ produces a quotient $B = M/\\mathrm{SO}(3)$ homeomorphic to a disk $D^2$, parameterized by the invariants $x = \\lVert\\mathbf{v}_1\\rVert^2$, $y = \\lVert\\mathbf{v}_3\\rVert^2$, $z = \\mathbf{v}_1\\cdot\\mathbf{v}_3$ subject to $x+y+z = 1/2$ and $z^2 \\le xy$. Over $B$ the paper asserts a stratified fiber bundle whose fibers are $S^2$ on the boundary, the lens space $L(4,1)$ where $\\lVert\\mathbf{v}_1\\rVert = \\lVert\\mathbf{v}_3\\rVert$, the real projective plane $\\mathbb{R}P^2$ where $\\mathbf{v}_1+\\mathbf{v}_3 = 0$, and a principal $\\mathrm{SO}(3)$ bundle elsewhere; Remark 3.12 notes that the plain map $M \\to B$ is not a fiber bundle, since the preimage of the vertical-stripe point is only two-dimensional, so the statement is cast in stratified form. Both Theorem 3.7 and Proposition 3.11 are presented as proof sketches, so the geometric scaffolding is asserted rather than fully derived in the text. From this, the Orthogonal Feature layer is defined by augmenting two seed templates (the vertical stripe $[1,0,-1]$ and the second difference $[1,-2,1]$, plus their negatives) and sampling $W_k = \\alpha R_k M_k$ with $R_k = \\exp(\\theta)$ drawn from the Lie algebra $\\mathfrak{so}(3)$. On SpeechBox, TIMIT, and LJSpeech phoneme tasks the paper reports that this layer matches or beats kernel-filter, circular-filter, and plain CNN baselines under clean and $\\mathrm{SNR}=20$ dB conditions, degrades at $\\mathrm{SNR}=0$ where kernel filters take over, and transfers to Speech Commands word classification and CIFAR-10 images.","pith_inferences":["My inference: the accuracy numbers do not actually depend on the disputed fiber-type labels, because the seed templates' orbits are spheres regardless of the stratum label; correcting the stratification would revise the theory without necessarily changing the reported accuracy.","My inference: the roughly 70% SpeechBox accuracy belongs to the relaxed zero-contrast variant that Section 4.3 keeps calling OF, not to the strictly orthogonal kernel set of Section 4.2, so any later comparison should pin down which OF variant a reported number refers to.","My inference: the paper's concession that an omitted audio-normalization step may explain the unexpected unfiltered rankings leaves open the possibility that the same confound affects the clean-audio OF-versus-KF margin, making a re-run with preprocessing fixed the cheapest firming-up test.","My inference: sampling kernels from each claimed stratum — boundary, lens-space, projective — and correlating the fiber type with task accuracy would test whether the topology itself does the work; the paper tests only two seeds and leaves that stratum-to-performance map unmeasured."],"forward_implications":["If the OF layer's accuracy holds, speech kernel design can start from a couple of topologically motivated seed templates plus a stochastic rotation sampler instead of free-form learned weights, shrinking the effective parameter space of early convolution layers.","The reported noise-dependent ranking — OF first at $\\mathrm{SNR}=20$ dB, KF first at $\\mathrm{SNR}=0$ dB — means the best kernel family depends on the noise regime, and accuracy averaged over noise levels can mask which architecture actually dominates.","The claimed description of $3\\times 3$ kernel space as $S^5$ with quotient disk $D^2$ gives kernel families a coordinate system, so future filters could be enumerated or interpolated by position in the quotient rather than sampled blindly.","The reported transfer to Speech Commands words and CIFAR-10 images, with parity to Klein-feature networks, supports the claim that the contrast-maximizing zero-sum constraint is not speech-specific even though spectrogram directionality motivated it."],"supporting_citations":[{"why":"the topological convolutional layer this work extends and the kernel-filter (KF) baseline it must beat on phoneme data","marker":"[5]"},{"why":"supplies the eight basic high-contrast patch vectors that the paper reduces, by zero-contrast and group-action arguments, to its two seed templates","marker":"[7]"},{"why":"the Klein-bottle analysis of natural-image patches that motivates treating local patches as topological objects","marker":"[2]"},{"why":"the theoretical framework for treating topological features as information-bearing structure in data, adopted as the paper's lens","marker":"[3]"},{"why":"the primary phoneme dataset (SpeechBox) for the OF-versus-baseline comparisons","marker":"[8]"},{"why":"TIMIT, the second phoneme dataset used to confirm the accuracy ranking across kernel families","marker":"[9]"},{"why":"LJSpeech, the third phoneme dataset and the cleanest of the three in the reported ranking","marker":"[10]"}],"fun_headline_variants":["Topology-inspired kernels sharpen phoneme recognition in quiet settings","Clean audio sees gains from topology-aware filter rotation","Orthogonal Feature layer lifts speech accuracy when noise is low","TDA-inspired kernels give speech recognition a clean-audio boost"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The construction assumes that the $\\mathrm{SO}(3)$ orbits of zero-sum, unit-norm $3\\times 3$ kernels split into exactly the four claimed symmetry types, and in particular that kernels of the form $[\\mathbf{v}, 0, -\\mathbf{v}]$ have a real-projective-plane orbit shape, even though the rotations that leave such a kernel fixed are precisely the circle of spins about $\\mathbf{v}$, which would make the orbit a sphere.","fun_headline_variants_meta":{"raw":{"variants":["Topology-inspired kernels sharpen phoneme recognition in quiet settings","Clean audio sees gains from topology-aware filter rotation","Orthogonal Feature layer lifts speech accuracy when noise is low","TDA-inspired kernels give speech recognition a clean-audio boost"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000497,"raw_usage":{"total_tokens":2537,"prompt_tokens":1147,"completion_tokens":1390,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":763,"completion_tokens_details":{"reasoning_tokens":1324}},"tokens_in":763,"tokens_out":1390,"duration_ms":13816,"temperature":1.0,"reasoning_tokens":1324,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:33:35.413713+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Two concrete checks settle the claim. Theoretically: take $A = (1/\\sqrt{6})\\, ([1,1,1]^{\\top},\\, 0,\\, [-1,-1,-1]^{\\top})$, the paper's own boundary point, compute its stabilizer in $\\mathrm{SO}(3)$, and observe that it is the circle of rotations about the all-ones axis, so the orbit is $S^2$ and the asserted $\\mathbb{R}P^2$ fiber for the $\\mathbf{v}_1+\\mathbf{v}_3 = 0$ stratum does not hold. Empirically: re-run the SpeechBox, TIMIT, and LJSpeech phoneme comparisons with audio normalization and all preprocessing held identical across OF, KF, CF, and plain CNN, and check whether OF's clean-audio margin over KF and CF persists; the paper's Section 5.1 concedes that an omitted normalization step may have altered earlier relative rankings, so this is a live test rather than a formality.","supporting_citations":[{"cited_title":"Topological convolutional layers for deep learning","cited_arxiv_id":null,"evidence_quote":"the topological convolutional layer this work extends and the kernel-filter (KF) baseline it must beat on phoneme data"},{"cited_title":"The nonlinear statistics of high-contrast patches in natural images","cited_arxiv_id":null,"evidence_quote":"supplies the eight basic high-contrast patch vectors that the paper reduces, by zero-contrast and group-action arguments, to its two seed templates"},{"cited_title":"On the local behavior of spaces of natural images","cited_arxiv_id":null,"evidence_quote":"the Klein-bottle analysis of natural-image patches that motivates treating local patches as topological objects"},{"cited_title":"Bradlow, A","cited_arxiv_id":null,"evidence_quote":"the primary phoneme dataset (SpeechBox) for the OF-versus-baseline comparisons"},{"cited_title":"Speech database development at mit: Timit and beyond","cited_arxiv_id":null,"evidence_quote":"TIMIT, the second phoneme dataset used to confirm the accuracy ranking across kernel families"},{"cited_title":"The lj speech dataset","cited_arxiv_id":null,"evidence_quote":"LJSpeech, the third phoneme dataset and the cleanest of the three in the reported ranking"}],"review_version":1}