Pith. sign in

REVIEW 1 major objections 5 minor 25 references

Hallucination Level of Artificial Intelligence Whisperer: Case Speech Recognizing Pantterinousut Rap Song

T0 review · 1 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Two speech-to-text engines trade wins on a Finnish rap song, with no clear champion and mixed results from stripping out the music.

desk verdict A self-aware, n=1 case study whose one quantitative claim is internally contradicted by its own table, so the preprocessing result is unsupported as printed. read the letter →

arxiv 2506.16174 v2 pith:36ROADB4 submitted 2025-06-19 cs.LG

classification cs.LG
keywords AIhallucinationautomaticspeechrecognitionFasterWhisperFinnishrapspeech-to-textLevenshteindistancestemseparationYouTubecaptions
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to measure how often two automatic speech-to-text systems—Whisper-derived Faster Whisper and YouTube's internal caption engine—hallucinate or mishear Finnish rap lyrics from a single song. The authors build an informal error metric that counts hallucinations as worse than mishearings and adjusts Levenshtein edit distance for phonemically close single-character swaps. Their verdict is that neither system is clearly better, and that preprocessing the audio with a stem splitter to isolate vocals improves some lyric extracts while worsening others. The point matters because Finnish is under-served by speech-to-text services, and rap over synth music is a stress case for automatic subtitling.

What carries the argument

The load-bearing object is the informal human-assisted error function of Section 6.1. For each difference between the transcribed line and the written Finnish lyrics, the first author first classifies the difference as a hallucination or a mishearing; hallucinations are penalized most heavily, mishearings are scored by Levenshtein string edit distance, and a single-character substitution with low phonemic distance (like b↔p or t↔d) counts as half an error. This metric decides every per-lyric verdict and the overall winner count. The other machinery is the additive signal model audio(t) = vocals(t) + instrumental(t), used to frame three preprocessing attempts—independent component analysis, center-channel cancellation, and LALAL.AI stem separation—of which only the last produced a usable vocal track.

What would settle it

Re-tally Section 8's per-lyric verdicts: the winners listed are B, A, draw, B, A, B, B, which is four wins for B and two for A—not the three B wins stated in the summary; recounting the table decides whether the headline count is a typo or a different counting rule.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that a large local Whisper-based model and YouTube's proprietary speech-to-text engine are roughly matched on Finnish rap: in the seven sampled lyrics, YouTube won twice, Faster Whisper won three times, and two lines were perfect for both. When vocals were separated from the instrumental with a stem splitter, the paper reports that pre-processed audio won three times and raw audio twice, but it also notes the result is not statistically meaningful and that preprocessing helps some lines and hurts others. The authors explicitly state that it is not clear that YouTube's algorithm performs better. The paper's own per-lyric table in Section 8 actually lists four wins for the pre-processed method and two for raw audio, with one draw, so the stated 3-out-of-7 summary is internally inconsistent; either way, the conclusion of no clear winner remains the intended claim.

Load-bearing premise

The whole scoreboard rests on one person's hand-labeled judgment of whether each wrong word is a hallucination or a mishearing, with no operational rule given for telling the two apart.

Editorial extensions

If this is right

  • If the claim holds, automatic subtitling of Finnish rap cannot yet rely on a single off-the-shelf model; hallucinations and mishearings will occur in roughly equal measure across engines.
  • Preprocessing by stem separation is not a guaranteed fix: it can remove background music artifacts but can also introduce new mishearings or hallucinations on other lines.
  • The metric's penalty structure matters: a system with fewer hallucinations can be declared the winner even when its total edit distance is larger, so hallucination rate, not just word accuracy, drives the comparison.
  • Because the sample is one song, the paper's own conclusion is that stronger claims would require more songs and an automated version of the error function.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: a natural extension the authors do not pursue is to replace the binary phonemic halving rule with a continuous phoneme-confusion score, such as one based on IPA feature distances, which would make the metric less sensitive to the judge's choices.
  • Editorial inference: the internal tally discrepancy (three vs four B wins) suggests that the counting rule may not be fully specified; a re-analysis that states exactly how multi-word hallucinations and multiple mishearings within one lyric are counted would decide whether the headline number is a typo.
  • Editorial inference: the same paired-comparison protocol could be run on a dozen Finnish rap tracks and scored with a McNemar-style test of the per-lyric winners, turning the authors' 'gut feeling' into a statistical statement.
  • Editorial inference: since preprocessing artifacts are a plausible cause of the mixed results, a testable extension would compare several stem splitters on the same track and check whether any splitter's vocal isolation quality predicts ASR improvement line by line.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 5 minor

Summary. The paper compares two automatic speech-to-text systems, Faster Whisper (referred to as 'Faster Whisperer') and YouTube's internal speech-to-text, on the task of transcribing a Finnish rap song, 'Pantterinousut'. The authors define an informal error metric that distinguishes hallucinations from mishearings, evaluate both systems on seven selected lyric fragments, and also test three preprocessing approaches (ICA, center cancellation, and a LALAL.AI stem splitter) aimed at isolating vocals before transcription. The main reported findings are that YouTube and Faster Whisper perform similarly on the raw audio (YouTube 2 wins, Faster Whisper 3 wins, 2 draws), that the LALAL.AI-preprocessed input changes the tally (raw audio 2 wins, preprocessed 3 wins, with a draw), and that, given the small sample, there is no statistical confidence in either comparison.

Significance. As a small, informal case study, the paper has some strengths: it is transparent about its methodology, includes the commands and code used, discloses the outputs for each lyric, and explicitly acknowledges the lack of statistical power. The authors also name the song and provide the reference URL, so the experiment is in principle reproducible. However, the contribution is limited by the subjective error metric, the very small sample size (one song, seven lyric fragments), and the judge being the first author, who states a prior 'gut feeling' in Section 3.3. The internal arithmetic inconsistency in Section 8 undermines the one quantitative summary. If the tally and metric issues are fixed, the paper could serve as a useful data point for informal comparisons of ASR on Finnish rap, but in its current form the reported counts are not trustworthy.

major comments (1)
  1. [Section 6.2, Lyrics #3 and #5] The reported Levenshtein distances are not consistent with the displayed strings. In Lyric #3, the paper reports Faster Whisperer mishearing 'sallii' as 'salliin' with distance 2, but that substitution is a single insertion. In Lyric #5, the printed distance totals for Faster Whisperer (3) and YouTube (4) appear too small: a direct edit count of 'Tordaita alas, niinku katu hauskaa' against 'Portaita alas, niin kuin katuhaukka' is considerably larger, even allowing for the paper's half-distance rule. Because the winner rule in Section 6.1 relies on these distances, the alignment used for each distance must be shown.
minor comments (5)
  1. [Throughout] The paper consistently writes 'Faster Whisperer' where the model is actually 'Faster Whisper'; the reference in [3] and the OpenAI Whisper reference in [6] use the correct name.
  2. [Section 7] Approaches #1 (ICA) and #2 (center cancellation) are described as failures, but no quantitative or qualitative evidence is given for that assessment; at least a brief description of what the resulting audio sounded like would help the reader.
  3. [Section 3.3] The phrase 'automatic text-to-speech translation' should read 'speech-to-text'.
  4. [Section 5 and Section 6.2] The 'Tordaita alas, niinku katu hauskaa' extraction is called an example of AI hallucination in Section 5, but in Section 6.2 Lyric #5 it is classified as 'hallucinates twice and mishears one word,' which is a contradiction in the labeling.
  5. [Throughout] There are several typos: 'Levenshstein' appears multiple times, 'Pantteri nousuista' in Lyric #7 of Section 8 is later spelled 'Pantteri nousuista' (inconsistent use of 'Pantteri' vs 'Pantteri'), and the reference to 'ref. 23' in Section 3.3 does not point to the caption reference.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the ASR outputs are external evidence and no parameter is fitted from them; the paper's claims are measurement summaries, not derivations from their own inputs.

full rationale

The paper's load-bearing claims are counts of ASR transcription wins derived by comparing two external systems' outputs (Faster Whisperer and YouTube speech-to-text) against a fixed reference text, the Finnish rap lyrics written by Mc Timo. Nothing in the reported pipeline fits a parameter to the ASR outputs and then re-predicts those same outputs; the error metric is a hand-applied Levenshtein-distance-plus-hallucination rule, but the transcribed strings themselves come from outside the paper and are not constructed from the metric. The subjective classification step (Section 6.1, step 1) and the author's prior 'gut feeling' (Section 3.3) create reproducibility and bias risks, but they do not make the argument circular: the ASR transcripts are not derived from the judge's classification. There are no load-bearing self-citations: the references are to software, tutorials, and external tools, not to prior papers by these authors that supply an unverified premise. The most serious quantitative defect is an internal arithmetic inconsistency in Section 8: the per-lyric verdicts listed above the summary give method B four wins, method A two wins, and one draw, whereas the text below says A won twice and B won three times. That is a correctness error, not a circularity, because the contradiction is between two statements of the same tally rather than between an output and an input. Under the requirement that circularity be exhibited as a specific reduction by construction, no such reduction is present, so the appropriate finding is no significant circularity.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

No fitted parameters; the only hand-set numeric quantity is the 0.5 distance-halving rule, which affects the winner counts. The measurement rests on the assumptions that the author's hallucination/mishearing classification is consistent and that edit distance captures perceptual error.

free parameters (1)
  • Phonemic-distance halving factor = 0.5 (edit distance halved)
    Introduced by hand in Section 6.1 (step 3); applied to substitutions like b/p and t/d. This factor flips the winner in Lyric #6 of the first comparison and Lyric #5 of the second, so the reported winner counts depend on it.
assumptions (3)
  • domain assumption A native Finnish-speaking judge can reliably classify each ASR error as hallucination or mishearing without an operational definition.
    Section 6.1, step 1; the first author serves as judge.
  • domain assumption Levenshtein edit distance, with halving for phonetically close single-character differences, is a valid error measure for comparing ASR outputs.
    Section 6.1, steps 2-3; asserted without validation.
  • domain assumption The lyrics provided by Mc Timo are the correct reference text.
    Abstract and Section 1 define this as the reference truth.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Hallucination Level of Artificial Intelligence Whisperer: Case Speech Recognizing Pantterinousut Rap Song." pith.science (2026). https://pith.science/paper/36ROADB4

@misc{pith2026250616174,
  author       = {Pith},
  title        = {Pith review of: Hallucination Level of Artificial Intelligence Whisperer: Case Speech Recognizing Pantterinousut Rap Song},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/36ROADB4}},
  note         = {Machine review of arXiv:2506.16174}
}
read the original abstract

All languages are peculiar. Some of them are considered more challenging to understand than others. The Finnish Language is known to be a complex language. Also, when languages are used by artists, the pronunciation and meaning might be more tricky to understand. Therefore, we are putting AI to a fun, yet challenging trial: translating a Finnish rap song to text. We will compare the Faster Whisperer algorithm and YouTube's internal speech-to-text functionality. The reference truth will be Finnish rap lyrics, which the main author's little brother, Mc Timo, has written. Transcribing the lyrics will be challenging because the artist raps over synth music player by Syntikka Janne. The hallucination level and mishearing of AI speech-to-text extractions will be measured by comparing errors made against the original Finnish lyrics. The error function is informal but still works for our case.

Figures

Figures reproduced from arXiv: 2506.16174 by the authors.

Figure 1
Figure 1. Start of the Pantterinousut video (from which we want to extract the Finnish rap [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Rules for selecting a suitable Whisper model, [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Architecture of OpenAI’s Whisper model, image by OpenAI (ref. 6.) As a reference, [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Spectrogram of audio of the Pantterinousut song (X is time and Y is frequency). [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 6
Figure 6. Figure 6: Applying Faster Whisper XXL model to the Pantterinousut song. [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: Idea of ICA in so-called cocktail party problem with two original signals, image by Enes Zvornicanin, ref 25. We have a stereo audio file split into left and right channels. Then, we assume that these two channels separated measure two independent signals with a linear…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

25 extracted references · 25 canonical work pages

  1. [1]

    Pantterinousut

    Mc Timo & Syntikka Janne (2025) . Pantterinousut. Internet url: https://www.youtube.com/watch?v=iq0hrThG96o (referenced on June 19th 2025)

  2. [2]

    Facebook pages

    Mc Timo & Syntikka Janne (2025) . Facebook pages. Internet url: https://www.facebook.com/profile.php?id=61576624451346 (referenced on June 19th 2025)

  3. [3]

    Faster Whisper standalone Wind ows version, https://github.com/Purfview/whisper- standalone-win (referenced on June 19 th 2025)

  4. [4]

    VLC Media Player with automatic AI subtitles, https://www.facebook.com/groups/726035349513216/posts/vlc-media-player-has- just-introduced-free-al-subtitles-with-real-time-translatio/999539558829459/ (referenced on June 19th 2025)

  5. [5]

    VLC Experimental Builds, https://nightlies.videolan.org/ (referenced on June 19th 2025)

  6. [6]

    OpenAI’s Whisper, https://openai.com/index/whisper/ (referenced on June 19th 2025)

  7. [7]

    LLPlayer, https://github.com/umlx5h/LLPlayer (referenced on June 19th 2025)

  8. [8]

    PotPlayer, https://potplayer.daum.net/ (referenced on June 19th 2025)

Show all 25 references
  1. [9]

    Omar Sanseviero, “Which Whisper To Use “ https://x.com/osanseviero/status/1725122881384776023 (referenced on June 19th 2025)

  2. [10]

    VLC Media Player, https://www.videolan.org/vlc/ (referenced on June 19th 2025)

  3. [11]

    Ans, B., Hérault, J., & Jutten, C. (1985). Architectures neuromimétiques adaptatives : Détection de primitives. Cognitiva 85 (volume 2, pp. 593-597). Paris: CESTA

  4. [12]

    LALAL.AI, https://www.lalal.ai/ http, (referenced on June 19th 2025)

  5. [13]

    Taskinen, S., Sirkiä, S., & Oja, H. (2007). Independent component analysis based on symmetrised scatter matrices. Comput. Statist. Data Anal., 51(10), pp. 5103- 5111

  6. [14]

    (2010) Contributions to independent component analysis, sensor array and complex valued signal processing

    Ollila, E. (2010) Contributions to independent component analysis, sensor array and complex valued signal processing. Helsinki University of Technology. Doctoral Dissertation

  7. [15]

    Complex-valued ICA based on a pair of generalized covariance matrices

    Ollila, E., Oja H., Koivunen V., (2008) “Complex-valued ICA based on a pair of generalized covariance matrices”. Computational Statistics & Data Analysis, volume 52, number 7, pp. 3789-3805

  8. [16]

    Benji Pugh and Caitlin Coffey, https://caitlincoffey.com/finalprojectml/ (referenced on June 19th 2025)

  9. [17]

    (2025) https://zaniboni.com/blog/post/noise-cancellation (referenced on June 19th 2025)

    Masella L. (2025) https://zaniboni.com/blog/post/noise-cancellation (referenced on June 19th 2025)

  10. [18]

    Audacity, https://www.audacityteam.org/ (referenced on June 19th 2025)

  11. [19]

    Hugging Face, https://huggingface.co/Finnish-NLP/wav2vec2-xlsr-300m-finnish-lm (referenced on June 15th 2025)

  12. [20]

    AaltoASR, https://github.com/aalto-speech/AaltoASR (referenced on June 19th 2025)

  13. [21]

    Fabian Berg, https://www.quora.com/Why-is-it-so-hard-to-hear-the-difference- between-p-b-and-t-d-Arent-people-just-saying-it-wrong-all-the-time (referenced on June 19th 2025)

  14. [22]

    Hyvärinen, P

    A. Hyvärinen, P. Ramkumar, L. Parkkonen and R. Hari, (2010). Independent component analysis of short-time Fourier transforms for spontaneous EEG/MEG analysis, NeuroImage 49(1): pp. 257-271

  15. [23]

    FourierICA, https://www.cs.helsinki.fi/group/neuroinf/code/fourierica/html/fourierica.html (referenced on June 19th 2025)

  16. [24]

    Why Captions are Critical for YouTube, https://scribie.com/blog/2019/02/why- captions-critical-youtube/ (referenced on June 19th 2025)

  17. [25]

    Enes Zvorničanin https://www.baeldung.com/cs/independent-component-analysis (referenced on June 19th 2025)

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.