Pith. sign in

REVIEW 3 major objections 3 minor 1 cited by

FilBench: Can LLMs Understand and Generate Filipino?

T0 review · 3 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper introduces FilBench, a Filipino-centric benchmark, and claims that the best tested model, GPT-4o, reaches only 72.23% accuracy, exposing a substantial gap in LLM performance on Filipino, Tagalog, and Cebuano tasks.

desk verdict FilBench is a plausible, useful addition to low-resource benchmark suites, but the abstract alone can't support the headline capability-gap numbers; worth sending to referees with the full methodology. read the letter →

arxiv 2508.03523 v1 pith:WJDYWOLD submitted 2025-08-05 cs.CL

classification cs.CL
keywords FilipinoNLPLLMbenchmarkmultilingualevaluationTagalogCebuanoreadingcomprehensionmachinetranslationculturalknowledge
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper seeks to measure how well large language models understand and generate Filipino, a language that is largely absent from mainstream LLM evaluation. To do this, it introduces FilBench, a benchmark built from Filipino NLP research priorities and spanning cultural knowledge, classical NLP, reading comprehension, and generation. The central finding is that the best model, GPT-4o, scores only 72.23%, and that models trained specifically for Southeast Asian languages do even worse, with SEA-LION v3 70B at 61.07%. The paper argues that this gap demonstrates the need for language-specific benchmarks to drive progress in Filipino NLP and to include Philippine languages in model development.

What carries the argument

The central object is FilBench itself, a curated set of Filipino-language tasks organized into four families: Cultural Knowledge, Classical NLP, Reading Comprehension, and Generation. It carries the argument by providing a single benchmark on which all 27 models are scored, allowing direct comparison and exposing task-specific weaknesses; the aggregate score is the evidence for the paper's claim that current LLMs are not yet proficient in Filipino.

What would settle it

An independent replication in which fluent Filipino speakers re-annotate a random sample of FilBench items and recompute the model scores would settle the claim: if the re-annotation finds high label disagreement or if human performance on the benchmark is no better than GPT-4o's 72.23%, the benchmark's difficulty and its implication of a model capability gap would be called into question.

Watch

Extended reading notes

Core claim

FilBench is a challenging benchmark for LLMs in Filipino, Tagalog, and Cebuano. Across 27 state-of-the-art models, the highest score is GPT-4o's 72.23%, and no model comes close to ceiling performance. Models specialized for Southeast Asian languages, such as SEA-LION v3 70B, underperform general-purpose models, reaching only 61.07%. The paper attributes this gap to weaknesses in reading comprehension and translation, and concludes that curated, language-specific evaluation is necessary to reveal and address these gaps.

Load-bearing premise

The entire conclusion rests on the assumption that FilBench's tasks, labels, and scoring accurately reflect real Filipino language use and NLP priorities; if the curation is flawed or the labels are noisy, the reported model scores and rankings would not reliably show how well LLMs handle Filipino.

Editorial extensions

If this is right

  • Current LLMs, including top general-purpose models, have a meaningful performance gap in Filipino-language tasks, which suggests that improvements are needed before reliable deployment in Filipino-speaking contexts.
  • The finding that SEA-LION v3 70B underperforms indicates that regionally specialized models do not automatically bring better performance in specific Philippine languages.
  • The benchmark provides a reusable evaluation suite for tracking progress in Filipino NLP, allowing future models to be measured against the same 72.23% ceiling.
  • Weaknesses in reading comprehension and translation point to concrete capability areas that model developers should target to improve Filipino language understanding.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A score of 72.23% on a carefully curated benchmark leaves substantial headroom, which suggests that fine-tuning on Filipino data, not just multilingual pretraining, may push models closer to ceiling performance.
  • Because the benchmark was designed to reflect Philippine NLP research trends, the reported rankings may also reflect the difficulty of tasks Filipinos actually care about, making them more informative than a generic multilingual test.
  • A natural next step would be to compare model scores with a human baseline on the same tasks; that would clarify whether the gap represents a real deficiency or merely hard questions.
  • The relative underperformance of SEA-LION models could be probed by testing whether it stems from training data coverage, tokenization, or evaluation protocol, which this paper does not isolate.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper introduces FilBench, a Filipino-centric benchmark covering Filipino, Tagalog, and Cebuano across categories such as Cultural Knowledge, Classical NLP, Reading Comprehension, and Generation. The authors evaluate 27 state-of-the-art LLMs and report that GPT-4o achieves the highest score of 72.23%, while the best Southeast Asian language-specific model, SEA-LION v3 70B, reaches only 61.07%. The abstract concludes that FilBench is challenging and demonstrates the need for language-specific LLM benchmarks. The manuscript under review consists solely of this abstract; no full text was provided with the submission.

Significance. If the benchmark is well-constructed and the evaluation protocol is sound, this work would make a valuable contribution to multilingual LLM evaluation and to Philippine NLP. The reported performance gap for a underrepresented language family is important, and the observation that region-specific models may underperform on Filipino is a potentially consequential finding. The paper would offer a concrete, reusable resource for tracking progress in Filipino language understanding and generation. However, the significance cannot currently be assessed because the abstract provides no information about task curation, annotation quality, evaluation settings, or scoring methodology. For these reasons, the work is potentially significant but presently unverified.

major comments (3)
  1. [Abstract] The central claim that 'FilBench is challenging, with the best model, GPT-4o, achieving only a score of 72.23%' is not interpretable without an explicit evaluation protocol. The abstract does not state whether all 27 LLMs were evaluated with identical prompts, decoding settings, few-shot examples, or scoring metrics. Without this information, the headline score cannot be compared across models, and it is impossible to determine whether the gap reflects Filipino-specific difficulty or artifacts of the evaluation setup.
  2. [Abstract] The claim that tasks were 'carefully curated' is not backed by any description of the curation criteria, data sources, annotation guidelines, inter-annotator agreement, or error-exclusion procedures. As the abstract stands, the benchmark's validity rests entirely on an unverifiable assertion. This is a load-bearing gap because noisy or unrepresentative tasks would undermine the capability rankings and the conclusion about LLM proficiency in Filipino.
  3. [Abstract] The abstract reports aggregated scores (e.g., 72.23% and 61.07%) without defining how these scores are computed. It is unclear whether they are averages over tasks, weighted by task difficulty, or computed with a specific metric (e.g., exact-match, F1, or human preference). Without the aggregation protocol, the numerical comparisons among models and the conclusion that 'several LLMs suffer from reading comprehension and translation capabilities' cannot be verified from the manuscript.
minor comments (3)
  1. [Abstract] The abstract lists four task categories but does not define them or state how many tasks constitute each category; adding this information would help the reader assess coverage and balance.
  2. [Abstract] The relationship among Filipino, Tagalog, and Cebuano in the benchmark is unspecified. The abstract should clarify whether each task is in a single language, whether all languages appear across tasks, and how language coverage was determined.
  3. [Abstract] The phrase 'the value of curating language-specific LLM benchmarks' is a general claim that would be strengthened by stating what concrete insights or actionable recommendations follow from the FilBench results.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity identified in the abstract-level derivation: FilBench scores are externally measured against fixed tasks.

full rationale

The abstract-level derivation chain is not circular. FilBench is a newly constructed benchmark whose task definitions are author-curated, but the paper's central quantitative claims (GPT-4o scoring 72.23%, SEA-LION v3 70B scoring 61.07%, and several LLMs underperforming on reading comprehension and translation) are obtained by evaluating external, independently trained LLMs on those fixed tasks. No parameter is fitted to the reported scores, no prediction is defined in terms of the benchmark's own outputs, and no uniqueness or self-citation argument is invoked to force a conclusion. The authors' statement that tasks are 'carefully curate[d]' reflects a design choice rather than a self-referential reduction: the benchmark's challenge level is an empirical finding about the evaluated models, not an input to the evaluation. The legitimate concern raised by the abstract is that the curation and scoring protocols are not disclosed, which affects interpretability and comparability of the reported numbers; that is a methodology/transparency risk, not circularity. Since the full text is unavailable and no equations or fitted parameters are visible, there is no quotable evidence of a circular step, and the honest finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

No free parameters or invented entities are introduced. The central result depends on domain assumptions about the representativeness of the curated tasks and the fairness of the evaluation setup.

assumptions (2)
  • domain assumption Task curation reflects the priorities and trends of Philippine NLP research.
    Abstract says tasks were 'carefully curated' to reflect these priorities, but gives no evidence or external standard for that reflection.
  • domain assumption The evaluation protocol is consistent across all 27 models.
    No prompts, decoding parameters, or scoring details are given in the abstract; a fair comparison is assumed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of FilBench: Can LLMs Understand and Generate Filipino?." pith.science (2026). https://pith.science/paper/WJDYWOLD

@misc{pith2026250803523,
  author       = {Pith},
  title        = {Pith review of: FilBench: Can LLMs Understand and Generate Filipino?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WJDYWOLD}},
  note         = {Machine review of arXiv:2508.03523}
}
read the original abstract

Despite the impressive performance of LLMs on English-based tasks, little is known about their capabilities in specific languages such as Filipino. In this work, we address this gap by introducing FilBench, a Filipino-centric benchmark designed to evaluate LLMs across a diverse set of tasks and capabilities in Filipino, Tagalog, and Cebuano. We carefully curate the tasks in FilBench to reflect the priorities and trends of NLP research in the Philippines such as Cultural Knowledge, Classical NLP, Reading Comprehension, and Generation. By evaluating 27 state-of-the-art LLMs on FilBench, we find that several LLMs suffer from reading comprehension and translation capabilities. Our results indicate that FilBench is challenging, with the best model, GPT-4o, achieving only a score of 72.23%. Moreover, we also find that models trained specifically for Southeast Asian languages tend to underperform on FilBench, with the highest-performing model, SEA-LION v3 70B, achieving only a score of 61.07%. Our work demonstrates the value of curating language-specific LLM benchmarks to aid in driving progress on Filipino NLP and increasing the inclusion of Philippine languages in LLM development.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KatotohananQA: Evaluating Truthfulness of Large Language Models in Filipino

    cs.CL 2025-09 conditional novelty 6.0 of 10

    KatotohananQA is a Filipino translation of TruthfulQA; seven LLMs scored 94.72% in English versus 83.87% in Filipino, with GPT-5 and GPT-5 mini showing the smallest gap.

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.