Pith. sign in

REVIEW 4 major objections 4 minor 6 cited by

Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models

T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper's central claim is that a single open recipe—curated Thai-English continued pretraining plus staged post-training—brings Thai text, vision, and audio models to the top of Thai-language benchmarks while preserving the base…

desk verdict Open Thai model family worth knowing about, but the headline benchmark story is circular. read the letter →

arxiv 2412.13702 v2 pith:4FALN3SA submitted 2024-12-18 cs.CL cs.AI

classification cs.CLcs.AI
keywords Thailanguagemodelscontinualpre-traininginstructiontuningfunctioncallinglongcontextmultimodalLLMspeech-to-speechAIsafetyclassifier
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports a family of Thai language models, from 1 to 70 billion parameters, built by continuing to train capable open base models on a mixture of Thai and English text and then applying a staged post-training recipe. The authors claim the resulting text models reach the best published results on Thai exams and Thai instruction-following while keeping English, math, coding, and long-context abilities largely intact. The same project adds a Thai vision model focused on document OCR and chart questions, an end-to-end audio model that can listen to speech and answer in parallel text and speech, and a lightweight safety classifier tuned to Thai cultural sensitivities. The significance, if the evaluations hold, is that a low-resource language can be brought to competitive model quality with an open, reproducible recipe rather than proprietary data.

What carries the argument

The load-bearing mechanism is the Typhoon2 training recipe. It starts with continual pre-training, where an already-trained English-centric base model is trained further on roughly 12 billion high-quality Thai tokens plus a 50 percent English mix to avoid forgetting; the Thai tokens are selected by four filters—a cultural-relevance classifier, a Thai quality fastText classifier, synthetic textbook-style augmentation, and high-educational-content filtering. Post-training then layers general SFT with a 3:7 Thai-to-English ratio, math and code SFT with a Thai-translated subset, long-context data at 15 percent of the mix, function-calling data at 5–10 percent, top-k logits distillation for 1B/3B models, and a DARE+linear merge with a newer instruction model for the 70B model. For audio, the machinery is an encoder-adapter-LLM front end (Whisper-type speech encoder plus audio-event encoder, aligned through a Q-Former) feeding a non-autoregressive speech decoder trained with CTC to output discrete speech units, which a unit vocoder turns into waveform; this lets text and speech be decoded in parallel from the same LLM hidden states.

What would settle it

Write a new Thai national-exam-style test with questions dated after the model's training cutoff, run the released Typhoon2-Text models and their base models on it, and compare; if the Typhoon2 advantage over the base models shrinks to near zero, the benchmark-driven improvement is largely an artifact of overlap with training data.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central claim is that a single adaptable recipe—domain-filtered continual pre-training on a Thai-English mix, followed by general and domain-specific supervised fine-tuning, long-context adaptation, function-calling data, logit distillation for small models, and DARE-plus-linear merging for the 70B model—produces Thai LLMs that outperform their base models and existing Thai-focused baselines across exams (ThaiExam, M3Exam), instruction following (IFEval-TH, MT-Bench-TH), math, coding, and function calling. The authors further claim the same approach transfers to other modalities: a vision model built on Qwen2-VL improves Thai OCR and chart QA without losing general captioning, and an audio model built on the Typhoon2 text model achieves end-to-end speech-to-speech interaction that beats an open baseline and is competitive with a proprietary system on Thai. The conclusion states that evaluation demonstrates superior performance across a majority of evaluated tasks, including math and reasoning.

Load-bearing premise

The load-bearing premise is that the ThaiExam and M3Exam scores used to guide every pre-training and post-training decision are honest measures of Thai language ability; the paper itself notes these scores sit far above typical human performance and may reflect contamination or overfitting, and if that is true the claimed gains may not appear in real Thai text.

Editorial extensions

If this is right

  • Typhoon2-Text models from 1B to 70B improve Thai exam and instruction-following performance over their base models, with the Qwen-based 7B variant reaching the top overall BFCL function-calling score (79.08% English, 75.12% Thai).
  • The 70B model's DARE+linear merge with a newer Llama instruction model raises IFEval and MT-Bench scores beyond the unmerged SFT model.
  • Llama-based Typhoon2 models hold longer contexts up to roughly 90,000 tokens and the Qwen-based model reaches 128,000 tokens.
  • Typhoon2-Vision improves Thai OCR and ChartQA over Qwen2-VL and prior Thai vision models, while the audio model generates text and speech in parallel and performs competitively with a proprietary system on Thai speech-to-speech.
  • Typhoon2-Safety, a small mDeBERTa-based classifier, reaches higher F1 than larger Llama-based guards on Thai-sensitive topics (about 88.5–88.7) and comparable English safety performance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: if ThaiExam and M3Exam contain contamination—a concern the paper itself raises—then the headline exam gains may overstate real Thai ability; a held-out, post-cutoff Thai exam would settle this.
  • Editorial inference: the same CPT-and-post-training recipe could plausibly transfer to other low-resource Southeast Asian languages, but the paper only demonstrates it for Thai, so the transfer claim remains untested.
  • Editorial inference: the audio architecture's parallel text-and-speech decoding should lower time-to-first-speech-token latency in conversational assistants; the paper does not report this latency directly, making it a measurable but unverified benefit.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This technical report describes Typhoon 2, a family of Thai-language LLMs built by continual pre-training of Llama and Qwen backbones. It covers Thai corpus construction and data-mixture selection, post-training for instruction following, math/code, long context, function calling, distillation and model merging, a Thai safety classifier, a vision model for Thai documents, and an end-to-end audio/speech model. The central claims are that Typhoon2-Text models reach state-of-the-art Thai exam and instruction-following performance while preserving base-model abilities, that Typhoon2-Safety provides strong Thai-specific content moderation, that Typhoon2-Vision improves Thai OCR and document VQA, and that Typhoon2-Audio enables simultaneous text and speech output in an end-to-end speech-to-speech system. The paper also releases model weights and hyperparameters.

Significance. If the headline results hold, the paper is a valuable resource for Thai NLP: it publicly releases a full model family across 1B-70B sizes, documents data-filtering and post-training recipes in unusual detail, and includes a genuinely interesting small safety classifier with high F1 on Thai-sensitive topics. The audio section is also substantial, with a concrete recipe for adapting an SALMONN-style encoder and Llama-Omni-style decoder to Thai. The authors are candid about several limitations, including possible contamination of ThaiExam/M3Exam and the English weakness of the audio model. However, the central text-model claim is currently supported mainly by benchmarks that the authors themselves used as development signals and that they concede may be contaminated, so the significance will be much stronger once a held-out Thai evaluation is provided.

major comments (4)
  1. [Sections 2.3, 2.5, and Table 2] The primary evidence for Typhoon2-Text superiority is circular. Section 2.3 states that each pretraining data source and mixture was selected by checking improvements on M3Exam and/or ThaiExam, and Section 2.5 explicitly concedes that scores on these datasets are 'highly above the average level of a typical Thai person' and 'can be attributed to contamination and saturation due to overfitting.' Table 2 then uses the same ThaiExam and M3Exam numbers as the headline evidence that Typhoon2-Llama-8B-base (ThaiExam 51.20) beats Llama3.1-8B (45.80). Because the model has been optimized on these exact instruments, these gains cannot be interpreted as general Thai language ability. Please add a genuinely held-out Thai evaluation (e.g., a Thai exam or Thai benchmark not used during development, with an n-gram overlap analysis against the pretraining corpus) and present it alongside Table 2.
  2. [Table 7] The code evaluation table contains a duplicated row that makes the comparison invalid as printed. The rows for Typhoon2-Llama3.1-8B-Instruct and Qwen2.5-7B-Instruct are identical across HumanEval-TH, HumanEval-EN, MBPP-TH, and MBPP-EN (58.5, 68.9, 60.8, 63.0 in each column). This cannot be correct for two different models unless the table was copied from one row to another. Please correct the table, rerun the evaluation if necessary, and report the actual numbers.
  3. [Section 4.5 and Table 25] The claim that 'Typhoon2-Qwen2-VL excels in key areas... compared to other competitive models' is selective and not supported by the full table. On OCRBench, Typhoon2-Qwen2-VL scores ROUGE-L/Accuracy of 64.38/49.60, which is below both Llama-3.2-11B (72.84/51.10) and Typhoon2-Llama-3.2-11B (81.20/71.70). Its average ROUGE-L across the eight benchmarks (62.77) is also lower than the Typhoon2-Llama-3.2 prototype (64.16). The Qwen-based model may still be the right release choice given its size and specific Thai OCR accuracy, but the stated rationale should be revised to name the metrics where it actually wins and to acknowledge the losses on OCRBench and average ROUGE-L.
  4. [Section 6 and Tables 15-16] The conclusion that 'our evaluation demonstrates superior performance across a majority of evaluated tasks' is not supported by the full evaluation tables. In Table 16, Qwen2.5-72B-Instruct outperforms Typhoon2-70B-Instruct on MTBench-EN (9.28 vs 8.85), GSM8K-EN (94.6 vs 93.4), HumanEval-EN (87.2 vs 83.5), MBPP-EN (90.5 vs 84.9), and FC-EN (77.9 vs 65.7). Similar gaps appear at 7-8B scale in Table 15, for example HumanEval-EN 81.1 for Qwen2.5-7B vs 68.9 for Typhoon2-Llama-8B. The conclusion should be scoped to the specific tasks and languages where Typhoon2 actually leads, with the trade-offs stated explicitly.
minor comments (4)
  1. [Table 12] The 'FuncCall' column for the 1B and 3B models is marked with '?' while Tables 8-9 report BFCL results for these sizes. Please clarify whether these models were trained on function-calling data and evaluated, or evaluated without that training stage.
  2. [Section 3.2.2] The Thai math and code test sets are described as translations made with an early Typhoon2 model and GPT-4o, but these translated test sets are not released. Reporting construction details and releasing the sets would improve reproducibility and allow readers to assess translation quality.
  3. [Table 25] The notation 'Accuracy x' is not explained in the table caption, and the 'Average(Accuracy)' row appears to ignore cells marked 'x'. Please clarify how missing values are handled in the average.
  4. [Section 3.1.2] The code-switching metric counts non-Thai characters and would penalize legitimate loanwords or code-mixed Thai-English usage that is natural in context; please state this limitation explicitly.

Circularity Check

1 steps flagged · score 6.0 of 10

ThaiExam/M3Exam are both the optimization target and the headline measurement: Section 2.3 selects data mixtures by these scores, Section 2.5 concedes contamination/overfitting, and Table 2 repurposes the same scores as evidence.

  1. fitted input called prediction [Section 2.3 (Data Mixture), Section 2.5 (Evaluation), Table 2, Section 6 (Conclusions)]
    "We examine this by performing CPT on the 1.5B Qwen2.5 (Yang et al., 2024a) model using a similar recipe to Blakeney et al. (2024) and verify that each of the data sources improves one of the metrics or scores based on M3Exam (Zhang et al., 2023b) and/or ThaiExam (Pipatanakul et al., 2023). ... the scores obtained using these datasets are highly above the average level of a typical Thai person. This can be attributed to contamination and saturation due to overfitting (Fourrier et al., 2024)."

    The data-mixture selection in Section 2.3 is driven by improvements on M3Exam and ThaiExam, and Section 2.5 explicitly calls these scores a 'development signal.' The same two benchmarks are then presented in Table 2 as the main 'Exam in Thai' evidence, and the conclusion claims 'superior performance across a majority of evaluated tasks.' The reported exam gains are therefore not an independent measurement of Thai ability; they are the objective used to choose the pretraining mixture. The paper's own admission of contamination and overfitting makes the circularity explicit. This is partial rather than total circularity because instruction-following, function-calling, math/code, and other evaluations (IFEval, BFCL, GSM8K, MATH, etc.) remain independent of the development signal.

full rationale

The paper contains no formal mathematical derivation, so most of its claims are empirical reports rather than theorems. The main circular step is the use of ThaiExam and M3Exam as both the development signal (Section 2.3) and the headline evaluation evidence (Table 2). Since data sources and mixtures were verified by checking improvements on these exact benchmarks, the later 'superior performance' on those benchmarks is partly forced by the selection procedure, not independently established. The paper itself concedes that scores are 'highly above the average level of a typical Thai person' and attributes this to contamination and overfitting. This does not invalidate the many independent evaluations (IFEval, MT-Bench, BFCL, GSM8K, MATH, HumanEval, MBPP, NIAH, TTS/ASR metrics), which support the model's capabilities on their own terms. It does mean the Thai-exam and general-Thai-language claims are partially circular and should be treated as a selection-on-the-test-set result rather than a validated generalization. The safety evaluation also relies on an in-house Thai Topic test set split from the training distribution, but that is a standard held-out split and is accompanied by external benchmarks, so it adds only mild self-reference, not circularity. Overall score 6 reflects the partial circularity of the central Thai-exam claim while acknowledging the independent content elsewhere.

Assumptions & free parameters 3 free parameters · 2 assumptions · 0 invented entities

The central claims rest on empirical choices rather than derivations. The free parameters are data mixture ratios and thresholds selected by experiments; the axioms are assumptions about base model quality and evaluation validity.

free parameters (3)
  • Thai-to-English SFT data ratio = 3:7
    Chosen empirically in Section 3.1.3 to optimize Thai performance; affects all instruction-tuned models.
  • DARE merge density and weight for Llama-3.3-70B = density 0.2, weight [0.4,0.4,0.0,0.0]
    Manually searched in Section 3.6; used for the 70B instruct model.
  • Harm threshold for safety labeling = 5 (score 1-10)
    Selected in Section 3.10.1 to binarize LLM-judged harm scores; affects the safety classifier training data.
assumptions (2)
  • domain assumption Base open models (Llama 3.1/3.2, Qwen2.5/2-VL) provide a sufficiently strong foundation for Thai adaptation.
    The entire approach builds on these models without validating whether an alternative base would change the conclusions.
  • domain assumption ThaiExam and M3Exam scores are meaningful proxies for Thai language ability.
    Section 2.5 states scores are above human averages and may be affected by contamination, yet they are used as the development signal for pre-training.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models." pith.science (2026). https://pith.science/paper/4FALN3SA

@misc{pith2026241213702,
  author       = {Pith},
  title        = {Pith review of: Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4FALN3SA}},
  note         = {Machine review of arXiv:2412.13702}
}
read the original abstract

This paper introduces Typhoon 2, a series of text and multimodal large language models optimized for the Thai language. The series includes models for text, vision, and audio. Typhoon2-Text builds on state-of-the-art open models, such as Llama 3 and Qwen2, and we perform continual pre-training on a mixture of English and Thai data. We employ post-training techniques to enhance Thai language performance while preserving the base models' original capabilities. We release text models across a range of sizes, from 1 to 70 billion parameters, available in both base and instruction-tuned variants. To guardrail text generation, we release Typhoon2-Safety, a classifier enhanced for Thai cultures and language. Typhoon2-Vision improves Thai document understanding while retaining general visual capabilities, such as image captioning. Typhoon2-Audio introduces an end-to-end speech-to-speech model architecture capable of processing audio, speech, and text inputs and generating both text and speech outputs.

Figures

Figures reproduced from arXiv: 2412.13702 by the authors.

Figure 1
Figure 1. Thai Pretraining Data Mixture 7 [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Evaluation of Typhoon2-Llama3.1-8B-Instruct on Needle-in-a-Haystack for both [PITH_FULL_IMAGE:figures/full_fig_p016_2.png] view at source ↗
Figure 3
Figure 3. Evaluation of Typhoon2-Llama3.1-70B-Instruct on Needle-in-a-Haystack for both [PITH_FULL_IMAGE:figures/full_fig_p016_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Evaluation of Typhoon2-Qwen2.5-7B-Instruct on Needle-in-a-Haystack for both [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]
Figure 5
Figure 5. Figure 5: Pipeline of Thai topic data generation Initially, we investigate direct text classification through LLM prompting (e.g., “classify the following text"). However, this approach is less effective than the scoring method and lacks the flexibility to adjust model behaviors…
Figure 6
Figure 6. Figure 6: Agentic Refine Ground Truth Framework We adhere to the outlined automated design framework specific to agentic systems as docu￾mented in Hu et al. (2024) for structuring each individual agent. Thereafter, our method 30 [PITH_FULL_IMAGE:figures/full_fig_p030_6.png]
Figure 7
Figure 7. Figure 7: An Example of Agentic Refine Ground Truth [PITH_FULL_IMAGE:figures/full_fig_p032_7.png]
Figure 8
Figure 8. Figure 8: Typhoon2-Audio End-to-End Model Architecture [PITH_FULL_IMAGE:figures/full_fig_p035_8.png]
Figure 9
Figure 9. Figure 9: Speech Instruction Following Data Creation Pipeline [PITH_FULL_IMAGE:figures/full_fig_p038_9.png]
Figure 10
Figure 10. Figure 10: Text Generation (left) and Speech Generation (right) [PITH_FULL_IMAGE:figures/full_fig_p041_10.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

    cs.CL 2026-04 conditional novelty 6.0 of 10

    Synthetic data for low-resource spoken language models creates a Stability-Expressivity Gap that DGSA and TDSC self-alignment close, enabling SOTA Thai TTS and first Lao zero-shot voice cloning.

  2. XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation

    eess.AS 2025-08 conditional novelty 6.0 of 10

    XEmoRAG synthesizes Thai speech with emotions cloned from Chinese reference audio by retrieving matching Thai prompts and aligning prosody with flow matching, outperforming baseline TTS in emotion similarity and intel...

  3. Mangosteen: An Open Thai Corpus for Language Model Pretraining

    cs.CL 2025-07 conditional novelty 6.0 of 10

    An open 47B-token Thai pre-training corpus and a Thai-adapted data cleaning pipeline, with ablations showing quality gains and an 8B model that improves on Thai benchmarks.

  4. Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR

    cs.SD 2025-05 conditional novelty 6.0 of 10

    EThai-ASR combines a self-refined Zipformer encoder with a Thai LLM and reports SOTA CER on Thai test sets plus a cosine-similarity frame pruning that gives 1.5-2.1x speedups in some modes.

  5. Adapting Language-Specific LLMs to a Reasoning Model in One Day via Model Merging -- An Open Recipe

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A Thai 70B model trained with an SFT-plus-DARE-merge recipe matches DeepSeek R1 on reasoning benchmarks while retaining most Thai language quality.

  6. Typhoon T1: An Open Thai Reasoning Model

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Structured long-thinking SFT turns a 3B Thai instruct model into a reasoning model that improves on several English benchmarks and can think in Thai, with a fully open recipe.

Reference graph

Works this paper leans on

133 extracted references · 6 canonical work pages · cited by 6 Pith papers

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

  2. [2]

    Thai LLM Leaderboard , 2024

    SCB 10X, VISTEC, and SEACrowd. Thai LLM Leaderboard , 2024. URL https://huggingface.co/spaces/ThaiLLM-Leaderboard/leaderboard

  3. [3]

    Phi-3 technical report: A highly capable language model locally on your phone

    Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al. Phi-3 technical report: A highly capable language model locally on your phone . arXiv preprint arXiv:2404.14219, 2024

  4. [4]

    Evolutionary Optimization of Model Merging Recipes , 2024

    Takuya Akiba, Makoto Shing, Yujin Tang, Qi Sun, and David Ha. Evolutionary Optimization of Model Merging Recipes , 2024. URL https://arxiv.org/abs/2403.13187

  5. [5]

    Ardila, M

    R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber. C ommon V oice: A M assively- M ultilingual S peech C orpus . In Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020), 2020

  6. [6]

    Christoph Auer, Maksym Lysak, Ahmed Nassar, Michele Dolfi, Nikolaos Livathinos, Panos Vagenas, Cesar Berrospi Ramis, Matteo Omenetti, Fabian Lindlbauer, Kasper Dinkla, Lokesh Mishra, Yusik Kim, Shubham Gupta, Rafael Teixeira de Lima, Valery Weber, Lucas Morin, Ingmar Meijer, Viktor Kuropiatnyk, and Peter W. J. Staar. Docling Technical Report , 2024. URL h...

  7. [7]

    Thonburian Whisper: Robust Fine-tuned and Distilled Whisper for T hai

    Zaw Htet Aung, Thanachot Thavornmongkol, Atirut Boribalburephan, Vittavas Tangsriworakan, Knot Pipatsrisawat, and Titipat Achakulvisut. Thonburian Whisper: Robust Fine-tuned and Distilled Whisper for T hai . In Mourad Abbas and Abed Alhakim Freihat (eds.), Proceedings of the 7th ICNLSP 2024, pp.\ 149--156, Trento, October 2024. Association for Computation...

  8. [8]

    Program Synthesis with Large Language Models , 2021

    Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. Program Synthesis with Large Language Models , 2021. URL https://arxiv.org/abs/2108.07732

Show all 133 references
  1. [9]

    L ong A lign: A Recipe for Long Context Alignment of Large Language Models

    Yushi Bai, Xin Lv, Jiajie Zhang, Yuze He, Ji Qi, Lei Hou, Jie Tang, Yuxiao Dong, and Juanzi Li. L ong A lign: A Recipe for Long Context Alignment of Large Language Models . In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.), Findings of the Association for Computation...

  2. [10]

    Cosmopedia , February 2024

    Loubna Ben Allal, Anton Lozhkov, Guilherme Penedo, Thomas Wolf, and Leandro von Werra. Cosmopedia , February 2024. URL https://huggingface.co/datasets/HuggingFaceTB/cosmopedia

  3. [11]

    Larsen, Sean Owen, and Jonathan Frankle

    Cody Blakeney, Mansheej Paul, Brett W. Larsen, Sean Owen, and Jonathan Frankle. Does your data spark joy? Performance gains from domain upsampling at the end of training . In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=vwIIAot0ff

  4. [12]

    IEMOCAP : Interactive emotional dyadic motion capture database

    Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan. IEMOCAP : Interactive emotional dyadic motion capture database . Language resources and evaluation, 2008

  5. [13]

    G iga S peech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

    Guoguo Chen, Shuzhou Chai, Guanbo Wang, Jiayu Du, Wei-Qiang Zhang, Chao Weng, Dan Su, Daniel Povey, Jan Trmal, Junbo Zhang, Mingjie Jin, Sanjeev Khudanpur, Shinji Watanabe, Shuaijiang Zhao, Wei Zou, Xiangang Li, Xuchen Yao, Yongqing Wang, Yujun Wang, Zhao You, and Zhiyong Yan....

  6. [14]

    M 3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation

    Jianlyu Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. M 3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation . In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Findings of the A...

  7. [15]

    Evaluating Large Language Models Trained on Code , 2021 b

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott G...

  8. [16]

    BEATs : audio pre-training with acoustic tokenizers

    Sanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu, Daniel Tompkins, Zhuo Chen, Wanxiang Che, Xiangzhan Yu, and Furu Wei. BEATs : audio pre-training with acoustic tokenizers . In Proceedings of the 40th International Conference on Machine Learning, 2023

  9. [17]

    Towards Robust Speech Representation Learning for Thousands of Languages

    William Chen, Wangyou Zhang, Yifan Peng, Xinjian Li, Jinchuan Tian, Jiatong Shi, Xuankai Chang, Soumi Maiti, Karen Livescu, and Shinji Watanabe. Towards Robust Speech Representation Learning for Thousands of Languages . In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds...

  10. [18]

    Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models

    Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou. Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models . arXiv preprint arXiv:2311.07919, 2023

  11. [19]

    Using Knowledge Distillation from Keyword Extraction to Improve the Informativeness of Neural Cross-lingual Summarization

    Nakhun Chumpolsathien. Using Knowledge Distillation from Keyword Extraction to Improve the Informativeness of Neural Cross-lingual Summarization . Master's thesis, Beijing Institute of Technology, 2020

  12. [20]

    Training Verifiers to Solve Math Word Problems , 2021

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training Verifiers to Solve Math Word Problems , 2021. URL https://arxiv.org/abs/2110.14168

  13. [21]

    Seamless: Multilingual Expressive and Streaming Speech Translation , 2023

    Seamless Communication, Loïc Barrault, Yu-An Chung, Mariano Coria Meglioli, David Dale, Ning Dong, Mark Duppenthaler, Paul-Ambroise Duquenne, Brian Ellis, Hady Elsahar, Justin Haaheim, John Hoffman, Min-Jae Hwang, Hirofumi Inaguma, Christopher Klaiber, Ilia Kulikov, Pengwei Li...

  14. [22]

    FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech

    Alexis Conneau, Min Ma, Simran Khanuja, Yu Zhang, Vera Axelrod, Siddharth Dalmia, Jason Riesa, Clara Rivera, and Ankur Bapna. FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech . arXiv preprint arXiv:2205.12446, 2022. URL https://arxiv.org/abs/2205.12446

  15. [23]

    Efficiently Adapting Pretrained Language Models To New Languages , 2023

    Zoltan Csaki, Pian Pawakapan, Urmish Thakker, and Qiantong Xu. Efficiently Adapting Pretrained Language Models To New Languages , 2023. URL https://arxiv.org/abs/2311.05741

  16. [24]

    Safe RLHF: Safe Reinforcement Learning from Human Feedback , 2023

    Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang. Safe RLHF: Safe Reinforcement Learning from Human Feedback , 2023. URL https://arxiv.org/abs/2310.12773

  17. [25]

    Ichigo: Mixed-Modal Early-Fusion Realtime Voice Assistant

    Alan Dao, Dinh Bach Vu, and Huy Hoang Ha. Ichigo: Mixed-Modal Early-Fusion Realtime Voice Assistant . arXiv preprint arXiv:2410.15316, 2024

  18. [26]

    Unleashing Reasoning Capability of LLMs via Scalable Question Synthesis from Scratch , 2024

    Yuyang Ding, Xinyu Shi, Xiaobo Liang, Juntao Li, Qiaoming Zhu, and Min Zhang. Unleashing Reasoning Capability of LLMs via Scalable Question Synthesis from Scratch , 2024. URL https://arxiv.org/abs/2410.18693

  19. [27]

    Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models , 2024

    Guanting Dong, Keming Lu, Chengpeng Li, Tingyu Xia, Bowen Yu, Chang Zhou, and Jingren Zhou. Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models , 2024. URL https://arxiv.org/abs/2406.13542

  20. [28]

    Sailor: Open Language Models for South-East Asia , 2024

    Longxu Dou, Qian Liu, Guangtao Zeng, Jia Guo, Jiahui Zhou, Wei Lu, and Min Lin. Sailor: Open Language Models for South-East Asia , 2024. URL https://arxiv.org/abs/2404.03608

  21. [29]

    Clotho: an Audio Captioning Dataset

    Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen. Clotho: an Audio Captioning Dataset . In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.\ 736--740, 2020. doi:10.1109/ICASSP40776.2020.9052990

  22. [30]

    Llama-omni: Seamless speech interaction with large language models

    Qingkai Fang, Shoutao Guo, Yan Zhou, Zhengrui Ma, Shaolei Zhang, and Yang Feng. Llama-omni: Seamless speech interaction with large language models . arXiv preprint arXiv:2409.06666, 2024

  23. [31]

    Fsd50k: an open dataset of human-labeled sound events

    Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra. Fsd50k: an open dataset of human-labeled sound events . IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30: 0 829--852, 2021

  24. [32]

    Open LLM Leaderboard v2

    Clémentine Fourrier, Nathan Habib, Alina Lozovskaya, Konrad Szafer, and Thomas Wolf. Open LLM Leaderboard v2 . https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard, 2024

  25. [33]

    Gemmeke, Daniel P

    Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. Audio Set: An ontology and human-labeled dataset for audio events . In 2017 IEEE International Conference on Acoustics, Speech and Signal Proces...

  26. [34]

    Arcee ' s M erge K it: A Toolkit for Merging Large Language Models

    Charles Goddard, Shamane Siriwardhana, Malikeh Ehghaghi, Luke Meyers, Vladimir Karpukhin, Brian Benedict, Mark McQuade, and Jacob Solawetz. Arcee ' s M erge K it: A Toolkit for Merging Large Language Models . In Franck Dernoncourt, Daniel Preo t iuc-Pietro, and Anastasia Shimo...

  27. [35]

    The Llama 3 Herd of Models , 2024

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...

  28. [36]

    Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks

    Alex Graves, Santiago Fern \'a ndez, Faustino Gomez, and J \"u rgen Schmidhuber. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks . In Proceedings of the 23rd ICML , pp.\ 369--376, 2006

  29. [37]

    Textbooks Are All You Need , 2023

    Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Y...

  30. [38]

    WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs , 2024

    Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri. WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs , 2024. URL https://arxiv.org/abs/2406.18495

  31. [39]

    DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing , 2023

    Pengcheng He, Jianfeng Gao, and Weizhu Chen. DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing , 2023. URL https://arxiv.org/abs/2111.09543

  32. [40]

    Distilling an end-to-end voice assistant without instruction training data

    William Held, Ella Li, Michael Ryan, Weiyan Shi, Yanzhe Zhang, and Diyi Yang. Distilling an end-to-end voice assistant without instruction training data . arXiv preprint arXiv:2410.02678, 2024

  33. [41]

    Measuring Mathematical Problem Solving With the MATH Dataset

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring Mathematical Problem Solving With the MATH Dataset . In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track...

  34. [42]

    The Benefit of Temporally-Strong Labels in Audio Event Classification

    Shawn Hershey, Daniel P W Ellis, Eduardo Fonseca, Aren Jansen, Caroline Liu, R Channing Moore, and Manoj Plakal. The Benefit of Temporally-Strong Labels in Audio Event Classification . In ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processi...

  35. [43]

    Hubert: Self-supervised speech representation learning by masked prediction of hidden units

    Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. Hubert: Self-supervised speech representation learning by masked prediction of hidden units . IEEE/ACM transactions on audio, speech, and language processing, 29...

  36. [44]

    Lo RA : Low-Rank Adaptation of Large Language Models

    Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lo RA : Low-Rank Adaptation of Large Language Models . In ICLR, 2022. URL https://openreview.net/forum?id=nZeVKeeFYf9

  37. [45]

    Automated Design of Agentic Systems , 2024

    Shengran Hu, Cong Lu, and Jeff Clune. Automated Design of Agentic Systems , 2024. URL https://arxiv.org/abs/2408.08435

  38. [46]

    Siming Huang, Tianhao Cheng, J. K. Liu, Jiaran Hao, Liuyihan Song, Yang Xu, J. Yang, J. H. Liu, Chenchen Zhang, Linzheng Chai, Ruifeng Yuan, Zhaoxiang Zhang, Jie Fu, Qian Liu, Ge Zhang, Zili Wang, Yuan Qi, Yinghui Xu, and Wei Chu. OpenCoder: The Open Cookbook for Top-Tier Code...

  39. [47]

    The VoiceMOS Challenge 2022

    Wen-Chin Huang, Erica Cooper, Yu Tsao, Hsin-Min Wang, Tomoki Toda, and Junichi Yamagishi. The VoiceMOS Challenge 2022 . Proc. Interspeech 2022, 2022

  40. [48]

    Hudson and Christopher D

    Drew A. Hudson and Christopher D. Manning. GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering , 2019. URL https://arxiv.org/abs/1902.09506

  41. [49]

    Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations , 2023

    Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa. Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations , 2023. URL https://arxiv.org/abs/2312.06674

  42. [50]

    BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset , 2023

    Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Chi Zhang, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset , 2023. URL https://arxiv.org/abs/2307.04657

  43. [51]

    Bag of Tricks for Efficient Text Classification

    Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. Bag of Tricks for Efficient Text Classification . In Mirella Lapata, Phil Blunsom, and Alexander Koller (eds.), Proceedings of the 15th Conference of the E uropean Chapter of the Association for Computational L...

  44. [52]

    DVQA: Understanding Data Visualizations via Question Answering

    Kushal Kafle, Scott Cohen, Brian Price, and Christopher Kanan. DVQA: Understanding Data Visualizations via Question Answering . In CVPR, 2018

  45. [53]

    LLMTest - Needle In A Haystack , 2023

    Gregory Kamradt. LLMTest - Needle In A Haystack , 2023. URL https://github.com/gkamradt/LLMTest_NeedleInAHaystack/blob/main/README.md. GitHub repository

  46. [54]

    A Diagram Is Worth A Dozen Images , 2016

    Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A Diagram Is Worth A Dozen Images , 2016

  47. [55]

    AudioCaps: Generating Captions for Audios in The Wild

    Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim. AudioCaps: Generating Captions for Audios in The Wild . In NAACL-HLT, 2019

  48. [56]

    Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

    Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis . Advances in neural information processing systems, 33: 0 17022--17033, 2020

  49. [57]

    Shamma, Michael S

    Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Fei-Fei Li. Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations , 2...

  50. [58]

    Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D

    Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V. Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D. Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Chris Wilhelm, Luc...

  51. [59]

    DataComp- LM : In search of the next generation of training sets for language models

    Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Yitzhak Gadre, Hritik Bansal, Etash Kumar Guha, Sedrick Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee F Chen, Suchin Gururangan, Mitchell Wortsman, Alon A...

  52. [60]

    BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models . In Proceedings of the 40th International Conference on Machine Learning, ICML'23. JMLR.org, 2023 a

  53. [61]

    Ryan, Kunat Pipatanakul, Potsawee Manakul, Hao Zhu, and Diyi Yang

    Minzhi Li, Will Held, Michael J. Ryan, Kunat Pipatanakul, Potsawee Manakul, Hao Zhu, and Diyi Yang. Talk Arena: Interactive Evaluation of Large Audio Models , 2024 b

  54. [62]

    Yodas: Youtube-Oriented Dataset for Audio and Speech

    Xinjian Li, Shinnosuke Takamichi, Takaaki Saeki, William Chen, Sayaka Shiota, and Shinji Watanabe. Yodas: Youtube-Oriented Dataset for Audio and Speech . In 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pp.\ 1--8. IEEE, 2023 b

  55. [63]

    Lawrence Zitnick, and Piotr Dollár

    Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár. Microsoft COCO: Common Objects in Context , 2015. URL https://arxiv.org/abs/1405.0312

  56. [64]

    Visual Instruction Tuning , 2023 a

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual Instruction Tuning , 2023 a . URL https://arxiv.org/abs/2304.08485

  57. [65]

    Is Your Code Generated by Chat GPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation

    Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and LINGMING ZHANG. Is Your Code Generated by Chat GPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation . In Thirty-seventh Conference on Neural Information Processing Systems, 2023 b . URL https://ope...

  58. [66]

    ToolACE: Winning the Points of LLM Function Calling , 2024 a

    Weiwen Liu, Xu Huang, Xingshan Zeng, Xinlong Hao, Shuai Yu, Dexun Li, Shuai Wang, Weinan Gan, Zhengying Liu, Yuanqing Yu, Zezhong Wang, Yuxian Wang, Wu Ning, Yutai Hou, Bin Wang, Chuhan Wu, Xinzhi Wang, Yong Liu, Yasheng Wang, Duyu Tang, Dandan Tu, Lifeng Shang, Xin Jiang, Rui...

  59. [67]

    MMBench: Is Your Multi-modal Model an All-around Player? , 2024 b

    Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, and Dahua Lin. MMBench: Is Your Multi-modal Model an All-around Player? , 2024 b . URL https://arxiv.org/abs/2307.06281

  60. [68]

    OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models , 2024 c

    Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang, Wenwen Yu, Chunyuan Li, Xucheng Yin, Cheng lin Liu, Lianwen Jin, and Xiang Bai. OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models , 2024 c . URL https://arxiv.org/abs/2305.07895

  61. [69]

    APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets

    Zuxin Liu, Thai Hoang, Jianguo Zhang, Ming Zhu, Tian Lan, Shirley Kokane, Juntao Tan, Weiran Yao, Zhiwei Liu, Yihao Feng, et al. APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets . arXiv preprint arXiv:2406.18518, 2024 d

  62. [70]

    WizardCoder: Empowering Code Large Language Models with Evol-Instruct

    Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. WizardCoder: Empowering Code Large Language Models with Evol-Instruct . In The Twelfth International Conference on Learning Representations, 2024. URL http...

  63. [71]

    Enhancing low-resource language and instruction following capabilities of audio language models

    Potsawee Manakul, Guangzhi Sun, Warit Sirichotedumrong, Kasima Tharnpipitchai, and Kunat Pipatanakul. Enhancing low-resource language and instruction following capabilities of audio language models . arXiv preprint arXiv:2409.10999, 2024

  64. [72]

    ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning , 2022

    Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning , 2022. URL https://arxiv.org/abs/2203.10244

  65. [73]

    DocVQA: A Dataset for VQA on Document Images

    Minesh Mathew, Dimosthenis Karatzas, R Manmatha, and CV Jawahar. DocVQA: A Dataset for VQA on Document Images. CoRR abs/2007.00398 (2020) . arXiv preprint arXiv:2007.00398, 2020

  66. [74]

    HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal , 2024

    Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal , 2024. URL https://arxiv.or...

  67. [75]

    OCR-VQA: Visual Question Answering by Reading Text in Images

    Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty. OCR-VQA: Visual Question Answering by Reading Text in Images . In ICDAR, 2019

  68. [76]

    PathummaLLM V 1.0.0 Release

    NECTEC . PathummaLLM V 1.0.0 Release . https://medium.com/nectec/pathummallm-v-1-0-0-release-6a098ddfe276, 2024

  69. [77]

    Basic Statistical Values of O-NET Test Results

    The National Institute of Educational Testing Service. Basic Statistical Values of O-NET Test Results . https://www.niets.or.th/th/content/view/11821, 2021

  70. [78]

    Scores report of the TPAT1 exam for thai medical schools admission

    Consortium of Thai Medical Schools. Scores report of the TPAT1 exam for thai medical schools admission . https://www9.si.mahidol.ac.th/cotmes_stat.html, 2023

  71. [79]

    Basic Statistical Report TGAT/TPAT Examination

    Council of University Presidents of Thailand. Basic Statistical Report TGAT/TPAT Examination . https://www.mytcas.com/stat/, 2023

  72. [80]

    Librispeech: an asr corpus based on public domain audio books

    Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. Librispeech: an asr corpus based on public domain audio books . In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), pp.\ 5206--5210. IEEE, 2015

  73. [81]

    Patil, Tianjun Zhang, Xin Wang, and Joseph E

    Shishir G. Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez. Gorilla: Large Language Model Connected with Massive APIs . arXiv preprint arXiv:2305.15334, 2023

  74. [82]

    The RefinedWeb Dataset for Falcon LLM : Outperforming Curated Corpora with Web Data Only

    Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Hamza Alobeidli, Alessandro Cappelli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. The RefinedWeb Dataset for Falcon LLM : Outperforming Curated Corpora with Web Data Only . In Thirty-seventh Co...

  75. [83]

    The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale , 2024

    Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale , 2024. URL https://arxiv.org/abs/2406.17557

  76. [84]

    Ya RN : Efficient Context Window Extension of Large Language Models

    Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. Ya RN : Efficient Context Window Extension of Large Language Models . In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=wHBfxhZu1u

  77. [85]

    PyThaiTTS , 2022

    Wannaphong Phatthiyaphaibun. PyThaiTTS , 2022. URL https://pythainlp.org/PyThaiTTS/

  78. [86]

    Typhoon: Thai large language models

    Kunat Pipatanakul, Phatrasek Jirabovonvisut, Potsawee Manakul, Sittipong Sripaisarnmongkol, Ruangsak Patomwong, Pathomporn Chokchainant, and Kasima Tharnpipitchai. Typhoon: Thai large language models . arXiv preprint arXiv:2312.13951, 2023

  79. [87]

    WangChanGLM — The Multilingual Instruction- Following Model , April 2023

    Charin Polpanumas, Wannaphong Phatthiyaphaibun, Patomporn Payoungkhamdee, Peerat Limkonchotiwat, Lalita Lowphansirikul, Can Udomcharoenchaikit, Titipat Achakulwisut, Ekapol Chuangsuwanich, and Sarana Nutanong. WangChanGLM — The Multilingual Instruction- Following Model , April...

  80. [88]

    Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

    Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux. Speech Resynthesis from Discrete Disentangled Self-Supervised Representations . In Proc. Interspeech 2021, 2021

  81. [89]

    Scaling Speech Technology to 1,000+ Languages , 2023

    Vineel Pratap, Andros Tjandra, Bowen Shi, Paden Tomasello, Arun Babu, Sayani Kundu, Ali Elkahky, Zhaoheng Ni, Apoorv Vyas, Maryam Fazel-Zarandi, Alexei Baevski, Yossi Adi, Xiaohui Zhang, Wei-Ning Hsu, Alexis Conneau, and Michael Auli. Scaling Speech Technology to 1,000+ Langua...

  82. [90]

    UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

    Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, and Hiroshi Saruwatari. UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022 . Proc. Interspeech 2022, 2022

  83. [91]

    Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020

    Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020. URL https://arxiv.org/abs/1910.01108

  84. [92]

    openthaigpt/thai-ocr-evaluation

    Suchut Sapsathien and Jillaphat Jaroenkantasima. openthaigpt/thai-ocr-evaluation . https://huggingface.co/datasets/openthaigpt/thai-ocr-evaluation, 2024. Available online at Hugging Face

  85. [93]

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y.K. Li, Y. Wu, and Daya Guo. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models , 2024. URL https://arxiv.org/abs/2402.03300

  86. [94]

    SlimPajama-DC: Understanding Data Combinations for LLM Training , 2024

    Zhiqiang Shen, Tianhua Tao, Liqun Ma, Willie Neiswanger, Zhengzhong Liu, Hongyi Wang, Bowen Tan, Joel Hestness, Natalia Vassilieva, Daria Soboleva, and Eric Xing. SlimPajama-DC: Understanding Data Combinations for LLM Training , 2024. URL https://arxiv.org/abs/2309.10818

  87. [95]

    SEA-LION (Southeast Asian Languages In One Network): A Family of Large Language Models for Southeast Asia

    AI Singapore. SEA-LION (Southeast Asian Languages In One Network): A Family of Large Language Models for Southeast Asia . https://github.com/aisingapore/sealion, 2024

  88. [96]

    Towards VQA Models That Can Read , 2019

    Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards VQA Models That Can Read , 2019. URL https://arxiv.org/abs/1904.08920

  89. [97]

    LLM Pruning and Distillation in Practice: The Minitron Approach , 2024

    Sharath Turuvekere Sreenivas, Saurav Muralidharan, Raviraj Joshi, Marcin Chochowski, Ameya Sunil Mahabaleshwarkar, Gerald Shen, Jiaqi Zeng, Zijia Chen, Yoshi Suhara, Shizhe Diao, Chenhan Yu, Wei-Chun Chen, Hayley Ross, Oluwatobi Olabiyi, Ashwath Aithal, Oleksii Kuchaiev, Danie...

  90. [98]

    RoFormer: Enhanced Transformer with Rotary Position Embedding , 2023

    Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. RoFormer: Enhanced Transformer with Rotary Position Embedding , 2023. URL https://arxiv.org/abs/2104.09864

  91. [99]

    C ross C heck GPT : U niversal H allucination R anking for M ultimodal F oundation M odels

    Guangzhi Sun, Potsawee Manakul, Adian Liusie, Kunat Pipatanakul, Chao Zhang, Phil Woodland, and Mark Gales. C ross C heck GPT : U niversal H allucination R anking for M ultimodal F oundation M odels . arXiv preprint arXiv:2405.13684, 2024

  92. [100]

    SALMONN : Towards Generic Hearing Abilities for Large Language Models

    Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun MA, and Chao Zhang. SALMONN : Towards Generic Hearing Abilities for Large Language Models . In The Twelfth International Conference on Learning Representations, 2024 a . URL https://openreview....

  93. [101]

    MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering , 2024 b

    Jingqun Tang, Qi Liu, Yongjie Ye, Jinghui Lu, Shu Wei, Chunhui Lin, Wanqing Li, Mohamad Fitri Faiz Bin Mahmood, Hao Feng, Zhen Zhao, Yanjie Wang, Yuliang Liu, Hao Liu, Xiang Bai, and Can Huang. MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering , 2024 b . ...

  94. [102]

    Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

    Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Austin Wang, Rob Fergus, Yann LeCun, and Saining Xie. Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs . arXi...

  95. [103]

    DART -Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving

    Yuxuan Tong, Xiwen Zhang, Rui Wang, Ruidong Wu, and Junxian He. DART -Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving . In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024 b . URL https://openreview.net/forum?id=zLU21oQjD5

  96. [104]

    iapp\_wiki\_qa\_squad , February 2021

    Kobkrit Viriyayudhakorn and Charin Polpanumas. iapp\_wiki\_qa\_squad , February 2021. URL https://doi.org/10.5281/zenodo.4539916

  97. [105]

    Thai Speech Emotion Dataset , 2021

    VISTEC. Thai Speech Emotion Dataset , 2021

  98. [106]

    MT-Bench Thai , 2024

    VISTEC . MT-Bench Thai , 2024. URL https://huggingface.co/datasets/ThaiLLM-Leaderboard/mt-bench-thai

  99. [107]

    airesearch/WangchanThaiInstruct , 2024

    Vistec. airesearch/WangchanThaiInstruct , 2024. URL https://huggingface.co/datasets/airesearch/WangchanThaiInstruct

  100. [108]

    CoVoST 2 and Massively Multilingual Speech Translation

    Changhan Wang, Anne Wu, Jiatao Gu, and Juan Pino. CoVoST 2 and Massively Multilingual Speech Translation . In Proc. Interspeech 2021, pp.\ 2247--2251, 2021. doi:10.21437/Interspeech.2021-2027

  101. [109]

    OpenChat: Advancing Open-source Language Models with Mixed-Quality Data

    Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, and Yang Liu. OpenChat: Advancing Open-source Language Models with Mixed-Quality Data . In The Twelfth International Conference on Learning Representations, 2024 a . URL https://openreview.net/forum?id=AOJyfhWYHf

  102. [110]

    Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. Qwen2-VL: Enhancing Vision-Language Model's Pe...

  103. [111]

    RedPajama: an Open Dataset for Training Large Language Models

    Maurice Weber, Daniel Y Fu, Quentin Gregory Anthony, Yonatan Oren, Shane Adams, Anton Alexandrov, Xiaozhong Lyu, Huu Nguyen, Xiaozhe Yao, Virginia Adams, Ben Athiwaratkun, Rahul Chalamala, Kezhen Chen, Max Ryabinin, Tri Dao, Percy Liang, Christopher Re, Irina Rish, and Ce Zhan...

  104. [112]

    Chain-of-Thought Prompting Elicits Reasoning in Large Language Models , 2023

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models , 2023. URL https://arxiv.org/abs/2201.11903

  105. [113]

    Magicoder: Empowering Code Generation with OSS-Instruct , 2024

    Yuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding, and Lingming Zhang. Magicoder: Empowering Code Generation with OSS-Instruct , 2024. URL https://arxiv.org/abs/2312.02120

  106. [114]

    DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

    Sang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du, Hanxiao Liu, Yifeng Lu, Percy Liang, Quoc V Le, Tengyu Ma, and Adams Wei Yu. DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining . In Thirty-seventh Conference on Neural Information Processing Systems, 2023. U...

  107. [115]

    Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement , 2024

    Wenda Xu, Guanglei Zhu, Xuandong Zhao, Liangming Pan, Lei Li, and William Yang Wang. Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement , 2024. URL https://arxiv.org/abs/2402.11436

  108. [116]

    TIES -Merging: Resolving Interference When Merging Models

    Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit Bansal. TIES -Merging: Resolving Interference When Merging Models . In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=xtaX3WyCj1

  109. [117]

    Qwen2 Technical Report , 2024 a

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze...

  110. [118]

    Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement , 2024 b

    An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li, Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, Keming Lu, Mingfeng Xue, Runji Lin, Tianyu Liu, Xingzhang Ren, and Zhenru Zhang. Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Se...

  111. [119]

    Gonzalez, and Bin Cui

    Ling Yang, Zhaochen Yu, Tianjun Zhang, Shiyi Cao, Minkai Xu, Wentao Zhang, Joseph E. Gonzalez, and Bin Cui. Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models , 2024 c . URL https://arxiv.org/abs/2406.04271

  112. [120]

    Griffiths, Yuan Cao, and Karthik Narasimhan

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of Thoughts: Deliberate Problem Solving with Large Language Models , 2023. URL https://arxiv.org/abs/2305.10601

  113. [121]

    Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch , 2024

    Le Yu, Bowen Yu, Haiyang Yu, Fei Huang, and Yongbin Li. Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch , 2024. URL https://arxiv.org/abs/2311.03099

  114. [122]

    OpenThaiGPT 1.5: A Thai-Centric Open Source Large Language Model

    Sumeth Yuenyong, Kobkrit Viriyayudhakorn, Apivadee Piyatumrong, and Jillaphat Jaroenkantasima. OpenThaiGPT 1.5: A Thai-Centric Open Source Large Language Model . arXiv preprint arXiv:2411.07238, 2024

  115. [123]

    LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

    Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu. LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech . In Proc. Interspeech 2019, 2019

  116. [124]

    Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities

    Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities . arXiv preprint arXiv:2305.11000, 2023 a

  117. [125]

    M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models

    Wenxuan Zhang, Mahani Aljunied, Chang Gao, Yew Ken Chia, and Lidong Bing. M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models . In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023 b ...

  118. [126]

    M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models , 2023 c

    Wenxuan Zhang, Sharifah Mahani Aljunied, Chang Gao, Yew Ken Chia, and Lidong Bing. M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models , 2023 c . URL https://arxiv.org/abs/2306.05179

  119. [127]

    SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages , 2024

    Wenxuan Zhang, Hou Pong Chan, Yiran Zhao*, Mahani Aljunied*, Jianyu Wang*, Chaoqun Liu, Yue Deng, Zhiqiang Hu, Weiwen Xu, Yew Ken Chia, Xin Li, and Lidong Bing. SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages , 2024. URL htt...

  120. [128]

    LLaMA Beyond English: An Empirical Study on Language Capability Transfer , 2024

    Jun Zhao, Zhihao Zhang, Luhui Gao, Qi Zhang, Tao Gui, and Xuanjing Huang. LLaMA Beyond English: An Empirical Study on Language Capability Transfer , 2024. URL https://arxiv.org/abs/2401.01055

  121. [129]

    Gonzalez, and Ion Stoica

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM -as-a-Judge with MT -Bench and Chatbot Arena . In Thirty-seventh Conference on Neural In...

  122. [130]

    Instruction-Following Evaluation for Large Language Models , 2023

    Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. Instruction-Following Evaluation for Large Language Models , 2023. URL https://arxiv.org/abs/2311.07911

  123. [131]

    @esa (Ref

    \@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...

  124. [132]

    \@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...

  125. [133]

    @open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.