REVIEW 10 cited by
HyperCLOVA X Technical Report
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
HyperCLOVA X Technical Report
read the original abstract
We introduce HyperCLOVA X, a family of large language models (LLMs) tailored to the Korean language and culture, along with competitive capabilities in English, math, and coding. HyperCLOVA X was trained on a balanced mix of Korean, English, and code data, followed by instruction-tuning with high-quality human-annotated datasets while abiding by strict safety guidelines reflecting our commitment to responsible AI. The model is evaluated across various benchmarks, including comprehensive reasoning, knowledge, commonsense, factuality, coding, math, chatting, instruction-following, and harmlessness, in both Korean and English. HyperCLOVA X exhibits strong reasoning capabilities in Korean backed by a deep understanding of the language and cultural nuances. Further analysis of the inherent bilingual nature and its extension to multilingualism highlights the model's cross-lingual proficiency and strong generalization ability to untargeted languages, including machine translation between several language pairs and cross-lingual inference tasks. We believe that HyperCLOVA X can provide helpful guidance for regions or countries in developing their sovereign LLMs.
Forward citations
Cited by 10 Pith papers
-
Ko-WideSearch: A Korean Breadth-Search Benchmark for Exhaustive Set Enumeration by Web Agents
Ko-WideSearch is a new Korean breadth-search benchmark spanning 16 categories and three difficulty tiers that evaluates web agents on full set membership plus per-item attributes, showing consistent gaps between set r...
-
Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit
LLM gender stereotyping across four languages spans roughly 2.5 times the human cross-country range on HEXACO-100, with translation altering specific stereotyped attributes and effects that can compound.
-
Discovering Lexical Gaps Using Embeddings from Multilingual LLMs
A framework extracts embeddings from Korean-English bilingual LLMs across thousands of spaces and uses similarity distributions plus logistic classifiers to identify lexical gaps with AUCs of 0.81 and 0.76.
-
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
K-MetBench shows LLMs have large gaps in interpreting meteorology diagrams and Korean-specific context, with smaller local models beating much larger global ones.
-
SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models
SCRIPT is a model-agnostic injection module that enhances Korean PLM embeddings with subcharacter compositional knowledge from Jamo, leading to better performance on NLU and NLG tasks and more linguistically coherent ...
-
Cooperative Memory Paging with Keyword Bookmarks for Long-Horizon LLM Conversations
Cooperative paging replaces evicted LLM context with keyword bookmarks and adds a recall tool, outperforming six other methods on the LoCoMo benchmark across four models with statistical significance.
-
Cooperative Memory Paging with Keyword Bookmarks for Long-Horizon LLM Conversations
Cooperative paging with keyword bookmarks and a recall() tool yields the highest answer quality among six long-context methods on LoCoMo across four LLMs.
-
CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield
CHERRY combines selective ground-truth token training, recurrent depth compression from 48 to 6 layers, and mixture-of-efficient-experts to achieve competitive loss with fewer parameters on a 1.8B Korean model.
-
CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield
Selective-pivot-token training plus layer-averaging-with-recurrence reportedly gives 2.5x parameter compression on a small Korean LLM, but the efficiency claim lacks its decisive controls and the abstract advertises r...
-
SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures
SemEval-2026 Task 7 presents a benchmark and two evaluation tracks for assessing LLMs on everyday knowledge in diverse languages and cultures without allowing training on the test data.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.