Split View: 한국어를 지원하는 오픈 모델과 토크나이저 비용
한국어를 지원하는 오픈 모델과 토크나이저 비용
- 한국어 지원이라는 말은 두 가지를 뜻합니다
- 한국어를 겨냥해 만든 모델들
- 다국어 모델의 한국어 표기를 확인하는 법
- 토크나이저가 한국어를 쪼개는 방식이 곧 비용입니다
- 토큰 수를 직접 재는 법
- 컨텍스트 길이는 토큰 기준이지 글자 기준이 아닙니다
- 라이선스가 특히 갈리는 구간
- 한국어 검색을 함께 설계할 때
- 직접 해보기
- 시리즈 안내
- 참고 자료
모델 정보는 2026-08-12에 Hugging Face 페이지에서 직접 확인했습니다. 모델 카드와 라이선스는 바뀔 수 있으니 사용 전 원본을 다시 확인하세요.
한국어 지원이라는 말은 두 가지를 뜻합니다
카드에 한국어가 적혀 있다는 사실은 두 가지 중 하나입니다. 학습 데이터에 한국어 비중이 유의미하게 들어갔다는 뜻이거나, 다국어 말뭉치에 한국어가 포함되어 있다는 뜻입니다. 앞의 경우와 뒤의 경우는 실제 품질이 크게 다르지만 카드 표기만으로는 구분되지 않습니다.
거꾸로, 카드에 한국어가 없다는 사실은 훨씬 분명한 신호입니다. meta-llama/Llama-3.1-8B-Instruct 카드가 공식적으로 나열하는 언어는 영어, 독일어, 프랑스어, 이탈리아어, 포르투갈어, 힌디어, 스페인어, 태국어이고 한국어는 여기에 없습니다. meta-llama/Llama-3.2-1B-Instruct도 같은 목록을 쓰며, 명시적으로 지원한다고 언급한 언어 외의 사용을 범위 밖으로 적어 둡니다.
한국어를 겨냥해 만든 모델들
| 저장소 | license | 파라미터 | 컨텍스트 | 카드에 적힌 언어 |
|---|---|---|---|---|
LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct | exaone | 임베딩 제외 6.98B | 32,768 | 영어, 한국어 |
LGAI-EXAONE/EXAONE-4.0-32B | exaone | 임베딩 제외 30.95B | 131,072 | 영어, 한국어, 스페인어 |
K-intelligence/Midm-2.0-Base-Instruct | MIT | 11.5B | 명시되어 있지 않음 | 영어, 한국어 |
K-intelligence/Midm-2.0-Mini-Instruct | MIT | 2.3B | 명시되어 있지 않음 | 영어, 한국어 |
kakaocorp/kanana-nano-2.1b-instruct | cc-by-nc-4.0 | 2.1B | 명시되어 있지 않음 | 영어, 한국어 |
naver-hyperclovax/HyperCLOVAX-SEED-Text-Instruct-1.5B | hyperclovax-seed | 1.5B | 16k | 명시되어 있지 않음 |
Bllossom/llama-3.2-Korean-Bllossom-3B | llama3.2 | 3B | 명시되어 있지 않음 | 한국어, 영어 |
upstage/SOLAR-10.7B-Instruct-v1.0 | cc-by-nc-4.0 | 약 10.7B | 명시되어 있지 않음 | 영어 |
몇 가지는 카드를 읽어야만 알 수 있는 사실입니다. K-intelligence/Midm-2.0-Mini-Instruct는 온디바이스 환경과 GPU 자원이 제한된 시스템에 최적화했다고 적고, 같은 계열의 K-intelligence/Midm-2.0-Base-Instruct는 한국 사회의 가치와 상식을 내재화한 한국 중심 AI를 지향한다고 밝힙니다. 두 카드 모두 학습이 영어와 한국어에 한정되어 다른 언어의 성능은 보장하지 않으며, 법률·의료·금융처럼 전문성이 필요한 분야에서 신뢰할 만한 조언을 보장하지 않는다고 적습니다.
naver-hyperclovax/HyperCLOVAX-SEED-Text-Instruct-1.5B는 접근하려면 연락처 공유에 동의해야 하고, 카드는 강화 학습 이전 단계의 작은 모델에서 긴 컨텍스트 처리에 취약한 점이 관찰되었다고 적습니다. upstage/SOLAR-10.7B-Instruct-v1.0은 이름 때문에 한국어 모델로 오해되기 쉽지만 카드가 적은 지원 언어는 영어이고, 단일 턴 대화 위주로 미세조정되어 여러 턴 대화에는 덜 적합하다고 명시합니다.
다국어 모델의 한국어 표기를 확인하는 법
한국어 전용 모델만 후보가 되는 것은 아닙니다. microsoft/Phi-4-mini-instruct 카드는 지원 언어 목록에 한국어를 명시적으로 포함하고, mistralai/Mistral-Small-24B-Instruct-2501도 한국어를 나열합니다. Qwen/Qwen3-8B은 100개 이상, google/gemma-3-4b-it은 140개 이상 언어를 지원한다고 적습니다.
여기서 갈리는 것은 표기의 구체성입니다. 언어를 하나하나 나열한 카드는 그 언어를 의식하고 평가했을 가능성이 높고, 개수만 적은 카드는 말뭉치에 포함되었다는 것 이상을 말하지 않습니다. 어느 쪽이든 우리 도메인 문장으로 직접 재는 과정을 대체하지는 못합니다.
토크나이저가 한국어를 쪼개는 방식이 곧 비용입니다
한국어를 쓸 때 성능 다음으로 중요한 것이 토큰 효율입니다. 이유는 단순합니다. API 요금도, 컨텍스트 한도도, 생성 속도도 전부 글자가 아니라 토큰을 기준으로 매겨지기 때문입니다.
영어는 대부분의 BPE 토크나이저에서 단어 하나가 토큰 한두 개로 처리됩니다. 한국어는 조사와 어미가 붙는 교착어라서 같은 어간이 문맥마다 다른 형태로 나타나고, 그 결과 같은 의미를 담은 문장이 더 많은 토큰으로 쪼개지는 경우가 흔합니다. 이 차이는 세 곳에서 동시에 비용이 됩니다. 요청당 입력 토큰이 늘고, 컨텍스트 창이 더 빨리 차고, 출력 토큰이 늘어 지연이 커집니다.
어휘 크기가 이 문제에 영향을 주기는 하지만 단순 비례하지는 않습니다. Qwen/Qwen3-8B의 config.json에는 vocab_size가 151936으로 들어 있습니다. 어휘가 크면 문장을 더 적은 토큰으로 표현할 여지가 생기지만, 그 어휘가 어떤 언어에 배분되었는지는 숫자만으로 알 수 없습니다. 결국 재 보는 것 말고는 방법이 없습니다.
토큰 수를 직접 재는 법
# 예시: 같은 문장을 여러 토크나이저로 쪼개 비교합니다
from transformers import AutoTokenizer
repos = [
"Qwen/Qwen3-8B",
"microsoft/Phi-4-mini-instruct",
"K-intelligence/Midm-2.0-Mini-Instruct",
]
samples = [
"연차 이월 규정은 취업규칙 제12조에 따라 운영됩니다.",
"Annual leave carryover follows Article 12 of the employment rules.",
]
for repo in repos:
tok = AutoTokenizer.from_pretrained(repo)
counts = [len(tok.encode(s, add_special_tokens=False)) for s in samples]
print(repo, "ko:", counts[0], "en:", counts[1], "vocab:", tok.vocab_size)
이 스크립트를 우리 서비스의 실제 문장 200개 정도로 돌리면, 후보 모델별 한국어 토큰 비용이 바로 나옵니다. 리더보드 점수보다 훨씬 빨리 얻을 수 있고, 훨씬 오래 유효한 숫자입니다.
게이트가 걸린 저장소는 이 코드가 접근 토큰 없이는 실패합니다. meta-llama 계열과 google/gemma-3-27b-it, naver-hyperclovax/HyperCLOVAX-SEED-Text-Instruct-1.5B가 모두 조건 동의 또는 연락처 공유를 요구한다고 페이지에 적고 있습니다.
컨텍스트 길이는 토큰 기준이지 글자 기준이 아닙니다
LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct의 32,768과 LGAI-EXAONE/EXAONE-4.0-32B의 131,072는 모두 토큰 수입니다. 같은 한도라도 토큰 효율이 다른 두 모델은 실제로 들어가는 한국어 분량이 다릅니다.
그래서 긴 문서를 다루는 서비스에서는 컨텍스트 한도를 비교할 때 반드시 앞 절의 측정을 함께 붙여야 합니다. 카드에 적힌 숫자가 더 크다고 해서 우리 문서가 더 많이 들어간다는 보장은 없습니다. 컨텍스트 길이가 아예 명시되어 있지 않은 모델도 있습니다. K-intelligence 계열과 kakaocorp/kanana-nano-2.1b-instruct 페이지에는 컨텍스트 길이가 적혀 있지 않습니다.
라이선스가 특히 갈리는 구간
한국어 모델은 라이선스 유형이 유난히 다양합니다. K-intelligence/Midm-2.0-Base-Instruct와 K-intelligence/Midm-2.0-Mini-Instruct는 MIT로 표기됩니다. kakaocorp/kanana-nano-2.1b-instruct는 cc-by-nc-4.0이고 카드가 비상업 조건을 명시합니다. upstage/SOLAR-10.7B-Instruct-v1.0도 cc-by-nc-4.0입니다.
EXAONE 계열은 자체 라이선스를 씁니다. LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct는 EXAONE AI Model License Agreement 1.1 - NC, LGAI-EXAONE/EXAONE-4.0-32B는 1.2 - NC로 표기되며, 후자의 페이지에는 EXAONE과 경쟁하는 모델 개발에 사용하는 것을 제한한다는 취지가 적혀 있습니다. naver-hyperclovax/HyperCLOVAX-SEED-Text-Instruct-1.5B는 hyperclovax-seed라는 자체 식별자를 쓰고 상업적 사용 조건은 페이지에 별도로 명시되어 있지 않습니다. Bllossom/llama-3.2-Korean-Bllossom-3B는 llama3.2를 따르며 카드에 상업적 이용이 가능하다는 문장이 적혀 있습니다.
라이선스 전문을 직접 읽고, 상업적 사용은 법무 검토를 거치세요. 이 글은 어떤 모델이 상업적으로 사용 가능한지 판단해 주지 않으며, 카드에 적힌 문장을 옮길 뿐입니다.
한국어 검색을 함께 설계할 때
한국어 RAG를 만든다면 생성 모델만으로는 절반입니다. 한국어를 겨냥해 추가 학습한 임베딩으로 nlpai-lab/KURE-v1, nlpai-lab/KoE5, dragonkue/BGE-m3-ko가 있고, 각각의 차원과 최대 입력 길이는 시리즈 3편에 정리해 두었습니다. 특히 nlpai-lab/KoE5는 기반 모델의 512 토큰 제약을 그대로 물려받으므로, 토큰 효율이 나쁜 조합에서는 한 문단도 다 들어가지 않을 수 있습니다.
직접 해보기
- LLM API 비용 계산기 — 토큰 수 측정 결과를 넣어 한국어 워크로드의 실제 비용을 계산합니다.
- LLM GPU 메모리(VRAM) 계산기 — 후보 모델이 보유 장비에 올라가는지 확인합니다.
- 프롬프트 엔지니어링 가이드 — 같은 지시를 더 적은 토큰으로 쓰는 연습을 합니다.
시리즈 안내
- 이전 글: 임베딩과 리랭커, RAG에서 실제로 중요한 것
- 다음 글: 음성 모델 고르기: STT와 TTS의 실전 기준
참고 자료
- 표의 모든 값은 2026-08-12에 해당 모델의 Hugging Face 페이지에서 직접 읽었습니다. 페이지에 없던 항목은 명시되어 있지 않음으로 적었습니다.
- 토큰 수 비교는 이 글에 수치를 싣지 않았습니다. 문장 구성에 따라 크게 달라지므로 각자의 데이터로 측정하는 편이 정확합니다.
- 라이선스 전문을 직접 읽고, 상업적 사용은 법무 검토를 거치세요.
Open Models That Support Korean, and the Cost of Tokenization
- Supporting Korean Means One of Two Things
- Models Built with Korean in Mind
- How to Check the Korean Line on a Multilingual Card
- How a Tokenizer Splits Korean Is the Cost
- How to Measure Token Counts Yourself
- Context Length Is Counted in Tokens, Not Characters
- Where Licenses Diverge Most
- When You Design Korean Retrieval Alongside
- Try It Yourself
- Series Navigation
- References
Model details were read directly from the Hugging Face pages on 2026-08-12. Model cards and licenses change, so check the original again before you use anything.
Supporting Korean Means One of Two Things
A card listing Korean means one of two things: that Korean made up a meaningful share of the training data, or that Korean was present somewhere in a multilingual corpus. The real quality gap between those two cases is large, and the card wording does not distinguish them.
The reverse, however, is a much clearer signal. The languages the meta-llama/Llama-3.1-8B-Instruct card officially lists are English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai — Korean is not among them. meta-llama/Llama-3.2-1B-Instruct uses the same list and marks use in languages beyond those explicitly referenced as out of scope.
Models Built with Korean in Mind
| Repository | license | Parameters | Context | Languages as stated |
|---|---|---|---|---|
LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct | exaone | 6.98B without embeddings | 32,768 | English, Korean |
LGAI-EXAONE/EXAONE-4.0-32B | exaone | 30.95B without embeddings | 131,072 | English, Korean, Spanish |
K-intelligence/Midm-2.0-Base-Instruct | MIT | 11.5B | Not stated | English, Korean |
K-intelligence/Midm-2.0-Mini-Instruct | MIT | 2.3B | Not stated | English, Korean |
kakaocorp/kanana-nano-2.1b-instruct | cc-by-nc-4.0 | 2.1B | Not stated | English, Korean |
naver-hyperclovax/HyperCLOVAX-SEED-Text-Instruct-1.5B | hyperclovax-seed | 1.5B | 16k | Not stated |
Bllossom/llama-3.2-Korean-Bllossom-3B | llama3.2 | 3B | Not stated | Korean, English |
upstage/SOLAR-10.7B-Instruct-v1.0 | cc-by-nc-4.0 | About 10.7B | Not stated | English |
Several facts only surface if you read the card. K-intelligence/Midm-2.0-Mini-Instruct states it is optimized for on-device environments and systems with limited GPU resources, and its sibling K-intelligence/Midm-2.0-Base-Instruct describes itself as Korea-centric AI that internalizes the values and commonsense reasoning of Korean society. Both cards state that training was limited to English and Korean so other languages are not guaranteed, and that the model is not guaranteed to give reliable advice in fields requiring professional expertise such as law, medicine, or finance.
naver-hyperclovax/HyperCLOVAX-SEED-Text-Instruct-1.5B requires agreeing to share contact information for access, and the card states that vulnerability to long-context handling was observed in smaller models prior to reinforcement training. upstage/SOLAR-10.7B-Instruct-v1.0 is easily mistaken for a Korean model because of its origin, but the language its card states is English, and it notes the model was fine-tuned primarily for single-turn conversation, making it less suitable for multi-turn chat.
How to Check the Korean Line on a Multilingual Card
Korean-focused models are not the only candidates. The microsoft/Phi-4-mini-instruct card explicitly includes Korean in its language list, and mistralai/Mistral-Small-24B-Instruct-2501 lists Korean as well. Qwen/Qwen3-8B states more than 100 languages and google/gemma-3-4b-it more than 140.
What separates them is the specificity of the claim. A card that enumerates languages one by one probably had them in view during evaluation; a card that states only a count says nothing beyond corpus inclusion. Neither replaces measuring with sentences from your own domain.
How a Tokenizer Splits Korean Is the Cost
After quality, the thing that matters most for Korean is token efficiency, for a simple reason: API pricing, context limits, and generation speed are all denominated in tokens rather than characters.
In most BPE tokenizers an English word costs one or two tokens. Korean is agglutinative — particles and endings attach to stems, so the same stem appears in different surface forms depending on context, and a sentence carrying the same meaning frequently splits into more tokens. That difference becomes cost in three places at once: input tokens per request rise, the context window fills faster, and output tokens grow, which raises latency.
Vocabulary size affects this but not proportionally. The config.json for Qwen/Qwen3-8B carries a vocab_size of 151936. A larger vocabulary creates room to express a sentence in fewer tokens, but how that vocabulary was allocated across languages cannot be read off the number. In the end there is no substitute for measuring.
How to Measure Token Counts Yourself
# Example: split the same sentence with several tokenizers and compare
from transformers import AutoTokenizer
repos = [
"Qwen/Qwen3-8B",
"microsoft/Phi-4-mini-instruct",
"K-intelligence/Midm-2.0-Mini-Instruct",
]
samples = [
"연차 이월 규정은 취업규칙 제12조에 따라 운영됩니다.",
"Annual leave carryover follows Article 12 of the employment rules.",
]
for repo in repos:
tok = AutoTokenizer.from_pretrained(repo)
counts = [len(tok.encode(s, add_special_tokens=False)) for s in samples]
print(repo, "ko:", counts[0], "en:", counts[1], "vocab:", tok.vocab_size)
Run this over about 200 real sentences from your own service and you get the Korean token cost of each candidate immediately. It is a number you can obtain much faster than a leaderboard score, and it stays valid much longer.
For gated repositories this code fails without an access token. The meta-llama family, google/gemma-3-27b-it, and naver-hyperclovax/HyperCLOVAX-SEED-Text-Instruct-1.5B all state on their pages that accepting conditions or sharing contact information is required.
Context Length Is Counted in Tokens, Not Characters
The 32,768 of LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct and the 131,072 of LGAI-EXAONE/EXAONE-4.0-32B are both token counts. Two models with the same limit but different token efficiency hold different amounts of actual Korean text.
So for a service handling long documents, any comparison of context limits has to carry the measurement from the previous section along with it. A bigger number on the card is no guarantee that more of your documents fit. Some models do not state a context length at all: the K-intelligence pages and kakaocorp/kanana-nano-2.1b-instruct do not.
Where Licenses Diverge Most
Korean models carry an unusually wide spread of license types. K-intelligence/Midm-2.0-Base-Instruct and K-intelligence/Midm-2.0-Mini-Instruct are marked MIT. kakaocorp/kanana-nano-2.1b-instruct is cc-by-nc-4.0 and its card states the non-commercial condition. upstage/SOLAR-10.7B-Instruct-v1.0 is cc-by-nc-4.0 as well.
The EXAONE family uses its own license. LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct is marked EXAONE AI Model License Agreement 1.1 - NC and LGAI-EXAONE/EXAONE-4.0-32B as 1.2 - NC, and the latter page states that the license restricts use toward developing models that compete with EXAONE. naver-hyperclovax/HyperCLOVAX-SEED-Text-Instruct-1.5B uses its own hyperclovax-seed identifier, and commercial terms are not separately stated on the page. Bllossom/llama-3.2-Korean-Bllossom-3B follows llama3.2, and its card contains a sentence saying commercial use is possible.
Read the full license text yourself and put commercial use through legal review. This post does not decide which model you may use commercially; it only reports what the cards say.
When You Design Korean Retrieval Alongside
For Korean RAG, the generation model is only half the job. Embeddings trained further for Korean include nlpai-lab/KURE-v1, nlpai-lab/KoE5, and dragonkue/BGE-m3-ko, whose dimensions and maximum input lengths are laid out in part 3 of this series. Note in particular that nlpai-lab/KoE5 inherits the 512-token limit of its base model, so with a poor token-efficiency pairing even a single paragraph may not fit.
Try It Yourself
- LLM API Cost Calculator — feed in your token measurements to price a Korean workload.
- GPU VRAM Calculator for LLMs — check whether the candidates fit the hardware you have.
- Prompt Engineering Guide — practice writing the same instruction in fewer tokens.
Series Navigation
- Previous: Embeddings and Rerankers: What Actually Matters in RAG
- Next: Choosing Speech Models: Practical Criteria for STT and TTS
References
- Every value in the tables was read directly from that model page on Hugging Face on 2026-08-12. Anything absent from the page is written as not stated.
- No token-count figures are published in this post. They vary heavily with sentence composition, so measuring on your own data is more accurate.
- Read the full license text yourself and put commercial use through legal review.