Skip to content

Split View: 음성 인식·합성 기술 리포트, 무엇을 읽을 것인가 — WER 한 숫자로는 안 보이는 것

|

음성 인식·합성 기술 리포트, 무엇을 읽을 것인가 — WER 한 숫자로는 안 보이는 것

들어가며

시리즈 4편은 음성입니다. 이 영역은 단일 숫자에 가장 심하게 속는 곳입니다. WER 하나로 모델을 고르면, 배포하고 나서 억양과 잡음과 고유명사에서 무너집니다.

재미있게도 이 문제를 가장 솔직하게 적어 둔 것이 최신 리포트 중 하나입니다. 뒤에 나올 Qwen3-ASR 초록에는 공개 벤치마크 점수가 거의 차이 나지 않아도 실제 환경에서는 품질 차이가 크게 드러날 수 있다는 문장이 있습니다. 이 시리즈가 하려는 말을 저자들이 대신 해 준 셈입니다.

논문 정보는 2026-08-12에 원문에서 직접 확인했습니다. 이 분야는 빠르게 바뀌므로 최신 상태는 직접 확인하세요.

이 글이 우열을 가리지 않는 이유

ASR 평가는 텍스트 정규화 규칙 하나로 WER이 바뀝니다. 숫자 표기, 약어, 구두점, 대소문자를 어떻게 처리하느냐가 점수에 직접 들어갑니다. TTS는 더합니다. 자연스러움 평가가 결국 사람 청취 실험이라 평가자 구성과 안내문에 따라 결과가 움직입니다.

그래서 이 글은 점수 대신 지연, 언어 커버리지, 스트리밍 가능 여부, 필요한 데이터 규모를 읽습니다. 이 값들은 하네스가 달라도 잘 변하지 않습니다.

음성에서 지금 실제로 다투는 것

[음성 리포트가 다투는 축]
 커버리지 : 몇 개 언어를, 얼마나 적은 데이터로
 지연     : 첫 소리가 나오기까지 몇 밀리초
 스트리밍 : 문장을 다 받고 시작하는가, 받으면서 시작하는가
 이중성   : 말하면서 듣는 전이중 대화가 되는가
 통합     : 인식·합성·이해를 한 모델에 넣을 것인가

읽을 만한 기술 리포트 열 편

Robust Speech Recognition via Large-Scale Weak Supervision (Whisper)

arxiv.org/abs/2212.04356 — 2022-12-06 등록.

68만 시간의 다국어·다중작업 지도로 학습해, 파인튜닝 없는 제로샷 전이만으로 이전의 완전 지도 방식과 경쟁한다고 보고한 논문입니다. 이후 거의 모든 ASR 리포트가 이 논문을 기준선으로 삼기 때문에, 최신 리포트를 읽기 전에 한 번은 원문을 봐야 합니다.

Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages

arxiv.org/abs/2511.09690 — 2025-11-12 등록.

자기지도 사전학습을 7B까지 키우고, 제로샷 일반화를 겨냥한 인코더-디코더 구조를 씁니다. 1,600개 이상 언어를 다루며 그중 500개 이상은 이전에 ASR가 제공되지 않던 언어라고 밝힙니다. 공동체가 소수의 샘플만으로 새 언어를 추가할 수 있게 한 점이 설계의 핵심입니다. 그리고 이 리포트는 자기 한계를 분명히 적습니다. 세계의 7,000개가 넘는 언어 중 대부분은 여전히 지원되지 않습니다.

Qwen3-ASR Technical Report

arxiv.org/abs/2601.21337 — 2026-01-29 등록, v2는 2026-01-30.

1.7B와 0.6B 두 모델에 정렬 도구를 더한 구성입니다. 52개 언어와 방언의 언어 식별과 인식을 지원하고, 작은 모델의 지연은 92밀리초까지 내려간다고 보고합니다. 여기서 인용할 가치가 있는 것은 점수가 아니라 저자들의 서술입니다. 공개 벤치마크에서 거의 차이가 없어 보이는 모델들이 실제 환경에서는 품질이 크게 갈릴 수 있다는 관찰입니다. Apache 2.0으로 공개했습니다.

Open ASR Leaderboard

arxiv.org/abs/2510.06961 — 2025-10-08 등록, v4는 2026-03-30.

이 목록에서 모델이 아닌 항목이고, ASR를 고르는 사람이라면 먼저 읽어야 할 항목입니다. 공개와 상용을 합쳐 86개 시스템을 12개 데이터셋에서 비교하며, WER과 RTFx를 서로 다른 구조와 툴킷에 걸쳐 표준화합니다. 저자들이 보고한 경향 하나는 유용합니다. 컨포머 인코더와 트랜스포머 디코더 조합이 오류율에서, CTC와 TDT 디코더가 효율 지표에서 유리하다는 것입니다. 코드와 데이터셋 로더를 공개해 재현이 가능하다는 점이 이 작업의 본질입니다.

Seed-TTS와 F5-TTS

arxiv.org/abs/2406.02430 (2024-06-04 등록)과 arxiv.org/abs/2410.06885 (2024-10-09 등록, v3는 2025-05-20).

앞의 것은 자기회귀 계열로, 화자 유사도와 자연스러움에서 실제 사람 음성 수준에 도달했다고 보고하며 자기 증류와 강화학습으로 견고성을 올립니다. 뒤의 것은 흐름 정합 기반 비자기회귀 계열로, 실시간 계수 0.15와 코드 스위칭, 속도 제어를 내세웁니다. F5-TTS 초록은 선행 모델의 느린 수렴과 낮은 견고성을 개선 동기로 명시합니다.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

arxiv.org/abs/2412.10117 — 2024-12-13 등록, v3는 2024-12-25.

유한 스칼라 양자화로 음성 토큰의 코드북 활용률을 높이고, 사전학습된 LLM을 백본으로 바로 쓰도록 구조를 단순화했습니다. 청크 단위 인과 흐름 정합 모델로 스트리밍과 비스트리밍을 함께 지원합니다. 문서 자체가 진행 중인 기술 보고서라고 밝히고 있으니, 수치를 확정된 결과로 읽지 않는 편이 좋습니다.

Qwen3-TTS Technical Report

arxiv.org/abs/2601.15621 — 2026-01-22 등록.

10개 언어, 500만 시간 이상의 음성으로 학습했고 3초 음성 복제와 설명 기반 제어를 지원합니다. 토크나이저를 두 개 둔 설계가 특징입니다. 하나는 의미 중심, 다른 하나는 초저지연 스트리밍용이며 첫 패킷까지 97밀리초를 보고합니다. Apache 2.0으로 공개했습니다.

Fish Audio S2 Technical Report

arxiv.org/abs/2603.08823 — 2026-03-09 등록, v2는 2026-03-11.

다화자 다중 턴 생성과 자연어 설명을 통한 지시 따르기 제어를 다룹니다. 캡셔닝과 품질 평가, 보상 모델링을 포함한 다단계 학습 파이프라인을 설명하며, 실시간 계수 0.195와 첫 오디오까지 100밀리초 미만을 보고합니다. 코드와 가중치, 파인튜닝 자료를 공개했습니다.

Moshi: a speech-text foundation model for real-time dialogue

arxiv.org/abs/2410.00037 — 2024-09-17 등록, v2는 2024-10-02.

문제 정의가 다른 리포트입니다. 부품을 이어 붙인 대화 파이프라인의 수 초 지연, 텍스트를 중간 표현으로 쓸 때의 정보 손실, 겹쳐 말하기와 끼어들기를 다루지 못하는 한계를 함께 지적합니다. 자기 음성과 사용자 음성을 병렬 스트림으로 모델링하고, 오디오 토큰 전에 텍스트 토큰을 예측하는 내적 독백 기법을 씁니다. 이론 지연 160밀리초, 실제 200밀리초를 보고합니다.

Qwen2.5-Omni와 MOSS-Audio

arxiv.org/abs/2503.20215 (2025-03-26 등록)과 arxiv.org/abs/2606.01802 (2026-06-01 등록, v3는 2026-06-05).

앞의 것은 텍스트와 이미지와 오디오와 비디오를 받아 텍스트와 음성을 스트리밍으로 내놓습니다. 오디오와 비디오의 시간 정렬을 위한 TMRoPE, 텍스트 생성과 음성 생성의 간섭을 줄이는 Thinker-Talker 구조가 핵심입니다. 뒤의 것은 음성과 환경음과 음악을 함께 이해하는 오디오 언어 모델로, 12.5Hz 시간 표현과 여러 인코더 깊이의 음향 정보를 디코더에 주입하는 방식, 시간 표식 삽입을 씁니다. 저자들은 이를 향후 음성 에이전트를 위한 이해 기반이라고 위치 짓습니다. 대화까지 함께 다루는 Step-Audio 2(arxiv.org/abs/2507.16632)도 같은 계열입니다.

저자들이 스스로 밝힌 한계

  • Omnilingual ASR은 세계 언어 대부분이 여전히 미지원이라고 초록에서 직접 밝히고 윤리적 고려도 함께 언급합니다.
  • Qwen3-ASR은 공개 벤치마크 점수 차이가 작아도 실제 환경 품질은 크게 다를 수 있다고 적습니다. 이 문장은 모든 ASR 선택에 적용해야 합니다.
  • CosyVoice 2는 문서 스스로를 진행 중인 작업으로 규정합니다.
  • F5-TTS는 선행 모델의 수렴 속도와 견고성 문제를 개선 동기로 명시합니다.
  • MOSS-Audio는 스스로를 완결된 시스템이 아니라 향후 음성 에이전트의 이해 기반으로 위치 짓습니다.
  • 음성 복제 기술은 사칭 위험을 동반합니다. 리포트가 성능을 말할 때 이 부분까지 말하는 경우는 드물다는 점을 기억하세요.

실무자는 무엇부터 읽어야 하나

ASR 모델을 고른다면 Open ASR Leaderboard부터 보세요. 개별 리포트의 자체 보고보다 표준화된 비교가 먼저입니다.

저자원 언어가 대상이라면 Omnilingual ASR입니다. 소수 샘플로 새 언어를 붙이는 설계가 실제 사용 시나리오를 바꿉니다.

실시간 대화를 만든다면 Moshi와 Qwen3-TTS를 붙여 읽으세요. 전이중 대화의 구조 문제와 첫 패킷 지연이라는 서로 다른 병목을 각각 다룹니다.

한 모델에 여러 양식을 묶으려 한다면 Qwen2.5-Omni와 MOSS-Audio가 설계 참고가 됩니다.

확인 방법과 시점

이 글에 등장하는 리포트는 모두 2026-08-12에 arXiv 초록 페이지를 직접 열어 확인했습니다. 열지 못한 후보는 인용하지 않았습니다. 지연과 실시간 계수를 포함한 모든 수치는 저자 보고이며 하드웨어와 설정에 따라 달라집니다.

직접 해보기

  • ML 학습 데이터 탐색기 — 음성 인식과 합성 모델이 각각 어떤 형태의 학습 데이터를 쓰는지 예시로 볼 수 있습니다.
  • 브라우저 AI 실험실 — 서버 없이 브라우저에서 모델을 돌려 보면, 리포트가 말하는 지연이라는 값이 무엇을 뜻하는지 감이 옵니다.

시리즈 이전 편: 비디오 생성·이해 기술 리포트, 무엇을 읽을 것인가

시리즈 다음 편: 이미지 생성·이해 기술 리포트, 무엇을 읽을 것인가

참고 자료

Speech Recognition and Synthesis Technical Reports: What to Read, and What a Single WER Hides

Introduction

Part four is speech. This is the domain where a single number misleads hardest. Pick a model by WER alone and it falls apart in production on accents, noise and proper nouns.

The most honest statement of that problem happens to sit in one of the newest reports. The Qwen3-ASR abstract, covered below, notes that ASR models can differ very little on open benchmark scores while showing significant quality differences in real-world scenarios. The authors said this series' thesis for me.

Paper details were verified directly from the primary sources on 2026-08-12. This field moves fast, so check the current state yourself.

Why this article does not declare a winner

In ASR evaluation, a single text normalization rule moves WER. How you handle numerals, abbreviations, punctuation and casing goes straight into the score. TTS is worse: naturalness ultimately comes from human listening tests, so results shift with the rater pool and the instructions they were given.

So this article reads latency, language coverage, streaming capability and required data scale instead of scores. Those values hold up across harnesses.

What speech reports are actually contesting

[The axes speech reports contest]
 Coverage  : how many languages, from how little data
 Latency   : milliseconds until the first sound
 Streaming : start after the full sentence, or while receiving it
 Duplex    : can it listen while speaking
 Unification: one model for recognition, synthesis and understanding

Ten technical reports worth reading

Robust Speech Recognition via Large-Scale Weak Supervision (Whisper)

arxiv.org/abs/2212.04356 — submitted 2022-12-06.

Trained on 680,000 hours of multilingual and multitask supervision, reporting that zero-shot transfer without any fine-tuning competes with prior fully supervised approaches. Nearly every ASR report since uses this as a baseline, so it is worth reading the original before the recent ones.

Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages

arxiv.org/abs/2511.09690 — submitted 2025-11-12.

Self-supervised pre-training scaled to 7B with an encoder-decoder architecture designed for zero-shot generalization. It covers more than 1,600 languages, over 500 of which the authors state were never before served by ASR. The design point is letting communities introduce an unserved language with only a handful of samples. And the report states its limit plainly: most of the world's 7,000-plus languages remain unsupported.

Qwen3-ASR Technical Report

arxiv.org/abs/2601.21337 — submitted 2026-01-29, v2 on 2026-01-30.

Two models, 1.7B and 0.6B, plus a timestamp alignment tool. Language identification and recognition for 52 languages and dialects, with latency down to 92 ms for the smaller model. The part worth quoting is not a score but the authors' own observation: models that look nearly identical on open benchmark scores can diverge substantially in real deployments. Released under Apache 2.0.

Open ASR Leaderboard

arxiv.org/abs/2510.06961 — submitted 2025-10-08, v4 on 2026-03-30.

Not a model, and the first thing anyone choosing an ASR system should read. It compares 86 open-source and proprietary systems across 12 datasets, standardizing WER and RTFx across differing architectures and toolkits. One reported trend is useful on its own: Conformer encoders with transformer decoders come out ahead on error rate, while CTC and TDT decoders lead on efficiency. Open-sourcing the code and dataset loaders so the comparison is reproducible is the substance of this work.

Seed-TTS and F5-TTS

arxiv.org/abs/2406.02430 (submitted 2024-06-04) and arxiv.org/abs/2410.06885 (submitted 2024-10-09, v3 on 2025-05-20).

The first is autoregressive, reporting speaker similarity and naturalness matching ground truth human speech, with self-distillation and reinforcement learning added for robustness. The second is non-autoregressive flow matching, foregrounding a 0.15 real-time factor, code switching and speed control. The F5-TTS abstract names slow convergence and low robustness in its predecessor as the motivation for its design.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

arxiv.org/abs/2412.10117 — submitted 2024-12-13, v3 on 2024-12-25.

Finite-scalar quantization to improve codebook utilization of speech tokens, plus a simplified architecture that lets a pre-trained LLM serve directly as the backbone. A chunk-aware causal flow matching model supports streaming and non-streaming synthesis together. The document describes itself as a tech report and work in progress, so its numbers should not be read as settled results.

Qwen3-TTS Technical Report

arxiv.org/abs/2601.15621 — submitted 2026-01-22.

Trained on over 5 million hours of speech across 10 languages, supporting 3-second voice cloning and description-based control. The distinguishing design is two speech tokenizers: one semantics-focused, the other for ultra-low-latency streaming, with a 97 ms first-packet emission time reported. Released under Apache 2.0.

Fish Audio S2 Technical Report

arxiv.org/abs/2603.08823 — submitted 2026-03-09, v2 on 2026-03-11.

Multi-speaker, multi-turn generation with instruction-following control through natural-language descriptions. The report describes a multi-stage training pipeline including captioning, quality assessment and reward modeling, and reports a 0.195 real-time factor with time-to-first-audio below 100 ms. Code, weights and fine-tuning resources are public.

Moshi: a speech-text foundation model for real-time dialogue

arxiv.org/abs/2410.00037 — submitted 2024-09-17, v2 on 2024-10-02.

A report with a different problem definition. It names three failures of component pipelines together: multi-second latency, information loss when text is the intermediate modality, and the inability to handle overlapping speech or interruptions. It models its own speech and the user's speech as parallel streams, and predicts text tokens before audio tokens in an Inner Monologue method. Reported latency is 160 ms theoretical, 200 ms in practice.

Qwen2.5-Omni and MOSS-Audio

arxiv.org/abs/2503.20215 (submitted 2025-03-26) and arxiv.org/abs/2606.01802 (submitted 2026-06-01, v3 on 2026-06-05).

The first takes text, images, audio and video and emits text and natural speech in a streaming manner, with TMRoPE aligning audio and video in time and a Thinker-Talker architecture separating text generation from speech generation to reduce interference. The second is an audio-language model covering speech, environmental sound and music understanding, using 12.5 Hz temporal representations, cross-layer feature injection that exposes the decoder to acoustic information from multiple encoder depths, and inserted time markers. The authors position it as an understanding foundation for future voice agents. Step-Audio 2 (arxiv.org/abs/2507.16632), which adds conversation, belongs to the same family.

Limitations the authors state themselves

  • Omnilingual ASR states directly in its abstract that most of the world's languages remain unsupported, and reflects on ethical considerations alongside that.
  • Qwen3-ASR writes that small differences in open benchmark scores can accompany large differences in real-world quality. Apply that sentence to every ASR selection you make.
  • CosyVoice 2 labels itself work in progress.
  • F5-TTS names its predecessor's convergence speed and robustness as the problems it set out to fix.
  • MOSS-Audio positions itself as an understanding foundation for future voice agents rather than a finished system.
  • Voice cloning carries impersonation risk. Reports rarely discuss that part when they discuss performance, which is worth remembering.

What a practitioner should read first

If you are choosing an ASR model, start with the Open ASR Leaderboard. A standardized comparison comes before any individual report's self-reported numbers.

If low-resource languages are your target, Omnilingual ASR. A design that attaches a new language from a handful of samples changes which use cases are feasible at all.

If you are building real-time conversation, read Moshi and Qwen3-TTS together. They address different bottlenecks: the structural problem of full-duplex dialogue, and first-packet latency.

If you want to fold several modalities into one model, Qwen2.5-Omni and MOSS-Audio are the design references.

How and when this was verified

Every report cited here was verified on 2026-08-12 by opening the arXiv abstract page directly. Candidates that could not be opened were not cited. All numbers, latency and real-time factors included, are as reported by the authors and vary with hardware and configuration.

Try it yourself

  • ML Training Data Explorer — see by example what training data recognition and synthesis models each consume.
  • Browser AI Lab — running a model in the browser with no server gives you a feel for what the latency figures in these reports actually mean.

Previous in the series: Video Generation and Understanding Technical Reports: What to Read

Next in the series: Image Generation and Understanding Technical Reports: What to Read

References