
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>Chaos and Order</title>
      <link>https://www.youngju.dev/blog</link>
      <description>천천히 올바르게. AI Researcher &amp; DevOps Engineer Youngju&#39;s blog. GPU/CUDA, LLM, MLOps, Kubernetes AI workloads, and data engineering — plus mindset essays on confidence, routines, health, and sport psychology.</description>
      <language>ko</language>
      <managingEditor>fjvbn2003@gmail.com (Youngju Kim)</managingEditor>
      <webMaster>fjvbn2003@gmail.com (Youngju Kim)</webMaster>
      <lastBuildDate>Wed, 12 Aug 2026 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://www.youngju.dev/tags/speech-to-text/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://www.youngju.dev/blog/ai/2026-08-12-huggingface-speech-models.en</guid>
    <title>Choosing Speech Models: Practical Criteria for STT and TTS</title>
    <link>https://www.youngju.dev/blog/ai/2026-08-12-huggingface-speech-models.en</link>
    <description>Unlike text models, speech models are chosen after the language coverage, audio length constraints, real-time requirement, and diarization need are already fixed. This post lays out the parameters, licenses, language coverage, and audio constraints read off open STT and TTS pages on 2026-08-12, explains how to decompose the phrase real-time, and shows why the hallucination and consent warnings on these cards are design constraints. Part 5 of the open model guide series.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai</category><category>huggingface</category><category>open-source-llm</category><category>speech-to-text</category><category>text-to-speech</category><category>asr</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai/2026-08-12-huggingface-speech-models.ja</guid>
    <title>音声モデルの選び方: STTとTTSの実践基準</title>
    <link>https://www.youngju.dev/blog/ai/2026-08-12-huggingface-speech-models.ja</link>
    <description>音声モデルはテキストモデルと違い、対応言語、音声長の制約、リアルタイム性、話者分離の要否が先に決まり、その後にモデルが決まります。この記事は2026-08-12に確認したSTTとTTSのオープンモデルのパラメータ、ライセンス、対応言語、音声制約を整理し、リアルタイムという語をどう分解すべきか、カードに書かれた幻覚と同意に関する警告がなぜ設計制約になるのかを説明します。オープンモデルガイドシリーズ第5回です。</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai</category><category>huggingface</category><category>open-source-llm</category><category>speech-to-text</category><category>text-to-speech</category><category>asr</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai/2026-08-12-huggingface-speech-models</guid>
    <title>음성 모델 고르기: STT와 TTS의 실전 기준</title>
    <link>https://www.youngju.dev/blog/ai/2026-08-12-huggingface-speech-models</link>
    <description>음성 모델은 텍스트 모델과 달리 언어 지원, 오디오 길이 제약, 실시간성, 화자 분리 여부가 먼저 정해지고 그다음에 모델이 결정됩니다. 이 글은 2026-08-12에 확인한 STT와 TTS 오픈 모델의 파라미터, 라이선스, 지원 언어, 오디오 제약을 정리하고, 실시간이라는 말을 어떻게 쪼개야 하는지, 카드에 적힌 환각과 동의 관련 경고가 왜 설계 제약인지를 설명합니다. 오픈 모델 가이드 시리즈 5편입니다.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai</category><category>huggingface</category><category>open-source-llm</category><category>speech-to-text</category><category>text-to-speech</category><category>asr</category>
  </item>

    </channel>
  </rss>
