
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>Chaos and Order</title>
      <link>https://www.youngju.dev/blog</link>
      <description>천천히 올바르게. AI Researcher &amp; DevOps Engineer Youngju&#39;s blog. GPU/CUDA, LLM, MLOps, Kubernetes AI workloads, and data engineering — plus mindset essays on confidence, routines, health, and sport psychology.</description>
      <language>ko</language>
      <managingEditor>fjvbn2003@gmail.com (Youngju Kim)</managingEditor>
      <webMaster>fjvbn2003@gmail.com (Youngju Kim)</webMaster>
      <lastBuildDate>Wed, 12 Aug 2026 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://www.youngju.dev/tags/tensor-parallel/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://www.youngju.dev/blog/ai-platform/2026-08-12-vllm-tuning-and-pitfalls.en</guid>
    <title>Inside vLLM (7) — Deployment Tuning, Common Pitfalls, and OOM Triage</title>
    <link>https://www.youngju.dev/blog/ai-platform/2026-08-12-vllm-tuning-and-pitfalls.en</link>
    <description>A practical order for tuning a vLLM deployment. Covers what gpu_memory_utilization actually sets, when to use tensor parallelism versus pipeline parallelism, how to choose quantization and a KV cache data type, and a diagnostic order for OOM that separates the startup stage from the operating stage, all checked against the official documentation. The final part of the Inside vLLM series.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>vllm</category><category>gpu</category><category>quantization</category><category>tensor-parallel</category><category>llm</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-platform/2026-08-12-vllm-tuning-and-pitfalls.ja</guid>
    <title>vLLM 内部構造 (7) — デプロイのチューニングとよくある落とし穴、OOM の切り分け順</title>
    <link>https://www.youngju.dev/blog/ai-platform/2026-08-12-vllm-tuning-and-pitfalls.ja</link>
    <description>vLLMのデプロイを実際にチューニングする手順を整理する。gpu_memory_utilizationが何を決めるのか、テンソル並列とパイプライン並列をいつ使うのか、量子化とKVキャッシュのデータ型をどう選ぶのか、そしてOOMが起きたときに起動段階と運用段階を分けて診断する手順まで、公式ドキュメントを基準に確認して整理した。vLLM内部構造シリーズ最終回。</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>vllm</category><category>gpu</category><category>quantization</category><category>tensor-parallel</category><category>llm</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-platform/2026-08-12-vllm-tuning-and-pitfalls</guid>
    <title>vLLM 내부 구조 (7) — 배포 튜닝과 흔한 함정, OOM 진단 순서</title>
    <link>https://www.youngju.dev/blog/ai-platform/2026-08-12-vllm-tuning-and-pitfalls</link>
    <description>vLLM 배포를 실제로 튜닝하는 순서를 정리했습니다. gpu_memory_utilization이 무엇을 정하는지, 텐서 병렬과 파이프라인 병렬을 언제 쓰는지, 양자화와 KV 캐시 자료형을 어떻게 고르는지, 그리고 OOM이 났을 때 기동 단계와 운영 단계를 나눠 진단하는 순서까지 공식 문서 기준으로 확인해 정리했습니다. vLLM 내부 구조 시리즈 마지막 편.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>vllm</category><category>gpu</category><category>quantization</category><category>tensor-parallel</category><category>llm</category>
  </item>

    </channel>
  </rss>
