
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>Chaos and Order</title>
      <link>https://www.youngju.dev/blog</link>
      <description>천천히 올바르게. AI Researcher &amp; DevOps Engineer Youngju&#39;s blog. GPU/CUDA, LLM, MLOps, Kubernetes AI workloads, and data engineering — plus mindset essays on confidence, routines, health, and sport psychology.</description>
      <language>ko</language>
      <managingEditor>fjvbn2003@gmail.com (Youngju Kim)</managingEditor>
      <webMaster>fjvbn2003@gmail.com (Youngju Kim)</webMaster>
      <lastBuildDate>Sun, 26 Jul 2026 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://www.youngju.dev/tags/regression-testing/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://www.youngju.dev/blog/llm/2026-07-26-llm-evaluation-without-vibes</guid>
    <title>LLM 평가를 감으로 하지 않는 법 — 표본 크기, 심사자 편향, CI 회귀 테스트</title>
    <link>https://www.youngju.dev/blog/llm/2026-07-26-llm-evaluation-without-vibes</link>
    <description>프롬프트를 고치고 &quot;좋아진 것 같다&quot;로 배포하는 팀은 조용히 쌓이는 회귀를 볼 방법이 없습니다. 평가를 어서션, 골든 데이터셋, LLM 심사자, 사람 평가의 네 층위로 나눠 각각의 비용과 신뢰도를 정리하고, 심사자 모델의 위치 편향과 길이 선호를 어떻게 상쇄하는지 다룹니다. 예시 20개로 낸 결론의 신뢰구간이 실제로 얼마나 넓은지 계산하고, 짝지은 비교가 필요 표본을 몇 분의 일로 줄이는지 코드로 보입니다. 온도 0도 완전히 재현되지 않는 환경에서 CI 회귀 테스트를 구성하는 방법까지 포함했습니다.</description>
    <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>llm</category><category>evaluation</category><category>llm-as-judge</category><category>statistics</category><category>regression-testing</category>
  </item>

    </channel>
  </rss>
