
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>Chaos and Order</title>
      <link>https://www.youngju.dev/blog</link>
      <description>천천히 올바르게. AI Researcher &amp; DevOps Engineer Youngju&#39;s blog. GPU/CUDA, LLM, MLOps, Kubernetes AI workloads, and data engineering — plus mindset essays on confidence, routines, health, and sport psychology.</description>
      <language>ko</language>
      <managingEditor>fjvbn2003@gmail.com (Youngju Kim)</managingEditor>
      <webMaster>fjvbn2003@gmail.com (Youngju Kim)</webMaster>
      <lastBuildDate>Sun, 09 Aug 2026 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://www.youngju.dev/tags/eval-driven-development/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://www.youngju.dev/blog/ai/2026-08-09-eval-driven-development-calibrate-the-judge.en</guid>
    <title>In Eval-Driven Development, the First Thing to Calibrate Is the Judge</title>
    <link>https://www.youngju.dev/blog/ai/2026-08-09-eval-driven-development-calibrate-the-judge.en</link>
    <description>The eval-driven development retrospective Airbnb Engineering published in July 2026 is less a plea to write the eval set first than a plea to earn the right to treat the grading model as an instrument. This post lays out the procedure for calibrating a judge model against a golden set of 50 to 100 examples, why agreement has to be measured with kappa rather than plain accuracy, and how an uncalibrated judge steers an entire team toward the wrong target — with runnable code.</description>
    <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai</category><category>llm</category><category>eval-driven-development</category><category>llm-as-judge</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai/2026-08-09-eval-driven-development-calibrate-the-judge.ja</guid>
    <title>評価駆動開発で最初に較正すべきなのは審査者です</title>
    <link>https://www.youngju.dev/blog/ai/2026-08-09-eval-driven-development-calibrate-the-judge.ja</link>
    <description>Airbnbエンジニアリングが2026年7月に公開した評価駆動開発の振り返りは、評価セットを先に書けという話ではなく、採点する側のモデルがまず計測器として認められなければならないという話に近いものです。この記事は、審査者モデルをゴールデンセット50〜100件で較正する手順、一致度を単純な正確度ではなくカッパで測るべき理由、そして較正されていない審査者がどうやってチーム全体を誤った方向へ最適化させてしまうのかを、実行可能なコードとともに整理します。</description>
    <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai</category><category>llm</category><category>eval-driven-development</category><category>llm-as-judge</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai/2026-08-09-eval-driven-development-calibrate-the-judge</guid>
    <title>평가 주도 개발에서 가장 먼저 보정해야 하는 것은 심사자입니다</title>
    <link>https://www.youngju.dev/blog/ai/2026-08-09-eval-driven-development-calibrate-the-judge</link>
    <description>Airbnb 엔지니어링이 2026년 7월에 공개한 평가 주도 개발 회고는 평가셋을 먼저 쓰라는 이야기가 아니라, 채점하는 모델부터 계측기로 인정받아야 한다는 이야기에 가깝습니다. 이 글은 심사자 모델을 골든셋 50~100개로 보정하는 절차, 일치도를 단순 정확도가 아니라 카파로 재야 하는 이유, 그리고 보정되지 않은 심사자가 어떻게 팀 전체를 잘못된 방향으로 최적화시키는지를 실행 가능한 코드와 함께 정리합니다.</description>
    <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai</category><category>llm</category><category>eval-driven-development</category><category>llm-as-judge</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai/2026-08-09-eval-driven-development-calibrate-the-judge.zh</guid>
    <title>评估驱动开发里，最先要校准的是评判者</title>
    <link>https://www.youngju.dev/blog/ai/2026-08-09-eval-driven-development-calibrate-the-judge.zh</link>
    <description>Airbnb 工程团队在 2026 年 7 月发布的评估驱动开发复盘，与其说是在讲「先把评测集写出来」，不如说是在讲「负责打分的那个模型必须先被承认为一台量具」。本文整理了用 50 到 100 条黄金集校准评判者模型的流程、一致度为什么要用 kappa 而不是简单准确率来度量，以及未经校准的评判者如何把整个团队优化到错误的方向上去 —— 并附可运行的代码。</description>
    <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai</category><category>llm</category><category>eval-driven-development</category><category>llm-as-judge</category><category>evaluation</category>
  </item>

    </channel>
  </rss>
