
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>Chaos and Order</title>
      <link>https://www.youngju.dev/blog</link>
      <description>천천히 올바르게. AI Researcher &amp; DevOps Engineer Youngju&#39;s blog. GPU/CUDA, LLM, MLOps, Kubernetes AI workloads, and data engineering — plus mindset essays on confidence, routines, health, and sport psychology.</description>
      <language>ko</language>
      <managingEditor>fjvbn2003@gmail.com (Youngju Kim)</managingEditor>
      <webMaster>fjvbn2003@gmail.com (Youngju Kim)</webMaster>
      <lastBuildDate>Wed, 12 Aug 2026 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://www.youngju.dev/tags/llm-as-a-judge/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-how-to-read-benchmarks.en</guid>
    <title>How to Read Leaderboards and Benchmarks: Why SOTA Has Such a Short Shelf Life</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-how-to-read-benchmarks.en</link>
    <description>Twelve benchmark methodology papers, each verified by opening the arXiv abstract page directly, assembled into a guide for reading leaderboard numbers. Data contamination, prompt format sensitivity, eval harness differences, self-reported scores, sampling asymmetry on leaderboards, LLM judge bias, and the statistical case for putting error bars on evals. The final part of the domain-by-domain technical report reading series, and a guide to reading the five that came before it.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>benchmark</category><category>evaluation</category><category>contamination</category><category>llm-as-a-judge</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-how-to-read-benchmarks.ja</guid>
    <title>リーダーボードとベンチマークの読み方 — SOTAの賞味期限はなぜ短いのか</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-how-to-read-benchmarks.ja</link>
    <description>ベンチマーク方法論の論文12本を、arXivの原文アブストラクトで直接確認し、リーダーボードの数値をどう読むべきかを整理しました。データ汚染、プロンプト形式への敏感さ、評価ハーネスの差、自己申告、リーダーボードの標本の偏り、LLM審査者の偏り、そして誤差棒をつける統計的アプローチまで扱います。領域別・最新技術レポートの読み方シリーズの最終回であり、先行する5回を読むための案内でもあります。</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>benchmark</category><category>evaluation</category><category>contamination</category><category>llm-as-a-judge</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-how-to-read-benchmarks</guid>
    <title>리더보드와 벤치마크를 읽는 법 — SOTA는 왜 유통기한이 짧은가</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-how-to-read-benchmarks</link>
    <description>벤치마크 방법론 논문 열두 편을 arXiv 원문 초록에서 직접 확인해, 리더보드 숫자를 어떻게 읽어야 하는지 정리했습니다. 데이터 오염, 프롬프트 형식 민감도, 평가 하네스 차이, 자체 보고, 리더보드의 표집 비대칭, LLM 심사자의 편향, 그리고 오차 막대를 붙이는 통계적 접근까지 다룹니다. 영역별 최신 기술 리포트 읽기 시리즈의 마지막 편이자, 앞선 다섯 편을 읽는 방법에 대한 안내입니다.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>benchmark</category><category>evaluation</category><category>contamination</category><category>llm-as-a-judge</category>
  </item>

    </channel>
  </rss>
