
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>Chaos and Order</title>
      <link>https://www.youngju.dev/blog</link>
      <description>천천히 올바르게. AI Researcher &amp; DevOps Engineer Youngju&#39;s blog. GPU/CUDA, LLM, MLOps, Kubernetes AI workloads, and data engineering — plus mindset essays on confidence, routines, health, and sport psychology.</description>
      <language>ko</language>
      <managingEditor>fjvbn2003@gmail.com (Youngju Kim)</managingEditor>
      <webMaster>fjvbn2003@gmail.com (Youngju Kim)</webMaster>
      <lastBuildDate>Sun, 09 Aug 2026 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://www.youngju.dev/tags/arc-agi/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://www.youngju.dev/blog/llm/2026-08-09-reasoning-effort-is-a-deployment-parameter.en</guid>
    <title>Reasoning Effort Is Not a Model Choice but a Per-Request Deployment Parameter</title>
    <link>https://www.youngju.dev/blog/llm/2026-08-09-reasoning-effort-is-a-deployment-parameter.en</link>
    <description>The DeepSeek V4 Flash 0731 results page published by ARC Prize carries not one score but three, one per reasoning effort level. This post computes what can actually be read out of those three numbers: that the same step up in effort buys 5 percentage points on the easy benchmark and 15 on the hard one, and therefore that effort should be treated as a value decided per request rather than as a model setting. The escalation structure and the conditions required for that structure to hold are laid out in code.</description>
    <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>llm</category><category>benchmark</category><category>arc-agi</category><category>inference</category><category>cost</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/llm/2026-08-09-reasoning-effort-is-a-deployment-parameter.ja</guid>
    <title>推論強度はモデル選択ではなくリクエスト単位のデプロイパラメータです</title>
    <link>https://www.youngju.dev/blog/llm/2026-08-09-reasoning-effort-is-a-deployment-parameter.ja</link>
    <description>ARC Prizeが公開したDeepSeek V4 Flash 0731の結果ページには、スコアがひとつではなく推論強度別に三つ載っています。この記事はその三つの数字から実際に読み取れることを計算します。同じ強度の引き上げが、易しいベンチマークでは5ポイントを、難しいベンチマークでは15ポイントを買ってくれるという事実と、ならば強度はモデル設定ではなくリクエストごとに決める値として扱うべきだという結論です。はしごを上る構造と、その構造が成り立つための条件をコードで整理しました。</description>
    <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>llm</category><category>benchmark</category><category>arc-agi</category><category>inference</category><category>cost</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/llm/2026-08-09-reasoning-effort-is-a-deployment-parameter</guid>
    <title>추론 강도는 모델 선택이 아니라 요청 단위 배포 파라미터입니다</title>
    <link>https://www.youngju.dev/blog/llm/2026-08-09-reasoning-effort-is-a-deployment-parameter</link>
    <description>ARC Prize가 공개한 DeepSeek V4 Flash 0731 결과 페이지에는 점수가 하나가 아니라 추론 강도별로 세 개 실려 있습니다. 이 글은 그 세 숫자에서 실제로 읽어 낼 수 있는 것을 계산합니다. 같은 강도 상향이 쉬운 벤치마크에서는 5퍼센트포인트를, 어려운 벤치마크에서는 15퍼센트포인트를 사 준다는 사실과, 그렇다면 강도를 모델 설정이 아니라 요청마다 결정하는 값으로 다뤄야 한다는 결론입니다. 사다리를 올리는 구조와 그 구조가 성립하기 위한 조건을 코드로 정리했습니다.</description>
    <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>llm</category><category>benchmark</category><category>arc-agi</category><category>inference</category><category>cost</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/llm/2026-08-09-reasoning-effort-is-a-deployment-parameter.zh</guid>
    <title>推理强度不是选模型，而是逐请求决定的部署参数</title>
    <link>https://www.youngju.dev/blog/llm/2026-08-09-reasoning-effort-is-a-deployment-parameter.zh</link>
    <description>ARC Prize 公开的 DeepSeek V4 Flash 0731 结果页上，分数不是一个，而是按推理强度列了三个。本文计算的是从这三个数字里真正能读出来的东西：同样一次强度上调，在简单基准上买到 5 个百分点，在困难基准上买到 15 个百分点；既然如此，强度就该被当成逐请求决定的值，而不是模型配置。文中用代码梳理了逐级上调的结构，以及这个结构成立所需要的条件。</description>
    <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>llm</category><category>benchmark</category><category>arc-agi</category><category>inference</category><category>cost</category>
  </item>

    </channel>
  </rss>
