
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>Chaos and Order</title>
      <link>https://www.youngju.dev/blog</link>
      <description>천천히 올바르게. AI Researcher &amp; DevOps Engineer Youngju&#39;s blog. GPU/CUDA, LLM, MLOps, Kubernetes AI workloads, and data engineering — plus mindset essays on confidence, routines, health, and sport psychology.</description>
      <language>ko</language>
      <managingEditor>fjvbn2003@gmail.com (Youngju Kim)</managingEditor>
      <webMaster>fjvbn2003@gmail.com (Youngju Kim)</webMaster>
      <lastBuildDate>Wed, 12 Aug 2026 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://www.youngju.dev/tags/technical-report/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-how-to-read-benchmarks.en</guid>
    <title>How to Read Leaderboards and Benchmarks: Why SOTA Has Such a Short Shelf Life</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-how-to-read-benchmarks.en</link>
    <description>Twelve benchmark methodology papers, each verified by opening the arXiv abstract page directly, assembled into a guide for reading leaderboard numbers. Data contamination, prompt format sensitivity, eval harness differences, self-reported scores, sampling asymmetry on leaderboards, LLM judge bias, and the statistical case for putting error bars on evals. The final part of the domain-by-domain technical report reading series, and a guide to reading the five that came before it.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>benchmark</category><category>evaluation</category><category>contamination</category><category>llm-as-a-judge</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-how-to-read-benchmarks.ja</guid>
    <title>リーダーボードとベンチマークの読み方 — SOTAの賞味期限はなぜ短いのか</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-how-to-read-benchmarks.ja</link>
    <description>ベンチマーク方法論の論文12本を、arXivの原文アブストラクトで直接確認し、リーダーボードの数値をどう読むべきかを整理しました。データ汚染、プロンプト形式への敏感さ、評価ハーネスの差、自己申告、リーダーボードの標本の偏り、LLM審査者の偏り、そして誤差棒をつける統計的アプローチまで扱います。領域別・最新技術レポートの読み方シリーズの最終回であり、先行する5回を読むための案内でもあります。</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>benchmark</category><category>evaluation</category><category>contamination</category><category>llm-as-a-judge</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-how-to-read-benchmarks</guid>
    <title>리더보드와 벤치마크를 읽는 법 — SOTA는 왜 유통기한이 짧은가</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-how-to-read-benchmarks</link>
    <description>벤치마크 방법론 논문 열두 편을 arXiv 원문 초록에서 직접 확인해, 리더보드 숫자를 어떻게 읽어야 하는지 정리했습니다. 데이터 오염, 프롬프트 형식 민감도, 평가 하네스 차이, 자체 보고, 리더보드의 표집 비대칭, LLM 심사자의 편향, 그리고 오차 막대를 붙이는 통계적 접근까지 다룹니다. 영역별 최신 기술 리포트 읽기 시리즈의 마지막 편이자, 앞선 다섯 편을 읽는 방법에 대한 안내입니다.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>benchmark</category><category>evaluation</category><category>contamination</category><category>llm-as-a-judge</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-image.en</guid>
    <title>Image Generation and Understanding Technical Reports: What to Read, and Why Design Beats Sample Images</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-image.en</link>
    <description>Eleven image generation and understanding technical reports, each verified by opening the arXiv abstract page directly. Rectified flow transformers, VAR, Emu3, SANA, Janus-Pro, FLUX.1 Kontext, the Qwen-Image line, Seedream 4.0, Z-Image and Mage-Flow, InternVL3, and Qwen-Image-Bench on the evaluation side: what each did that was new, and which limitations the authors stated. No rankings. Part 5 of the domain-by-domain technical report reading series.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>image-generation</category><category>diffusion-transformer</category><category>multimodal</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-image.ja</guid>
    <title>画像生成・理解の技術レポート、何を読むか — サンプル画像ではなく設計を読む</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-image.ja</link>
    <description>画像生成と理解の技術レポート11本を、arXivの原文アブストラクトで直接確認して整理しました。整流フロートランスフォーマーとVAR、Emu3、SANA、Janus-Pro、FLUX.1 Kontext、Qwen-Image系列、Seedream 4.0、Z-ImageとMage-Flow、InternVL3、そして評価側のQwen-Image-Benchまで、何を新しく行い著者がどんな限界を述べたかを読みます。順位は扱いません。領域別・最新技術レポートの読み方シリーズ第5回。</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>image-generation</category><category>diffusion-transformer</category><category>multimodal</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-image</guid>
    <title>이미지 생성·이해 기술 리포트, 무엇을 읽을 것인가 — 샘플 이미지 대신 설계를 읽기</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-image</link>
    <description>이미지 생성과 이해 기술 리포트 열한 편을 arXiv 원문 초록에서 직접 확인해 정리했습니다. 정류 흐름 트랜스포머와 VAR, Emu3, SANA, Janus-Pro, FLUX.1 Kontext, Qwen-Image 계열, Seedream 4.0, Z-Image와 Mage-Flow, InternVL3, 그리고 평가 쪽 Qwen-Image-Bench까지 각각 무엇을 새로 했고 저자들이 어떤 한계를 밝혔는지 읽습니다. 순위는 다루지 않습니다. 영역별 최신 기술 리포트 읽기 시리즈 5편.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>image-generation</category><category>diffusion-transformer</category><category>multimodal</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-llm.en</guid>
    <title>Text LLM Technical Reports: What to Read, and How to Read Design Decisions Instead of Rankings</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-llm.en</link>
    <description>Nine text-LLM technical reports, each verified by opening the arXiv abstract page directly. What DeepSeek-V3, DeepSeek-R1, Qwen3, Gemma 3, Olmo 3, Kimi K2, MiniMax-01, Mellum2 and s1 each did that was new, and which limitations the authors themselves stated. This article does not rank models: leaderboards keep moving and report numbers are almost always self-reported. Part 1 of the domain-by-domain technical report reading series.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>llm</category><category>moe</category><category>reasoning</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-llm.ja</guid>
    <title>テキストLLMの技術レポート、何を読むか — 順位ではなく設計判断を読む</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-llm.ja</link>
    <description>テキストLLMの技術レポート9本を、arXivの原文アブストラクトで直接確認して整理しました。DeepSeek-V3とDeepSeek-R1、Qwen3、Gemma 3、Olmo 3、Kimi K2、MiniMax-01、Mellum2、s1がそれぞれ何を新しく行い、著者自身がどんな限界を述べているかを読みます。モデルの順位は扱いません。リーダーボードは動き続け、レポートの数値はほぼ自己申告だからです。領域別・最新技術レポートの読み方シリーズ第1回。</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>llm</category><category>moe</category><category>reasoning</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-llm</guid>
    <title>텍스트 LLM 기술 리포트, 무엇을 읽을 것인가 — 순위 대신 설계 결정 읽기</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-llm</link>
    <description>텍스트 LLM 기술 리포트 9편을 arXiv 원문 초록에서 직접 확인해 정리했습니다. DeepSeek-V3와 DeepSeek-R1, Qwen3, Gemma 3, Olmo 3, Kimi K2, MiniMax-01, Mellum2, s1이 각각 무엇을 새로 했고 저자들이 어떤 한계를 스스로 밝혔는지 읽습니다. 모델 순위는 다루지 않습니다. 리더보드는 계속 바뀌고 리포트의 점수는 대부분 자체 보고이기 때문입니다. 영역별 최신 기술 리포트 읽기 시리즈 1편.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>llm</category><category>moe</category><category>reasoning</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-ocr-document.en</guid>
    <title>OCR and Document Understanding Technical Reports: What to Read, and Why Parsing Is Not Finished</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-ocr-document.en</link>
    <description>Ten OCR and document understanding technical reports, each verified by opening the arXiv abstract page directly. From Donut and Nougat through GOT-OCR2.0, olmOCR, DeepSeek-OCR and its successor, GLM-OCR, Qianfan-OCR, MinerU2.5-Pro and HunyuanOCR-1.5: what each did that was new, and which limitations the authors stated. No rankings here, because document parsing scores are unusually sensitive to benchmark version and scoring rules. Part 2 of the domain-by-domain technical report reading series.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>ocr</category><category>document-ai</category><category>vision-language-model</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-ocr-document.ja</guid>
    <title>OCR・文書理解の技術レポート、何を読むか — パースはなぜまだ終わっていないのか</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-ocr-document.ja</link>
    <description>OCRと文書理解の技術レポート10本を、arXivの原文アブストラクトで直接確認して整理しました。DonutとNougatからGOT-OCR2.0、olmOCR、DeepSeek-OCRとその後継、GLM-OCR、Qianfan-OCR、MinerU2.5-Pro、HunyuanOCR-1.5まで、それぞれ何を新しく行い著者がどんな限界を述べたかを読みます。順位は扱いません。文書パースのスコアはベンチマークのバージョンと採点方法にとりわけ敏感だからです。領域別・最新技術レポートの読み方シリーズ第2回。</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>ocr</category><category>document-ai</category><category>vision-language-model</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-ocr-document</guid>
    <title>OCR·문서 이해 기술 리포트, 무엇을 읽을 것인가 — 파싱은 왜 아직 안 끝났나</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-ocr-document</link>
    <description>OCR과 문서 이해 기술 리포트 열 편을 arXiv 원문 초록에서 직접 확인해 정리했습니다. Donut과 Nougat에서 GOT-OCR2.0, olmOCR, DeepSeek-OCR과 그 후속, GLM-OCR, Qianfan-OCR, MinerU2.5-Pro, HunyuanOCR-1.5까지 각각 무엇을 새로 했고 저자들이 어떤 한계를 밝혔는지 읽습니다. 순위는 다루지 않습니다. 문서 파싱 점수는 벤치마크 버전과 채점 방식에 특히 민감하기 때문입니다. 영역별 최신 기술 리포트 읽기 시리즈 2편.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>ocr</category><category>document-ai</category><category>vision-language-model</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-speech.en</guid>
    <title>Speech Recognition and Synthesis Technical Reports: What to Read, and What a Single WER Hides</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-speech.en</link>
    <description>Ten speech recognition and synthesis technical reports, each verified by opening the arXiv abstract page directly. Whisper, Omnilingual ASR, Qwen3-ASR, the Open ASR Leaderboard, Seed-TTS and F5-TTS, CosyVoice 2, Qwen3-TTS, Fish Audio S2, Moshi, Qwen2.5-Omni and MOSS-Audio: what each did that was new, and which limitations the authors stated. Instead of rankings: latency, language coverage and streaming constraints. Part 4 of the domain-by-domain technical report reading series.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>speech-recognition</category><category>text-to-speech</category><category>audio-language-model</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-speech.ja</guid>
    <title>音声認識・合成の技術レポート、何を読むか — WERという一つの数字では見えないもの</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-speech.ja</link>
    <description>音声認識と合成の技術レポート10本を、arXivの原文アブストラクトで直接確認して整理しました。WhisperとOmnilingual ASR、Qwen3-ASR、Open ASR Leaderboard、Seed-TTSとF5-TTS、CosyVoice 2、Qwen3-TTS、Fish Audio S2、Moshi、Qwen2.5-OmniとMOSS-Audioが何を新しく行い、著者がどんな限界を述べたかを読みます。順位ではなく遅延、言語カバレッジ、ストリーミングの制約を見ます。領域別・最新技術レポートの読み方シリーズ第4回。</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>speech-recognition</category><category>text-to-speech</category><category>audio-language-model</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-speech</guid>
    <title>음성 인식·합성 기술 리포트, 무엇을 읽을 것인가 — WER 한 숫자로는 안 보이는 것</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-speech</link>
    <description>음성 인식과 합성 기술 리포트 열 편을 arXiv 원문 초록에서 직접 확인해 정리했습니다. Whisper와 Omnilingual ASR, Qwen3-ASR, Open ASR Leaderboard, Seed-TTS와 F5-TTS, CosyVoice 2, Qwen3-TTS, Fish Audio S2, Moshi, Qwen2.5-Omni와 MOSS-Audio가 각각 무엇을 새로 했고 저자들이 어떤 한계를 밝혔는지 읽습니다. 순위 대신 지연, 언어 커버리지, 스트리밍 제약을 봅니다. 영역별 최신 기술 리포트 읽기 시리즈 4편.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>speech-recognition</category><category>text-to-speech</category><category>audio-language-model</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-video.en</guid>
    <title>Video Generation and Understanding Technical Reports: What to Read, and Why Constraints Beat Demos</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-video.en</link>
    <description>Ten video generation and understanding technical reports, each verified by opening the arXiv abstract page directly. CogVideoX, Movie Gen, HunyuanVideo, LTX-Video, Wan, Seedance 1.0 and 2.0, plus Qwen2.5-VL, VideoLLaMA 3 and HY-Himmel on the understanding side: what each did that was new, and which limitations the authors stated. No rankings here. Instead: duration, resolution, VRAM and token budgets, the constraints that decide what you can actually build. Part 3 of the domain-by-domain technical report reading series.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>video-generation</category><category>diffusion-transformer</category><category>video-understanding</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-video.ja</guid>
    <title>動画生成・理解の技術レポート、何を読むか — デモではなく制約を読む</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-video.ja</link>
    <description>動画生成と理解の技術レポート10本を、arXivの原文アブストラクトで直接確認して整理しました。CogVideoX、Movie Gen、HunyuanVideo、LTX-Video、Wan、Seedance 1.0と2.0、そして理解側のQwen2.5-VL、VideoLLaMA 3、HY-Himmelまで、何を新しく行い著者がどんな限界を述べたかを読みます。順位は扱わず、長さ・解像度・トークン予算といった実際の制約を見ます。領域別・最新技術レポートの読み方シリーズ第3回。</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>video-generation</category><category>diffusion-transformer</category><category>video-understanding</category><category>evaluation</category>
  </item>

  <item>
    <guid>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-video</guid>
    <title>비디오 생성·이해 기술 리포트, 무엇을 읽을 것인가 — 데모가 아니라 제약을 읽기</title>
    <link>https://www.youngju.dev/blog/ai-papers/2026-08-12-sota-video</link>
    <description>비디오 생성과 이해 기술 리포트 열 편을 arXiv 원문 초록에서 직접 확인해 정리했습니다. CogVideoX, Movie Gen, HunyuanVideo, LTX-Video, Wan, Seedance 1.0과 2.0, 그리고 이해 쪽의 Qwen2.5-VL, VideoLLaMA 3, HY-Himmel까지 무엇을 새로 했고 저자들이 어떤 한계를 밝혔는지 읽습니다. 순위는 다루지 않고, 대신 길이·해상도·토큰 예산 같은 실제 제약을 봅니다. 영역별 최신 기술 리포트 읽기 시리즈 3편.</description>
    <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    <author>fjvbn2003@gmail.com (Youngju Kim)</author>
    <category>ai-papers</category><category>paper-review</category><category>technical-report</category><category>video-generation</category><category>diffusion-transformer</category><category>video-understanding</category><category>evaluation</category>
  </item>

    </channel>
  </rss>
