<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Benchmark on Blowing in the wind</title>
    <link>https://zheng-bobo.github.io/tags/benchmark/</link>
    <description>Recent content in Benchmark on Blowing in the wind</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-CN</language>
    <lastBuildDate>Sat, 26 Sep 2026 23:00:00 +0200</lastBuildDate>

  <atom:link href="https://zheng-bobo.github.io/tags/benchmark/index.xml" rel="self" type="application/rss+xml" />


    <item>
      <title>大模型与 AI Agent Benchmark 全景指南：它们究竟在评估什么</title>
      <link>https://zheng-bobo.github.io/post/llm-agent-benchmarks-guide/</link>
      <pubDate>Sat, 26 Sep 2026 23:00:00 +0200</pubDate>

      <guid>https://zheng-bobo.github.io/post/llm-agent-benchmarks-guide/</guid>
      <description>&lt;p&gt;看模型发布报告时，我们经常会遇到一长串名字：ARC-E、ARC-C、MMLU、GPQA、GSM8K、HumanEval、SWE-bench、GAIA、WebArena、OSWorld……它们的分数不能直接横向比较，因为每个 benchmark 测量的能力、输入形式、可用工具和评分方式都不同。&lt;/p&gt;

&lt;p&gt;本文整理一张从“大模型答题”到“Agent 完成真实任务”的评测地图，并解释每个常见 benchmark 究竟是什么、适合测什么，以及不能从分数中推出什么结论。&lt;/p&gt;</description>
    </item>

  </channel>
</rss>