パターン評価ラボ
小さなテストセットを異なるエージェントパターンと異なるモデルで並べて実行します。各実行では実際の出力、レイテンシ、トークン数、コストが表示され、AIジャッジが期待する答えと照らして品質を採点し、比較全体をJSONやCSVでエクスポートできます。パターン選びを感覚ではなく計測値に基づかせましょう。
実際の比較
同じタスクを4通りで
1つの問いに対して2つのモデルと2つのパターンを、推測ではなく実測しました。レイテンシは実時間、トークン数は提供元の値、コストは実際に請求された金額です。
タスク: A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost?
| モデル | パターン | 回答 | レイテンシ | トークン | コスト |
|---|---|---|---|---|---|
| anthropic/claude-sonnet-5 | Direct | $0.05 | 3,065 ms | 63 | $0.000174 |
| anthropic/claude-sonnet-5 | Chain of Thought | $0.05 | 4,870 ms | 366 | $0.0031 |
| google/gemini-2.5-flash | Direct | $0.05 | 458 ms | 45 | $0.000024 |
| google/gemini-2.5-flash | Chain of Thought | $0.05 | 1,257 ms | 276 | $0.00058 |
4通りすべてが正解しました。つまりこのタスクでは、推論パターンは時間とコストを増やしただけでした。最も高い実行は最も安い実行の約128倍かかっています。パターン選択を推測せず実測すべき理由であり、このラボが出力の隣にコストとレイテンシを表示する理由でもあります。
計測日 2026-08-06
仕事ぶりを見せるジャッジ
このラボは、あなたが与える期待する答えを根拠に、AIジャッジで品質も採点します。この記録された比較では、両方の実行が正しい数値を計算しましたが、ジャッジはそれでも差を付けました。タスクは数値を最初に答えるよう求めており、それを後回しにした実行は指示遵守の点を失いました。
タスク: A retail API rate limiter allows 120 requests per minute per key. A batch job needs to send 4,500 requests. What is the minimum time in minutes to send them all, and how should the job pace itself? Answer with the number first.
期待する答え: 37.5 minutes (4500 / 120), pacing at or under 2 requests per second.
| モデル | パターン | 品質 | 指示遵守 | レイテンシ | コスト |
|---|---|---|---|---|---|
| openai/gpt-5-mini | Chain of Thought | 5.00/5 | 5/5 | 14,220 ms | $0.0021 |
| anthropic/claude-haiku-4-5 | Chain of Thought | 4.67/5 | 3/5 | 4,688 ms | $0.0023 |
減点された実行へのジャッジの所見: “The calculation and pacing advice are correct and clearly explained, but the answer is not given first as instructed; it only appears at the end after several intermediate steps.”
一対一の判定は位置バイアスを検査します。順序を入れ替えてジャッジに2回尋ね、この僅差のペアでは選択が変わったため、ラボは勝者をでっち上げずに引き分けと報告しました。
anthropic/claude-sonnet-5が期待する答えを基準に1〜5で採点。計測日 2026-08-10