検証: Jevの「校正された確率」を実測: ECEはHaikuと互角、差は別の所に出た

00:01:27 ・ source: https://zenn.dev/mskbhd/articles/lab-682-jev-ece-692

Transcript

hostenToday we test Jev, a model that returns typed judgments with calibrated probabilities instead of generating text.
guestja今日はJevというモデルを検証します。校正された確率を返すことが売りで、文章は生成しません。
hostenThe key finding: ECE shows no difference from Claude Haiku, but Jev wins on accuracy and confidence discrimination.
guestja最大の発見は、ECE(校正誤差)ではHaikuと差がなかったことです。ただ、正答率と確信度の判別力ではJevが勝りました。
hostenOn hard multiplication problems, Jev drops to 63% accuracy but lowers confidence when unsure—4-digit checks at 0.9+ were all correct.
guestja4桁掛け算の検算で、Jevは正答率63%に落ちても確信度をちゃんと下げます。信頼度0.9以上の4件はすべて正解でした。
hostenSpeed is roughly 3× faster than Haiku from Tokyo, costing 63 times less—but keep-alive connections are essential.
guestja速度は東京からで約3倍速く、コストは約63分の1です。ただしHTTPのkeep-aliveが必須です。
hostenFor speech-end detection, Jev's probabilities never exceeded 0.6, so the fixed threshold 0.5 underperformed—tuning per domain is required.
guestja発話終端判定では、Jevの確率は0.6を超えず、固定的なしきい値0.5では固定待ちに負けました。ドメインごとの調整が必須です。