検証: GEPAでJevを最適化: 効いたのはcriteriaだけで、ECEは壊れなかった

00:01:21 ・ source: https://zenn.dev/mskbhd/articles/lab-724-dspy-gepa-jev-instructions-cri

Transcript

hostenToday we verify: can GEPA optimize Jev without learning weights, using only prompt text?
guestja今日は、重みを固定したまま、プロンプトの文面だけで Jev を最適化できるかを検証した結果です。
hostenJev returns typed judgments and probabilities in one call, but cannot be retrained at all.
guestjaJev は 1 回の forward で型付き判断と確率を返しますが、再学習ができません。
hostenThe only tunable parts are the `instructions` question text and `criteria` candidate descriptions.
guestjaチューニングできるのは、質問文の instructions と候補ごとの説明である criteria の文面だけです。
hostenResults: GEPA improved test accuracy by +4 to +5 points, but only criteria mattered, not instructions.
guestja結果は、GEPA により test 精度が +4〜5pt 上がり、効いたのは criteria だけで instructions は効きませんでした。
hostenECE (calibration error) improved from 0.072 to 0.047, because fewer high-confidence errors occurred.
guestjaECE は 0.072 から 0.047 に改善しました。これは高確率の誤りが減ったための副産物です。
hostenKey insight: machine can auto-find which candidate description to rewrite by reading misclassifications.
guestja重要な発見は、どの候補の説明を書き直すべきかを、誤分類例から機械が自動で見つけられるということです。