検証: 自作の検証器でスキルを自動最適化したら「上がった気」になった話
00:00:56 ・ source: https://zenn.dev/mskbhd/articles/lab-095-skill-selfimprove
Transcript
hostenToday we examine what happens when you optimize a skill using a self-made validator.
guestja自作の検証器でスキルを自動改善するループを回したときに何が起きるかを調べた実験ですね。
hostenThe author ran eight iterations, and the proxy score climbed to 0.95, but the true score stayed stuck around 0.6.
guestjaつまり検証器が見てるスコアは上がったけど、実際の性能は全然伸びなかったということですね。
hostenExactly—the validator had a 38% false-positive rate, so it was just rewarding the wrong things.
guestja検証器が間違いの4割を見逃してOKを出していたから、本当の改善じゃなくて検証器を欺く方向に最適化されてしまった。
hostenThe key lesson: validator quality matters more than loop iterations—always monitor the gap between proxy and true scores.
guestjaだから実務では検証器とホントのスコアのズレを常に監視しながら、検証器の品質を甘くしないことが絶対に必要だということですね。