検証: Ralph Loop vs /goal: context disposal の検証
00:01:32 ・ source: https://zenn.dev/mskbhd/articles/lab-117-claude-code-goal
Transcript
hostenToday we explore whether Claude Code's /goal feature truly replaces the Ralph Loop pattern.
guestja今日は、Claude Codeの/goalがRalph Loopというパターンを本当に置き換えられるかを検証する記事です。
hostenThe Ralph Loop has three parts: state externalization, context disposal, and external testing—but /goal only ships one.
guestjaRalph Loopは「状態保存」「毎回新しいコンテキストで起動」「外部テスト」の3つから成りますが、/goalはテスト部分だけを取ったんですね。
hostenSimulations show /goal matches Ralph on short tasks, but fails badly on long ones—20-step tasks lose 11% accuracy.
guestjaシミュレーション結果では、短いタスクなら/goalとRalph Loopは同等ですが、20ステップの長いタスクでは精度が11%落ちます。
hostenThe problem is context rot: rotten context accumulates, and the evaluator itself becomes blind to errors.
guestja原因はコンテキストロットで、時間とともに蓄積された悪いコンテキストを見ると、評価器自体もエラーに気づけなくなります。
hostenFresh context restart is the only way to prevent agents from tricking themselves with reward hacking.
guestja毎回新しいコンテキストで起動することが、エージェント自身が騙される報酬ハッキングを防ぐ唯一の方法です。
hostenFor long production tasks, consider adding your own context disposal—periodic restarts or checkpoint-based session switching.
guestja本番で長時間動くエージェントなら、定期的な再起動やチェックポイント機能を自分で入れることが現実的な対策になります。