検証: 「ループを積めば賢くなる」Loopcraftをローカル12Bで実測したら効いたのは1タスクだけだった

00:01:05 ・ source: https://zenn.dev/mskbhd/articles/lab-012-loopcraft

Transcript

hostenToday we examine whether stacking loops like Writer-Reviewer cycles actually makes AI smarter.
guestja今日は、ライター→レビュアー→修正というループを重ねると本当にAIが賢くなるのか、という仮説を検証した話です。
hostenThe author tested this on a local 12B model with three coding tasks: palindrome, calculator, and brackets.
guestja著者がローカルの12Bモデルで回文判定・四則演算・括弧判定の3つのタスクを試した結果、ループが有効だったのは実は1つだけでした。
hostenThe results showed loops helped calc jump from 46.7% to 100%, but palindrome and brackets were already perfect without loops.
guestja難しい四則演算タスクではループで46.7%から100%に改善しましたが、簡単なタスクはループなしで既に完璧だったため、時間が無駄になっただけです。
hostenSo the lesson is: loops only help with hard tasks, not easy ones, making selective use more practical than stacking loops everywhere.
guestjaつまり『ループを積めば賢くなる』は半分本当で、難しいタスクだけループを使う選別的なアプローチが現実的だという結論ですね。