検証: 3モデル合議制でハルシネーションは減るのか?検証してみた

00:01:24 ・ source: https://zenn.dev/mskbhd/articles/lab-137-dexmond-orchestrator-3

Transcript

hostenToday we verify whether three models voting together actually reduces hallucinations.
guestja今日は、3つのAIモデルが合議制で判断すると、ハルシネーション(幻覚)が本当に減るのかを検証した実験について見ていきます。
hostenThe author tested GPT-4o-mini, Claude-3-haiku, and Gemini-2.5-flash with thirty questions.
guestja3つのモデルを使って、存在しないイベント質問や計算問題など30問のテストを行いました。
hostenThe results showed that consensus voting helped with fake events but failed at math.
guestja結果として、存在しないイベントを聞く質問では合議制が有効でしたが、計算問題では複数モデルが同じ誤りをして逆効果になりました。
hostenFor hallucination-prone questions, consensus voting improved accuracy from 44% to 67%.
guestja特にハルシネーションしやすい質問では、単一モデルの44%から合議制で67%に正解率が向上しました。
hostenHowever, schema validation and majority voting work best when models fail differently, not identically.
guestjaただし、3モデル合議制が効果を発揮するのは、モデル同士が異なる誤りをする場合であり、同じ知識不足で同じ間違いをすると逆効果になります。
hostenThe key insight is that consensus spreads risk rather than eliminating hallucinations completely.
guestjaつまり、合議制はハルシネーションを完全になくすのではなく、リスクを分散させるという設計思想だということです。