検証: RAGレイテンシ、Meta REFRAGで30倍速くなる?検証
00:01:28 ・ source: https://zenn.dev/mskbhd/articles/lab-025-refrag-rag
Transcript
hostenToday we verify Meta's REFRAG paper—does it really speed up RAG latency 30 times?
guestjaMetaのREFRAG論文について、本当にRAGのレイテンシが30倍速くなるのか検証していきます。
hostenThe key problem: we retrieve 80 chunks but only 5 to 10 are actually useful for the answer.
guestjaつまり80個のチャンクから検索してきても、実際に役立つのは5~10個だけで、残りの70個は無駄なアテンション計算を消費しているということです。
hostenREFRAG solves this by compressing each chunk to a single vector, then selectively expanding only important ones.
guestjaREFRAGはこの問題を、各チャンクを埋め込みに圧縮してRLで重要度を判定し、重要なチャンクだけを展開することで解決します。
hostenThe verification shows that 98.6 percent of attention weights across chunks are nearly zero and wasteful.
guestja検証実験でチャンク間のアテンション行列を見ると、全体の98.6%がほぼゼロの無駄な計算だったのです。
hostenWith compression rate 32 and 10 chunks selected, theory predicts 40-fold speedup in TTFT.
guestja圧縮率を32にして10チャンクだけを選択した場合、理論上は40倍のTTFT削減が期待できます。
hostenIn practice, the real challenge is getting inference engines and training infrastructure to support REFRAG at scale.
guestjaただし実運用では、推論エンジンが埋め込み入力に対応し、大規模な学習基盤が必要になるのが課題です。