検証: 人間には読めないのにLLMは読める「圧縮文」をLLMLingua-2で試した
00:01:20 ・ source: https://zenn.dev/mskbhd/articles/lab-001-llmlingua
Transcript
hostenToday we test whether LLMs can read text compressed so heavily that humans cannot.
guestja今日は、人間には読めないほど圧縮された文章をLLMが本当に理解できるのか検証します。
hostenThe researcher used LLMLingua-2 to compress a meeting transcript and measured accuracy loss.
guestja研究者がLLMLingua-2を使って会議録を圧縮し、正答率の低下を測定しました。
hostenAt 50% compression, the model kept 80% accuracy; beyond that, performance dropped sharply.
guestja圧縮率50%では精度80%を維持しましたが、それ以上圧縮すると急激に低下しました。
hostenSo the claim of 99.5% information retention didn't hold in this shorter, fact-dense text.
guestjaつまり『99.5%の情報保持』という主張は、この短くて事実が密集した文章では再現されませんでした。
hostenHowever, for longer and more redundant RAG contexts, the compression method might preserve more information.
guestjaただし、もっと長く冗長なRAGコンテキストなら、圧縮法はより多くの情報を保つ可能性があります。
hostenAt 50% compression, the cost drops by half while keeping useful accuracy for real applications.
guestja50%圧縮なら処理コストが半減し、実運用に十分な精度が保たれるので、導入の価値があります。