検証: Qwen3-0.6Bで自作Jev: 固定ヘッドよりラベルを読ませる方が強かった
00:01:45 ・ source: https://zenn.dev/mskbhd/articles/lab-704-llmjev
Transcript
hostenToday we explore building a classification model by swapping Qwen3-0.6B's language head with a typed prediction head.
guestjaQwen3-0.6Bの言語生成ヘッドを分類ヘッドに載せ替えて、型付きの判断を1回の forward で返すモデルを作った実験です。
hostenThe key finding: reading label descriptions as text beats fixed classification heads in accuracy.
guestja選択肢を文章として読ませる方式(クロスエンコーダ風)が、固定ヘッドよりも精度で勝りました。意図分類で0.950対0.911です。
hostenHowever, removing the language head saved almost nothing because embeddings are tied to the output layer.
guestjaパラメータ削減を狙ってlm_headを外しても、入力埋め込みと重みを共有しているため、総パラメータ数は5.96億のまま変わりませんでした。
hostenOn GPU with batching, the cost is only +4ms; on CPU it becomes seventeen times slower.
guestjaGPUでバッチ処理すれば24msから29msへの+4msで済みますが、CPUでは0.44秒から7.3秒へと17倍になり実用的ではありません。
hostenThe main limitation: unseen labels achieve 0.83 accuracy versus Claude Haiku's 0.99.
guestja学習時に無かったラベルへの対応力では0.83対0.99と大きく差が出ており、小型モデルの0.6Bと530件の学習では埋まりませんでした。
hostenFor fixed-type tasks on GPU, self-built pair LoRA works well; for dynamic types, buy the specialized model.
guestja型が固定されていてGPUがあれば自作のpair_LoRA方式で十分競えますが、型が頻繁に増える場合は専用モデルJevを使う価値が残ります。