
Decision 2.0 を Strands Decider と同じ 20 問で M4 Mac の MPS で試してみた
はじめに
Decision 2.0 は、vLLM Semantic Router の判定専用モデルのシリーズです。Hugging Face のリポジトリは 2026年9月28日から29日にかけて作成されました。Kai-0.6B から Vega-27B まで 6 サイズが公開されています。
このうち Eos-0.8B、Sol-2B、Nox-4B の 3 つを、先行記事の Strands Decider 2B と同じ 20 問・同じ判定文で動かしました。
Kev・Strands Decider との違い
| 観点 | Decision 2.0 | Kev | Strands Decider |
|---|---|---|---|
| サイズ | Kai-0.6B / Eos-0.8B / Sol-2B / Nox-4B / Lux-9B / Vega-27B | 0.8B / 4B / 9B / 27B | 2B |
| 元のモデル | Qwen3(Kai)、Qwen3.5(Eos・Sol・Nox・Lux)、Qwen3.8(Vega) | Qwen3.5(0.8B・4B・9B)、Qwen3.8(27B) | Qwen3.5-2B |
| 学習方式 | サイズで異なる。Kai・Eos・Sol・Lux は全パラメータを fine-tune、Nox は rank-128 の LoRA を本体に統合、Vega は rank-512 の LoRA アダプタ | 0.8B・4B・9B は凍結した本体に rank-16 の LoRA。27B は全パラメータを fine-tune | rank-16 の LoRA |
| 出力部 | 各選択肢の端点ベクトルと問い全体のベクトルを、共有の bilinear MLP で読み出す | pointer head(選択肢の末尾トークンを、問いの最終トークンに対してスコア) | pointer head(torso が出した各選択肢の答えをスコア) |
| 呼び出し | Hugging Face の AutoModel(trust_remote_code=True)と system_one()、または pipeline("decision")。パッケージに HTTP サーバーはなく、dtype はランタイム固定で選べない |
POST /v1/systemone(TypeSafe の System One API 互換)。Mac では MLX で動く |
strands-decider serve の POST /v1/systemone(同 API 互換)と ask |
| 入力長の上限(モデルカード・設定上の値) | Kai 8,192、Eos〜Lux 16,384、Vega 32,768 トークン | state 65,536 トークン(検証済みは 0.8B〜9B が 8,192、27B が 65,536) | max_length 4,096(設定ファイル) |
Kev と Strands Decider の詳細は、それぞれのモデルカードと告知を参照してください。
試した条件
判定文は、Kiro Crew の model.route と同じ instructions を使います。Crew の判定文は criteria が null のため、先行記事では Decider が HTTP 422 で拒否しました。Eos-0.8B も、推論を実行せず invalid_question を返しました。そのため、先行記事と同じく tier の説明を criteria として補った B_desc を使いました。state は {"message": 本文} の形で渡しました。日本語 20 問も先行記事と同一です。各条件で、最初に 1 問をウォームアップとして動かし、計測から除外しました。
判定文 B_desc(prompts.py)
INSTR_B = (
"How hard is this request for an AI coding assistant? Answer one of: "
"simple = a short, local, mechanical request answerable in one step -- a rename, a lookup, a small edit, a direct factual question; "
"medium = ordinary work over a few files or steps -- write a function, explain some code, fix a bug whose cause is already named; "
"complex = work needing a plan, a trade-off or reasoning across a whole system -- design, architecture, a diagnosis with no named cause, a risky refactor. "
"Being wrong in the two directions does not cost the same. A turn put in a lower tier than it needs gets less room to work in and a worse answer; "
"a turn put higher only costs more. So answer simple only when the request is self-contained and you can see the whole of it -- "
"a request that is a plan, a judgement call, several instructions at once, or that names work you cannot see, is not simple however short it is."
)
TIER_DESC = {
"simple": "a short, local, mechanical request answerable in one step -- a rename, a lookup, a small edit, a direct factual question",
"medium": "ordinary work over a few files or steps -- write a function, explain some code, fix a bug whose cause is already named",
"complex": "work needing a plan, a trade-off or reasoning across a whole system -- design, architecture, a diagnosis with no named cause, a risky refactor",
}
実行環境は Mac mini(M4 Pro、24GB、macOS 27.0.1)で、Python 3.12.14、torch 2.14.1、transformers 5.18.0 を使いました。20 問の実行回数は、3 モデルとも 3 回です。
モデルは commit を固定しました(Eos-0.8B は 3594047d、Sol-2B は 23cbe9f9、Nox-4B は 25e8f67d)。結果の表にある Decider と Kev-4B の値は、同じ Mac で計測した先行記事のログの値です。
動かし方
uv で仮想環境を作り、依存を入れます。
uv venv --python 3.12
uv pip install torch==2.14.1 transformers==5.18.0 safetensors==0.8.0
prompts.py に上の判定文を保存し、次の Python コードから呼び出しました。
from transformers import AutoModel
from prompts import INSTR_B, TIER_DESC
model = AutoModel.from_pretrained(
"vllm-sr/Decision-2.0-Eos-0.8B",
trust_remote_code=True,
revision="3594047d69f476f1d01cf84c593e213fc3a4dfe0",
device="mps",
)
model.eval()
text = "このリポジトリで使っている Node.js のバージョンは?"
result = model.system_one(
state={"message": text},
questions={"tier": {"type": "choice", "instructions": INSTR_B, "criteria": TIER_DESC}},
)
trust_remote_code=True は Hugging Face 上の Python コードを実行するため、revision に commit を指定して固定しています。
ロードのたびに incorrect regex pattern … fix_mistral_regex=True というトークナイザの警告が出ました。警告の影響を確かめるため、ランタイム内部のトークナイザ(runtime.backend.tokenizer)を fix_mistral_regex=True で読み直したものに差し替えました。Eos-0.8B を M4 Pro の MPS で B_desc の 20 問で再実行したところ、差し替え前と比べて tier は変わらず、確率の差は最大 0.046 でした。結果の表の値は差し替え前のものです。
結果
判定
「想定内」は tier が「許容」列に入った問数、「完全一致」は想定回答と同じ tier を返した問数です。Eos-0.8B の判定は、3 回の実行で同一でした。
| 設問 | 想定 | 許容 | Decider 2B | Kev-4B | Eos-0.8B | Eos と Decider が同じ |
|---|---|---|---|---|---|---|
| P01 | simple | simple, medium | simple | simple | medium | × |
| P02 | medium | medium, simple | simple | simple | medium | × |
| P03 | simple | simple | simple | simple | simple | ○ |
| P04 | medium | medium, complex | complex | simple | complex | ○ |
| P05 | medium | medium, complex | complex | simple | medium | × |
| P06 | complex | complex, medium | complex | complex | medium | × |
| P07 | simple | simple, medium | simple | simple | simple | ○ |
| P08 | simple | simple, medium | simple | simple | simple | ○ |
| P09 | complex | complex, medium | complex | complex | complex | ○ |
| P10 | simple | simple, medium | simple | simple | simple | ○ |
| P11 | medium | medium, complex | medium | simple | complex | × |
| P12 | simple | simple, medium | simple | simple | medium | × |
| P13 | complex | complex | complex | complex | complex | ○ |
| P14 | complex | complex, medium | medium | medium | complex | × |
| P15 | medium | medium, simple | medium | complex | medium | ○ |
| P16 | medium | medium, simple | medium | simple | medium | ○ |
| P17 | complex | complex | complex | complex | complex | ○ |
| P18 | complex | complex, medium | medium | medium | medium | ○ |
| P19 | medium | medium, complex | complex | medium | medium | × |
| P20 | complex | complex | complex | complex | complex | ○ |
| 対象 | 想定内 | 完全一致 | 想定より低い / 高い |
|---|---|---|---|
| Decider 2B(先行) | 20/20 | 14 | 3 / 3 |
| Kev-4B(先行) | 16/20 | 12 | 7 / 1 |
| Eos-0.8B | 20/20 | 14 | 2 / 4 |
| Sol-2B | 19/20 | 16 | 0 / 4 |
| Nox-4B | 16/20 | 12 | 7 / 1 |
Eos-0.8B が想定回答と完全一致した件数は、Decider と同じ 14 問です。一方、Eos-0.8B と Decider が同じ tier を返したのは 20 問中 12 問でした。Sol-2B と Nox-4B も同じ条件で動かしました。想定内の件数はサイズが大きくなっても増えず、Nox-4B の想定外 4 件はすべて simple の判定です。
応答時間
Decider と Kev-4B の値は、今回とは別の先行記事のセッションで、HTTP サーバーへリクエストを送りクライアント側で測ったものです。Kev-4B は MLX(bf16)で動かしています。Eos-0.8B・Sol-2B・Nox-4B は、プロセス内で system_one() の呼び出しを測りました。
| 対象 | 平均 | 中央値 | 最小〜最大 | 20 問合計 | 回数 |
|---|---|---|---|---|---|
| Decider 2B(先行) | 376〜381ms | 327〜335ms | 320〜752ms | 約 7.6 秒 | 3 |
| Kev-4B(先行、MLX) | 604ms | 526ms | 439〜1,222ms | 12.1 秒 | 2 |
| Eos-0.8B | 479ms | 435ms | 365〜997ms | 9.4〜9.8 秒 | 3 |
| Sol-2B | 645ms | 587ms | 508〜1,338ms | 12.8〜13.1 秒 | 3 |
| Nox-4B | 1,794ms | 1,628ms | 1,366〜3,470ms | 34.6〜37.2 秒 | 3 |
Eos-0.8B・Sol-2B・Nox-4B の平均・中央値・最小〜最大は、3 回分(60 応答)をまとめた値です。Decider の平均・中央値・20 問合計は、3 回の実行ごとの値の範囲です。Kev-4B は 2 回とも同じ値でした。Eos-0.8B は Decider より平均で約 100ms 長く、Sol-2B と Nox-4B は Eos-0.8B よりさらに長くなりました。
メモリ
footprint は macOS の footprint コマンドの Physical footprint で、ユニファイドメモリ上の GPU 分を含みます。RSS は ps の常駐メモリです。Eos-0.8B・Sol-2B・Nox-4B の値は、括弧外がプロセス終了時、括弧内が計測中のピークです。Decider の値は、先行記事のサーバープロセスで評価終了直後に測ったものです。
| 対象 | footprint | RSS | MPS 割当 |
|---|---|---|---|
| Decider 2B(先行) | 約 5.1〜5.2GB | 0.9GB | — |
| Kev-4B(先行、MLX) | 10GB | 2.2GB | — |
| Eos-0.8B | 5.3GB(5.96GB) | 1.1GB(3.6GB) | 2.9GB |
| Sol-2B | 9.7GB(11GB) | 0.6GB(8.6GB) | 7.2GB |
| Nox-4B | 19GB(21GB) | 0.6GB(10.3GB) | 16.1GB |
footprint は Eos-0.8B と Decider で同程度でした。24GB の Mac mini で、Nox-4B の footprint はピークで 21GB に達しました。
まとめ
Decision 2.0 には 0.6B〜27B の 6 サイズがあり、Kev と同様にサイズを選べます。
今回の 20 問では、判定精度で Strands Decider 2B に匹敵する結果を Eos-0.8B と Sol-2B で得ることができました。
差が出たのは速度とメモリです。Eos-0.8B は Decider より約 100ms 遅く、メモリは同程度でした。Sol-2B は応答がさらに遅く、メモリも多く使いました。
Decision 2.0 は、先行記事で取り上げた Jev をはじめとする判定専用モデルの中でも、OSS 実装でサイズ選択の幅が広いモデルの 1 つです。今後も注目していきたいと思います。









