[Strands Decider] Ran it on an M5 Mac (16GB) and measured speed and memory for MPS, MLX, and CPU

[Strands Decider] Ran it on an M5 Mac (16GB) and measured speed and memory for MPS, MLX, and CPU

I tried Strands Decider 2B on an M5 Mac. I measured speed and memory using two calling methods, ask and serve, and three execution methods: MPS, MLX, and CPU.
2026.10.06

This page has been translated by machine translation. View original

Hi, I'm Kema.

Strands Decider 2B is a small model that performs only three types of judgments: choosing one option from a list, returning a Yes/No probability, and evaluating in stages.
Since it's 2B, it can run on a local Mac.
In this article, I tested two calling methods, ask and serve, and three execution methods, MPS, MLX, and CPU, on an M5 Mac (16GB memory), and measured speed and memory usage.

0. Prerequisites

Item Details
Machine Apple M5, 16GB memory, macOS 26.6.2
Python 3.12.13 (venv created with uv 0.11.3)
strands-decider PyPI version 0.1.0 (torch 2.14.1, transformers 5.18.0), MLX is from repository 75c9fd3
Model StrandsAgents/strands-decider-2B-hobson-v19

1. Installation

mkdir -p ~/work/tools/strands-decider && cd ~/work/tools/strands-decider
uv venv sd-venv --python 3.12
VIRTUAL_ENV=sd-venv uv pip install strands-decider

The model is fetched from Hugging Face on the first call.

2. Asking once with ask

sd-venv/bin/strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 \
  --state "Payment has been failing for 3 consecutive days. Please handle this urgently!" \
  --choice "Which team should handle this?=Billing,Sales,Store" \
  --noul "Is urgency being conveyed"
# Example output
noul_0 noul = 0.916
choice_0 -> Sales (confidence 0.417)
    Sales                      0.611
    Store                      0.345
    Billing                    0.043
  • noul_0 noul = 0.916: Probability of answering Yes to "Is urgency being conveyed"

  • choice_0 -> Sales: The option with the highest probability. confidence indicates how concentrated the probabilities are — 0 if all options are equal, 1 if concentrated on one

The urgency is being recognized, but "Sales" was selected as the responsible team.
This is because --choice in ask can only pass option names, and there is no way to convey what "Billing" is responsible for.

Since ask loads the model on every call, each call took 9–11 seconds.
If you need to make judgments repeatedly, use serve.

3. Running as a resident process with serve

serve starts an HTTP server that accepts POST /v1/systemone.
The request format is the same as TypeSafe's Jev.
The default port is 8000, but since another tool was using it in my environment, I set it to 8765.

sd-venv/bin/strands-decider serve StrandsAgents/strands-decider-2B-hobson-v19 --port 8765
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
  "state "Payment has been failing for 3 consecutive days. Please handle this urgently!",
  "questions": {
    "team": {"type": "choice", "instructions": "Which team should handle this",
             "criteria": {"Billing": "Responsible for billing and payments", "Sales": "Responsible for new contracts and deals", "Store": "Responsible for store operations"}},
    "urgent": {"type": "noul", "instructions": "Is urgency being conveyed"}
  }
}'
// Example output (second request)
{"model":"strands-decider-2B-hobson-v19",
 "answers":{"team":{"type":"choice","choice":"Billing","probabilities":{"Billing":0.9251,"Sales":0.0228,"Store":0.0521},"confidence":0.8877},
            "urgent":{"type":"noul","noul":0.9163}},
 "usage":{"input_tokens":161,"output_tokens":2},"latency_ms":189.03}

When descriptions are written in criteria, "Billing" (probability 0.93) was selected.
The first request took about 1.1 seconds, and the second took about 0.19 seconds (189ms).

4. Running with MLX

On Apple Silicon Macs, you can use MLX (Apple's machine learning framework) with --device mlx.
According to the official README, it is 1.4–1.6x faster than MPS, but it is not included in PyPI version 0.1.0 and needs to be installed from the repository.

On an Apple-silicon Mac, --device mlx runs the model through MLX, 1.4 to 1.6x faster than MPS. It needs the mlx extra, which ships with the next release; until then, install from a clone with pip install -e ".[mlx]".

Source: README.md | strands-labs/strands-decider | GitHub

cd ~/work/tools/strands-decider
git clone https://github.com/strands-labs/strands-decider.git src
uv venv mlx-venv --python 3.12
cd src && VIRTUAL_ENV=../mlx-venv uv pip install -e ".[mlx]"
cd .. && mlx-venv/bin/strands-decider serve StrandsAgents/strands-decider-2B-hobson-v19 --port 8765 --device mlx

When I had MPS and MLX solve 240 short Japanese questions, all 240 returned the same answers, and the average difference in confidence was 0.004. It appears that the confidence does not change depending on where it runs.

5. Speed and Memory

I sent requests with varying input lengths 30 times each (10 times each for CPU).
The body text was fictional Japanese, with different content for each request.
Values are the median latency_ms returned by the server.

Input (characters / tokens) MPS MLX CPU
50 chars / 118 105ms 45ms 1,763ms
200 chars / 209 145ms 65ms -
500 chars / 385 266ms 115ms 2,827ms
1,000 chars / 690 437ms 185ms -
2,000 chars / 1,290 865ms 345ms 5,695ms
4,000 chars / 2,488 1,765ms 625ms -
8,000 chars / 4,096**(truncated at limit)** 3,015ms 1,072ms -
Number of questions per request (500 chars) MPS MLX CPU
1 question 266ms 115ms 2,827ms
3 questions 393ms 169ms 7,665ms
5 questions 524ms 217ms 8,004ms
Item MPS MLX CPU
Time from startup to ready 9.4s 16.3s 15.7s
First request 1.1s 3.2s 4.2s
Memory (footprint, at end of measurement) ~5.8GB ~5.0GB ~7.9GB

The footprint is the value from macOS's footprint command, which includes unified memory used by the GPU.

  • Speed: MLX is approximately 2.2–2.8x faster than MPS, exceeding the official 1.4–1.6x. CPU is approximately 7–17x slower than MPS

  • Input length: Time is roughly proportional to the number of tokens. Japanese is approximately 0.6 tokens per character, and the 4,096-token limit corresponds to about 6,500 characters including the question text. Exceeding this does not result in an error — the tail end is truncated

  • Number of questions: The increase per question is about 65ms for MPS and about 25ms for MLX. Since the body text is processed only once, sending multiple questions about the same text together is faster

6. Summary

By running as a resident process with serve, short single-question judgments could be made in about 0.1 seconds.
However, since it uses approximately 5–6GB of memory while running, there are trade-offs with other apps on a 16GB Mac.

References


AI白書2026 配布中

クラスメソッドが独自に行なったAI診断調査をもとに、企業のAI活用の現在地を調査レポートとしてまとめました。企業規模別の活用度傾向に加え、規模を超えてAI活用を進める企業に共通する取り組みまで、自社の現在地を捉えるためのヒントにぜひ。

AI白書2026

無料でダウンロードする

Share this article

DevelopersIO 2026