![[Strands Decider] Ran it on an M5 Mac (16GB) and measured speed and memory for MPS, MLX, and CPU](https://images.ctfassets.net/ct0aopd36mqt/348VnC3440CoUtpS4pd5NS/8a7182ab96a0851a72a88d40328811b3/aws.png?w=3840&fm=webp)
[Strands Decider] Ran it on an M5 Mac (16GB) and measured speed and memory for MPS, MLX, and CPU
This page has been translated by machine translation. View original
Hi, I'm Kema.
Strands Decider 2B is a small model that performs only three types of judgments: choosing one option from a list, returning a Yes/No probability, and evaluating in stages.
Since it's 2B, it can run on a local Mac.
In this article, I tested two calling methods, ask and serve, and three execution methods, MPS, MLX, and CPU, on an M5 Mac (16GB memory), and measured speed and memory usage.
0. Prerequisites
| Item | Details |
|---|---|
| Machine | Apple M5, 16GB memory, macOS 26.6.2 |
| Python | 3.12.13 (venv created with uv 0.11.3) |
| strands-decider | PyPI version 0.1.0 (torch 2.14.1, transformers 5.18.0), MLX is from repository 75c9fd3 |
| Model | StrandsAgents/strands-decider-2B-hobson-v19 |
1. Installation
mkdir -p ~/work/tools/strands-decider && cd ~/work/tools/strands-decider
uv venv sd-venv --python 3.12
VIRTUAL_ENV=sd-venv uv pip install strands-decider
The model is fetched from Hugging Face on the first call.
2. Asking once with ask
sd-venv/bin/strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 \
--state "Payment has been failing for 3 consecutive days. Please handle this urgently!" \
--choice "Which team should handle this?=Billing,Sales,Store" \
--noul "Is urgency being conveyed"
# Example output
noul_0 noul = 0.916
choice_0 -> Sales (confidence 0.417)
Sales 0.611
Store 0.345
Billing 0.043
-
noul_0 noul = 0.916: Probability of answering Yes to "Is urgency being conveyed" -
choice_0 -> Sales: The option with the highest probability.confidenceindicates how concentrated the probabilities are — 0 if all options are equal, 1 if concentrated on one
The urgency is being recognized, but "Sales" was selected as the responsible team.
This is because --choice in ask can only pass option names, and there is no way to convey what "Billing" is responsible for.
Since ask loads the model on every call, each call took 9–11 seconds.
If you need to make judgments repeatedly, use serve.
3. Running as a resident process with serve
serve starts an HTTP server that accepts POST /v1/systemone.
The request format is the same as TypeSafe's Jev.
The default port is 8000, but since another tool was using it in my environment, I set it to 8765.
sd-venv/bin/strands-decider serve StrandsAgents/strands-decider-2B-hobson-v19 --port 8765
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
"state "Payment has been failing for 3 consecutive days. Please handle this urgently!",
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this",
"criteria": {"Billing": "Responsible for billing and payments", "Sales": "Responsible for new contracts and deals", "Store": "Responsible for store operations"}},
"urgent": {"type": "noul", "instructions": "Is urgency being conveyed"}
}
}'
// Example output (second request)
{"model":"strands-decider-2B-hobson-v19",
"answers":{"team":{"type":"choice","choice":"Billing","probabilities":{"Billing":0.9251,"Sales":0.0228,"Store":0.0521},"confidence":0.8877},
"urgent":{"type":"noul","noul":0.9163}},
"usage":{"input_tokens":161,"output_tokens":2},"latency_ms":189.03}
When descriptions are written in criteria, "Billing" (probability 0.93) was selected.
The first request took about 1.1 seconds, and the second took about 0.19 seconds (189ms).
4. Running with MLX
On Apple Silicon Macs, you can use MLX (Apple's machine learning framework) with --device mlx.
According to the official README, it is 1.4–1.6x faster than MPS, but it is not included in PyPI version 0.1.0 and needs to be installed from the repository.
On an Apple-silicon Mac,
--device mlxruns the model through MLX, 1.4 to 1.6x faster than MPS. It needs themlxextra, which ships with the next release; until then, install from a clone withpip install -e ".[mlx]".
Source: README.md | strands-labs/strands-decider | GitHub
cd ~/work/tools/strands-decider
git clone https://github.com/strands-labs/strands-decider.git src
uv venv mlx-venv --python 3.12
cd src && VIRTUAL_ENV=../mlx-venv uv pip install -e ".[mlx]"
cd .. && mlx-venv/bin/strands-decider serve StrandsAgents/strands-decider-2B-hobson-v19 --port 8765 --device mlx
When I had MPS and MLX solve 240 short Japanese questions, all 240 returned the same answers, and the average difference in confidence was 0.004. It appears that the confidence does not change depending on where it runs.
5. Speed and Memory
I sent requests with varying input lengths 30 times each (10 times each for CPU).
The body text was fictional Japanese, with different content for each request.
Values are the median latency_ms returned by the server.
| Input (characters / tokens) | MPS | MLX | CPU |
|---|---|---|---|
| 50 chars / 118 | 105ms | 45ms | 1,763ms |
| 200 chars / 209 | 145ms | 65ms | - |
| 500 chars / 385 | 266ms | 115ms | 2,827ms |
| 1,000 chars / 690 | 437ms | 185ms | - |
| 2,000 chars / 1,290 | 865ms | 345ms | 5,695ms |
| 4,000 chars / 2,488 | 1,765ms | 625ms | - |
| 8,000 chars / 4,096**(truncated at limit)** | 3,015ms | 1,072ms | - |
| Number of questions per request (500 chars) | MPS | MLX | CPU |
|---|---|---|---|
| 1 question | 266ms | 115ms | 2,827ms |
| 3 questions | 393ms | 169ms | 7,665ms |
| 5 questions | 524ms | 217ms | 8,004ms |
| Item | MPS | MLX | CPU |
|---|---|---|---|
| Time from startup to ready | 9.4s | 16.3s | 15.7s |
| First request | 1.1s | 3.2s | 4.2s |
| Memory (footprint, at end of measurement) | ~5.8GB | ~5.0GB | ~7.9GB |
The footprint is the value from macOS's footprint command, which includes unified memory used by the GPU.
-
Speed: MLX is approximately 2.2–2.8x faster than MPS, exceeding the official 1.4–1.6x. CPU is approximately 7–17x slower than MPS
-
Input length: Time is roughly proportional to the number of tokens. Japanese is approximately 0.6 tokens per character, and the 4,096-token limit corresponds to about 6,500 characters including the question text. Exceeding this does not result in an error — the tail end is truncated
-
Number of questions: The increase per question is about 65ms for MPS and about 25ms for MLX. Since the body text is processed only once, sending multiple questions about the same text together is faster
6. Summary
By running as a resident process with serve, short single-question judgments could be made in about 0.1 seconds.
However, since it uses approximately 5–6GB of memory while running, there are trade-offs with other apps on a 16GB Mac.

