Ran the OSS model "Kev", a judgment-specialized AI Jev-compatible model, on MLX with an M4 Mac, and tested the judgment accuracy of the 4B version

Ran the OSS model "Kev", a judgment-specialized AI Jev-compatible model, on MLX with an M4 Mac, and tested the judgment accuracy of the 4B version

Ran "Kev," an OSS model compatible with specialized AI Jev, on M4 Pro/macOS 27 via MLX. The 4B model, too heavy for M1/Docker, responded in 0.26s and outperformed 0.8B in AI-DLC scope selection.
2026.09.27

This page has been translated by machine translation. View original

Introduction

"Kev" is an OSS model with a Jev-compatible API specialized for judgment, supporting MLX that can be used on Apple Silicon Macs with M1 or later.

https://github.com/jaredpalmer/kev

Previously, I ran Kev on an M1 Mac with Docker / CPU. Since the Docker Linux container cannot use the Mac's GPU, MLX was unavailable and execution was CPU-based. Only the low-end 0.8B model could run at a practical speed, so I was unable to verify judgment accuracy.

https://dev.classmethod.jp/articles/kev-oss-jev-compatible-docker/

This time, I set up an environment with Apple M4 Pro / macOS 27 and will introduce the results of trying the 4B model with MLX.

Setting Up the Environment

Version Information

  • Mac mini (M4 Pro, 24GB memory)
  • macOS 27.0
  • uv 0.12.19
  • Python 3.13.15
  • mlx 0.32.2 / mlx-lm 0.31.3
  • Kev jaredpalmer/kev-4b (release_date: 2026-09-24)

Installation

Install and launch Kev with uv. The uv binary, Python, venv, and Hugging Face model cache are all placed in the user space.

# Install uv (if not already installed)
curl -LsSf https://astral.sh/uv/install.sh | sh
# Clone Kev and install dependencies + MLX into venv
git clone --depth 1 https://github.com/jaredpalmer/kev.git
cd kev
uv sync --extra serve --python 3.13
# Start server (weights are downloaded on first run only)
uv run python -m kev.serve --run jaredpalmer/kev-4b --port 8009

If the following line appears in the startup log, the MLX backend is running on the Mac's GPU (mps).

serving jaredpalmer/kev-4b (...) on mps via mlx (bfloat16) 127.0.0.1:8009

GET /v1/models also returned 200 OK, with backend as mlx, device as mps, and dtype as bfloat16.

Evaluation

Memory and Response Time by Model

I started 0.8B, 4B, and 9B by changing the model name in --run, and measured the first and second response times when sending the same Japanese request (the same as the previous article, asking about the responsible department for a support ticket, escalation, and degree of anger) after the server started. Memory is the measured value of Physical footprint read with vmmap --summary.

Model Physical footprint (peak) First time Second time
Kev-0.8B 2.1G (3.1G) 205.7ms 45.0ms
Kev-4B 8.6G (15.6G) 1200.9ms 262.1ms
Kev-9B 17.4G (30.6G) 30244.8ms 483.0ms

The Physical footprint of 17.4G for 9B matches "Kev-9B needs about 17 GB" in the Kev README. The peak of 30.6G exceeds the 24GB of installed memory on the Mac, which is likely the reason why the first run took 30244.8ms. Both 0.8B and 4B had footprint and peak within 24GB, and the first run was within 1.2 seconds.

Trying the Choice API

Kev's choice API is an API that selects one option from candidates.

The selection targets are scopes from AI-DLC (awslabs/aidlc-workflows v2.10.0). A scope is a category that determines which stages to execute for each type of work, with 11 options including bugfix, feature, and enterprise. The AI-DLC itself automatically determines the scope from a /aidlc <free text> request using keywords, but the question this time is whether having Kev select by meaning can choose the appropriate scope name from a request statement.

https://github.com/awslabs/aidlc-workflows

The Auto (Jev) mode added in Kiro Crew 0.7 interprets the meaning of a message from many skills and selects one. Since the problem structure of "selecting one option from candidates based on a request statement" is the same, I thought that if this works, we can also expect accuracy in skill selection by Jev in Auto (Jev).

Request and Response

The choice API sends a question with type: choice to /v1/systemone. I listed 12 choices in criteria, adding "not applicable (none)" to the 11 scopes, and put the request statement in state. One entry for a scope looks like this:

curl -s http://127.0.0.1:8009/v1/systemone -H "content-type: application/json" -d "$(cat scope-request.json)"
{
  "state": "The login page returns a 500 error for users whose email contains an apostrophe. Fix it.",
  "model": "kev-latest",
  "questions": {
    "scope": {
      "type": "choice",
      "instructions": "Which AI-DLC scope should handle this request?",
      "criteria": {
        "aidlc-bugfix": "Fix a specific bug",
        "aidlc-classic": "V1-style ceremony through Inception and Construction - the implicit default",
        "aidlc-enterprise": "Regulated enterprise feature, full audit trail",
        "aidlc-express": "Lightest run: requirements to deploy, no design pass, no reviewers",
        "aidlc-feature": "Full lifecycle for new features, practical depth",
        "aidlc-infra": "Infrastructure changes",
        "aidlc-mvp": "Skip operations, ship the core",
        "aidlc-poc": "Prove feasibility fast",
        "aidlc-refactor": "Clean up existing code",
        "aidlc-security-patch": "CVE response",
        "aidlc-workshop": "Facilitated group session with mandatory gates",
        "none": "None of these apply to this request."
      }
    }
  }
}

For this request statement, Kev-4B selected aidlc-bugfix. The probability was 0.63, the runner-up was none at 0.14, and the response time was 268.5ms. Since the correct label I assigned was also aidlc-bugfix, this one item is correct.

I prepared request statements of the same format in both English and Japanese, and assigned correct labels myself. Only the request statement (state) was in Japanese, while the criteria and instructions remained in English. I tested two versions of criteria descriptions: a short version as shown above, and a detailed version with keywords and the first sentence of the body text added. In addition to scopes, I also tested AI-DLC skills (4 items: knowledge, outcomes-pack, replay, session-cost + none, for 5 choices) using the same procedure.

There are two sets of request statements.

Clear requests are those where the words in the request statement closely match the description of the correct scope.

  • "A CVE-2026-1234 has been found in the logging library we depend on. Please address it." → Correct answer: aidlc-security-patch (description: "CVE response")
  • "We are adding a new payment method under PCI DSS audit. All decisions require an audit trail." → Correct answer: aidlc-enterprise

Requests that don't apply to any scope, such as asking about the weather, are also mixed in as correct answers for none. There are 50 items total: 15 scope items × 2 languages + 10 skill items × 2 languages.

Confusing requests are those that mix words from other scopes.

  • "Migrate to the new load balancer and address the 502 errors occurring there." → Correct answer: aidlc-infra. "Address the 502" pulls toward aidlc-bugfix
  • "We are adding multi-factor authentication to the admin console. It's subject to SOC 2 audit sampling." → Correct answer: aidlc-enterprise. "We are adding" pulls toward aidlc-feature

There are 20 items total: 8 scope items × 2 languages + 2 skill items × 2 languages.

Results

Number of errors combining both criteria versions, with accuracy in parentheses. The denominator is 100 for clear requests and 40 for confusing requests.

Model Clear requests (/100) Confusing requests (/40)
Kev-0.8B 23 (77%) 20 (50%)
Kev-4B 0 (100%) 10 (75%)
Kev-9B 1 (99%) 12 (70%)

0.8B is clearly inferior to 4B and 9B. Of the 23 cases that 0.8B got wrong in clear requests, 4B got all 23 correct, and 9B got 22 correct. Most of 0.8B's errors were missed cases where it selected none, and it only incorrectly selected a different scope twice in clear requests.

There was no clear difference between 4B and 9B. Of the 10 cases where 4B got wrong in confusing requests, 9B also got 8 of them wrong, and 6 of those were the same wrong answer. The only case where 9B alone got wrong was the multi-factor authentication request, where it selected aidlc-feature in all 4 combinations of English/Japanese and both criteria versions. 4B got all 4 correct.

Requests That Both 4B and 9B Got Wrong

There are 4 types of request statements. Language and criteria versions that produced the same result are combined.

Request statement Correct answer 4B 9B 0.8B
Please rewrite the payment module cleanly. While you're at it, add two new items. (Japanese, version B) refactor feature feature refactor
Migrate to the new load balancer and address the 502 errors occurring there. (Japanese, versions A/B) infra none none bugfix
Tidy up the CI configuration and cut the build time while you are at it. (English, version A) refactor infra infra infra
For the record, please tell me what was decided and why. (Japanese/English, versions A/B) replay (skill) none / outcomes-pack none / knowledge / outcomes-pack none / outcomes-pack

The first item is a request that I myself judged to be ambiguous between refactor and feature. The fourth item has close descriptions between the skill replay and outcomes-pack, which would be confusing even for humans. The cases that can be called model errors are the second and third, and in both cases, both models got the correct answer in other languages or criteria versions.

Requests where only 0.8B got wrong (4B and 9B correct, 16 types)
Request statement Correct answer 0.8B
We are adding a new payment method under PCI DSS audit. All decisions require an audit trail. (Japanese, version B) enterprise feature
Please take us from requirements to deployment as quickly as possible. You can skip the design pass and reviewers. (Japanese, version B) express none
Prove whether on-device inference can stay under 200 ms. Feasibility only, throw the code away after. (English, version A) poc none
Run a facilitated session with the whole team on Tuesday, with a mandatory gate at each step. (English, version B) workshop none
Make things better. (English, version A) none bugfix
Please import the PDFs from the team drive so that agents can cite them. (Japanese/English, version A) knowledge none
Please register and sync the internal design documents in the catalog. (Japanese/English, version A) knowledge none
The workflow is done. Please create handover materials so the team can operate it themselves. (Japanese, version A) outcomes-pack none
Please write the closing handover document so ownership can be transferred to the platform team. (Japanese/English, versions A/B) outcomes-pack none
Please summarize what happened in this session for stakeholders who were not present. (Japanese/English, versions A/B) replay none
Please tell me the time taken for this workflow, the number of stages run, and the number of sensor firings. Numbers only. (Japanese, version A) session-cost none
Please show me the cost view of the current workflow. (English versions A/B, Japanese version A) session-cost none
We will ship the first version of the dashboard quickly. We'll skip monitoring for now. (Japanese, version B) mvp infra
We will hold a training day for new employees, walking through the workflow together. (Japanese, version A) workshop none
Please tidy up the CI configuration and cut the build time while you're at it. (Japanese versions A/B, English version B) refactor infra
Please summarize this workflow for handover. Please include the time taken for each stage. (Japanese/English, versions A/B) outcomes-pack none

Summary

I was able to use Kev-4B with MLX on a Mac with M4 Pro / 24GB. When using the choice API to select AI-DLC scopes, 4B was clearly better than 0.8B, and there was no difference between 9B and 4B. Even when the request statement was in Japanese, the number of errors for 4B was the same as in English, with no significant degradation in answer quality.

Also, while the response time for 4B was 26 to 93 seconds on M1 / Docker / CPU, it was a practical speed of 1.2 seconds for the first run and 0.26 seconds for the second run on M4 Pro / MLX. Therefore, for cases where you don't want to send input data to an external API during inference, Kev + MLX seems to be a strong candidate as an execution environment for judgment-specialized AI.

In the future, I would like to evaluate Auto (Jev) supported in Kiro Crew 0.7 (a mode where Jev judges which model and skills to use based on the prompt), as well as compare it with the original Jev, CLM, and other alternatives.

https://dev.classmethod.jp/articles/clm-contrastive-language-model-switchyard-judge/

Share this article

DevelopersIO 2026