
Claude Code v2.1.269 Major Updates - Addition of claude plugin eval and fix for permission rule application omissions
This page has been translated by machine translation. View original
This is Ishikawa from the Cloud Business Division. Claude Code v2.1.269 (released 2026-09-11) has been released. With 98 changes, this is another high-volume release following the previous one. The standout feature of this release is the concentration of new features and security fixes around plugins. Today, I tried out the claude plugin eval command, which scores and quantifies whether plugins (skills) are actually working.
The previous update article is here.
Update Summary
v2.1.269 includes 98 changes. Counting the CHANGELOG entries, there are 66 Fixed, 14 Added, 11 Improved, 6 Changed, and 1 Removed. The targets are 58 in the CLI itself, with the remainder being 23 in the VSCode extension, 10 in Claude Tag (Claude in Slack), and 7 in Claude Code on the web.
Notable Updates
Plugins can be scored with claude plugin eval
claude plugin eval has been added. It runs a plugin's eval suite against Claude Code and provides reproducible results with scores in JSON and HTML reports.
Those distributing plugins or skills internally will find it useful as a mechanism for re-scoring with the same cases each time an update is made.
Diffs of files changed by Bash commands are included in results
When the Bash tool handles file editing, diffs of files changed by Bash commands are now included in the execution results (setting bashEditDiffEnabled).
Those using Bash-based editing will find the reduced effort worthwhile, as they can now confirm what was rewritten within the results themselves.
Fixed permission rule application gaps
Two fixes were made around permissions.
- An issue where deny/ask rules starting with
!were being applied outside the configuration source they were written in. The applicable rules are now only applied within their own configuration source, and negation with a standalone!is ignored - An issue where
Edit()deny rules and write path checks were not being applied to files written by Bash'steecommand.Bash(tee:*)allow rules no longer permit writes outside the working directory
Teams enforcing operations with deny rules may have had paths that were not working as intended, so we consider it worth prioritizing this upgrade.
Fixed permissions around plugin archive extraction
Three issues were fixed: extracted plugin archives for sessions being readable by other users on the same machine, extracted files inheriting world-writable bits from the archive, and old files remaining even after re-extraction.
For those using Claude Code on shared machines or development servers with multiple users, this is considered a high-priority fix.
Fixed issue where Japanese prompt suggestions were not displayed
An issue was fixed where prompt suggestions were not displayed for Japanese, Chinese, Thai, and other text written without spaces between words. Filtering has also been improved: suggestions with mixed character types or single-word suggestions are retained, while meta/evaluative text is excluded in the same way as English.
For those writing prompts in Japanese, this fix will work directly.
Session stalling and prompt cache fixes
Two fixes were made that affect long sessions.
- An issue where a session could not recover from "Prompt is too long" when there were no completed past exchanges that could be summarized and auto-compaction could not be performed (mainly Agent SDK sessions handling very large prompts)
- An issue where the prompt cache was partially invalidated in the turn following an auto-resume after a response was cut off by the output token limit
This will benefit those running long sessions or handling large prompts with the Agent SDK.
Update Details
New Features
/output-style [name]has been added, enabling listing and switching of output styles. It can also be used via Remote Control and in cloud and other headless sessionsOTEL_METRICS_INCLUDE_REPOSITORYhas been added, attachingvcs.*repository attributes to OpenTelemetry metrics and events. When used together withOTEL_LOG_TOOL_DETAILS,vcs.ref.head.*is attached to commit eventsCLAUDE_CODE_WORKFLOW_MAX_CONCURRENT_AGENTS(1–256) has been added, allowing you to increase the maximum number of concurrent agents per Workflow tool executionCLAUDE_CODE_GATEWAY_MODEL_DISCOVERY_TIMEOUT_MShas been added, allowing you to extend the timeout (default 3 seconds) for LLM gateway/v1/modelsdiscovery- [VSCode] An agent map has been added, accessible from the "N agents" pill in the footer, allowing you to open cards per sub-agent, stop agents, and open read-only transcripts
- [VSCode] A Hooks dialog has been added to the command menu, enabling viewing of hooks and adding, editing, and deleting them in user, project, and local settings. Managed, plugin, and session hooks are read-only
- [VSCode] A permission rules dialog has been added, enabling listing of permission rules and adding and deleting them in user, project, and local settings. Launch option, session-only, and managed rules are read-only
- [VSCode] Live progress rows for running sub-agents are now displayed below tool call groups in the Focus view
- [VSCode] A Cancel button has been added to the account switching screen, allowing you to return to the session with the current account
- [Claude Code on the web] In cloud sessions, you can now cancel queued messages before Claude reads them. They can be removed from the queue, or pressing Esc or the up key returns the text to the message box
- [Claude Tag] In admin settings GitHub installation, a confirmation dialog is now shown before Connect all and Disconnect. This is described as a safeguard against accidentally making organization-wide changes
Improvements
- Responsiveness in long sessions has improved, as the entire conversation is no longer reprocessed to create summaries of collapsed tool uses on every transcript update
- Keyboard support over SSH and in unrecognized terminals has improved, enabling Shift+Enter and Ctrl+Shift shortcuts in terminals that respond to kitty keyboard queries (such as foot, Alacritty 0.16 and later)
- The
/diffpanel now opens in a rendered state without an intermediate loading state - A guide for
/focushas been added to spinner hints, allowing you to switch to a view showing only the prompt, a one-line work summary, and the response - The "Unknown skill" error for the Skill tool now shows the full name when the name without the plugin prefix matches exactly one plugin skill
/ultrareview --postnow directly posts a PR comment when results are ready and displays the comment link, without starting a second cloud session to post- Database reads for Artifact data saved to the session scratchpad no longer pause for working folder approval
- In first-party sessions with telemetry disabled,
alwaysLoadMCP servers that complete connection mid-conversation are now available from the next turn without a tool discovery round trip - Skills synced from claude.ai in cloud sessions are now named
anthropic-skills:<name>, matching Claude Desktop. They can still be referenced by the unprefixed name if nothing else shares that name - [VSCode] For documents and messages written for audiences other than the user, Claude now writes text directed at that reader and explicitly identifies the reader at the beginning of the reply
- [VSCode] Screen reader and keyboard accessibility has been improved for the slash command menu, @ mention menu, output style selection, send/stop buttons, permission and question cards, and the onboarding checklist
- [VSCode] Claude Code items have been removed from the session tab right-click menu and the "..." menu in the editor title bar, as they could not operate on the tab that was right-clicked
- [Claude Tag] Scheduled routines in Slack channels can now reply to existing threads instead of always posting new top-level messages
- [Claude Tag] Loading times for the admin settings page and Slack channel selection have improved. The effect is especially significant for organizations with many channels or those with multiple connected workspaces
Fixes
The major fixes beyond those covered in the Notable Updates section are as follows.
- Fixed terminal key input regression: Issues where F1, F2, and F4 did not work in kitty-protocol-compatible terminals, Delete did not work in st, Alt+arrow was treated as Escape in rxvt-unicode, and Shift+symbol inputted unshifted characters in WezTerm were fixed (regression from v2.1.247)
- Fixed reduced cache reuse on session resume: After interrupting during Claude's thinking and resuming, the way previous context was resent changed, sometimes reducing prompt cache reuse rates
- Fixed CLAUDE.md attribution rules being overwritten: Attribution reminder was overwriting CLAUDE.md and memory rules prohibiting attribution in commits and pull requests. Lines specified in managed settings continue to apply
- Fixed missing
permission_denialsrecords: Read, Edit, and Write calls blocked by path-scope deny rules were missing frompermission_denialsin--output-format stream-jsonresults - [Claude Tag] Fixed incorrect sharing scope display: The sharing banner and sharing dialog for sessions started from Slack indicated that links could be opened organization-wide. They now indicate the Slack channel scope
- [Claude Tag] Fixed responses with unapproved models: Claude was accepting model switches to models not enabled by the organization and silently responding with a fallback model. It now declines the switch and informs users that an administrator can enable it
- Fixed lingering LSP servers: Plugin LSP servers that do not accept
shutdownparameters (such as rust-analyzer) were continuing to run after the session ended.exitis now sent even ifshutdownfails - Fixed
/insightsfailures: It was failing in Bedrock, Vertex, Foundry, and gateway configurations where accounts cannot reach the default Opus model. These environments now use the session model - Fixed
/goalsilent stops: Execution was silently stopping after API errors, network disconnections, or token limits. It now retries with backoff or pauses with an explanation including waiting for usage limit resets - Fixed
/btwhallucinated tool calls: Responses sometimes included non-existent tool calls and their outputs. Claude is now instructed not to include these in side questions, and when they appear, they are shown as unexecuted - Fixed organization plugins not loading: Organization plugins enabled in managed settings were not loading in headless sessions and Claude Desktop (since the version bundling this CLI). They will load from the next session
- Fixed background PowerShell commands stopping on Windows: PowerShell tool commands sent to the background were being stopped when Claude Code exited
- [Claude Code on the web] Fixed duplicate execution of scheduled routines: An issue where one-time routines ran twice after temporary server errors, and an issue where executions using sub-agents were treated as finishing early, causing retries to be skipped or duplicate executions to start, were both fixed
- A quietly appreciated fix: On slow connections (SSH, browser-based terminals), an issue where strings like
22cand terminal color/version responses were being entered into the prompt at startup was fixed. Since this was the kind of bug that occurred at every startup over SSH, it quietly makes a noticeable difference - In addition, numerous minor bugs have been fixed in areas including full-screen redraw, returning from external editors, MCP server reconnection, the VSCode extension session list, and Claude Tag notifications
Trying out claude plugin eval
claude plugin eval is a command that scores and quantifies whether plugins (skills) are actually working, using test cases. Today's Claude Plugin Eval took some effort to validate. But once tamed, it should become a powerful ally.
claude plugin eval is a command that scores how well a plugin (skill) is working using test cases. It runs the same cases with and without the plugin, and the difference (Δ) represents the plugin's contribution. If both scores are high, it means the case was passing even without the plugin.
Terminology
Since the terminology is quite unfamiliar, I've summarized it here. These are terms that appear when creating and running cases.
Terms for composing a suite
| Term | What it refers to |
|---|---|
| eval (evaluation) | Scoring a plugin's behavior with test cases. While claude plugin validate examines file syntax, eval examines the results of actually running it |
| suite | The complete set of cases under evals/. One suite per plugin |
| case | A single test. One directory evals/<name>/ equals one case. Composed of one prompt and several graders |
| prompt | The request sent to Claude for that case. The body of prompt.md. Do not write the skill name; use the phrasing a user would actually type |
| grader | A single pass/fail judgment on an execution result. One file graders/<name>.md equals one grader |
| rubric | Text describing the scoring criteria for an llm grader. Lists PASS/FAIL conditions |
| judge | The model called by an llm grader for scoring. Specified with --judge-model. Should be different from the model being scored (using the same model leads to lenient self-scoring) |
| weight | A grader's point value. Assign the primary check a 1 and supplementary checks 0.5, for example |
Terms related to execution
| Term | What it refers to |
|---|---|
| run | Executing a case once. claude -p runs once in a disposable working directory |
runs |
The setting for how many times to run each case. Default is 3. Results from a single run can vary, so multiple runs are used |
| arm | A group of executions. The with arm = the group run with the plugin loaded; the without arm = the group run without the plugin |
| ablation | The setting for running a comparison without the plugin. --ablation with-without for 2 arms, none for the with arm only |
| baseline | The without arm. The reference line for comparison |
| Δ (delta) | The with score − the without score. This is the plugin's contribution and the number to focus on in an eval. Not the absolute score itself |
| pilot | A trial run before the full run. Runs just 1 run to verify that the graders work as intended and that Δ is produced. claude plugin eval init automatically performs this at the end of the dialog |
| full run | The production run. Runs all cases × 2 arms with the default runs (3). For 7 cases: 7 × 3 × 2 = 42 runs |
| threshold | The passing bar for a case. Default 1.0. If any case falls below it, the command returns exit 1 (a report, not an execution failure) |
| fires | When a skill is invoked. Measured with the tool_used: Skill grader |
Step 0: Preparing for verification
This appears when run from an empty directory or project root. This time, we create a plugin root with plugin.json (or .claude-plugin/plugin.json) and a skills folder with SKILL.md.
Project root
├── .claude-plugin
│ └── plugin.json
└── skills
└── commit-message
└── SKILL.md
Running the script below creates the skill for this verification in one go.
mkdir -p plugin-eval && cd plugin-eval
mkdir -p .claude-plugin skills/commit-message
cat > .claude-plugin/plugin.json <<'JSON'
{
"name": "commit-style",
"version": "0.1.0",
"description": "Draft commit messages in Conventional Commits format"
}
JSON
cat > skills/commit-message/SKILL.md <<'MD'
---
name: commit-message
description: Use when the user asks for a commit message for a change they describe.
---
# Commit message
Receives a description of changes and returns a single line commit message in Conventional Commits format.
- Format is `<type>(<scope>): <subject>`
- type is chosen from feat / fix / docs / refactor / test / chore
- subject is in English imperative form, within 72 characters, with no trailing period
- Return only the single commit message line, without any explanatory text, preamble, or body
MD
When you want to evaluate an already-used skill, move to ~/.claude/skills/<skill-name>/ and run the same command (it will be resolved as a plugin in the skills directory).
Step 2: Creating cases
Run from the plugin root (the plugin-eval directory you cd'd into in Step 1).
claude plugin eval init
Running the command opens an interactive session where Claude asks questions and you provide answers. Claude writes the cases and graders.

It proceeds through several confirmation gates, so answer the questions to continue. Claude drafts the cases and graders, runs a pilot to verify behavior, and writes them to evals/<case-name>/. The estimated cost for a full run is also presented there.

Step 3: Running it
The default is 3 runs with the plugin and 3 runs without, for a total of 6 runs per case. If you just want to confirm it works first, drop it to 1 run each with --runs 1 (since results from a single run can vary, return to the default 3 when making judgments about quality).

The full run completed and the suite is finished. The init dialog ends here. Add the final step (fix weaknesses and re-measure) to your procedure.
Exit the session with /exit and return to the shell. Under evals/plugin, 7 cases, FOLLOWUPS.md, and 3 generations of results remain.
plugin
├── 01-ja-prose
│ ├── graders
│ │ ├── conventional-type-prefix.md
│ │ ├── no-japanese.md
│ │ ├── no-wrapper.md
│ │ ├── one-line-only.md
│ │ ├── skill-fired.md
│ │ └── subject-quality.md
│ └── prompt.md
├── 02-ja-bullets
│ ├── graders
│ │ ├── conventional-type-prefix.md
│ │ ├── no-japanese.md
│ │ ├── no-wrapper.md
│ │ ├── one-line-only.md
│ │ ├── skill-fired.md
│ │ └── subject-quality.md
│ └── prompt.md
├── 03-diff-paste
│ ├── graders
│ │ ├── conventional-type-prefix.md
│ │ ├── no-japanese.md
│ │ ├── no-wrapper.md
│ │ ├── one-line-only.md
│ │ ├── skill-fired.md
│ │ └── subject-quality.md
│ └── prompt.md
├── 04-en-oneline
│ ├── graders
│ │ ├── conventional-type-prefix.md
│ │ ├── no-japanese.md
│ │ ├── no-wrapper.md
│ │ ├── one-line-only.md
│ │ ├── skill-fired.md
│ │ └── subject-quality.md
│ └── prompt.md
├── 05-ja-casual
│ ├── graders
│ │ ├── conventional-type-prefix.md
│ │ ├── no-japanese.md
│ │ ├── no-wrapper.md
│ │ ├── one-line-only.md
│ │ ├── skill-fired.md
│ │ └── subject-quality.md
│ └── prompt.md
├── 06-neg-explain
│ ├── graders
│ │ ├── answers-the-question.md
│ │ ├── not-a-bare-commit.md
│ │ └── skill-not-used.md
│ └── prompt.md
├── 07-neg-typelist
│ ├── graders
│ │ ├── answers-the-question.md
│ │ ├── not-a-bare-commit.md
│ │ └── skill-not-used.md
│ └── prompt.md
├── FOLLOWUPS.md
└── results
├── 2026-09-12T02-55-47-422Z
│ ├── aggregate-result.json
│ └── report.html
├── 2026-09-12T03-58-41-015Z
│ ├── aggregate-result.json
│ └── report.html
├── 2026-09-12T04-04-49-220Z
│ ├── aggregate-result.json
│ └── report.html
└── 2026-09-12T04-29-53-821Z
├── aggregate-result.json
└── report.html
Open the HTML report at evals/results/2026-09-12T04-29-53-821Z/report.html. It contains cards per case, grader judgments per run, judge votes, and the full text the judge saw. It has more information than the terminal summary, so be sure to check it.

Finally, address the items recorded in FOLLOWUPS.md. Fix the weaknesses found at the end and re-measure the effect as needed.
Thoughts on claude plugin eval
Until now, writing SKILL.md and CLAUDE.md was a matter of "write it and see what happens," with quality dependent on the author's intuition. Eval feels like a tool that brings testing and baseline comparison to that process. In particular, the comparison against a no-plugin baseline is a concept absent from unit testing — it's closer to an A/B test mindset.
The biggest takeaway from running it end-to-end was seeing in numbers the obvious truth that what you write is not guaranteed to be followed, and what you don't write is certainly not followed. Even if you write "return only one line," fences still appear. Conversely, there are things vanilla Claude can do even without being told. Being able to separate these two cases is what gives eval its value.
On the other hand, the upfront effort for skill development increases. What was once "just write it and distribute" now adds case design, grader design, and pilot iteration. I see this effort as justified only for things intended for distribution. If a skill is only for personal use, if it doesn't work, only you are affected. If a skill distributed internally turns out to do nothing at all, it ends up consuming everyone's tokens continuously. Being able to verify that numerically is where the practical value lies.
The impression is that the barrier to quickly evaluating a Skill/Plugin is fairly high. Those actually trying it out should refer to the following documentation.
Closing
On the plugin front, a new feature in the form of an evaluation mechanism and a security fix for extracted archive permissions were introduced simultaneously. For those distributing plugins internally, we recommend prioritizing the permission fix before setting up evaluations.
For those enforcing operations with permission rules or using Claude Code on shared machines, why not update and give it a try?
References
