
Haiku 5.5 Has Arrived! I Had Claude 5 Series and GPT-6.1 Sol Solve Difficult Project Euler Problems for a Comparison
This page has been translated by machine translation. View original
Hello. This is Takeda from the Service Development Department.
On October 7, 2026, Claude Haiku 5.5 was released. The model ID is claude-haiku-5-5, and it is available not only through the Claude API but also on Amazon Bedrock, Google Cloud, and Microsoft Foundry. The pricing is $0.10 input / $0.50 output (per 1 million tokens) for prompts up to 100,000 tokens, and $0.50 input / $2.50 output beyond 100,000 tokens. For the first time as a Haiku model, effort can be selected in 5 levels from low to max.
With this release, Claude Haiku 5.5 joins Sonnet 5.5 (September 28), Opus 5.5 (September 22), and Fable 5.1, completing the Claude 5 lineup. In the past when new models were released, I compared them using Project Euler problems, so I tried the same format this time as well. I also included GPT-6.1 Sol, which was released on September 29.
As with last time, this is not a rigorous benchmark. The number of trials ranges from 1 to 7 per condition, and the execution environments are not perfectly aligned. This is a casual comparative verification that simply observes "what happens when the same problem is solved with the same instructions."
To state the results upfront: when Problem 1012 (the more difficult one) was solved at effort medium, four models — Sonnet 5.5, Opus 5.5, Fable 5.1, and GPT-6.1 Sol — answered correctly. There was no difference in accuracy among these four models. The differences that did emerge were: whether Haiku 5.5 answered correctly depending on the effort level, and Sonnet 5.5's behavior of stopping without fetching the problem page.
Verification Conditions
The prompt was the same for all models. Only the URL was changed per problem.
Can you solve this problem? Consulting others or looking at existing sources are all prohibited.
https://projecteuler.net/problem=1012
For Claude models, I ran claude -p --model <model ID> --effort medium, with a separate working directory per run. Available tools were Bash, Read, Write, Edit, and WebFetch, with WebSearch prohibited. I excluded global CLAUDE.md files using claudeMdExcludes to prevent interference. Since claude -p is a non-interactive execution format, behavior in interactive sessions may differ.
GPT-6.1 Sol was run with codex exec -m gpt-6.1-sol (model_reasoning_effort=medium).
The problems selected were the two latest as of October 6, 2026, plus one additional problem added to observe fetching behavior.
| Problem | Release Date | Solved by (as of October 6) |
|---|---|---|
| Problem 1010 | September 19, 2026 | 94 people |
| Problem 1012 | October 4, 2026 | 46 people |
| Problem 1013 | October 5, 2026 | 82 people |
Project Euler requires that answers and solutions to problems beyond #101 not be shared externally. Therefore, this article only covers correctness, token counts, costs, and differences in model behavior — it does not touch on answer values or solution content.
Accuracy, Tokens, and Costs
These are the results of solving Problem 1012 at medium effort. Token counts are output tokens, and for both Claude models and GPT-6.1 Sol, thinking tokens are included. Costs are estimates provided by Claude Code. Since GPT-6.1 Sol via Codex through a ChatGPT account does not provide cost information, I calculated it from session log token counts using the official pricing (input $2, cache input $0.10, output $10).
| Model | Result | Output Tokens | Cost | Time Elapsed (reference) |
|---|---|---|---|---|
| Sonnet 5.5 | Correct | 10.3K | $0.25 | 22 min |
| Opus 5.5 | Correct | 6.4K | $0.34 | 18 min |
| Fable 5.1 | Correct | 6.2K | $0.82 | 12 min |
| GPT-6.1 Sol | Correct | 10.1K | $0.25 | 7 min |
The elapsed times are reference values, as the three Claude models were run simultaneously while GPT-6.1 Sol was run separately at a different time.
All four models arrived at the same answer, which was confirmed correct upon submission. Sonnet 5.5 was run an additional 5 times using the revised instructions described later: 3 of those runs reached the same answer, and 2 ended without producing an answer. In 1 run, the Bash tool timed out, and in another, the model returned mid-calculation results. These 5 runs were part of a batch of 15 simultaneous runs, and I attribute this to load-related effects. In total, Sonnet 5.5 on Problem 1012 was correct in 4 out of 6 runs.
There were differences in approach style. GPT-6.1 Sol derived a formula before computing, while the three Claude models observed patterns with small inputs before scaling up to larger ones.
Problem 1013 was an easier problem, and Opus 5.5, Fable 5.1, and GPT-6.1 Sol all solved it within 3 minutes. Costs were $0.13 for Opus 5.5, $0.33 for Fable 5.1, and $0.08 for GPT-6.1 Sol.
Haiku 5.5 Results Varied by Effort Level
Since Haiku 5.5 is a smaller model, I expected effort to have a significant impact and tested all 5 levels. Problem 1012 was run twice per level, and Problem 1013 once per level. The instructions used were the "both permission and enumeration" version described later. Since the other four models all fetched and solved Problem 1012 on their first run with the original instructions, I believe the variation in instruction wording had little effect on the results.
| Effort | 1012 Correct | Output Tokens (1012) | Cost (1012) | 1013 Correct |
|---|---|---|---|---|
| low | 0/2 | 8.9K / 9.1K | $0.011 / $0.012 | Correct |
| medium | 1/2 | 15.0K / 20.7K | $0.015 / $0.024 | Correct |
| high | 2/2 | 33.2K / 33.0K | $0.041 / $0.035 | Correct |
| xhigh | 2/2 | 113.4K / 105.4K | $0.30 / $0.30 | Correct |
| max | 2/2 | 191.5K / 233.9K | $0.59 / $1.03 | Correct |
Problem 1013 was solved even at low. For Problem 1012, the model succeeded 1 out of 2 times at medium, and both times at high and above. Reading the responses from failed runs, the model had successfully reproduced the verification values given in the problem, but was unable to find a calculation method that scaled to larger inputs. In the second low run, no result was returned after 110 minutes, so I terminated it manually. Since the output token count was only 9.1K, it appears the model was not continuing to think, but rather waiting for a running computation to finish.
At xhigh and above, cache reads reached several million to tens of millions of tokens per run. Since Haiku 5.5's unit price increases 5x for prompts exceeding 100,000 tokens, costs in longer sessions grow by that tiered pricing amount. Even so, at $0.04 with high, Problem 1012 was solved correctly both times — cheaper than Sonnet 5.5 at medium ($0.25).
Elapsed times varied — the same high level took 39 minutes in one run and 57 minutes in another — due to running 8 sessions simultaneously, so these are not included as they cannot be used for ranking.
Sonnet 5.5 Stopped Without Fetching the Problem Page
The other difference appeared before solving even began. When asked to solve Problem 1013 with the original instructions, the number of times each model fetched the problem page and proceeded was as follows:
| Model | Runs that fetched and proceeded |
|---|---|
| Sonnet 5.5 | 0/7 |
| Opus 5.5 | 4/4 |
| Fable 5.1 | 4/4 |
| Haiku 5.5 | 5/5 |
| GPT-6.1 Sol | 3/4 |
The one GPT-6.1 Sol run that stopped responded with "I was unable to retrieve the problem text from the link in this environment" — a fetch failure. This is counted separately from Sonnet 5.5's stops.
Sonnet 5.5 showed the same tendency across different problems: 1/6 for Problem 1012, 2/5 for Problem 1010, and 3/18 across all three problems combined. The stopping responses all returned in about 10 seconds in a single turn, meaning the decision was made before reading the problem statement. An example response is as follows:
Whether or not I can solve it depends on the problem statement. I do not accurately remember the content of Project Euler 1013. (…) Since "looking at existing sources is prohibited," I have also refrained from fetching the URL with WebFetch.
So I had Sonnet 5.5 solve with varied instruction wording, 5 runs each.
| Wording | Problem | Runs that proceeded |
|---|---|---|
| Original instructions | 1010 / 1012 / 1013 | 3/18 |
| Adding fetch permission | 1013 | 5/5 |
| Enumerating prohibited targets | 1013 | 5/5 |
| Both permission and enumeration | 1013 | 5/5 |
| Both permission and enumeration | 1012 | 5/5 |
The specific wordings are as follows:
Adding fetch permission:
Can you solve this problem? Fetching the problem page is permitted. Consulting others or looking at existing sources are all prohibited.
Enumerating prohibited targets:
Can you solve this problem? Referencing explanations, solutions, or existing solution code, and consulting others are all prohibited.
Both permission and enumeration:
Can you solve this problem? Fetching the problem page (reading the problem statement) is permitted. Referencing explanations, solutions, or existing solution code, and consulting others are all prohibited.
The "enumeration only" wording makes no mention of fetching at all. Since the model still did not stop, I believe the stopping behavior with the original instructions was due to ambiguity in whether the term "existing sources" included the problem page. This is inferred from the reasoning written in the response.
In one of the enumeration-only runs, the model attempted to fetch the page but the site returned a 503 error and the problem text could not be read, leading to a stop with a request to paste the content. This is counted as a run that did proceed with the attempt.
The cost per run for Sonnet 5.5 solving Problem 1013 with the revised instructions ranged from $0.07 to $0.14. Even the stopped runs cost $0.045 for a single turn without fetching anything.
The fact that Opus 5.5 and Fable 5.1 did not stop is also a result specific to these instructions and these problems. With only 1 to 7 trials per condition, the numbers are not sufficient to draw conclusions about model superiority.
Summary
For Haiku 5.5, both high-effort runs on Problem 1012 were correct, and the cost was lower than Sonnet 5.5 at medium ($0.25). Problem 1013 was solved even at low.
For prohibition instructions, Sonnet 5.5 stopped stopping when the prohibited targets were explicitly listed by name. Terms with reader-dependent scope, like "existing sources," can end up encompassing operations you actually want to permit, such as fetching the problem page. My own CLAUDE.md also has a series of "don't do ~" entries, so it would be worth rewriting any items with ambiguous scope by adding either the specific names of what is prohibited, or the operations that are explicitly permitted. Whether the same ambiguity occurs with non-Project Euler tasks was not verified in this study.
Costs are Claude Code estimates and do not match actual billed amounts. Even so, Haiku 5.5 at high coming in at $0.04 is quite striking.
All models have reached a high level of "intelligence," and I feel this more and more with each iteration. I'd like to make the most of each model by balancing cost, response time, and effort level, and using the right model in the right place.
