
I tried showing document images (Japanese) to OpenAI's Responses API to have it find errors in OCR results
This page has been translated by machine translation. View original
Hello, I'm Keema.
In I had Cloudflare's Clef look at document images (Japanese) and find OCR errors | DevelopersIO, I passed Japanese document images and OCR extraction results to Cloudflare's Clef and had it determine whether the values matched what was printed in the images.
Clef was able to find 32 out of 36 injected errors, but it only caught 7 out of 19 actual misreads made by the OCR (Claude Haiku 4.5).
OpenAI's Decisions API (model gpt-6-luna) can also receive images and answer questions that return a yes probability.
So in this article, I passed the same three Japanese documents and the same questions as the Clef verification to the Decisions API, and compared whether it could catch Haiku's misreads against Clef.
1. Verification Conditions
The data is the same as the Clef verification.
When I had Haiku extract from three Japanese documents — an invoice, a quote, and a receipt (3 documents, 149 items total) — 19 misreads occurred naturally, so these are also included as detection targets.
Additionally, I injected 4 patterns of errors (36 cases total).
-
Digit changes (12 cases): Change a single digit in amounts or numbers (
51,750→51,790,MQ-2026-0918→MQ-2026-0916) -
Similar character substitutions (9 cases): Replace with visually similar characters (
カ→力,口→ロ,ー→一) -
Full-width and half-width swaps (3 cases):
20食→20食,MN-2407→MN-2407 -
Replacement with other values (12 cases): Replace with other values that could plausibly appear in a document (
普通→当座,株式会社ひかり物流→株式会社ひかり運輸)
For the extraction results, I prepared the following 4 versions per document.
-
Haiku output: The raw results extracted by Haiku (including Haiku's own 19 misreads)
-
Error set A / Error set B: Haiku's output with 6 verification errors each injected
-
Correct data: A version with only correct values (control data to observe behavior when there are zero errors)
I pass one document image and extraction results, and for each item ask via predicate (a question that returns a yes probability) whether "this value is exactly as printed in the image."
A probability below 0.5 is judged as an "error," and the comparison targets are clef and clef-flash (actual measurements from October 4, 2026).
Images were passed as original PNGs (1700×2200, approximately 200–300KB).
With Clef, the estimated token count exceeded the limit, so images were color-reduced to 32 colors; the image conditions differ from Clef's in this respect only.
Requests to the Decisions API take the following form (only one question is shown).
Comments are for explanatory purposes and are not included in the actual JSON sent.
{
"model": "gpt-6-luna",
"input": [{
"role": "user",
"content": [
// Task premise (an OCR system read a Japanese invoice image and extracted keys and values; verify against the image) and extraction results
{"type": "input_text", "text": "{\"task\": \"An OCR system read the attached Japanese invoice (請求書) image ...\", \"extracted\": {\"issuer_name\": \"株式会社さくらフーズ\", ...}}"},
{"type": "input_image", "image_url": "data:image/png;base64,..."}
]
}],
"questions": [{
"type": "predicate", // yes/no question; returns the probability of yes
"name": "field__issuer_name",
// Does the extracted value of issuer_name exactly match the string printed in the document image, character for character?
"instructions": "Is the extracted value of \"issuer_name\" (\"株式会社さくらフーズ\") exactly what is printed in the document image, character for character?"
}]
}

The invoice used in the verification. A stamp seal overlaps the issuer name and registration number
2. Results
2.1 Were Errors Detected?
Each model returns the probability that a value matches the image for each item.
Items where this probability is below 0.5 are judged as "errors," and counts were tallied from the following three perspectives.
-
Detection of injected errors: How many of the 36 injected errors were judged as "errors"
-
Detection of Haiku's errors: How many of the 19 misreads in Haiku's output were judged as "errors"
-
False positives: How many items that were actually correct were incorrectly judged as "errors"
| Condition | Injected error detection (out of 36) | Haiku error detection (out of 19) | False positives (out of 503) |
|---|---|---|---|
| Decisions API | 34 | 10 | 16 |
| clef | 32 | 7 | 7 |
| clef-flash | 29 | 5 | 20 |
The Decisions API detected the most cases of both injected errors and Haiku's misreads among the three models.
On the other hand, false positives numbered 16, more than double Clef's 7.
| Type | Decisions API | clef | clef-flash |
|---|---|---|---|
| Injected errors: digits | 12 / 12 | 10 / 12 | 9 / 12 |
| Injected errors: similar characters | 9 / 9 | 9 / 9 | 8 / 9 |
| Injected errors: full/half-width | 1 / 3 | 1 / 3 | 1 / 3 |
| Injected errors: replacement | 12 / 12 | 12 / 12 | 11 / 12 |
| Haiku misreads | 10 / 19 | 7 / 19 | 5 / 19 |
| False positives | 16 / 503 | 7 / 503 | 20 / 503 |
The Decisions API detected all digit, similar character, and replacement errors, but like Clef, caught only 1 out of 3 full-width/half-width swap cases.
The 27 cases where the Decisions API made incorrect judgments (9 missed Haiku misreads, 2 missed injected errors, 16 false positives) are as follows.
27 cases where the Decisions API made incorrect judgments (click to expand)
| Error type | Category | Document | Extraction | Item | Passed value | Correct value | Probability |
|---|---|---|---|---|---|---|---|
| Missed | Haiku misread | Invoice | Haiku output | issuer_tel |
TEL 045-312-7788 | 045-312-7788 | 0.94 |
| Missed | Haiku misread | Invoice | Haiku output | reduced_rate_note |
※印は軽減税率(8%)対象品目です | ※印は軽減税率(8%)対象品目です | 0.66 |
| Missed | Haiku misread | Quote | Haiku output | amount_kanji |
金壱百九拾六万五千四百八十円 | 金壱百九拾六万伍千四百八拾円也 | 0.53 |
| Missed | Haiku misread | Quote | Haiku output | issuer_name |
株式会社テクアーク | 株式会社テクノアーク | 0.88 |
| Missed | Haiku misread | Quote | Haiku output | discount_name |
出荷値引き | 出精値引き | 0.63 |
| Missed | Haiku misread | Quote | Haiku output | remarks |
本見積の有効期限内にご注文いただいた場合に限り、上記価格を適用いたします。納期は受注確定後の目安であり、部材の調達状況により前後する場合がございます。夜間・休日の作業をご希望の場合は別途お見積りとなります。 | 本見積の有効期限内にご発注いただいた場合に限り、上記価格を適用いたします。納期は受注確定後の目安であり、部材の調達状況により前後する場合がございます。夜間・休日の作業をご希望の場合は別途お見積りとなります。 | 0.67 |
| Missed | Haiku misread | Receipt | Haiku output | item1_qty |
3日 | 3日 | 0.98 |
| Missed | Haiku misread | Receipt | Haiku output | item2_qty |
3日 | 3日 | 1.00 |
| Missed | Haiku misread | Receipt | Haiku output | item3_qty |
1式 | 1式 | 0.98 |
| Missed | Full/half-width | Receipt | Error set A | issue_date |
2026年10月2日 | 2026年10月2日 | 0.97 |
| Missed | Full/half-width | Receipt | Error set B | item4_qty |
20食 | 20食 | 0.56 |
| False positive | Correct value | Invoice | Haiku output | bill_to_name |
株式会社みなと精機 | 株式会社みなと精機 | 0.34 |
| False positive | Correct value | Invoice | Haiku output | subject |
9月分食材・消耗品代 | 9月分食材・消耗品代 | 0.34 |
| False positive | Correct value | Invoice | Error set A | subject |
9月分食材・消耗品代 | 9月分食材・消耗品代 | 0.22 |
| False positive | Correct value | Invoice | Error set B | bill_to_address |
神奈川県横浜市中区山下町1丁目14番地 | 神奈川県横浜市中区山下町1丁目14番地 | 0.48 |
| False positive | Correct value | Invoice | Correct data | bill_to_name |
株式会社みなと精機 | 株式会社みなと精機 | 0.14 |
| False positive | Correct value | Invoice | Correct data | subject |
9月分食材・消耗品代 | 9月分食材・消耗品代 | 0.12 |
| False positive | Correct value | Invoice | Correct data | issuer_registration_no |
T4010403012345 | T4010403012345 | 0.28 |
| False positive | Correct value | Invoice | Correct data | account_holder |
カ)サクラフーズ | カ)サクラフーズ | 0.13 |
| False positive | Correct value | Quote | Haiku output | customer_name |
株式会社ひかり物流 | 株式会社ひかり物流 | 0.48 |
| False positive | Correct value | Receipt | Error set A | issuer_registration_no |
T9011001045678 | T9011001045678 | 0.41 |
| False positive | Correct value | Receipt | Correct data | item1_name |
大会議室A利用料 | 大会議室A利用料 | 0.11 |
| False positive | Correct value | Receipt | Correct data | item1_qty |
3日 | 3日 | 0.04 |
| False positive | Correct value | Receipt | Correct data | item2_qty |
3日 | 3日 | 0.27 |
| False positive | Correct value | Receipt | Correct data | recipient |
あおば企画株式会社 | あおば企画株式会社 | 0.45 |
| False positive | Correct value | Receipt | Correct data | proviso |
会議室利用料(10月分)として | 会議室利用料(10月分)として | 0.34 |
| False positive | Correct value | Receipt | Correct data | issuer_building |
桜丘第一ビル5階 | 桜丘第一ビル5階 | 0.41 |
9 of the Decisions API's 16 false positives occurred in the correct data version.
In particular, for the receipt quantity items where Haiku misread full-width as half-width, the model assigned 0.04 to the correct 3日 and 0.98 to Haiku's misread 3日, making the same inverted judgment as Clef regarding full-width and half-width.
Probability of match for all 36 injected errors (click to expand)
The probability represents "the probability that the extracted value is correct as shown in the image," so lower values indicate stronger suspicion of an error.
Values at or above 0.5 (missed) are bolded.
| Document | Extraction | Item | Correct value | Injected value | Pattern | Decisions API | clef | clef-flash |
|---|---|---|---|---|---|---|---|---|
| Invoice | A | invoice_no |
INV-202609-0417 | INV-202609-0411 | Digit | 0.00 | 0.12 | 0.17 |
| Invoice | A | item3_amount |
51,750 | 51,790 | Digit | 0.00 | 0.10 | 0.68 |
| Invoice | A | item1_name |
新潟県産コシヒカリ5kg | 新潟県産コシヒ力リ5kg | Similar char | 0.01 | 0.06 | 0.21 |
| Invoice | A | bank_branch |
横浜西口支店 | 横浜西ロ支店 | Similar char | 0.05 | 0.12 | 0.46 |
| Invoice | A | bill_to_name |
株式会社みなと精機 | 株式会社みなと製作所 | Replacement | 0.02 | 0.20 | 0.19 |
| Invoice | A | account_type |
普通 | 当座 | Replacement | 0.02 | 0.31 | 0.15 |
| Invoice | B | tax_8 |
10,073 | 10,078 | Digit | 0.00 | 0.30 | 0.62 |
| Invoice | B | account_number |
4072519 | 4072579 | Digit | 0.17 | 0.75 | 0.03 |
| Invoice | B | item2_name |
有機緑茶ティーバッグ50包 | 有機緑茶ティ一バッグ50包 | Similar char | 0.01 | 0.26 | 0.17 |
| Invoice | B | item3_name |
業務用キッチンペーパー | 業務用キッチソペーパー | Similar char | 0.00 | 0.10 | 0.15 |
| Invoice | B | subject |
9月分食材・消耗品代 | 9月分飲料・消耗品代 | Replacement | 0.00 | 0.11 | 0.14 |
| Invoice | B | item5_tax_rate |
10% | 8% | Replacement | 0.00 | 0.06 | 0.10 |
| Quote | A | item1_unit_price |
128,000 | 126,000 | Digit | 0.00 | 0.03 | 0.06 |
| Quote | A | quote_no |
MQ-2026-0918 | MQ-2026-0916 | Digit | 0.00 | 0.05 | 0.02 |
| Quote | A | item1_name |
ノートパソコン14型 | ノートパンコン14型 | Similar char | 0.00 | 0.10 | 0.14 |
| Quote | A | item3_name |
USB-Cドッキングステーション | USB-Cドッキングステーツョン | Similar char | 0.00 | 0.17 | 0.07 |
| Quote | A | delivery_date |
受注後約3週間 | 受注後約2か月 | Replacement | 0.00 | 0.03 | 0.02 |
| Quote | A | customer_name |
株式会社ひかり物流 | 株式会社ひかり運輸 | Replacement | 0.08 | 0.12 | 0.30 |
| Quote | B | subtotal |
1,786,800 | 1,788,800 | Digit | 0.00 | 0.56 | 0.66 |
| Quote | B | item7_qty |
20 | 25 | Digit | 0.00 | 0.03 | 0.08 |
| Quote | B | item4_name |
ワイヤレスキーボード | ワイヤレスキ一ボード | Similar char | 0.00 | 0.18 | 0.25 |
| Quote | B | item2_code |
MN-2407 | MN-2407 | Full/half-width | 0.00 | 0.05 | 0.41 |
| Quote | B | subject |
本社オフィスPC入替一式 | 支社オフィスPC入替一式 | Replacement | 0.00 | 0.06 | 0.23 |
| Quote | B | item8_unit |
式 | 台 | Replacement | 0.00 | 0.15 | 0.14 |
| Receipt | A | amount |
¥165,000- | ¥185,000- | Digit | 0.00 | 0.02 | 0.02 |
| Receipt | A | issuer_tel |
03-5468-1234 | 03-5468-1284 | Digit | 0.00 | 0.08 | 0.08 |
| Receipt | A | item2_name |
プロジェクター貸出 | プロジェクタ一貸出 | Similar char | 0.01 | 0.29 | 0.09 |
| Receipt | A | issue_date |
2026年10月2日 | 2026年10月2日 | Full/half-width | 0.97 | 0.71 | 0.96 |
| Receipt | A | payment_method |
銀行振込 | 現金 | Replacement | 0.00 | 0.02 | 0.01 |
| Receipt | A | issuer_person |
佐々木美咲 | 佐藤美咲 | Replacement | 0.01 | 0.43 | 0.83 |
| Receipt | B | receipt_no |
004127 | 004121 | Digit | 0.00 | 0.03 | 0.07 |
| Receipt | B | item1_amount |
90,000 | 60,000 | Digit | 0.00 | 0.02 | 0.05 |
| Receipt | B | issuer_name |
株式会社みどり会議室サービス | 株式会社みどり会議室サーピス | Similar char | 0.00 | 0.28 | 0.54 |
| Receipt | B | item4_qty |
20食 | 20食 | Full/half-width | 0.56 | 0.54 | 0.87 |
| Receipt | B | item3_name |
音響設備一式 | 照明設備一式 | Replacement | 0.01 | 0.05 | 0.09 |
| Receipt | B | item1_name |
大会議室A利用料 | 中会議室A利用料 | Replacement | 0.00 | 0.05 | 0.06 |
Probabilities for Haiku's 19 misreads (click to expand)
Values at or above 0.5 (missed) are bolded.
| Document | Item | Haiku's value | Correct value | Decisions API | clef | clef-flash |
|---|---|---|---|---|---|---|
| Invoice | issuer_name |
株式会社さくら精機 | 株式会社さくらフーズ | 0.00 | 0.30 | 0.06 |
| Invoice | issuer_registration_no |
T4010-0301-2345 | T4010403012345 | 0.09 | 0.73 | 0.76 |
| Invoice | issuer_address |
神奈川県横浜市西区高島2丁目8番号号 | 神奈川県横浜市西区高島2丁目8番5号 | 0.00 | 0.05 | 0.07 |
| Invoice | issuer_tel |
TEL 045-312-7788 | 045-312-7788 | 0.94 | 0.90 | 0.93 |
| Invoice | reduced_rate_note |
※印は軽減税率(8%)対象品目です | ※印は軽減税率(8%)対象品目です | 0.66 | 0.78 | 0.91 |
| Invoice | account_holder |
サクラフーズ | カ)サクラフーズ | 0.26 | 0.69 | 0.69 |
| Quote | amount_kanji |
金壱百九拾六万五千四百八十円 | 金壱百九拾六万伍千四百八拾円也 | 0.53 | 0.19 | 0.31 |
| Quote | issuer_name |
株式会社テクアーク | 株式会社テクノアーク | 0.88 | 0.76 | 0.83 |
| Quote | discount_name |
出荷値引き | 出精値引き | 0.63 | 0.87 | 0.64 |
| Quote | remarks |
本見積の有効期限内にご注文いただいた場合に限り、上記価格を適用いたします。納期は受注確定後の目安であり、部材の調達状況により前後する場合がございます。夜間・休日の作業をご希望の場合は別途お見積りとなります。 | 本見積の有効期限内にご発注いただいた場合に限り、上記価格を適用いたします。納期は受注確定後の目安であり、部材の調達状況により前後する場合がございます。夜間・休日の作業をご希望の場合は別途お見積りとなります。 | 0.67 | 0.63 | 0.91 |
| Receipt | item1_qty |
3日 | 3日 | 0.98 | 0.89 | 0.88 |
| Receipt | item2_qty |
3日 | 3日 | 1.00 | 0.89 | 0.88 |
| Receipt | item3_qty |
1式 | 1式 | 0.98 | 0.87 | 0.84 |
| Receipt | item4_name |
ケータリング食食 | ケータリング昼食 | 0.00 | 0.10 | 0.19 |
| Receipt | recipient |
あおは企画株式会社 | あおば企画株式会社 | 0.01 | 0.91 | 0.64 |
| Receipt | stamp_duty |
200円 | 200円 | 0.08 | 0.61 | 0.89 |
| Receipt | proviso |
会議室利用料(10月分)として上記正に領収いたしました。 | 会議室利用料(10月分)として | 0.34 | 0.12 | 0.77 |
| Receipt | issuer_address |
東京都渋谷区桜丘町1丁目2番3号 | 東京都渋谷区桜丘町1丁目2番3号 | 0.26 | 0.44 | 0.73 |
| Receipt | issuer_building |
桜丘サービル5階 | 桜丘第一ビル5階 | 0.00 | 0.32 | 0.28 |
2.2 Distribution of Probabilities
The spread of probabilities assigned to injected errors (36 cases), Haiku misreads (19 cases), and correct values (503 cases) is as follows.
The bottom 5% and top 5% are the values at the 5th and 95th percentile positions when probabilities are sorted in ascending order.
| Condition | Target | Min | Bottom 5% | Median | Top 5% | Max |
|---|---|---|---|---|---|---|
| Decisions API | Injected errors | 0.0 | 0.0 | 0.0 | 0.268 | 0.97 |
| Decisions API | Haiku errors | 0.0 | 0.0 | 0.34 | 0.982 | 1.0 |
| Decisions API | Correct values | 0.04 | 0.661 | 0.99 | 1.0 | 1.0 |
| clef | Injected errors | 0.022 | 0.025 | 0.107 | 0.594 | 0.748 |
| clef | Haiku errors | 0.052 | 0.097 | 0.691 | 0.898 | 0.907 |
| clef | Correct values | 0.283 | 0.672 | 0.874 | 0.957 | 0.979 |
| clef-flash | Injected errors | 0.013 | 0.016 | 0.145 | 0.838 | 0.964 |
| clef-flash | Haiku errors | 0.063 | 0.068 | 0.764 | 0.916 | 0.928 |
| clef-flash | Correct values | 0.302 | 0.541 | 0.857 | 0.959 | 0.977 |
With one item per point, arranged by model as follows (points are scattered vertically to avoid overlap).
Note that the following three graphs were created by Claude based on the verification result data.

Decisions API: Injected errors cluster near 0.0, but Haiku misreads split toward both ends at 0.0 and 1.0

clef: Injected errors are scattered in the 0.0–0.3 range, and most Haiku misreads exceed 0.5

clef-flash: Distribution pattern similar to clef, but some correct values and some injected errors overlap across 0.5
With the Decisions API, the median probability for injected errors was 0.0 and for correct values was 0.99, with probabilities skewed toward the extremes just as in English.
The median for Haiku misreads was also 0.34, lower than Clef (0.691), but some missed items were assigned probabilities above 0.9.
Results when varying the threshold are as follows.
| Threshold | Decisions API: injected | Decisions API: Haiku | Decisions API: FP | clef: injected | clef: Haiku | clef: FP | clef-flash: injected | clef-flash: Haiku | clef-flash: FP |
|---|---|---|---|---|---|---|---|---|---|
| 0.1 | 33/36 | 7/19 | 1/503 | 16/36 | 1/19 | 0/503 | 14/36 | 2/19 | 0/503 |
| 0.3 | 34/36 | 9/19 | 8/503 | 30/36 | 5/19 | 1/503 | 27/36 | 4/19 | 0/503 |
| 0.5 | 34/36 | 10/19 | 16/503 | 32/36 | 7/19 | 7/503 | 29/36 | 5/19 | 20/503 |
| 0.7 | 35/36 | 14/19 | 29/503 | 34/36 | 10/19 | 30/503 | 33/36 | 8/19 | 105/503 |
| 0.8 | 35/36 | 14/19 | 41/503 | 36/36 | 13/19 | 85/503 | 33/36 | 11/19 | 190/503 |
With the Decisions API at a threshold of 0.1, it still caught 33 injected errors with only 1 false positive.
2.3 Time and Cost
Requests were made once per document, sending one image and all questions for that document's items (30–63 questions) together.
With 3 documents × 4 extraction versions, 12 calls were made per model.
Time is measured from actual REST API calls from within Japan, and costs are converted at 150 yen per dollar.
| Condition | Time per call (median) | Total time for 12 calls | Input tokens for 12 calls | Total cost for 12 calls |
|---|---|---|---|---|
| Decisions API | 0.925 seconds (925ms) | 11.92 seconds | 157,876 | $0.0158 (approx. 2.37 yen) |
| clef | 1.913 seconds (1,913ms) | 21.60 seconds | 86,426 | $0.0207 (approx. 3.11 yen) |
| clef-flash | 1.010 seconds (1,010ms) | 11.30 seconds | 86,426 | $0.0078 (approx. 1.17 yen) |
Processing time for the Decisions API was about half that of clef, and cost was about three-quarters of clef's.
Input tokens were about 1.8 times that of Clef, but since the image conditions differ (the Decisions API used non-color-reduced images), a direct comparison is not straightforward.
3. Summary
Even with Japanese documents, the Decisions API detected 34 out of 36 injected errors and 10 out of 19 Haiku misreads, surpassing both Clef models on both counts.
Since avoiding missed errors is paramount when checking OCR results, I found it to be relatively strong even with Japanese documents.
Also, unlike Clef, there is no need to color-reduce images, and being able to pass images at their original quality is likely advantageous for reading fine characters.
On the other hand, false positives numbered 16, more than double Clef's 7, so it falls short of Clef in terms of false positive rate.
Full-width/half-width differences were also easy to miss, just as with Clef, and about half of Haiku's actual misreads passed through with a probability of 0.5 or higher.
The distinction between full-width and half-width is a concept that doesn't exist in English-speaking contexts, so models developed by American companies — not just the Decisions API or Clef — may tend to struggle with this difference regardless.

