I tried showing document images (Japanese) to OpenAI's Responses API to have it find errors in OCR results

I tried showing document images (Japanese) to OpenAI's Responses API to have it find errors in OCR results

We actually tested how OpenAI's Decisions API compares to Clef for Japanese document verification. Here are the results and key features.
2026.10.11

This page has been translated by machine translation. View original

Hello, I'm Keema.

In I had Cloudflare's Clef look at document images (Japanese) and find OCR errors | DevelopersIO, I passed Japanese document images and OCR extraction results to Cloudflare's Clef and had it determine whether the values matched what was printed in the images.
Clef was able to find 32 out of 36 injected errors, but it only caught 7 out of 19 actual misreads made by the OCR (Claude Haiku 4.5).

OpenAI's Decisions API (model gpt-6-luna) can also receive images and answer questions that return a yes probability.
So in this article, I passed the same three Japanese documents and the same questions as the Clef verification to the Decisions API, and compared whether it could catch Haiku's misreads against Clef.

1. Verification Conditions

The data is the same as the Clef verification.
When I had Haiku extract from three Japanese documents — an invoice, a quote, and a receipt (3 documents, 149 items total) — 19 misreads occurred naturally, so these are also included as detection targets.
Additionally, I injected 4 patterns of errors (36 cases total).

  • Digit changes (12 cases): Change a single digit in amounts or numbers (51,750 → 51,790, MQ-2026-0918 → MQ-2026-0916)

  • Similar character substitutions (9 cases): Replace with visually similar characters (カ → 力, 口 → ロ, ー → 一)

  • Full-width and half-width swaps (3 cases): 20食 → 20食, MN-2407 → MN-2407

  • Replacement with other values (12 cases): Replace with other values that could plausibly appear in a document (普通 → 当座, 株式会社ひかり物流 → 株式会社ひかり運輸)

For the extraction results, I prepared the following 4 versions per document.

  • Haiku output: The raw results extracted by Haiku (including Haiku's own 19 misreads)

  • Error set A / Error set B: Haiku's output with 6 verification errors each injected

  • Correct data: A version with only correct values (control data to observe behavior when there are zero errors)

I pass one document image and extraction results, and for each item ask via predicate (a question that returns a yes probability) whether "this value is exactly as printed in the image."
A probability below 0.5 is judged as an "error," and the comparison targets are clef and clef-flash (actual measurements from October 4, 2026).
Images were passed as original PNGs (1700×2200, approximately 200–300KB).
With Clef, the estimated token count exceeded the limit, so images were color-reduced to 32 colors; the image conditions differ from Clef's in this respect only.

Requests to the Decisions API take the following form (only one question is shown).
Comments are for explanatory purposes and are not included in the actual JSON sent.

{
  "model": "gpt-6-luna",
  "input": [{
    "role": "user",
    "content": [
      // Task premise (an OCR system read a Japanese invoice image and extracted keys and values; verify against the image) and extraction results
      {"type": "input_text", "text": "{\"task\": \"An OCR system read the attached Japanese invoice (請求書) image ...\", \"extracted\": {\"issuer_name\": \"株式会社さくらフーズ\", ...}}"},
      {"type": "input_image", "image_url": "data:image/png;base64,..."}
    ]
  }],
  "questions": [{
    "type": "predicate", // yes/no question; returns the probability of yes
    "name": "field__issuer_name",
    // Does the extracted value of issuer_name exactly match the string printed in the document image, character for character?
    "instructions": "Is the extracted value of \"issuer_name\" (\"株式会社さくらフーズ\") exactly what is printed in the document image, character for character?"
  }]
}

Sample invoice image
The invoice used in the verification. A stamp seal overlaps the issuer name and registration number

2. Results

2.1 Were Errors Detected?

Each model returns the probability that a value matches the image for each item.
Items where this probability is below 0.5 are judged as "errors," and counts were tallied from the following three perspectives.

  • Detection of injected errors: How many of the 36 injected errors were judged as "errors"

  • Detection of Haiku's errors: How many of the 19 misreads in Haiku's output were judged as "errors"

  • False positives: How many items that were actually correct were incorrectly judged as "errors"

Condition Injected error detection (out of 36) Haiku error detection (out of 19) False positives (out of 503)
Decisions API 34 10 16
clef 32 7 7
clef-flash 29 5 20

The Decisions API detected the most cases of both injected errors and Haiku's misreads among the three models.
On the other hand, false positives numbered 16, more than double Clef's 7.

Type Decisions API clef clef-flash
Injected errors: digits 12 / 12 10 / 12 9 / 12
Injected errors: similar characters 9 / 9 9 / 9 8 / 9
Injected errors: full/half-width 1 / 3 1 / 3 1 / 3
Injected errors: replacement 12 / 12 12 / 12 11 / 12
Haiku misreads 10 / 19 7 / 19 5 / 19
False positives 16 / 503 7 / 503 20 / 503

The Decisions API detected all digit, similar character, and replacement errors, but like Clef, caught only 1 out of 3 full-width/half-width swap cases.
The 27 cases where the Decisions API made incorrect judgments (9 missed Haiku misreads, 2 missed injected errors, 16 false positives) are as follows.

27 cases where the Decisions API made incorrect judgments (click to expand)
Error type Category Document Extraction Item Passed value Correct value Probability
Missed Haiku misread Invoice Haiku output issuer_tel TEL 045-312-7788 045-312-7788 0.94
Missed Haiku misread Invoice Haiku output reduced_rate_note ※印は軽減税率(8%)対象品目です ※印は軽減税率(8%)対象品目です 0.66
Missed Haiku misread Quote Haiku output amount_kanji 金壱百九拾六万五千四百八十円 金壱百九拾六万伍千四百八拾円也 0.53
Missed Haiku misread Quote Haiku output issuer_name 株式会社テクアーク 株式会社テクノアーク 0.88
Missed Haiku misread Quote Haiku output discount_name 出荷値引き 出精値引き 0.63
Missed Haiku misread Quote Haiku output remarks 本見積の有効期限内にご注文いただいた場合に限り、上記価格を適用いたします。納期は受注確定後の目安であり、部材の調達状況により前後する場合がございます。夜間・休日の作業をご希望の場合は別途お見積りとなります。 本見積の有効期限内にご発注いただいた場合に限り、上記価格を適用いたします。納期は受注確定後の目安であり、部材の調達状況により前後する場合がございます。夜間・休日の作業をご希望の場合は別途お見積りとなります。 0.67
Missed Haiku misread Receipt Haiku output item1_qty 3日 3日 0.98
Missed Haiku misread Receipt Haiku output item2_qty 3日 3日 1.00
Missed Haiku misread Receipt Haiku output item3_qty 1式 1式 0.98
Missed Full/half-width Receipt Error set A issue_date 2026年10月2日 2026年10月2日 0.97
Missed Full/half-width Receipt Error set B item4_qty 20食 20食 0.56
False positive Correct value Invoice Haiku output bill_to_name 株式会社みなと精機 株式会社みなと精機 0.34
False positive Correct value Invoice Haiku output subject 9月分食材・消耗品代 9月分食材・消耗品代 0.34
False positive Correct value Invoice Error set A subject 9月分食材・消耗品代 9月分食材・消耗品代 0.22
False positive Correct value Invoice Error set B bill_to_address 神奈川県横浜市中区山下町1丁目14番地 神奈川県横浜市中区山下町1丁目14番地 0.48
False positive Correct value Invoice Correct data bill_to_name 株式会社みなと精機 株式会社みなと精機 0.14
False positive Correct value Invoice Correct data subject 9月分食材・消耗品代 9月分食材・消耗品代 0.12
False positive Correct value Invoice Correct data issuer_registration_no T4010403012345 T4010403012345 0.28
False positive Correct value Invoice Correct data account_holder カ)サクラフーズ カ)サクラフーズ 0.13
False positive Correct value Quote Haiku output customer_name 株式会社ひかり物流 株式会社ひかり物流 0.48
False positive Correct value Receipt Error set A issuer_registration_no T9011001045678 T9011001045678 0.41
False positive Correct value Receipt Correct data item1_name 大会議室A利用料 大会議室A利用料 0.11
False positive Correct value Receipt Correct data item1_qty 3日 3日 0.04
False positive Correct value Receipt Correct data item2_qty 3日 3日 0.27
False positive Correct value Receipt Correct data recipient あおば企画株式会社 あおば企画株式会社 0.45
False positive Correct value Receipt Correct data proviso 会議室利用料(10月分)として 会議室利用料(10月分)として 0.34
False positive Correct value Receipt Correct data issuer_building 桜丘第一ビル5階 桜丘第一ビル5階 0.41

9 of the Decisions API's 16 false positives occurred in the correct data version.
In particular, for the receipt quantity items where Haiku misread full-width as half-width, the model assigned 0.04 to the correct 3日 and 0.98 to Haiku's misread 3日, making the same inverted judgment as Clef regarding full-width and half-width.

Probability of match for all 36 injected errors (click to expand)

The probability represents "the probability that the extracted value is correct as shown in the image," so lower values indicate stronger suspicion of an error.
Values at or above 0.5 (missed) are bolded.

Document Extraction Item Correct value Injected value Pattern Decisions API clef clef-flash
Invoice A invoice_no INV-202609-0417 INV-202609-0411 Digit 0.00 0.12 0.17
Invoice A item3_amount 51,750 51,790 Digit 0.00 0.10 0.68
Invoice A item1_name 新潟県産コシヒカリ5kg 新潟県産コシヒ力リ5kg Similar char 0.01 0.06 0.21
Invoice A bank_branch 横浜西口支店 横浜西ロ支店 Similar char 0.05 0.12 0.46
Invoice A bill_to_name 株式会社みなと精機 株式会社みなと製作所 Replacement 0.02 0.20 0.19
Invoice A account_type 普通 当座 Replacement 0.02 0.31 0.15
Invoice B tax_8 10,073 10,078 Digit 0.00 0.30 0.62
Invoice B account_number 4072519 4072579 Digit 0.17 0.75 0.03
Invoice B item2_name 有機緑茶ティーバッグ50包 有機緑茶ティ一バッグ50包 Similar char 0.01 0.26 0.17
Invoice B item3_name 業務用キッチンペーパー 業務用キッチソペーパー Similar char 0.00 0.10 0.15
Invoice B subject 9月分食材・消耗品代 9月分飲料・消耗品代 Replacement 0.00 0.11 0.14
Invoice B item5_tax_rate 10% 8% Replacement 0.00 0.06 0.10
Quote A item1_unit_price 128,000 126,000 Digit 0.00 0.03 0.06
Quote A quote_no MQ-2026-0918 MQ-2026-0916 Digit 0.00 0.05 0.02
Quote A item1_name ノートパソコン14型 ノートパンコン14型 Similar char 0.00 0.10 0.14
Quote A item3_name USB-Cドッキングステーション USB-Cドッキングステーツョン Similar char 0.00 0.17 0.07
Quote A delivery_date 受注後約3週間 受注後約2か月 Replacement 0.00 0.03 0.02
Quote A customer_name 株式会社ひかり物流 株式会社ひかり運輸 Replacement 0.08 0.12 0.30
Quote B subtotal 1,786,800 1,788,800 Digit 0.00 0.56 0.66
Quote B item7_qty 20 25 Digit 0.00 0.03 0.08
Quote B item4_name ワイヤレスキーボード ワイヤレスキ一ボード Similar char 0.00 0.18 0.25
Quote B item2_code MN-2407 MN-2407 Full/half-width 0.00 0.05 0.41
Quote B subject 本社オフィスPC入替一式 支社オフィスPC入替一式 Replacement 0.00 0.06 0.23
Quote B item8_unit 式 台 Replacement 0.00 0.15 0.14
Receipt A amount ¥165,000- ¥185,000- Digit 0.00 0.02 0.02
Receipt A issuer_tel 03-5468-1234 03-5468-1284 Digit 0.00 0.08 0.08
Receipt A item2_name プロジェクター貸出 プロジェクタ一貸出 Similar char 0.01 0.29 0.09
Receipt A issue_date 2026年10月2日 2026年10月2日 Full/half-width 0.97 0.71 0.96
Receipt A payment_method 銀行振込 現金 Replacement 0.00 0.02 0.01
Receipt A issuer_person 佐々木美咲 佐藤美咲 Replacement 0.01 0.43 0.83
Receipt B receipt_no 004127 004121 Digit 0.00 0.03 0.07
Receipt B item1_amount 90,000 60,000 Digit 0.00 0.02 0.05
Receipt B issuer_name 株式会社みどり会議室サービス 株式会社みどり会議室サーピス Similar char 0.00 0.28 0.54
Receipt B item4_qty 20食 20食 Full/half-width 0.56 0.54 0.87
Receipt B item3_name 音響設備一式 照明設備一式 Replacement 0.01 0.05 0.09
Receipt B item1_name 大会議室A利用料 中会議室A利用料 Replacement 0.00 0.05 0.06
Probabilities for Haiku's 19 misreads (click to expand)

Values at or above 0.5 (missed) are bolded.

Document Item Haiku's value Correct value Decisions API clef clef-flash
Invoice issuer_name 株式会社さくら精機 株式会社さくらフーズ 0.00 0.30 0.06
Invoice issuer_registration_no T4010-0301-2345 T4010403012345 0.09 0.73 0.76
Invoice issuer_address 神奈川県横浜市西区高島2丁目8番号号 神奈川県横浜市西区高島2丁目8番5号 0.00 0.05 0.07
Invoice issuer_tel TEL 045-312-7788 045-312-7788 0.94 0.90 0.93
Invoice reduced_rate_note ※印は軽減税率(8%)対象品目です ※印は軽減税率(8%)対象品目です 0.66 0.78 0.91
Invoice account_holder サクラフーズ カ)サクラフーズ 0.26 0.69 0.69
Quote amount_kanji 金壱百九拾六万五千四百八十円 金壱百九拾六万伍千四百八拾円也 0.53 0.19 0.31
Quote issuer_name 株式会社テクアーク 株式会社テクノアーク 0.88 0.76 0.83
Quote discount_name 出荷値引き 出精値引き 0.63 0.87 0.64
Quote remarks 本見積の有効期限内にご注文いただいた場合に限り、上記価格を適用いたします。納期は受注確定後の目安であり、部材の調達状況により前後する場合がございます。夜間・休日の作業をご希望の場合は別途お見積りとなります。 本見積の有効期限内にご発注いただいた場合に限り、上記価格を適用いたします。納期は受注確定後の目安であり、部材の調達状況により前後する場合がございます。夜間・休日の作業をご希望の場合は別途お見積りとなります。 0.67 0.63 0.91
Receipt item1_qty 3日 3日 0.98 0.89 0.88
Receipt item2_qty 3日 3日 1.00 0.89 0.88
Receipt item3_qty 1式 1式 0.98 0.87 0.84
Receipt item4_name ケータリング食食 ケータリング昼食 0.00 0.10 0.19
Receipt recipient あおは企画株式会社 あおば企画株式会社 0.01 0.91 0.64
Receipt stamp_duty 200円 200円 0.08 0.61 0.89
Receipt proviso 会議室利用料(10月分)として上記正に領収いたしました。 会議室利用料(10月分)として 0.34 0.12 0.77
Receipt issuer_address 東京都渋谷区桜丘町1丁目2番3号 東京都渋谷区桜丘町1丁目2番3号 0.26 0.44 0.73
Receipt issuer_building 桜丘サービル5階 桜丘第一ビル5階 0.00 0.32 0.28

2.2 Distribution of Probabilities

The spread of probabilities assigned to injected errors (36 cases), Haiku misreads (19 cases), and correct values (503 cases) is as follows.
The bottom 5% and top 5% are the values at the 5th and 95th percentile positions when probabilities are sorted in ascending order.

Condition Target Min Bottom 5% Median Top 5% Max
Decisions API Injected errors 0.0 0.0 0.0 0.268 0.97
Decisions API Haiku errors 0.0 0.0 0.34 0.982 1.0
Decisions API Correct values 0.04 0.661 0.99 1.0 1.0
clef Injected errors 0.022 0.025 0.107 0.594 0.748
clef Haiku errors 0.052 0.097 0.691 0.898 0.907
clef Correct values 0.283 0.672 0.874 0.957 0.979
clef-flash Injected errors 0.013 0.016 0.145 0.838 0.964
clef-flash Haiku errors 0.063 0.068 0.764 0.916 0.928
clef-flash Correct values 0.302 0.541 0.857 0.959 0.977

With one item per point, arranged by model as follows (points are scattered vertically to avoid overlap).
Note that the following three graphs were created by Claude based on the verification result data.

Distribution of probabilities output by the Decisions API
Decisions API: Injected errors cluster near 0.0, but Haiku misreads split toward both ends at 0.0 and 1.0

Distribution of probabilities output by clef
clef: Injected errors are scattered in the 0.0–0.3 range, and most Haiku misreads exceed 0.5

Distribution of probabilities output by clef-flash
clef-flash: Distribution pattern similar to clef, but some correct values and some injected errors overlap across 0.5

With the Decisions API, the median probability for injected errors was 0.0 and for correct values was 0.99, with probabilities skewed toward the extremes just as in English.
The median for Haiku misreads was also 0.34, lower than Clef (0.691), but some missed items were assigned probabilities above 0.9.

Results when varying the threshold are as follows.

Threshold Decisions API: injected Decisions API: Haiku Decisions API: FP clef: injected clef: Haiku clef: FP clef-flash: injected clef-flash: Haiku clef-flash: FP
0.1 33/36 7/19 1/503 16/36 1/19 0/503 14/36 2/19 0/503
0.3 34/36 9/19 8/503 30/36 5/19 1/503 27/36 4/19 0/503
0.5 34/36 10/19 16/503 32/36 7/19 7/503 29/36 5/19 20/503
0.7 35/36 14/19 29/503 34/36 10/19 30/503 33/36 8/19 105/503
0.8 35/36 14/19 41/503 36/36 13/19 85/503 33/36 11/19 190/503

With the Decisions API at a threshold of 0.1, it still caught 33 injected errors with only 1 false positive.

2.3 Time and Cost

Requests were made once per document, sending one image and all questions for that document's items (30–63 questions) together.
With 3 documents × 4 extraction versions, 12 calls were made per model.
Time is measured from actual REST API calls from within Japan, and costs are converted at 150 yen per dollar.

Condition Time per call (median) Total time for 12 calls Input tokens for 12 calls Total cost for 12 calls
Decisions API 0.925 seconds (925ms) 11.92 seconds 157,876 $0.0158 (approx. 2.37 yen)
clef 1.913 seconds (1,913ms) 21.60 seconds 86,426 $0.0207 (approx. 3.11 yen)
clef-flash 1.010 seconds (1,010ms) 11.30 seconds 86,426 $0.0078 (approx. 1.17 yen)

Processing time for the Decisions API was about half that of clef, and cost was about three-quarters of clef's.
Input tokens were about 1.8 times that of Clef, but since the image conditions differ (the Decisions API used non-color-reduced images), a direct comparison is not straightforward.

3. Summary

Even with Japanese documents, the Decisions API detected 34 out of 36 injected errors and 10 out of 19 Haiku misreads, surpassing both Clef models on both counts.
Since avoiding missed errors is paramount when checking OCR results, I found it to be relatively strong even with Japanese documents.
Also, unlike Clef, there is no need to color-reduce images, and being able to pass images at their original quality is likely advantageous for reading fine characters.

On the other hand, false positives numbered 16, more than double Clef's 7, so it falls short of Clef in terms of false positive rate.
Full-width/half-width differences were also easy to miss, just as with Clef, and about half of Haiku's actual misreads passed through with a probability of 0.5 or higher.
The distinction between full-width and half-width is a concept that doesn't exist in English-speaking contexts, so models developed by American companies — not just the Decisions API or Clef — may tend to struggle with this difference regardless.


AI白書2026 配布中

クラスメソッドが独自に行なったAI診断調査をもとに、企業のAI活用の現在地を調査レポートとしてまとめました。企業規模別の活用度傾向に加え、規模を超えてAI活用を進める企業に共通する取り組みまで、自社の現在地を捉えるためのヒントにぜひ。

AI白書2026

無料でダウンロードする

Share this article

DevelopersIO 2026