I tried having Cloudflare's Clef classify document types

I tried having Cloudflare's Clef classify document types

We tested if Cloudflare Workers AI's new multimodal model "Clef" can auto-classify invoices and quotes from images alone, without OCR, sorting 28 fictional documents into 7 types.
2026.10.04

This page has been translated by machine translation. View original

Hello, I'm Keema.

When sorting received documents like invoices, quotations, and purchase orders, decision models that can only accept text (such as TypeSafe's Jev) required OCR to extract text first.
Clef, released on Cloudflare Workers AI on October 1, 2026, can accept images as input, so it might be possible to classify documents using only images without going through OCR.

So, I passed 28 fictional Japanese and English documents to Clef and had it classify them into 7 categories, then checked the accuracy and confidence scores.

1. Test Conditions

Item Details
Model clef (27B), clef-flash (9B)
Input 1 document image (PNG, 1240×1754). No document text is passed
Documents 28 fictional documents. No scan skew, fading, or handwriting
Question 1 choice question. 8 options: 7 categories plus other
How called REST API, 1 image per request, once per model

choice is a question type that returns the selected option, the probability for each option, and a confidence score (Changelog: Introducing Clef | Cloudflare Docs).

Probability and confidence are different values.

  • Probability: the likelihood that the document is of that type. The 8 options sum to 1, and the highest-scoring type is selected as the answer.

  • Confidence: how decisively that answer was chosen (0–1).

For example, in this test, one purchase order had invoice at 0.435 and quotation at 0.3625 nearly tied, the answer came out as invoice, but confidence dropped to 0.2384.
The calculation method for confidence is not documented by Cloudflare, so this article uses the returned values as-is.
Whether to process an answer automatically or send it to a human for review is decided based on this confidence score, which reflects the degree of hesitation.

The request format is as follows (only 2 of the 8 criteria are shown).

{
  "model": "clef",
  "images": ["data:image/png;base64,..."],
  // state: premise for the decision (the attached image is one page of a business document; determine what kind of document it is)
  "state": "The attached image is one page of a business document. Decide what kind of document it is.",
  "questions": {
    "doc_type": {
      "type": "choice", // select one option
      // instructions: instruction for classification (which type of business document is shown in the image)
      "instructions": "Which kind of business document is shown in the image?",
      "criteria": {
        // invoice: a document requesting payment for goods or services already provided (with amount due and payment deadline)
        "invoice": "Invoice / bill: the seller requests payment for goods or services already provided, with an amount due and a payment deadline.",
        // quotation: a document proposing prices before an order is placed (usually with a validity period)
        "quotation": "Quotation / estimate: the seller proposes prices before an order is placed, usually with a validity period."
      }
    }
  }
}

2. Data

The 7 categories are invoice, quotation, receipt, delivery note, purchase order, contract, and resume.
For each category, 2 documents were prepared — one in Japanese only and one in English only — with the second document including misleading elements such as the following.

  • No title, a small title, or an alternative name (Estimate, Packing Slip, 発注書, 職務経歴書)

  • Reference to another document type in the body text (quotation number in an invoice body, reference to purchase order and invoice in a quotation's notes)

  • A resume containing business document terminology (work experience handling invoices and purchase orders)

English invoice without a title
An example of an added misleading element. No title, and the body references a quotation number.

3. Results

The following table shows the characteristics of all 28 documents and each model's answer, probability, and confidence score (incorrect answers are shown in bold).

Clef's choice type returns two values in addition to the selected answer: "probability" and "confidence." "Probability" is the predicted probability assigned to the selected answer, where all options sum to 1.0. "Confidence," on the other hand, is a value from 0 to 1 indicating how confident the model itself is in that judgment.

Correct Language Document characteristics clef answer Probability Confidence clef-flash answer Probability Confidence
Invoice Japanese Standard (title: 請求書) Invoice 0.974 0.9416 Invoice 0.9497 0.8884
Invoice Japanese No title. Body references quotation number Invoice 0.9769 0.948 Invoice 0.9751 0.9439
Invoice English Standard (title: INVOICE) Invoice 0.981 0.957 Invoice 0.9611 0.9131
Invoice English No title. Body references an accepted quotation number Invoice 0.9832 0.9619 Invoice 0.9727 0.9386
Quotation Japanese Standard (title: 御見積書) Quotation 0.9671 0.9265 Quotation 0.9599 0.9105
Quotation Japanese Title is small 「お見積り」 in top right only. References purchase order and invoice Quotation 0.9405 0.8698 Invoice 0.7331 0.5175
Quotation English Standard (title: QUOTATION) Quotation 0.9782 0.9507 Quotation 0.974 0.9416
Quotation English Title is the alternative name "Estimate." States "This is not an invoice" and references purchase order Quotation 0.975 0.9437 Quotation 0.9571 0.9044
Receipt Japanese Standard (title: 領収書, description field, revenue stamp column) Receipt 0.9608 0.9126 Receipt 0.9046 0.7957
Receipt Japanese No title, POS receipt (amount tendered, change) Receipt 0.9822 0.9597 Receipt 0.9603 0.9115
Receipt English Standard (title: RECEIPT, PAID IN FULL) Receipt 0.9871 0.9707 Receipt 0.9775 0.9491
Receipt English Title is the alternative name "Payment Confirmation." Table shows invoice number and billed amount Receipt 0.9862 0.9687 Receipt 0.9557 0.9017
Delivery Note Japanese Standard (title: 納品書, receipt stamp column) Delivery Note 0.9288 0.8453 Delivery Note 0.9441 0.8766
Delivery Note Japanese Small title 「納品書(控)」. References purchase order number, no amount column Delivery Note 0.9746 0.9429 Delivery Note 0.9706 0.9338
Delivery Note English Standard (title: DELIVERY NOTE, recipient signature field) Delivery Note 0.9824 0.9601 Delivery Note 0.9733 0.9399
Delivery Note English Title is the alternative name "Packing Slip." References customer PO number, no amount column Delivery Note 0.9797 0.9542 Delivery Note 0.9443 0.8768
Purchase Order Japanese Standard (title: 注文書) Purchase Order 0.9599 0.9106 Purchase Order 0.9685 0.9293
Purchase Order Japanese Title: 発注書. References quotation, delivery note, and invoice Invoice 0.435 0.2384 Purchase Order 0.975 0.9436
Purchase Order English Standard (title: PURCHASE ORDER) Purchase Order 0.9817 0.9586 Purchase Order 0.9761 0.9461
Purchase Order English No title. Only "Please supply the following items" and an approver field Purchase Order 0.9781 0.9506 Purchase Order 0.9768 0.9477
Contract Japanese Standard (non-disclosure agreement) Contract 0.9822 0.9597 Contract 0.9628 0.9167
Contract Japanese Service agreement. Compensation clause references quotation and invoice Contract 0.9798 0.9543 Contract 0.9555 0.9009
Contract English Standard (NON-DISCLOSURE AGREEMENT) Contract 0.9852 0.9666 Contract 0.9548 0.8993
Contract English Master Services Agreement. Section headings include Purchase Orders and Invoices and Payment Contract 0.9858 0.9679 Contract 0.9728 0.9388
Resume Japanese Standard (close to JIS format 「履歴書」, photo field) Resume 0.9566 0.9036 Resume 0.9647 0.9211
Resume Japanese 「職務経歴書」 instead of 「履歴書」. Body contains terms like quotation, purchase order, and invoice Resume 0.9542 0.8985 Resume 0.9488 0.8865
Resume English Small "RESUME." Body contains terms invoices, purchase orders, receipts Resume 0.9774 0.949 Resume 0.9714 0.9358
Resume English No title. Researcher CV with only name and headings (paper title contains "invoices") Resume 0.958 0.9065 Resume 0.9624 0.916

3.1 Accuracy

Condition clef clef-flash
Overall 27 / 28 27 / 28
Japanese 13 / 14 13 / 14
English 14 / 14 14 / 14
Standard documents (1st of each pair) 14 / 14 14 / 14
Documents with misleading elements (2nd of each pair) 13 / 14 13 / 14

Both clef and clef-flash correctly classified 27 out of 28 documents, getting all English documents and all standard documents right.
The one document each missed was a misleading Japanese document, and the two models did not miss the same document.

3.2 Misclassified Documents

Model Document (correct) Answer and its probability Probability assigned to correct type Confidence
clef Purchase order form (Purchase Order) Invoice 0.435 Purchase Order 0.066 0.2384
clef-flash Quotation (Quotation) Invoice 0.7331 Quotation 0.199 0.5175

Purchase order that clef answered as invoice
The purchase order that clef answered as invoice. The body references three other document types: quotation, delivery note, and invoice.

Quotation that clef-flash answered as invoice
The quotation that clef-flash answered as invoice. The title is only a small 「お見積り」 in the top right.

clef was torn between invoice (probability 0.435) and quotation (probability 0.3625) for the purchase order, likely pulled by the other document names mentioned in the body text.
clef-flash misclassified the quotation with the small title as invoice, but clef correctly identified the same image as a quotation.

In both misclassifications, the confidence was clearly lower than for correctly classified documents, and both were caught reliably by the threshold described in the next section.

3.3 Confidence Score Distribution

Item clef clef-flash
Confidence of the 27 correct documents (min / median) 0.8453 / 0.9506 0.7957 / 0.9167
Confidence of the 1 incorrect document 0.2384 0.5175
Documents sent to human review at threshold < 0.8 1 2
Errors remaining above threshold ≥ 0.8 0 0

Across all 56 responses, the confidence score was lower than the probability for every answer.
Setting the threshold at 0.8, clef sends only the 1 incorrect document to a human, while clef-flash sends the 1 incorrect document plus 1 correct document to a human, leaving no errors in the automatically processed documents.
With a workflow where only documents with confidence below 0.8 are reviewed by a human, all misclassifications in these 28 documents would be caught.

3.4 Time and Cost

Time is the measured round-trip time of REST API calls from Japan, and the exchange rate used is 150 yen per dollar.

Model Round-trip time (median) Input tokens (per image) Cost (28 images)
clef 0.976 seconds (976ms) 1,367 $0.0092 (approx. 1.38 yen)
clef-flash 0.3535 seconds (354ms) 1,367 $0.0034 (approx. 0.52 yen)

clef-flash responded in about one-third the time of clef, and the cost was also reduced to about 40%.

4. Summary

Based on testing with these 28 documents, Clef was able to classify document types using images alone without OCR, and the misclassified documents could be identified by their low confidence scores.

In terms of which to use: clef is better suited when you want to minimize the number of documents sent for manual review, while clef-flash is better when you want to reduce per-document processing time.

References


AI白書2026 配布中

クラスメソッドが独自に行なったAI診断調査をもとに、企業のAI活用の現在地を調査レポートとしてまとめました。企業規模別の活用度傾向に加え、規模を超えてAI活用を進める企業に共通する取り組みまで、自社の現在地を捉えるためのヒントにぜひ。

AI白書2026

無料でダウンロードする

Share this article

DevelopersIO 2026