I tried having Cloudflare's Clef classify document types
This page has been translated by machine translation. View original
Hello, I'm Keema.
When sorting received documents like invoices, quotations, and purchase orders, decision models that can only accept text (such as TypeSafe's Jev) required OCR to extract text first.
Clef, released on Cloudflare Workers AI on October 1, 2026, can accept images as input, so it might be possible to classify documents using only images without going through OCR.
So, I passed 28 fictional Japanese and English documents to Clef and had it classify them into 7 categories, then checked the accuracy and confidence scores.
1. Test Conditions
| Item | Details |
|---|---|
| Model | clef (27B), clef-flash (9B) |
| Input | 1 document image (PNG, 1240×1754). No document text is passed |
| Documents | 28 fictional documents. No scan skew, fading, or handwriting |
| Question | 1 choice question. 8 options: 7 categories plus other |
| How called | REST API, 1 image per request, once per model |
choice is a question type that returns the selected option, the probability for each option, and a confidence score (Changelog: Introducing Clef | Cloudflare Docs).
Probability and confidence are different values.
-
Probability: the likelihood that the document is of that type. The 8 options sum to 1, and the highest-scoring type is selected as the answer.
-
Confidence: how decisively that answer was chosen (0–1).
For example, in this test, one purchase order had invoice at 0.435 and quotation at 0.3625 nearly tied, the answer came out as invoice, but confidence dropped to 0.2384.
The calculation method for confidence is not documented by Cloudflare, so this article uses the returned values as-is.
Whether to process an answer automatically or send it to a human for review is decided based on this confidence score, which reflects the degree of hesitation.
The request format is as follows (only 2 of the 8 criteria are shown).
{
"model": "clef",
"images": ["data:image/png;base64,..."],
// state: premise for the decision (the attached image is one page of a business document; determine what kind of document it is)
"state": "The attached image is one page of a business document. Decide what kind of document it is.",
"questions": {
"doc_type": {
"type": "choice", // select one option
// instructions: instruction for classification (which type of business document is shown in the image)
"instructions": "Which kind of business document is shown in the image?",
"criteria": {
// invoice: a document requesting payment for goods or services already provided (with amount due and payment deadline)
"invoice": "Invoice / bill: the seller requests payment for goods or services already provided, with an amount due and a payment deadline.",
// quotation: a document proposing prices before an order is placed (usually with a validity period)
"quotation": "Quotation / estimate: the seller proposes prices before an order is placed, usually with a validity period."
}
}
}
}
2. Data
The 7 categories are invoice, quotation, receipt, delivery note, purchase order, contract, and resume.
For each category, 2 documents were prepared — one in Japanese only and one in English only — with the second document including misleading elements such as the following.
-
No title, a small title, or an alternative name (Estimate, Packing Slip, 発注書, 職務経歴書)
-
Reference to another document type in the body text (quotation number in an invoice body, reference to purchase order and invoice in a quotation's notes)
-
A resume containing business document terminology (work experience handling invoices and purchase orders)

An example of an added misleading element. No title, and the body references a quotation number.
3. Results
The following table shows the characteristics of all 28 documents and each model's answer, probability, and confidence score (incorrect answers are shown in bold).
Clef's choice type returns two values in addition to the selected answer: "probability" and "confidence." "Probability" is the predicted probability assigned to the selected answer, where all options sum to 1.0. "Confidence," on the other hand, is a value from 0 to 1 indicating how confident the model itself is in that judgment.
| Correct | Language | Document characteristics | clef answer | Probability | Confidence | clef-flash answer | Probability | Confidence |
|---|---|---|---|---|---|---|---|---|
| Invoice | Japanese | Standard (title: 請求書) | Invoice | 0.974 | 0.9416 | Invoice | 0.9497 | 0.8884 |
| Invoice | Japanese | No title. Body references quotation number | Invoice | 0.9769 | 0.948 | Invoice | 0.9751 | 0.9439 |
| Invoice | English | Standard (title: INVOICE) | Invoice | 0.981 | 0.957 | Invoice | 0.9611 | 0.9131 |
| Invoice | English | No title. Body references an accepted quotation number | Invoice | 0.9832 | 0.9619 | Invoice | 0.9727 | 0.9386 |
| Quotation | Japanese | Standard (title: 御見積書) | Quotation | 0.9671 | 0.9265 | Quotation | 0.9599 | 0.9105 |
| Quotation | Japanese | Title is small 「お見積り」 in top right only. References purchase order and invoice | Quotation | 0.9405 | 0.8698 | Invoice | 0.7331 | 0.5175 |
| Quotation | English | Standard (title: QUOTATION) | Quotation | 0.9782 | 0.9507 | Quotation | 0.974 | 0.9416 |
| Quotation | English | Title is the alternative name "Estimate." States "This is not an invoice" and references purchase order | Quotation | 0.975 | 0.9437 | Quotation | 0.9571 | 0.9044 |
| Receipt | Japanese | Standard (title: 領収書, description field, revenue stamp column) | Receipt | 0.9608 | 0.9126 | Receipt | 0.9046 | 0.7957 |
| Receipt | Japanese | No title, POS receipt (amount tendered, change) | Receipt | 0.9822 | 0.9597 | Receipt | 0.9603 | 0.9115 |
| Receipt | English | Standard (title: RECEIPT, PAID IN FULL) | Receipt | 0.9871 | 0.9707 | Receipt | 0.9775 | 0.9491 |
| Receipt | English | Title is the alternative name "Payment Confirmation." Table shows invoice number and billed amount | Receipt | 0.9862 | 0.9687 | Receipt | 0.9557 | 0.9017 |
| Delivery Note | Japanese | Standard (title: 納品書, receipt stamp column) | Delivery Note | 0.9288 | 0.8453 | Delivery Note | 0.9441 | 0.8766 |
| Delivery Note | Japanese | Small title 「納品書(控)」. References purchase order number, no amount column | Delivery Note | 0.9746 | 0.9429 | Delivery Note | 0.9706 | 0.9338 |
| Delivery Note | English | Standard (title: DELIVERY NOTE, recipient signature field) | Delivery Note | 0.9824 | 0.9601 | Delivery Note | 0.9733 | 0.9399 |
| Delivery Note | English | Title is the alternative name "Packing Slip." References customer PO number, no amount column | Delivery Note | 0.9797 | 0.9542 | Delivery Note | 0.9443 | 0.8768 |
| Purchase Order | Japanese | Standard (title: 注文書) | Purchase Order | 0.9599 | 0.9106 | Purchase Order | 0.9685 | 0.9293 |
| Purchase Order | Japanese | Title: 発注書. References quotation, delivery note, and invoice | Invoice | 0.435 | 0.2384 | Purchase Order | 0.975 | 0.9436 |
| Purchase Order | English | Standard (title: PURCHASE ORDER) | Purchase Order | 0.9817 | 0.9586 | Purchase Order | 0.9761 | 0.9461 |
| Purchase Order | English | No title. Only "Please supply the following items" and an approver field | Purchase Order | 0.9781 | 0.9506 | Purchase Order | 0.9768 | 0.9477 |
| Contract | Japanese | Standard (non-disclosure agreement) | Contract | 0.9822 | 0.9597 | Contract | 0.9628 | 0.9167 |
| Contract | Japanese | Service agreement. Compensation clause references quotation and invoice | Contract | 0.9798 | 0.9543 | Contract | 0.9555 | 0.9009 |
| Contract | English | Standard (NON-DISCLOSURE AGREEMENT) | Contract | 0.9852 | 0.9666 | Contract | 0.9548 | 0.8993 |
| Contract | English | Master Services Agreement. Section headings include Purchase Orders and Invoices and Payment | Contract | 0.9858 | 0.9679 | Contract | 0.9728 | 0.9388 |
| Resume | Japanese | Standard (close to JIS format 「履歴書」, photo field) | Resume | 0.9566 | 0.9036 | Resume | 0.9647 | 0.9211 |
| Resume | Japanese | 「職務経歴書」 instead of 「履歴書」. Body contains terms like quotation, purchase order, and invoice | Resume | 0.9542 | 0.8985 | Resume | 0.9488 | 0.8865 |
| Resume | English | Small "RESUME." Body contains terms invoices, purchase orders, receipts | Resume | 0.9774 | 0.949 | Resume | 0.9714 | 0.9358 |
| Resume | English | No title. Researcher CV with only name and headings (paper title contains "invoices") | Resume | 0.958 | 0.9065 | Resume | 0.9624 | 0.916 |
3.1 Accuracy
| Condition | clef | clef-flash |
|---|---|---|
| Overall | 27 / 28 | 27 / 28 |
| Japanese | 13 / 14 | 13 / 14 |
| English | 14 / 14 | 14 / 14 |
| Standard documents (1st of each pair) | 14 / 14 | 14 / 14 |
| Documents with misleading elements (2nd of each pair) | 13 / 14 | 13 / 14 |
Both clef and clef-flash correctly classified 27 out of 28 documents, getting all English documents and all standard documents right.
The one document each missed was a misleading Japanese document, and the two models did not miss the same document.
3.2 Misclassified Documents
| Model | Document (correct) | Answer and its probability | Probability assigned to correct type | Confidence |
|---|---|---|---|---|
| clef | Purchase order form (Purchase Order) | Invoice 0.435 | Purchase Order 0.066 | 0.2384 |
| clef-flash | Quotation (Quotation) | Invoice 0.7331 | Quotation 0.199 | 0.5175 |

The purchase order that clef answered as invoice. The body references three other document types: quotation, delivery note, and invoice.

The quotation that clef-flash answered as invoice. The title is only a small 「お見積り」 in the top right.
clef was torn between invoice (probability 0.435) and quotation (probability 0.3625) for the purchase order, likely pulled by the other document names mentioned in the body text.
clef-flash misclassified the quotation with the small title as invoice, but clef correctly identified the same image as a quotation.
In both misclassifications, the confidence was clearly lower than for correctly classified documents, and both were caught reliably by the threshold described in the next section.
3.3 Confidence Score Distribution
| Item | clef | clef-flash |
|---|---|---|
| Confidence of the 27 correct documents (min / median) | 0.8453 / 0.9506 | 0.7957 / 0.9167 |
| Confidence of the 1 incorrect document | 0.2384 | 0.5175 |
| Documents sent to human review at threshold < 0.8 | 1 | 2 |
| Errors remaining above threshold ≥ 0.8 | 0 | 0 |
Across all 56 responses, the confidence score was lower than the probability for every answer.
Setting the threshold at 0.8, clef sends only the 1 incorrect document to a human, while clef-flash sends the 1 incorrect document plus 1 correct document to a human, leaving no errors in the automatically processed documents.
With a workflow where only documents with confidence below 0.8 are reviewed by a human, all misclassifications in these 28 documents would be caught.
3.4 Time and Cost
Time is the measured round-trip time of REST API calls from Japan, and the exchange rate used is 150 yen per dollar.
| Model | Round-trip time (median) | Input tokens (per image) | Cost (28 images) |
|---|---|---|---|
| clef | 0.976 seconds (976ms) | 1,367 | $0.0092 (approx. 1.38 yen) |
| clef-flash | 0.3535 seconds (354ms) | 1,367 | $0.0034 (approx. 0.52 yen) |
clef-flash responded in about one-third the time of clef, and the cost was also reduced to about 40%.
4. Summary
Based on testing with these 28 documents, Clef was able to classify document types using images alone without OCR, and the misclassified documents could be identified by their low confidence scores.
In terms of which to use: clef is better suited when you want to minimize the number of documents sent for manual review, while clef-flash is better when you want to reduce per-document processing time.

