A small decision model sorts documents from page one
A 9-billion-parameter vision model read only the first page of each PDF and named the kind of document 58 times out of 60. It also spotted the scanned pages, ran in under a second, and gave the same answer every time.
If you look after a large pile of PDFs, the first question is rarely "is this accessible?" It is "what is this?" A form needs different work from a flyer, and a scanned page needs different work from both. Sorting that pile by hand does not scale. This note reports a test of a small decision model that does the sorting from the first page alone, and does it well.
The model is clef-flash, a 9-billion-parameter vision model served through Ollama. It is one of the smallest openly released models that reads an image and returns a decision rather than text, and the smaller of the two in its family. Given one rendered page and a typed question, it named the document kind correctly for 58 of 60 PDFs, or 96.7 percent. It flagged scanned pages correctly for 67 of 68. Each answer took under a second on a laptop, and a second run returned the same answers.
What a decision model is
Most language models write. A decision model answers. You give it a question with a fixed set of options and it returns a probability for each option instead of a paragraph. That changes how you can use it:
- The output is a number, so you can set a threshold, route on it, and count it.
- There is no prose to parse and no format to drift.
- A low probability is useful information. It tells you the model is unsure, which is the moment to send a document to a person.
- Runs are repeatable at temperature zero.
For document triage those properties matter more than raw size. The 27B model in the same family was tried first and dropped. It was slower and heavier, and the 9B model was accurate enough.
The task
The documents were 75 public PDFs collected from one organization's websites. Each PDF was rendered at page one only, at 150 dots per inch with the long edge at 1,200 pixels. The filename was withheld so the model could not read the answer off it.
The model answered one multiple-choice question with nine options. Each option came with a one-line definition, which is the whole "prompt" for that question. The chart shows how the 60 scored documents fell across the kinds and how many the model named correctly. The lighter bar is the number of scored documents of that kind; the darker bar is the number the model got right.
| Kind | What it means | Model chose it | Correct, of scored |
|---|---|---|---|
| General document | A document that fits no more specific kind | 9 | 6 of 6 |
| Policy or procedure | A governing policy, rule, or step-by-step procedure | 12 | 10 of 10 |
| Form or application | A fillable form, application, or template to complete and submit | 15 | 13 of 14 |
| Handbook, manual, or guide | A handbook, manual, or how-to guide | 13 | 11 of 12 |
| Report, plan, or proposal | A report, plan, or proposal | 4 | 4 of 4 |
| Map, flyer, or brochure | A map, flyer, or promotional brochure | 9 | 8 of 8 |
| Meeting agenda or minutes | A meeting agenda or minutes | 1 | 1 of 1 |
| Research or grant | A research or grant document | 5 | 5 of 5 |
| Other | None of the above | 0 | 0 of 0 |
"Model chose it" counts all 68 answered documents, including the eight that were later excluded or had no label. The model never used "other", and it was most sure of itself on forms and flyers, where a page has a shape you can see at a glance.
It also answered three yes or no questions: is this student-facing, is it technical, and is it a scanned page.
What the question looks like
A decision model takes a state and a schema of typed questions. Here the state is one page image, and the schema has one choice question and three yes or no questions. This is the request, with the image left out:
{
"model": "clef-flash",
"state": [{"type": "image", "image": "<page 1, rendered at 150 dpi>"}],
"questions": {
"document_kind": {
"type": "choice",
"instructions": "You are looking at page 1 of a document published on an organization's website. Classify the document by its kind, what sort of document it is as a genre, not its topic. Judge only from the page shown.",
"criteria": {
"general_document": "A general document that fits no more specific kind.",
"policy_or_procedure": "A governing policy, rule, or step-by-step procedure.",
"form_or_application": "A fillable form, application, or template to complete and submit.",
"handbook_manual_or_guide": "A handbook, manual, or how-to guide.",
"report_plan_or_proposal": "A report, plan, or proposal.",
"map_flyer_or_brochure": "A map, flyer, or promotional brochure.",
"meeting_agenda_or_minutes": "A meeting agenda or minutes.",
"research_or_grant": "A research or grant document.",
"other": "None of the above."
}
},
"is_scanned": {
"type": "noul",
"instructions": "Is this page a scan or photograph of a physical page (softened edges, skew, shadow, paper texture, halftone noise), rather than rendered from a digital source with clean typography?"
}
}
}
And this is the shape of the answer for one form. There is no text to parse. Every option gets a probability, and the yes or no question returns one number:
{
"answers": {
"document_kind": {
"choice": "form_or_application",
"probabilities": {
"form_or_application": 0.88,
"general_document": 0.04,
"handbook_manual_or_guide": 0.02,
"policy_or_procedure": 0.02,
"other": 0.02,
"map_flyer_or_brochure": 0.01,
"report_plan_or_proposal": 0.01,
"research_or_grant": 0.01,
"meeting_agenda_or_minutes": 0.00
}
},
"is_scanned": {"noul": 0.07}
}
}
The word "noul" is the API's name for a yes or no question. The number is the probability of yes.
Results
| Measure | Result |
|---|---|
| Document kind, against the corrected reference | 58 of 60, 96.7 percent |
| Document kind, against the labels as first written | 49 of 63, 77.8 percent |
| Scanned page detection, against a model-free rule | 67 of 68, 98.5 percent |
| Audience and format questions, where a label existed | 13 of 18, 72.2 percent |
| Model errors or empty answers | 0 |
| Warm response time per page | about 0.6 seconds |
| Memory while loaded | about 14 GB |
| Same answers on a second run | yes |
Seven of the 75 files never downloaded and five carried no kind label. Those twelve are reported separately. They are not in any denominator above.
The probability the model attached to its chosen kind is worth a look, because it is the number you would route on.
| Probability of the chosen kind | Answers | Correct, of scored |
|---|---|---|
| 0.30 to 0.49 | 15 | 10 of 10 |
| 0.50 to 0.69 | 17 | 13 of 15 |
| 0.70 to 0.89 | 23 | 22 of 22 |
| 0.90 to 1.00 | 13 | 13 of 13 |
In this sample the unsure answers were all right, and the two misses sat in the 0.50 to 0.69 band. So a threshold on the kind question would not have caught the two mistakes here unless you set it above 0.7, which would also send a quarter of the pile to a person. Treat the number as a signal to look, not as a guarantee, and set the threshold on your own data.
The scanned-page check deserves a sentence. A rule with no model in it, fewer than 25 characters of text on page one plus a single image covering most of the page, agreed with the model on 67 of 68 documents. The one miss was the only fully scanned PDF in the set. The model called it born-digital with a probability of 0.07. That is a low number, and a threshold would have sent it for review.
Why the two agreement numbers differ
The labels came from a content catalog. They were written to describe documents for people browsing a site, not to test a model, and several were loose. The model's answer was often more specific than the label: "form" where the label said "general document", or "policy" where the label said "guide".
So the 14 disagreements were re-checked. Three larger models each saw the page image plus two pages of text, saw the catalog label, did not see the small model's answer, and voted on one question: is this document clearly one kind? Eleven were, and the majority kind became the reference. Three were not clearly any one kind and left the scoring. The catalog label survived that review on 2 of the 14.
Scored against that corrected reference, the model agreed on 58 of 60. On the 49 documents nobody disputed it agreed on all 49. On the 11 re-labelled documents it agreed on 9.
Where it was wrong
Two documents remain in disagreement. In both, the model said "policy or procedure". One is an instruction sheet the reference calls a guide. The other is a request form the reference calls a form. Both sit on a boundary between kinds, and a person could argue either way.
The audience and format questions were weaker, at 13 of 18. The sample is small and the labels for those questions were sparse, so treat that number as a first look rather than a result.
What you can do with this
- Sort an inbox of PDFs by kind before anyone opens one. Forms, flyers, minutes, and reports each have a known remediation path.
- Pull scanned pages out first. They need text recognition before anything else, and the model finds them from the page image.
- Use the probability as a triage signal. Route answers under a threshold you choose to a person. In this test the unsure answers were right, so pick the threshold from your own disagreements, not from this note.
- Run it on a laptop. A 9B model at about 14 GB fits on ordinary hardware, and sub-second answers mean a pile of thousands is an afternoon, not a project.
Limits
- The model saw page one only. A long document can change kind after its cover page.
- The corrected reference came from a model panel, not from people. It is a repeatable second opinion, frozen so later runs compare against the same file.
- The panel reviewed only the 14 disagreements. The other 49 labels were not re-checked.
- 75 documents from one source is a small sample. The numbers show the model is worth testing on your own pile, not that it will score the same there.
Try it on your own pile
Take 50 of your own PDFs, write down the kind you believe each one is, and run the same question. Score against your labels, then read every disagreement before you trust either side. The agent test corpus on this site comes with a scorer that reports precision and recall per rule and lists each miss by page, which is a good shape to copy for your own scoring.
Sources
- huggingface.co/cloudflare/clef-flash, the model card and Apache 2.0 weights for the 9B model
- huggingface.co/cloudflare/clef, the 27B model in the same family
- ollama.com, the library page for clef-flash and the question format
- a11y.world, the agent test corpus and its scorer