← yayster.com

Local AI · October 2026

Why Your Local LLM's 75% Accuracy Is Worse Than It Sounds

A 4-billion-parameter model sorted a month of bookkeeping in 23.0 seconds on a CPU, with no GPU, no cloud and no API key. It got 75% of the rows right. Then I weighted the errors by dollars instead of rows, and the number became 29.5%.

* * *

Every tutorial that puts a language model on a spreadsheet quotes the same metric: it got n percent of the rows right. We did that too. The number was 75% — 27 of 36 transactions correctly categorised, offline, in 23 seconds, by a model small enough to run on a CPU.

Then we graded it the way a bookkeeper would, by money rather than by row. $2,096.88 of the $7,114.41 that left the account was in the wrong category. Twenty-nine and a half percent. Same run, same output, same model — a completely different verdict, because the errors do not land randomly.

Row accuracy flatters a model, and it flatters it hardest on exactly the data you care about: the big charges.

This page is the whole method — the install, the prompt, the run, the grading — plus the two failures that actually taught us something, and the one-line check that caught the expensive one. The transactions are synthetic; our real books are not going on the internet. Everything else is a real run.

* * *

The setup: a 4B model, a CPU, and no network

The model is qwen3.5:4b — 4.7 billion parameters, quantised to Q4_K_M, Apache-2.0 licensed. It runs under ollama with the GPU explicitly disabled rather than merely absent:

options = {
    "num_gpu": 0,        # not "no GPU available" — actively off
    "num_thread": 4,     # see "the config trap" below. This line is the whole post.
    "temperature": 0,
}

Read the licence yourself before you build on a model — ollama show qwen3.5:4b prints it, and “it's open source” on a forum is not a licence. Apache-2.0 is why this one can go in something commercial.

The prompt, and the one thing that must be a closed list

No framework, no agent loop, no vector store. The whole prompt:

You are a bookkeeper categorising business bank transactions.

Allowed categories (use these exact strings, nothing else):
  <15 fixed categories>

Rules:
- Positive amounts are money received -> "Income".
- Answer for EVERY numbered line, in order.
- Output ONLY lines of the form: <number>|<category>
- No explanation, no header, no blank lines.

The categories are a closed list, and that is not a style choice: a category the model invents is one your accounting software has never heard of, and you will find out at the worst possible moment. temperature: 0 for the same reason — you want the same answer twice from the same configuration.

Dry-run five rows before you run the month. Ours: 5 rows, 3.9 seconds, 0.78 s/row. That is enough to catch a prompt that returns prose, or a model that ignores the output format, before you have spent a minute per row finding out.

The real run: 36 rows, 23.0 seconds, on a CPU

Warmed, four threads, no GPU: 36 rows in 23.0 seconds, 0.64 seconds per row, 15.8 output tokens/sec. Not sped up, not trimmed.

That is the headline most write-ups stop at, and it is genuinely good: a month of books, offline, on hardware you already own, in less time than it takes to open the spreadsheet. One caveat we state out loud rather than bury — this was measured on a server CPU, not a laptop. A laptop will be slower.

The config trap: the same job at 169.84 seconds per row

The first time we ran this, it took eight and a half minutes to do three rows. Same box, same model, same prompt.

ollama sizes its thread pool from the machine's core count, not the container's. A 4-vCPU container on a 16-core host asks for 16 threads, and the threads fight each other for four cores. The fix is the num_thread line above, and it has to be passed per request — an environment variable will not do it.

The two numbers, with the caveats that belong to them, because without the caveats the comparison is not honest:

default, cold:   16 threads, 3 rows, 509.5s wall → 169.84 s/row
                 (includes 46.89s of one-time model load)
tuned, warmed:    4 threads, 36 rows, 23.0s wall →   0.64 s/row

Those runs differ in three ways at once — thread count, warm vs cold, and 3 rows against 36 — so the only column that compares them fairly is seconds per row, and on that column it is 169.84 against 0.64. We are not going to print a single multiplier as the headline, because a multiplier that quietly contains a 46.89-second model load is the kind of number that gets repeated without its footnotes.

“Local models are too slow” is usually this, unmeasured.

One more thing that fell out of it: temperature was 0 in both runs, and they disagreed with each other. Determinism is per-configuration, not a property of temperature: 0 on its own.

Now grade it — and grade it by money

27 of 36 rows correct. 75%. Here is what the nine wrong ones actually were:

$600.00  VENMO PAYMENT J MARTINEZ  → "Personal"      (really: Contractors)
$450.00  WEWORK 1120 MARKET        → "Software"      (really: Rent & Utilities)
$329.99  AMZN Mktp US*8H31LK022    → "Software"      (really: Equipment)
$289.00  MARRIOTT BONVOY ATLANTA   → "Meals"         (really: Travel)
$240.00  PAYPAL *FIVERR SELLER     → "Income"        (really: Contractors)
 $88.32  HOME DEPOT #4501          → "Office"        (really: Equipment)
 $64.18  AMZN Mktp US*2K4L9WQ83    → "Software"      (really: Office Supplies)
 $23.40  WALGREENS #6621           → "Office"        (really: Personal)
 $11.99  SPOTIFY USA               → "Software"      (really: Personal)

Look at the order. The errors are sorted by size and the expensive ones are at the top, which is not a coincidence: a $600 Venmo payment and a $450 coworking charge are exactly the rows whose description does not say what they are. “WeWork” reads like a SaaS company. A hotel chain reads like a meal. Meanwhile the rows the model gets right are the obvious small ones.

Worth splitting honestly, because not all of it is the model's fault: of the $2,096.88, $1,002.40 sits on rows where the description does contain enough to get it right, and $1,094.48 sits on rows where nobody could — a bare $600 Venmo to a person, an Amazon order that could be equipment or stationery. The second half is not an accuracy problem. It is a this row needs a human and always will problem, and the right design sends it to one instead of guessing.

Money received, for completeness: $6,950.00, all of it correct. The model is good at the easy direction.

Then 13 regexes beat it in 0.0006 seconds

Before concluding anything about the model, we wrote the dumbest possible baseline: thirteen regular expressions over merchant names, plus one line that files any positive amount as money received, with everything unmatched held for review.

rules layer:   27 of 36 rows filed in 0.0006 seconds
               (25 by a merchant regex, 2 by the positive-amount rule)
model called:  0 times
held for a human: 9 rows — $2,634.64, 37% of total spend

The same 27 rows. Four orders of magnitude faster than “fast”. Zero model calls. The difference is that the regex layer knows what it does not know: it files what matches and refuses the rest, where the model answered all 36 with equal confidence and was wrong on nine of them.

If you are automating a categorisation job, write the regexes first. They will handle the boring majority, they cost nothing, and they make it obvious how much of the problem is actually left.

So what is the model actually for?

It is for writing the rules, once — which is a completely different job from filing the rows, and the only one in this pipeline it is genuinely better at than a regex.

Fed 34 distinct merchants from the history, it proposed a rules table in 26.6 seconds: 21 marked AUTO (confident enough to file without asking) and 13 marked ASK. No invented categories. Run once at setup; the table it writes then runs in microseconds forever.

And in that table was the single worst error of the whole project:

AMZN Mktp US  →  Income   [AUTO]

validator: REJECTED — 2 row(s), $394.17 LEAVING the account

$394.17 of costs, booked as revenue, marked AUTO, with no human check. Not a rounding error — a sign error, in the direction that invents income you then pay tax on. And note where it happened: not in the run everyone grades, but in the quiet one-off setup step that writes the thing that runs forever.

The one-line guardrail

The check that caught it, in full:

Money leaving the account can never be Income.

That is a sign check. It knows nothing about bookkeeping, it cannot be reasoned with, and it rejected the bad rule instantly: 1 of 21 proposals rejected, 20 passed through to a human, the entire $394.17 caught.

A bigger model does not fix this. The problem was never capability — it was that nothing was checking. A 70B model marking the same rule AUTO produces the same wrong books, slightly more fluently.

One inference, flagged as an inference rather than a measurement: the same check pointed at the first run's output would also have flagged the $240.00 Fiverr row, which was money leaving the account labelled Income for exactly the same reason. We measured the rejection on the rules table; that one is us reasoning, and we would rather say so than let it read as a second result.

The shape that actually ships

Putting the pieces in the order that survives contact with real books:

The model writes the draft. It never signs off. That one sentence is the difference between this being useful and this being a liability, and it is the part the demos leave out — because the demo ends at 23 seconds and the liability shows up at tax time.

What is real here, and what is not

The transactions are synthetic — invented data with realistic merchant strings. Our real books are not going on the internet, and we would rather say that than imply we published a month of company spend.

Everything else is a real, uncut run: the model, the prompt, the timings, and every mistake. The 23.0 seconds is not sped up and the 36-row run is not trimmed. The $394.17 is what the model actually wrote. Every figure on this page comes out of a run file written by the command that produced it — the artifact and the number are one execution, not a transcription of one.

We are also not going to pretend the model did the bookkeeping. It did a fast first pass, got the cheap rows right, got the expensive rows wrong, and was most dangerous in the step nobody inspects. That is a useful tool with a known failure mode, which is a different and more honest claim than “AI does your books.”

We build these: one scoped job, tested, handed over with the guardrails already in it. If you have a repetitive task you would rather hand off than spend a weekend on, message us — a human reads every message.

And if you would rather build it yourself: genuinely fine. That is why the numbers are in here.

This is Yayster. We build local-AI and automation pipelines, and we document them while we do it — every guide here is something we ran.

Reach us on Telegram → t.me/yaysterllc_official

YouTube · X/Twitter