DE EN

Blog

On classification, Jev and TypeSafe. And why the Mittelstand can ignore the model debate

My former mentor Ashley wrote an article this week that I recommend to anyone who thinks of AI not as chat but as part of a process. It is about Jev, the model from TypeSafe AI that writes no text but answers questions with a typed answer and a probability. Ashley is building an assistant for actuaries on top of it. The sentence that stuck with me sits in the middle of the article, not in the headline: “Code owns the control flow.”

The wheel is not left to chance. A workflow needs deterministic rules, the model answers only the specific questions that need pattern matching, and the program assembles the answers. The opposite does not hold in processes without supervision: an agent that decides for itself what to do next has, in every loop, a chance to drift, and then a great deal of software gets written to keep it on course.

Ashley writes from the perspective of a reinsurer. Our customers are by now mostly in the German Mittelstand, above all together with our partner Kendox. And there, “the code” means something very concrete.

The code is the workflow, and the company already has it

Anyone running a document management system has long since made the important decisions. The invoice workflow says who checks an invoice, who approves it, above which amount a second approver is needed and when it goes to the ERP. The mailroom workflow says which letter belongs to which file and who receives it. The file says what complete means. And all of it is logged, because an auditor will ask in three years.

Kendox InfoShare is such a system: invoice intake with a configurable approval workflow, mailroom, files, hand-over to SAP or Microsoft Dynamics, around 1,500 installations in the German-speaking countries. The code in Ashley’s sentence is, here, a set of rules that a head of accounting or a municipal treasurer laid down, often years ago, and that everyone in the building knows. These rules are the most valuable part of the process. They encode how the company decides, and they are exactly what gets lost in an AI project that starts with an agent instead of with the workflow.

What we bring to scale together with Kendox

Concretely, at joint customers it looks like this. feld.ai is natively integrated into InfoShare, users keep working in the interface they know, and the workflow calls us when it needs an answer:

  • Automated invoice intake. Which invoice is this, from whom, with which line items, and where does it belong.
  • Routing incoming mail to the right file. The least glamorous job in the whole process, and the one that costs the most time when a person does it.
  • Checking files for completeness and proposing follow-up questions. Is something missing that the rules say should be there, and who needs to be asked what.
  • Redaction. Which passage in the document is personal data and must not leave the building.
  • Checking costs against contracts. Does the line item match what was agreed, and where can you see it. Approval stays in the workflow, with the person the rules say is responsible.
  • Automated quotation drafting. Which request is this, what was quoted for something similar before, which items belong in.

All of these are classification questions. We answer them on our own infrastructure, with large language models, with fine-tuned small models and with whatever else helps, and we answer them so that it works in the process, not so that it looks good in a demo. Every answer comes with a probability and with the place in the original document it came from. The workflow decides what happens next: above a threshold it continues, below it the case goes to the person the rules name. The installation runs as a standalone instance in the environment Kendox chooses.

That is exactly the architecture Ashley describes, only nobody has to build it from scratch. What is new about Jev is that this design now has a name, and that an answer becomes cheap and fast enough to ask three times instead of once. What is not new is the idea that the model takes over the process. It does not.

An example: mail, file, follow-up

A letter arrives, say an objection against a wastewater fee notice. The first question is not “what does it say” but: which file does it belong to? Once the file is assigned, the workflow asks its questions, the same ones that stand in every set of working instructions an experienced clerk leaves behind: Is a document missing? Is the addressee correct? Was the deadline met? Is the surface area substantiated, and by what?

Each of these questions gets an answer with a probability and a source reference. “Deadline met: 99 percent, page 2, delivery note” is a relief. “Area substantiated: 71 percent, annex 3” is an instruction: look here, and that is the follow-up question the workflow proposes right away. The clerk no longer sees every question, only the two she is actually needed for. And the promise someone made on the phone years ago is in no file; no model sees it, which is why one question stays in the workflow that only a person can answer: is something missing here that is not in the document?

And the auditor? A decision assembled from ten named, individually inspectable answers is more auditable than a paragraph generated after the decision. The auditor has to take the probability on trust. The delivery note he can read. And the log sits where it belongs: in the DMS, next to the file, not in a chat history.

The weak spot was never the model

The weak spot in these projects was never that one more model was missing. Of course the models have to get better, and of course the next round is welcome (hello, TypeSafe). If the answer to “which file” and “was the deadline met” gets cheaper and faster, we ask it more often, and that is a good thing. The design behind Jev has already been reproduced several times; which model answers the questions in two years, nobody knows. What remains is the workflow: the questions the department wrote down, the threshold the management signed, the log. That belongs to the company and survives every model change.

The real weak spot is the same as it has been forever: does the human actually know what is what? Which invoice is an exception, which letter belongs to which file, what does complete mean, who may approve, and who notices when the machine is wrong. Add to that what no model comparison ever includes: the colleague who has looked at every invoice for fifteen years has to want to see only the exceptions from now on. The treasurer has to sign the threshold. The IT lead has to set up the sample that checks whether 95 percent really is 95 percent. A model is not Father Christmas delivering ROI on request. Reality, unfortunately, is more real.

That is why the Mittelstand can largely ignore the model debate. The first question is not “which model” but: who in the building knows what is what, and where is it written down? The workflow that encodes it is the most important rule of the whole AI project, and usually it already exists. What the integration into a DMS looks like on our side is on the page for software partners.

The difference between theory and practice is, as a wise person once put it, smaller in theory than in practice. And because practice has the better examples: if you know cases where even people find it hard to make “trivial” process decisions (Is the invoice right? Who is next? What does the customer actually want?), I would be grateful to hear them, on LinkedIn or by email.

Source: Ashley Hirst, LinkedIn

Share Share on LinkedIn

Get new blog posts delivered to your inbox

Occasional notes on sovereign AI and document automation. No spam.

Talk to the founder Request a demo All posts
Talk to the founder