Jev as a Pre-Filter in the Ticket Workflow: Fewer Model Calls, Fewer False Alarms

Advanced Topics 2 · Series: Getting Started with n8n
Up to version 0.10, every ticket in the workflow went through a local language model. A round robin split the tickets between Ollama running qwen2.5 and Qwen3-8B on MLX, as set up in Article 9. The model picked the category and decided whether a ticket triggers the pager. In my test of Jev I announced that I would put TypeSafe’s decision model in front of it. Jev answers typed questions with probabilities. When Jev is sure, the ticket no longer needs a language model.
That worked, but not on the first attempt. The measurement needed labels of its own, the priority from my matrix contradicted Jev’s P1 suspicion, and the pager fired falsely on 32 of 100 tickets, all of them from the language model branch. At the end there is the bill for 1,000 tickets a day with cloud models instead of local ones. With DeepSeek V4 Flash behind Jev it comes to 67 dollars a year. But first things first.
Jev decides first, the language model only when in doubt
The pre-classification is a sub-workflow of its own with four nodes. A Code node builds the request from the ticket, an HTTP node sends it to Jev via OpenRouter, and a second Code node evaluates the answer. Jev receives four questions in German about the same ticket:
| Question | Type | Answers |
|---|---|---|
| Category | Choice | sap-basis, sap-functional, infrastruktur, cloud, security-pki, sonstiges |
| Impact | Choice | single person, group, many or all |
| Urgency | Choice | later, today, now |
| P1 | Noul (yes/no) | probability that a P1 incident is present |
The Code node picks the route from the answers. If the category confidence is below 0.9, the ticket goes to the language model. The same applies if the P1 probability lies between 0.5 and 0.9, so neither a clear yes nor a clear no. In all other cases Jev decides.
In the main workflow the pre-classification sits right after validation. An IF node sends confident tickets straight to the pipeline and uncertain ones into the existing round robin, which stays unchanged. If Jev fails, the ticket also ends up with the language model. The HTTP node has a two-second timeout for that.
Jev costs next to nothing. A request with four questions averages 935 input tokens, and 1,000 tickets cost 0.039 dollars (own measurement via OpenRouter, September 2026).
No yardstick without labels of my own
The test dataset from Article 3 has one category per ticket, which comes from the generator and can be measured against. That is not true for priority. The generator set the priority field at random. A ticket about a complete production outage can carry low. Measuring against these values would only have shown how well a model guesses the dice.
So I had the 100 tickets rated from P1 to P4 again, by Claude Opus 5.5 in a Claude Code session, ticket by ticket with a justification and following a written rubric. What counts are the facts described, not the tone. P1 means a production system or central service is down for many users, a security incident has occurred, or a central business process has stopped. Whatever the ticket does not say is not assumed. The result was 8 P1, 31 P2, 49 P3 and 12 P4. 15 tickets are marked as borderline cases where two levels are defensible. For comparison, Opus 5.5 and Fable 5.1 had already rated the tickets once before, independently of each other and with the same rubric. So the labels do not come from operations experts, and the borderline cases are where I would check them first.
These labels are used for evaluation only. The measurement script sends neither them nor the random generator priority to the workflow. The pipeline from Article 6 also pages when the submitter sets critical. With the random priority in the request, the pager would have fired more often by chance than because of the classification.
The priority contradicted the P1 suspicion
The Code node computes a priority from Jev’s answers on impact and urgency. The levels critical, high, medium and low correspond to P1 through P4. In the first version, the combination “many or all” and “now” set the priority to critical. The P1 suspicion, however, came from the separate P1 question. The two drifted apart. In 20 cases the priority said critical, but the P1 suspicion was set for only 3 of them. Against the labels, 13 of these 20 tickets were not P1 at all.
The fix separates the two paths. critical now comes only from the P1 question, and the matrix sets high at most. After that, the priority critical appeared in all 100 cases exactly when the P1 suspicion was set as well.
A second weakness remained. The matrix rated too high, and 36 P3 tickets ended up as P2. I therefore replaced it with the classic ITIL matrix, which weights impact and urgency equally:
| Impact ↓ / Urgency → | later | today | now |
|---|---|---|---|
| single person | low | low | medium |
| group | low | medium | high |
| many or all | medium | high | high |
critical, meaning P1, still comes only from the P1 question. Against the labels, exact agreement rose from 50 percent after the critical fix (40 originally) to 63 percent, and all 100 tickets are at most one level off. The price is 6 P2 tickets that now pass as P3. A matrix I had fitted to exactly these 100 tickets reached 67 percent. I discarded it because it memorizes the test data and proves nothing on other tickets.
32 false alarms from the language model
The bigger finding concerned the pager. Version 0.10 raised an alarm on 48 tickets that were not P1. With Jev in front it was 32, and all 32 came from the path the language model still decides. Both versions caught all eight real P1.
The cause was the prompt from Article 6. It sets p1_suspected when a ticket “indicates” a production outage, when data loss “threatens” or a month-end close is “at risk”. On top of that, a clause treats “URGENT” in the subject combined with concrete damage as P1. The field itself is a boolean. A ticket that is serious but not P1 has no answer other than yes or no. When in doubt, the models chose yes, especially when the ticket was worded loudly.
To check whether this came from the size of the local models, I sent the same prompt to Claude Haiku 4.5 and Opus 5.5 (via the subscription with claude -p, without tools and with the classifier prompt as system prompt). Opus produced 30 false alarms, Haiku 50, and qwen2.5 with 7 billion parameters 31. With this prompt, model size makes no difference.
That changes with stricter wording. A variant that only counts conditions that have actually occurred as P1 brought false alarms down to 0 for Opus and 1 for Haiku. In exchange, both missed three or four of the eight real P1. The local models stayed at 26 and 41 false alarms.
The solution was a scale instead of a switch. The new prompt asks for a severity field with levels P1 to P4, described by the same facts as in my rubric, and states explicitly that possible consequences do not count as having occurred. That gives a serious ticket a place below P1. In the offline comparison, false alarms for qwen2.5 dropped from 31 to 16, with all eight P1 still caught.
One gap remained in the model itself. qwen2.5 set severity to P1 on 13 tickets but p1_suspected on 24, although the prompt equates the two. So in the workflow the model no longer decides about the pager. A Code node after the LLM chain does:
// Pager nur bei severity P1: kleine Modelle setzen p1_suspected sonst auch bei P2
const o = $json.output ?? {};
const p1 = o.severity === 'P1';
return { json: { ...$json, output: { ...o, p1_suspected: p1, p1_reason: p1 ? (o.p1_reason ?? '') : '' } } };
The rest of the pipeline keeps reading p1_suspected and stays unchanged.

The results in the live workflow
All values come from runs with the 100 tickets against the running workflow, each without sending a priority. Jev ran via OpenRouter, the language models locally on the Mac.
| v0.10, LLM only | v0.12 with Jev, first version | v0.12 with Jev, fixed | |
|---|---|---|---|
| Category correct | 96 | 98 | 97 |
| Tickets sent to the language model | 100 | 45 | 47 |
| Real P1 paged | 8 of 8 | 8 of 8 | 7 of 8 |
| False alarms | 48 | 32 | 16 |
| Jev priority exact | n/a | 40 % | 63 % |
| Sum of response times | 399 s | 269 s | 243 s |
| Median when Jev decides | n/a | 0.93 s | 0.80 s |
| Median with language model | 2.62 s | 3.57 s | 3.57 s |
Jev decides a little over half of the tickets on its own and needs less than a second for it, measured across the whole workflow including the pipeline and the SAP enrichment. The sum of response times drops by 39 percent. False alarms fell to a third. Of that drop, 16 come from Jev and 15 from the new prompt, which works without Jev too. The ticket that now slips through is TKT-0020, a certificate whose key was probably stolen and which has already been revoked. Jev passed it on with a P1 probability of 0.59, and the language model rated it below P1. It is marked as a borderline case, and Fable 5.1 had rated it P2 in the comparison.
The tickets that reach the language model take longer than the average of all tickets used to. Most of that is the Jev call that runs first and takes just under a second at the median. Whether the uncertain tickets also take longer for the language model itself, I did not measure separately.
Eight cloud models and a router on the same prompt
Locally, a model call costs no tokens, but it costs hardware and power. To compare with cloud models, I sent the new prompt with the same 100 tickets via OpenRouter to six small models, plus the OpenRouter Auto Router, which picks the model per request. I excluded Anthropic models, including in the Auto Router, because I had already measured Haiku and Opus through the subscription. Two large models were added as a reference for the cost calculation. As in the workflow, the pager fires on severity P1. OpenRouter reports the cost with each request.
| Model | Category | Level exact | Real P1 paged | False alarms | Median | Dollars per 1,000 tickets |
|---|---|---|---|---|---|---|
| gpt-oss-20b (OpenAI) | 99 | 49 | 8/8 | 23 | 7.8 s | 0.08 |
| Ministral 8B (Mistral) | 98 | 34 | 8/8 | 28 | 1.0 s | 0.09 |
| Qwen3-8B, no reasoning (Alibaba) | 99 | 48 | 7/8 | 21 | 1.7 s | 0.17 |
| GPT-6 Luna (OpenAI) | 99 | 80 | 4/8 | 0 | 2.6 s | 0.21 |
| DeepSeek V4 Flash (DeepSeek) | 99 | 76 | 8/8 | 4 | 5.1 s | 0.25 |
| Auto Router (OpenRouter) | 99 | 67 | 7/8 | 7 | 3.6 s | 0.32 |
| Gemini 3.1 Flash Lite (Google) | 99 | 78 | 7/8 | 6 | 1.0 s | 0.38 |
| GPT-6 Sol (OpenAI) | 97 | 79 | 2/8 | 0 | 2.8 s | 3.85 |
| Gemini 3.1 Pro Preview (Google) | 99 | 82 | 7/8 | 1 | 8.5 s | 13.65 |
Own measurement, September 26, 2026, prices as listed by OpenRouter on that day.
Every model gets the category right almost every time. The differences are in the pager. The smallest models behave like the local ones, finding all P1 and firing falsely 21 to 28 times, more often than the local setup with 16. GPT-6 Luna and GPT-6 Sol are the opposite, with hardly any false alarm but half or more of the real P1 missed. For a pager, that is the worse direction.
DeepSeek V4 Flash finds all eight P1 with four false alarms, which makes it the only model that does both well. It reasons before answering, hence the five seconds. Gemini 3.1 Flash Lite comes in close behind and answers in one second.
The Auto Router sent 84 requests to DeepSeek V4 Flash and 16 to Gemini 2.5 Flash. It costs more than DeepSeek directly and scores slightly worse. For a fixed task with a fixed prompt, automatic selection does not pay off.
Anyone sending tickets to cloud models passes their content on to the respective provider. OpenRouter can restrict routing to providers with suitable data policies. For customer data, that setting comes before the cost comparison.
1,000 tickets a day on API tokens
For the projection I calculated each model twice. Once it gets all tickets. Once Jev sits in front, and the model only gets the 47 tickets Jev passed on in the live run. Jev’s cost is then included for all 1,000 tickets. Category and pager for the combined setup are assembled from the actual answers of both sides.
| Model | LLM only, per year | with Jev, per year | Savings | with Jev: P1 / false alarms |
|---|---|---|---|---|
| gpt-oss-20b | $31 | $30 | 2 % | 8/8 / 19 |
| Ministral 8B | $34 | $30 | 11 % | 8/8 / 25 |
| Qwen3-8B | $61 | $43 | 29 % | 7/8 / 17 |
| GPT-6 Luna | $77 | $53 | 31 % | 5/8 / 0 |
| DeepSeek V4 Flash | $92 | $67 | 27 % | 8/8 / 4 |
| Auto Router | $117 | $75 | 35 % | 7/8 / 7 |
| Gemini 3.1 Flash Lite | $139 | $80 | 42 % | 7/8 / 6 |
| GPT-6 Sol | $1,404 | $694 | 51 % | 3/8 / 0 |
| Gemini 3.1 Pro Preview | $4,982 | $2,482 | 50 % | 7/8 / 1 |
1,000 tickets per day, 365 days, prices as of September 26, 2026. Jev alone costs 14 dollars a year in this setup.
The savings stay below the 53 percent of tickets Jev takes off the language model. The tickets left for the model are the difficult ones, and with reasoning models they cost more tokens than the average. With the cheapest models, Jev nearly eats up its own advantage, leaving 2 percent for gpt-oss-20b. The more expensive the model, the closer the savings get to half.
In absolute numbers, classification with small models is cheap. DeepSeek V4 Flash costs less than 100 dollars a year for 1,000 tickets a day, 67 dollars with Jev in front. Anyone using a large model because it is supposed to be more reliable at paging pays thousands a year instead. The table shows that this does not pay off. Gemini 3.1 Pro gets the level right slightly more often and fires falsely less often, but misses one P1 that DeepSeek V4 Flash finds. At fifty times the cost that is no gain, and GPT-6 Sol misses most P1.
Time shows a clearer difference than money. Locally, version 0.10 needed 4.0 seconds per ticket on average, which comes to 67 minutes of compute time for 1,000 tickets a day. With Jev it is 2.4 seconds and 40 minutes. With DeepSeek V4 Flash in the cloud, half of the tickets skip the five seconds of reasoning because Jev handles them in under a second.
Limits of the measurement
The numbers come from 100 synthetic tickets and from labels a language model set following a fixed rubric. Eight P1 is not many. A single ticket shifts the detection rate by 12.5 percentage points, and the 15 borderline cases decide whether a model looks cautious or careless. I did not fit the ITIL matrix to the data. The new prompt, however, grew out of exactly these false alarms, and the cloud comparison uses it as well. Both need to be measured again on real tickets.
The cloud prices apply to September 26, 2026. Some of the models have only been available for a few weeks, and prices and behavior can change. The projection assumes that real tickets are about as long as the test tickets and that Jev decides a similar share on its own.
Two signals for the borderline cases
The workflow now decides a little over half of the tickets without a language model, raises a third as many false alarms as before, and finds seven of eight P1. What remains open is the pager in the borderline zone. Jev’s P1 suspicion and the language model’s severity are two independent signals. Whether combining them catches the missed ticket without pushing false alarms back up is the next measurement. A pre-filter saves calls. False alarms only went down once the model stopped deciding about the pager itself.
The test scripts, the labels with their rubric and the sub-workflow are in the demo repo n8n-einstieg under the tag v0.12.
← Advanced Topics 1: What an Agent Remembers
Sources
- Jev Put to the Test: A Model That Decides Instead of Writing, my own article with the basics of Jev
- TypeSafe: Primitives, question types Choice, Score and Noul
- OpenRouter: Jev, access via OpenRouter
- OpenRouter: Auto Router, restricting model selection with
allowed_models - OpenRouter: Provider Routing, filtering by data policies
- OpenRouter: Models and pricing, as of September 26, 2026