Laya Put to the Test: The Open Jev Alternative on Your Own Mac

Since September 19, TypeSafe AI has offered Jev, a model that writes no text but decides between fixed answer options and returns probabilities. Jev only runs as a hosted API. One day earlier, a project appeared on PyPI that implements the same idea in the open. Laya from Convai Innovations is licensed under Apache 2.0, runs locally, and by its own account speaks the same protocol as Jev.
I measured Laya on a MacBook Pro with an M3 Max against the same 200 banking queries I used to test Jev a week ago. On top of that came 100 German translations and a run with the answer options in a different order.
An Encoder Instead of a Language Model
Laya is not a language model in the usual sense. It does not generate anything word by word. It reads the input and the answer options in a single pass and scores each option. Under the hood sits an encoder, a model that turns text into a representation of its meaning instead of producing new text. On top of it sits a small decision head that scores the options.
| Checkpoint | Base | Parameters | Context |
|---|---|---|---|
english | ModernBERT-large | 421M | 512 tokens |
multilingual | mmBERT-base | 322M | 1,024 tokens, up to 8,192 |
typed-decisions | ModernBERT-large, fine-tuned | 421M | 1,024 tokens |
A router picks the checkpoint based on the script and language of the input. English text goes to english, everything else to multilingual. The third checkpoint is fine-tuned on four synthetic workflows and is only used when explicitly requested (according to the README and model card).
Development moves fast. Between the first version on September 18 and version 0.3.24 on October 2, there were 32 releases on PyPI.
What the Vendor Promises
Laya openly compares itself with Jev in its own documentation. The following figures come from Convai Innovations, measured on a Tesla T4 GPU, and are not independent measurements.
| Claim (vendor) | Laya | Jev |
|---|---|---|
| Latency per question | 32.8 ms | 236 to 276 ms |
| Calibration error (ECE) | 0.081 | 0.246 |
Accuracy typed-decisions | 0.766 | 0.727 |
The expected calibration error (ECE) measures how far the reported probabilities deviate from the actual hit rate. The smaller it is, the more a statement like “90 percent sure” can be trusted.
The documentation is not fully consistent with itself. The README lists 48 of 51 measured languages, the model card 45 of 51. The model card also admits that the base checkpoints are weak without fine-tuning and reach only 0.362 on its own decision task. Their probabilities are said to be overconfident, and the order of the options influences the answer.
The Test Setup
The task is the same as in the Jev test. From PolyAI’s public Banking77 dataset come 200 real customer queries, 20 from each of ten categories, drawn with a fixed random seed. The measurement script reads categories and descriptions directly from the Jev script, so both models get exactly the same question.
Laya ran locally in version 0.3.24 with PyTorch 2.14.1 on macOS 27.2. I asked each query three times, with the options in original order, reversed, and shifted by five positions. All 600 queries went to the English checkpoint, and none was truncated. Model loading time is not included in the latency figures.
Ten Percentage Points Behind Jev
| Model | Hits (original order) | Median latency |
|---|---|---|
| Jev 1.13 (API, Sept 25) | 187 of 200 (93.5%) | 380 ms including network |
Laya english, M3 Max GPU | 167 of 200 (83.5%) | 35 ms |
Laya english, M3 Max CPU | 167 of 200 (83.5%) | 180 ms |
Laya stays ten percentage points behind Jev. The latencies are not directly comparable, because Jev’s figure includes the round trip through OpenRouter while Laya’s is pure compute time on the local machine. For operations, that is exactly the difference. On the Mac’s GPU, Laya answers in about 35 milliseconds without a query leaving the device, and it costs nothing per call. Jev cost €0.02 per 1,000 queries in the test.
The most common mistake is the same for both models. Eleven times, Laya classifies a query where the customer’s own balance did not go up after a transfer as “transfer not received by recipient”. Jev mixes up the same two categories seven times. The descriptions are close in meaning, and this is where the task itself reaches its limits.
Option Order Changes the Result
With reversed order, Laya gets 161 of 200 right, with shifted order 160. For 37 of the 200 queries, the answer changes only because the options are sorted differently. Laya answers just 151 queries correctly in all three orders.
That matches what Laya documents itself, and a small comparison test by Runtime Weekly, in which Laya decided differently depending on the order in five of 30 support cases. Anyone deploying Laya should test the option order as well. Laya offers an option_order parameter for running through several orders.
Useful Probabilities Despite the Warnings
The reported probabilities are more useful than the documentation’s warnings suggest. 139 answers had a probability of at least 0.9, and 133 of those were correct. For correct answers, the probability of the chosen option averaged 0.93, for wrong ones 0.61.
That suggests a practical approach. Answers below a threshold go to a human, the rest are processed automatically. On my 200 queries, a threshold of 0.9 would have handled about 70 percent automatically, with a hit rate of 96 percent. For production use, the threshold would have to be set on your own data.
Laya itself reports one limitation when loading the English checkpoint. For questions with eleven or more options, the checkpoint ships invalid temperature values, and confidence for those entries should be treated as uncalibrated. My task had ten options and, as I read the message, should not be affected.
Clearly Weaker in German
For the German queries, I used the 100 translations from the Jev test, with English categories as before.
| Model | German queries | Same 100 in English |
|---|---|---|
| Jev 1.13 | 95 of 100 | 95 of 100 |
| Laya, automatic routing | 72 of 100 | 86 of 100 |
Laya, forced to english | 39 of 100 | 86 of 100 |
Jev loses nothing in translation. Laya drops from 86 to 72 hits. The router sent 94 of the 100 German queries to the multilingual checkpoint, and six stayed with the English one. Without the router, the English checkpoint falls to 39 hits. The figures confirm what the documentation says about the English checkpoint, and they show that the multilingual checkpoint also stays noticeably behind the English result in German. At 23 milliseconds, it is at least the faster of the two.
Jev Code Works Unchanged
Laya ships an HTTP server that mirrors Jev’s /v1/systemone endpoint. I sent the unchanged request from my Jev test script to the local server, including the model name typesafe/jev-1.13. The response came back with the same fields (choice, probabilities, confidence) and the correct category.
pip install "laya[serve]"
LAYA_PORT=8765 LAYA_MODELS=english,multilingual laya-serve
If you use Jev today, you can try Laya by changing only the base URL. On first start, the server downloads all three checkpoints by default. LAYA_MODELS limits that to the ones you need.
Where Laya Is the Better Choice
Laya does not replace Jev when the last few percentage points matter or when the input is in German. It is a serious option when data must not leave the building, when cost per query counts, or when decisions have to be made within a few milliseconds. For pre-sorting English queries, with a threshold for the uncertain cases, the quality is sufficient. Testing the option order and setting the threshold on your own data remain mandatory.
The project is two weeks old and moving fast. The figures here apply to version 0.3.24 and may look different in a month.
Sources
- Laya, source code and README: github.com/NandhaKishorM/laya
- Laya, model card: huggingface.co/convaiinnovations/laya
- Laya on PyPI: pypi.org/project/laya
- PolyAI, Banking77 dataset: github.com/PolyAI-LDN/task-specific-datasets
- Runtime Weekly, Kev and Laya compared: github.com/Runtime-weekly/runtime-tutorials
- Jev Put to the Test (comparison figures from Sept 25, 2026): rotecodefraktion.de
- Own measurement: Laya 0.3.24, checkpoint revision
55cf4c4, PyTorch 2.14.1, MacBook Pro M3 Max, macOS 27.2, October 3, 2026