Jev vs Laya: Hosted API or Open-Source Model?
Last checked · Independent guide, not affiliated with TypeSafe AI
Jev is TypeSafe's closed, hosted decision model; Laya is a free Apache-2.0 model (421M parameters) that answers the same questions on your own hardware. On identical test questions, Laya beat Jev on a 4-option news task (96 vs 86 of 100) and ran far faster locally, while Jev won easily with 77 options (78 vs 41). Laya suits few-option or private jobs; Jev suits many options and long inputs.
Laya is the open-source model people compare Jev with most. It appeared on GitHub on September 18, 2026, three days after Jev, by Nandakishor M of Convai Innovations, whose Hacker News post “I built non-autoregressive decision models with RL a year ago” reached 1,351 points. A week later its GitHub repository had about 23,700 stars.
Laya does not copy Jev’s weights, which are not public. It is a separate, much smaller model that answers the same kinds of questions in the same format.
At a glance
Section titled “At a glance”| Jev | Laya | |
|---|---|---|
| Made by | TypeSafe AI | Nandakishor M / Convai Innovations |
| Released | September 15, 2026 | September 18, 2026 |
| Access | Hosted API only (TypeSafe, OpenRouter, Vercel, Cloudflare) | Download and run yourself: pip install laya |
| License | Proprietary | Apache-2.0 |
| Model size | Not published | 421M (English, ModernBERT-large); 322M (multilingual, mmBERT-base) |
| Price | $0.042 per million input tokens, output free | Free; you pay for your own hardware |
| Question types | Noul, Choice, Score | Noul, Choice, Score, same JSON shape |
| Input length | About 32k tokens of state plus the longest question | 512 tokens (English checkpoint); 1,024, up to 8,192, on the multilingual one |
| Options per Choice | Up to 255 | Limited by a shared token budget for option text; accuracy drops beyond about 20 options |
| Languages | Best in English | Router sends non-English text to a multilingual checkpoint covering 100+ languages |
| Fine-tuning | Not offered | Yes, with a notebook that runs on free Kaggle GPUs |
| Data leaves your machine | Yes | No |
Our head-to-head test
Section titled “Our head-to-head test”On September 25, 2026 we ran both models on the same 200 questions from two public datasets: 100 news articles from the AG News test set (choose one of 4 topics) and 100 customer messages from the Banking77 test set (choose one of 77 intents). Each model got exactly the same instruction and option descriptions, one Choice question per item. Jev was jev-1.13.0 on TypeSafe’s API. Laya was version 0.3.20 with its default router, which picked the English checkpoint for every item, run on an Apple M2 Max laptop. We did not tune or calibrate either model.
| Result | Jev | Laya |
|---|---|---|
| AG News, 4 options: correct out of 100 | 86 | 96 |
| Banking77, 77 options: correct out of 100 | 78 | 41 |
| Median time per question, AG News | 523 ms (API call from East Asia, includes network) | 34 ms (laptop GPU); 169 ms on CPU only |
| Median time per question, Banking77 | 552 ms (same) | 54 ms (laptop GPU) |
| Average probability of the chosen option when wrong, AG News | 0.87 | 0.65 |
| Average probability of the chosen option when wrong, Banking77 | 0.72 | 0.84 |
What the numbers say:
- Few, clearly different options: Laya is at least as good. On AG News, Laya got 12 items right that Jev missed, and Jev got 2 that Laya missed. Laya was also less sure of itself when it was wrong, which is what you want from a probability.
- Many similar options: Jev is far ahead. On Banking77, Jev got 39 items right that Laya missed, and Laya got only 2 that Jev missed. Laya was also confidently wrong: its wrong answers averaged 0.84 probability. Laya’s own README explains why: all option descriptions share a fixed token budget, so with 77 options each label gets only a few tokens.
- Speed depends on where the model runs. Jev’s times include a trip across the Pacific from our test machine; independent US-based tests cited in Laya’s README measured a median of about 240 to 280 ms. Laya runs next to your code, so there is no network at all.
These are small samples on two tasks with one wording, so treat them as a direction rather than a ranking. They match the pattern in Laya’s own README, which reports Laya ahead on AG News and on emotion labels, and Jev far ahead on Banking77. Our full scripts, samples and raw outputs are kept in our test records.
What Laya’s own benchmarks claim
Section titled “What Laya’s own benchmarks claim”Laya’s README compares it with Jev on several public and synthetic tasks, and is unusually open about the limits:
- Its Jev numbers come from third-party published runs, not its own Jev calls, so prompts and sample sizes differ.
- The headline typed-decisions score (0.766 against Jev’s 0.727) comes from a separately fine-tuned checkpoint. The README says the base checkpoints score below a majority-class baseline on that benchmark, and that all of the capability there comes from fine-tuning.
- Its best calibration figure (0.081 expected calibration error) is after fitting temperatures on your domain; as shipped, the README calls the checkpoints over-confident.
- It lists where Jev leads: more than about 20 options, matching full probability distributions, and raw calibration before any fitting.
Switching between them
Section titled “Switching between them”Laya ships a server, laya-serve, that speaks the same POST /v1/systemone protocol as TypeSafe’s API. Existing Jev code can call it by changing the base URL, for example with the official Python SDK’s base_url setting. Three things differ when you move code across, according to Laya’s documentation:
- Options per question share a token budget, so long or numerous option descriptions get cut. Jev accepts up to 255 options.
- Score levels all need descriptions; a
nulllevel is rejected with 422. confidenceis calculated differently, so a threshold tuned on Jev does not transfer. Laya suggests gating on itsanswer_confidencefield instead.
Because both accept the same request, running your own labeled examples through each is the fastest way to decide. See Your first Jev call for the request format.
Which should you use?
Section titled “Which should you use?”| Your situation | Better fit |
|---|---|
| A few well-separated options (topic, sentiment, yes/no checks) at high volume | Laya, if you can host it |
| Dozens of similar labels, such as support intents or product categories | Jev |
| Long documents, logs or conversations as the state | Jev (32k tokens against Laya’s 512 to 8,192) |
| Data that must not leave your servers | Laya |
| Non-English text | Try Laya’s multilingual checkpoint; Jev is strongest in English |
| You have labeled data from your own domain | Either: fine-tune Laya, or use the data to set thresholds on Jev |
| No machine learning infrastructure at all | Jev |
Laya is one of more than a dozen open-source Jev alternatives; the others, including Kev, SemIf and Bespoke Nimble, are compared on Jev alternatives. For Jev’s own weak spots, see Jev limitations.