Skip to content

Jev vs Laya: Hosted API or Open-Source Model?

Last checked · Independent guide, not affiliated with TypeSafe AI

ANSWER

Jev is TypeSafe's closed, hosted decision model; Laya is a free Apache-2.0 model (421M parameters) that answers the same questions on your own hardware. On identical test questions, Laya beat Jev on a 4-option news task (96 vs 86 of 100) and ran far faster locally, while Jev won easily with 77 options (78 vs 41). Laya suits few-option or private jobs; Jev suits many options and long inputs.

Laya is the open-source model people compare Jev with most. It appeared on GitHub on September 18, 2026, three days after Jev, by Nandakishor M of Convai Innovations, whose Hacker News post “I built non-autoregressive decision models with RL a year ago” reached 1,351 points. A week later its GitHub repository had about 23,700 stars.

Laya does not copy Jev’s weights, which are not public. It is a separate, much smaller model that answers the same kinds of questions in the same format.

Jev Laya
Made by TypeSafe AI Nandakishor M / Convai Innovations
Released September 15, 2026 September 18, 2026
Access Hosted API only (TypeSafe, OpenRouter, Vercel, Cloudflare) Download and run yourself: pip install laya
License Proprietary Apache-2.0
Model size Not published 421M (English, ModernBERT-large); 322M (multilingual, mmBERT-base)
Price $0.042 per million input tokens, output free Free; you pay for your own hardware
Question types Noul, Choice, Score Noul, Choice, Score, same JSON shape
Input length About 32k tokens of state plus the longest question 512 tokens (English checkpoint); 1,024, up to 8,192, on the multilingual one
Options per Choice Up to 255 Limited by a shared token budget for option text; accuracy drops beyond about 20 options
Languages Best in English Router sends non-English text to a multilingual checkpoint covering 100+ languages
Fine-tuning Not offered Yes, with a notebook that runs on free Kaggle GPUs
Data leaves your machine Yes No

On September 25, 2026 we ran both models on the same 200 questions from two public datasets: 100 news articles from the AG News test set (choose one of 4 topics) and 100 customer messages from the Banking77 test set (choose one of 77 intents). Each model got exactly the same instruction and option descriptions, one Choice question per item. Jev was jev-1.13.0 on TypeSafe’s API. Laya was version 0.3.20 with its default router, which picked the English checkpoint for every item, run on an Apple M2 Max laptop. We did not tune or calibrate either model.

Result Jev Laya
AG News, 4 options: correct out of 100 86 96
Banking77, 77 options: correct out of 100 78 41
Median time per question, AG News 523 ms (API call from East Asia, includes network) 34 ms (laptop GPU); 169 ms on CPU only
Median time per question, Banking77 552 ms (same) 54 ms (laptop GPU)
Average probability of the chosen option when wrong, AG News 0.87 0.65
Average probability of the chosen option when wrong, Banking77 0.72 0.84
Jev vs Laya: accuracy and speedAG News with 4 options: Jev 86, Laya 96 correct out of 100. Banking77 with 77 options: Jev 78, Laya 41. Median time per question: Jev 523 milliseconds including the network trip from East Asia, Laya 34 milliseconds on a laptop GPU.Jev (API)Laya (local)AG NEWS · 4 OPTIONS · CORRECT OF 100Jev86Laya96BANKING77 · 77 OPTIONS · CORRECT OF 100Jev78Laya41MEDIAN TIME PER QUESTION (AG NEWS) · LOWER IS BETTERJev523 msLaya34 ms
Same questions and option descriptions for both. Jev time includes the network trip from East Asia; Laya ran on an M2 Max GPU.

What the numbers say:

  • Few, clearly different options: Laya is at least as good. On AG News, Laya got 12 items right that Jev missed, and Jev got 2 that Laya missed. Laya was also less sure of itself when it was wrong, which is what you want from a probability.
  • Many similar options: Jev is far ahead. On Banking77, Jev got 39 items right that Laya missed, and Laya got only 2 that Jev missed. Laya was also confidently wrong: its wrong answers averaged 0.84 probability. Laya’s own README explains why: all option descriptions share a fixed token budget, so with 77 options each label gets only a few tokens.
  • Speed depends on where the model runs. Jev’s times include a trip across the Pacific from our test machine; independent US-based tests cited in Laya’s README measured a median of about 240 to 280 ms. Laya runs next to your code, so there is no network at all.

These are small samples on two tasks with one wording, so treat them as a direction rather than a ranking. They match the pattern in Laya’s own README, which reports Laya ahead on AG News and on emotion labels, and Jev far ahead on Banking77. Our full scripts, samples and raw outputs are kept in our test records.

Laya’s README compares it with Jev on several public and synthetic tasks, and is unusually open about the limits:

  • Its Jev numbers come from third-party published runs, not its own Jev calls, so prompts and sample sizes differ.
  • The headline typed-decisions score (0.766 against Jev’s 0.727) comes from a separately fine-tuned checkpoint. The README says the base checkpoints score below a majority-class baseline on that benchmark, and that all of the capability there comes from fine-tuning.
  • Its best calibration figure (0.081 expected calibration error) is after fitting temperatures on your domain; as shipped, the README calls the checkpoints over-confident.
  • It lists where Jev leads: more than about 20 options, matching full probability distributions, and raw calibration before any fitting.

Laya ships a server, laya-serve, that speaks the same POST /v1/systemone protocol as TypeSafe’s API. Existing Jev code can call it by changing the base URL, for example with the official Python SDK’s base_url setting. Three things differ when you move code across, according to Laya’s documentation:

  1. Options per question share a token budget, so long or numerous option descriptions get cut. Jev accepts up to 255 options.
  2. Score levels all need descriptions; a null level is rejected with 422.
  3. confidence is calculated differently, so a threshold tuned on Jev does not transfer. Laya suggests gating on its answer_confidence field instead.

Because both accept the same request, running your own labeled examples through each is the fastest way to decide. See Your first Jev call for the request format.

Your situation Better fit
A few well-separated options (topic, sentiment, yes/no checks) at high volume Laya, if you can host it
Dozens of similar labels, such as support intents or product categories Jev
Long documents, logs or conversations as the state Jev (32k tokens against Laya’s 512 to 8,192)
Data that must not leave your servers Laya
Non-English text Try Laya’s multilingual checkpoint; Jev is strongest in English
You have labeled data from your own domain Either: fine-tune Laya, or use the data to set thresholds on Jev
No machine learning infrastructure at all Jev

Laya is one of more than a dozen open-source Jev alternatives; the others, including Kev, SemIf and Bespoke Nimble, are compared on Jev alternatives. For Jev’s own weak spots, see Jev limitations.

Sources

  1. Laya on GitHub (README, benchmarks, Jev-compatible server)
  2. Laya model card (Hugging Face)
  3. I built non-autoregressive decision models with RL a year ago (Hacker News) (Sep 19, 2026)
  4. Models: pricing and limits (TypeSafe docs)
  5. AG News dataset (Hugging Face)
  6. Banking77 dataset (PolyAI on GitHub)