Is Jev Deterministic? What Repeat Calls Show
Last checked · Independent guide, not affiliated with TypeSafe AI
Not strictly. In our tests, identical requests returned the same choice every time, but probabilities could move by a few hundredths: one borderline yes/no answer ranged from 0.54 to 0.60 over 20 identical calls. TypeSafe says it designs Jev for consistency (similar answers for similar meaning) rather than determinism. Wording changes the answer far more than repetition does.
Deterministic means the same input always gives exactly the same output. Jev does not promise that, and in our tests it did not quite deliver it. But the variation is small, and it is not the variation you should worry about most.
What TypeSafe says
Section titled “What TypeSafe says”TypeSafe answers this question in the FAQ on its homepage. It separates determinism, returning the identical result for an identical input, from consistency, making similar decisions when the meaning stays the same even if the wording changes. It says the second matters more, and that “Jev is designed for consistency.” It does not claim determinism.
Our repeat-call tests
Section titled “Our repeat-call tests”On September 25, 2026 we sent requests to TypeSafe’s API with the model pinned to jev-1.13.0, repeating each one exactly.
| Test | Result |
|---|---|
| Clear message, 3 questions, sent 20 times | Same choice (billing, probability 1.0) and same yes/no value (0.99) every time. The Score came back as 1.99 in 15 runs and 2.0 in 5 |
| Borderline yes/no question, sent 20 times | Ranged from 0.54 to 0.60; most often 0.57 or 0.58 |
| The same borderline question three times in one request | 0.58, 0.57 and 0.59 |
| Choice on the borderline message, sent 20 times | Same choice and same probabilities every time |
| Question order reversed within the request | No change |
| Choice options listed in reverse order, 5 runs each way | Small, consistent shift: the winning option moved from 0.95–0.96 to 0.96–0.97 |
jev-latest, jev-preview and jev-1.13.0 |
0.56, 0.57 and 0.59, inside the same noise band; all three currently resolve to jev-1.13.0 |
| The borderline question worded three ways | 0.56, 0.66 and 0.26 |
Probabilities come back rounded to two decimal places, so a difference of 0.01 may be just a rounding boundary. Differences of 0.06, as in the borderline question, are not.
jev-1.13.0, Sep 25, 2026. Wordings: A "need a reply today?", B "answer before the end of today?", C "same-day response required?"What others found
Section titled “What others found”- “Jev’s Architecture Unmasked”, an independent analysis based on thousands of API calls, also saw small differences between repeated identical requests, including between duplicate questions in one request. It suggests ordinary causes such as numerical kernels, dynamic batching or routing between servers, and notes this does not mean Jev samples text.
- A Chinese classification test posted to TypeSafe’s GitHub ran 130 items three times and got zero changed decisions: 127 of 130 correct in each run. At the level of the final choice, that matches what we saw.
- “64 Tiny Benchmarks for Jev” prompted a Hacker News commenter to complain that Jev “isn’t even slightly deterministic”. In practice the choices were stable; the probabilities were not bit-for-bit identical.
Why the wording matters more
Section titled “Why the wording matters more”The largest swing in our tests did not come from repeating a request. It came from rephrasing it. “Does this message need a reply today?” scored 0.56, “Should someone answer this message before the end of today?” scored 0.66, and “Is a same-day response required for this message?” scored 0.26. The third wording asks a stricter question (“required”), and Jev treated it that way.
That is TypeSafe’s consistency argument seen from the other side: Jev responds to meaning, and small changes of meaning in your instructions move the answer much more than run-to-run noise does. See Noul: wording changes the answer.
How to build around it
Section titled “How to build around it”- Don’t put a threshold right where answers land. If a decision flips between 0.49 and 0.51, repeat calls can flip it. Leave a band around the threshold and send those cases to review, as TypeSafe’s confidence-gated routing pattern suggests. See Confidence.
- Freeze your wording. Treat instructions and option descriptions like code: keep them in version control and re-test when you change them.
- Pin the version. Use
jev-1.13.0rather thanjev-latestso a new model release does not change answers under you. See Jev changelog. - Cache when you need reproducibility. If an audit trail must show the same answer twice, store the response (with its
modelfield and request ID) instead of calling again. - Average duplicates for borderline calls. Asking the same question two or three times in one request costs only a few extra tokens, because the state is paid for once, and the average is steadier than a single value.
Is an LLM more deterministic?
Section titled “Is an LLM more deterministic?”Not usually. Even with the temperature set to zero, hosted LLM APIs rarely guarantee identical output, for the same batching and hardware reasons. The difference is that an LLM’s variation shows up as different text you have to parse, while Jev’s shows up as a slightly different number in a fixed format. For the underlying trade-offs, see Jev vs LLMs.
Sources
- TypeSafe AI homepage FAQ: Is Jev deterministic?
- Jev's Architecture Unmasked: repeat-call noise (archerhume) (Sep 17, 2026)
- Independent evaluation on Chinese classification: 3 runs, zero flips (typesafe-ai/skills #3) (Sep 19, 2026)
- 64 Tiny Benchmarks for Jev (Hacker News discussion) (Sep 20, 2026)
- Models and aliases (TypeSafe docs)