Skip to content

Is Jev Deterministic? What Repeat Calls Show

Last checked · Independent guide, not affiliated with TypeSafe AI

ANSWER

Not strictly. In our tests, identical requests returned the same choice every time, but probabilities could move by a few hundredths: one borderline yes/no answer ranged from 0.54 to 0.60 over 20 identical calls. TypeSafe says it designs Jev for consistency (similar answers for similar meaning) rather than determinism. Wording changes the answer far more than repetition does.

Deterministic means the same input always gives exactly the same output. Jev does not promise that, and in our tests it did not quite deliver it. But the variation is small, and it is not the variation you should worry about most.

TypeSafe answers this question in the FAQ on its homepage. It separates determinism, returning the identical result for an identical input, from consistency, making similar decisions when the meaning stays the same even if the wording changes. It says the second matters more, and that “Jev is designed for consistency.” It does not claim determinism.

On September 25, 2026 we sent requests to TypeSafe’s API with the model pinned to jev-1.13.0, repeating each one exactly.

Test Result
Clear message, 3 questions, sent 20 times Same choice (billing, probability 1.0) and same yes/no value (0.99) every time. The Score came back as 1.99 in 15 runs and 2.0 in 5
Borderline yes/no question, sent 20 times Ranged from 0.54 to 0.60; most often 0.57 or 0.58
The same borderline question three times in one request 0.58, 0.57 and 0.59
Choice on the borderline message, sent 20 times Same choice and same probabilities every time
Question order reversed within the request No change
Choice options listed in reverse order, 5 runs each way Small, consistent shift: the winning option moved from 0.95–0.96 to 0.96–0.97
jev-latest, jev-preview and jev-1.13.0 0.56, 0.57 and 0.59, inside the same noise band; all three currently resolve to jev-1.13.0
The borderline question worded three ways 0.56, 0.66 and 0.26

Probabilities come back rounded to two decimal places, so a difference of 0.01 may be just a rounding boundary. Differences of 0.06, as in the borderline question, are not.

Repeating a request versus rewording itThe same borderline yes/no request sent 20 times returned values between 0.54 and 0.60. The same question asked three times in one request returned 0.57 to 0.59. The same question in three wordings returned 0.26, 0.56 and 0.66.0.000.250.500.751.000.5 THRESHOLDSame request, sent 20 times0.58: 5 of 20 runs0.56: 4 of 20 runs0.57: 5 of 20 runs0.54: 2 of 20 runs0.59: 2 of 20 runs0.60: 1 of 20 runs0.55: 1 of 20 runs0.54 – 0.60 (bar height = runs)Same question 3× in one request0.57 – 0.59Same question, 3 wordingsCAB
Borderline yes/no question, jev-1.13.0, Sep 25, 2026. Wordings: A "need a reply today?", B "answer before the end of today?", C "same-day response required?"
  • “Jev’s Architecture Unmasked”, an independent analysis based on thousands of API calls, also saw small differences between repeated identical requests, including between duplicate questions in one request. It suggests ordinary causes such as numerical kernels, dynamic batching or routing between servers, and notes this does not mean Jev samples text.
  • A Chinese classification test posted to TypeSafe’s GitHub ran 130 items three times and got zero changed decisions: 127 of 130 correct in each run. At the level of the final choice, that matches what we saw.
  • “64 Tiny Benchmarks for Jev” prompted a Hacker News commenter to complain that Jev “isn’t even slightly deterministic”. In practice the choices were stable; the probabilities were not bit-for-bit identical.

The largest swing in our tests did not come from repeating a request. It came from rephrasing it. “Does this message need a reply today?” scored 0.56, “Should someone answer this message before the end of today?” scored 0.66, and “Is a same-day response required for this message?” scored 0.26. The third wording asks a stricter question (“required”), and Jev treated it that way.

That is TypeSafe’s consistency argument seen from the other side: Jev responds to meaning, and small changes of meaning in your instructions move the answer much more than run-to-run noise does. See Noul: wording changes the answer.

  • Don’t put a threshold right where answers land. If a decision flips between 0.49 and 0.51, repeat calls can flip it. Leave a band around the threshold and send those cases to review, as TypeSafe’s confidence-gated routing pattern suggests. See Confidence.
  • Freeze your wording. Treat instructions and option descriptions like code: keep them in version control and re-test when you change them.
  • Pin the version. Use jev-1.13.0 rather than jev-latest so a new model release does not change answers under you. See Jev changelog.
  • Cache when you need reproducibility. If an audit trail must show the same answer twice, store the response (with its model field and request ID) instead of calling again.
  • Average duplicates for borderline calls. Asking the same question two or three times in one request costs only a few extra tokens, because the state is paid for once, and the average is steadier than a single value.

Not usually. Even with the temperature set to zero, hosted LLM APIs rarely guarantee identical output, for the same batching and hardware reasons. The difference is that an LLM’s variation shows up as different text you have to parse, while Jev’s shows up as a slightly different number in a fixed format. For the underlying trade-offs, see Jev vs LLMs.

Sources

  1. TypeSafe AI homepage FAQ: Is Jev deterministic?
  2. Jev's Architecture Unmasked: repeat-call noise (archerhume) (Sep 17, 2026)
  3. Independent evaluation on Chinese classification: 3 runs, zero flips (typesafe-ai/skills #3) (Sep 19, 2026)
  4. 64 Tiny Benchmarks for Jev (Hacker News discussion) (Sep 20, 2026)
  5. Models and aliases (TypeSafe docs)