Has anyone wired Jev into their automation suite yet?

Hey all,

A few people have already put Jev, TypeSafe AI’s new decision model, into test loops. Jason Arbon saw a 131 ms median round trip in a Playwright loop, plus a 0.49 vs 0.47 near-tie and a false flag on a normal add-to-bag step.

The fits we keep seeing: picking the next action in an exploratory agent, “did the page move forward?” checks next to hard assertions, and judging UI text a string match can’t (“does this error tell the user how to fix it?”). TypeSafe’s own docs say it’s weak at math, dates, and anything visual.

Curious where you land:

  1. Where in your suite would a probability beat a hard assertion, if anywhere?
  2. Would you run a semantic check as a hard fail, a warning, or shadow mode only?
  3. Has anyone tried it against a labeled set of their own past failures?

For transparency: Katalon team here. There’s no Jev integration in Studio and no partnership with TypeSafe. We just want to know where it’s useful for testers.

Its again back to invitation only , have been trying to get it since a few days now

not wired it in yet, still on the waitlist like Monty. but on the question itself, the only place i would let a probability win over a hard assertion is the “soft” stuff we currently don’t assert at all. error message quality, does the empty state make sense, is the toast telling the user what to do. today those are either skipped or a brittle string match that breaks on every copy change.

i would never run it as a hard fail at first. shadow mode for a few weeks, log the score next to the normal pass/fail, then look at where it disagreed with us. if the disagreements are mostly “it was right and we missed it” then promote it to a warning. hard fail only for checks where a wrong flag costs less than a missed bug.

and yes a labeled set of our own past failures is the only benchmark i would trust. vendor numbers on someone else’s app don’t tell me much about mine.