What would you need to trust Jev with failure triage?

Hi everyone,

Jev from TypeSafe AI went into early access on September 15. It doesn’t write text. It answers typed questions (pick an option, give a score, or true/false) and returns probabilities.

That makes it a natural fit for test management calls like triaging failed runs, routing defects, or scoring severity. But a few things give us pause:

  • The benchmarks so far are vendor-run.
  • It can’t return malformed output, but it can still pick the wrong answer with confidence.
  • There’s no reasoning trace, which is hard to defend in a release review.

Over to you:

  1. Would you let AI label a failure without a person checking it? What confidence would you need?
  2. What would you need logged to trust it in an audit?
  3. If you piloted it, which decision would you start with?

We’re writing about Jev for QA and want to hear what practitioners think first.

No, I would not trust an AI model specially Jev to handle failure triage completely without human oversight right out of the box—especially in a strict release review environment.