Hi everyone,
Jev from TypeSafe AI went into early access on September 15. It doesn’t write text. It answers typed questions (pick an option, give a score, or true/false) and returns probabilities.
That makes it a natural fit for test management calls like triaging failed runs, routing defects, or scoring severity. But a few things give us pause:
- The benchmarks so far are vendor-run.
- It can’t return malformed output, but it can still pick the wrong answer with confidence.
- There’s no reasoning trace, which is hard to defend in a release review.
Over to you:
- Would you let AI label a failure without a person checking it? What confidence would you need?
- What would you need logged to trust it in an audit?
- If you piloted it, which decision would you start with?
We’re writing about Jev for QA and want to hear what practitioners think first.