Workbench

Email SaaS · lead qualification

50 labeled examples · Binary decision · Interactive demo

Use your own data
You’re exploring an example. Scores are illustrative, not a live Jev benchmark. Every metric below is calculated from these scores.
WHAT HAPPENED ON THIS TEST SET

5 examples need a closer look

45/ 50MATCHED YOUR LABELS
45 matched 5 disagreed

A strong score was wrong here. Moving the line alone may not resolve it.

These counts describe only this test set.

WHAT SHOULD I DO?

Review the decision question and rerun these same examples.

Decision tested“Is this person currently experiencing an email marketing or deliverability problem our product could solve?”
THE DECISION LINE

See where the line falls

Scores left of the line become NO. Scores on or right become YES.

Some score overlap
Why 0.61?

22 of 25 expected YES scored at or above 0.61. 23 of 25 expected NO scored below it.

JevBench chose this line because it best balanced caught YES examples and mistakes on this set.

Some expected YES and NO scores overlap, so moving the line may trade one kind of mistake for another.

EXPECTED YES
22 caught 3 missed
EXPECTED NO
23 correctly rejected 2 wrong YES

Choose your tradeoff

Move the line. Watch the two kinds of mistakes change.

← Catch more YESMore wrong YES may get through.Reduce wrong YES →More real YES may be missed.
WRONG YES2Jev said YES when you expected NO.
MISSED YES3Jev said NO when you expected YES.
Advanced metrics & recommendation method
Precision 92%Recall 88%F1 90%Accuracy 90%TP / FP 22 / 2TN / FN 23 / 3

Evaluated: 50/50. Precision is the share of Jev's YES decisions that matched your labels. Recall is the share of expected YES examples it caught.

Balanced maximizes F1, then recall, then prefers a cutoff near 0.50. Fewer wrong alerts maximizes precision among cutoffs with at least 5 accepted examples and 20% recall. Catch more positives maximizes recall among cutoffs with at least 50% precision. We test 0.05–0.95 in steps of 0.05, endpoints, and observed score breakpoints. A score equal to the cutoff counts as YES.

These are estimates on the same data used to choose the threshold. They are not guarantees of future accuracy. Observed score breakpoints are included so the recommendation is not limited to round numbers.

THE MOST USEFUL PART

See where Jev gets it wrong

Every mistake is a clue to a better decision.

5 mistakes
WRONG YESA customer asked us to write a guide about spam.
EXPECTEDNO
JEV SAIDYES
YES SCORE0.64
YOUR LINE0.61

Jev scored this 0.64. Your 0.61 line put it on the YES side, 0.03 away from the line.

Raising the line enough to correct this also flips 1 currently correct example.

MISSED YESI wish our email reports made more sense.
EXPECTEDYES
JEV SAIDNO
YES SCORE0.43
YOUR LINE0.61

Jev scored this 0.43. Your 0.61 line put it on the NO side, 0.18 away from the line.

Lowering the line enough to correct this also flips 3 currently correct examples.

MISSED YESIs there a way to stop paying for unsubscribed contacts?
EXPECTEDYES
JEV SAIDNO
YES SCORE0.35
YOUR LINE0.61

Jev scored this 0.35. Your 0.61 line put it on the NO side, 0.26 away from the line.

Lowering the line enough to correct this also flips 3 currently correct examples.

WRONG YESI am building an alternative to Mailchimp.
EXPECTEDNO
JEV SAIDYES
YES SCORE0.91
YOUR LINE0.61

Jev scored this 0.91. Your 0.61 line put it on the YES side, 0.30 away from the line.

Strong-score mistake: Jev scored this far beyond the line. This is a score-based flag, not calibrated certainty.

Raising the line enough to correct this also flips 18 currently correct examples.

MISSED YESEmail is becoming a headache for our team.
EXPECTEDYES
JEV SAIDNO
YES SCORE0.18
YOUR LINE0.61

Jev scored this 0.18. Your 0.61 line put it on the NO side, 0.43 away from the line.

Lowering the line enough to correct this also flips 13 currently correct examples.

The score determines which side of the line Jev chooses. Your labels tell JevBench whether that choice matched.

Do the scores separate?

The two answer groups share part of the score range.

Expected YES Expected NO
SOME OVERLAPA single line may trade wrong YES decisions for missed YES examples.
0.00 · More likely NO0.501.00 · More likely YES
Read score counts as a table
Score intervalYESNO
0.00–0.0502
0.05–0.1003
0.10–0.1503
0.15–0.2013
0.20–0.2503
0.25–0.3003
0.30–0.3503
0.35–0.4010
0.40–0.4510
0.45–0.5001
0.50–0.5500
0.55–0.6001
0.60–0.6512
0.65–0.7020
0.70–0.7510
0.75–0.8030
0.80–0.8540
0.85–0.9050
0.90–0.9541
0.95–1.0020

Found the tradeoff that works for you?

Export these results, then validate your cutoff on a fresh set of examples.

Read the threshold guide