ONE SAMPLE · 29 SEP 2026

The problem was not the buying decision.

Fifty texts were labeled twice. First: does this person currently have a problem the product could solve? Second: are they looking to buy?

The labels changed

Current problem: 22 YES, 25 NO, 3 unclear. Buying intent: 10 YES, 34 NO, 6 unclear. 16 of the 50 texts changed.

Two of those texts, in plain language:

The scores stayed the same

Jev was asked the problem question. At a line of 0.50, those scores against the problem labels were 0 false YES and 3 missed YES. The same scores against the buying-intent labels were 8 false YES and 4 missed YES.

Moving the line cannot turn the problem question into the buying question. The texts and the scores did not change. The human decision did.

What this does not show

This is a small, constructed set. It is not a universal Jev benchmark. It is not a measure of production accuracy. It does not prove that a question and a label disagree in every project. One medical-coding workflow found the same kind of mismatch independently. That is one more case, not a survey.

JevBench does not automatically say whether a disagreement comes from the question, the label, missing context, or the model.

Open the related Playground example or try your own decision.