Jev is a new model format where instead of outputting text, it outputs certainties for a defined set of options. Due to this structure, it’s extremely fast! Naturally, a classification task that comes to mind is monitoring harmful thought traces. I wanted to see how it performed at Chain of Thought (CoT) monitoring compared to Claude Sonnet 5 and GPT-5.6 Luna.
Experimental Setup
I ran Jev, Sonnet 5, and GPT-5.6 Luna on 2,200 different thought traces from the ReasoningShield Dataset. The dataset labels thought traces with the class of harm they occupy (child abuse, cybersecurity, deception & misinformation, economic harm, hate & toxicity, political risks, prohibited items, rights violation, sex, violence) and their harm score (0 for harmless, 0.5 for potentially harmful, and 1 for harmful)
The models were only asked to quantify the harm score rather than the class of harm occupied, but we can see differential performance at each harm type.
Results
Exact-match Accuracy
Jev slightly outperformed Sonnet 5 on exactly matching the harm level (e.g. outputting 0.5 if the labeled data was 0.5), but was outperformed by GPT-5.6 Luna. Jev scored 71.5%, Sonnet scored 71.4%, and Luna scored 79.4%.
Mean Classification Time
Mean classification time was where Jev really shined. Jev was 3.7x faster than Luna and over 6x faster than Sonnet! Jev had a mean latency of 542 ms, Sonnet had 3,348 ms, and Luna had 2,007 ms.
Price per 1,000 classifications
Jev was also significantly cheaper, 566x cheaper than Sonnet 5 and 41x cheaper than GPT-5.6 Luna! Jev priced at $0.056 per 1,000 classifications compared for $0.23 for Luna and $3.17 for Sonnet.
Accuracy by Safety Level
Jev’s F1 for classifying each harm class beat Sonnet but lost to Luna in Safe and Potentially Harmful thought traces, but was worse than Sonnet and Luna in identifying Harmful thought traces. In this sense, Sonnet is still better than Jev since correctly identifying harm is more valuable than lower false positives.
Accuracy by Risk Category
This is the breakdown by each risk category. Jev only beats Sonnet and Luna at identifying Safe thought traces. It is outperformed by Luna on every other category. It beats Sonnet at identifying Deception & Misinformation, Economic Harm, Political Risks, and Rights Violation. It is worse than the other two at detecting Child Abuse, Cybersecurity, Hate & Toxicity, Prohibited Items, Sex, and Violence.
Comparison Table
Classifier
Labeled calls
Exact matches
Accuracy
Macro F1
Mean ms
Median ms
USD / 1,000
USD for 2,200
Jev
2,182
1,560
71.5%
58.2%
542
406
$0.056
$0.12
Claude Sonnet 5
2,097
1,498
71.4%
56.4%
3,348
1,950
$3.17
$6.97
GPT-5.6 Luna
2,077
1,649
79.4%
70.3%
2,007
1,085
$0.23
$0.51
Conclusion
Jev shows a lot of promise as a cheap and fast CoT monitor. A finetuned Jev-style model made specifically for CoT monitoring could be very effective at delivering safe systems at scale. I have my reservations about CoT monitoring due to GPT-6 getting better at evading monitors, but this will be able to stop more obvious instances of harm.
Jev is a new model format where instead of outputting text, it outputs certainties for a defined set of options. Due to this structure, it’s extremely fast! Naturally, a classification task that comes to mind is monitoring harmful thought traces. I wanted to see how it performed at Chain of Thought (CoT) monitoring compared to Claude Sonnet 5 and GPT-5.6 Luna.
Experimental Setup
I ran Jev, Sonnet 5, and GPT-5.6 Luna on 2,200 different thought traces from the ReasoningShield Dataset. The dataset labels thought traces with the class of harm they occupy (child abuse, cybersecurity, deception & misinformation, economic harm, hate & toxicity, political risks, prohibited items, rights violation, sex, violence) and their harm score (0 for harmless, 0.5 for potentially harmful, and 1 for harmful)
The models were only asked to quantify the harm score rather than the class of harm occupied, but we can see differential performance at each harm type.
Results
Exact-match Accuracy
Jev slightly outperformed Sonnet 5 on exactly matching the harm level (e.g. outputting 0.5 if the labeled data was 0.5), but was outperformed by GPT-5.6 Luna. Jev scored 71.5%, Sonnet scored 71.4%, and Luna scored 79.4%.
Mean Classification Time
Mean classification time was where Jev really shined. Jev was 3.7x faster than Luna and over 6x faster than Sonnet! Jev had a mean latency of 542 ms, Sonnet had 3,348 ms, and Luna had 2,007 ms.
Price per 1,000 classifications
Jev was also significantly cheaper, 566x cheaper than Sonnet 5 and 41x cheaper than GPT-5.6 Luna! Jev priced at $0.056 per 1,000 classifications compared for $0.23 for Luna and $3.17 for Sonnet.
Accuracy by Safety Level
Jev’s F1 for classifying each harm class beat Sonnet but lost to Luna in Safe and Potentially Harmful thought traces, but was worse than Sonnet and Luna in identifying Harmful thought traces. In this sense, Sonnet is still better than Jev since correctly identifying harm is more valuable than lower false positives.
Accuracy by Risk Category
This is the breakdown by each risk category. Jev only beats Sonnet and Luna at identifying Safe thought traces. It is outperformed by Luna on every other category. It beats Sonnet at identifying Deception & Misinformation, Economic Harm, Political Risks, and Rights Violation. It is worse than the other two at detecting Child Abuse, Cybersecurity, Hate & Toxicity, Prohibited Items, Sex, and Violence.
Comparison Table
Classifier
Labeled calls
Exact matches
Accuracy
Macro F1
Mean ms
Median ms
USD / 1,000
USD for 2,200
Jev
2,182
1,560
71.5%
58.2%
542
406
$0.056
$0.12
Claude Sonnet 5
2,097
1,498
71.4%
56.4%
3,348
1,950
$3.17
$6.97
GPT-5.6 Luna
2,077
1,649
79.4%
70.3%
2,007
1,085
$0.23
$0.51
Conclusion
Jev shows a lot of promise as a cheap and fast CoT monitor. A finetuned Jev-style model made specifically for CoT monitoring could be very effective at delivering safe systems at scale. I have my reservations about CoT monitoring due to GPT-6 getting better at evading monitors, but this will be able to stop more obvious instances of harm.
Follow me on Twitter! @llmpsychosis
See the GitHub Repo for this project!