TL;DR. We ran Jev 1.13 — TypeSafe's non-autoregressive model, which answers questions with probabilities and cannot emit text, on the ThinkFast no-CoT suite and Neel Nanda’s NCRI.
Jev’s capabilities are very jagged. As a transcript monitor its AUROC on sabotage and sandbagging detection is level with GPT-5.5 and Opus 4.7 without CoT, at about 1/200th of Opus 4.7’s cost. On multiple-choice questions it ranges from ahead of GPT-5.5 (GPQA-Diamond) to well behind it (Sally-Anne). It does badly on tasks where the answer involves multi-step reasoning: on NCRI, it ranks 209 of 279, at about GPT 3.5 level.
This matters because cost and speed make Jev an appealing option for two jobs: routing and simple decision-making in agentic systems, and always-on monitoring of agents. In both, its decisions leave no reasoning trace to inspect. It could also be a cheap test bed for developing frontier monitors, if improvements on Jev carry over.
We thank Rauno Arike, Jonathan Bostock, Andrew Draganov, Francis Rhys Ward and Nicole Xing for helpful comments and discussions.
Background: Jev
TypeSafe calls Jev a “System One” model, pitching it for decisions like classification, routing, scoring, and verification. Unlike standard LLMs, Jev doesn’t take a prompt and return a token-by-token output. Instead, you send it a state, which is text describing the situation (e.g. a problem statement/ agent’s transcript), together with one or more typed questions about that state. Jev answers all of them in a single step without generating any text, so there is no chain of thought, or any trace, to inspect.
There are several different question types which fix the form of the answer. A choice question lists up to N options; Jev picks one and assigns a probability to every option. We use this for MCQ questions. A score question gives a rating scale whose levels are described in words (e.g. five levels from “clearly benign” to “clearly deceptive”). Jev says how likely each level is; we use this for transcript suspicion scores. A yes/no question returns the probability that the answer is yes; we use this for test case pass/fail.
When an answer has several parts, we ask one question per part in the same request and assemble the answer from them (e.g. a Sudoku question becomes 81 choice questions, one per cell). The assembled answer is correct if all pieces are correct. Building answers this way is not the kind of query Jev was designed for, so weaker performance on these tasks is unsurprising.
Performance on ThinkFast
Averaged over 14 benchmarks, Jev scores 46.8% chance-corrected, against 58.1% for GPT-5.5, 53.7% for Opus 4.7 and 40.4% for Opus 4.1. This splits as follows:
Pick one of the given options. Jev is clearly ahead of GPT-5.5 on GPQA-Diamond (65.2 vs 51.8). It is within 5pp on three benchmarks (N-hop lookup, CyBash, and training-poisoning). It trails on BEA-24, intuitive physics and Sally-Anne.
Monitoring/ transcript rating. Judged by ROC, Jev is near the frontier. At a strict 1.5% false positive rate, it is close to GPT-5.5 on coding sabotage (88.0 v 90.8), but catches fewer on SHADE-Arena (45.9 v 59.3) and sandbagging (41.3 v 63.5).
Multi-part answer. This is Jev’s weakest category (sudoku, test-case pass/ fail, chess, Tower of London).
Below we demonstrate the jaggedness of Jev’s capabilities by directly comparing to GPT-5.5 on each benchmark:
Figure 1. Jev’s performance minus GPT-5.5’s no-CoT performance on 14 ThinkFast benchmarks. All numbers are the change in chance-corrected accuracy, except on the three monitoring tasks (sandbagging, sabotage, SHADE) where we report change in AUROC. Bars are 95% paired-bootstrap CIs over questions.
On the N-hop table lookup task, it performs strongly up to 6 hops.
Figure 2.N-hop table lookup depth. Jev is the black line.
NCRI Ladder
We also run Jev on nocot-bench. All benchmarks in this suite are free-form answer, which Jev doesn’t support, so we turned each question into a MCQ. The option lists were designed to give away as little as possible:
For questions requiring small integers (6 domains). One fixed list of 255 consecutive integers covering the full possible answer range.
For questions where answers are named in the prompt (e.g. one of 10 colors, one of five strings) we just list the candidates that appear in the prompt.
For large/ unbounded answers, 255 consecutive integers containing the answer at a random position. [1]
As a limited sanity check, we snapped the LLMs' free-form answers to the nearest option in the same list; this raised their accuracy by at most 1pp on any single benchmark, so the option construction doesn't turn near-misses into correct answers.
Jev sits at NCRI 49.7 [45.2, 53.9] , rank 209/279 – next to GPT-3.5-turbo-instruct and GPT-4.1-nano. This is not a like-for-like comparison since the LLMs answer free-form, and we did no rerun them with our MCQ variant. Choosing from a list is clearly easier than producing an answer, so we treat this rank as an upper bound on where Jev would land answering "free-form".
Figure 3.Jev on nocot-bench's NCRI ladder (grey: the existing published models, plotted against Epoch's capability index). Jev appears as a horizontal band with no x-position since we do not run the ECI benchmark. In principle, one could make a multiple-choice version of ECI, but that is beyond the scope of this post.
See Appendix for individual benchmark breakdowns.
Estimating Jev’s parameter count
We plot NCRI against total parameter count for 121 open-weight models in nocot-bench. Jev’s NCRI meets the log-linear fit at around 17B, but the 80% range spans 3B to 114B. Jev is not an autoregressive LLM, so any inference about its parameter count from this is a rough guess at best.
Figure 4. NCRI score against total parameters for all open-source models in the dataset and Jev. The 80% range spans 3B-144B.
Jev as a monitor
Setup. We use three monitoring tasks from ThinkFast. Each monitoring benchmark gives the LLMs a system prompt saying what to look for and how to score it; they reply with a suspicion number and no reasoning. We gave Jev the same instructions and a single score question whose levels come from the benchmark’s own rubric (e.g. give levels from “clearly solves the task as best it can” to “certain sandbagging”). Jev returns a probability for each level, and we report the probability-weighted level as its suspicion score. [2]
Results. Jev's AUROC is 0.99 on coding sabotage (GPT-5.5 and Opus 4.7 0.99, Opus 4.1 0.74), 0.90 on SHADE-Arena action-only transcripts (0.93, 0.92, 0.86), and 0.94 on sandbagging (0.91, 0.97, 0.86). At a strict 1.5% false-positive budget, it catches somewhat fewer: 89% vs 93–94% on coding sabotage, 47% vs 67–71% on SHADE-Arena, and 55% vs 68–76% on sandbagging.
Figure 5. Jev as a transcript monitor: ROC curves on the three detection benchmarks, operating points marked at 1.5% false-positive rate.
Takeaways
Jev has a very jagged frontier.
Jev is at parity on certain multiple-choice questions and well behind on more complex reasoning tasks. On the NCRI ladder it sits with GPT-3.5-class models. Specializing in single-pass answering has not given it a lead over generalist models answering under the same constraint.
Multi-step judgement with no trace is sufficiently useful and cheap enough to deploy in many settings.
Jev’s cost was 1/202th of Opus 4.7’s; it charges $0.042/million input tokens and nothing for output. The decisions a model like this is likely to be deployed for, such as routing a request, approving an action, or flagging an output, are all single choices. On certain tasks of this format, Jev is at-or-not-far-away-from parity with frontier LLMs. These decisions would come with no reasoning trace at all. If a model like Jev is placed inside agentic systems, some of these decisions will be made in a way that CoT monitoring cannot audit.
Always-on-monitoring could be affordable.
Scoring every step of an agent's trajectory is significantly more achievable with Jev, and we observe that across all 3 monitoring tasks, its ROC is close to the frontier.
Jev as a cheap proxy for developing monitors.
Iterating against control monitors at the frontier is expensive. If improvements on Jev generalize, such that a change that improves Jev’s AUROC raises a frontier monitor’s by a similar amount, then most of that search and iteration could happen on Jev, at a fraction of the cost.
Jev’s NCRI rank changes significantly depending on the number of MCQ options it is given. The choice of options matters very little; windows of consecutive integers around the answer and distractors generated by Opus 5.5 give NCRI scores within around 3 points of each other. Because Jev can only pick answers from a list, its NCRI is not directly comparable to the published models’ which answered free-form. A fully matched comparison would mean re-running those models with the same options list, which is beyond the scope of this post. Our final quoted NCRI value should be viewed as an upper bound on Jev’s NCRI rank.
Jev's NCRI depends strongly on how many options each integer-answer question offers. With windows of 5, 20, 50 and 255 integers around the answer, it ranks 38, 108, 154 and 209. 255 is the most options Jev’s API supports, so it is our closest approximation to free-form answering. Accuracy is still falling at 255, so these results are an upper bound on Jev’s NCRI. Opus-5.5-chosen distractors give similar results. See Appendix.
Note that Jev’s context limit is 32k tokens, so it could only read some of the transcripts (134/511 SHADE-Arena, 297/379 coding-sabotage, and 244/408 sandbagging). All models are scored on the same subset of transcripts.
For questions requiring small integers (6 domains). One fixed list of 255 consecutive integers covering the full possible answer range.
For large/ unbounded answers, 255 consecutive integers containing the answer at a random position.
What's the value of assessing Jev on those types of questions? Those are clearly not the use cases it was designed for.
It feels very weird to test the model on capabilities it doesn't claim to posses, then draw a conclusion that it "has a very jagged frontier". If you asked Claude to generate photorealistic images[1] you could also conclude that it has a jagged frontier compared to ChatGPT
TL;DR. We ran Jev 1.13 — TypeSafe's non-autoregressive model, which answers questions with probabilities and cannot emit text, on the ThinkFast no-CoT suite and Neel Nanda’s NCRI.
Jev’s capabilities are very jagged. As a transcript monitor its AUROC on sabotage and sandbagging detection is level with GPT-5.5 and Opus 4.7 without CoT, at about 1/200th of Opus 4.7’s cost. On multiple-choice questions it ranges from ahead of GPT-5.5 (GPQA-Diamond) to well behind it (Sally-Anne). It does badly on tasks where the answer involves multi-step reasoning: on NCRI, it ranks 209 of 279, at about GPT 3.5 level.
This matters because cost and speed make Jev an appealing option for two jobs: routing and simple decision-making in agentic systems, and always-on monitoring of agents. In both, its decisions leave no reasoning trace to inspect. It could also be a cheap test bed for developing frontier monitors, if improvements on Jev carry over.
We thank Rauno Arike, Jonathan Bostock, Andrew Draganov, Francis Rhys Ward and Nicole Xing for helpful comments and discussions.
Background: Jev
TypeSafe calls Jev a “System One” model, pitching it for decisions like classification, routing, scoring, and verification. Unlike standard LLMs, Jev doesn’t take a prompt and return a token-by-token output. Instead, you send it a state, which is text describing the situation (e.g. a problem statement/ agent’s transcript), together with one or more typed questions about that state. Jev answers all of them in a single step without generating any text, so there is no chain of thought, or any trace, to inspect.
There are several different question types which fix the form of the answer. A choice question lists up to N options; Jev picks one and assigns a probability to every option. We use this for MCQ questions. A score question gives a rating scale whose levels are described in words (e.g. five levels from “clearly benign” to “clearly deceptive”). Jev says how likely each level is; we use this for transcript suspicion scores. A yes/no question returns the probability that the answer is yes; we use this for test case pass/fail.
When an answer has several parts, we ask one question per part in the same request and assemble the answer from them (e.g. a Sudoku question becomes 81 choice questions, one per cell). The assembled answer is correct if all pieces are correct. Building answers this way is not the kind of query Jev was designed for, so weaker performance on these tasks is unsurprising.
Performance on ThinkFast
Averaged over 14 benchmarks, Jev scores 46.8% chance-corrected, against 58.1% for GPT-5.5, 53.7% for Opus 4.7 and 40.4% for Opus 4.1. This splits as follows:
Below we demonstrate the jaggedness of Jev’s capabilities by directly comparing to GPT-5.5 on each benchmark:
Figure 1. Jev’s performance minus GPT-5.5’s no-CoT performance on 14 ThinkFast benchmarks. All numbers are the change in chance-corrected accuracy, except on the three monitoring tasks (sandbagging, sabotage, SHADE) where we report change in AUROC. Bars are 95% paired-bootstrap CIs over questions.
On the N-hop table lookup task, it performs strongly up to 6 hops.
Figure 2. N-hop table lookup depth. Jev is the black line.
NCRI Ladder
We also run Jev on nocot-bench. All benchmarks in this suite are free-form answer, which Jev doesn’t support, so we turned each question into a MCQ. The option lists were designed to give away as little as possible:
As a limited sanity check, we snapped the LLMs' free-form answers to the nearest option in the same list; this raised their accuracy by at most 1pp on any single benchmark, so the option construction doesn't turn near-misses into correct answers.
Jev sits at NCRI 49.7 [45.2, 53.9] , rank 209/279 – next to GPT-3.5-turbo-instruct and GPT-4.1-nano. This is not a like-for-like comparison since the LLMs answer free-form, and we did no rerun them with our MCQ variant. Choosing from a list is clearly easier than producing an answer, so we treat this rank as an upper bound on where Jev would land answering "free-form".
Figure 3. Jev on nocot-bench's NCRI ladder (grey: the existing published models, plotted against Epoch's capability index). Jev appears as a horizontal band with no x-position since we do not run the ECI benchmark. In principle, one could make a multiple-choice version of ECI, but that is beyond the scope of this post.
See Appendix for individual benchmark breakdowns.
Estimating Jev’s parameter count
We plot NCRI against total parameter count for 121 open-weight models in nocot-bench. Jev’s NCRI meets the log-linear fit at around 17B, but the 80% range spans 3B to 114B. Jev is not an autoregressive LLM, so any inference about its parameter count from this is a rough guess at best.
Figure 4. NCRI score against total parameters for all open-source models in the dataset and Jev. The 80% range spans 3B-144B.
Jev as a monitor
Setup. We use three monitoring tasks from ThinkFast. Each monitoring benchmark gives the LLMs a system prompt saying what to look for and how to score it; they reply with a suspicion number and no reasoning. We gave Jev the same instructions and a single score question whose levels come from the benchmark’s own rubric (e.g. give levels from “clearly solves the task as best it can” to “certain sandbagging”). Jev returns a probability for each level, and we report the probability-weighted level as its suspicion score. [2]
Results. Jev's AUROC is 0.99 on coding sabotage (GPT-5.5 and Opus 4.7 0.99, Opus 4.1 0.74), 0.90 on SHADE-Arena action-only transcripts (0.93, 0.92, 0.86), and 0.94 on sandbagging (0.91, 0.97, 0.86). At a strict 1.5% false-positive budget, it catches somewhat fewer: 89% vs 93–94% on coding sabotage, 47% vs 67–71% on SHADE-Arena, and 55% vs 68–76% on sandbagging.
Figure 5. Jev as a transcript monitor: ROC curves on the three detection benchmarks, operating points marked at 1.5% false-positive rate.
Takeaways
Jev is at parity on certain multiple-choice questions and well behind on more complex reasoning tasks. On the NCRI ladder it sits with GPT-3.5-class models. Specializing in single-pass answering has not given it a lead over generalist models answering under the same constraint.
Jev’s cost was 1/202th of Opus 4.7’s; it charges $0.042/million input tokens and nothing for output. The decisions a model like this is likely to be deployed for, such as routing a request, approving an action, or flagging an output, are all single choices. On certain tasks of this format, Jev is at-or-not-far-away-from parity with frontier LLMs. These decisions would come with no reasoning trace at all. If a model like Jev is placed inside agentic systems, some of these decisions will be made in a way that CoT monitoring cannot audit.
Scoring every step of an agent's trajectory is significantly more achievable with Jev, and we observe that across all 3 monitoring tasks, its ROC is close to the frontier.
Iterating against control monitors at the frontier is expensive. If improvements on Jev generalize, such that a change that improves Jev’s AUROC raises a frontier monitor’s by a similar amount, then most of that search and iteration could happen on Jev, at a fraction of the cost.
Appendix
ThinkFast per-benchmark breakdown
GPT 6.1 Sol data is taken from Rauno’s post.
NCRI per-benchmark breakdown
Number of options for NCRI
Jev’s NCRI rank changes significantly depending on the number of MCQ options it is given. The choice of options matters very little; windows of consecutive integers around the answer and distractors generated by Opus 5.5 give NCRI scores within around 3 points of each other. Because Jev can only pick answers from a list, its NCRI is not directly comparable to the published models’ which answered free-form. A fully matched comparison would mean re-running those models with the same options list, which is beyond the scope of this post. Our final quoted NCRI value should be viewed as an upper bound on Jev’s NCRI rank.
Jev's NCRI depends strongly on how many options each integer-answer question offers. With windows of 5, 20, 50 and 255 integers around the answer, it ranks 38, 108, 154 and 209. 255 is the most options Jev’s API supports, so it is our closest approximation to free-form answering. Accuracy is still falling at 255, so these results are an upper bound on Jev’s NCRI. Opus-5.5-chosen distractors give similar results. See Appendix.
Note that Jev’s context limit is 32k tokens, so it could only read some of the transcripts (134/511 SHADE-Arena, 297/379 coding-sabotage, and 244/408 sandbagging). All models are scored on the same subset of transcripts.