decision.host
首页 / 模型 / Jev 1.13

Jev 1.13

TypeSafe AI ChoiceNoulScore curateddiopenrouters1mbsom

第一个 System One 决策模型,三原语的定义者。文本输入,返回类型化概率,不生成文本。

S1MB Task Avg
59.6
第 4 名 · 覆盖 137/137(100%)
Decision Index Full
60.1
第 3 名 · public 58.0
校准 ECE
0.074
越低越好 · Brier 0.356

模型信息

厂商
TypeSafe AI
类别
托管 API
参数规模
未公开
权重
闭源 / 未公开
许可证
闭源 需自行核对
输入模态
文本
决策原语
ChoiceNoulScore
上下文
32,000 token
输入价格
$0.042/M tok
延迟
70–500 ms(厂商口径)
可微调
不支持
微调方法
官方明确:不对客户数据做微调或 LoRA,所有账号共用同一权重
训练方式
hosted API
OpenRouter 模型 ID
typesafe/jev-1.13
OpenRouter 计费
输入 $0.042/M tok · 输出 免费
上架状态
Generally available

Decision Index 领域得分

知识与推理
71.6
语言理解
82.5
检索与分类
66.8
工具与自动化
94.6
艺术与人类品味
57.0

S1MB · 137 个基准

基准类型该模型全场最佳 全场均值对比来源
open jev / browser control v1 Choice 100.0 100.0 67.4 来源 ↗
open jev / customer control v1 Noul 100.0 100.0 58.0 来源 ↗
open jev / email selection control v1 Noul 100.0 100.0 40.6 来源 ↗
open jev / ir control v1 Noul 100.0 100.0 47.3 来源 ↗
open jev / mailroom control v1 Choice 100.0 100.0 82.9 来源 ↗
open jev / painting geometry v1 Choice 100.0 100.0 46.2 来源 ↗
open jev / phone extraction control v1 Noul 100.0 100.0 50.4 来源 ↗
open jev / sponsor segment control v1 Choice 100.0 100.0 52.2 来源 ↗
s1mb generalization contextual choice Choice 100.0 100.0 82.6 来源 ↗
s1mb generalization diverse noul Noul 100.0 100.0 59.2 来源 ↗
dbpedia Choice 98.9 98.9 83.5 来源 ↗
open jev / citation control v1 Choice 98.5 100.0 40.9 来源 ↗
open jev / context retention control v1 Noul 98.0 100.0 19.0 来源 ↗
s1mb generalization contextual noul Noul 98.0 98.0 62.0 来源 ↗
s1mb generalization diverse choice Choice 97.4 100.0 69.5 来源 ↗
openbookqa Choice 97.2 97.2 53.7 来源 ↗
arc Choice 97.1 98.5 64.4 来源 ↗
ruletaker Noul 96.0 96.0 27.7 来源 ↗
s1mb generalization diverse score Score 95.3 95.3 42.8 来源 ↗
open jev / vizdoom basic v1 Noul 94.1 100.0 39.0 来源 ↗
hatecheck Noul 94.0 98.0 47.5 来源 ↗
open jev / vizdoom basic v1 Score 93.9 100.0 29.6 来源 ↗
bbq Choice 93.3 98.3 53.8 来源 ↗
open jev / mailroom control v1 Noul 92.5 100.0 62.5 来源 ↗
laya / enron spam Noul 92.0 100.0 54.6 来源 ↗
clinc Choice 91.0 99.3 57.3 来源 ↗
open jev / silent failure control v1 Noul 90.0 100.0 38.8 来源 ↗
arct Choice 89.4 89.4 39.4 来源 ↗
open jev / entity alignment control v1 Noul 88.9 100.0 35.9 来源 ↗
open jev / drone control v1 Noul 88.2 100.0 50.2 来源 ↗
quartz Choice 88.1 88.1 49.1 来源 ↗
scone Choice 88.0 92.0 34.2 来源 ↗
massive Choice 87.4 96.7 64.5 来源 ↗
s1mb generalization contextual score Score 87.0 92.7 42.1 来源 ↗
winobias Choice 87.0 91.3 30.1 来源 ↗
gretel pii Choice 86.4 94.5 43.1 来源 ↗
miqa Choice 86.4 95.5 61.9 来源 ↗
scitail Noul 86.0 98.0 53.0 来源 ↗
open jev / reasoning control v1 Score 85.8 90.7 27.7 来源 ↗
argument quality Choice 85.7 95.9 37.4 来源 ↗
lonli Choice 85.3 86.9 39.2 来源 ↗
open jev / reasoning control v1 Choice 85.0 95.0 28.4 来源 ↗
open jev / email selection control v1 Choice 84.3 100.0 23.4 来源 ↗
open jev / amount extraction control v1 Choice 84.3 100.0 55.0 来源 ↗
hwu64 Choice 84.0 92.5 56.8 来源 ↗
babi nli Noul 84.0 84.0 37.0 来源 ↗
impli Noul 84.0 88.0 46.7 来源 ↗
open jev / ir control v1 Choice 84.0 88.0 44.8 来源 ↗
open jev / workflow controls v1 / invoice processing Noul 83.9 100.0 34.6 来源 ↗
mtop Choice 83.4 86.9 58.6 来源 ↗
creak Noul 82.0 88.0 38.8 来源 ↗
laya / ag news Choice 80.6 94.0 75.4 来源 ↗
defeasible nli Choice 80.4 80.4 32.8 来源 ↗
open jev / painting geometry v1 Noul 80.0 100.0 27.2 来源 ↗
sdoh nli Noul 80.0 92.0 60.8 来源 ↗
open jev / workflow controls v1 / security incidents Noul 79.8 100.0 37.8 来源 ↗
open jev / ir control v1 Score 79.0 98.6 10.1 来源 ↗
gsm8k Choice 78.6 78.6 23.5 来源 ↗
open jev / customer control v1 Choice 78.3 82.6 64.3 来源 ↗
open jev / amount extraction control v1 Noul 78.1 100.0 33.3 来源 ↗
aqua rat Choice 77.5 77.5 10.7 来源 ↗
banking77 Choice 75.8 93.7 54.6 来源 ↗
few nerd Choice 74.5 86.4 43.9 来源 ↗
qasper Noul 74.0 88.0 41.6 来源 ↗
snli Choice 72.6 98.4 46.0 来源 ↗
open jev / workflow controls v1 / agent trace observability Noul 72.3 100.0 27.8 来源 ↗
hans Noul 70.0 100.0 43.3 来源 ↗
synthetic relevance / nanocoir Noul 69.9 73.0 33.6 来源 ↗
crows pairs Choice 68.8 70.8 30.4 来源 ↗
paws Noul 68.0 88.0 45.7 来源 ↗
open jev / workflow controls v1 / customer service Noul 66.0 100.0 41.6 来源 ↗
ethics Noul 65.8 68.6 23.9 来源 ↗
robust lr Choice 65.0 66.7 18.8 来源 ↗
followir robust04 Noul 64.5 90.9 46.9 来源 ↗
canttalk Noul 64.0 86.0 43.9 来源 ↗
ethos Noul 64.0 76.0 43.7 来源 ↗
synthetic relevance / nanobeir / nanonq Noul 62.3 67.2 39.5 来源 ↗
scicite Choice 61.3 71.0 44.0 来源 ↗
open jev / phone extraction control v1 Choice 60.9 100.0 32.4 来源 ↗
synthetic relevance / nanobeir / nanomsmarco Noul 60.8 62.5 37.3 来源 ↗
open jev / customer control v1 Score 58.9 99.3 36.7 来源 ↗
synthetic relevance / nanobeir / nanohotpotqa Noul 58.8 70.8 28.4 来源 ↗
cladder Noul 58.0 94.0 17.7 来源 ↗
ud ewt Choice 56.7 56.7 13.2 来源 ↗
synthetic relevance / nanobeir / nanohotpotqa Score 54.5 63.6 13.9 来源 ↗
contract nli Choice 54.5 76.0 31.8 来源 ↗
synthetic relevance / nanobeir / nanofiqa2018 Noul 53.1 54.0 28.6 来源 ↗
open jev / reasoning control v1 Noul 52.2 100.0 25.3 来源 ↗
aegis2 Noul 51.8 62.3 32.8 来源 ↗
go emotions Noul 48.5 61.3 38.3 来源 ↗
temporal nli Choice 48.4 80.7 18.3 来源 ↗
synthetic relevance / nanobeir / nanoarguana Noul 47.2 48.2 15.6 来源 ↗
logical entailment Noul 46.0 72.0 12.2 来源 ↗
synthetic relevance / nanobeir / nanotouche2020 Noul 44.6 47.0 26.6 来源 ↗
open jev / painting geometry v1 Score 44.2 100.0 15.2 来源 ↗
stepgame Choice 44.0 82.1 17.8 来源 ↗
winowhy Noul 44.0 62.0 18.8 来源 ↗
synthetic relevance / nanobeir / nanodbpedia Score 43.7 44.9 17.6 来源 ↗
laya / typed decisions Noul 42.7 47.8 24.4 来源 ↗
synthetic relevance / nanobeir / nanomsmarco Score 42.4 52.6 17.8 来源 ↗
synthetic relevance / nanobeir / nanodbpedia Noul 42.1 52.1 29.0 来源 ↗
synthetic relevance / nanobeir / nanonfcorpus Noul 41.5 46.2 25.1 来源 ↗
patent similarity Score 40.5 48.3 17.7 来源 ↗
laya / typed decisions Score 39.5 63.5 26.5 来源 ↗
boardgameqa Choice 39.1 62.5 15.6 来源 ↗
spartqa Choice 38.7 38.7 4.4 来源 ↗
synthetic relevance / nanobeir / nanofiqa2018 Score 37.9 46.7 15.2 来源 ↗
civil comments Noul 35.4 36.7 21.4 来源 ↗
synthetic relevance / nanobeir / nanoarguana Score 35.4 36.0 7.7 来源 ↗
ethics Choice 33.3 83.3 20.9 来源 ↗
esci Choice 32.9 46.6 13.0 来源 ↗
lexcomp Noul 32.0 52.0 22.5 来源 ↗
laya / typed decisions Choice 31.8 34.5 18.9 来源 ↗
sgd Choice 30.5 96.6 24.4 来源 ↗
synthetic relevance / nanobeir / nanonfcorpus Score 30.5 38.6 10.7 来源 ↗
synthetic relevance / nanocoir Score 28.7 44.3 4.3 来源 ↗
synthetic relevance / nanobeir / nanoscidocs Noul 28.3 41.7 17.3 来源 ↗
synthetic relevance / nanobeir / nanoscidocs Score 28.3 36.7 7.2 来源 ↗
synthetic relevance / nanobeir / nanonq Score 27.4 49.4 9.3 来源 ↗
followir core17 Noul 27.2 46.6 21.2 来源 ↗
followir news21 Noul 26.7 42.7 14.9 来源 ↗
synthetic relevance / nanobeir / nanotouche2020 Score 25.9 44.3 20.6 来源 ↗
wiqa Choice 25.4 58.7 19.3 来源 ↗
tracie Noul 24.0 34.0 8.7 来源 ↗
fol nli Choice 22.9 54.1 10.4 来源 ↗
open jev / snake v1 Noul 21.8 100.0 9.0 来源 ↗
hh rlhf Choice 17.9 23.1 3.1 来源 ↗
open jev / tic tac toe v1 Choice 9.8 48.5 3.5 来源 ↗
open jev / snake v1 Choice 9.3 83.5 8.1 来源 ↗
argument quality Noul 8.4 23.9 5.4 来源 ↗
plane Noul 8.0 100.0 11.1 来源 ↗
open jev / entity alignment control v1 Score 6.5 100.0 8.1 来源 ↗
corr2cause Choice 0.0 91.7 3.8 来源 ↗
nlsat Noul 0.0 22.0 2.1 来源 ↗
open jev / drone control v1 Choice 0.0 100.0 4.8 来源 ↗
open jev / drone control v1 Score 0.0 99.2 1.1 来源 ↗
poem sentiment Choice 0.0 57.1 4.0 来源 ↗

Decision Index · 53 个基准

基准类型该模型全场最佳 全场均值对比来源
MiniWoB++ 混合(Choice 100.0 100.0 100.0 来源 ↗
ARC-Easy 混合(Choice 99.3 99.5 87.3 来源 ↗
ARC-Challenge 混合(Choice 97.8 98.2 79.4 来源 ↗
Email spam 混合(Choice 97.4 97.4 97.4 来源 ↗
BFCL 混合(Choice 95.8 98.8 80.5 来源 ↗
Phishing gradient 混合(Choice 94.6 94.6 94.6 来源 ↗
HellaSwag 混合(Choice 94.5 98.5 73.0 来源 ↗
OpenBookQA 混合(Choice 94.0 94.0 94.0 来源 ↗
BBH 混合(Choice 92.9 92.9 58.3 来源 ↗
WinoGrande 混合(Choice 92.0 97.5 69.8 来源 ↗
MMLU 混合(Choice 91.7 91.9 64.4 来源 ↗
BPoMP 混合(Choice 90.6 97.0 73.8 来源 ↗
Support tickets 混合(Choice 90.5 90.5 90.5 来源 ↗
CLINC150 混合(Choice 89.3 97.4 67.8 来源 ↗
API-Bank 混合(Choice 88.2 93.1 56.3 来源 ↗
CommonsenseQA 混合(Choice 87.5 87.5 87.5 来源 ↗
FinEntity 混合(Choice 87.0 97.1 73.8 来源 ↗
NLI4CT 混合(Choice 84.1 86.2 69.3 来源 ↗
MMLU-Pro 混合(Choice 82.7 82.7 42.3 来源 ↗
When2Call 混合(Choice 81.0 91.7 57.8 来源 ↗
RouterBench 混合(Choice 79.9 80.1 73.2 来源 ↗
BANKING77 混合(Choice 79.7 94.1 68.7 来源 ↗
GPQA Diamond 混合(Choice 78.3 78.3 37.7 来源 ↗
RAGTruth 混合(Choice 76.5 86.0 52.1 来源 ↗
ANLI 混合(Choice 74.8 98.6 54.0 来源 ↗
CRUXEval 混合(Choice 73.0 87.9 50.7 来源 ↗
HoVer 混合(Choice 72.8 89.4 63.8 来源 ↗
CLadder 混合(Choice 72.6 97.7 62.5 来源 ↗
GSM8K 混合(Choice 71.8 83.5 36.0 来源 ↗
ContractNLI 混合(Choice 71.7 86.5 59.9 来源 ↗
New Yorker 混合(Choice 70.1 82.0 52.9 来源 ↗
MuSR 混合(Choice 66.1 86.2 56.0 来源 ↗
ToolRet 混合(Choice 65.3 69.1 55.1 来源 ↗
cfcolor 混合(Choice 64.7 70.2 57.9 来源 ↗
VAST 混合(Choice 64.6 82.0 50.1 来源 ↗
PhishNChips 混合(Choice 62.5 99.9 62.1 来源 ↗
Humicroedit 混合(Choice 61.9 75.1 57.2 来源 ↗
Amazon ESCI 混合(Choice 55.2 61.6 40.9 来源 ↗
Home appliances 混合(Choice 52.3 98.9 23.2 来源 ↗
iSarcasmEval 混合(Choice 50.5 71.0 38.4 来源 ↗
BRIGHT 混合(Choice 47.5 50.9 36.1 来源 ↗
Habermas 混合(Choice 45.9 71.8 42.2 来源 ↗
SGD 混合(Choice 43.0 73.4 47.4 来源 ↗
ACOS 混合(Choice 29.5 57.3 14.6 来源 ↗
SATA-Bench 混合(Choice 26.4 37.2 20.3 来源 ↗
HLE 混合(Choice 20.4 20.4 12.3 来源 ↗
POP909 混合(Choice 18.1 74.6 12.6 来源 ↗
ChessBench 混合(Choice 17.2 40.2 14.2 来源 ↗
RTFM 混合(Choice 17.0 17.0 17.0 来源 ↗
ScienceWorld 混合(Choice 6.7 6.7 6.7 来源 ↗
Boxoban 混合(Choice 0.0 0.0 0.0 来源 ↗
Hanabi 混合(Choice 0.0 0.0 0.0 来源 ↗
Codenames 混合(Choice 0.0 0.0 0.0 来源 ↗

来自 systemonemodels.org 的详细介绍

What Jev is

Jev is TypeSafe AI’s System One model, a classification model that reads a block of text and answers typed questions about it. You supply the state, which is the text you want judged, and the questions, which you define in your own code. It answers each one and stops. It does not write a reply, produce code, or explain itself. Questions come in three shapes, Choice, Score and Noul, each walked through in Choice, Score and Noul. Several questions in one request are evaluated against the same state in parallel. TypeSafe says it trained the model with Reinforcement Learning for Calibrated Decisions, so the probabilities are meant to match real outcome rates rather than match what a human rater prefers to read. No paper, dataset or training recipe has been published, which RLCD explained goes through. TechCrunch reported on 18 September 2026 that “Almeida says Jev is trained exclusively on synthetic data using a technique he calls ‘reinforcement learning from calibrated decisions.’” That is the only public statement about the training data, and TechCrunch renders the acronym as “from calibrated decisions” where TypeSafe’s own docs say “for calibrated decisions.”

Jev guides and comparisons

New to Jev: What is Jev AI? is the plain-English explainer. Getting started: how to get Jev access and using Jev with Hermes. Cost and speed: Jev speed and pricing and Jev benchmarks. Internals: how Jev is built and RLCD explained. What people build: Jev game demos. Other options: Jev alternatives, plus Jev vs Clef, Jev vs pplx-decider, Jev vs Claude, Jev vs GPT, Jev vs BERT and Jev vs GLiNER.

Jev 1.13: the current version and its aliases

jev-1.13.0 is the only published version. Two aliases point at it, jev-latest and jev-preview, and the docs say why: “jev-preview currently points to the same model as jev-latest. There is no preview build available right now.” There is no model changelog. TypeSafe’s llms.txt lists SDK changelogs only, and no version above 1.13.0 appeared in the docs, the blog, the press or search when this page was checked on 20 September 2026. Pinning jev-1.13.0 rather than an alias keeps the published weakness list matched to the model you are actually calling.

What Jev 1.13 does badly

TypeSafe publishes a jaggedness page per version. The one for 1.13 was last reviewed on 17 September 2026 and lists nine failure modes. Literal reading: the model “answers the question you wrote, not the one you meant.” Math and numbers: counting, numeric representations, and arithmetic attempted through a Score are all weak. The page ships a runnable Python sample that counts by asking one Noul per item instead. Date and time comparison: Jev “reads dates as text, not as ordered quantities.” Indirection: “instructions carrying double negatives or complex indirection are answered less reliably.” Large state full of irrelevant detail: accuracy falls as the state grows with content unrelated to the decision. TypeSafe’s own name for the cause is context rot. Adversarial content: “state is data, and jev-1.13 does not treat it as hostile by default.” Screening user-supplied text is your job, which is what the prompt injection screen recipe is for. Contradictory instructions and criteria: answers degrade when the instructions and the criteria pull in different directions. Common-sense structural invariants: the page works two numbers through, an idea that scores 0.22 as a Noul but comes back at 0.99 when asked as a Choice, and a Noul plus its negation summing to 1.19 instead of 1. TypeSafe’s line is that “there are many reasons that P(noul) and 1 - P(not noul) may not be directly comparable.” Generation: “jev-1.13 is not trained to generate text.” The page closes by telling you to avoid System Two tasks, which it describes as work with “more layers of indirection.”

How to call Jev

One POST to https://api.typesafe.ai/v1/systemone with an Authorization: Bearer header. Both SDKs read the key from TYPESAFE_API_KEY when you construct the default client. The quickstart’s curl example, unchanged: curl -X POST https://api.typesafe.ai/v1/systemone \ -H "Authorization: Bearer $TYPESAFE_API_KEY" \ -H "Content-Type: application/json" \ -d @- <<'EOF' { "state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.", "model": "jev-latest", "questions": { "urgency": { "type": "noul", "instructions": "Does this message express urgency?" } } } EOF The Python SDK, typesafe-sdk, installs with uv add typesafe-sdk or pip install typesafe-sdk and needs Python 3.10 or newer. from typesafe_sdk import Choice, TypeSafeClient with TypeSafeClient() as client: response = client.system_one( state={"document": "I was charged twice. Please fix this ASAP."}, questions={ "category": Choice( instructions="What is this ticket about?", criteria={"billing": None, "technical": None, "other": None}, ), }, ) print(response.choices["category"].choice) Check that last line against the SDK you installed. TypeSafe’s quickstart page reads answers back as response.answers["department"].choice while the Python SDK’s own page uses response.choices["category"].choice. Both are TypeSafe’s text, and the two disagree. Python is on 0.7.0, released 18 September 2026. That release swapped serialisation from msgspec to pydantic, which the release notes flag as a breaking change, and gave system_one() a response_model argument that takes a Pydantic model. The JavaScript SDK, @typesafe-ai/sdk, installs with npm install @typesafe-ai/sdk and needs Node.js 20 or newer. It was still on 0.6.0 on 20 September 2026, with no 0.7.0 release published. import { choice, TypeSafeClient } from "@typesafe-ai/sdk"; const client = new TypeSafeClient(); const response = await client.systemOne({ state: { document: "I was charged twice. Please fix this ASAP." }, questions: { category: choice("What is this ticket about?", { billing: n

Jev on OpenRouter

OpenRouter has listed Jev since 18 September 2026 as typesafe/jev-1.13, plus ~typesafe/jev-latest, an alias that points at the newest version. TypeSafe is the only provider. The price matches TypeSafe’s: $0.042 per million input tokens, with free output. It is not a chat model there. OpenRouter serves it through its Decisions API, POST https://openrouter.ai/api/alpha/decisions, and warns that chat completions SDKs will not work with it. The request takes the same state and questions as TypeSafe’s API, and an OpenRouter key replaces the TypeSafe key. OpenRouter’s own SDK calls it through openrouter.alpha.decisions.create(). The alpha path means OpenRouter still treats the endpoint as unfinished. The OpenRouter integration page has more.

Languages and data handling

English first. The docs say “English is the primary training language and where accuracy is currently best. Other languages, including CJK scripts, are handled but not equally well.” They add that you should test on your own content before putting a non-English workload through the model. Input is text only, as a string, a JSON object, or an array of text values. On data, TypeSafe states that “Jev is not trained on customer requests or responses” and that the model is “not fine-tuned or LoRA-adapted with customer data,” because “the same weights serve every account.” Zero data retention is offered to enterprise customers under a data processing agreement. None of that has been independently audited.

Access today: open sign-up, console and API key

There is no waitlist. TypeSafe removed it on 20 September 2026, five days after launch, posting on X that “Jev is now available to everyone. No waitlist.” Sign in at console.typesafe.ai with Google or an email code, create a key, and call the API. Getting access and making your first call covers the steps. The weights are not released, so there is no self-hosted option and no local build.

What it is not for

Anything that ends in prose. Drafting, summarising and code generation sit outside what the model was trained to do, and TypeSafe says so plainly. Arithmetic, counting and date comparison belong in your code rather than in a question. So does anything that needs two facts chained together, because indirection is one of the nine documented weak spots. Unscreened user input is the other thing you handle yourself. The model reads the state as data and has no defence against instructions hidden in it, so a separate screen sits in front. The vendor speed and cost claims are unpacked in Jev speed and pricing, and what is known about the internals is in Jev architecture.

Specifications

Question types Choice Score Noul Max Choice options255 Score levelsUp to 10 Questions per callNot documented Total context64,000 tokens State budget32,000 tokens Rate limit100,000 tokens/sec and 40 requests/sec as of 2026-09-30, adjusting dynamically EndpointPOST https://api.typesafe.ai/v1/systemone SDKsPython: typesafe-sdkTypeScript: @typesafe-ai/sdk 64k tokens across the whole request. The state plus the longest single question must fit in 32k of that. OpenRouter lists Jev with a 32,000-token context, which is that second limit, not the whole request. A Score needs at least 2 levels. The maximum number of questions in one call is not documented. On rate limits the docs warn that the published numbers can change without notice while TypeSafe serves launch demand, and that higher limits come with custom and enterprise plans.

Versions

jev-1.13.0, 15 Sep 2026, The only published version. The aliases jev-latest and jev-preview both resolve to it, and the docs say there is no preview build available right now. TypeSafe ships a per-version jaggedness page listing what this version does badly. Release notes

Use cases

What people use Jev for, one page per pattern. Workflow controlStarter