decision.host
首页 / 模型 / gliner2-large-v1

gliner2-large-v1

Fastino Labs ChoiceNoulScore 开源权重 340M(活跃 355M) curateddis1mbsom

340M 编码器,既做类型化分类、也抽 span 与关系,并支持跨决策规则约束。这类「传统信息抽取模型转向类型化决策」的代表。

S1MB Task Avg
7.4
第 86 名 · 覆盖 137/137(100%)

模型信息

厂商
Fastino Labs
类别
开源权重
参数规模
340M(活跃 355M)
基座模型
fastino/gliner2-large-v1
权重
可下载(开源)
许可证
Apache-2.0 可商用
输入模态
文本
决策原语
ChoiceNoulScore
延迟
38-167 ms
可微调
未确认
训练方式
full fine-tune
上架状态
Generally available

数据来源与链接

本站聚合自: curated、di、s1mb、som。 分数与链接均指向原始出处。

S1MB · 137 个基准

基准类型该模型全场最佳 全场均值对比来源
open jev / mailroom control v1 Choice 91.7 100.0 82.9 来源 ↗
s1mb generalization contextual choice Choice 70.4 100.0 82.6 来源 ↗
dbpedia Choice 59.5 98.9 83.5 来源 ↗
banking77 Choice 59.0 93.7 54.6 来源 ↗
laya / ag news Choice 58.2 94.0 75.4 来源 ↗
massive Choice 54.5 96.7 64.5 来源 ↗
s1mb generalization contextual noul Noul 54.0 98.0 62.0 来源 ↗
open jev / amount extraction control v1 Choice 52.9 100.0 55.0 来源 ↗
mtop Choice 39.5 86.9 58.6 来源 ↗
open jev / customer control v1 Score 39.1 99.3 36.7 来源 ↗
miqa Choice 38.6 95.5 61.9 来源 ↗
s1mb generalization diverse choice Choice 37.9 100.0 69.5 来源 ↗
hwu64 Choice 35.1 92.5 56.8 来源 ↗
scicite Choice 33.9 71.0 44.0 来源 ↗
arct Choice 31.9 89.4 39.4 来源 ↗
followir robust04 Noul 26.0 90.9 46.9 来源 ↗
few nerd Choice 25.8 86.4 43.9 来源 ↗
s1mb generalization diverse noul Noul 20.0 100.0 59.2 来源 ↗
ethos Noul 18.0 76.0 43.7 来源 ↗
laya / enron spam Noul 18.0 100.0 54.6 来源 ↗
s1mb generalization contextual score Score 17.8 92.7 42.1 来源 ↗
arc Choice 17.6 98.5 64.4 来源 ↗
go emotions Noul 17.0 61.3 38.3 来源 ↗
ethics Choice 16.7 83.3 20.9 来源 ↗
gretel pii Choice 16.6 94.5 43.1 来源 ↗
clinc Choice 15.9 99.3 57.3 来源 ↗
s1mb generalization diverse score Score 15.7 95.3 42.8 来源 ↗
laya / typed decisions Noul 14.0 47.8 24.4 来源 ↗
sgd Choice 13.6 96.6 24.4 来源 ↗
ethics Noul 13.0 68.6 23.9 来源 ↗
open jev / mailroom control v1 Noul 12.5 100.0 62.5 来源 ↗
lexcomp Noul 12.0 52.0 22.5 来源 ↗
civil comments Noul 11.3 36.7 21.4 来源 ↗
laya / typed decisions Choice 11.3 34.5 18.9 来源 ↗
quartz Choice 9.5 88.1 49.1 来源 ↗
hatecheck Noul 8.0 98.0 47.5 来源 ↗
sdoh nli Noul 8.0 92.0 60.8 来源 ↗
openbookqa Choice 6.9 97.2 53.7 来源 ↗
open jev / browser control v1 Choice 6.5 100.0 67.4 来源 ↗
plane Noul 6.0 100.0 11.1 来源 ↗
tracie Noul 6.0 34.0 8.7 来源 ↗
winowhy Noul 6.0 62.0 18.8 来源 ↗
creak Noul 4.0 88.0 38.8 来源 ↗
open jev / ir control v1 Choice 4.0 88.0 44.8 来源 ↗
open jev / silent failure control v1 Noul 4.0 100.0 38.8 来源 ↗
aegis2 Noul 2.4 62.3 32.8 来源 ↗
stepgame Choice 2.4 82.1 17.8 来源 ↗
winobias Choice 2.2 91.3 30.1 来源 ↗
synthetic relevance / nanobeir / nanofiqa2018 Noul 2.1 54.0 28.6 来源 ↗
babi nli Noul 2.0 84.0 37.0 来源 ↗
nlsat Noul 2.0 22.0 2.1 来源 ↗
paws Noul 2.0 88.0 45.7 来源 ↗
synthetic relevance / nanobeir / nanonq Noul 1.2 67.2 39.5 来源 ↗
open jev / reasoning control v1 Score 1.0 90.7 27.7 来源 ↗
synthetic relevance / nanobeir / nanonfcorpus Noul 0.3 46.2 25.1 来源 ↗
aqua rat Choice 0.0 77.5 10.7 来源 ↗
argument quality Choice 0.0 95.9 37.4 来源 ↗
argument quality Noul 0.0 23.9 5.4 来源 ↗
bbq Choice 0.0 98.3 53.8 来源 ↗
boardgameqa Choice 0.0 62.5 15.6 来源 ↗
canttalk Noul 0.0 86.0 43.9 来源 ↗
cladder Noul 0.0 94.0 17.7 来源 ↗
contract nli Choice 0.0 76.0 31.8 来源 ↗
corr2cause Choice 0.0 91.7 3.8 来源 ↗
crows pairs Choice 0.0 70.8 30.4 来源 ↗
defeasible nli Choice 0.0 80.4 32.8 来源 ↗
esci Choice 0.0 46.6 13.0 来源 ↗
fol nli Choice 0.0 54.1 10.4 来源 ↗
followir core17 Noul 0.0 46.6 21.2 来源 ↗
followir news21 Noul 0.0 42.7 14.9 来源 ↗
gsm8k Choice 0.0 78.6 23.5 来源 ↗
hans Noul 0.0 100.0 43.3 来源 ↗
hh rlhf Choice 0.0 23.1 3.1 来源 ↗
impli Noul 0.0 88.0 46.7 来源 ↗
laya / typed decisions Score 0.0 63.5 26.5 来源 ↗
logical entailment Noul 0.0 72.0 12.2 来源 ↗
lonli Choice 0.0 86.9 39.2 来源 ↗
open jev / amount extraction control v1 Noul 0.0 100.0 33.3 来源 ↗
open jev / citation control v1 Choice 0.0 100.0 40.9 来源 ↗
open jev / context retention control v1 Noul 0.0 100.0 19.0 来源 ↗
open jev / customer control v1 Choice 0.0 82.6 64.3 来源 ↗
open jev / customer control v1 Noul 0.0 100.0 58.0 来源 ↗
open jev / drone control v1 Choice 0.0 100.0 4.8 来源 ↗
open jev / drone control v1 Noul 0.0 100.0 50.2 来源 ↗
open jev / drone control v1 Score 0.0 99.2 1.1 来源 ↗
open jev / email selection control v1 Choice 0.0 100.0 23.4 来源 ↗
open jev / email selection control v1 Noul 0.0 100.0 40.6 来源 ↗
open jev / entity alignment control v1 Noul 0.0 100.0 35.9 来源 ↗
open jev / entity alignment control v1 Score 0.0 100.0 8.1 来源 ↗
open jev / ir control v1 Noul 0.0 100.0 47.3 来源 ↗
open jev / ir control v1 Score 0.0 98.6 10.1 来源 ↗
open jev / painting geometry v1 Choice 0.0 100.0 46.2 来源 ↗
open jev / painting geometry v1 Noul 0.0 100.0 27.2 来源 ↗
open jev / painting geometry v1 Score 0.0 100.0 15.2 来源 ↗
open jev / phone extraction control v1 Choice 0.0 100.0 32.4 来源 ↗
open jev / phone extraction control v1 Noul 0.0 100.0 50.4 来源 ↗
open jev / reasoning control v1 Choice 0.0 95.0 28.4 来源 ↗
open jev / reasoning control v1 Noul 0.0 100.0 25.3 来源 ↗
open jev / snake v1 Choice 0.0 83.5 8.1 来源 ↗
open jev / snake v1 Noul 0.0 100.0 9.0 来源 ↗
open jev / sponsor segment control v1 Choice 0.0 100.0 52.2 来源 ↗
open jev / tic tac toe v1 Choice 0.0 48.5 3.5 来源 ↗
open jev / vizdoom basic v1 Noul 0.0 100.0 39.0 来源 ↗
open jev / vizdoom basic v1 Score 0.0 100.0 29.6 来源 ↗
open jev / workflow controls v1 / agent trace observability Noul 0.0 100.0 27.8 来源 ↗
open jev / workflow controls v1 / customer service Noul 0.0 100.0 41.6 来源 ↗
open jev / workflow controls v1 / invoice processing Noul 0.0 100.0 34.6 来源 ↗
open jev / workflow controls v1 / security incidents Noul 0.0 100.0 37.8 来源 ↗
patent similarity Score 0.0 48.3 17.7 来源 ↗
poem sentiment Choice 0.0 57.1 4.0 来源 ↗
qasper Noul 0.0 88.0 41.6 来源 ↗
robust lr Choice 0.0 66.7 18.8 来源 ↗
ruletaker Noul 0.0 96.0 27.7 来源 ↗
scitail Noul 0.0 98.0 53.0 来源 ↗
scone Choice 0.0 92.0 34.2 来源 ↗
snli Choice 0.0 98.4 46.0 来源 ↗
spartqa Choice 0.0 38.7 4.4 来源 ↗
synthetic relevance / nanobeir / nanoarguana Noul 0.0 48.2 15.6 来源 ↗
synthetic relevance / nanobeir / nanoarguana Score 0.0 36.0 7.7 来源 ↗
synthetic relevance / nanobeir / nanodbpedia Noul 0.0 52.1 29.0 来源 ↗
synthetic relevance / nanobeir / nanodbpedia Score 0.0 44.9 17.6 来源 ↗
synthetic relevance / nanobeir / nanofiqa2018 Score 0.0 46.7 15.2 来源 ↗
synthetic relevance / nanobeir / nanohotpotqa Noul 0.0 70.8 28.4 来源 ↗
synthetic relevance / nanobeir / nanohotpotqa Score 0.0 63.6 13.9 来源 ↗
synthetic relevance / nanobeir / nanomsmarco Noul 0.0 62.5 37.3 来源 ↗
synthetic relevance / nanobeir / nanomsmarco Score 0.0 52.6 17.8 来源 ↗
synthetic relevance / nanobeir / nanonfcorpus Score 0.0 38.6 10.7 来源 ↗
synthetic relevance / nanobeir / nanonq Score 0.0 49.4 9.3 来源 ↗
synthetic relevance / nanobeir / nanoscidocs Noul 0.0 41.7 17.3 来源 ↗
synthetic relevance / nanobeir / nanoscidocs Score 0.0 36.7 7.2 来源 ↗
synthetic relevance / nanobeir / nanotouche2020 Noul 0.0 47.0 26.6 来源 ↗
synthetic relevance / nanobeir / nanotouche2020 Score 0.0 44.3 20.6 来源 ↗
synthetic relevance / nanocoir Noul 0.0 73.0 33.6 来源 ↗
synthetic relevance / nanocoir Score 0.0 44.3 4.3 来源 ↗
temporal nli Choice 0.0 80.7 18.3 来源 ↗
ud ewt Choice 0.0 56.7 13.2 来源 ↗
wiqa Choice 0.0 58.7 19.3 来源 ↗

同族条目(4)

型号参数Decision Index S1MBVision榜单
gliner2-large-v1 当前 340M(活跃 355M) — 7.4 — s1mb
GLiNER2.5-Decide 486M 9.2 — — di
gliner2.5-base-v1 194M(活跃 95M) — 6.4 — s1mb
gliner2.5-multi-v1 287M(活跃 95M) — 3.3 — s1mb

来自 systemonemodels.org 的详细介绍

What GLiNER2.5-Decide is

GLiNER2.5-Decide is Fastino Labs’ System One model: an encoder-based decision model rather than a generative one. It evaluates a set of user-defined typed questions and rules against a piece of text in one pass, then jointly decodes the answers, returning probability distributions and confidence scores. Beyond typed Choice, Score and Noul questions, Fastino describes it returning character-level spans, relations, and structured records, and applying implications, exclusions, cardinality limits, and ordinal bounds across related decisions in one call. The model card is direct about scope: “This release is not a general-purpose model. It does not reason, explain, or answer open questions.”

How it is built

The Hugging Face model card lists GLiNER2.5-Decide as a 340M-parameter model fine-tuned from Fastino’s own gliner2-large-v1, a DeBERTa-v3-large encoder. The Hugging Face API’s safetensors metadata for the published weights reports 486,444,053 parameters, which does not match the 340M figure. Neither source explains the gap. It loads through the gliner2 Python package (AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")) and accepts label sets at runtime, so a new task needs no retraining or prompt template.

What it’s good at, per Fastino’s own benchmark

Fastino evaluated the model on “Fast Decisions,” an internal suite of 17 datasets covering intent routing, triage, sentiment and content understanding. This is Fastino’s own benchmark, not an independently run one. The blog post and company X post give GLiNER2.5-Decide 60.1% average accuracy, ahead of JevK5 (57.5%), SemIf (56.4%) and Laya (46.6%), leading on 9 of 17 datasets. The Hugging Face card’s own results table gives slightly different figures for the same run: 60.2% and JevK5 at 57.6%. Fastino highlights support intent classification at 75.3%, 18.6 points ahead of the next model, and banking intent at 64.3%, 8.6 points ahead. The GitHub repository behind it, fastino-ai/GLiNER2, had 2,169 stars and was last pushed on 24 September 2026, the day of release.

What an independent benchmark shows

The Decision Index 0.2.1 by multimodalart, updated 28 September 2026, is a broader test than Fastino’s. On a chance-corrected score where 0 is random guessing and 100 is perfect, averaged over 38 benchmarks in five weighted areas, GLiNER2.5-Decide scores 11.21 against Jev’s 57.91. That is the best result among the models under 500M parameters it tested, but far behind the larger open models. JevK5, which Fastino’s suite ranks below GLiNER2.5-Decide, scores 38.81 here. Its expected calibration error, the average gap between stated confidence and actual accuracy, is 0.088 against Jev’s 0.074.

Access and running it

Weights are Apache 2.0 on Hugging Face with no waitlist. Fastino says the model is “efficient enough to run locally on consumer-grade CPUs or in air-gapped environments,” and co-founder George Maloney gives CPU latency around 167ms and GPU latency of 38 to 47ms for short requests. Fastino also runs hosted inference and fine-tuning at http://agent.fastino.ai, for calling from inside a coding agent; no pricing for that service appeared in the sources checked.

What it’s not for

It does not generate text, reason through multi-step problems, or answer open-ended questions; the model card says so outright. No context-window limit, maximum Choice count, or per-call question cap is documented anywhere checked, so those limits are unknown rather than unlimited.

Specifications

Question types Choice Score Noul Max Choice optionsNot documented Score levelsNot documented Questions per callNot documented Total contextNot documented State budgetNot documented Rate limitNone for self-hosted use. Fastino's hosted inference and fine-tuning API at http://agent.fastino.ai did not have published rate limits in the sources checked. Endpointhttp://agent.fastino.ai SDKsPython: gliner2 Neither the blog post, the Hugging Face model card, nor the GitHub repo README documented a fixed context window, a maximum number of Choice options, or a cap on questions per call as of this check.

Versions

fastino/GLiNER2.5-Decide, 24 Sep 2026, 340M-parameter encoder fine-tuned from fastino/gliner2-large-v1, a DeBERTa-v3-large base per the Hugging Face model card. The Hugging Face API's safetensors metadata reports 486,444,053 total parameters for the published weights, which does not match the 340M figure stated in the blog post and model card; this discrepancy is unresolved as of this check. Apache 2.0 license. Release notes

Use cases

What people use GLiNER2.5-Decide for, one page per pattern. Workflow controlStarter

Support inbox triage with System One models

Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket. Choice Score Noul Workflow controlIntermediate

Intent and model routing with System One models

One Jev call reads an incoming request and returns its intent as a label plus a difficulty rating on a scale you wrote. Your router reads both numbers and picks the handler: deterministic code, a cheap model, an expensive one, or a human queue. Choice Score Real-time and agentsIntermediate

Agent routing and skill selection with System One models

An agent choosing from a long skill roster reads one truncated line per entry and often loads the wrong thing. Jev ranks every entry in one request and separately answers whether any skill applies at all, so the agent gets a short hint instead of a guess. Choice Noul Data and operationsStarter