decision.host
首页 / 模型 / convaiinnovations/laya-multilingual

convaiinnovations/laya-multilingual

Convai Innovations ChoiceNoulScore 开源权重 421M(活跃 370M) curateddis1mbsom

端侧首选。400M 级 ModernBERT 编码器,Apache-2.0,有 MLX / CoreML / C++ / ONNX 多套本地运行时,M3 Max 上短决策只需 5–14 ms。代价是 S1MB 上分数偏低(13–15 分)。

S1MB Task Avg
9.3
第 80 名 · 覆盖 137/137(100%)

模型信息

厂商
Convai Innovations
类别
端侧小模型
参数规模
421M(活跃 370M)
基座模型
ModernBERT-large(英文)/ mmBERT-base(多语言)
权重
可下载(开源)
许可证
Apache-2.0 可商用
输入模态
文本
决策原语
ChoiceNoulScore
延迟
32.8–39.5 ms(官方);MLX 短决策 7–14 ms(M3 Max)
可微调
可以
微调方法
CLI: laya-train --data tickets.csv --out ./ft;RLCD 循环 + 温度校准,输出 train_report.json(含 acc / ECE / Brier)
微调硬件
Kaggle 2×T4 或 MPS/CPU 笔记本均可
训练方式
full fine-tune
上架状态
Generally available

S1MB · 137 个基准

基准类型该模型全场最佳 全场均值对比来源
laya / enron spam Noul 98.0 100.0 54.6 来源 ↗
laya / ag news Choice 86.6 94.0 75.4 来源 ↗
sdoh nli Noul 68.0 92.0 60.8 来源 ↗
ethos Noul 66.0 76.0 43.7 来源 ↗
open jev / painting geometry v1 Choice 63.2 100.0 46.2 来源 ↗
clinc Choice 60.0 99.3 57.3 来源 ↗
open jev / mailroom control v1 Choice 58.3 100.0 82.9 来源 ↗
dbpedia Choice 57.3 98.9 83.5 来源 ↗
go emotions Noul 55.5 61.3 38.3 来源 ↗
open jev / amount extraction control v1 Choice 49.0 100.0 55.0 来源 ↗
mtop Choice 48.6 86.9 58.6 来源 ↗
banking77 Choice 43.2 93.7 54.6 来源 ↗
s1mb generalization contextual choice Choice 42.1 100.0 82.6 来源 ↗
paws Noul 40.0 88.0 45.7 来源 ↗
hwu64 Choice 36.2 92.5 56.8 来源 ↗
synthetic relevance / nanobeir / nanomsmarco Noul 36.0 62.5 37.3 来源 ↗
synthetic relevance / nanobeir / nanonq Noul 35.2 67.2 39.5 来源 ↗
scitail Noul 32.0 98.0 53.0 来源 ↗
massive Choice 32.0 96.7 64.5 来源 ↗
s1mb generalization diverse choice Choice 31.3 100.0 69.5 来源 ↗
synthetic relevance / nanobeir / nanotouche2020 Noul 30.5 47.0 26.6 来源 ↗
hatecheck Noul 30.0 98.0 47.5 来源 ↗
synthetic relevance / nanobeir / nanonfcorpus Noul 28.3 46.2 25.1 来源 ↗
followir robust04 Noul 26.2 90.9 46.9 来源 ↗
synthetic relevance / nanobeir / nanohotpotqa Noul 24.8 70.8 28.4 来源 ↗
snli Choice 24.2 98.4 46.0 来源 ↗
synthetic relevance / nanobeir / nanodbpedia Noul 24.0 52.1 29.0 来源 ↗
synthetic relevance / nanobeir / nanofiqa2018 Noul 22.4 54.0 28.6 来源 ↗
civil comments Noul 20.4 36.7 21.4 来源 ↗
impli Noul 20.0 88.0 46.7 来源 ↗
synthetic relevance / nanocoir Noul 16.1 73.0 33.6 来源 ↗
lexcomp Noul 16.0 52.0 22.5 来源 ↗
plane Noul 16.0 100.0 11.1 来源 ↗
open jev / customer control v1 Score 15.2 99.3 36.7 来源 ↗
synthetic relevance / nanobeir / nanoscidocs Noul 14.9 41.7 17.3 来源 ↗
scicite Choice 14.5 71.0 44.0 来源 ↗
argument quality Choice 14.3 95.9 37.4 来源 ↗
winowhy Noul 14.0 62.0 18.8 来源 ↗
followir core17 Noul 13.5 46.6 21.2 来源 ↗
fol nli Choice 13.1 54.1 10.4 来源 ↗
defeasible nli Choice 13.0 80.4 32.8 来源 ↗
aegis2 Noul 12.6 62.3 32.8 来源 ↗
open jev / snake v1 Choice 12.4 83.5 8.1 来源 ↗
synthetic relevance / nanobeir / nanoarguana Noul 10.1 48.2 15.6 来源 ↗
logical entailment Noul 10.0 72.0 12.2 来源 ↗
s1mb generalization diverse noul Noul 10.0 100.0 59.2 来源 ↗
open jev / customer control v1 Noul 9.9 100.0 58.0 来源 ↗
sgd Choice 8.5 96.6 24.4 来源 ↗
tracie Noul 8.0 34.0 8.7 来源 ↗
arc Choice 7.3 98.5 64.4 来源 ↗
gretel pii Choice 6.8 94.5 43.1 来源 ↗
miqa Choice 6.8 95.5 61.9 来源 ↗
babi nli Noul 6.0 84.0 37.0 来源 ↗
quartz Choice 4.8 88.1 49.1 来源 ↗
argument quality Noul 4.3 23.9 5.4 来源 ↗
arct Choice 4.3 89.4 39.4 来源 ↗
esci Choice 2.7 46.6 13.0 来源 ↗
ethics Noul 2.6 68.6 23.9 来源 ↗
laya / typed decisions Choice 2.4 34.5 18.9 来源 ↗
hans Noul 2.0 100.0 43.3 来源 ↗
s1mb generalization contextual noul Noul 2.0 98.0 62.0 来源 ↗
s1mb generalization contextual score Score 1.5 92.7 42.1 来源 ↗
laya / typed decisions Noul 0.8 47.8 24.4 来源 ↗
open jev / reasoning control v1 Noul 0.4 100.0 25.3 来源 ↗
aqua rat Choice 0.0 77.5 10.7 来源 ↗
bbq Choice 0.0 98.3 53.8 来源 ↗
boardgameqa Choice 0.0 62.5 15.6 来源 ↗
canttalk Noul 0.0 86.0 43.9 来源 ↗
cladder Noul 0.0 94.0 17.7 来源 ↗
contract nli Choice 0.0 76.0 31.8 来源 ↗
corr2cause Choice 0.0 91.7 3.8 来源 ↗
creak Noul 0.0 88.0 38.8 来源 ↗
crows pairs Choice 0.0 70.8 30.4 来源 ↗
ethics Choice 0.0 83.3 20.9 来源 ↗
few nerd Choice 0.0 86.4 43.9 来源 ↗
followir news21 Noul 0.0 42.7 14.9 来源 ↗
gsm8k Choice 0.0 78.6 23.5 来源 ↗
hh rlhf Choice 0.0 23.1 3.1 来源 ↗
laya / typed decisions Score 0.0 63.5 26.5 来源 ↗
lonli Choice 0.0 86.9 39.2 来源 ↗
nlsat Noul 0.0 22.0 2.1 来源 ↗
open jev / amount extraction control v1 Noul 0.0 100.0 33.3 来源 ↗
open jev / browser control v1 Choice 0.0 100.0 67.4 来源 ↗
open jev / citation control v1 Choice 0.0 100.0 40.9 来源 ↗
open jev / context retention control v1 Noul 0.0 100.0 19.0 来源 ↗
open jev / customer control v1 Choice 0.0 82.6 64.3 来源 ↗
open jev / drone control v1 Choice 0.0 100.0 4.8 来源 ↗
open jev / drone control v1 Noul 0.0 100.0 50.2 来源 ↗
open jev / drone control v1 Score 0.0 99.2 1.1 来源 ↗
open jev / email selection control v1 Choice 0.0 100.0 23.4 来源 ↗
open jev / email selection control v1 Noul 0.0 100.0 40.6 来源 ↗
open jev / entity alignment control v1 Noul 0.0 100.0 35.9 来源 ↗
open jev / entity alignment control v1 Score 0.0 100.0 8.1 来源 ↗
open jev / ir control v1 Choice 0.0 88.0 44.8 来源 ↗
open jev / ir control v1 Noul 0.0 100.0 47.3 来源 ↗
open jev / ir control v1 Score 0.0 98.6 10.1 来源 ↗
open jev / mailroom control v1 Noul 0.0 100.0 62.5 来源 ↗
open jev / painting geometry v1 Noul 0.0 100.0 27.2 来源 ↗
open jev / painting geometry v1 Score 0.0 100.0 15.2 来源 ↗
open jev / phone extraction control v1 Choice 0.0 100.0 32.4 来源 ↗
open jev / phone extraction control v1 Noul 0.0 100.0 50.4 来源 ↗
open jev / reasoning control v1 Choice 0.0 95.0 28.4 来源 ↗
open jev / reasoning control v1 Score 0.0 90.7 27.7 来源 ↗
open jev / silent failure control v1 Noul 0.0 100.0 38.8 来源 ↗
open jev / snake v1 Noul 0.0 100.0 9.0 来源 ↗
open jev / sponsor segment control v1 Choice 0.0 100.0 52.2 来源 ↗
open jev / tic tac toe v1 Choice 0.0 48.5 3.5 来源 ↗
open jev / vizdoom basic v1 Noul 0.0 100.0 39.0 来源 ↗
open jev / vizdoom basic v1 Score 0.0 100.0 29.6 来源 ↗
open jev / workflow controls v1 / agent trace observability Noul 0.0 100.0 27.8 来源 ↗
open jev / workflow controls v1 / customer service Noul 0.0 100.0 41.6 来源 ↗
open jev / workflow controls v1 / invoice processing Noul 0.0 100.0 34.6 来源 ↗
open jev / workflow controls v1 / security incidents Noul 0.0 100.0 37.8 来源 ↗
openbookqa Choice 0.0 97.2 53.7 来源 ↗
patent similarity Score 0.0 48.3 17.7 来源 ↗
poem sentiment Choice 0.0 57.1 4.0 来源 ↗
qasper Noul 0.0 88.0 41.6 来源 ↗
robust lr Choice 0.0 66.7 18.8 来源 ↗
ruletaker Noul 0.0 96.0 27.7 来源 ↗
s1mb generalization diverse score Score 0.0 95.3 42.8 来源 ↗
scone Choice 0.0 92.0 34.2 来源 ↗
spartqa Choice 0.0 38.7 4.4 来源 ↗
stepgame Choice 0.0 82.1 17.8 来源 ↗
synthetic relevance / nanobeir / nanoarguana Score 0.0 36.0 7.7 来源 ↗
synthetic relevance / nanobeir / nanodbpedia Score 0.0 44.9 17.6 来源 ↗
synthetic relevance / nanobeir / nanofiqa2018 Score 0.0 46.7 15.2 来源 ↗
synthetic relevance / nanobeir / nanohotpotqa Score 0.0 63.6 13.9 来源 ↗
synthetic relevance / nanobeir / nanomsmarco Score 0.0 52.6 17.8 来源 ↗
synthetic relevance / nanobeir / nanonfcorpus Score 0.0 38.6 10.7 来源 ↗
synthetic relevance / nanobeir / nanonq Score 0.0 49.4 9.3 来源 ↗
synthetic relevance / nanobeir / nanoscidocs Score 0.0 36.7 7.2 来源 ↗
synthetic relevance / nanobeir / nanotouche2020 Score 0.0 44.3 20.6 来源 ↗
synthetic relevance / nanocoir Score 0.0 44.3 4.3 来源 ↗
temporal nli Choice 0.0 80.7 18.3 来源 ↗
ud ewt Choice 0.0 56.7 13.2 来源 ↗
winobias Choice 0.0 91.3 30.1 来源 ↗
wiqa Choice 0.0 58.7 19.3 来源 ↗

同族条目(2)

型号参数Decision Index S1MBVision榜单
convaiinnovations/laya-multilingual 当前 421M(活跃 370M) — 9.3 — s1mb
convaiinnovations/laya-typed-decisions 421M(活跃 370M) 4.4 13.4 — s1mb, di

来自 systemonemodels.org 的详细介绍

What Laya is

Laya is a System One model from Convai Innovations, an Indian company led by CEO Nandakishor M, who publishes the code. It reads a state, which can be text, an email, a ticket or a JSON document, and answers typed questions about it: Choice, Score and Noul, the same three shapes Jev uses. Every question in a call is answered in one forward pass, and nothing is generated, so there is no output to parse. Unlike Jev, Laya is not a hosted API. It ships as Apache 2.0 weights on Hugging Face and a Python package, pip install laya, and runs on your own GPU or CPU.

How it is built

Laya is an encoder with a decision head on top. The English checkpoint fully fine-tunes ModernBERT-large (395M parameters) and adds a head trained from scratch: two transformer layers, a scorer that reads one [MASK] token per option, and an act-or-escalate output. That comes to 421M parameters. The multilingual checkpoint uses mmBERT-base and comes to 322M. Options are defined per request, so a new question schema needs no retraining. Training uses what Convai also calls RLCD. The reward is a strictly proper scoring rule, so reporting honest probabilities earns the most reward, and updates use REINFORCE with a group-mean baseline. TypeSafe has not published its recipe. Convai has, along with a fine-tuning notebook that takes 4 to 5 hours on Kaggle’s free pair of T4 GPUs. A Router picks between the checkpoints by detecting the script and language before the forward pass. The English checkpoint fails on other scripts without losing confidence: on Khmer it scored 0.000 at 0.952 confidence.

What it is good at

Its strengths are speed and ownership. Convai measures 32.8ms for one question on a T4, and you can keep, inspect and retrain the weights. After fitting a temperature per question type on held-out data, its calibration error falls from 0.466 to 0.081. The package ships ready-made question sets for model routing, prompt guardrails, moderation and ticket triage. The Decision Index 0.2.1 by multimodalart, updated 28 September 2026, tests Laya zero-shot across a broad suite. On a chance-corrected score where 0 is random guessing and 100 is perfect, averaged over 38 benchmarks in five weighted areas, Laya scores 6.04 against Jev’s 57.91, close to random guessing. That fits Convai’s own advice to specialise it on your task rather than use it zero-shot.

What it is not for

The base checkpoints are close to chance on the typed-decisions benchmark without fine-tuning: 0.362 against a 0.318 random baseline. Convai’s own README calls Laya “a fast base to specialise, not a zero-shot decision engine.” Score questions are its weakest type, and large label sets need extra configuration. Probabilities are over-confident until you fit temperatures on your own data. Convai’s comparisons with Jev set its own numbers against Jev figures published by third parties, not measured in the same run, so treat the head-to-head as indicative.

Specifications

Question types Choice Score Noul Max Choice optionsNot documented Score levelsNot documented Questions per callNot documented Total context512 tokens State budget320 tokens Rate limitNone. It runs on your own hardware. SDKsPython: laya The English checkpoint defaults to 512 tokens, 192 of them reserved for the question and its options, leaving about 320 for the state. laya-multilingual and laya-typed-decisions default to 1,024 tokens with 256 for options. The mmBERT encoder under the multilingual checkpoint accepts up to 8,192 if you raise max_len. There is no hard cap on Choice options, but they share the option budget: 77 labels get about 3 to 4 tokens each, and Banking77 accuracy falls to 0.425. Raise head_max_len or shortlist with predict_shortlist for large label sets.

Versions

convaiinnovations/laya, 18 Sep 2026, English checkpoint at the repo root. ModernBERT-large encoder, 421M parameters, 512-token context. Release notes convaiinnovations/laya-typed-decisions, 18 Sep 2026, The same 421M English model fine-tuned on the typed-decisions benchmark's training split, with a 1,024-token context. Release notes convaiinnovations/laya-multilingual, 19 Sep 2026, mmBERT-base encoder, 322M parameters, 1,024-token context, for text outside English. Release notes

Use cases

What people use Laya for, one page per pattern. Workflow controlStarter

Support inbox triage with System One models

Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket. Choice Score Noul Workflow controlIntermediate

Intent and model routing with System One models

One Jev call reads an incoming request and returns its intent as a label plus a difficulty rating on a scale you wrote. Your router reads both numbers and picks the handler: deterministic code, a cheap model, an expensive one, or a human queue. Choice Score Safety and qualityIntermediate

LLM guardrails with System One models

Put one Jev request in front of an LLM and one behind it. Yes/no questions return the probability that each hazard holds, a Score rates how much harm complying would do, and your thresholds turn those numbers into pass, review, block, or a crisis path. Noul Score Workflow controlIntermediate

Confidence-gated actions with System One models

Jev returns a confidence value from 0 to 1 alongside every Choice and Score answer. Your code treats it as a separate axis: act automatically when it's high, confirm or flag when it's middling, hand the decision to a person when it's low. Riskier actions get higher bars. Choice Score

Examples built with Laya

The most-starred and most-viewed entries in the directory. Browse all examples.