Skip to content
JevDirectory.org
CommunityPractices & Patterns2 starsVerified 2026-09-22

jev-agent-failure-benchmark

A benchmark that runs Jev on all 6,257 text traces of Who&When Pro to attribute agent failures, scoring it with the official pinned scorer against the paper's GPT-5.4, Claude Sonnet 4.6, GLM-5, and Qwen3.5-122B baselines.

Category
Practices & Patterns
Published by
Community
Author
TokenTrim
Added
2026-09-22
Tagscommunitypythonbenchmarksevaluationagents

Highlights

  • Jev outperformed GPT-5.4 on every axis at about $1.28 total: Who 73.4, When 76.4, and error-type F1 23.7, the best in the field.
  • Jev answers three typed Choice questions per trace (responsible agent, decisive step, error mode), each with a calibrated probability distribution.
  • The official whowhen_eval scorer is pinned at commit 14369dcb so rows are directly comparable with the paper's numbers.
  • Who and When are constrained-choice for Jev while the LLM baselines free-generate, so the like-for-like axis is What, the error type over a shared taxonomy.
  • Ground-truth labels never enter Jev's input, and a leakage test enforces it; runs are resumable and record failures rather than dropping them.

Quickstart

bash
uv venv && uv pip install -e ".[dev]"
mkdir -p data
REV=0bd196c8a040841c4ae167ab33cc8151de246f1f
curl -L "https://huggingface.co/datasets/Leoxx/whowhen_pro/resolve/$REV/data/text.jsonl" -o data/text.jsonl
curl -L "https://huggingface.co/datasets/Leoxx/whowhen_pro/resolve/$REV/taxonomy.yaml" -o data/taxonomy.yaml
jevbench sample --n 300 --out results/run/sample.json
jevbench run --sample results/run/sample.json
jevbench report

Watch out

Apache-2.0 code; the CC-BY-4.0 Who&When Pro dataset is not included and must be fetched, failures are injected by a controlled pipeline rather than being natural incidents, and running needs Python with uv plus a TYPESAFE_API_KEY.

More like this

2GitHub stars
A reproducible harness that scores Jev Ultrafast research-browser runs over 11 baseline cases plus human and quant stress suites with CoS-locked QC grades, regenerating the field note and trace notebooks offline.
Practices & Patterns#community#python#evaluation
208GitHub stars
A self-hosted implementation of TypeSafe's Jev System One API powered by the 400M-parameter GLiFormer encoder: it serves choice, score, and noul and drops into the official typesafe-sdk via TYPESAFE_BASE_URL, but trails Jev on reasoning-heavy tasks.
Practices & Patterns#community#python#open-models
Communityjeff
104GitHub stars
A personal-assistant agent with 100 mocked tools that measures how many steps a task takes when the LLM picks the tool versus when Jev picks it before every model step.
Practices & Patterns#community#typescript#agents
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Inferring Jev's internals from 1,000 calls

Jevの内部アーキテクチャを推測している技術記事(Jev’s Architecture Unmasked)からメモ。 ・本記事はJevのAPIを約1万回の呼び出して、内部構造を推測したもの ・従来の言語モデルを用いた分類やルーティングでは、トークンを1文字ずつ逐次生成するために膨大な無駄な計算コストが発生していた。 Show more

Reply

The open System One roundup

Jev 发布没几天,开源社区已经开始疯狂复刻了🔥 最值得推荐的五个模型: 1、Laya 421M:原生决策模型,支持 Mac 2、Decider-2B:最像 Jev,基于 Qwen3.5 3、NanoJev 0.6B:专门的 Decision Head 4、Reflex:Qwen3.5 + Direct Logits 5、System-One 4B:专门做概率校准 Show more

小墨同学
小墨同学
@xiaomovps

Jev 刚发布没几天,开源社区就出现了同款🔥 Decider-2B模型,是基于 Qwen3.5-2B 做了特殊调整 它和 Jev 模型是一样的 只做选择 评分和判断 不是文本类的 LLM 模型 但两者还是有几个明显区别: 1、模型 Jev:闭源 System One Model Decider:Qwen3.5-2B,约 1.9B 参数,Apache 2.0 开源 2、价格

Image
Reply

The launch post

Trading bot, one decision per block

Classifying 1,500 real emails

this model is actually insane at email classification i tested it on 1500 of my own emails to see how well it works and I am blown away

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

Fast browser use with Stagehand

we built blazing fast computer/browser use with Jev + @Stagehanddev. this task cost $0.001 and executed at near instant speed (in a remote browser btw) the loop: observe the page, send a11y tree as state + actions as questions, Jev decides the next action, then Stagehand Show more

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply