Skip to content
JevDirectory.org
CommunityPractices & Patterns104 starsVerified 2026-09-22

jev-eval-agent

A personal-assistant agent with 100 mocked tools that measures how many steps a task takes when the LLM picks the tool versus when Jev picks it before every model step.

Category
Practices & Patterns
Published by
Community
Author
vinilana
Added
2026-09-22
Tagscommunitytypescriptagentsevaluationbenchmarksopenroutertool-calling

Highlights

  • AGENT_MODE switches between llm-direct, where the LLM sees all 100 tools each step, and jev-classifier, where Jev picks before every model step.
  • Confidence-gated routing: if Jev picks respond_to_user but done is below JEV_DONE_THRESHOLD (0.5), the reply is blocked and the next-best tool is exposed.
  • Six tasks (three basic, three complex) include look-alike distractor tools; calling one counts as a wrong tool in the report.
  • Per-run metrics include modelSteps, toolCalls, distractorCalls, completion, efficiency, token usage, and estimated OpenRouter cost.

Quickstart

bash
cp .env.example .env.local     # OPENROUTER_API_KEY and TYPESAFE_API_KEY
npm install
EVAL_MODEL=deepseek-v4.1-flash npm run eval

Watch out

No license file, so reuse terms are unclear. Needs Node.js, an OPENROUTER_API_KEY and a TYPESAFE_API_KEY, and the benchmark prompts are in Portuguese.

More like this

91GitHub stars
A comparison arena that runs the same batch of review comments through Jev and DeepSeek, showing processing time, cost, and per-label results with CSV/Excel import, replay, and offline reports.
Practices & Patterns#community#javascript#benchmarks
Communityjev-arena
38GitHub stars
Side-by-side benchmark of TypeSafe Jev, Qwen 3.8 27B on Cerebras, and a local Needle 3 across seven synthetic workloads, recording validated outputs, mistakes, latency, tokens, and estimated cost. Raw exports and per-scene limitations are published.
Practices & Patterns#community#typescript#benchmarks
A benchmark that runs Jev on all 6,257 text traces of Who&When Pro to attribute agent failures, scoring it with the official pinned scorer against the paper's GPT-5.4, Claude Sonnet 4.6, GLM-5, and Qwen3.5-122B baselines.
Practices & Patterns#community#python#benchmarks
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

A million judged questions

Inferring Jev's internals from 1,000 calls

Jevの内部アーキテクチャを推測している技術記事(Jev’s Architecture Unmasked)からメモ。 ・本記事はJevのAPIを約1万回の呼び出して、内部構造を推測したもの ・従来の言語モデルを用いた分類やルーティングでは、トークンを1文字ずつ逐次生成するために膨大な無駄な計算コストが発生していた。 Show more

Reply

The open System One roundup

Jev 发布没几天,开源社区已经开始疯狂复刻了🔥 最值得推荐的五个模型: 1、Laya 421M:原生决策模型,支持 Mac 2、Decider-2B:最像 Jev,基于 Qwen3.5 3、NanoJev 0.6B:专门的 Decision Head 4、Reflex:Qwen3.5 + Direct Logits 5、System-One 4B:专门做概率校准 Show more

小墨同学
小墨同学
@xiaomovps

Jev 刚发布没几天,开源社区就出现了同款🔥 Decider-2B模型,是基于 Qwen3.5-2B 做了特殊调整 它和 Jev 模型是一样的 只做选择 评分和判断 不是文本类的 LLM 模型 但两者还是有几个明显区别: 1、模型 Jev:闭源 System One Model Decider:Qwen3.5-2B,约 1.9B 参数,Apache 2.0 开源 2、价格

Image
Reply

The launch post

Trading bot, one decision per block

Classifying 1,500 real emails

this model is actually insane at email classification i tested it on 1500 of my own emails to see how well it works and I am blown away

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply