Ask Jev anything. Give it a try at askjev.ai It won't answer. It will judge. Let's see if we can get to 1 million questions. @typesafeai 🤝 @convex work great together. @hmartenjoyer @CompleteSkeptic @justKDeng @mikeysee
jev-eval-agent
A personal-assistant agent with 100 mocked tools that measures how many steps a task takes when the LLM picks the tool versus when Jev picks it before every model step.
- Category
- Practices & Patterns
- Published by
- Community
- Author
- vinilana
- Added
- 2026-09-22
Highlights
- AGENT_MODE switches between llm-direct, where the LLM sees all 100 tools each step, and jev-classifier, where Jev picks before every model step.
- Confidence-gated routing: if Jev picks respond_to_user but done is below JEV_DONE_THRESHOLD (0.5), the reply is blocked and the next-best tool is exposed.
- Six tasks (three basic, three complex) include look-alike distractor tools; calling one counts as a wrong tool in the report.
- Per-run metrics include modelSteps, toolCalls, distractorCalls, completion, efficiency, token usage, and estimated OpenRouter cost.
Quickstart
cp .env.example .env.local # OPENROUTER_API_KEY and TYPESAFE_API_KEY
npm install
EVAL_MODEL=deepseek-v4.1-flash npm run evalWatch out
No license file, so reuse terms are unclear. Needs Node.js, an OPENROUTER_API_KEY and a TYPESAFE_API_KEY, and the benchmark prompts are in Portuguese.
More like this
From the community
Posts from builders shipping with Jev right now.
A million judged questions
Inferring Jev's internals from 1,000 calls
Jevの内部アーキテクチャを推測している技術記事(Jev’s Architecture Unmasked)からメモ。 ・本記事はJevのAPIを約1万回の呼び出して、内部構造を推測したもの ・従来の言語モデルを用いた分類やルーティングでは、トークンを1文字ずつ逐次生成するために膨大な無駄な計算コストが発生していた。 Show more
The open System One roundup
Jev 发布没几天,开源社区已经开始疯狂复刻了🔥 最值得推荐的五个模型: 1、Laya 421M:原生决策模型,支持 Mac 2、Decider-2B:最像 Jev,基于 Qwen3.5 3、NanoJev 0.6B:专门的 Decision Head 4、Reflex:Qwen3.5 + Direct Logits 5、System-One 4B:专门做概率校准 Show more
Jev 刚发布没几天,开源社区就出现了同款🔥 Decider-2B模型,是基于 Qwen3.5-2B 做了特殊调整 它和 Jev 模型是一样的 只做选择 评分和判断 不是文本类的 LLM 模型 但两者还是有几个明显区别: 1、模型 Jev:闭源 System One Model Decider:Qwen3.5-2B,约 1.9B 参数,Apache 2.0 开源 2、价格
The launch post
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x Show more
Trading bot, one decision per block
I built a trading bot with Jev! Jev decides if it should "buy" or "sell", given the price feed of an asset pair, and executes real trades. It uses Monad to place the orders on Kuru's on-chain order book in every 300ms block. Demo link → jev-trader.vercel.app
Classifying 1,500 real emails
this model is actually insane at email classification i tested it on 1500 of my own emails to see how well it works and I am blown away
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x



