Ran @typesafeai's Jev against an existing classifier eval that previously used Gemini 2.5 Flash Lite. It won both on quality (saturated the eval) and speed (6x)
An early-access test of TypeSafe's Jev: calibrated judgments for half a cent
An independent trial running Jev on 24 Norwegian resource-tax hearing responses with eleven typed questions each, then comparing agreement, calibration, latency and cost against DeepSeek V4.1 Flash.
- Category
- Practices & Patterns
- Published by
- Community
- Author
- —
- Added
- 2026-09-22
Highlights
- Jev answered eleven questions per document for 24 documents at a half cent total, median 0.32s versus 2.7s for DeepSeek with reasoning off.
- On the ordered substance Score it beat DeepSeek 19 of 24 against 14, the one clear accuracy difference in the trial.
- Calibration on 192 argument judgments was directionally right: 0% yes in the 0.0-0.1 bin and 98% in the 0.9-1.0 bin, slightly underconfident.
- Rewriting short draft questions with careful qualifiers cut agreement 0.89 to 0.86 and raised calibration error from 0.040 to 0.116.
- Norwegian runs about 2.06 characters per token, making Jev's 32k-token state limit roughly 64,000 characters.
Watch out
Twenty-four documents is a first look, not a benchmark, and reference labels come from Claude Fable 5.1, so results measure agreement with a frontier model.
More like this
From the community
Posts from builders shipping with Jev right now.
Beating Gemini Flash Lite on an eval
Browser Use Ultrafast, powered by Jev
Breaking: Browser Use + Jev = Ultrafast ⚡ Findings flights took 7s and cost only $0.0039 🤯 > new action space every step > DOM state space > small LLM fallback to type (this video is at 1x speed btw) Built a tiny open source browser agent. try it below ↓
A really smart switch statement
hype-free explanation of jev: jev does not replace gpt / claude jev is just a *really* smart switch statement like if 2016 ml classifiers got 2026 levels of intelligence it's a new* type of tool that will make a lot of workloads insanely fast, cheap, and accurate * = and by Show more
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x
When a designer gets Jev
when a designer gets access to Jev
Full Jev video tutorial
Full Jev Tutorial What it is, how you can build with it and what new applications it can unlock → 0:00 Intro → 0:34 Jev explained → 4:06 API setup → 5:59 Demo 1: Voice-controlled browser → 11:33 Demo 2: AI memory → 17:27 Demo 3: YouTube predictor
The case against Jev-scored compaction
This is a terrible compaction strategy that fundamentally doesn't understand how compaction and context management work. Seems like a lot of people are confused so let's break this down. 1. Compaction isn't a filter The role of compaction is to clean up history to keep the Show more
found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant



