I made a DuckDB extension where you can use @typesafeai 's Jev to do quick classification of rows in any csv/parquet file or duckdb table about 10sec for 1k rows ~ better than using an LLM, way more ergonomic than a classifier game-changing for data analysis!
OpenJev (Verdict)
An encoder-side reproduction of Jev built on a retrained GLiClass ModernBERT base (151M): it evaluates multiple typed questions in one non-autoregressive forward pass, returning choices, ordinal scores, and probabilities under 35 ms.
- Category
- Practices & Patterns
- Published by
- Community
- Author
- Heman10x-NGU
- Added
- 2026-09-22
Highlights
- On 231 public JevBench tasks, inference-only fixes raised standard-tier accuracy from 62.5% to 69.4% and cut hard-tier ECE from 0.298 to 0.118.
- The weights are byte-identical to the published checkpoint heman10x/rlcd-modernbert-151m; the v1.4 changes are all inference-engine fixes.
- Candidate labels are framed as NLI hypothesis sentences ('It is {description}') to match the GLiClass backbone's pretraining.
- The context budget was cut from 1024 to 512 tokens after training on states under 71 tokens, avoiding out-of-distribution positional drift.
- Hard-tier accuracy stayed at 36.9% across the update, while probability fidelity rose 10 points to 72.8.
Quickstart
pip install -e .Watch out
Licensed as Other (NOASSERTION), so reuse terms need checking; requires Python 3.10+, and the headline numbers are self-reported on the project's own benchmark slice.
More like this
From the community
Posts from builders shipping with Jev right now.
Classifying rows in DuckDB
A playable 16-judgment demo
typesafe's jev is fun! live demo you can play with: typesafe-demo.val.run
AI multiple choice, not essay writing
WTF is Jev by @typesafeai? Here’s the tl;dr ELI5: Think AI multiple choice, not AI essay writing. It doesn’t chat. It makes decisions your software can act on: “Spam or not?” “Which tool should this agent use?” “Does this need a human?” The exciting part: roughly 200x faster Show more
Screening agent actions with Jev
Tested TypeSafe’s Jev (no-text, probability-only model) as an AI agent safety monitor. Checking each action first worked well caught most attacks with almost no false blocks, and much faster than Gemini.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x
Cua's small System One models
1/ Introducing CUA-S1: a family of System One Models, small, specialized, and built for computer use. Today we're open-sourcing CUA-S1-FORMS, the first in the family: github.com/trycua/cua
A 706K-parameter form filler
cua open sourced a 706k param model that fills a whole form in one 50ms pass the llm agent doing the same form took 23 turns and 39.6 seconds the specialists are going to eat the generalists from the bottom
1/ Introducing CUA-S1: a family of System One Models, small, specialized, and built for computer use. Today we're open-sourcing CUA-S1-FORMS, the first in the family: github.com/trycua/cua
