typesafe's jev is fun! live demo you can play with: typesafe-demo.val.run
jev-search-rerank-eval
A graded relevance evaluation of Jev as a reranker: 9,831 labelled pairs from 164 queries over a 33,047-item skills catalog, comparing Jev score reranks with BM25, bge-m3, and rank fusion.
- Category
- Practices & Patterns
- Published by
- Community
- Author
- zhuyansen
- Added
- 2026-09-22
Highlights
- Jev alone does not beat a good embedding ranker: jev-score(bge-m3@30) is +0.012 NDCG@10 [-0.013, +0.037] and -0.028 under LLM-only labels.
- The fusion rrf(bge-m3, jev@30) is the best system at 0.864 NDCG@10, +0.090 over bge-m3, and still +0.064 under LLM-only labels.
- Jev's edge under its own labels (+0.053) is measurable judge circularity: it reads +0.012 on merged labels and -0.028 on LLM-only labels.
- On a weak lexical list Jev is the better reranker: jev-score(ash@30) beats bge-m3(ash@30) by +0.060, or +0.037 under LLM-only labels.
- Total API spend was about $2.6, and committed frozen artifacts let scoring and the robustness analysis reproduce without API access.
Quickstart
uv venv --python 3.12 .venv && uv pip install --python .venv/bin/python -e ".[dev]"
.venv/bin/python -m jse robustness # re-scores cached runs, no API key neededWatch out
MIT-licensed and not affiliated with TypeSafe; the ash scoring functions are copied verbatim from @agentskillshub/cli v0.4.0, and catalog data belongs to its respective authors.
More like this
From the community
Posts from builders shipping with Jev right now.
A playable 16-judgment demo
AI multiple choice, not essay writing
WTF is Jev by @typesafeai? Here’s the tl;dr ELI5: Think AI multiple choice, not AI essay writing. It doesn’t chat. It makes decisions your software can act on: “Spam or not?” “Which tool should this agent use?” “Does this need a human?” The exciting part: roughly 200x faster Show more
Screening agent actions with Jev
Tested TypeSafe’s Jev (no-text, probability-only model) as an AI agent safety monitor. Checking each action first worked well caught most attacks with almost no false blocks, and much faster than Gemini.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x
Cua's small System One models
1/ Introducing CUA-S1: a family of System One Models, small, specialized, and built for computer use. Today we're open-sourcing CUA-S1-FORMS, the first in the family: github.com/trycua/cua
A 706K-parameter form filler
cua open sourced a 706k param model that fills a whole form in one 50ms pass the llm agent doing the same form took 23 turns and 39.6 seconds the specialists are going to eat the generalists from the bottom
1/ Introducing CUA-S1: a family of System One Models, small, specialized, and built for computer use. Today we're open-sourcing CUA-S1-FORMS, the first in the family: github.com/trycua/cua
Navigating Neo4j with Jev
Jev 这个 waitlist 还是很给力的,昨天申请,今天就能用上。 给已经拿到 API、但还不知道怎么玩的人整理了一份 Awesome Jev,目前我能确认到的 Jev 项目基本都在这里: 1. jev-ultrafast Browser Use 做的高速浏览器 Agent。Jev Show more
前 OpenAI 研究员 Diogo Almeida 创办的 TypeSafe AI 推出新模型 Jev。它有点像一个能读懂自然语言的超级分类器,不生成文本,只返回选项、分数和概率,专门给软件做判断。 普通大模型需要一个 token 一个 token 往外生成,Jev 则可以并行给出多个结果。TypeSafe 还用新的 RLCD
