Skip to content
JevDirectory.org
CommunityPractices & Patterns6 starsVerified 2026-09-22

jev-search-rerank-eval

A graded relevance evaluation of Jev as a reranker: 9,831 labelled pairs from 164 queries over a 33,047-item skills catalog, comparing Jev score reranks with BM25, bge-m3, and rank fusion.

Category
Practices & Patterns
Published by
Community
Author
zhuyansen
Added
2026-09-22
Tagscommunitypythonsearchrerankingevaluationbenchmarks

Highlights

  • Jev alone does not beat a good embedding ranker: jev-score(bge-m3@30) is +0.012 NDCG@10 [-0.013, +0.037] and -0.028 under LLM-only labels.
  • The fusion rrf(bge-m3, jev@30) is the best system at 0.864 NDCG@10, +0.090 over bge-m3, and still +0.064 under LLM-only labels.
  • Jev's edge under its own labels (+0.053) is measurable judge circularity: it reads +0.012 on merged labels and -0.028 on LLM-only labels.
  • On a weak lexical list Jev is the better reranker: jev-score(ash@30) beats bge-m3(ash@30) by +0.060, or +0.037 under LLM-only labels.
  • Total API spend was about $2.6, and committed frozen artifacts let scoring and the robustness analysis reproduce without API access.

Quickstart

bash
uv venv --python 3.12 .venv && uv pip install --python .venv/bin/python -e ".[dev]"
.venv/bin/python -m jse robustness   # re-scores cached runs, no API key needed

Watch out

MIT-licensed and not affiliated with TypeSafe; the ash scoring functions are copied verbatim from @agentskillshub/cli v0.4.0, and catalog data belongs to its respective authors.

More like this

5GitHub stars
Reranking benchmark that gave Jev, Cohere Rerank 4 Pro, ZeroEntropy zerank-2, and DeepSeek the same thirty BM25 candidates across eight English datasets, publishing saved responses, scoring code, and paired-bootstrap intervals. Jev's rubric scored 0.692 nDCG@10 against Cohere Pro's 0.691.
Practices & Patterns#community#python#search
208GitHub stars
A self-hosted implementation of TypeSafe's Jev System One API powered by the 400M-parameter GLiFormer encoder: it serves choice, score, and noul and drops into the official typesafe-sdk via TYPESAFE_BASE_URL, but trails Jev on reasoning-heavy tasks.
Practices & Patterns#community#python#open-models
Communityjeff
74GitHub stars
Independent cross-model benchmark for Jev-class decision models, running 534 frozen cases per complete entrant with scoring code and a four-axis score of accuracy, calibration, latency and cost.
Practices & Patterns#community#python#benchmarks
Communityjevbench
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

A playable 16-judgment demo

AI multiple choice, not essay writing

Screening agent actions with Jev

Tested TypeSafe’s Jev (no-text, probability-only model) as an AI agent safety monitor. Checking each action first worked well caught most attacks with almost no false blocks, and much faster than Gemini.

Image
Image
Image
Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

Cua's small System One models

A 706K-parameter form filler

cua open sourced a 706k param model that fills a whole form in one 50ms pass the llm agent doing the same form took 23 turns and 39.6 seconds the specialists are going to eat the generalists from the bottom

Cua
Cua
@trycua

1/ Introducing CUA-S1: a family of System One Models, small, specialized, and built for computer use. Today we're open-sourcing CUA-S1-FORMS, the first in the family: github.com/trycua/cua

Image
Reply

Navigating Neo4j with Jev

Jev 这个 waitlist 还是很给力的,昨天申请,今天就能用上。 给已经拿到 API、但还不知道怎么玩的人整理了一份 Awesome Jev,目前我能确认到的 Jev 项目基本都在这里: 1. jev-ultrafast Browser Use 做的高速浏览器 Agent。Jev Show more

Image
思维怪怪
思维怪怪
@0xLogicrw

前 OpenAI 研究员 Diogo Almeida 创办的 TypeSafe AI 推出新模型 Jev。它有点像一个能读懂自然语言的超级分类器,不生成文本,只返回选项、分数和概率,专门给软件做判断。 普通大模型需要一个 token 一个 token 往外生成,Jev 则可以并行给出多个结果。TypeSafe 还用新的 RLCD

Reply