Skip to content
JevDirectory.org
CommunityPractices & Patterns0 starsVerified 2026-09-22

Zero-shot email classification with TypeSafe's Jev

An exploratory study of zero-shot ham, spam, and phishing classification with Jev: enriched email context raised main-set accuracy from 93.62% to 98.64%, within 0.23 points of a trained TF-IDF baseline.

Category
Practices & Patterns
Published by
Community
Author
bitnovus
Added
2026-09-22
Tagscommunitypythonevaluationclassificationbenchmarks

Highlights

  • On 9,886 messages, adding link destinations, Reply-To, and attachment metadata raised accuracy from 93.62% to 97.98% with the question unchanged.
  • Phishing recall rose from 85.71% to 98.43%, while legitimate messages flagged as phishing stayed at one.
  • With evidence-focused wording the main set reached 98.64%, 0.23 points behind a TF-IDF logistic regression trained on about 4,600 labeled messages per fold.
  • On 853 phishing messages from 2024-25, enriched Jev caught 95.31% versus 75.26% for regression trained on older mail.

Quickstart

bash
export TYPESAFE_API_KEY=YOUR_API_KEY
uv run python experiments/jev-context/evaluate.py --limit 3

Watch out

MIT-licensed. Results are exploratory on jev-1.13.0 and public corpora that may have leaked into pretraining, and dataset licenses are separate from the code license.

More like this

208GitHub stars
A self-hosted implementation of TypeSafe's Jev System One API powered by the 400M-parameter GLiFormer encoder: it serves choice, score, and noul and drops into the official typesafe-sdk via TYPESAFE_BASE_URL, but trails Jev on reasoning-heavy tasks.
Practices & Patterns#community#python#open-models
Communityjeff
74GitHub stars
Independent cross-model benchmark for Jev-class decision models, running 534 frozen cases per complete entrant with scoring code and a four-axis score of accuracy, calibration, latency and cost.
Practices & Patterns#community#python#benchmarks
Communityjevbench
An experiment comparing Jev with GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6 as evaluators of five frozen weather-agent runs, measuring pass-or-fail accuracy against human labels plus variance, cost, and latency.
Practices & Patterns#community#python#evaluation
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Beating Gemini Flash Lite on an eval

Browser Use Ultrafast, powered by Jev

A really smart switch statement

hype-free explanation of jev: jev does not replace gpt / claude jev is just a *really* smart switch statement like if 2016 ml classifiers got 2026 levels of intelligence it's a new* type of tool that will make a lot of workloads insanely fast, cheap, and accurate * = and by Show more

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

When a designer gets Jev

Full Jev video tutorial

The case against Jev-scored compaction

This is a terrible compaction strategy that fundamentally doesn't understand how compaction and context management work. Seems like a lot of people are confused so let's break this down. 1. Compaction isn't a filter The role of compaction is to clean up history to keep the Show more

tamara
tamara
@tamarajtran

found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant

Reply