Ran @typesafeai's Jev against an existing classifier eval that previously used Gemini 2.5 Flash Lite. It won both on quality (saturated the eval) and speed (6x)
Zero-shot email classification with TypeSafe's Jev
An exploratory study of zero-shot ham, spam, and phishing classification with Jev: enriched email context raised main-set accuracy from 93.62% to 98.64%, within 0.23 points of a trained TF-IDF baseline.
- Category
- Practices & Patterns
- Published by
- Community
- Author
- bitnovus
- Added
- 2026-09-22
Highlights
- On 9,886 messages, adding link destinations, Reply-To, and attachment metadata raised accuracy from 93.62% to 97.98% with the question unchanged.
- Phishing recall rose from 85.71% to 98.43%, while legitimate messages flagged as phishing stayed at one.
- With evidence-focused wording the main set reached 98.64%, 0.23 points behind a TF-IDF logistic regression trained on about 4,600 labeled messages per fold.
- On 853 phishing messages from 2024-25, enriched Jev caught 95.31% versus 75.26% for regression trained on older mail.
Quickstart
export TYPESAFE_API_KEY=YOUR_API_KEY
uv run python experiments/jev-context/evaluate.py --limit 3Watch out
MIT-licensed. Results are exploratory on jev-1.13.0 and public corpora that may have leaked into pretraining, and dataset licenses are separate from the code license.
More like this
From the community
Posts from builders shipping with Jev right now.
Beating Gemini Flash Lite on an eval
Browser Use Ultrafast, powered by Jev
Breaking: Browser Use + Jev = Ultrafast ⚡ Findings flights took 7s and cost only $0.0039 🤯 > new action space every step > DOM state space > small LLM fallback to type (this video is at 1x speed btw) Built a tiny open source browser agent. try it below ↓
A really smart switch statement
hype-free explanation of jev: jev does not replace gpt / claude jev is just a *really* smart switch statement like if 2016 ml classifiers got 2026 levels of intelligence it's a new* type of tool that will make a lot of workloads insanely fast, cheap, and accurate * = and by Show more
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x
When a designer gets Jev
when a designer gets access to Jev
Full Jev video tutorial
Full Jev Tutorial What it is, how you can build with it and what new applications it can unlock → 0:00 Intro → 0:34 Jev explained → 4:06 API setup → 5:59 Demo 1: Voice-controlled browser → 11:33 Demo 2: AI memory → 17:27 Demo 3: YouTube predictor
The case against Jev-scored compaction
This is a terrible compaction strategy that fundamentally doesn't understand how compaction and context management work. Seems like a lot of people are confused so let's break this down. 1. Compaction isn't a filter The role of compaction is to clean up history to keep the Show more
found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant



