Skip to content
JevDirectory.org
CommunityPractices & Patterns2 starsVerified 2026-09-22

jev-sec-bench

Blind security benchmarks for Jev on public corpora: 662 prompt-injection messages and 200 matched vulnerable-code pairs, with raw per-sample results and a TUI dashboard.

Category
Practices & Patterns
Published by
Community
Author
Gaurav-Gosain
Added
2026-09-22
Tagscommunitygobenchmarkssecurityevaluationcalibration

Highlights

  • Prompt-injection run at a plain 0.50 cut: 96.5% accuracy, 95.1% recall, ROC-AUC 0.9927, ECE 0.0588, 22.7s for 662 messages at p50 325 ms.
  • Adding deployment context to the state raised injection recall from 74.9% to 95.1% without any threshold tuning.
  • In 200 matched code pairs the vulnerable half scored above its secure twin 178 times (89.0%), including 43/46 SQL injection pairs.
  • The audit found at least 38% of apparent false positives are corpus label errors, counting only what a regex can prove.

Quickstart

bash
export TYPESAFE_API_KEY=YOUR_API_KEY
go run ./cmd/jev-sec-bench -bench all
go run ./cmd/jev-tui

Watch out

MIT-licensed. Needs Go and a TYPESAFE_API_KEY; both corpora are public and may be in training data, results are a single run, and thresholds should be validated on your own labelled data.

More like this

A public benchmark comparing Jev with Claude Haiku 4.5 on whether an email agent should click a link, over 2,000 PhishNChips emails, with calibration, latency, cost, and five signal questions.
Practices & Patterns#community#python#security
74GitHub stars
Independent cross-model benchmark for Jev-class decision models, running 534 frozen cases per complete entrant with scoring code and a four-axis score of accuracy, calibration, latency and cost.
Practices & Patterns#community#python#benchmarks
Communityjevbench
An experiment comparing Jev with GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6 as evaluators of five frozen weather-agent runs, measuring pass-or-fail accuracy against human labels plus variance, cost, and latency.
Practices & Patterns#community#python#evaluation
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Jev lands on OpenRouter

700 leads scored for $0.09

Beating Gemini Flash Lite on an eval

Browser Use Ultrafast, powered by Jev

A really smart switch statement

hype-free explanation of jev: jev does not replace gpt / claude jev is just a *really* smart switch statement like if 2016 ml classifiers got 2026 levels of intelligence it's a new* type of tool that will make a lot of workloads insanely fast, cheap, and accurate * = and by Show more

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

When a designer gets Jev