Skip to content
JevDirectory.org
OfficialPractices & PatternsDocsVerified 2026-09-22

Workflow evals

Published evaluations of four automation workflows, security incidents, agent trace observability, invoice processing, and customer service, comparing Jev and frontier LLMs as structured workflows versus single prompts.

Category
Practices & Patterns
Published by
TypeSafe AI
Author
Added
2026-09-22
Tagsofficialbenchmarksevaluationmethodologycalibration

Highlights

  • Covers four workflows: security incidents, agent trace observability, invoice processing, and customer service.
  • Scores models as workflows of Noul, Choice, and Score questions against consensus labels from GPT-6 Astra and Claude Fable 5.1 at high thinking.
  • Reports Jev averaging 67.8% accuracy at $0.0004 and 0.4s per case across the four workflows.
  • Shows every model is more accurate, cheaper, and faster in a workflow than with the same policy as a prompt, on average.
  • Charts accuracy against cost and time on log scales, with up and to the left marked better.

Watch out

Vendor-run evals whose reference labels come from two frontier models; harness and label correctness is assumed rather than independently audited.

More like this

74GitHub stars
Independent cross-model benchmark for Jev-class decision models, running 534 frozen cases per complete entrant with scoring code and a four-axis score of accuracy, calibration, latency and cost.
Practices & Patterns#community#python#benchmarks
Communityjevbench
An experiment comparing Jev with GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6 as evaluators of five frozen weather-agent runs, measuring pass-or-fail accuracy against human labels plus variance, cost, and latency.
Practices & Patterns#community#python#evaluation
38GitHub stars
Side-by-side benchmark of TypeSafe Jev, Qwen 3.8 27B on Cerebras, and a local Needle 3 across seven synthetic workloads, recording validated outputs, mistakes, latency, tokens, and estimated cost. Raw exports and per-scene limitations are published.
Practices & Patterns#community#typescript#benchmarks
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

A Claude session from 1M to 86K tokens

This is actually insane. This uses @typesafeai Jev model, as a plugin in Claude to review all the un-nesseasary tool calls, and it takes 1s to run! Like, literally, 1 second to take my Claude session from nearly 1M to ... 86K tokens! 😮 Ask your claude to install it and be  Show more

Image
Image
tamara
tamara
@tamarajtran

found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant

Reply

Vercel's fx safety reviewer, 18x faster

We're seeing extraordinary results from @typesafeai. Default mode in 𝚏𝚡 is auto, with a safety reviewer analyzing every command. That reviewer runs on GPT Luna today. Jev is up to 18x faster (p95) *and* more accurate. It's coming to @vercel AI Gateway and likely new default.

Pranit
Pranit
Vercel
@fazxes

We benchmarked fx auto mode (safety) classifier with @typesafeai's Jev. tl;dr: ~5-18x faster and more accurate than 𝚐𝚙𝚝-𝟻.𝟼-𝚕𝚞𝚗𝚊, our current top choice

Image
Reply

Jev lands on OpenRouter

700 leads scored for $0.09

Beating Gemini Flash Lite on an eval

Browser Use Ultrafast, powered by Jev