Skip to content
JevDirectory.org
CommunityPractices & PatternsVerified 2026-09-20

Jev + WebMCP Breaks a Benchmark

The WebMCP benchmark author reports that Jev paired with the small Mercury 2.5 model solved 100% of tasks at roughly 112x lower model cost than GPT-6 Astra with computer use.

Category
Practices & Patterns
Published by
Community
Author
idan levin
Added
2026-09-20

The post

Tagscommunityxbenchmarkbrowser

Highlights

  • Jev + Mercury 2.5 solved 100% of the WebMCP tasks in the run.
  • Reported roughly 112x lower model cost than GPT-6 Astra with computer use.
  • Shared by the benchmark's own authors rather than a third party.

Watch out

A single benchmark configuration; reproduce it before quoting the multiple.

More like this

Vercel CTO Malte Ubl reports that Jev beat an existing classifier eval previously run on Gemini 2.5 Flash Lite, saturating the eval on quality and running 6x faster.
Practices & Patterns#community#x#evaluation
Community
Yuchen Jin's take on Jev for LLM-as-a-judge work: 20-200x faster and 40-400x cheaper, with a screenshot of the model picking which candidate is AGI after a prompt sweep under $0.10.
Practices & Patterns#community#x#evaluation
Community
A widely read counterpoint to Jev-based context compaction: compaction is reconstruction, not filtering, and scoring tool calls to drop history misunderstands context management.
Practices & Patterns#community#x#compaction
Community
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Full Jev video tutorial

The case against Jev-scored compaction

This is a terrible compaction strategy that fundamentally doesn't understand how compaction and context management work. Seems like a lot of people are confused so let's break this down. 1. Compaction isn't a filter The role of compaction is to clean up history to keep the Show more

tamara
tamara
@tamarajtran

found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant

Reply