Skip to content
JevDirectory.org
OfficialSites & GuidesArticleVerified 2026-09-22

The Bitterest Lesson

A TypeSafe blog essay arguing that in ML the order that matters is doing the right task, then data, then compute, then algorithms, using the InstructGPT result as its example.

Category
Sites & Guides
Published by
TypeSafe AI
Author
Added
2026-09-22
Tagsofficialarticlemodelstraining

Highlights

  • Extends Sutton's compute-beats-algorithms lesson: doing the right task > data > compute > algorithms.
  • Cites InstructGPT, where GPT-2-sized models over 100x smaller than GPT-3 beat it once trained on the right task.
  • Says scaling pre-training would need roughly GPT-7 level to beat that baseline and GPT-9 to beat InstructGPT on GPT-3.
  • Published September 10, 2026, with footnotes linking to the original version on The Complete Skeptic.
  • Argues picking the right task often requires leaving ML to study users, products, and organizations.

Watch out

A persuasive essay rather than new experimental work; the InstructGPT comparison is read from the original paper's figure 31.

More like this

TypeSafe AI's product site for Jev, its first System One Model: typed decisions with calibrated confidence, performance and pricing claims, a FAQ, and links to the docs, console, workflow evals, and launch post.
Sites & GuidesDocs#official#docs#models
Official
TypeSafe's case for machine-native composable AI: intelligence as a dependable primitive that software branches on, rather than an assistant that keeps humans in the loop.
Sites & GuidesArticle#official#article#architecture
Official
Wikipedia's article on Jev: TypeSafe AI, the September 15, 2026 early-access release of jev-1.13.0, the Choice, Score, and Noul primitives, RLCD training, and the unpublished weights and technical paper.
Sites & GuidesArticle#community#article#models
Community
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Instant compaction with Jev

A Claude session from 1M to 86K tokens

This is actually insane. This uses @typesafeai Jev model, as a plugin in Claude to review all the un-nesseasary tool calls, and it takes 1s to run! Like, literally, 1 second to take my Claude session from nearly 1M to ... 86K tokens! 😮 Ask your claude to install it and be  Show more

Image
Image
tamara
tamara
@tamarajtran

found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant

Reply

Vercel's fx safety reviewer, 18x faster

We're seeing extraordinary results from @typesafeai. Default mode in 𝚏𝚡 is auto, with a safety reviewer analyzing every command. That reviewer runs on GPT Luna today. Jev is up to 18x faster (p95) *and* more accurate. It's coming to @vercel AI Gateway and likely new default.

Pranit
Pranit
Vercel
@fazxes

We benchmarked fx auto mode (safety) classifier with @typesafeai's Jev. tl;dr: ~5-18x faster and more accurate than 𝚐𝚙𝚝-𝟻.𝟼-𝚕𝚞𝚗𝚊, our current top choice

Image
Reply

Jev lands on OpenRouter

700 leads scored for $0.09

Beating Gemini Flash Lite on an eval