Skip to content
JevDirectory.org
CommunityPractices & Patterns3.2k starsVerified 2026-09-22

Kev

Apache-licensed, locally runnable Jev-style decision models (0.8B, 4B, 9B on Qwen3.5) with released weights, training code, a System One-compatible server, frozen eval suites, and a playground.

Category
Practices & Patterns
Published by
Community
Author
jaredpalmer
Added
2026-09-22
Tagscommunitypythonopen-modelsbenchmarkscalibration

Highlights

  • Weights for Kev-0.8B, Kev-4B, and Kev-9B are released on Hugging Face with training code and evaluation data.
  • Its API matches TypeSafe's System One, so the official Python SDK can point at a local server; requests mix noul, choice, and score questions.
  • Runs on CUDA, ROCm, and Apple Silicon; the 4B and 9B fit a 32 GB Mac in bf16, and a web playground permutes option order.
  • On its frozen new-source test set the author reports Kev-9B at 0.852 versus 0.857 for hosted Jev, but Jev's training data is unknown.
  • Fine-tuning gaps are documented: the first Kev-9B dropped deadline questions to 0.72 and KEV_DATE_FACTS=1 recovers 0.90.

Quickstart

bash
git clone https://github.com/jaredpalmer/kev.git && cd kev
uv sync --extra serve
KEV_DTYPE=bf16 uv run --extra serve python -m kev.serve --run jaredpalmer/kev-4b --port 8009

Watch out

Apache-2.0, with the Qwen bases under the same license. Not a controlled architecture comparison because Jev's training data is unknown; needs Python 3.12+ and uv, and calibration and option order should be tested on your own data.

More like this

3.6kGitHub stars
An independent baseline that reads typed option probabilities straight from a frozen Qwen3.5-4B's logits in one forward pass, reproducing Jev's interface pattern with open models rather than Jev's undisclosed model or training.
Practices & Patterns#community#python#open-models
CommunitySemIf
296GitHub stars
An independent open reproduction of the System One model class: Qwen3.5-based 2B and 35B mixture-of-experts models that return typed Choice, Score, and Noul probabilities in one forward pass, with nothing distilled from Jev.
Practices & Patterns#community#python#open-models
Communitydecider
260GitHub stars
A ModernBERT-based open decision model that evaluates Choice, Score, and Noul schemas in one non-autoregressive pass; Verdict 2.0 reports 77.10% accuracy on 2,000 held-out enterprise decisions.
Practices & Patterns#community#python#open-models
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

The launch post

Trading bot, one decision per block

Classifying 1,500 real emails

this model is actually insane at email classification i tested it on 1500 of my own emails to see how well it works and I am blown away

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

Fast browser use with Stagehand

we built blazing fast computer/browser use with Jev + @Stagehanddev. this task cost $0.001 and executed at near instant speed (in a remote browser btw) the loop: observe the page, send a11y tree as state + actions as questions, Jev decides the next action, then Stagehand Show more

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

LLM-as-a-judge, sped up

Jev has spoken. It picked which model is AGI. 20–200x faster. 40–400x cheaper. This could make things like LLM-as-a-judge insanely fast and nearly free. (I tried a bunch of prompts and still didn’t burn through $0.10.)

Image
Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

Instant compaction with Jev