Skip to content
JevDirectory.org
CommunityRepos & SDKs376 starsVerified 2026-09-22

tax-doc-classifier

Text-only classifier for 261 federal tax forms that sends each PDF page's text to Jev, which returns a probability over forms and page kinds behind a confidence gate; no model is trained.

Category
Repos & SDKs
Published by
Community
Author
kyotofin
Added
2026-09-22
Tagscommunitytypescriptclassificationextractiondata

Highlights

  • One request per page returns a distribution over 261 IRS forms and 7 page kinds; blank pages are answered without a call.
  • The classifier is data: criteria.json describes each form and is generated from the IRS's own PDFs; no model is trained or hosted.
  • The author reports 0 wrong labels on both corpora and 38 low-confidence pages (5.05%) on the blank-form corpus; page-level results are not committed.
  • Against the previous Sonnet pipeline: $0.00115 vs $0.039 per page (34x cheaper) and about 0.5 s vs 3.3 s warm latency.
  • Five corporate or foreign forms absorb their schedules through a second small question; ids follow kebab-cased MeF naming.

Quickstart

ts
import { classifyPage, jevBackend, pdfPageLines } from 'tax-doc-classifier'
import criteria from 'tax-doc-classifier/data/criteria.json' with { type: 'json' }

const lines = await pdfPageLines('return.pdf', 3)
const r = await classifyPage(lines, { backend: jevBackend(), criteria })

Watch out

Apache-2.0 with IRS-derived data under DATA-LICENSE.md. Requires Node 20+, poppler's pdftotext and pdfinfo for the PDF helpers, and a TYPESAFE_API_KEY; text-only, English-only and federal forms only.

More like this

45GitHub stars
Experimental Hono router that matches an incoming HTTP request to a plain-language route description: one yes/no Jev question per description is evaluated in a single call, and the first registered route above the threshold wins.
Repos & SDKs#community#typescript#routing
13GitHub stars
A Python library that adds .jev accessors to pandas and Polars: ask a natural-language question per row and get labels, scores, and full probability distributions back.
Repos & SDKs#community#python#data
Communityjevframe
10GitHub stars
An OpenTelemetry exporter wrapper that scores each log record with Jev for diagnostic value, priority, and routing, letting low-value events skip an LLM-analysis branch while every record stays in your archive.
Repos & SDKs#community#typescript#routing
Communityjevlogs
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Inferring Jev's internals from 1,000 calls

Jevの内部アーキテクチャを推測している技術記事(Jev’s Architecture Unmasked)からメモ。 ・本記事はJevのAPIを約1万回の呼び出して、内部構造を推測したもの ・従来の言語モデルを用いた分類やルーティングでは、トークンを1文字ずつ逐次生成するために膨大な無駄な計算コストが発生していた。 Show more

Reply

The open System One roundup

Jev 发布没几天,开源社区已经开始疯狂复刻了🔥 最值得推荐的五个模型: 1、Laya 421M:原生决策模型,支持 Mac 2、Decider-2B:最像 Jev,基于 Qwen3.5 3、NanoJev 0.6B:专门的 Decision Head 4、Reflex:Qwen3.5 + Direct Logits 5、System-One 4B:专门做概率校准 Show more

小墨同学
小墨同学
@xiaomovps

Jev 刚发布没几天,开源社区就出现了同款🔥 Decider-2B模型,是基于 Qwen3.5-2B 做了特殊调整 它和 Jev 模型是一样的 只做选择 评分和判断 不是文本类的 LLM 模型 但两者还是有几个明显区别: 1、模型 Jev:闭源 System One Model Decider:Qwen3.5-2B,约 1.9B 参数,Apache 2.0 开源 2、价格

Image
Reply

The launch post

Trading bot, one decision per block

Classifying 1,500 real emails

this model is actually insane at email classification i tested it on 1500 of my own emails to see how well it works and I am blown away

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

Fast browser use with Stagehand

we built blazing fast computer/browser use with Jev + @Stagehanddev. this task cost $0.001 and executed at near instant speed (in a remote browser btw) the loop: observe the page, send a11y tree as state + actions as questions, Jev decides the next action, then Stagehand Show more

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply