Together AI releases Tev1-4B-experimental and $17 fine-tuning guide
TypeSafe AI's Jev decision model
Together AI Releases $17 Jev-Like Classifier, Tev1-4B
The decision-model category is drawing third-party implementation work within days of Jev's launch, with Together AI publishing a reproducible fine-tuning recipe at low cost. Bias concerns about opaque probability outputs remain unresolved.
The full picture
Together AI released together/Tev1-4B-experimental, a Jev-like classifier built on Qwen3.5 4B and hosted on Together's serverless platform, with a blog post showing how to fine-tune a similar model for approximately $17. This follows TypeSafe AI's September 15 release of Jev, a decision model that accepts text input and returns typed probability scores for classification-style questions rather than generating prose. TypeSafe founder Diogo Almeida, who helped develop the research behind ChatGPT, built Jev using a training method called Reinforcement Learning for Calibrated Decisions. TypeSafe claims Jev answers in 70 to 500 milliseconds and is 20 to 200 times faster and 40 to 400 times cheaper than comparable LLM workflows, priced at $0.042 per million input tokens with free output. On September 19, TypeSafe released JevBench, a composite benchmark combining Intelligence, Calibration, Speed, and Cost via geometric mean for bounded software decision tasks. Jev 1.13.0 leads the JevBench composite over GPT-5.6 Luna, which records substantially higher hard-case accuracy; Jev's advantages in latency, calibration, and cost carry it to the top position. JevBench results are specific to narrow typed-decision workloads and do not indicate general capability comparisons. Simon Willison covered Jev on September 21, flagging bias risk from its opacity, citing a city-ranking experiment where Jev rated Cupertino top and East Palo Alto bottom with no way to inspect which signals drove the scores. Community experimentation produced a character-by-character chat model, a left-pad implementation, and a 2048 game driver. Open-weight recreations including Kev, based on Qwen, appeared within a week of launch.
How it developed
TypeSafe AI launched Jev on September 15, a model that takes unstructured text and returns numeric confidence scores for decisions instead of prose.
On September 19, TypeSafe released JevBench, a composite benchmark of intelligence, calibration, speed, and cost where Jev 1.13.0 leads GPT-5.6 Luna despite its lower hard-case accuracy. Simon Willison flagged bias risk on September 21, citing a ranking experiment where Jev scored Cupertino top and East Palo Alto bottom with no inspectable rationale.
Simon Willison covers Jev, flagging bias risk; community recreations appear
TypeSafe releases JevBench composite benchmark
TypeSafe AI releases Jev decision model
Sources
Want this in your inbox?
I send one email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free