Ideas from the work, sent by email. Sign up for the newsletter →
runpoint.
Jev reads text and makes a decision. Opus 5.5 and GPT-6 Sol sent our team back to Claude.
 
runpoint.FIELD NOTES

September 25, 2026  ·  Sam Gaddis

Jev is a new kind of model

One release this month introduced a new type of model. Two others sent our whole team back to Claude.

Overhead view of a chaotic pile of blank pages beside three neat stacks, one marked with an orange tab.

Many inputs, a few fixed answers.

The most useful new model we started using this month can’t write a sentence.

Jev, from a startup called TypeSafe AI, came out on September 15. Every language model we’ve used before reads text and writes text back. Jev reads text and makes a decision. TypeSafe calls it a “System One” model, after Daniel Kahneman’s term for fast, intuitive thinking. You give it some text and a fixed set of possible answers, and it can:

  • Answer yes or no. “Is this email a sales lead?”
  • Pick from a list. “Which of these six queues should this ticket go to?”
  • Give a score. “How urgent is this request, on a scale you describe?”

Each answer comes with a confidence score. It’s very fast and very cheap, and the answer is always one of the options you gave it, so there’s nothing to parse or clean up afterward.

That sounds narrow, and it is. But a lot of business work turns out to be classification once you look at it closely. Until this month we sent most of it to small language models like OpenAI’s Luna. For that kind of work, Jev is a big deal, and it’s now our first choice.

We’ve already built it into a legal application for one of our clients. Legal review is full of questions with a fixed set of answers, which is exactly the kind of question Jev is built for.

It has limits. TypeSafe says Jev can’t hallucinate, which is true only in the sense that it answers from your list. It can still pick the wrong option. Its documentation says it struggles with numbers, dates and adversarial inputs, and Simon Willison pointed out that a score with no explanation makes bias hard to catch. Test it on examples where you already know the right answer before you trust it with anything that matters.

If you have work like this and want to talk through whether Jev fits, schedule a call with us.

Why we’re back on Claude

OpenAI released GPT-6 Astra, its new flagship, on September 3. We were excited about it, and it’s a strong model. It’s also expensive. Within days most of us were hitting usage limits, and our costs climbed fast.

Anthropic released Claude Opus 5.5 on September 22. Anthropic says it:

  • performs at the level of its top model, Fable 5.1, on most work
  • costs much less to run than Opus 5
  • writes output more than 30% faster
  • comes with higher five-hour usage limits on paid plans
A two-color print of a fuel gauge with its needle just above empty.

Running low.

After the last three weeks, the higher limits matter to us as much as any benchmark. On Anthropic’s own coding tests, Opus 5.5 also beats Astra and costs less per task. Those are the vendor’s numbers.

Our own experience is only a few days old, but the whole team has already moved its default work back to Claude. It’s the first time we’ve gone back since we moved to OpenAI. Our early read is that Anthropic has the best model again.

GPT-6 Sol is a price cut

Two identical white mugs. The left has a whole cream tag; the right has an orange tag torn in half.

Same mug, smaller tag.

OpenAI released GPT-6 Sol the same day. It’s much cheaper than GPT-5.6 Sol, and OpenAI says it makes about half as many factual mistakes.

But Artificial Analysis found it scores about the same as GPT-5.6 Sol on its intelligence and coding indexes, a little better on some tests and worse on others. It’s a cheaper version of roughly the same model. We had hoped for more.

What we’d do this week:

  • If a task ends in a label, a yes or no, or a score, try Jev before a language model.
  • If you moved to Astra and the limits hurt, run the same work through Opus 5.5 and compare.

Either way, test with your own hardest examples. Anthropic says Sonnet 5.5 and Haiku 5.5 are coming within weeks, so some of this may change again before the next issue.

Have you pointed Jev at anything yet? Reply and tell me what you tried.

Sam Gaddis • Partner, Runpoint

512.695.6294  /  runpoint.ai  /  Videos

The Runpoint Letter

Get the next issue

One idea, one number, one thing you can do Monday. Three minutes.

Sign up for Field Notes →