Welcome to session 50! This one was postponed a week, from 26 September to 3 October. Bala ran it himself, on a model that had been out for about three weeks and had already changed how he builds.
Bala is the founder of Applied AI and ZORP. He has been running Jev in production for two weeks, and the numbers in this session come from those workloads.
Session overview
Most AI writes text. You ask a question, get a paragraph back, and then write code to parse it and hope the format holds. Jev does not write anything. You give it state and a question, and it picks one of the answers you offered, with a probability attached. Your code gets a value it can branch on.
The name comes from Jevons Paradox: build bigger roads and traffic goes up, because more people use them. TypeSafe's bet is the same for AI judgments. Make a decision cost a fraction of a cent, and people will use far more of them.
Bala frames it with Kahneman's Thinking, Fast and Slow. Large reasoning models are System 2: they think, deliberate and then answer. Jev is System 1, the instant call you make while driving without stopping to reason about the brake. Most steps in a business workflow are that kind of call: which team gets this ticket, is this an invoice, should this email go to finance.
The live demo is Relay, a voice-controlled email inbox Bala built in a couple of hours the day before. Jev decides what each spoken command means, and an LLM only writes the reply drafts.
Key takeaways
- Jev chooses, it never generates: Ask "is poha a type of food?" and you get
trueat 97% in 136 milliseconds. The same question to Gemini took around 8 to 11 seconds and returned a paragraph. - Three kinds of answer: Choice picks from a set ("which team handles this ticket?"), Score places something on a scale ("how frustrated is this customer?"), and Noul is a yes/no with a probability ("is the customer asking for a refund?").
- Every call is state plus questions: State is the context, such as the customer's message and the channel it came in on. Questions are what you want decided. One state can carry ten questions, answered in parallel.
- It is fast because there is no token loop: An LLM writes one token, feeds it back, then writes the next. Jev does one pass through the network and outputs probabilities, and repeat questions on the same state hit the cache.
- Close to Opus at a tiny fraction of the cost: In a bake-off on 179 WhatsApp conversations, Jev scored 73% label accuracy against 76% for Opus and 59% for Haiku, measured on the 106 that were hand-labelled. Per million calls that is about $91 against $24,900 for Opus.
- Put Jev in front of your LLM, not instead of it: Use Jev for routing and branching, then hand off to an LLM for extraction, writing and code. On one production workload this cut spend from about $50 to half a dollar over two weeks.
- Gate on confidence: Relay sends anything under 75% back to the user as a question. The probability is your signal for when to ask a human.
- Prompt engineering before fine-tuning: Start with good evals and context. Fine-tuning only pays off at real scale, once you have collected enough clean data.
Topics covered
What Jev is, and why it is named after a paradox
Jevons Paradox, TypeSafe, and System 1 against System 2 thinking. Why the large models have moved towards reasoning, and why that makes them slower and costlier for decisions that do not need it.
Jev against an LLM, live
The same question goes to Gemini in AI Studio and to Jev in the TypeSafe playground. Thousands of tokens and seconds of waiting on one side, a single probability in 136 milliseconds on the other.
Choice, Score, Noul, and how a call is framed
The three output types walked through with a support-ticket example, followed by state and question: what goes in each, and why separating them lets Jev answer many questions about the same context at once.
Why Jev is fast
Karthik asks what stops an LLM from doing the same thing. Bala's answer covers the KV cache, the token-by-token loop and Jev's single forward pass. Laya, an open-source System 1 model, comes up as the one whose architecture you can actually read.
Jev in production, and the bake-off
Routing in front of LLMs, the $50 to half-a-dollar workload, and the WhatsApp classification bake-off against Opus and Haiku covering latency, cost and accuracy. Then a tour of what people are building with Jev, including Laya playing Snake in real time.
Live demo: a voice-controlled inbox
Relay, running on a seeded demo inbox. Press space, say "open the email from Mira", and Jev returns the action (open), the scope (one email) and the target (Mira Nair). Say "tell Mira, sorry I'm late, it is approved", and Jev picks reply while Gemini extracts the message and drafts it. Bala covers the architecture: a 75% confidence gate, a fixed list of actions, and about 40 cents a day for 1,000 voice commands.
Guardrails with decision models
Checking whether a prompt or tool call is malicious, out of policy or destructive, without adding LLM latency to every call.
Q&A highlights
- Can Jev help an LLM make better decisions, and how would you wire it into Claude?
- Are there other System 1 models? How does Jev compare with re-ranker models?
- How do you trust the numbers Jev gives you at scale?
- What would slow it down or break it?
- When Jev is not sure, how does your application know?
- How do you use Jev inside Claude Code? (Hooks and skills. Bala uses Jev for compaction.)
- When is fine-tuning worth it, compared with prompt and context engineering?
- Will Jev-style models handle tool calling on behalf of frontier models?
Here's the entire recording of the session.