# The Cost of Deciding

Quick context, because this moved fast.

On September 15, TypeSafe came out of stealth with Jev: a model that never writes a word. It reads a situation, answers a typed question, and hands you a probability. Two weeks later OpenAI announced a Decisions API doing the same thing. Two days after that, Cloudflare and AWS shipped open-weight versions under Apache 2.0, one of them small enough to run on a laptop.

Four shops. Three weeks. Two closed, two open.

Kahneman called the fast, automatic half of thinking System 1. The snap judgment before you have words for it. This is System 1 for machines.

## What got cheap

Not intelligence. **Judgment.**

Here's the difference. Until this month, asking an AI "is this change safe?" meant making a big model write a paragraph first. A second or two. A few cents. And a confident tone with no real probability behind it.

So you asked rarely. At the end. Once.

Now a typed verdict with a confidence number costs $0.00004 at Jev's list price. On open weights running on your own GPU, it rounds to zero. Latency is tens to a few hundred milliseconds. Fast enough to sit inside the request, not after it.

When a decision costs that, you stop rationing it. You ask a thousand times inside one workflow instead of once at the end. That's the whole unlock.

## What I built

I didn't read about this. I built one.

I trained a small classifier on Modal, pulled the weights down, and wired it into the harness around my coding agents. It answers one question, thousands of times a day: is this change safe to let through? The first verdict came back before I'd finished reading the diff.

I'd assumed training the model was the hard part. **It was the easy fifth.**

The hard part started the next morning:

*   **What happens when it's unsure?** I chose fail closed. The agent stops and asks.
    
*   **Where does each verdict go?** An audit log: what was asked, what was answered, how sure it was. A decision you can't replay is a decision you can't defend.
    
*   **How do I know it's right?** I ended up building a CLI and a labelling UI just to go through its calls by hand and feed the misses back in.
    
*   **How does it know what matters in this codebase?** That's the harness, feeding it context before it decides.
    

Four problems. None of them in the model. All of them took longer than the model.

## The catch

Here's the part nobody puts on the slide. Cheap decisions don't remove the hard work. They move it.

**A wrong decision that is cheap and fast is just a wrong decision at scale.**

The first month's numbers back this up. Anthropic made a classifier gate the default in Claude Code in August, and their own write-up says it still misses 17% of risky actions in their test set. An independent test found Jev's confidence scores well off out of the box, and nearly perfect after someone refit them on real data. Nobody ships that refit. And the first public failure wasn't a model failure at all: a trading tool wired a decision model in as a gate and let trades through whenever the model was unsure or the account ran out of credits. Fail open. Never backtested.

The model did exactly what it was asked. The scaffolding was the bug.

## Why open matters

If deciding is about to happen at that volume inside everything, the judgment layer becomes the most important code you run. You don't want to rent your conscience from a closed vendor.

Neither closed model lets you fine-tune on your own data today. The open ones shipped weights, a training recipe, and a way to fine-tune on your own traffic where it already lives. Within days Ollama added an endpoint for decision models. That's a standard forming, not a product.

Not ideology. Leverage. When the thing making the calls is inspectable, you can see why it decided, fix it when it's wrong, and run it where your data lives. No regulator requires open weights. They require logs, oversight and the ability to explain. Inspectable is just the cheapest way to get all three.

I run engineering at a regulated fintech. I've spent years being asked, after the fact, to show which decision fired and why. That's the only reason I noticed where the hard part was.

## Where I land

Optimistic. Not naive.

The magic is real. Decisions got cheap this month, and cheap decisions rewire what software can do. But the upside doesn't fall out of the model. It falls out of the discipline underneath it.

Try one thing this week. Take one expensive human check you make today and turn it into a cheap automated one that runs every time instead of sometimes. Then watch what it frees you to attempt next.

The models keep getting the headlines. The scaffolding keeps getting the results. I know which one I'd rather be building.

So: can you still stand behind every decision your systems make, once they fire a thousand times a minute?

If you can, this era belongs to you.
