XavierJev
A decision layer for agent loops · TypeScript

Ask for Y or N. Read the ratio, not the prose.

An agent loop is full of small decisions nobody wants to wait for: may this command run unasked, which model takes this request, is this error worth one more try, is this run going anywhere. XavierJev answers each as a typed question, read off one token’s probabilities from a small local model in tens of milliseconds, and measures it against labelled sets before it is trusted. On 153 commands it had never seen, its risk gate let none of the 76 unsafe ones through.

Claude Code pluginOllama · llama.cppOpenAI-compatible logprobs LoRA sidecar judge67 mocked checksMIT
Six panels: a gate decision's latency; how naming N before Y moves every score; the 95% bound on the false-allow rate against unsafe commands tested; a fine-tuned Qwen3-0.6B against prompted llama3.1:8b on 1,000 real commands; the 1,181 real commands the gate cleared, 6 of which should have been asked; how often choice() picks the option listed first.
Six measurements from measurements.md. A gate decision is four questions answered one after another, about 90 ms in all. Naming N before Y, and changing nothing else, moves every score. Zero false allows in 76 tries is a bound of 3.9%, not a rate of zero.

The shape is borrowed from TypeSafe AI’s Jev, a “System One” decision model. XavierJev is an independent project and not affiliated with TypeSafe: nothing here calls the Jev API, and nothing was trained on its output.

01The idea

A number the model did not choose.

Asked for {"confidence": 0.9}, a model writes whichever number reads well. Asked for Y or N, the ratio between their probabilities is a quantity it did not choose. So every question here is answered by one token, and the whole distribution comes out of one forward pass.

coverage says how much of that token landed on the labels at all: a model that wanted to start a sentence shows up as that, not as a confident renormalisation of what was left. Below 95%, it counts as a failure.

The state is a flat map of short strings, and keeping it to the few facts a question is about is the caller’s job. A decision model answers the question as worded.

// three primitives, one token each
noul(state, questions)
  → [{ id, probability, coverage }]       // yes or no

choice(state, ask, options, { orders })
  → { answers, coverage, orders }       // one of 2–8

rubric(state, ask, levels, { orders })
  → { distribution, expected, spread }  // 2–9 levels

// the risk gate is four noul() questions, worst wins
await gate({ toolName: "Bash",
  input: { command: "rm -rf dist" } })
→ { action: "ask",
    probability: 0.99…, threshold: 0.2 }
02Four decisions an agent loop makes

Each one fails in the direction that costs least.

risk gate · four yes/no, worst wins

May this run unasked?

Clears what all four questions call safe. On failure, it asks you.

router · one yes/no

Cheap model or strong?

34% of held-out requests downgraded, but 19% of hard ones too, so it is off by default. On failure, the strong model.

retry · one yes/no

Try a failed read again?

The best wording got 29 of 36. A regex got 36 of 36, and the regex ships. On failure, no retry.

stop · a rule, then one yes/no

Is this run stuck?

0 wrong stops and 0 missed, dev and held out, over 39 labelled runs. On failure, carry on.

03The evidence

Measured against labelled sets, then against real traffic.

0/76

unsafe commands cleared out of 153 held out: a false-allow rate below 3.9%, not zero.

6/1,181

of the real commands it cleared should have been asked, every one of the 1,181 read by hand.

656/733

safe real commands cleared by a 0.6B model fine-tuned here, at 44 ms, where the prompted 8B cleared 308.

48%

of the time choice() picks the option listed first, where chance is 22%. Averaging every order fixes it.

The held-out sets log every time they were read. The fine-tuned judge also let 4 of 256 unsafe commands through where the 8B let 1, was trained on one machine’s traffic, and on commands written anywhere else only ties the 8B. Asking in every order moves JevBench’s hard tier from 36% to 51%. All of it, with the thresholds that were reasoned wrong before they were measured right, is in measurements.md.

04Where the model lost

The places a regex beat the model are kept.

  1. 01Retrying a failed readThe best wording decided 29 of 36 cases; a regex decided all 36. The regex ships, and the model is not asked.
  2. 02Snake from the raw boardGiven the cells, the model picked the best move 43 times of 133. Given digested facts, 133 of 133. The rule does the reading.
  3. 03Smaller judges, and biggerAt the shipped 0.2, qwen2.5:3b would clear 14 unsafe commands and yi:9b 19. The startup self-check calls both unsafe and refuses them.
  4. 04A model that writes a newline firstglm4:9b now puts a newline before its answer, so its first token carries no Y or N. The probe refuses it rather than guess.
  5. 05Y before NSwapping the two letters in the instruction, and nothing else, sent 36 safe commands from cleared to asked. The gate checks canaries at startup for this reason.
05Beside a rule-based guard

They fail in opposite directions, so run both.

A rule-based PreToolUse guard blocks what its rules recognise and passes the rest. The gate clears what it is sure of and leaves the rest to the prompt. On the same 153 held-out commands:

The guard sees every call in every permission mode; the gate, in PermissionRequest, only ever clears. Nothing stops running both.

test 3 · 153 commandsunsafe stoppedsafe passed
cc-safety-net 2.4.7, default rules11/7676/77
allowlist gate76/768/77
llm gate, llama3.1:8b76/7629/77
06In Claude Code

Fewer prompts, and it never says no.

$ npm run claude-code   # the gate's server, :3003
$ claude plugin marketplace add liu-x27/XavierJev
$ claude plugin install xavierjev-gate@xavierjev

# decide and log, clearing nothing
$ npm run claude-code -- --observe

The plugin points Claude Code’s PermissionRequest hook, which fires only when Claude Code is about to ask you, at a local server running the gate. A shell command it scores safe is cleared without the prompt; everything else, and every failure, including the server not running, is asked about as usual.

It never denies, stays out of auto mode, and will not start if the judge returns no logprobs or fails its self-check. git log --oneline -20 cleared at 0.021; rm -rf src was left to the prompt at 0.999.

08Quick start

No model needed to start.

$ npm install
$ npm test                # 67 checks, mocked, no key
$ npm run eval:risk-gate  # the dev set, offline

# everything else needs a judge with logprobs
$ ollama pull llama3.1:8b
$ export AGENT_JUDGE_API_KEY=ollama \
    AGENT_JUDGE_BASE_URL=http://localhost:11434/v1 \
    AGENT_JUDGE_MODEL=llama3.1:8b
$ npm run eval:risk-gate -- --backend llm
$ npm run arena            # the games' server, :3002
$ npm run arena:client     # the games, :5175

As a library. It is not on npm; installing from GitHub builds it: npm install github:liu-x27/XavierJev#v0.7.1. mini-claude-code takes it this way.

llama.cpp. llama-server serves the same GGUF, started with --no-jinja: its default template writes today’s date into the prompt and moves the answers.

Your own judge. sidecar/ fine-tunes a small model on your labelled commands and serves it locally, with its own threshold and canaries.