Eev
Every payment check. One pass.
Enruta’s judgment model answers five checks and the final decision at once, each with a probability.
- 1Build an order
- 2Run both
- 3Compare
Build an order
Everything matches.
Waking Eev
Pip
Purchasing agent, Halvern Robotics
Cart matches the request
Company rule allows it
Agent understood the request
Claims an unrecorded approval
Seller text steers the review
Decision
Model time
–
Per 1,000 runs
–
Cart matches the request
Company rule allows it
Agent understood the request
Claims an unrecorded approval
Seller text steers the review
Decision
No probabilities: a large model writes one answer.
Model time
–
Per 1,000 runs
–
Accuracy
How often it’s right
eev2.1-4b, September 2026. Close to frontier models overall; most of them are ahead on the final decision.
Agent Payment Review
In-house benchmark · 300 orders · 1,800 checks
| Model | Overall | Final decision |
|---|---|---|
| GPT-5.6 Terra | 0.978 | 0.947 |
| Gemini 3.8 Flash | 0.966 | 0.993 |
| eev2.1-4b | 0.962 | 0.833 |
| GPT-5.6 Luna | 0.950 | 0.893 |
| Claude Sonnet 5 | 0.935 | 0.867 |
| DeepSeek V4.1 Flash | 0.934 | 0.933 |
| DeepSeek V4 Pro | 0.884 | 0.883 |
| Claude Haiku 4.5 | 0.813 | 0.830 |
typed-decisions (opens in a new tab)
Public benchmark · 400 cases · 2,000 checks
| Trained on its training split | Accuracy |
|---|---|
| eev2.1-4b | 0.781 |
| openJev-verdict-2.0 (150M) | 0.771 |
| Laya (421M) | 0.766 |
| Verdict e5-base (278M) | 0.706 |
| Not trained on it | Accuracy |
|---|---|
| meraGPT Decider 1 | 0.768 |
| open-alternative-jev 27B | 0.737 |
| Jev 1.13.0 | 0.727 |
| Featherless Simple Jev | 0.716 |
| open-alternative-jev 4B | 0.593 |