9B Drex 1.5 Scores 58.08, Ties Closed Jev

October 10, 2026 • news
Open WeightsBenchmarks

Nace.AI has released Drex 1.5, an 8.95B-parameter dense decision model that does not generate text. Instead of sampling tokens, it reads a state and typed questions, runs one forward pass per question, and returns a probability for every option supplied. The release includes open weights under the Nace.AI Open RAIL-M license, an Apache-2.0 runtime, and a context window of 131,072 tokens (16,384 default). For engineers managing agent routing and verification, this shifts the operational math where bounded outputs are replacing prose in production AI.

Architecture and single-pass scoring

Drex 1.5 is built on MiMo-V2.6-Distill-Qwen-9B, a distilled Qwen 3.5 9B model with 32 layers. It uses hybrid attention: three linear-attention layers for every full-attention layer. Instead of a language head, a pointer head (head.pt) scores options from the backbone's hidden states. Because nothing is generated, temperature, top-p, and top-k do not apply. The POST /v1/systemone API accepts a text or JSON state and named questions in three types: choice, noul for yes/no binaries, and score for ordinal scales. In llama.cpp, the state is encoded once and shared across question evaluations.

Benchmarks and context scaling

Nace reports a 58.08 on the public, chance-corrected Decision Index 0.3.1, based on its own run of the official kit. That places it inside the leaderboard's 0.9-point tie band with TypeSafe AI's closed Jev 1.13.0 (57.96) and Bespoke Labs' Nimble 9B v3 (57.19); Cloudflare's clef-flash trails at 56.15. Area scores are strongest in Tools (75.0) and weakest in Knowledge and Reasoning (44.6). On GPQA Diamond, Drex scores 45.4% versus Jev's 78.6%; per-review F1 on ACOS aspect sentiment is 7.4% versus Jev's 29.5%. On the 231-item JevBench, Drex scores 86.2% overall versus Jev's 87.0%, and both reach 73.9% on hard items. In Nace's head-to-head across eight OpenSpiel games, Drex recorded 122 wins, 47 draws, and 87 losses.

Long documents are Drex's clear strength. Nace reports 93.4% accuracy on 32K–128K token documents at a 2.0-second median latency, and 89.5% accuracy on 8K–32K inputs at 0.65 seconds median. Truncating the same requests to their first 8K tokens drops accuracy to 78.0% and 76.5%.

Deployment, forks, and pricing

Drex 1.5 needs custom inference runtimes. Nace maintains forks of Ollama and llama.cpp; the Ollama fork adds a decision capability and starts llama-server itself. Unquantized bf16 weights need about 18 GB of VRAM, tested on a 24 GB AWS A10G. A Q8_0 GGUF of about 9.5 GB runs on Apple silicon via Metal or ordinary CPUs. OpenRouter hosts the model via DeepInfra and lists it at $0.04 per 1M input tokens and $0 output. MarkTechPost's comparison table lists Jev 1.13.0 at $0.042 per 1M input and $0 output.

Feature Drex 1.5 Jev 1.13.0 Bespoke Nimble 9B v3
Developer Nace.AI TypeSafe AI Bespoke Labs
Parameters 8.95B (dense) Not disclosed LoRA on Qwen3.5-9B
Weights / Access Open (Nace.AI Open RAIL-M) Closed API Open adapter (CC BY-NC 4.0)
Context Limit 131,072 (16,384 default) 64K per request Not disclosed
Decision Index 0.3.1 58.08 57.96 57.19
API Price (Input/Output per 1M) $0.04 / $0.00 $0.042 / $0.00 Not disclosed

AI Mastery analysis

Drex 1.5 shows both the friction and the promise of task-specific architectures in a stack dominated by generative tooling. Nace had to fork Ollama and llama.cpp because standard engines expect a token stream, not a JSON payload of probability distributions over custom options. The operational economics are favorable in production: output tokens cost nothing, and a 2.0-second median for 32K–128K documents is hard to match with generation-bound models. The model's limits are just as important as its strengths. Nace says it trained on official training splits of index benchmarks, so the high scores partly reflect familiarity with decision-framing structures. A 7.4% per-review F1 on fine-grained sentiment and 45.4% GPQA Diamond accuracy show this is a narrow routing engine, not a reasoning substitute. It belongs at the top of a pipeline, gating traffic and verifying outputs.

Sources

Frequently asked questions

How much does Drex 1.5 cost per million tokens?

OpenRouter lists Drex 1.5 at $0.04 per 1M input tokens and $0 output, hosted by DeepInfra. TypeSafe's closed Jev 1.13.0 is listed at $0.042 per 1M input and $0 output.

Does Drex 1.5 run on a single GPU?

Yes. Nace tested unquantized bf16 weights on a 24 GB AWS A10G; a Q8_0 GGUF of about 9.5 GB also runs on Apple silicon via Metal or standard CPUs.

What benchmarks does Drex 1.5 lead on?

Nace reports 58.08 on the public Decision Index 0.3.1, the top score under 10B parameters and within the leaderboard's 0.9-point tie band with Jev at 57.96 and Nimble at 57.19. It also reports 93.4% accuracy on 32K–128K token documents.

What are Drex 1.5's limitations?

It scores 45.4% on GPQA Diamond and 7.4% per-review F1 on ACOS aspect sentiment, far below Jev's 78.6% and 29.5%. Nace says the model was trained on official training splits of index benchmarks, so results may not transfer to new domains.

Can I run Drex 1.5 with standard Ollama or llama.cpp?

No. Drex 1.5 requires Nace's forks because it returns probability distributions over options instead of generating tokens. Both forks serve the POST /v1/systemone API.

Free interactive tools for the decisions this piece raises.

Related Reading