--- license: apache-2.0 language: [en, ko, zh, ja, multilingual] library_name: transformers pipeline_tag: text-generation tags: - darwin - darwin-v9 - darwin-jgos - vidraft - final-bench - qwen - qwen3.5 - qwen3_5_moe - moe - mixture-of-experts - sparse-moe - 397b - a17b - hybrid-attention - linear-attention - long-context - 262k-context - fp8 - w8a8 - compressed-tensors - quantized - reasoning - reasoning-model - thinking - chain-of-thought - cot - math - science - stem - code - agentic - tool-calling - function-calling - ztc - zero-token-confidence - confidence-estimation - uncertainty-quantification - hallucination-detection - calibration - self-verification - selective-prediction - pre-action-gating - agent-safety - llm-router - guardrails - gpqa - gpqa-diamond - mmlu-pro - benchmark - eval-results - greedy - korean - english - bilingual - multilingual-llm - vllm - sglang - openai-compatible - multi-gpu - h100 model-index: - name: Darwin-397B-ZTC results: - task: {type: text-generation, name: Graduate-Level Reasoning} dataset: {type: Idavidrein/gpqa, name: GPQA Diamond, config: gpqa_diamond, split: train} metrics: - {type: accuracy, value: 93.43, name: "Accuracy (greedy, single-sample)", verified: false} --- # Darwin-397B-ZTC ### 397B Mixture-of-Experts built on Qwen 3.5 · **FP8** · GPQA Diamond **93.43 %** · **ZTC on board** `reasoning` · `MoE` · `FP8` · `262K long context` · `Korean + English` · `hallucination detection` · `tool calling`

**Half the footprint, GPQA Diamond 93.43 %. And this model stops itself before it acts on an answer it is about to get wrong.** --- ## 🧬 The Darwin Family

**Darwin** is [VIDRAFT](https://vidraft.net)'s measurement-driven reasoning model family — roughly **20 official models**, **400+ community derivatives**, and a standing place among the top open models on GPQA. --- ## 🧬 Darwin — transplanting the experts that work A large MoE model is made of hundreds of **experts**. **Darwin V9** selects the experts that perform best across several high-performing models, transplants them onto a base backbone, and fuses them with trust-weighted evolutionary merging. **Nothing is trained from scratch — proven capability is grafted on.** That is why the same method holds across every model size. | Model | Scale | GPQA Diamond | |:---|:---|:---:| | Darwin-9B-NEG | 9B | 84.3 | | Darwin-27B-Opus | 27B dense | 86.9 | | Darwin-36B-Opus | 36B MoE | 88.4 | | Darwin-28B-Opus | 28B | 88.89 | | Darwin-28B-REASON | 28B + DELPHI | 89.39 | | Darwin-398B-JGOS | 397B MoE (bf16) | 90.9 | | **Darwin-397B-ZTC** | **397B MoE (FP8)** | **93.43** | ### Lineage | Role | | | |:---|:---|:---| | **Base** | `Qwen/Qwen3.5-397B-A17B` | 397B MoE backbone, ~17B active — Apache-2.0 | | **Darwin V9** | expert transplant + trust-weighted evolutionary merging | this is where the model becomes Darwin | | **Precision** | compressed-tensors W8A8 FP8 | 418.7 GB | | **ZTC** | zero-token confidence readout | ships in `ztc/` | - **Darwin V9** — evolutionary FFN/expert transplant and trust-weighted merging onto large MoE backbones - **FINAL Bench** — VIDRAFT's evaluation framework - **Four-layer Pre-AGI roadmap** — Darwin → AETHER → PROMETHEUS → HEPHAESTUS --- ## 🏛️ ZTC — it knows **before it answers** Until now there were two ways to find out whether a model is about to be wrong. Both of them only work **after the answer already exists**. | Existing approach | Limitation | |:---|:---| | **Ask the model in words** | Costs extra tokens, adds latency, and models are badly overconfident | | **Attach an external judge model** | **Two models to operate** · re-reads the entire answer · **degrades on long outputs** · 🔴 **arrives too late — the answer is already produced** | **ZTC is a third path. It reads the model's own internal state once, before generation begins.** | | External judge model | **ZTC** | |:---|:---|:---| | When | **After** the answer | **Before it starts** | | Extra model | Required (two to operate) | **None (one)** | | Extra generated tokens | Re-processes prompt + answer | **0** | | Added latency | A second inference pass | **0.52 ms** — 0.003 % of generation cost | | Long answers, long trajectories | **Degrades as length grows** | **Length-independent** | --- ### 📊 Measured — on this model **① It judges its own answers** (PubMedQA, 539 items, 146 incorrect) | | AUROC | |:---|:---:| | Self-reported confidence (asked in words) | 0.7646 | | **ZTC (internal-state readout)** | **0.8801** | | **Gain** | **+0.1155** | Permutation null control: **z = 13.31** — shuffle the labels and the signal disappears. **② It judges other models' answers** (Korean KMMLU, 400 items — law, math, biology, history) | Judge | AUROC | |:---|:---:| | **Darwin-397B-ZTC** | **0.8228** (z = 9.66) | | Qwen3.5-27B | 0.8171 | | Qwen3.5-9B | 0.7297 | | Qwen3.5-4B | 0.7284 | | Open-source 4B judge model | 0.6844 | *Single domain, random folds. The ladder under the harder leaderboard protocol reads 0.7364 / 0.7282 / 0.6506 / 0.6360 — see the section below.* **Same 400 items, same conditions: +0.138 over the open-source judge model.** --- ### 📊 Independent leaderboard — 2,018 items, leave-one-domain-out The **Typed Decision Leaderboard** scores answer verifiers from several vendors on one identical item set with identical labels: | System | AUC | |:---|:---:| | **Darwin-397B-ZTC** | **0.7364** | | JEV (TypeSafe AI) | 0.7350 | | ZTC-Judge-27B | 0.7282 | | GPT-5.2 asked directly | 0.7148 | | open-jev 4B | 0.6844 | | *Answer length and formatting only* | *0.6223* | | Patronus Lynx 8B | 0.5179 | | *The answering model's own stated confidence* | *0.5000* | **First place — and the gap to second is 0.0014, with a 95% interval of −0.019 to +0.032.** Under the board's own rule an interval containing zero yields no rank, so this model and JEV are **not statistically separable**. That is stated here for the same reason it is stated there. **Per domain, against the surface baseline in the same domain:** | Domain | Baseline | **Darwin-397B-ZTC** | ZTC-Judge-27B | |:---|:---:|:---:|:---:| | Professional exams (law · math · biology) | 0.7138 | **0.8660** | 0.8462 | | Biology & medicine | 0.5908 | **0.7433** | 0.7154 | | Disaster & safety procedures | 0.5949 | **0.7319** | 0.6961 | | Scientific reasoning | 0.7272 | 0.6287 | **0.7410** | | General multi-step reasoning | 0.5420 | 0.6072 | **0.6172** | | **Size-weighted mean** | **0.6223** | **0.7364** | 0.7282 | 🔴 **On scientific reasoning the 27B model beats this one by 0.11.** A model fourteen times smaller wins that column. It is printed rather than dropped, because the ladder only means something if the places it inverts are visible. **Self-readout.** Given only the question, this model answers on its own and the same forward pass tells whether it was right: **0.7572** (3 domains, 1,595 items). Verifiers that see only text from outside a model cannot do this at all. **Protocol.** Every figure comes from a domain the probe never saw; hyper-parameters are selected inside the training domains only; scores are computed per domain and then size-weighted. Pooling all items into a single AUC inflates the result, because score scales differ between domains. --- ### Known limitation — the answering-model mixture matters - **Sensitive to which model wrote the answer.** The probe is fitted on answers from four models. Adding 1,772 answers from a single additional model shifted the mixture and **lowered** the size-weighted score from 0.7278 to 0.7177 — professional exams rose to 0.8575 while every other domain fell. Treat "works on any model's output" as a design goal, not a measured guarantee: if your generator differs sharply from the training mixture, measure before relying on the number. ### 🔎 Two measurements, two protocols — do not mix them | | Section above (PubMedQA / KMMLU) | Leaderboard | |:---|:---|:---| | Items | 539 self-judged · 400 other-judged | 2,018, five domains | | Split | random folds | **held-out domain** | | Result | 0.8801 · 0.8228 | **0.7364** | Leave-one-domain-out is far harsher than random folds, which is why the numbers differ. **Quote 0.7364 when comparing against other systems**; the higher figures describe an easier protocol. --- ### 📦 The probe ships with this model | File | | |:---|:---| | `ztc/ztc_probe_darwin397b.npz` | **45 KB** — the confidence readout for this model | | `ztc/usage.py` | minimal, runnable example | ```python z = np.load("ztc/ztc_probe_darwin397b.npz") s = ((h - z["mu"]) / z["sd"]) @ z["w"] # h = last-token hidden state, 4096-dim p = 1 / (1 + np.exp(-(z["cal_A"] * (s - z["s_mean"]) / z["s_std"] + z["cal_B"]))) ``` One matrix product. No second model, no extra tokens, no network call. The probe is specific to this model's hidden space (4096-dim) and does not transfer to others. --- ### 🤖 Why this is decisive for agents — **after-the-fact report vs. pre-action stop** In an agent loop the expensive thing is not tokens. It is **actions**. Files get edited, APIs get called, payments go through, mail leaves the building. ``` External judge : [generate] → [tool runs] → [cost, time, side effects] → [judge] → "that was wrong" ZTC : [read state, 0.52 ms] → stop here if risky → the action never happens ``` **In front of an irreversible action, an after-the-fact verdict is an incident report.** ### Patterns | Pattern | Behaviour | |:---|:---| | **Tool-call gating** | Low confidence → do not call the tool, ask a human instead | | **Model routing** | Send only the low-confidence queries to a larger model or external API | | **Retry budgeting** | Spend multi-sample decoding only on the steps that wobble | | **Long-trajectory monitoring** | Agent trajectories run to tens of thousands of tokens — **length-independent, so it can stay on at every step** | | **Selective prediction** | Withhold a risky answer and return "I don't know" | ### Gate deployment, measured | Metric | Before | After | |:---|:---:|:---:| | Gate accuracy | 71.3 % | **93.3 %** | | Incorrect answers blocked | 40.7 % | **74.1 %** | | Expensive-path calls | 42 % | **17 %** | At effectively zero cost it can stay on for **every** request. **Use cases** — hallucination detection · uncertainty quantification · confidence calibration · selective prediction · routing risky queries upstream · **pre-action gating for agents** --- ## 🏆 GPQA Diamond 93.43 % | Model | GPQA Diamond | |:---|:---:| | **Darwin-397B-ZTC** | **93.43** | | GPT5.2 | 92.4 | | Gemini-3 Pro | 91.9 | | Qwen3.5-397B-A17B | 88.4 | | Claude 4.5 Opus | 87.0 | ``` GPQA Diamond, all 198 items · greedy · single sample · no test-time engine ``` *Comparison figures: Qwen3.5-397B-A17B official model card.* --- ## API — drop-in for an existing JEV integration The endpoint takes the same request shape and returns the same response shape, so switching an existing integration is a URL change. ```bash POST /v1/evaluate Authorization: Bearer {"model": "vidraft/ztc", "state": {"question": "...", "answer": "..."}, "questions": {"correct": {"type": "boolean", "instructions": "Is the ANSWER factually correct?"}}} ``` ```json {"model": "vidraft/ztc-judge-397b", "answers": {"correct": { "probability": 0.1043, "verdict": "review", "score": -0.72, "position": 0.268, "band": "low", "action": "hold_or_escalate", "measured": { "band_accuracy": 0.485, "base_accuracy": 0.748, "if_lowest_20pct_dropped": 0.814, "escalate_gain_at_20pct_budget": 0.0134, "do_not": "resample_same_model", "why_not": "measured: fixes 6.7% of wrong answers, breaks 13.1% of right ones"}}}, "usage": {"generated_tokens": 0}} ``` `type` accepts `boolean` and `noul`. Existing clients read `answers..probability` and ignore the rest; the additional fields are there for clients that want to act on the score rather than merely record it. **0.19 s per call, zero generated tokens.** ### What `probability` means The raw score is unbounded. The shipped calibration maps it to P(answer is correct), fitted **leave-one-domain-out** — the mapping never sees the domain it is applied to. | | Expected calibration error | |---|---| | **ZTC-Judge-27B (after calibration)** | **0.0245** | | JEV, as shipped | 0.0381 | | JEV, after the same calibration | 0.0261 | | Laya-Multilingual, as shipped | 0.4985 | | Laya-Typed-Decisions, as shipped | 0.2641 | Measured on the same 2,018 items. **ZTC and JEV are effectively tied on calibration; the difference of 0.0016 is not meaningful.** Figures published elsewhere for these systems were measured on other test sets and do not reproduce here. 🔴 Calibration is uneven across domains: 0.0225 on biology & medicine, but 0.2941 on scientific reasoning and 0.2381 on general reasoning. Treat the probability as reliable in the first case and as a ranking signal only in the other two. ## What to do when the score is low The score alone is not actionable, so the response carries the measured consequence of each choice. | Band | Share | **Actual accuracy of answers in this band** | Recommended | |---|---|---|---| | `low` | 20% | **48.5%** | hold, or escalate to a stronger model | | `mid` | 40% | 73.2% | escalate if budget allows | | `high` | 40% | **89.6%** | accept | **Three things that work, measured on 2,018 items:** | Action | Effect | |---|---| | **Drop the lowest-scoring 20%** | accuracy of what remains: **74.8% → 81.4%** | | **Escalate the lowest 20% to a stronger model** | **+1.34 pp** end-to-end | | **Send the lowest 20% to human review** | catches **47.2%** of all errors — 2.4× random | ### Generate several candidates and let the verifier pick When the score is low, the most effective next step is not to escalate — it is to **produce more candidate answers and select between them**. Measured on 346 questions with five candidates each (1,730 candidates), all scored through this endpoint: | Policy | Accuracy | Output tokens | vs. one attempt | |---|---|---|---| | One attempt | 49.13% | 72 | — | | Majority vote over 3 | 46.82% | 212 | **−2.31 pp** | | Majority vote over 5 | 46.53% | 353 | **−2.60 pp** | | **Pick best of 5 with this model** | **51.45%** | 353 | **+2.31 pp** | | Pick best of 5 with JEV | 53.47% | 353 | +4.34 pp | | Pick one of 5 at random | 46.82% | 353 | −2.31 pp | | *Oracle — any correct candidate counts* | *63.87%* | *353* | *+14.74 pp* | **The same five candidates swing by 5 points depending on how one is chosen.** Majority voting is worse than not resampling at all: when a model prefers a wrong answer, more samples make that wrong consensus more certain. A verifier that ranks the candidates is what turns extra samples into accuracy. **Spend the budget only where it is needed.** Generating extra candidates only for low-scoring first attempts captures most of the gain at a fraction of the cost: | Triggered on | Accuracy | Output tokens | vs. one attempt | |---|---|---|---| | 10% of items | 49.71% | 80 | +0.58 pp | | **30% of items** | **50.87%** | **132** | **+1.73 pp** | | 100% of items | 51.45% | 353 | +2.31 pp | **At a 30% trigger rate you get three quarters of the benefit for 1.8× the tokens**, where always generating costs 4.9× for 1.3× the benefit. *Scope: one generator (GPT-4o-mini), one item set, five candidates. The oracle row shows the headroom that remains — a correct candidate is present far more often than any policy recovers it.* **One thing that does not work:** 🔴 **Do not take a majority vote over resamples.** Measured: five resamples with majority voting score **46.53%** where a single attempt scores **49.13%**. More candidates make a wrong consensus more certain unless something picks between them — see the table above. *Escalation pays for itself through precision, not recall. Re-answering repairs about 38% of wrong answers and damages about 30% of right ones, so a gate is only worth its budget if it mostly calls answers that are actually wrong.* --- ## ⚙️ Specifications | Item | Value | |:---|:---| | Architecture | `Qwen3_5MoeForConditionalGeneration` | | Parameters | **397 B total / 17 B active** (512 experts, 10 routed + 1 shared per token) | | Layers · hidden | 60 · 4096 | | Attention | Hybrid (45 linear + 15 full attention layers) | | **Precision** | **FP8** (compressed-tensors W8A8) | | Size on disk | **418.7 GB** | | Context | **262,144 tokens** | | License | apache-2.0 | --- ## 🚀 Quickstart ### Serving with vLLM (4 × H100 80GB) ```bash vllm serve FINAL-Bench/Darwin-397B-ZTC \ --served-model-name darwin-397b \ --tensor-parallel-size 1 --pipeline-parallel-size 4 \ --gpu-memory-utilization 0.92 --max-model-len 262144 \ --cpu-offload-gb 20 --enforce-eager --trust-remote-code \ --reasoning-parser qwen3 --enable-auto-tool-choice \ --port 8000 ``` ### SGLang ```bash python -m sglang.launch_server --model-path FINAL-Bench/Darwin-397B-ZTC \ --port 8000 --tp-size 8 --context-length 262144 ``` ### Chat Completions (OpenAI-compatible) ```python from openai import OpenAI c = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") r = c.chat.completions.create( model="darwin-397b", messages=[{"role": "user", "content": "Why is the Riemann hypothesis hard?"}], temperature=0.0, max_tokens=8192, ) m = r.choices[0].message print(m.reasoning_content) # thinking trace print(m.content) # final answer ``` ### 🛠️ Tool calling ```python tools = [{ "type": "function", "function": { "name": "get_weather", "description": "Get the current weather for a city", "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}, }, }] r = c.chat.completions.create( model="darwin-397b", tools=tools, messages=[{"role": "user", "content": "What's the weather in Paris?"}], ) print(r.choices[0].message.tool_calls) ``` ### 🤖 Agents and coding CLIs The endpoint is OpenAI-compatible, so **existing tooling connects unchanged.** opencode — `~/.config/opencode/opencode.json` ```json { "$schema": "https://opencode.ai/config.json", "provider": { "darwin": { "npm": "@ai-sdk/openai-compatible", "name": "Darwin (local)", "options": { "baseURL": "http://localhost:8000/v1", "apiKey": "EMPTY" }, "models": { "darwin-397b": { "name": "Darwin-397B-ZTC" } } } } } ``` Any OpenAI-compatible client (Cline, Continue, Aider, …) ```bash export OPENAI_BASE_URL=http://localhost:8000/v1 export OPENAI_API_KEY=EMPTY export OPENAI_MODEL=darwin-397b ``` --- ## 🎯 Intended use - Graduate-level STEM reasoning (GPQA, science qualifying exams) - Mathematics and long multi-step chains of thought - Code generation and debugging - 🤖 **Agent workflows** — ZTC blocks irreversible tool calls **before** they run - **Bilingual Korean + English reasoning** (Chinese and Japanese supported) - **Work where a wrong answer is expensive** — ZTC filters risky answers before they ship ## 🔗 Links - 🌐 **[vidraft.net](https://vidraft.net)** — VIDRAFT - 🤗 **[FINAL-Bench](https://huggingface.co/FINAL-Bench)** — all models - 📱 **[POCKET](https://huggingface.co/collections/FINAL-Bench/pocket-models-6a618ee5d23eafb7e185a5c6)** — on-device line that runs on phones and GPU-less PCs ## 📚 Citation ```bibtex @misc{darwin397b_ztc_2026, title = {Darwin-397B-ZTC: FP8 Mixture-of-Experts with Zero-Token Confidence}, year = {2026}, url = {https://vidraft.net}, note = {Base: Qwen/Qwen3.5-397B-A17B} } ```