LocalFlash: Qwen3.8-27B + DFlash2 serving configuration for Apple Silicon

This repository documents a measured, reproducible deployment recipe (not new weights): Qwen3.8-27B at MLX 4-bit, accelerated by the DFlash 2 block-diffusion drafter under the z-lab oMLX fork, tuned for long-context coding-agent workloads on an M4 Max / 64 GB.

Measured results (M4 Max, 64 GB, macOS 27.0)

Metric llama.cpp baseline This configuration
Decode (median) ~15 tok/s 809 tok/s
TTFT @ 32.5k fresh prompt 435 s every turn 486 s once
TTFT @ cached prefix turn โ€” 8โ€“16 s
Context window 32k 262k native
Needle recall @ 70k โ€” 5/5 ordered

Raw measurement records: GitHub repo โ†’ results/raw/.

Files

No weight files are hosted here โ€” use the upstream artifacts:

Engine configuration

Place as ~/.omlx/model_settings.json:

{
  "version": 1,
  "models": {
    "Qwen3.8-27B-4bit": {
      "dflash_enabled": true,
      "dflash_draft_model": "/Users/<you>/Models/mlx/Qwen3.8-27B-DFlash2",
      "dflash_draft_quant_enabled": true,
      "dflash_draft_quant_weight_bits": 4,
      "dflash_draft_quant_activation_bits": 16,
      "dflash_draft_quant_group_size": 64,
      "dflash_block_size": 5,
      "dflash_verify_mode": null,
      "dflash_in_memory_cache": true,
      "display_name": "Qwen3.8-27B 4bit + DFlash2"
    }
  }
}

Notes: block size โ‰ค5 per z-lab guidance for quantized targets; adaptive verify; L1 prefix cache is what turns multi-turn agent sessions from minutes-per-turn into seconds-per-turn.

Integrity requirement

Hash-verify both models against the hub manifests before first serve. Two of our three shards arrived corrupt after resumed downloads โ€” with valid safetensors headers and correct sizes. The model loads and then emits fluent gibberish. bench/hf_fetch.py in the GitHub repo wraps download + verification.

Sampling (per Qwen3.8 model card)

Thinking mode: temperature=1.0, top_p=0.95, top_k=20; reasoning effort xhigh/medium/low via the server's chat-template kwargs.

Citation

@misc{localflash2026,
  title={Losing the Draft, Keeping the Speed},
  author={chengyixu},
  year={2026},
  url={https://github.com/chengyixu/qwen38-dflash2-bench}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ChengyiX/Qwen3.8-27B-DFlash2-LocalFlash-M4Max

Base model

Qwen/Qwen3.8-27B
Finetuned
(212)
this model