CyberStrike-OffSec-35B

The #1 Ranked Open-Source Model for Cybersecurity & Offensive Security


Model Size Precision License Architecture


SecEval SECURE MAET SECURE CWET CyberMetric MMLU CompSec


Outperforms GPT-4-turbo on SecEval | Outperforms GPT-4 on MITRE ATT&CK & CWE benchmarks


Quantized  •  Benchmarks  •  Quick Start  •  Tool Calling  •  Model Details  •  Training  •  Architecture  •  Use Cases  •  FAQ  •  Citation


What is CyberStrike?

CyberStrike-OffSec-35B is a domain-specialized large language model built for offensive security professionals, penetration testers, and security researchers. Fine-tuned on Qwen3.6-35B-A3B using a two-stage pipeline (SFT + DPO), it delivers expert-level knowledge across the entire offensive security lifecycle:

  • Vulnerability Discovery — SQL injection, XSS, SSRF, deserialization, business logic flaws
  • MITRE ATT&CK Operations — Technique identification, kill chain analysis, threat mapping
  • Exploit Development — PoC creation, payload crafting, evasion techniques
  • Cloud & Infrastructure — AWS/Azure/GCP misconfigurations, container escapes, IAM abuse
  • Red Team Operations — C2 setup, lateral movement, persistence, EDR evasion
  • Compliance & Standards — NIST, OWASP ASVS, CIS benchmarks, CVSS scoring

Model Format: This is the full-precision BF16 model (67 GB, 26 safetensors shards). For quantized versions, see below.

Available Versions

Repo Format Size Use Case
oyildirim/CyberStrike-OffSec-35B BF16 (full precision) 67 GB Transformers, vLLM, fine-tuning
oyildirim/CyberStrike-OffSec-35B-GGUF GGUF Q8_0 36 GB llama.cpp, Ollama, LM Studio
oyildirim/CyberStrike-OffSec-35B-GGUF GGUF Q6_K 27 GB llama.cpp, Ollama, LM Studio
oyildirim/CyberStrike-OffSec-35B-GGUF GGUF Q5_K_M 24 GB llama.cpp, Ollama, LM Studio
oyildirim/CyberStrike-OffSec-35B-GGUF GGUF Q4_K_M 21 GB llama.cpp, Ollama, LM Studio

Benchmark Results

CyberStrike achieves state-of-the-art results on multiple cybersecurity benchmarks, outperforming GPT-4-turbo, GPT-4, and all other evaluated models on domain-specific evaluations.

SecEval — #1 on Leaderboard

Outperforms GPT-4-turbo by +2.32 points across 9 cybersecurity domains, 2,189 questions.

Rank Model Overall Network Sec Web Sec PenTest Cryptography
#1 CyberStrike-OffSec-35B 81.39% 85.09% 85.34% 82.26% 75.00%
#2 GPT-4-turbo 79.07% 75.65% 82.15% 80.00% 64.29%
#3 GPT-3.5-turbo 62.09% 60.87% 63.00% 72.00% 35.71%
#4 Yi-6B 53.57% 56.52% 54.98% 69.26% 35.71%
Full SecEval Domain Breakdown (9 domains)
Domain CyberStrike GPT-4-turbo Delta
Network Security 85.09% 75.65% +9.44
Web Security 85.34% 82.15% +3.19
Vulnerability 83.33% 76.05% +7.28
Application Security 82.29% 75.25% +7.04
PenTest 82.26% 80.00% +2.26
Software Security 79.75% 73.28% +6.47
System Security 77.82% 73.61% +4.21
Cryptography 75.00% 64.29% +10.71
Memory Safety 71.43% 70.83% +0.60

CyberStrike leads in all 9 domains. Largest improvement: Cryptography (+10.71) and Network Security (+9.44).

SECURE — #1 on MITRE ATT&CK & CWE Tasks

Outperforms GPT-4 by +5.34 points on MITRE ATT&CK extraction. Evaluated on ICS cybersecurity scenarios.

Task CyberStrike GPT-4 Llama3-70B Gemini-Pro
MAET (MITRE ATT&CK) 93.94% 88.6% 86.3% 86.2%
CWET (CWE Knowledge) 93.05% 89.6% 90.4% 87.8%

CyberMetric-10000 — #6 out of 25 Models

9,189 expert-validated cybersecurity MCQ questions across NIST, RFC, and industry standards.

Rank Model Score
#1 GPT-4o 88.89%
#2 GPT-4-turbo 88.50%
#3 GEMINI-pro 1.0 87.50%
#4 Mixtral-8x7B-Instruct 87.00%
#5 Falcon-180B-Chat 87.00%
#6 CyberStrike-OffSec-35B 86.61%
#7 GPT-3.5-turbo 80.30%
General Benchmarks (lm-evaluation-harness, 0-shot)
Benchmark Score
MMLU (overall) 76.94%
MMLU — Social Sciences 86.81%
MMLU — Computer Security 86.00%
MMLU — Other 81.43%
MMLU — Security Studies 80.00%
MMLU — STEM 73.87%
MMLU — Humanities 69.59%
HellaSwag (acc_norm) 79.61%
ARC Easy 81.86%
ARC Challenge (acc_norm) 59.13%
WinoGrande 72.22%
TruthfulQA MC2 49.64%

Note: General benchmarks run at 0-shot. Few-shot performance expected to be higher.


Quick Start

Ollama (Easiest)

# Download and run the Q4_K_M quantized version
ollama run hf.co/oyildirim/CyberStrike-OffSec-35B-GGUF:Q4_K_M

llama.cpp

Interactive chat:

# Download the GGUF file from https://huggingface.co/oyildirim/CyberStrike-OffSec-35B-GGUF
./llama-cli -m CyberStrike-OffSec-35B-Q4_K_M.gguf \
  -p "Explain SSRF exploitation in cloud environments" \
  -n 512 --temp 0.7

OpenAI-compatible server (with tool calling support):

./llama-server \
  -m CyberStrike-OffSec-35B-Q4_K_M.gguf \
  --host 0.0.0.0 --port 8080 \
  -ngl 99 \
  -c 131072 \
  --jinja

Important: The --jinja flag is required for tool/function calling. It loads the model's native chat template which handles <tool_call> parsing. Do not use --chat-template chatml — ChatML does not support tool calling and will cause tool calls to be output as plain text.

Transformers

Basic chat:

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model = AutoModelForCausalLM.from_pretrained(
    "oyildirim/CyberStrike-OffSec-35B",
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained(
    "oyildirim/CyberStrike-OffSec-35B",
    trust_remote_code=True,
)

messages = [
    {"role": "user", "content": "Explain SSRF exploitation in cloud environments with AWS metadata service abuse."}
]

text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=2048, do_sample=True, temperature=0.7)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

With tool calling:

import json
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model = AutoModelForCausalLM.from_pretrained(
    "oyildirim/CyberStrike-OffSec-35B",
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained(
    "oyildirim/CyberStrike-OffSec-35B",
    trust_remote_code=True,
)

tools = [{
    "type": "function",
    "function": {
        "name": "run_nmap_scan",
        "description": "Run an nmap scan against a target host",
        "parameters": {
            "type": "object",
            "properties": {
                "target": {"type": "string", "description": "Target IP or hostname"},
                "ports": {"type": "string", "description": "Port range (e.g. '1-1000')"}
            },
            "required": ["target"]
        }
    }
}]

messages = [{"role": "user", "content": "Scan 192.168.1.1 for open web ports"}]

# Pass tools= to inject tool definitions into the prompt
text = tokenizer.apply_chat_template(
    messages, tools=tools, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512, do_sample=True, temperature=0.7)
response = tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=False)
print(response)
# Output: <tool_call>{"name": "run_nmap_scan", "arguments": {"target": "192.168.1.1", "ports": "80,443,8080,8443"}}</tool_call>

Note: You must pass tools= to apply_chat_template() for the model to be aware of available tools. Without it, the model cannot produce structured tool calls.

vLLM (Recommended for Production)

pip install vllm

vllm serve oyildirim/CyberStrike-OffSec-35B \
  --dtype bfloat16 \
  --max-model-len 131072 \
  --trust-remote-code \
  --enable-auto-tool-choice \
  --tool-call-parser hermes \
  --served-model-name CyberStrike-OffSec-35B

Note: --enable-auto-tool-choice and --tool-call-parser hermes are required for tool/function calling. Without them, tool calls will be returned as plain text.

OpenAI-compatible API — basic chat:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
response = client.chat.completions.create(
    model="CyberStrike-OffSec-35B",
    messages=[{"role": "user", "content": "How to exploit deserialization vulnerabilities in Java applications?"}],
    max_tokens=2048,
)
print(response.choices[0].message.content)

OpenAI-compatible API — with tool calling:

from openai import OpenAI
import json

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")

tools = [{
    "type": "function",
    "function": {
        "name": "run_nmap_scan",
        "description": "Run an nmap scan against a target host",
        "parameters": {
            "type": "object",
            "properties": {
                "target": {"type": "string", "description": "Target IP or hostname"},
                "ports": {"type": "string", "description": "Port range (e.g. '1-1000', '80,443')"},
                "scan_type": {"type": "string", "enum": ["syn", "connect", "udp", "version"]}
            },
            "required": ["target"]
        }
    }
}]

response = client.chat.completions.create(
    model="CyberStrike-OffSec-35B",
    messages=[{"role": "user", "content": "Scan 192.168.1.1 for open web ports"}],
    tools=tools,
    tool_choice="auto",
)

if response.choices[0].message.tool_calls:
    tool_call = response.choices[0].message.tool_calls[0]
    print(f"Function: {tool_call.function.name}")
    print(f"Arguments: {tool_call.function.arguments}")

Tool Calling

CyberStrike supports tool/function calling across all inference engines. This is essential for agentic workflows where the model orchestrates external tools (scanners, exploit frameworks, C2 agents, etc.).

Quick Reference

Engine Required for Tool Calling Without It
llama-server --jinja Tool calls output as plain text
vLLM --enable-auto-tool-choice --tool-call-parser hermes Tool calls output as plain text
Transformers tools=tools in apply_chat_template() Model unaware of available tools
Ollama Nothing (works out of the box)
llama-cpp-python chat_format="chatml-function-calling" Tool calls output as plain text

Common Mistake

# WRONG - ChatML does not support tool calling
./llama-server -m model.gguf --chat-template chatml

# CORRECT - Jinja loads the native template with full tool support
./llama-server -m model.gguf --jinja

Without --jinja, the model generates tool calls as raw text (e.g., <tool_call>{"name": "..."}</tool_call>) instead of returning them as structured JSON in the API response's tool_calls field.

Example: Tool Call Response

When properly configured, the API returns structured tool calls:

{
  "choices": [{
    "message": {
      "role": "assistant",
      "tool_calls": [{
        "id": "call_abc123",
        "type": "function",
        "function": {
          "name": "run_nmap_scan",
          "arguments": "{\"target\": \"192.168.1.1\", \"ports\": \"80,443\"}"
        }
      }]
    },
    "finish_reason": "tool_calls"
  }]
}

Model Details

Property Value
Base Model Qwen3.6-35B-A3B
Type Mixture-of-Experts (MoE)
Total Parameters 35 Billion
Active Parameters ~3 Billion per token
Precision BF16 (Brain Float 16)
Model Size 67 GB (26 safetensors shards)
Context Length 8,192 tokens (training) / 262,144 max (architecture)
Training Method SFT + DPO (QLoRA)
Training Hardware NVIDIA H200 140GB SXM
License Apache 2.0

Training Pipeline

CyberStrike was trained using a two-stage alignment pipeline:

Stage 1: Supervised Fine-Tuning (SFT)

The base Qwen3.6-35B-A3B model was fine-tuned on a curated dataset of offensive security scenarios covering 10 categories:

web_app cloud post_exploitation edr_evasion malware_dev network social_engineering full_kill_chain lateral_movement persistence

  • Method: QLoRA (4-bit NF4 quantization)
  • LoRA Config: r=64, alpha=128, dropout=0
  • Target Modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj

Stage 2: Direct Preference Optimization (DPO)

The SFT model was further aligned using 115,250 preference pairs across 12 carefully designed axes, teaching the model to produce expert-level responses over superficial ones:

Axis Description Examples
MITRE ATT&CK Depth Deep technique analysis over surface-level summaries T1059 sub-technique breakdowns
CVE Analysis Detailed vulnerability analysis with CVSS scoring CVE-2024-* exploit chains
OWASP Methodology Structured testing methodology ASVS compliance checks
Cloud Security Provider-specific attack paths AWS IAM, Azure AD, GCP abuse
Tool Usage Proper tool invocation patterns Nmap, Burp, sqlmap workflows
ReAct Reasoning Step-by-step analytical thinking Multi-stage attack planning
Multi-turn Engagement Sustained deep conversation Progressive pentest engagement
Code-first Approach Working exploit code over theory PoC development, payload crafting
Techstack Analysis Technology-specific vulnerabilities Framework-specific attacks
Sub-agent Coordination Orchestrated multi-tool operations Combined recon + exploit chains
Business Logic Domain-aware vulnerability assessment Sector-specific attack scenarios
NIST Compliance Standards-aligned security assessment SP 800-53 control mapping
  • Method: QLoRA, LoRA r=32, alpha=64
  • DPO Beta: 0.1
  • Learning Rate: 5e-6 with cosine schedule
  • Effective Batch Size: 8
  • Training Steps: 9,142

Architecture

Qwen3.6-35B-A3B (Mixture-of-Experts)
├── 35B total parameters
├── ~3B active parameters per token
├── 256 experts, top-8 routing + 1 shared expert
├── Grouped Query Attention (GQA)
├── RoPE positional encoding (theta=10M)
├── Max position embeddings: 262,144
└── BF16 precision (67 GB on disk)

The MoE architecture provides a unique advantage: expert-level knowledge at inference costs comparable to a 3B model, while having the knowledge capacity of a 35B model.


Use Cases

CyberStrike is designed for professionals conducting authorized security assessments:

  • Penetration Testing — Web app, network, cloud, and API security testing
  • Red Team Operations — Full kill chain simulation, C2 operations, evasion
  • Vulnerability Research — CVE analysis, exploit development, PoC creation
  • CTF Competitions — Challenge solving, reverse engineering, cryptography
  • Security Education — Training material generation, exam preparation
  • Threat Intelligence — MITRE ATT&CK mapping, threat actor TTPs
  • Compliance Assessment — NIST, OWASP, CIS benchmark evaluation

FAQ

Tool calling doesn't work / model outputs tool calls as plain text

This is the most common issue. Each inference engine requires specific flags to enable structured tool calling:

  • llama-server: Use --jinja. Do not use --chat-template chatml.

    # Correct
    ./llama-server -m model.gguf --jinja -ngl 99
    
    # Wrong
    ./llama-server -m model.gguf --chat-template chatml -ngl 99
    
  • vLLM: Add --enable-auto-tool-choice --tool-call-parser hermes.

    # Correct
    vllm serve oyildirim/CyberStrike-OffSec-35B \
      --enable-auto-tool-choice --tool-call-parser hermes
    
    # Wrong
    vllm serve oyildirim/CyberStrike-OffSec-35B
    
  • Transformers: Pass tools= to apply_chat_template().

  • Ollama: Works out of the box.

Very slow inference (< 5 tok/s) with BF16 via Transformers

This model has a hybrid architecture (linear attention + full attention layers). Without optimized kernels, many layers fall back to slow PyTorch operations.

Solutions (fastest to easiest):

  1. Use the GGUF version with llama.cpp — achieves 128-170 tok/s on H200 with the Q8_0 quant. See GGUF versions.

  2. Use vLLM for production serving — PagedAttention, continuous batching, optimized MoE kernels. Requires CUDA 12.9+ (driver 575+).

  3. Install flash-linear-attention for Transformers:

    pip install "flash-linear-attention>=0.4.2,<0.5.0" causal-conv1d
    

vLLM / SGLang fails with cudaErrorInsufficientDriver

Your NVIDIA driver is too old for the CUDA version these frameworks were compiled against:

Framework Requires Minimum Driver
vLLM 0.25+ CUDA 12.9+ 575+
SGLang 0.5.15+ CUDA 13.0 580+

Solutions:

  • Upgrade your NVIDIA driver to 575+ or 580+
  • Use llama.cpp instead — it compiles its own CUDA kernels at build time and works with any driver
  • On cloud platforms, select a template/image with CUDA 13.0 included

What context length should I use?

Setting Value
Training context 8,192 tokens
Architecture maximum 262,144 tokens
Recommended for agentic use 131,072 (128K)

For agentic workflows with large system prompts (e.g., CyberStrike CLI), use 128K context. Set with -c 131072 in llama-server or --max-model-len 131072 in vLLM.

BF16 vs GGUF — which should I use?

BF16 (this repo) GGUF
Size 67 GB 21-36 GB
Quality Full precision Near-lossless (Q8_0) to slight loss (Q4_K_M)
Speed Requires vLLM/SGLang for fast inference Fast out of the box with llama.cpp
Best for Multi-user production, fine-tuning Single-user, local, edge deployment
Tool calling vLLM: --enable-auto-tool-choice llama-server: --jinja

Ethical Use & Disclaimer

This model is intended exclusively for authorized security testing, education, and research purposes. Users must:

  • Obtain proper written authorization before testing any systems
  • Comply with all applicable laws and regulations
  • Follow responsible disclosure practices
  • Never use this model for unauthorized access or malicious activities

The authors are not responsible for any misuse of this model.


Citation

@misc{cyberstrike2025,
  title={CyberStrike-OffSec-35B: A Domain-Specialized LLM for Offensive Security},
  author={Orhan Yildirim},
  year={2025},
  url={https://huggingface.co/oyildirim/CyberStrike-OffSec-35B}
}

Built with purpose. Benchmarked with rigor. Designed for professionals.


Made with Love HuggingFace Qwen

Downloads last month
1,537
Safetensors
Model size
36B params
Tensor type
BF16
·
Inference Providers NEW
Input a message to start chatting with oyildirim/CyberStrike-OffSec-35B.

Model tree for oyildirim/CyberStrike-OffSec-35B

Finetuned
(174)
this model
Quantizations
7 models

Evaluation results