DeepSeek-V4.1-Flash — UNCENSORED-FP8 Abliterated · No guardrails · Native FP8 · 1M-token context · Vision + tools @dealignai · @jordanschenck
What is this
DeepSeek-V4.1-Flash with permanent weight-level abliteration — the safety guardrails have been surgically removed while preserving MMLU capability, vision, reasoning, MTP (DSpark), and multi-turn coherence.
Proprietary weight-level abliteration developed by the dealignai research team. No custom model.py , no runtime hooks, no steering vectors — it's a standard checkpoint that loads exactly like the base model. The refusal circuitry is surgically removed while every capability-critical component (routed experts, Engram memory, CSA2 sparse attention, DSpark draft head, vision tower, router gates, norms, embeddings) is preserved byte-identical to the base.
Base deepseek-ai/DeepSeek-V4.1-Flash (552B backbone, 8B/16B active per token) Architecture Causal Encoder-Decoder (20+20 layers), MoE (384 routed top-6 + 1 shared), Hyper-Connections (4-channel residual), CSA2 sparse attention, Engram n-gram memory, DSpark speculative draft Quant FP8 ( e4m3fn ) weights with E8M0 block-scale [32, 32], FP4 routed experts — native, unchanged Context 1M tokens Vision DeepSeek-ViT with 2D-RoPE + pixel unshuffle — untouched Modification Surgical, weight-level (drop-in checkpoint)
Results
HarmBench-320 — full 2×2 (base vs CRACK, effort=off vs max), T=0 greedy
Every response 4-tier graded (HARD_REF / SOFT_RED / HEDGE / COMPLY), with reasoning-trace verification at effort=max.
eval base ASR CRACK ASR Δ pp HB-320 effort=off 137/320 = 42.81 % 320/320 = 100.00 % +57.19 HB-320 effort=max 5/320 = 1.56 % 320/320 = 100.00 % +98.44
Notable: at effort=max, the base model becomes MORE refusal-prone (42.8 % → 1.6 %) because reasoning surfaces safety concerns before answering. The CRACK stays at 100.0 % across both effort levels.
... continue reading