Gemma 4 31B uncensored · format family
59 GB → 20 GBNVFP4 checkpoint · vLLM and SGLang validated
- Result
BF16 evaluation: 0/686 effective refusals across four datasets. The 20 GB NVFP4 build, reduced from 59 GB, was validated in vLLM and SGLang. GGUF: 12–31 GB with 0–1/100 hard refusals across quant levels.
Evidence: BF16 (opens in a new tab)NVFP4 (opens in a new tab)GGUF (opens in a new tab)
Problem and decision
- Problem
Remove refusal behavior from a 30.7B multimodal model without losing a deployable format set.
- My contribution & decision
Applied weight-level abliteration, then carried the same model lineage through BF16, ModelOpt NVFP4, and a GGUF K-quant ladder. For multimodal NVFP4 serving, the vision tower, embedder, and lm_head stayed in BF16.
