Paper of the Week — Not All 4-bit Quantizers Are Equal: Deployment-Time Mitigation of PII Leakage in Fine-Tuned Small Language Models 4-bit quantization choice leaks fine-tuned PII at different rates — GPTQ-style methods outperform GGUF on privacy, independent of bit width. 2026-09-24T12:00:00.000Z Paper of the Week Paper of the Week researchpapersarxivpractical-ai

Paper of the Week — Not All 4-bit Quantizers Are Equal: Deployment-Time Mitigation of PII Leakage in Fine-Tuned Small Language Models

4-bit quantization choice leaks fine-tuned PII at different rates — GPTQ-style methods outperform GGUF on privacy, independent of bit width.

Weekly One research paper, broken down for people who build things.

Not All 4-bit Quantizers Are Equal: Deployment-Time Mitigation of PII Leakage in Fine-Tuned Small Language Models

Cristhian Kapelinski, Diego Kreutz. Published 2026-08-02. arXiv:2609.25014

One sentence summary

The quantization method you choose at deployment time — not just the bit width — meaningfully changes how much private training data a fine-tuned model will leak under extraction attacks.

Why this paper

Every team shipping a fine-tuned model to production is compressing it. The default assumption is that GGUF and GPTQ are interchangeable at the same bit width — this paper shows that assumption is wrong for privacy-sensitive workloads, and the fix costs nothing extra.

What they did

The authors fine-tuned small language models (SLMs) on datasets containing PII, then compressed them to 4 bits using multiple quantization methods. They then ran standard membership inference and extraction attacks against each compressed variant to measure how much private information survived into the deployable artifact. The key variable wasn’t the bit width — it was whether the quantizer used a calibration dataset to tune its rounding decisions.

Key findings

  • Methods that optimize rounding on a small calibration sample (GPTQ-style) retain significantly less extractable PII than post-hoc round-to-nearest methods (GGUF/GGML)
  • The gap persists across multiple model families and PII categories — it’s not an artifact of one architecture
  • Bit width alone does not predict leakage; two 4-bit models can have meaningfully different privacy profiles
  • Calibration-based quantization effectively acts as a weak unlearning step, disrupting memorized verbatim sequences without explicit forgetting algorithms
  • The privacy benefit comes essentially for free: calibration-based quantizers are already the default in many inference pipelines

Why it matters for practitioners

If your team fine-tunes on customer data and ships a compressed model — to an edge device, a private deployment, or a third-party integration — your quantization choice is a privacy decision you’re probably not treating as one. This paper gives you a concrete, zero-overhead lever: switch from GGUF to a calibration-based quantizer (GPTQ, AWQ) and you reduce extraction attack surface without changing bit width, model architecture, or inference speed.

What you can use today

  • Swap GGUF exports for GPTQ or AWQ variants when deploying fine-tuned models trained on private data — the auto-gptq and llm-compressor libraries both support calibration-based 4-bit quantization with a small calibration dataset you control
  • Run a quick extraction probe on your current compressed model before shipping: the pii-masking-200k dataset on Hugging Face gives you a ready-made PII corpus to test against
  • Treat quantization method as a field in your model card and deployment checklist alongside the usual safety and license metadata