Skip to content

Repository files navigation

FlagOS-Compressor

FlagOS-Compressor converts and quantizes local Hugging Face safetensors checkpoints. It supports BF16 conversion, selected-layer quantization, calibrated quantization, and output validation.

Why use FlagOS-Compressor

  • Convert supported low-precision checkpoint weights to BF16.
  • Quantize selected model modules with MSE, GPTQ, AWQ, or native AutoRound workflows.
  • Inspect source checkpoints and validate generated artifacts from one CLI.

Install

Install the project from a source checkout in editable mode:

pip install -e .

See the installation guide for prerequisites and optional extras.

Quick start

Inspect a checkpoint, quantize selected weights, and validate the output:

flagos-compressor inspect --input /path/to/model
flagos-compressor quantize \
  --input /path/to/model \
  --output /path/to/output \
  --select linear \
  --bits 8 \
  --strategy group \
  --group-size 128 \
  --backend cuda
flagos-compressor validate --input /path/to/output

Run quantization with --dry-run before writing a large output directory. Structural validation does not replace runtime or performance testing. See the output validation guide for the complete verification workflow.

Documentation and help

Start with the documentation portal.

About

No description, website, or topics provided.

Resources

Stars

5 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages