FlagOS-Compressor converts and quantizes local Hugging Face safetensors checkpoints. It supports BF16 conversion, selected-layer quantization, calibrated quantization, and output validation.
- Convert supported low-precision checkpoint weights to BF16.
- Quantize selected model modules with MSE, GPTQ, AWQ, or native AutoRound workflows.
- Inspect source checkpoints and validate generated artifacts from one CLI.
Install the project from a source checkout in editable mode:
pip install -e .See the installation guide for prerequisites and optional extras.
Inspect a checkpoint, quantize selected weights, and validate the output:
flagos-compressor inspect --input /path/to/modelflagos-compressor quantize \
--input /path/to/model \
--output /path/to/output \
--select linear \
--bits 8 \
--strategy group \
--group-size 128 \
--backend cudaflagos-compressor validate --input /path/to/outputRun quantization with --dry-run before writing a large output directory. Structural validation does not replace runtime or performance testing. See the output validation guide for the complete verification workflow.
Start with the documentation portal.
- Tutorial for a complete first run
- How-to guides for installation, quantization, calibration, model families, and troubleshooting
- Reference for CLI, formats, selectors, recipes, outputs, and compatibility
- Troubleshooting guide when a command fails