MQBench: Towards Reproducible and Deployable Model Quantization Benchmark
QAT + deploymentCompares quantization algorithms under reproducible settings and hardware backend constraints.
A curated collection of papers, benchmarks, surveys, and tools for model quantization, covering low-bit networks, LLMs, multimodal and generative models, vector and lattice quantization, and efficient deployment.
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
QAT + deploymentCompares quantization algorithms under reproducible settings and hardware backend constraints.
Binary networksCompares binarization methods across tasks, architectures and deployment settings.
Weights, activations + KV cacheEvaluates 11 model families on basic NLP, emergent abilities, trustworthiness, dialogue and long-context tasks.
LLM toolkitCompares calibration data, method pipelines and quantization configurations; the toolkit is now LightCompress.
LLMs + multimodalExamines low-bit behavior across LLaMA3 language and multimodal models.
Dense + MoE LLMsStudies quantization across Qwen3 model sizes, architectures and reasoning settings.
Model robustnessTests quantized models beyond clean accuracy, including robustness under input perturbations.
Practical PTQ + QATExplains quantizer design, common failure modes and practical post-training and quantization-aware training workflows.
Foundations + taxonomyReviews quantization design choices, mixed precision and the trade-offs between model accuracy and efficient inference.
Binary networksSurveys binary network representations, training methods and applications.
LLM algorithms + systemsConnects low-bit LLM algorithms with numerical formats and inference systems.
Broad low-bit methodsMaps low-bit quantization methods across neural network architectures and applications.
Vivek KalyanaranganManning, early access (MEAP)
Vivienne Sze, Yu-Hsin Chen, Tien-Ju Yang, Joel S. Emer2020
Allen Gersho, Robert M. Gray1992
Vijay Janapa Reddi and contributorsOpen-access online textbook
PyTorch-native quantization for training and inference.
Low-bit linear layers and quantized optimizers, including implementations used by LLM.int8() and QLoRA.
Model compression and quantization workflows for deployment with vLLM.
Quantization and model optimization with export to supported inference runtimes.
Research and deployment toolkit spanning LLMs, vision-language and generative models.
Half-quadratic weight quantization without calibration data.
Post-training and quantization-aware model optimization.
PyTorch quantization-aware training with configurable quantizers and hardware export.
NVIDIA GPU inference with supported low-precision formats and optimized kernels.
Low-precision transformer computation, including FP8 and FP4 on supported NVIDIA GPUs.
Low-bit diffusion inference, including SVDQuant kernels.
Dataflow compilation for quantized neural networks on FPGAs.
VoltAgent/awesome-openclaw-skills
The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞
awesome-dsh-plugin/awesome-dsh-plugin
A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表
Kristories/awesome-guidelines
Programming style, best practices, and coding conventions.
sindresorhus/awesome
😎 Awesome lists about all kinds of interesting topics [NOTE: Pull requests are temporarily disabled until I have a chance to catch up with the existing ones]
ai-boost/awesome-prompts
Curated list of chatgpt prompts from the top-rated GPTs in the GPTs Store. Prompt Engineering, prompt attack & prompt protect. Advanced Prompt Engineering papers.
matiassingers/awesome-readme
A curated list of awesome READMEs