Hugging Face Transformers Now Runs Llama.cpp's GGUF Quants Natively | Smart Chunks
Hugging Face just let Transformers load llama.cpp's GGUF files directly, shrinking an 8.42GB model to 2.74GB without leaving PyTorch. smartchunks.com/transformer... Hugging Face Transformers now natively supports llama.cpp GGUF quantized models, reusing ggml kernels for near-native speed on Apple Silicon and letting developers fine-tune quantized checkpoints in PyTorch.