Transformers can now run GGUF llama.cpp checkpoints directly using ggml Metal kernels on Mac
Original titleRun GGUF models directly with transformers. This work brings ggml's Metal kernels to the transformers ecosystem, increasing compatibility...
AISummary
Hugging Face's transformers library can now load the same GGUF checkpoints used by llama.cpp, with fast local inference on Mac powered by ggml's Metal kernels. The author, Georgi Gerganov, says the work brings those kernels into the transformers ecosystem to increase compatibility and performance.
Source: Georgi Gerganov · x.comPublished · added here