llama.cpp now runs Kev-4B GGUF decision models on-device
Original titleDecision models now run on device in llama.cpp. Free, fast, private!
AISummary
llama.cpp can now run decision models on-device, according to Clément Delangue of Hugging Face. He says the setup is free, fast, and private, and gives the command llama serve -hf ggml-org/Kev-4B-GGUF to start it.
Source: Clément Delangue · x.comPublished · added here