52 points | by zxdc7896 15 hours ago ago
2 comments
> https://github.com/darshi1337/apogee/blob/main/MODELS.md#loc...
> Unlike Ollama, [llama.cpp] serves one model at a time: the GGUF you launched it with.
FWIW: llama-server's supported multiple models for a while now:
https://github.com/ggml-org/llama.cpp/tree/master/tools/serv...
Ohh thank you for your suggestion. I will try implementing it this weekend.
> https://github.com/darshi1337/apogee/blob/main/MODELS.md#loc...
> Unlike Ollama, [llama.cpp] serves one model at a time: the GGUF you launched it with.
FWIW: llama-server's supported multiple models for a while now:
https://github.com/ggml-org/llama.cpp/tree/master/tools/serv...
Ohh thank you for your suggestion. I will try implementing it this weekend.