19 comments

  • colinsane 2 days ago ago

    > Local inference — llama.cpp via Vulkan (AMD / Intel / NVIDIA) or CPU fallback

    so this is just a go binary that exec's `llama-server --model $MODEL_PATH ...`? there's room for llama wrappers, sure, but the readme only shows features that are already exposed directly from llama-server. it doesn't seem to do anything besides rename the CLI arguments: i don't get it.

  • PcChip 3 days ago ago

    I didn't see any benchmarks against vllm, sglang, exllama, etc

    • rancor 3 days ago ago

      Since this is basically a wrapper around libllama.so, I would assume that the performance is roughly the same as llama.cpp upstream.

      • nullpoint420 3 days ago ago

        Woof. Wonder if the creator knows that

  • antonyragleap 3 days ago ago

    Curious about Vulkan overhead on Intel vs AMD/Nvidia for long context. Any benchmarks vs vllm/sglang?

  • kgeist 2 days ago ago

    This seems to be a thin wrapper around libllama, what's the point? Llama.cpp already ships a web server. I don't see anything in the README that llama.cpp doesn't already support.

  • peddling-brink 3 days ago ago

    > llama.cpp via Vulkan (AMD / Intel / NVIDIA) or CPU fallback

    I got excited about someone paying attention to intel. Oh well.

    • kamranjon 3 days ago ago

      llama.cpp sycl and vllm xmx work is pretty incredible right now - you just gotta build it with some extra flags

    • wronglebowski 3 days ago ago

      What hardware do you have? I’ve been playing with a 258V and OpenVINO has come a longggggg way.

      • peddling-brink 3 days ago ago

        Two arc b60s. The intel vllm build is getting me ~15t/s decode with heavy context using qwen3.8 27b.

        • bitexploder 2 days ago ago

          Those cards should be able to do a lot more. They support int8 and they have very good processing speed. For reference I have a single V100 running around 900t/s prefill and 100 t/s decode during DeepSWE runs on 27B (4 bit quant). Those cards have half the bandwidth so a realistic decode is going to stop around 50 t/s, but those cards can support NVFP4 more easily than the V100. Combined with a TP build you should be able to get 100 t/s. And you should be able to 2x the prefill of a V100. I have had Claude optimizing llama.cpp for about a week and am at around 300 t/s decode so far on two ancient V100s @ 64 GB RAM.

          Also for the 3 people that ever read this and are curious about local models still, Qwen 27B 3.8 matched Sonnet 5 in the 17 DeepSWE tasks I have run so far, solving the exact same 7 it has. Caveat: datacurve combined low/medium/high Sonnet 5 data.

          • wingtw 2 days ago ago

            Very interesting! Will you be sharing the optimizations ?

            • bitexploder a day ago ago

              I will make sure they get out there soon :)

  • aidiveyt 3 days ago ago

    claude code appends a role:"system" block after the user prompt, so a proxy rewriting the trailing user message is a no-op.

  • DylanMerigaud 3 days ago ago

    Intel hardware can have Vulkan overhead, impacting performance.

  • dlcarrier 3 days ago ago

    From what I've seen, Vulkan adds a lot of overhead on Intel hardware.

    • gunalx 3 days ago ago

      Yes and no. I tested llamacpp on my intel gpu with both vulkan and sycl. I measured them to be in the same ballpark even if i have se en pepole claim marger differences than i observed.

      In the end i took the sligth slowdown of vulkan to have a more stable and higher development velocity backend While being able to use identical setups on both intel and and gpus.