52fcf No.2116
fr just found a breakdown comparing vllm, sglang, and llama. cpp using an rtx pro 6000. the dev even included an FP8 pass on the top performer to see if the speedup holds up.
it's pretty wild how much difference the serving stack makes when you have
raw csvs and a repro script to verify everything. does anyone else think
llama. cpp is still too slow for production workloads like this?
https://dev.to/conatusai/qwen3-8b-on-workstation-blackwell-vllm-vs-sglang-vs-llamacpp-plus-an-fp8-pass-325c