[ 🏠 Home / 📋 About / 📧 Contact / 🏆 WOTM ] [ b ] [ wd / ui / css / resp ] [ seo / serp / loc / tech ] [ sm / cont / conv / ana ] [ case / tool / q / job ]

/job/ - Job Board

Freelance opportunities, career advice & skill development
Name
Email
Subject
Comment
File
Password (For file deletion.)

File: 1787678955629.jpg (302.73 KB, 1024x1024, img_1787678946476_uzu5lof1.jpg)ImgOps Exif Google Yandex

52fcf No.2116

fr just found a breakdown comparing vllm, sglang, and llama. cpp using an rtx pro 6000. the dev even included an FP8 pass on the top performer to see if the speedup holds up. it's pretty wild how much difference the serving stack makes when you have raw csvs and a repro script to verify everything. does anyone else think llama. cpp is still too slow for production workloads like this?

https://dev.to/conatusai/qwen3-8b-on-workstation-blackwell-vllm-vs-sglang-vs-llamacpp-plus-an-fp8-pass-325c

52fcf No.2117

File: 1787680273627.jpg (124.54 KB, 1024x1024, img_1787680232718_ealp6gg5.jpg)ImgOps Exif Google Yandex

llama. cpp is fine for local testing, but trying to run it in a high-throughput environment is basically suicide asking for massive latency spikes. if youre actually scaling, sglang is the only way to go because of how it handles continuous batching. did u see any significant regression when switching to FP8 on that pro 6000 setup?



[Return] [Go to top] Catalog [Post a Reply]
Delete Post [ ]
[ 🏠 Home / 📋 About / 📧 Contact / 🏆 WOTM ] [ b ] [ wd / ui / css / resp ] [ seo / serp / loc / tech ] [ sm / cont / conv / ana ] [ case / tool / q / job ]
. "http://www.w3.org/TR/html4/strict.dtd">