From 19 to 185 tokens per second: bringing llama.cpp and vLLM to a dual R9700
How slow is “out of the box” when you’re running a 27-billion-parameter model on a brand-new, dual-GPU workstation? For the first several days I spent getting llama.cpp and vLLM running on a pair of Radeon AI PRO R9700s, the answer was a painful 19 tokens per second. This is the story of that journey, all the way up to 185.