Alex K. Miller

Economic Development Data Scientist

From 19 to 185 tokens per second: bringing llama.cpp and vLLM to a dual R9700

How slow is “out of the box” when you’re running a 27-billion-parameter model on a brand-new, dual-GPU workstation? For the first several days I spent getting llama.cpp and vLLM running on a pair of Radeon AI PRO R9700s, the answer was a painful 19 tokens per second. This is the story of that journey, all the way up to 185.

Read More

Why does IATI validation enforce element order?

If you’ve ever had to write software that imports or exports IATI data, you may have noticed that the order of the XML elements is considered a critical part of the schema validation. Why was this change made historically, and why does it persist today?

Read More

Welcome to my blog

Hello and welcome to the first post on my blog. I’m intending this blog to be a place to host some of my previous and current writings, documenting some interesting technical challenges I’ve faced as well as the solutions I’ve designed to overcome them.

Read More