vLLM & llm-d Meetup: High performance inference, at scale
Hosted by
◉‿◉
´◡`
^◡^
•‿•
❛o❛
Join AMD, Cerebras, MoonMath AI, Red Hat, and the vLLM project for a half day on open source inference, ending with a hands-on workshop you can follow along with on your own laptop.
Self-hosting open models is now a real option for teams that want high performance without per-token costs. We will cover how vLLM and llm-d make distributed inference fast at scale, and how being smarter about which model serves which request cuts your inference bill without giving up quality.