Partiful logo
Get the app
Login

Self-Host Frontier AI: Running Open-Weight Models on AWS - #LATechWeek

Friday, Oct 16
1:00pm
Hosted by
´◡`
ツ
0/253 spots left
The open-weight model revolution is here — frontier models matching or beating closed APIs on benchmarks, and you can run them on your own infrastructure. But hosting them in production isn't as simple as pulling weights from HuggingFace. How much GPU memory do you actually need? What happens to throughput as context windows grow? When does managed infrastructure save you money vs. burn it? And how do you go from "it runs on my machine" to serving real production traffic? Join us for a 90-minute deep dive where we'll cut through the hype and show you exactly how to deploy the latest open-weight frontier models — from choosing the right hardware to production-grade serving at scale. What You'll Learn 🏗️ The GPU Math That Matters — Understand the memory tradeoffs that determine whether your deployment handles 1 user or 100 concurrent users. Context length vs. throughput — and how to think about it 🗺️ A Clear Decision Framework — From fully managed endpoints to self-hosted clusters — which path fits your model, your traffic pattern, and your team's ops appetite 💰 The Cost Math Your CFO Will Ask About — Managed APIs vs. self-hosted: when does convenience justify the premium? We'll walk through real break-even scenarios 🔧 Day-0 Model Support — How to run the newest models the day they drop, without waiting for official platform support 🌐 Data Sovereignty & Control — Why some teams are moving off API providers and back to self-hosted — and when it actually makes sense Models We'll Cover We'll walk through real deployment scenarios using today's top open-weight models — from lightweight models that fit on a single GPU to the largest frontier models pushing hardware limits: GLM • Kimi K3 • Gemma • DeepSeek • MiniMax M3 (and whatever drops between now and October) Covering the full spectrum — from models any team can deploy affordably to trillion-parameter beasts that require serious infrastructure planning. Who Should Attend • ML/AI Engineers — Building inference pipelines and choosing between serving frameworks • Platform Teams — Evaluating managed vs. self-managed GPU infrastructure • CTOs & Founders — Deciding whether to self-host or use API providers • Anyone who's Googled "how much VRAM do I need for \[model\]" in the last month Ideal for teams actively evaluating self-hosted inference — whether for data sovereignty, cost control, customization, or all three. What You'll Walk Away With ✅ A decision framework for choosing the right deployment path for any open model ✅ Hardware selection guidance — how to right-size your GPU setup ✅ Key considerations for production-grade model serving ✅ Cost comparison: managed APIs vs. self-hosted vs. GPU cloud providers ✅ Direct access to AWS Solutions Architects who deploy these models with startups weekly Event Details 📅 Date: October 16, 2026 1:00 PM PT (LA Tech Week) ⏱️ Duration: 90 minutes 📍 Location: 2450 Colorado Ave, Santa Monica, CA 90404 👥 💬 Bring: Your toughest "how do I host X model?" questions Hosted By AWS Startups — Built by Solutions Architects who deploy open-weight models with startups every week. This isn't slides about what's theoretically possible — it's what we ship in production. This event is a part of #LATechWeek—a week of events hosted by VCs and startups to bring together the tech ecosystem. Learn more at [www.tech-week.com](http://www.tech-week.com).

Guest List

253 on the list
>.<
:◗
-‿-

Restricted Access

Verify your phone number to view event details and activity
Sign in
Partiful logo
Home
Explore
Create
Send a card
Log in
Partiful logo
Explore eventsCreate a free event
HelpBlogCareersAboutGet the app