Join the Network
← Inferact
Inferact · Hiring

Member of Technical Staff, Production Site Reliability Engineer

San FranciscoOn SiteFull Time$200K – $400K • Offers Equity

Overview

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for a Site Reliability Engineer to make Inferact's vLLM-powered inference services dependable in production. You'll work alongside the engineers building the platform, bringing a reliability perspective to how systems are designed, shipped, and operated. This is a hands-on engineering role for someone who thinks about failure before launch, simplifies operations, and writes software that makes reliable inference possible at scale.

You'll help own reliability across the production lifecycle, from release safety and observability to incident response and recovery. You'll define SLOs around availability, latency, and successful inference requests; build monitoring that helps engineers pinpoint failures; and automate recurring operational work. When incidents occur, you'll drive mitigation, coordinate escalation, and lead post-mortems that result in concrete prevention work. You'll also partner with engineering on capacity planning and operational readiness so the platform can grow without becoming harder to operate.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

Interested in This Role?

Apply at Inferact

You'll head to Inferact's own careers page.