Sr. ML Infrastructure Engineer
Company Description
Today, when you go to your doctor and get referred to a specialist, your doctor sends out a referral and tells you, “They’ll be in touch soon.” So you wait. And wait. Sometimes days, weeks, or even months. Why? Because too often providers are overwhelmed with the painstakingly tedious work required to get paid by insurance companies. Powered by proprietary models, Tennr handles the complex paperwork that gets patients through the door and providers paid, helping operators get patients the right care, at the right time, in the right setting.
Role Description
We're looking for a founding Sr. ML Infrastructure Engineer with a strong background in distributed systems, cloud architecture, and building systems that scale. In this role, you'll own the infrastructure that powers Tennr's AI-driven healthcare platform - the training and inference pipelines, data systems, and deployment tooling that let our models handle growing traffic and an expanding product surface.
Our ML team builds in-house, proprietary VLMs, LLMs, and other models purpose-built for hard problems in healthcare. You don't need deep ML experience to thrive here - what matters is a strong cloud infrastructure foundation and the interest to grow into the ML side. If you think in systems, care about reliability, and want to expand into ML infrastructure, this is a rare chance to build foundational systems from the ground up.
Responsibilities
Architect, build, and scale the infrastructure behind our ML training, inference, and data pipelines.
Design resilient systems for model deployment, evaluation, and monitoring that stay reliable as traffic grows.
Own observability across the stack-logging, metrics, tracing, and alerting.
Troubleshoot production issues and continuously improve performance and efficiency.
Collaborate with ML engineers, backend engineers, and cross-functional teams to integrate models cleanly with data pipelines and products.
Candidate Qualifications
4+ years building and scaling infrastructure in production-distributed systems, cloud platforms, or data engineering.
Strong backend software engineering fundamentals, with proficiency in Python and TypeScript.
Hands-on experience with AWS and PostgreSQL.
Solid grasp of observability, reliability, and production incident response.
Comfortable with ambiguity and high ownership; you move fast and drive projects from idea to production in a startup environment.
Interested in growing into ML infrastructure - prior ML experience is not required.
Nice to have: exposure to ML frameworks (PyTorch, TensorFlow), inference tooling (vLLM, TensorRT, Triton), or GPU orchestration and optimization.
Why Tennr?
Drive Impact: one of our company values is Cowboy, meaning you set the pace. You won’t just talk about things, you’ll get them done. And feel the impact.
Develop Operational Expertise: learn the inner workings of scaling systems, tools, and infrastructure
Innovate with Purpose: we’re not just doing this for fun (although we do have a lot of fun). At Tennr, you’ll join a high-caliber team maniacally focused on reducing patient delays across the U.S. healthcare system.
Build Relationships: collaborate and connect with like-minded, driven individuals in our Hudson Square office 4 days/week
Free lunch! Plus a pantry full of snacks.
Benefits
Beautiful new office at 345 Hudson Street
Unlimited PTO
100% paid employee health benefit options
Employer-funded 401(k) match
Competitive parental leave