← Back to Connect Jobs
CONNECT VERIFIED

Staff Machine Learning Systems Engineer

Jobgether

US · remote · Full-time

Jobgether

Accountabilities:: Design and implement end-to-end model lifecycle patterns and MLOps capabilities covering data preparation, model management, experiment tracking, deployment, and related workflows. Lead zero-to-one development of graph machine learning infrastructure and codebases that abstract common patterns and enable scalable model development. Collaborate with machine learning engineers to optimize model performance, training duration, memory utilization, and GPU costs across large distributed environments. Optimize batch data processing and data pipelines using technologies such as Apache Beam, Apache Spark, Ray Data, and cloud data warehouse services. Architect pipelines capable of building and maintaining massive graph data structures containing billions of nodes and tens of billions of edges. Build and evolve cloud-based infrastructure supporting machine learning platforms, with an emphasis on scalability, reliability, performance, and operational efficiency. Administer and integrate MLOps tools for experiment tracking, model serving, and model registries. Develop solutions that improve the overall machine learning development lifecycle and make platform capabilities easier and more efficient for engineering teams to use. Partner with cross-functional technical teams to understand infrastructure needs, identify bottlenecks, and deliver scalable solutions. Establish and promote engineering patterns that improve platform reliability, developer productivity, model iteration, and ease of use. Provide technical leadership across complex infrastructure initiatives and help shape the long-term architecture of machine learning systems. Requirements: 8+ years of experience working in machine learning infrastructure, including model training and model deployment environments. Hands-on experience optimizing machine learning systems, including memory profiling, GPU profiling, training performance, and resource utilization. Deep experience with cloud technologies used to support ML platforms, including Google Cloud Platform services such as BigQuery and Google Cloud Storage. Experience with infrastructure-as-code tools such as Terraform. Hands-on experience administering and integrating MLOps platforms for experiment tracking, model serving, and model registries, such as MLflow or Weights & Biases. Strong proficiency in Python and familiarity with machine learning frameworks such as PyTorch and TensorFlow. Deep experience with distributed machine learning and data processing frameworks, including Ray and Kubernetes. Strong understanding of scalable, reliable, high-performance platform architecture and the machine learning development lifecycle. Ability to advocate effectively for platform users and translate their needs into intuitive, scalable infrastructure solutions. Strong organizational, communication, and collaboration skills, with the ability to work effectively across technical teams. Experience with graph databases such as Neo4j, JanusGraph, or TigerGraph is highly valued. Experience with graph neural networks and graph ML frameworks such as PyTorch Geometric or Deep Graph Library is highly valued. Ability to operate effectively in a technically complex, rapidly evolving environment and take ownership of large-scale infrastructure initiatives. Benefits: Base salary range of $230,000–$322,000 USD, with final compensation determined by factors including skills, experience, and relevant credentials. Eligibility for equity in the form of restricted stock units. Potential eligibility for commission depending on the position offered. Medical, dental, and vision insurance for U.S.-based employees. 401(k) program with employer matching. Generous paid vacation and time-off benefits. Paid parental leave. Remote work opportunity within the United States. Opportunities to work on large-scale machine learning infrastructure and technically challenging systems. Opportunities for significant technical ownership and professional growth. For select roles and locations, interviews may be recorded, transcribed, and summarized using AI; candidates can opt out of these processes before scheduled interviews. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1

PythonAWSKubernetesMachine LearningAITerraformGitUI
FREE MATCHED JOB ALERTS

Get jobs like this without searching manually.

Tell Connect what you want once. We will use your preferences to surface matching opportunities and invite you into your free career workspace.