About this role
Accountabilities:: Architect, implement, and optimize data platforms and pipelines designed for LLMs, RAG, and advanced AI agentic systems at Exabyte scale. Drive the adoption and deployment of agentic workflows and agent-harnessing techniques to support autonomous, data-driven capabilities. Design highly scalable, fault-tolerant, secure, and cost-effective data solutions that enable rapid iteration without compromising engineering quality. Develop production-ready code with strong attention to performance, maintainability, testing, and operational reliability. Provide technical leadership in data modeling, normalization, semantic cataloging, and data architecture for AI and machine learning workloads. Establish MLOps and DataOps best practices for LLM platforms, including monitoring, observability, automated recovery, and service reliability. Own the end-to-end lifecycle of critical data services, including development, testing, deployment, monitoring, and continuous optimization. Collaborate with data scientists, product managers, and engineering teams to transform research prototypes into robust, production-ready services. Lead technical workshops, design reviews, and knowledge-sharing initiatives while mentoring engineers and strengthening organizational expertise in AI platform technologies. Champion DevSecOps practices and engineering standards across large-scale distributed data environments. Identify opportunities to improve platform performance, reliability, scalability, and developer productivity through new technologies and engineering practices. Requirements Master’s degree or PhD in Computer Science, Data Engineering, or a related STEM discipline, or equivalent practical experience. 10+ years of progressive experience in Data Engineering or Platform Engineering, including at least 3 years architecting and building AI/ML or Data Science platforms at massive scale. 3+ years of experience in a Principal or Staff-level engineering capacity, with demonstrated technical leadership and mentorship experience. Hands-on expertise with LLM engineering, including fine-tuning, prompt engineering, deployment, RAG, and agentic workflow development. Proven experience designing and delivering large-scale distributed systems, including sharding, partitioning, concurrency, and fault-tolerant architectures. Expert-level proficiency in Python or JVM-based technologies, with a strong ability to write clean, performant, maintainable, and well-tested production code. Deep experience with distributed data processing frameworks such as Spark, Dask, or Flink. Strong knowledge of cloud platforms such as AWS, GCP, or OCI and their associated data services. Expertise with containerization and orchestration technologies including Docker and Kubernetes. Experience with messaging and streaming technologies such as Kafka or Pulsar. Familiarity with data warehousing and orchestration platforms such as Snowflake, BigQuery, Airflow, and Kubeflow. Experience with MLOps technologies such as MLflow, SageMaker, or Vertex AI. Familiarity with agentic AI frameworks such as LangChain or LlamaIndex. Strong understanding of engineering practices including peer code reviews, resilient architecture, comprehensive testing, and secure development methodologies. Demonstrated ability to use AI technologies to improve decision-making, automate workflows, increase efficiency, and support measurable business outcomes. Strong communication and collaboration skills, with the ability to influence technical direction and mentor engineers across teams. Direct experience deploying and managing LLMs in production is a plus. Experience in cybersecurity, intelligence, or highly regulated industries is a plus. Contributions to open-source data or AI/ML projects are a plus. Benefits CAD $210,000–$320,000 annual base salary for Canadian-based employment, plus variable/incentive compensation, equity, and benefits. Flexible remote work opportunities. Comprehensive health and wellness programs supporting physical and mental wellbeing. Competitive vacation and holiday programs to support time away and recharge. Paid parental and adoption leave. Professional development and continuous learning opportunities at all career levels. Employee networks, geographic communities, and volunteer opportunities to build professional connections. Opportunities to work on large-scale AI, data engineering, and cybersecurity technologies. Equity opportunities as part of the overall compensation package. Retirement and financial benefits as applicable. Inclusive workplace practices and support for employees with disabilities. Canadian employment requires legal entitlement to work in Canada and may include applicable background checks. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1
- LOCATION
- Remote · US, United States
- WORK MODE
- Remote
- JOB TYPE
- Full-time
- POSTED
- Sep 26, 2026
Source: Jobgether Careers