Start Your Search Here

Job Search

EPAM Systems

The Netherlands / Global

MLOps Engineer

  • Hybrid

Job Description

We're looking for an MLOps Engineer to join our team in Amsterdam or Rijswijk, Netherlands, in a hybrid working mode. In this role, you will build, deploy and maintain production-ready machine learning solutions with a strong focus on MLOps practices including CI/CD, model serving, monitoring and robust cloud infrastructure. You will also contribute to extending these capabilities toward LLMOps and agentic AI workflows, enabling areas such as LLM applications, RAG pipelines, model evaluation and observability for enterprise environments. The position involves close collaboration with engineering and data science teams as well as advisory engagement with clients on best practices in AI infrastructure and operational scalability. This is an opportunity to deliver impactful AI capabilities while working at the intersection of modern AI and enterprise systems.

Responsibilities

Build and maintain platform components for ML model training, deployment, serving and monitoring

Develop and optimize CI/CD pipelines for machine learning workflows

Implement and support model lifecycle management, including registries and observability tooling

Design and manage scalable, secure deployments using containerization and Kubernetes

Enable secure, reusable and automated workflows to enhance ML developer productivity

Extend platform capabilities to support LLMOps, RAG and agentic AI workloads

Collaborate with engineering teams to improve reliability, automation and operational maturity

Apply governance and compliance standards across AI operations

Participate in presales and client-facing sessions to translate requirements into scalable solutions

Advocate cloud best practices for reliability, scalability and cost optimization

Requirements

Bachelor’s or Master’s degree in Computer Science, Engineering or related discipline

Experience in delivering machine learning or MLOps systems into production environments

Proficiency in Python for building services, APIs, scripts and CI/CD automation

Working knowledge of modern MLOps stacks including experiment tracking and artifact management

Hands-on experience with orchestration tools (e.g., Kubeflow, Apache Airflow, Metaflow or Prefect)

Demonstrated skills with Docker, Kubernetes and distributed deployments

Practical knowledge of Infrastructure-as-Code (Terraform) and a major cloud provider (AWS, Azure or GCP)

Familiarity with ML model serving, scaling and monitoring frameworks in production

Strong communications skills to convey technical decisions and engage with clients effectively

Nice to have Background deploying Generative AI solutions, LLM inference pipelines or agentic AI systems

Experience with feature stores, vector databases and retrieval-augmented generation approaches

Knowledge of AI governance, security and compliance for regulated sectors

Familiarity with advanced observability and tracing solutions, such as OpenTelemetry or Langfuse

Consulting or enterprise architecture experience in large-scale AI programs

Understanding of FinOps strategies for managing GPU/CPU costs in cloud environments

Certifications in cloud technologies (AWS, Azure, GCP) or Kubernetes (CKA/CKAD)

Expertise in securing and operationalizing ML/LLM/agent-based systems for enterprise readiness

#J-18808-Ljbffr

Toepassen Now

Similar Opportunities

View all jobs