EPAM Systems
The Netherlands / Global
MLOps Engineer
- Hybrid
The Netherlands / Global
We're looking for an MLOps Engineer to join our team in Amsterdam or Rijswijk, Netherlands, in a hybrid working mode. In this role, you will build, deploy and maintain production-ready machine learning solutions with a strong focus on MLOps practices including CI/CD, model serving, monitoring and robust cloud infrastructure. You will also contribute to extending these capabilities toward LLMOps and agentic AI workflows, enabling areas such as LLM applications, RAG pipelines, model evaluation and observability for enterprise environments. The position involves close collaboration with engineering and data science teams as well as advisory engagement with clients on best practices in AI infrastructure and operational scalability. This is an opportunity to deliver impactful AI capabilities while working at the intersection of modern AI and enterprise systems.
Responsibilities
Build and maintain platform components for ML model training, deployment, serving and monitoring
Develop and optimize CI/CD pipelines for machine learning workflows
Implement and support model lifecycle management, including registries and observability tooling
Design and manage scalable, secure deployments using containerization and Kubernetes
Enable secure, reusable and automated workflows to enhance ML developer productivity
Extend platform capabilities to support LLMOps, RAG and agentic AI workloads
Collaborate with engineering teams to improve reliability, automation and operational maturity
Apply governance and compliance standards across AI operations
Participate in presales and client-facing sessions to translate requirements into scalable solutions
Advocate cloud best practices for reliability, scalability and cost optimization
Requirements
Bachelor’s or Master’s degree in Computer Science, Engineering or related discipline
Experience in delivering machine learning or MLOps systems into production environments
Proficiency in Python for building services, APIs, scripts and CI/CD automation
Working knowledge of modern MLOps stacks including experiment tracking and artifact management
Hands-on experience with orchestration tools (e.g., Kubeflow, Apache Airflow, Metaflow or Prefect)
Demonstrated skills with Docker, Kubernetes and distributed deployments
Practical knowledge of Infrastructure-as-Code (Terraform) and a major cloud provider (AWS, Azure or GCP)
Familiarity with ML model serving, scaling and monitoring frameworks in production
Strong communications skills to convey technical decisions and engage with clients effectively
Nice to have Background deploying Generative AI solutions, LLM inference pipelines or agentic AI systems
Experience with feature stores, vector databases and retrieval-augmented generation approaches
Knowledge of AI governance, security and compliance for regulated sectors
Familiarity with advanced observability and tracing solutions, such as OpenTelemetry or Langfuse
Consulting or enterprise architecture experience in large-scale AI programs
Understanding of FinOps strategies for managing GPU/CPU costs in cloud environments
Certifications in cloud technologies (AWS, Azure, GCP) or Kubernetes (CKA/CKAD)
Expertise in securing and operationalizing ML/LLM/agent-based systems for enterprise readiness
#J-18808-Ljbffr
The Netherlands / Global
The Netherlands / Global
The Netherlands / Global
The Netherlands / Global
The Netherlands / Global
The Netherlands / Global