Start Your Search Here

Job Search

Nebul

The Netherlands / Global

Infrastructure Engineer

Job Description

Direct message the job poster from Nebul

People Power Business! Mastering Talent, Recruitment, and Culture Integration.

Join Nebul

Nebul is a leader in sovereign‑hybrid cloud solutions, combining the security of private cloud infrastructure with the scalability of global hyperscalers. Rooted in European values of privacy, security, and compliance, Nebul enables businesses to harness AI with confidence. Our infrastructure powers large‑scale AI workloads, simulations, and real‑time analytics. If working on high‑performance infrastructure excites you, this is your opportunity.

What You’ll Be Doing

As an

Infra Engineer – GPU Datacenter & Kubernetes , you’ll be central to building and operating Nebul’s performance‑critical infrastructure for AI. You will:

Deploy, maintain, and scale Kubernetes clusters optimized for GPU workloads

Integrate and manage NVIDIA GPU Operators, plugins, MIG configurations, and GPU scheduling logic

Automate infrastructure (compute, storage, networking) with Terraform, Ansible, Helm or equivalent tools

Optimize GPU resource utilization, minimize fragmentation, and ensure high throughput

Instrument clusters with observability tooling (Prometheus, DCGM, Grafana, OpenTelemetry)

Ensure secure multi‑tenant usage, RBAC, network policies, and isolation between workloads

Collaborate with AI/ML, product, and security teams to understand workload needs and align infrastructure strategy

Evolve architecture over time and contribute to platform roadmap and design decisions

Support performance tuning, incident response, capacity planning, and upgrades

Optionally, mentor others or lead small infrastructure projects or pods as the platform matures

Key Responsibilities

Architect and operate

GPU‑accelerated Kubernetes clusters

with high availability and performance

Build and maintain custom controllers, operators, or scheduling extensions to support NVIDIA features

Enforce multi‑tenant security and resource isolation via RBAC, namespaces, network policies, and policy engines

Monitor GPU & cluster health; build dashboards, alerts, telemetry, and feedback loops

Tune systems, detect bottlenecks, and iterate on GPU, networking, storage performance

Drive capacity planning and cost efficiency

Evaluate and integrate new GPU and infrastructure technologies

Document architecture, operational playbooks, and best practices

What You Bring

Solid experience in running Kubernetes clusters in production, especially for GPU workloads

Deep familiarity with NVIDIA GPU integration: GPU Operator, device plugins, MIG, scheduling logic

Proficiency with Linux, networking, container runtimes, and GPU toolkits

Experience in monitoring, telemetry, and observability stacks

Understanding of AI/ML or HPC workload patterns and how they stress infrastructure

Good programming/scripting skills in Go or Python (for tooling, controllers)

Ownership mindset, reliability focus, strong communication, ability to work across functions

Bonus Points If You Have

Kernel or driver‑level GPU/compute knowledge

Experience with schedulers like Slurm, Volcano, or custom scheduling plugins

Contributions to open‑source infrastructure or GPU‑Kubernetes ecosystem

Experience with hybrid‑cloud or multi‑cloud GPU orchestration

Exposure to storage optimization for AI workloads (NVMe, Lustre, Ceph, NVMf)

Eligibility & Application Information

We welcome non‑native Dutch speakers. To apply:

Valid work permit in the Netherlands

Reside in the Netherlands and be able to travel to the office near The Hague

Nebul does

not

offer relocation assistance or sponsorship.

Referrals increase your chances of interviewing at Nebul by 2x.

Mid‑Senior level | Full‑time | Data Infrastructure and Analytics, IT System Custom Software Development

Leiden, South Holland, Netherlands

#J-18808-Ljbffr

Toepassen Now

Similar Opportunities

View all jobs