25.08.2026
Bewerbungsfrist : 20.09.2026
The Cluster of Excellence "Machine Learning - New Perspectives for Science" together with the Tübingen AI Center at the University of Tübingen offers a position as
HPC Systems Engineer – Security Architecture & Cluster Administration
(m/f/d, E13 TV-L, 100%)
The position is available in the team of the Machine Learning Science Cloud and runs until 31st December 2032.
About Us
The ML Cloud team designs, runs and safeguards high-performance computing infrastructure tailored to machine learning workloads. Our state of the art systems are distributed across four data centers for the Tübingen AI Center, the Cluster of Excellence ’Machine Learning: New Perspectives for Science’ and the Hertie Institute for AI in Brain Health. Every day, researchers unleash thousands of compute jobs on our systems, training frontier-scale neural networks and running experiments.
The Opportunity
Traditionally, HPC environments prioritized only performance and open access, with security as an afterthought. As our clusters grow and handle increasingly sensitive research data, security becomes mission-critical. We are looking for an HPC System Engineer who combines hands-on cluster operations with a security mindset. You will actively shape the security architecture of a live, production HPC / ML environment, keep our systems hardened, and build the processes and infrastructure that protect researchers, their data and their compute against a fast-evolving threat landscape.
Your Tasks
- Design and operate our HPC clusters across four data centers, including scheduler (SLURM), parallel filesystems, networks and accelerators, ensuring high availability and throughput for research workloads
- Conceive and establish the security architecture of the Machine Learning Science Cloud and harden the HPC environment
- Evolve the automated provisioning and configuration of heterogeneous compute, storage and network nodes (e.g. image-based provisioning, node lifecycle)
- Run patch and vulnerability management - risk assessment across heterogeneous systems
- Build and operate logging, monitoring and intrusion detection, and integrate HPC telemetry into both operational dashboards and incident-response workflows
- Lead incident response for the clusters: detection, containment, forensic support and post-incident review
- Automate operations and security policy as code (Ansible/IaC)
- Advice researchers on efficient, secure cluster usage (job scheduling, data handling, access workflows) and derive requirements for our further roadmap from their scientific workloads
Your Profile
- Masters degree in Computer Science or a related field
- In-depth IT security knowledge: system hardening, network security, applied cryptography and IAM - with the ability to derive architectural decisions from a threat model, not only to apply given baselines
- Solid Linux experience in production environments (RHEL/Almalinux/Ubuntu)
- Hands-on HPC background: Slurm, parallel file systems (Weka, Lustre, Ceph), GPU workloads and high-speed networks (InfiniBand, 400G Ethernet)
- Experience with virtualization for management-plane and infrastructure services (Proxmox)
- Strong scripting and automation skills (Bash, Python, Ansible) and experience with configuration management / Infrastructure-as-Code
- A plus: security frameworks (ISO 27001, BSI Grundschutz), container security (Apptainer/Singularity/Docker) or offensive-security fundamentals
- Independent, structured working style and good communication in English; German is a plus
- A collaborative, user-facing mindset – comfortable supporting and advising researchers and translating their needs into platform design
What We Offer
- Technically deep, architecturally open work that directly enables cutting-edge machine-learning research
- Flexible working hours and the option to work partially from home
- Working in an English-speaking, international team of HPC experts
- A small, senior team with flat hierarchy where responsibility is split by domain
- Ownership of a technical domain in a production environment of real scale
- Professional development, conference attendance and real influence on our roadmap
Interested? We'd love to hear from you!
Please submit your application, including CV and relevant transcripts by September 20, 2026 via our application portal careers.tuebingen.ai. For questions, please contact Simon Kreuzer at
simon.kreuzerspam [email protected] . The position is available from now on. Hiring is done by the Central Administration of the University of Tübingen. Severely disabled persons will be given preferential consideration if equally qualified. The University of Tübingen is committed to equity and diversity and actively promotes equal opportunities. The position is divisible.