About the role
Job Description
Key responsibilities
* Build and operate Kubernetes environments that host AI engineering tools, internal model gateways, retrieval components, workflow services, CI/CD runners, and documentation services. * Implement GitOps and Infrastructure as Code patterns for reproducible provisioning, configuration, policy enforcement, platform upgrades, and disaster recovery readiness. * Manage private registries, package mirrors, secrets, identity integration, network segmentation, storage classes, backup routines, and controlled connectivity models. * Provide observability for engineering workloads, including metrics, logs, traces, GPU and CPU utilization, service health, cost signals, and operational runbooks. * Work with software, security, and architecture teams to ensure the platform supports AI-assisted SDLC workflows without creating uncontrolled data exposure or audit gaps.
Examples of market tools, models, and platform components expected
* Platform tooling such as Kubernetes, Helm, Terraform, Ansible, ArgoCD, Crossplane, GitLab runners, Jenkins agents, private registries, and internal package mirrors. * AI platform components such as vLLM, Ollama, OpenAI-compatible gateways, Qdrant or similar vector stores, Open WebUI, Continue-compatible endpoints, and workflow services. * Observability and operations stacks such as Prometheus, Grafana, Loki, OpenTelemetry, ELK/OpenSearch, Alertmanager, SRE runbooks, and incident management tooling. * Security and governance components such as Vault, Keycloak, network policies, RBAC, admission controls, image scanning, SBOM tooling, and audit logging. * Infrastructure awareness covering GPU-backed nodes, CPU-only fallback, storage performance, network isolation, proxy patterns, on-premise environments, and dedicated landing zones.
Qualifications
* 5+ years in SRE, platform engineering, DevOps, cloud infrastructure, or operations roles with strong Kubernetes and Linux expertise. * Proven experience building and operating production-grade engineering platforms with GitOps, Infrastructure as Code, observability, and operational runbooks. * Hands-on skills in Terraform, Ansible, Helm, Python or shell scripting, CI/CD runners, private registries, and secure configuration management. * Good understanding of networking, storage, secrets, access control, monitoring, backup, disaster recovery, and operational hardening in high-security environments. * Comfortable supporting AI-enabled engineering workloads in sovereignty-driven contexts where isolation, controlled data handling, reliability, and auditability are mandatory.
Additional Information
What do we offer you?
Work environment & flexibility
* International, dynamic and collaborative environment. * T-Social: social initiatives (sports, community, health, ...). * Hybrid work model (remote/on-site). * Flexible working hours.
Growth & development
* Customized training: access to Coursera to learn whatever you want, whenever you want. * Weekly language classes (English & German). * International Mentoring Sessions & Experience Days.
Compensation & benefits
* Flexible compensation plan (health insurance, meal vouchers, childcare, transport). * Telemedicine. * Life and accident insurance. * Social fund.
Wellbeing & time off
* 26+ working days of vacation per year. * Free access to specialist services (medical, legal, wellness). * 100% salary coverage during medical leave.
And many more advantages of being part of T-Systems!
If you are looking for a new challenge, do not hesitate to send us your CV! Please send CV in English. Join our team!
T-Systems Iberia will only process the CVs of candidates who meet the requirements specified for each offer.