Dein persönlicher KI-Karriere-Agent
Senior Site Reliability Engineer, DGX Cloud (m/w/x)
Operating GPU workloads on Kubernetes across AWS, GCP, Azure, OCI, and private clouds. Expert-level Kubernetes administration and infrastructure automation skills required. Consulting on service launches and tool development.
Deine Match-Analyse
Warum dieser Job zu dir passt
Mögliche Lücken
Tipps für deine Bewerbung
Anforderungen
- BS in Computer Science or related technical field, or equivalent experience
- 10+ years of experience operating production services
- Expert-level Kubernetes administration, containerization, and microservices architecture
- Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet)
- Proficiency in at least one high-level programming language (e.g., Python, Go)
- In-depth Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards
- Proficient SRE principles, encompassing SLOs, SLIs, error budgets, and incident handling
- Experience building and operating observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.
- Operating GPU-accelerated clusters with KubeVirt in production
- Applying generative-AI techniques to reduce operational toil
- Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions
Aufgaben
- Build, implement, and support large-scale Kubernetes clusters for performance, monitoring, and alerting
- Define SLOs/SLIs, monitor error budgets, and streamline reporting
- Support service launches through consulting, tool development, capacity management, and reviews
- Maintain live services by measuring and monitoring system health
- Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds
- Scale systems sustainably through automation and evolve them for reliability and velocity
- Lead triage and root-cause analysis of high-severity incidents
- Practice balanced incident response and blameless postmortems
- Participate in on-call rotation for production services
Berufserfahrung
Ausbildung
Sprachen
Tools & Technologien
Von Nejo automatisch aufbereitet
Nejo hat diesen Job automatisch von der Website des Unternehmens CH01 NVIDIA Switzerland AG erfasst und die Informationen auf Nejo mit Hilfe von KI für dich aufbereitet. Trotz sorgfältiger Analyse können einzelne Informationen unvollständig oder ungenau sein. Bitte prüfe immer alle Angaben in der Originalanzeige! Inhalte und Urheberrechte der Originalanzeige liegen beim ausschreibenden Unternehmen.
Zur Originalanzeige bei CH01 NVIDIA Switzerland AGÜber das Unternehmen
NVIDIA has been defining computer graphics, PC gaming, and accelerated computing for more than 25 years.
Nejo bewertet ihn, und bringt ihn danach gemeinsam mit dir in Bestform.
Noch nicht perfekt?
- Open Systems AGSenior Site Reliability Engineer (m/w/x)Vollzeitnur vor OrtSeniorZürich
- NVIDIA Switzerland AGSenior DevOps Engineer (m/w/x)Vollzeitnur vor OrtSeniorZürich
- Proton Technologies AG.Site Reliability Engineer (m/w/x)Vollzeitnur vor OrtKeine AngabeGenf, Zürich
- NVIDIASenior GPU Networking Architect (m/w/x)Vollzeitnur vor OrtSeniorZürich
- NVIDIA Switzerland AGSenior HPC and AI Network Software Architect (m/w/x)Vollzeitnur vor OrtSeniorZürich