Lead impactful DevOps projects optimizing GPU-accelerated container platforms. Collaborate with dynamic teams to drive innovation in high-performance computing. Enjoy comprehensive benefits and professional growth opportunities.
Devops
in Information Technology PermanentJob Detail
Job Description
DevOps Overview
- The DevOps role focuses on designing and optimizing GPU-accelerated container platforms for high-performance workloads across hybrid or on-prem environments.
- Collaborate with HPC, ML, and DevOps teams to ensure multi-tenant, high-throughput cluster performance.
- Drive observability and telemetry integrations using Prometheus, Grafana, DCGM Exporter, and OpenTelemetry.
- Implement secure multi-user and multi-namespace GPU isolation with RBAC and policy enforcement.
- Maintain CI/CD pipelines for Kubernetes infrastructure using GitOps tools like ArgoCD and FluxCD.
- Contribute to infrastructure-as-code practices using Terraform, Helm, and Kustomize.
- Participate in performance tuning, incident response, and production readiness reviews.
- Enjoy comprehensive benefits, including employer-paid medical coverage, 401(k) matching, and generous paid time off.
DevOps Key Responsibilities & Duties
- Architect and operate Kubernetes clusters optimized for GPU workloads, leveraging NVIDIA GPU Operator and DCGM.
- Develop and maintain custom Kubernetes operators and controllers to automate infrastructure services.
- Integrate NVIDIA device plugins, Multi-Instance GPU (MIG), and GPU sharing features into scheduling layers.
- Optimize GPU utilization and job placement through scheduler extensions like kube-scheduler plugins.
- Collaborate with teams to ensure high-performance cluster operations and multi-tenant efficiency.
- Implement telemetry integrations using tools like Prometheus and OpenTelemetry.
- Ensure secure GPU isolation with RBAC and policy enforcement mechanisms.
- Maintain CI/CD pipelines for Kubernetes infrastructure using GitOps methodologies.
DevOps Job Requirements
- Bachelor of Science (BS) degree in a relevant field is required.
- Minimum of 10 years of experience in Kubernetes and NVIDIA GPU ecosystems.
- Proficiency in Go or Python for Kubernetes operator and controller development.
- Deep understanding of Kubernetes internals, including CRDs, RBAC, and scheduler extensions.
- Experience with GPU-intensive workloads such as LLMs and scientific computing.
- Hands-on expertise with Helm, Kustomize, and GitOps workflows.
- Familiarity with CNI plugins, NVIDIA CNI, and Multus.
- Legally authorized to work in the United States without sponsorship.
- ShareAustin:
