Drive innovation by optimizing GPU-accelerated Kubernetes platforms for machine learning. Collaborate on impactful projects in advanced infrastructure development. Elevate your career in a dynamic, technology-driven environment.
Platform Engineer
in Information Technology PermanentJob Detail
Job Description
Overview
- Lead the development and optimization of scalable Kubernetes platforms tailored for high-performance machine learning workloads.
- Design and deploy GPU-accelerated clusters across hybrid cloud and bare-metal infrastructures to support advanced computing needs.
- Implement robust CI/CD pipelines to streamline software delivery and enhance operational efficiency.
- Develop and maintain observability systems to monitor infrastructure performance and ensure reliability.
- Collaborate with cross-functional teams to refine development workflows and optimize tooling.
- Apply infrastructure-as-code practices to ensure consistency and scalability in system configurations.
- Manage core infrastructure components, including networking, storage, and system setups.
- Ensure compliance with security standards and implement measures to mitigate risks.
Key Responsibilities & Duties
- Deploy and manage Kubernetes clusters optimized for GPU-accelerated machine learning workloads.
- Enhance system performance through proactive monitoring and optimization strategies.
- Design and operate automated CI/CD pipelines to support efficient software delivery.
- Develop observability tools to monitor system health and generate actionable insights.
- Collaborate with development teams to refine tooling and improve deployment processes.
- Implement infrastructure-as-code using tools such as Terraform, Helm, and Ansible.
- Manage networking, storage, and system configurations for high-performance environments.
- Ensure adherence to stringent security standards and protocols.
Job Requirements
- Bachelor's degree in Computer Science or a related field is required.
- Minimum of 3 years, preferably 5 years, of experience in platform engineering or DevOps.
- Proficiency in Python and Bash for scripting and automation tasks.
- Strong expertise in Kubernetes administration and GPU environments.
- Experience with CI/CD pipeline design and observability tools.
- Advanced knowledge of Linux systems and infrastructure-as-code practices.
- Ability to manage complex systems and ensure compliance with security standards.
- Preferred experience with machine learning orchestration tools and build toolchains.
- ShareAustin:
