Job Description
Role: HPC (High-Performance Computing) Consultant Location: Remote Duration: Fulltime Interview Type: Video Must Have: Kubernetes + Slurm + NVIDIA GPU + AI/ML Infrastructure + Terraform + Python + AWS (EKS/FSx Lustre) + HPC Storage + Monitoring/SRE We are looking for engineers with expertise across Kubernetes, cloud infrastructure, HPC platforms, GPU computing, Terraform, and automation. Depending on experience, candidates may be considered for Platform Engineer, Kubernetes Engineer, HPC Engineer, Cloud Infrastructure Engineer, DevOps Engineer, AI Infrastructure Engineer, or Site Reliability Engineer (SRE) roles. Key Responsibilities Design, deploy, operate, and support large-scale Kubernetes platforms across AWS, GCP, CoreWeave, OCI, and other cloud environments. Manage Kubernetes cluster lifecycle activities including provisioning, scaling, node pool management, upgrades, troubleshooting, and performance optimization. Support AI/ML and HPC workloads, including GPU-enabled compute infrastructure for training and inference environments. Provision and automate cloud and infrastructure resources using Terraform and CI/CD pipelines. Troubleshoot Kubernetes scheduling, networking, storage, and platform reliability issues. Implement monitoring, observability, alerting, SLIs/SLOs, and incident response processes. Collaborate with Networking, Security, Storage, AI/ML, Data Engineering, and Application teams. Develop automation and operational tooling using Python and cloud-native technologies. Participate in production support, root cause analysis, capacity planning, and platform optimization initiatives. Required Skills Kubernetes & Container Platforms Kubernetes (EKS, GKE, AKS, OpenShift, CoreWeave CKS) Cluster Lifecycle Management Node Pool Management Scheduler Troubleshooting CNI Troubleshooting Networking Policies RBAC Helm Autoscaling Rolling Upgrades Cloud Infrastructure AWS (EC2, S3, IAM, VPC, EKS, EFS, FSx for Lustre) Google Cloud Platform (GCP) OCI (Oracle Cloud Infrastructure) Azure (Preferred) Multi-Cloud Infrastructure Infrastructure as Code & Automation Terraform Infrastructure as Code (IaC) CI/CD Pipelines GitHub Actions / Jenkins / GitLab CI Ansible (Preferred) Programming & Scripting Python Bash/Shell Scripting Automation Development REST API Integration HPC & GPU Infrastructure (Preferred) High-Performance Computing (HPC) Slurm NVIDIA GPU Platforms GPU Scheduling CUDA Distributed Computing AI/ML Infrastructure Monitoring & Reliability Prometheus Grafana Datadog Splunk Cloud Monitoring SLI/SLO Management Incident Response Root Cause Analysis (RCA) Preferred Experience Experience with any of the following is highly desirable: AI/ML Infrastructure Platforms Kubeflow KServe Ray MLflow vLLM Vector Databases Distributed Training Platforms CoreWeave AWS ParallelCluster FSx for Lustre Lustre WekaFS InfiniBand RDMA AI Training & Inference Workloads The pay range for this role is USD $180,000- $200,000 per annum including any bonuses or variable pay. Tech Mahindra also offers benefits like medical, vision, dental, life, disability insurance and paid time off (including holidays, parental leave, and sick leave, as required by law). Ask our recruiters for more details on our Benefits package. The exact offer terms will depend on the skill level, educational qualifications, experience and location of the candidate. AI tools may assist in the recruitment process; however, all hiring decisions are made by the recruitment team based on a comprehensive evaluation of candidates. "Tech Mahindra is an Equal Employment Opportunity employer. We promote and support a diverse workforce at all levels of the company. All qualified applicants will receive consideration for employment without regard to race, religion, color, sex, age, national origin, or disability. All applicants will be evaluated solely on the basis of their ability, competence, and performance of the essential functions of their positions with or without reasonable accommodations. Reasonable accommodations also are available in the hiring process for applicants with disabilities. Candidates can request a reasonable accommodation by contacting the company ADA Coordinator at ."