Staff/Sr. ML Infrastructure / Platform Engineer
TrendAI · Taipei
Job description
About the role
We are building a production‑grade, GPU‑accelerated large‑language‑model (LLM) serving platform that powers multiple AI products at enterprise scale. You will design, build, and operate the infrastructure that serves large language models, from raw Kubernetes cluster management to multi‑GPU inference optimization and autoscaling.
Key responsibilities
- Design, build, and operate multi‑model LLM serving infrastructure.
- Tune autoscaling policies to balance GPU cost and latency service‑level agreements.
- Operate production Kubernetes clusters with NVIDIA GPU nodes.
- Handle GPU node lifecycle, including driver installation and updates.
- Write and maintain Terraform/Terragrunt modules for AWS and GCP environments.
- Package platform components and model deployments as Helm charts.
- Maintain a monitoring stack (Prometheus, Grafana) and create dashboards for GPU utilization, cache occupancy, latency and cost per token.
- Set up alerting for SLA violations and out‑of‑memory events.
Required profile
- Proven experience operating production Kubernetes clusters with GPU nodes.
- Strong background in infrastructure‑as‑code using Terraform/Terragrunt on AWS or GCP.
Required skills
- Kubernetes
- NVIDIA GPU
- Terraform
- Terragrunt
- AWS
- GCP
- Helm
- Prometheus
- Grafana
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in 台湾.
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Une question sur cette offre ?
Posez-la ici : vous recevrez le récapitulatif de l'offre par e-mail, tout de suite.
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
TrendAI
Taipei