Back to all jobs
T

Staff/Sr. ML Infrastructure / Platform Engineer

Trendmicro

Taipei · tw 11h ago

Job description

Join Trend ‧ Join New Generation 趨勢科技 - 全球雲端資安領航者 / 全亞洲最大軟體公司 / 企業版圖橫跨五大洲 / 趨勢全球研發基地在台灣 =============================================================== About the Role We are building a production-grade, GPU-accelerated LLM serving platform that powers multiple AI products at enterprise scale. You will be responsible for designing, building, and operating the infrastructure that serves large language models — from raw Kubernetes cluster management to multi-GPU inference optimization and autoscaling. Required Qualifications Model Serving & Inference Operate multi-model LLM serving infrastructure Tune autoscaling policies to balance GPU cost and latency SLAs Kubernetes & GPU Infrastructure Operate production K8s clusters with NVIDIA GPU nodes Handle GPU node lifecycle: NVIDIA driver setup Infrastructure as Code Write and maintain Terraform/Terragrunt modules for AWS/GCP cloud Package platform components and model deployments as Helm charts Manage multi-environment configurations Observability & Performance Maintain monitoring stack: Prometheus , Grafana , Build dashboards for GPU utilization, KV cache occupancy, TTFT/ITL latency, and cost per token Set up alerting for SLA violations and OOM events Bonus Skills These are not required, but candidates with these skills will stand out. LoRA / PEFT fine-tuning workflows MLflow for experiment tracking, model registry, and automated adapter deployment Experience building LoRA adapter CI/CD pipelines (training → registry → serving) Experience with alternative inference frameworks such as SGLang or NVIDIA NIM , including deep Parameter Tuning for Continuous Batching , KV Cache management, and Speculative Decoding . =============================================================== 連結智慧 守護世界 --- Connected Intelligence for Securing a Connected World

Similar open jobs