System Team Lead – Edge Infrastructure & Observability
Berry AI · District de Neihu
Job description
About the role
Berry AI is looking for a System Team Lead to own and scale the hybrid cloud‑edge infrastructure that powers its AI‑driven operations platform for QSR restaurants. You will lead a growing team, stay hands‑on with the technology, and ensure the fleet of in‑store edge servers and cameras runs reliably at scale.
Key responsibilities
- Lead, mentor, and hire for the system team while maintaining a hands‑on engineering role.
- Own the on‑prem edge fleet of thousands of in‑store servers and cameras, connected via resilient VPN/mesh networks.
- Build internal tools, scripts, and APIs that enable self‑service installation and troubleshooting for customer‑site technicians.
- Manage the self‑hosted observability platform (Prometheus, Grafana, Mimir, Loki, VMAgent, Vector) with SLO dashboards and automated alerting.
- Drive automation using Ansible/AWX and serverless AWS services (SNS, SQS, Lambda) to keep operational load flat as the fleet grows.
- Run core infrastructure at the Taipei headquarters, including servers, networking, virtualization, Kubernetes, storage, and self‑hosted services.
- Partner with product, ML, and customer success teams to align roadmap and enable product velocity.
- Lead on‑call rotation, incident response, and turn recurring issues into durable fixes.
- Own security and compliance operations (ISO 27001, vulnerability management, code scanning, disaster‑recovery drills).
Required profile
- Proven experience leading an infrastructure, platform, or SRE team while remaining hands‑on.
- 5+ years in systems, infrastructure, or DevOps engineering with strong Linux (Ubuntu) administration.
- Deep networking knowledge (TCP/IP, DNS, VLANs, VPN, firewalls).
- Extensive experience operating large‑scale fleets.
- Fluent communication in Mandarin and English.
Required skills
- Linux (Ubuntu)
- TCP/IP, DNS, VLAN, VPN, firewall configuration
- Terraform, Pulumi
- Prometheus, Grafana, Mimir, Loki, VMAgent, Vector
- Ansible, AWX
- AWS SNS, SQS, Lambda
- Bash, Python, Go
- Kubernetes, virtualization, storage, device monitoring
- Container registry, authentication, reverse proxy
- ISO 27001, vulnerability management, code security scanning, disaster‑recovery procedures
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in 台湾.
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Berry AI
District de Neihu