Jobiglo

No results.

Senior/Staff AI Model Optimization Architect – Hsinchu/Taipei

Qualcomm · Taipei

New
Senior 🇬🇧 English
PyTorch torch.compile TorchDynamo ONNX Python Triton Transformer architectures Distributed systems ML accelerators

Job description

About the role

Qualcomm's Cloud AI team is looking for a Staff‑level AI Model Optimization Architect to lead end‑to‑end transformation and optimization of large language, vision, diffusion and multimodal models on Qualcomm inference accelerators. The role works closely with compiler, performance and accuracy teams to deliver efficient inference across a range of workloads.

Key responsibilities

  • Architect and deliver model optimization strategies that transform PyTorch models for efficient inference on Qualcomm accelerators.
  • Drive graph capture and deployment using PyTorch, ONNX, and torch.compile, including model rewrites and graph‑level transformations.
  • Design and implement fusion kernels using DSL‑based approaches (e.g., Triton) to enable fused operations and performance‑critical algorithmic rewrites.
  • Partner with compiler, performance and accuracy teams to co‑design lowering strategies, kernel fusion, layout decisions and runtime integration.
  • Profile and optimize LLM/VLM/diffusion inference for throughput and latency across batch sizes, sequence lengths and serving modes.

Required profile

  • Expert level expertise in PyTorch and inference‑focused model optimization with strong Python engineering skills.
  • Deep understanding of transformer architectures, attention mechanisms, MoEs and performance trade‑offs.
  • Strong foundation in computer architecture, ML accelerators and distributed systems.
  • Proven ability to lead cross‑functional technical efforts and influence design decisions.
  • MS in Computer Science, Machine Learning, Computer Engineering, Electrical Engineering or equivalent; or higher degree with relevant experience.

Required skills

  • PyTorch, torch.compile, TorchDynamo.
  • ONNX and graph capture workflows.
  • Python programming.
  • DSL‑based kernel development (e.g., Triton).
  • Transformer model optimization, KV‑cache management, continuous batching.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Qualcomm.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

Apply now →

By continuing, you accept our terms of use.

Already have an account? Login

Une question sur cette offre ?

Posez-la ici : vous recevrez le récapitulatif de l'offre par e-mail, tout de suite.

💬 Chat with us on Telegram

Published 18小时前

Expires 1个月后

5 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Qualcomm

Taipei