Huawei Canada

Research Engineer - MLLM Serving Optimization

Check with seller / month
Vancouver, British Columbia, Canada IT Engineer & Developer Active
Actively Hiring Vancouver Full Time
Advertisement
Verified Listing
Direct Apply — No Agent
Your Data is Safe
Trusted by 5 Lakh+ Jobseekers

Job at a Glance

Category
IT Engineer & Developer
Location
Vancouver, British Columbia, Canada
Salary
Check with seller
Job Type
Full Time
Company
Huawei Canada
Status
Open & Active

Job Description

Job description
Huawei Canada has an immediate permanent opening for a researcher.

About the team:

The Intelligent Cloud Infrastructure Lab aims to innovate technologies, algorithms, systems, and platforms for next-generation cloud infrastructure. The lab addresses scalability, performance, and resource utilization challenges in existing cloud services while preparing for future challenges with appropriate technologies and architectures. Additionally, the lab aims to understand industry dynamics and technology trends to create a robust ecosystem.

About the job:
• Design, implement, and optimize a high-performance serving platform for MLLMs.
• Integrate SOTA open-source serving frameworks such as vLLM, sglang, or lmdeploy.
• Develop techniques for efficient resource utilization and low-latency inference for MLLMs in serverless environments.
• Optimize memory usage, scalability, and throughput of the serving platform.
• Conduct experiments to evaluate and benchmark MLLM serving performance..
• Contribute novel ideas to improve serving efficiency and publish findings when applicable.

The base salary for this position ranges from $100,000 to $190,000 depending on education, experience and demonstrated expertise

Job requirements

About the ideal candidate:
• Bachelor’s degree or higher in Computer Science, Electrical and Computer Engineering (ECE), or a related field.
• Experience with one or more SOTA LLM serving frameworks such as vLLM, sglang, or lmdeploy.
• Strong proficiency in PyTorch.
• Familiarity with distributed systems, serverless architectures, and cloud computing platforms.
• Experience with inference optimization for large-scale AI models.
• Familiarity with multimodal architectures and serving requirements.
• Previous experience in deploying AI platforms on cloud services.
Ready to take the next step?

Don't wait — new applications are being reviewed daily.

Login & Apply Free
Job Safety Alert Real jobs on Jobsiya are always free. Never pay for an interview and never share bank or OTP details. Report this job →