Presentation
Research Scientist-LLM System
SessionJob Postings
DescriptionWe are looking for a Research Scientist to design and optimize Huawei’s large model infrastructure, enabling highly efficient training, inference acceleration, and scalable deployment. This role focuses on building cutting-edge system architectures and engineering platforms that empower large model research and production at scale. You will work at the intersection of AI infrastructure, distributed systems, and high-performance computing, driving innovations that accelerate Huawei’s foundation model ecosystem.
Responsibilities:
1.Lead system-level optimizations for large model training, including but not limited to parallel strategies (PP/TP/CP), memory optimization, communication optimization, and fault tolerance in large-scale training clusters.
2.Drive inference acceleration for large models, including but not limited to fused operator optimization, parallel strategy optimization, memory optimization, communication optimization, and other parallel or mathematical optimization techniques.
3.Design, develop, and maintain large model training platforms and toolchains to provide simple, efficient, and scalable training systems for algorithm teams. This includes but is not limited to automated evaluation systems, job monitoring systems, and visualization tools.
4.Build and maintain large model inference and deployment platforms, including automated pipelines, one-click deployment systems, and data feedback systems, to enable rapid and cost-effective internal and external deployment of large models.
Responsibilities:
1.Lead system-level optimizations for large model training, including but not limited to parallel strategies (PP/TP/CP), memory optimization, communication optimization, and fault tolerance in large-scale training clusters.
2.Drive inference acceleration for large models, including but not limited to fused operator optimization, parallel strategy optimization, memory optimization, communication optimization, and other parallel or mathematical optimization techniques.
3.Design, develop, and maintain large model training platforms and toolchains to provide simple, efficient, and scalable training systems for algorithm teams. This includes but is not limited to automated evaluation systems, job monitoring systems, and visualization tools.
4.Build and maintain large model inference and deployment platforms, including automated pipelines, one-click deployment systems, and data feedback systems, to enable rapid and cost-effective internal and external deployment of large models.
Location
Shanghai, Shenzhen, Hong Kong, Singapore, Canada, Beijing, Hangzhou, etc.
In-Person:
Onsite
Description of Position:
PhD; Full-Time; Desired Level of Education: Master’s degree
Other Skills or Experience:
1. Master’s degree or above in Computer Science, Electrical Engineering, or a related engineering discipline, with at least 3+ years of industry experience.
2. Solid understanding of large model AI infrastructure and system design; proficient in PP/TP/CP parallel strategies.
3. Hands-on experience with large model training and inference acceleration frameworks such as PyTorch, Megatron, vLLM, and SGLang.
4. Strong knowledge of cluster-level training optimization and inference deployment strategies on GPU/NPU hardware.
·
·
2025-08-01
Event Type
Job Posting
TimeSunday, 10 August 20258:00am - 9:00am PDT
LocationWest Building, Exhibit Hall C
Session TimeSunday, 10 August 20258:00am - 9:00am PDT
LocationWest Building, Exhibit Hall C