Bachelor's degree or above in Computer Science, Artificial Intelligence, Natural Language Processing (NLP), Machine Learning, Distributed Systems, or a related technical field.
Hands-on experience delivering or contributing to end-to-end LLM pre-training projects.
Proven experience participating in the pre-training of models with 7B+ parameters.
Strong hands-on expertise with distributed training frameworks such as Megatron-LM, DeepSpeed, or FSDP.
Practical experience working with large-scale training environments involving 64+ GPUs.
Strong understanding of large-scale pre-training data pipelines, including data cleaning, deduplication, quality filtering, tokenization, data mixing, and data quality optimization.
Experience designing or optimizing data pipelines for large-scale LLM training.
Strong ability to analyze training loss, gradients, convergence, and training stability.
Proven experience troubleshooting and resolving issues in large-scale distributed training environments.
Familiarity with long-context training and context-extension techniques, including
RoPE scaling, NTK-aware interpolation, and YaRN.
Preferred Qualifications
Experience with 70B+ parameter model pre-training.
Experience with Mixture-of-Experts (MoE) model pre-training.
Experience optimizing large-scale GPU clusters, distributed training systems, and AI training infrastructure.
Publications in top-tier AI/ML conferences such as NeurIPS, ICML, ICLR, ACL, or EMNLP, particularly in areas related to LLM pre-training, model architecture, or training optimization.
Tell employers what skills you have
Machine Learning Optimization Distributed Processing Salesforce Training Training Analysis End User Training Training Course Development Natural Language Processing Artificial Intelligence Model Training Computer Science Data Engineering Training Strategy GPU Programming SharePoint training Training Delivery
Create a job alert for this search
Large Language Model Pre-training Engineer • D13 Macpherson, Braddell, SG