{"slug": "ml-systems-performance-engineer-mfu-higgsfield", "title": "ML Systems Performance Engineer (MFU) — Higgsfield", "summary": "Higgsfield AI, a generative AI company with $500M in annual revenue run rate and 25M+ users, is hiring an ML Systems Performance Engineer (MFU) for its Almaty, Kazakhstan office. The role focuses on optimizing distributed training performance, including MFU, tokens/sec/GPU, and scaling efficiency, with responsibilities spanning profiling, parallelism strategies, and CUDA/Triton kernel development. The position offers a competitive USD salary, equity, and relocation support, and requires on-site work five days per week.", "body_md": "# ML Systems Performance Engineer (MFU)\n\n- Salary\n- Not published\n- Location\n- Almaty, Kazakhstan\n- Work type\n- On-site\n- Posted\n- today\n\n[Apply on company site (opens in new tab)](https://jobs.ashbyhq.com/higgsfieldai/82db7018-fa8a-46fd-a5e1-71c5d9f67fcc/application)\n\nWhy work at Higgsfield AI?\n\nHiggsfield AI is the fastest-scaling generative AI company in history, hitting $500M in annual revenue run rate, 25M+ users worldwide, 6M+ generations per day, and powering 390 of Fortune 500 brands. We're building at the absolute frontier of AI-powered video creation and next-generation creative tools. Joining Higgsfield means becoming part of a high-impact team shaping the future of AI-native experiences, at a company that isn't just moving fast, but rewriting what fast looks like.\n\nWhat you will do\n\n- Profile end-to-end training runs and identify bottlenecks across compute, memory, communication, storage, and orchestration.\n- Define, measure, and improve MFU, tokens/sec/GPU, scaling efficiency, training goodput, and GPU uptime.\n- Optimize distributed training and model-sharding strategies, including data, tensor, pipeline, context, and expert parallelism.\n- Improve collective communication through topology-aware placement and compute/communication overlap.\n- Develop or integrate optimized CUDA and Triton kernels\n- Optimize data loading, preprocessing, sequence packing, and checkpointing so that I/O does not leave accelerators idle.\n- Diagnose distributed hangs фтв performance regressions.\n- Improve fault tolerance for long-running training jobs.\n\nWhat we are looking for\n\n- Strong experience running and optimizing multi-GPU or multi-node training.\n- Experience with PyTorch Distributed or an equivalent training framework.\n- Understanding of GPU architecture, including memory hierarchy, Tensor Cores\n- Understanding of collective communication, cluster topology, and distributed-training bottlenecks.\n- Experience with distributed parallelism technologies such as FSDP, DeepSpeed, Megatron-LM, TorchTitan, or similar.\n- Ability to debug complex performance and reliability problems across multiple layers of the training stack.\n\nNice to have\n\n- CUDA, Triton or GPU-kernel development experience.\n- Experience with NCCL, MPI, UCX, RDMA, InfiniBand, RoCE, GPUDirect, NVLink, or NVSwitch.\n- Experience training Mixture-of-Experts, multimodal, or reinforcement-learning models.\n- Knowledge of PyTorch internals, torch.compile, XLA, ML compilers, or custom operators.\n- Experience with mixed-precision training, including BF16, FP8, or FP4.\n\nWhat We Offer\n\n- Competitive base salary in USD, based on your experience, skills, and the scope of the role.\n- Equity participation through the company’s stock option program, giving you the opportunity to share in Higgsfield’s long-term growth.\n- Relocation support to Almaty for candidates moving from another city or country.\n- A highly collaborative, fast-paced environment where you can work directly with experienced leaders and have a meaningful impact on the product and company.\n- Opportunities for professional growth, ownership, and career development as the company scales.\n- Company-provided equipment, meals, transportation, or other office benefits.\n\nThis is a fully on-site role based in our Almaty office. Our team works from the office five days per week for the full working day. We believe in-person collaboration is an important part of how we move quickly, solve complex problems, and build strong teams.", "url": "https://wpnews.pro/news/ml-systems-performance-engineer-mfu-higgsfield", "canonical_source": "https://frontierroles.com/jobs/higgsfieldai-ml-systems-performance-engineer-mfu-0cc0bf/", "published_at": "2026-08-24 06:22:01+00:00", "updated_at": "2026-08-25 03:44:24.374567+00:00", "lang": "en", "topics": ["machine-learning", "ai-infrastructure", "ai-research"], "entities": ["Higgsfield AI", "PyTorch Distributed", "FSDP", "DeepSpeed", "Megatron-LM", "TorchTitan", "CUDA", "Triton"], "alternates": {"html": "https://wpnews.pro/news/ml-systems-performance-engineer-mfu-higgsfield", "markdown": "https://wpnews.pro/news/ml-systems-performance-engineer-mfu-higgsfield.md", "text": "https://wpnews.pro/news/ml-systems-performance-engineer-mfu-higgsfield.txt", "jsonld": "https://wpnews.pro/news/ml-systems-performance-engineer-mfu-higgsfield.jsonld"}}