AI/ML Solution Architect
doghouse recruitment·ICT & data·thuiswerken mogelijk·geplaatst 8 juli 2026·nog niet gecontroleerd
Vacaturetekst
AI/ML Solutions Architect – Distributed Training & GPU Infrastructure
Location: Remote from anywhere in the U.S. or Canada
Total compensation up to ~$500k (base + variable), depending on level and experience
Join a fast-moving AI infrastructure team working on the cutting edge of large-scale ML workloads. This role is ideal for engineers who enjoy solving deep technical challenges in distributed training, multi-GPU systems, and scalable AI inference infrastructure. You will work directly with AI-focused clients, helping them get the most out of modern GPUs (H100, B200, etc.) and ML frameworks such as PyTorch (and JAX in some environments).
Work alongside senior AI and infrastructure engineers building large-scale GPU platforms. As part of the customer solutions team, you will:
Design and validate production-grade distributed training (primary) and large-scale inference architectures on large GPU clusters, typically tens to thousands of GPUs
Work hands-on with customers to debug, optimize, and scale ML workloads across multi-node GPU environments
Act as a technical authority on GPU performance, networking, and schedulers , making trade-offs at scale and translating customer needs into concrete platform requirements
Collaborate closely with engineering, product, and R&D to influence roadmap decisions based on real-world ML workloads
This is a hands-on, technical role ; you are expected to work directly in customer environments, not only advise at a high level
Hands-on experience designing and operating production-grade, multi-node GPU workloads for training or inference
Strong background in distributed deep learning (PyTorch Distributed, DeepSpeed) on GPU clusters
Deep understanding of GPU architecture and interconnects (H100/A100 class, NVLink, InfiniBand)
Experience with Kubernetes or Slurm and performance tuning using GPU profiling and monitoring tools
This role is not a fit if your experience is limited to single-node training, high-level AI strategy, or non-production research environments. We are looking for engineers and architects who thrive at the intersection of AI workloads and large-scale infrastructure.
Meer bij doghouse recruitment
- Utrecht
Utrecht | €85.000 | MedTech company transitioning from legacy environment | Scale up | Replatforming towards AWS, Terraform, Gitlab, ECS --> EKS | Multi-tentant Saas migration Do you want to work for…
- Utrecht
|Azure/Terraform/Kubernetes|Utrecht|115k No Relocation possible Ready to take end-to-end ownership of complex cloud infrastructure for leading clients? Do you get energy from both designing and…
Business (ERP) Consultant | SaaS, Finance & Vastgoed (80k/Hybride/Leaseauto/Bonus) Ben jij een ervaren Business (ERP) Consultant die finance en software moeiteloos met elkaar verbindt? Voor een…
- Amsterdam
As a C++ / Rust Developer in either the Pircing or Market Making team, you will design and build high-performance systems and plug-ins for order placement or market making infrastructure, around core…
- Utrecht
Vacancy: AWS Consultant | Utrecht Area Tech: AWS | CloudFormation | Python | CI/CD | Kubernetes/EKS For the main consultancy company in the Netherlands when it comes to AWS, we are looking for…
Cloud Solutions Architect – AI/ML Cloud Infrastructure – Remote (US/Canada) – Up to $320K OTE (higher for principals) Our client is a cloud technology company driving the next generation of AI…
En nog 39 vacatures bij deze werkgever; die vind je via de lijst.
Bron
Wij zagen deze vacature op de website van de werkgever. Daar staat de actuele tekst; wijzigingen na 30 juli 2026 zien wij pas bij de volgende controle. Wij hebben deze vacature na het vinden nog niet opnieuw gecontroleerd, dus hij kan inmiddels vervuld zijn.