Senior Solution Engineer – GPU & AI Infrastructure at Civo

4DayWeek Company Unknown Publicerat 11 augusti 2026
full_timeremotesenior
Responsibilities:** **Solution Design & Architecture** - System Design Documents: Author comprehensive High-Level Design (HLD) and Low-Level Design (LLD) documentation for enterprise-scale GPU supercomputing clusters. - Bill of Materials (BOM): Generate detailed BOMs covering compute nodes, NVLink switches, network fabrics, transceivers/cabling, liquid/air cooling requirements, power distribution, and high-performance storage. - GPU Cluster Topology: Architect scale-up (NVLink/NVSwitch) and scale-out network topologies (Fat-Tree, Rail-Optimized) for NVIDIA Blackwell platforms, specifically B300 and GB300NVL rack-scale architectures. - Fabric & Networking Engineering: Design high-throughput, low-latency networking architectures utilizing both InfiniBand (e.g., NDR/X800) and RoCE / RoCEv2 (e.g., NVIDIA Spectrum-X / Spectrum-4) with lossless Ethernet mechanisms (PFC, ECN, Adaptive Routing). - Multi-Tenant & Deployment Models: Deliver tailored architectures for both Bare-Metal (Slurm, OpenMPI, bare-metal provisioning) and Cloud-Native / Kubernetes environments (NVIDIA GPU Operator, Network Operator, Run:ai, KubeFlow). - Storage Integration: Architect high-bandwidth parallel storage solutions utilizing GPUDirect Storage (GDS) and enterprise AI file systems (e.g., VAST Data). **Technical Sales Support & Customer Engagement** - Partner with Civo’s sales and commercial teams as the technical lead for high-value AI infrastructure opportunities. - Engage directly with customer CTOs, Chief AI Officers, infrastructure leads, and ML engineers to evaluate technical requirements, compute sizing, and fabric choices. - Lead deep-dive architectural workshops and technical presentations on Civo's bare-metal GPU and managed Kubernetes offerings. - Produce precise technical proposals and lead responses to complex RFPs/RFIs regarding AI infrastructure. **Proof-of-Concept (PoC) & Benchmarking** - Architect and oversee Proof-of-Concept (PoC) deployments to validate real-world performance for customer workloads. - Benchmark cluster performance using industry-standard tools (NCCL tests, GPUDirect RDMA latency/bandwidth, MLPerf, Megatron-LM benchmarks). - Address network congestion, fabric routing, and thermal/power optimization during validation phases. **Product & Ecosystem Collaboration** - Serve as the bridge between enterprise AI clients, hardware vendors (NVIDIA, network OEMs), and Civo’s internal platform engineering team. - Provide continuous feedback to product teams on market trends, hardware platform demands, and feature requirements for AI/GPU orchestration. * * * **Key Results/Objectives:** - Technical Wins: Achieve high technical win rates on large-scale AI/GPU cluster sales opportunities. - Design Excellence: Successfully deliver complete, peer-reviewed HLDs, LLDs, and BOMs within target deal timelines. - Customer Satisfaction: Achieve successful PoC completion and sign-off for enterprise clients scaling AI workloads on Civo infrastructure. * * * **Requirements:** #### **Experience & Core Qualifications** - 5+ years in a Solution Architecture, Systems Engineering, or Technical Pre-Sales role focused on high-performance cloud, HPC, or AI infrastructure. - Bachelor’s degree in Computer Science, Electrical Engineering, Systems Engineering, or equivalent practical experience. #### **Technical Expertise** - NVIDIA GPU Architecture: Deep hands-on knowledge of NVIDIA HGX/DGX platforms, NVLink/NVSwitch fabrics, and Blackwell architectures (B300, GB300NVL, GB200 NVL72/NVL36). - High-Speed Networking: Expert-level knowledge of cluster fabric topologies: - InfiniBand: Quantum-2 / Quantum-X800, Subnet Management, Adaptive Routing. - RoCE / RoCEv2: Spectrum-X / Spectrum-4 Ethernet switches, PFC, ECN, RoCE configuration, and optimization. - GPU Direct Technologies: GPUDirect RDMA (GDR) and GPUDirect Storage (GDS). - Orchestration & Platforms: Proficiency in deploying and optimizing GPU workloads on: - Kubernetes: Container networking (CNI), NVIDIA GPU Operator, RDMA Shared Device Plugin, MPI Operator. - Bare-Metal: Slurm, Ansible, Terraform, PyTorch/NCCL environment tuning. - Documentation Skills: Demonstrated experience creating enterprise-grade HLDs, LLDs, network rack diagrams, and itemized BOMs. - Power & Thermal Awareness: Familiarity with high-density datacenter environments, liquid cooling technologies (Direct-to-Chip, CDU/liquid loop setups), and power delivery constraints for 100kW+ per rack deployments. #### **Soft Skills** - Strong technical leadership and presentation skills, with the ability to articulate complex network and hardware tradeoffs to executive stakeholders. - Problem-solving mindset capable of diagnosing complex hardware-software interaction bottlenecks in distributed training/inference setups. #### **Location** - Must be UK based. * * * **Nice to Have:** - NVIDIA Certified Professional: AI Infrastructure (NCP-AII). - NVIDIA Certified Professional: AI Networking (NCP-AIN). - NVIDIA Certified Professional: InfiniBand (NCP-IB). - NVIDIA Certified Associate / Professional: AI Workload Deployment & Cloud Native. * * * **Why Join Civo?** - Competitive compensation and benefits package. - 4-day week company (unless attending an event). - Uncapped holiday. - Remote work environment with flexibility and autonomy. - Collaborative and inclusive culture that values diversity and creativity. - Opportunity to work with a dynamic and innovative team in the fast-growing cloud industry. Senior Solution Engineer – GPU & AI Infrastructure Civo Apply nowSave

Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.

Visa jobbet och ansök

Ursprunglig annons: 4dayweek.io