Ryan Crawford
Sr HPC Operations Engineer @ Lambda
About
TS/SCI-Cleared Senior HPC Engineer with expertise in designing, deploying, and optimizing mission-critical GPU clusters and hybrid cloud HPC solutions for defense-sector enterprises. Proven leader in end-to-end cluster installations, validation, and storage integration, delivering 30+ clusters for some of the largest and most prestigious companies in the world. Specializing in InfiniBand networking, NVIDIA GPU ecosystems, and performance tuning for AI/ML workloads. Combines deep technical acumen in HPC DevOps and Site Reliability Engineering (SRE) with experience leading multi-site teams to ensure operational excellence. Skill set: Many different Linux OS’s including Redhat, CentOS, Rocky, Debian, and Ubuntu. Network architecture, including HDR and NDR Infiniband configuration and topologies, Mellanox managed and unmanaged switches, UFM, OpenSM, high speed ethernet. Lustre storage systems. Vast Storage systems. Highly experienced SLURM resource, acting as lead Slurm expert for first SLURM production cluster in HPCMP. Deep knowledge of slurm accounting, qos’s, Lua plugins, advanced configuration of reservations. Deep knowledge of installing, configuring multiple compilers and communication softwares such as AOCC, GCC, Intel Compilers, Intel MPI, OpenMPI, HPC-X, UCX Vast knowledge of different cluster management tools. Ansible, CMU, Clusterware, MAAS. Node imaging, cluster monitoring, highly focused on cluster repeatability. Installed and provided monitoring through Nagios, Thruk, and Zabbix. Spearheaded custom monitoring scripts during periods of system testing and acceptance. Hardware expert in all aspects of HPC clusters. Parts internal to nodes such as AMD and Intel CPU’s, motherboards, NIC’s, RAM, PCI-Eboards. Power related troubleshooting related to in rack transformers, power shelves, UPS’s, power supplies, power feeds from facility to rack. Installed and troubleshooting cooling capabilities related to Cooling distribution units, MotivAir cooling doors, direct to chip liquid cooling. Cabling procedures.
United States
Greater New Orleans Region
Computer & Network Security
AI, HPC, GPU, VAST, Intelliflash, SHARP, HPC-X, OpenMPI, UCX, NVLink, Unmanaged Switches, Managed Switches, NCCL, FIO, NVBandwidth, DGX, HGX, H100, H200, B200
Experience

Sr HPC Operations Engineer
- Led GB300 new product introduction, serving as the SME for system setup, configuration, validation, - and benchmarking. - Directed NVL72 NVLink Switch product introduction, including imaging, configuration, validation, and production deployment. - Headed XDR InfiniBand and NVIDIA UFM initiatives, driving next-generation interconnect adoption across large-scale GPU clusters. - Spearheaded HPC engineering efforts for the first GA deployment of B200 GPU clusters, establishing scalable architectures optimized for large AI/ML workloads. - Contributed to the company’s first Managed Slurm offering, enabling automated workload orchestration and hybrid-cloud HPC operations. - Led validation and installation of the largest HPC cluster deployment in 2024 for a major enterprise customer, ensuring on-time delivery and high performance—resulting in multiple follow-on system deployments. - Served as InfiniBand SME for cloud-scale HPC environments supporting 10,000+ GPUs. - Pioneered NVIDIA SHARP implementation across cloud-based HPC clusters to accelerate collective communication performance. - Led end-to-end deployment of 6+ GPU clusters and played a major role in 20+ cluster installations, including on-prem DGX systems for healthcare institutions, government agencies, and academic research organizations. - Defined and validated cluster software and firmware stacks, including OS (RHEL/CentOS), NVIDIA drivers, CUDA, OFED, InfiniBand HCA firmware, switch firmware, and system BIOS. - Designed and executed cluster validation and stress-testing frameworks across 12+ deployments using tools such as NCCL, NVBandwidth, FIO, DCGMI, GPU_burn, and MLPerf benchmarks. - Drove engineering initiatives to integrate third-party cloud HPC solutions, improving interoperability and scalability across hybrid infrastructure.

HPC Team Lead
Stennis Space Center
- HPC Team lead for 2 DOD customers, overseeing 6 engineers maintaining 4 systems, 2 of which have landed in the top half of the Top500 - Primary Slurm technical expert for DOD sites first Slurm clusters - Installed 2 Nvidia DGX’s into a commercial vendor environment, installed all hardware and software, including the configuring additional 2 infiniband switches. - Successfully led team through 3 major DOD system security audits - Planned and mentored team through RHEL 8.8 OS upgrades on all clusters - Headed planning,installment, and integration of an additional 80 node rack into 1500 node cluster. - Headed benchmarks of new installation, including build config of benchmarks such as GAMESS, HYCOM, LINPAK, MPI Ping Pong, and Storage Benchmarks MDTEST and IOR - Installed and configured Robinhood for Lustre, for advanced storage metrics - Assisted in the application of Hot Pool functionality on Lustre file system, allowing for flash storage and hard drives to be in same file system and data to be moved from to the other when utilized or created. - Maintaining and configuring multiple versions of openmpi ranging from 4.1.6 to 5.0.1 with use for aocc,gcc,intel, nvhpc compilers

Sr HPC and Storage Engineer
- Lead On-Site Engineer for install, implementation, and maintenance of DOD HPC Cluster (8.2 petaFLOPS, 176,128 compute cores, 382 TB of memory, and 26 petabytes of storage) 1 - Configured and Installed Mellanox InfiniBand high speed network across 157 IB switches and 4000+ links - Administered a DragonFly Plus topology through UFM - including SHARP implementation - Headed First HPCMP Program & Site Slurm implementation, advanced QoS’s, plugins, accounting. Integration with program Software License server and User Management system. - Primary POC and Liaison between Customer and Vendor relation - Led Contract Acceptance Testing in the minimum possible window of Completion. Completing System Effectiveness Level Testing and System Capabilities Testing in 34 days with an expected completion time of 60 days. - Assisted in the install of 3 Cooling Distribution Units and 37 MotivAir cooling doors - Primary resource for hardware failure diagnosis and replacement - includes AMD CPU’s, Gigabyte Motherboards, NVIDIA A100’s and A40’s, Mellanox switches, Bel Power PDU’s, Lustre Controllers and SFA’s. - Utilized Ansible and Clusterware for node imaging and cluster configuration - Oversaw 2 RHEL upgrades from versions 8.4 to 8.5 to 8.7 clusterwide. - Lead engineer responsible for strict system requirements including 98% uptime, 75%+ of system utilization, minimal system interrupts and node down time, over 30 day periods. - Trained and Led 1 Jr. Linux Admin for first time administration and maintenance of HPC Cluster

HPC DevOps
- Developed and managed HPC Log Management Tool, used to diagnose day to day critical hardware and software issues - Authored HPC System Report Application, identifying file misconfigurations, Lustre status, Infiniband and Fabric Issues - Primary military liason and tech lead for 2 DOD HPC systems - Utilization of CMU - Cluster management utility - Configured and installed Ganglia monitoring tool across multiple MPP clusters - Accomplished Nagios and Thruk script automation checks - Experienced in configuring and troubleshooting SLURM - Familiar with Configuration Management tools such as BitBucket, Salt, Git - Experienced in in Imaging nodes using cluster software - Prepared and Instructed 8 week Python Training program for 12 suboordinates - Supervised and Managed 2 airman, mentored through their Career Progression certifying them in over 262 job tasks
Ryan Crawford's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.


