Raul Perez Pelaez
Research Scientist @ Universidad Autónoma de Madrid
About
I accelerate compute-intensive workloads by combining deep CUDA expertise with modern software engineering practices. Over the past decade, I have built and optimized GPU libraries that support drug discovery, fluid dynamics, and deep‑learning frameworks used by thousands. What I do • Design CUDA/C++ libraries from scratch. e.g., UAMMD for complex‑fluid simulation. • Maintain and distribute GPU‑heavy packages through channels such as conda-forge. • Engineer robust CI/CD with tools such as GitHub Actions, CMake, and Conda so every commit ships with tests, docs, and binaries in a reproducible manner (GitOps). • Build GPU‑accelerated web apps (React/Astro + WebGPU/WebAssembly) and advanced OpenGL visualizers (compute shaders, SSAO, PBR). • Containerize and orchestrate workloads via Docker / Compose / Kubernetes. How I work – Emphasis on automated testing, documentation, and reproducible builds. – Cross‑disciplinary: comfortable jumping from low‑level CUDA kernels to TypeScript front‑ends or LLM tooling. – Community-minded: conference organiser, university lecturer, and PhD mentor. I thrive where physics, ML, and high‑performance engineering intersect. Seeking roles that enable me to push GPU hardware to its limits, whether that’s cores in an exascale cluster or GPUs in the cloud. Let’s connect if you need someone to turn “could be faster” into “is 50× faster.” I excel at writing scalable, testable, documented software following healthy community standards.
Spain
Madrid
Computer Software
High Performance Computing (HPC), Slurm Workload Manager, GitOps, complex fluids, Conference Organization, Event Planning, Project Management, DevOps, SYCL, conda, WebAssembly, WebGPU, Docker compose, Docker, jax, CMake, Nsight Compute, Continuous Integration and Continuous Delivery (CI/CD), GitHub, Triton
Experience

Research Scientist
Madrid, Comunidad de Madrid, España
• Developed several GPU-ready open source tools for high-performance scientific computing (C++/CUDA with Python wrappers). • Set up GitHub actions with custom runners for automatic testing and benchmarking in a local GPU cluster. • Developed new massively parallel algorithms for biosensor modelling, published in scientific journals. • Wrote a high-performance web application (WebAssembly via C++ with Emscripten, WebGPU, Astro/React) for a digital twin product. • Developed and implemented algorithms for research in liquid memory storage. • Mentored two PhD students.

Conference Organizer
Freelance
Founder and chair: First international DoNEV conference. • Conceived and organized an international meeting (CECAM flagship event) on non-equilibrium dynamics (60 + attendees from 10 countries). https://www.cecam.org/workshop-details/dynamics-of-non-equilibrium-variables-multiscale-multiphysics-applications-of-fluctuating-hydrodynamics-1408

Research Director
Barcelona, Cataluña, España
Provide core GPU libraries and tooling to research groups, industry, and OSS projects, mainly for AI-assisted molecular modeling and drug-discovery. Actions • Implemented custom torch.compile-ready, Autograd-enabled, CUDA ops in TorchMD‑Net; profiled & refactored neighbor search path. • Wrote and maintained conda‑forge recipes for OpenMM, TorchMD‑Net, OpenMM‑Torch, CUTLASS, pyg‑lib, NNPOPS, and more. • Integrated CUDA Graphs into OpenMM‑Torch for smoother ML‑augmented MD runs. • Built CI/CD matrices (GitHub Actions) with GPU, CPU, and cross‑platform wheels, plus auto‑generated docs. • Ported Neural Network Potential ML models to JAX, TinyGrad, and TensorFlow. • Wrote a Slack App integrating a local LLM model acting as an assistant in the internal workspace. The backend LLM model offered functionality such as tool usage, RAG, model hot-swap, paper summarization, and more. The LLM ran in an in-house GPU computing cluster, capable of hosting sub-60b parameter models (Llama family, Mistral,...). • Finetuned LLM models for custom workflows in drug discovery. Impact • TorchMD‑Net inference time reduced from hours to minutes (50 × speed‑up). • conda‑forge packages adopted by multiple bio‑physics labs; eased deployment on HPC clusters.

Predoctoral researcher in the area of fluid physics and complex systems
Madrid y alrededores, España
I specialize in the development of software for complex fluid simulations, implementing parallel code that leverages the GPU and CPU. • Conceptualized and developed Universally Adaptable Multiscale Molecular Dynamics (UAMMD), a complex fluid simulation toolkit with state of the art performance, written in CUDA/C++.
Education
Raul Perez Pelaez's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.




