Vaibhav Bhembre

Vaibhav Bhembre

Senior Infrastructure Engineer @ CoreWeave

About

I have designed and architected systems on and within cloud infrastructure for over 10 years. I am extremely passionate about building software that directly improves productivity and lives of individuals. My recent focus has been solving hard challenges related to large scale distributed systems, starting at the order of tens of petabytes and 100s of nodes, while enforcing strong performance and reliability guarantees through all the phases of growth. In past I have scaled systems to process and store data at the rate of nearly 3 billion events per day. Other than working on technical solutions, I enjoy mentoring engineers helping impart engineering best practices and effective debugging and problem solving methodologies. I have worked deeply on nearly all major system languages as I have described in my projects but at this moment I am most interested in writing software in Go because of its value proposition and my experience with it. I also possess substantial experience operating and customizing internals of Ceph distributed storage system. I coauthored a book to share my learnings on it that has successfully sold hundreds of copies till date. Additionally I have written pieces that interacted with and modified Qemu/KVM and Libvirt. I have built production services on top of stateful systems such as Consul, Vault, MariaDB, MySQL, PostgreSQL, MongoDB and RocksDB, including queueing systems such as Kafka, ActiveMQ and RabbitMQ. I have integrated CI environments such as Concourse, Drone and Jenkins into production.

Country

United States

City

New York

Industry

Computer Software

Skill

Fluentd, OPA, Calico, Concourse CI, Go (Programming Language), Kubernetes, Rancher Kubernetes Management, Argo, CockroachDB, gRPC, C, C++, PHP, Java, JavaScript, Distributed Systems, Perl, CSS, jQuery, MongoDB

Experience

CoreWeave

Senior Infrastructure Engineer

CoreWeave

LinkedIn
2023-5 - Present · 3 yrs 5 mos
DigitalOcean

Staff Engineer

DigitalOcean

LinkedIn
2021-3 - 2023-5 · 2 yrs 3 mos

New York, New York, United States

Wrote and published architecture design docs and RFCs to help kick-off an org-wide move to bare-metal container orchestration using Kubernetes. The proposals described mechanisms for supporting granular team-based RBAC for groups of namespaces and clusters, aiding workload placement using cluster-labels, thus extending the notion of “nodes as cattle, clusters as pets” to even “clusters as cattle”. Led design discussions with product engineering, networking, infrastructure, and capacity management teams to plan out bare-metal container orchestration rollout across all of DigitalOcean’s geographically distributed datacenters. The rollout was successfully completed with zero downtime within the required time commitments. Championed OLTP architecture evolution to solve DigitalOcean’s globally distributed dataset challenges, comparing several solutions before finalizing on CockroachDB. I published RFCs detailing the need and benefits, and worked on writing CockroachDB automation for on-premise deployment using its Kubernetes Operator. I pushed several improvements upstream helping fix outstanding deficiencies affecting our use-cases. I worked with product teams to onboard a large set of their workloads onto the managed Kubernetes platform. I helped accelerate adoption by writing in-depth documentation, providing tutorials, and conducting synchronous walkthroughs as and when needed. Built a proof-of-concept to integrate safe deployments for services running on bare-metal kubernetes touted as “unbreakable prod” with zero-touch integration, aiding seamless transition to this new deployment model, for platform consumers. In addition to basic Kubernetes primitives, this mechanism tapped into gRPC passthrough Ingress support for API-functionality-level blue-green, staged canary rollout, pre/post-promotion analysis, threshold-based automatic rollbacks and A/B testing.

DigitalOcean

Senior Software Engineer II

DigitalOcean

LinkedIn
2018-7 - 2021-3 · 2 yrs 9 mos

Greater New York City Area

Led effort to migrate manual deployments for a set of storage control-plane services to use Concourse CI for fully automated deployments with sane defaults but necessary overrides in place. Pushed the teams to increase reliance on synthetics, smoke tests and canaries to determine forward progress when in the golden path which improved developer productivity and service uptime substantially. Wrote an internal Go service to safely assist lifecycle operations for nodes containing Ceph data and index OSDs. This became a critical tool in our toolset for managing our customer-facing multi-Petabyte cluster node augments and decommission. Node augment process was reduced to just 1 person-hour per augment compared to 7 person-days. The project was eventually open-sourced to help the Ceph community: https://github.com/digitalocean/archimedes. Published design proposal for implementing a global auto-remediation service to analyze and mitigate faulty behavior in production for externally-written applications that could not have a short-term fix. I helped guide my juniors, using this as a mentoring opportunity, to build and ship this service in a relatively short amount of time. In lieu of a dedicated Product Manager, I stepped up to present an extended Block Storage roadmap to the leadership that included a strong pitch for offering a FileSystem solution showcasing considerable ARR wins.

DigitalOcean

Senior Software Engineer

DigitalOcean

LinkedIn
2016-11 - 2018-6 · 1 yr 8 mos

Greater New York City Area

I was responsible for leading technical aspects of the product ranging from backend architecture design on top of Ceph, development and code/design review to benchmarking and hardware selection in order to build and ship Block Storage, DigitalOcean's first storage integration. Along with proposing design RFC for the Block Storage implementation, I worked on writing components atop management and backend layers, automating operational work, building performance testing tools, and led technical decision-making for feature rollouts in addition to projects aimed at long-term system stability and planning capacity. I interfaced with all necessary teams including product, sales, marketing, support and datacenter operations to drive expectations leading up to the launch and subsequent feature rollouts alongside Storage Engineering Manager. I was appointed as Tech Lead of Storage team after launch. I was handed over the charge to create and lead a new team for DigitalOcean’s 2nd major storage-related product offering after flagship Droplet, Spaces Object Storage. The product was launched on schedule in October ‘17 and with desired feature-set. Achieved Q4 company OKR of acquiring 10,000 Spaces MAU within weeks of launch. My responsibilities included educating Storage engineers and System administrators on internals of backend engine for Block and Object Storage by writing tutorials and conducting knowledge sharing sessions. I took out time to coauthor “Learning Ceph” book based on years of experience of operating Ceph in production for the benefit of a larger audience. My team launched Object Storage successfully across US, Europe and Asia data centers within 3 months, working through various challenges that spanned bugs at scale, unique design, custom hardware setup and supplier issues. I was awarded Starfish stock grant as a result that recognizes outstanding performers and talent. I was also featured in builtinnyc.com article based on 7 best engineers in NYC.

DigitalOcean

Software Engineer

DigitalOcean

LinkedIn
2014-1 - 2016-10 · 2 yrs 10 mos

Greater New York City Area

I joined the Compute team that owned the management of online infrastructure operating on over 9000 machines across 5 data centers spanning US and Europe. I assisted in version 1.5 (v1.5) rewrite of the backend event subsystem layer that acted as the core of our cloud. The v1.5 rewrite was a major improvement that finally surfaced IPv6 from our stack, added ability to make networking changes on a droplet live and considerably reduced snapshot and backup service times. SGP1 datacenter was the first to receive all these changes and then we embarked on another big project to retrofit these changes live on existing regions. 7 new data centers were launched in quick succession with v1.5. During my tenure in Compute, I wrote new low-level services to optimize virtualization experience and to support existing services in production. I wrote code in Perl and C, working primarily on QEMU-KVM, Libvirt and our internal cloud core infrastructure. I took over cross-region Image Transfers project, architected and implemented an end-to-end design whose objective was to move images seamlessly and securely using SRX tunnels across any of our data centers on-demand. Towards the end of the year we started transitioning towards Go as our primary runtime for all services. I worked as a part of internal GoTeam guild to create new practices for smooth transition towards Go. The tasks I performed involved writing build and CI integrations for our monorepo, detect dependencies on internal libraries and improve dependency testing times, documenting new engineering workflows and mentoring engineers on our use of Go. I undertook another project that involved writing a metrics aggregation service to perform analysis and monitoring of Droplet workloads. This service was initially implemented using sharded MySQL databases. We experimented with InfluxDB for storing certain usage-related time-series metrics and using cluster message queues based on RabbitMQ to send aggregated stats payload.

Marigold

Lead Reliability Engineer

Marigold

LinkedIn
2013-8 - 2014-1 · 6 mos

Greater New York City Area

My primary set of goals were to build and lead a core Reliability team with an objective of instrumenting and enhancing Sailthru’s overall system architecture built on CentOS and Ubuntu distributions. An added responsibility was to build powerful communication primitives for improving cross-team workflows that affected all teams company-wide. In my short stint in this role, one of my major achievements was speeding up global campaign generation time by adding a disk-backed caching layer for sends. The enhancements that were rolled out offered nearly 40x performance boost for batched processing of incoming impressions. This was deemed as a big win as it helped resolve a large chunk of client-facing issues that were growing with scale and impacting reliability of our services.

Marigold

Lead Backend Engineer

Marigold

LinkedIn
2012-8 - 2013-8 · 1 yr 1 mo

Greater New York City Area

One of my major projects was to revamp Sailthru’s legacy Task-Management system to handle increasing load for up to 3B impressions/day. Me and my team were able to redo the complete Task-Management system to operate in a distributed manner and handle dispatch, control and monitoring of backend/frontend processors much efficiently. I took on the task of redeveloping the foundation of internal API service used engineering-wide by creating an intermediary built on Facebook's Thrift protocol. The objective was to serve previews and facilitate finer control over all types of email delivery by establishing clear contracts between all services within engineering. I led the team of six engineers to fine-tune and improve the SLA of email delivery and report generation. We were able to produce a game changing effort to generate campaign statistics in real-time moving away from periodic and on-demand processing using high-performance message queue ActiveMQ. This helped us scale our impressions collector that was designed for campaigns sent to 10K users to now effectively process campaigns that went out to 4MM users (x400 more users). I oversaw the code conversion of principal QueryBuilder component from PHP to Java. The motivation was tapping into multi-threading primitives that Java offered out of the box. This project was completed much before its slated deadline so we extended the scope to expose QueryBuilder service via internal API that our clients could use. I was responsible for creating a complete backend release workflow. My team was the first to implement and promote usage of a continuous integration system that was then adopted engineering-wide. I mentored devs across the board on applying best coding practices (including choice of optimal data-structures and algorithms), conducted developer bootcamps for refining code reviews, organized hackathons, and actively encouraged discussion sessions for proposing new enhancements to various parts of our system.

Marigold

Developer

Marigold

LinkedIn
2011-8 - 2012-8 · 1 yr 1 mo

Greater New York City Area

My first project involved writing an internal UI toolset for non-devs within the company to access and modify client/account level information, perform timely actions on campaign sends and batch processing jobs, and tracking resources accurately for billing among other things. This functionality was written in PHP. While working on this project I found a way to significantly improve productivity of teams interacting with this service by writing an IRC bot in Perl for developers and customer success teams to access and update our dataset, and importantly to be notified of all the live changes instantly. I spend my first year supporting systems and automation for enhancing migrations of our backend services that were initially running on thousands of AWS EC2 instances, to our colocation facilities in New York. The automation was performed via Chef and Puppet. Worked on schema design and query optimization for our primary datasets that were stored in MongoDB in lieu of a full­-time DBA. I managed replica-set reconfiguration, failure analysis, benchmarking and monitoring for large distributed database clusters, at an order of 100s of Terabytes of user-profiles and their activities, across two colocation facilities. In quest to achieve big latency gains I implemented an internal enhancement into mongo-php driver (v1.2.6) for linux kernel version 3.1.10 to force pick only certain seeds first for queries and retrying other nodes only if they were non-responsive. These changes that were rolled out were written in C. The initial deployment system was hand-written from scratch and buggy. I was tasked with rewriting our whole deployment framework using Fabric in Python to enable modular and incremental deploys for over 5 different sets of repositories.

University at Buffalo

Graduate Research Assistant

University at Buffalo

LinkedIn
2010-6 - 2011-5 · 1 yr

Buffalo/Niagara, New York Area

Coauthored the design of a new and efficient incremental algorithm for solving partially-dynamic all-pairs shortest path (APSP) problem. The previously known optimal approach computed both incremental and decremental APSP in O(n^2 * logn) amortized time. Our novel approach optimized the incremental part by improving batched node additions to complete within theoretical lower bound of an APSP problem i.e. O(n^2). I also wrote the new incremental approach along with conventional dynamic-APSP algorithms for generating a comparative diagnostic based on SNA Centrality metrics (Closeness, Betweenness, Degree), across both C and Java runtimes. This newly developed algorithm was published at peer-reviewed journal "Social Network Analysis and Mining". This work was supported by the National Science Foundation via grant number ICES-1216082.

Intrack Inc.

Software Engineer

Intrack Inc.

LinkedIn
2007-8 - 2008-5 · 10 mos

I developed a product selection server application for assisting users with analyzing, comparing and evaluating several mainstream solutions in the E. H. & S. (Environmental Health and Safety) domain. I was responsible for working directly with clients in order to understand their requirements and features that would be later built into the application itself. These features were initially written using .NET 2.0 and integrated with ASP.NET and MSSQL db, and then later ported to .NET 3.0. I also worked on extending 3CMS - 'Church Management System' to support AJAX-control additions using Microsoft ATLAS (AJAX style server-side controls). My other responsibilities included researching new technologies and coordinating with teams including QA, system analysts, deployment and client professionals.

Education

Buffalo State University

Buffalo State University

LinkedIn

Computer Science

2009 - 2011 · 2 yrs

I was tasked with architecting a reliable, scalable and highly available distributed file system to be deployed over NSF-funded PetaShare, a wide-area data sharing project spanning 10+ institutions across Louisiana and Buffalo. It took me nearly 3 months to complete this project. This system was written on top of rule-based iRODs engine which is one of the oldest (20+ years) distributed databases out there. This filesystem was written in C++. As a part of my second semester project, my team devised an advanced graph algorithm to significantly parallelize LAP using CUDA tool-chain aimed towards effectively solving transportation and graph-matching problems. For my systems course, I chose to write a performance benchmark for Neutron, a recovery-based extension over embedded TinyOS kernel. I was able to successfully measure the performance of kernel-level versus application-level recovery unit reboots against an in-house fault-injection testbed. This project was coded in C.

University of Mumbai

University of Mumbai

LinkedIn

Computer Engineering

2003 - 2007 · 4 yrs

I created an advanced version of a static Chatbot which was able to converse with a user on a wide range of topics called i-S.O.C.H. (intelligent Speech Oriented Cognitive Hyperbot). A 2D visual was constructed for improved user engagement. Information extraction was done by writing APIs for MIT CSAIL's START project. Information deemed important was selectively and locally cached in AIML format, and was continuously updated with every conversation. The bot had speech capabilities. The element of speech synthesis was built with the help of Microsoft Speech API SDK. Adobe Flash (then Macromedia) was used for achieving lip-sync on the 2D bot visual. Separately, as my primer into Java, I designed a set of applets for a Quiz application(DB), Paint-brush application(UI), and an 'asteroids' game(collision detection + alpha blending + particle effects + reflections). This was submitted as a part of my third-year project.

Vaibhav Bhembre's Contact Information

Email

******@***.com

Phone

(**) *** ****

Find the Right Leads
Find Verified Contact Data

Try with: Jensen Huang @ nvidia.com Click to autofill
LeadContact awards, five-star ratings, and GDPR compliance badges

What LeadContact does well

Find verified emails, phone numbers, and decision-makers with 98% accuracy.

Find Leads

Find Leads

Find the right people by company, role, industry, location, and more.

925M+ professional profiles

Find Leads
Find Emails

Find Emails

Access verified email addresses for your target contacts.

657M+ emails

Find Emails
Find Phone Numbers

Find Phone Numbers

Get cross-validated phone data from multiple top sources.

239M+ phone numbers

Find Phone Numbers

More Accurate. Lower Cost.

Find contact data in 1 tool with 98% accuracy

LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.

LeadContact Logo
Competitor Tools

All these = $289 per month

Great conversations start with the right contact.

It’s time to find yours.