Thomas Wentworth

Thomas Wentworth

Principal AI/ML Manager @ Aetna, a CVS Health Company

About

Experienced data scientist and manager with a passion for solving problems.

Country

United States

City

Greater Boston

Industry

Internet

Skill

Technical Leadership, Written Communication, NoSQL, Research and Development (R&D), Pattern Recognition, Team Building, Neural Networks, Predictive Modeling, Solution Architecture, Cloud Infrastructure, Real World Evidence (RWE), Real World Data (RWD), Model Development, Data Management, Product Launch, Business Insights, Cross-functional Team Leadership, Generative AI, Data Strategies, GPT-4

Experience

Aetna, a CVS Health Company

Principal AI/ML Manager

Aetna, a CVS Health Company

LinkedIn
2025-10 - Present · 1 yr
ImportGenius

Technical Consultant

ImportGenius

LinkedIn
2025-10 - Present · 1 yr
S&P Global

Associate Director of Data Science

S&P Global

LinkedIn
2021-8 - 2025-3 · 3 yrs 8 mos

Cambridge, Massachusetts, United States

I lead a cross-functional team of data scientists and data analysts focused on improving data quality, extracting insights, and delivering scalable solutions across the Panjiva platform. Our work spans entity resolution, data imputation, and error detection/correction, powering core features of the product. Our work spans critical areas including classification, data value correction, and scalable data infrastructure. We apply advanced techniques to classify maritime shipment content using free-text fields, leading to a patented HS-Code imputation system that significantly outperformed prior neural network models. We correct unlabeled data errors by triangulating across multiple data sources, statistical signals, and qualitative inputs to drive measurable improvements in data reliability. We build and maintain reusable libraries to support model development, entity matching, and high-performance querying. To support entity resolution at scale, we develop tools that map noisy, unstructured entity mentions to standardized company profiles. This includes innovations like Tomenshtein (a Smith-Waterman-inspired similarity metric robust to truncation and translocation) and Min_Hashr (a min-hash-based fuzzy text indexer). Several of these tools are patent-pending and serve as foundational components of our platform’s data infrastructure. In addition to team leadership, I remain a hands-on contributor—personally driving multiple key projects such as our HS-Code parser and classifier, as well as being heavily involved in most other projects on the team.

S&P Global

Senior Lead Data Scientist

S&P Global

LinkedIn
2020-4 - 2021-8 · 1 yr 5 mos

Cambridge, Massachusetts, United States

Wayfair

Manager III (Data Science)

Wayfair

LinkedIn
2018-9 - 2020-1 · 1 yr 5 mos

Boston, Massachusetts

I led a team of up to 6 data scientists focused on estimating shipping costs for Pricing. Our main project was PriceShip, an algorithm that leverages simulation, multiple standard and home-made machine learning algorithms, and internal tools to produce shipping estimates that are more accurate, more granular, aware of business initiatives (such as our growing logistics network), and up-to-date. PriceShip was also built for horizontal scalability, running in docker on a Kubernetes cluster with ZeroMQ connecting the various components.

Wayfair

Data Science Manager II

Wayfair

LinkedIn
2018-2 - 2018-9 · 8 mos

Greater Boston Area

I managed a small team responsible for estimating cost and revenue inputs for pricing at Wayfair. The team owns 4 main projects: 1) PriceShip: Described above. 2) LineShip: This is a historical-based shipping cost algorithm to estimate the average shipping cost for all of our products at Wayfair. The model was designed to handle a wide variety of product types, shipping carriers, sparse data and an ever evolving shipping network. 3) Cost and Revenue Adjustments: This is an algorithm used by pricing to align all non-shipping cost and revenue estimates with actuals. 4) ORCA: This project provides a clean dataset of all of our historical orders that is catered to the needs of pricing and is now used as a "source of truth" for over a dozen of projects across Pricing. By using ORCA, Data Scientists and business users across pricing have access to a dataset that is accurate and easy to understand without needing to understand, and write code to correct for the various nuances of its source tables.

Wayfair

Data Science Engineer II

Wayfair

LinkedIn
2017-2 - 2018-2 · 1 yr 1 mo

Greater Boston Area

I continued developing and expanding LineShip (described above) to all of Wayfair US and CA pricing and developed the Cost and Revenue Adjustment model (described above). I also created and maintained a repository of shared python functions and classes for the team.

Wayfair

Data science engineer I

Wayfair

LinkedIn
2016-6 - 2017-2 · 9 mos

I developed the shipping cost model LineShip (described above) and deployed in the pricing of all US company revenue.

Variantyx

Advanced Development Associate

Variantyx

LinkedIn
2015-3 - 2016-5 · 1 yr 3 mos

Ashton MA

I developed algorithms for the detection of large heterozygous insertion or deletion mutations and paternity tests from whole exome sequencing data.

Koch Institute of Integrative Cancer Research

Postdoc

Koch Institute of Integrative Cancer Research

2014-8 - 2015-8 · 1 yr 1 mo

Cambridge, MA

I wrote algorithms to measure how RNA localization is affected by muscleblind.

North Carolina State University

PhD Graduate

North Carolina State University

LinkedIn
2011-1 - 2014-8 · 3 yrs 8 mos

Raleigh-Durham, North Carolina Area

Research under the advisement of Dr. Ilse Ipsen We are researching randomized sampling from matrices with orthonormal columns and associated matrix properties. Our results include bounds on the condition number of a sampled matrix, perturbation bounds on an associated matrix property called “leverage scores,” and an algorithm to generate test matrices with specified leverage scores. Our work has applications to a randomized least squares solver called Blendenpik. Blendenpik (which was originally described by Rokhlin and Tygert and then further examined by Avron et. al) can be much faster than more traditional deterministic methods for solving a least squares problem but comes with the trade-off of having a small failure probability and is most useful for problems involving massive amounts of data. In this application, our condition number bounds would give the probability that the solution is computed correctly and quickly. Our perturbation bound on leverage scores gives the necessary conditions for the leverage scores to be stable with respect to small matrix perturbations. Finally, our matrix generation algorithm provides researchers with a method to generate test matrices.

North Carolina State University

Instructor for MA 116: Introduction to Matlab

North Carolina State University

LinkedIn
2012-8 - 2013-5 · 10 mos

Raleigh-Durham, North Carolina Area

I taught two semesters of an introductory programing course in Matlab to approximately 70 students, most of who have had no prior programing experience. I focused on basic programing concepts and techniques with the goal of getting my students to write complete programs and scripts to accomplish a task as opposed to short, single-line code segments. With the exception of the textbook, I created all of the class materials for this course, including lesson plans, class lectures, practice problems, projects and tests.

North Carolina State University

REU Graduate Mentor

North Carolina State University

LinkedIn
2011-5 - 2011-7 · 3 mos

I worked with and lead a group of four students under the group’s faculty advisor Dr. Ilse Ipsen. The main goal of the program is to introduce undergraduates to research with the hope of publishing a paper. The students in our group had a particularly large variation in their abilities and, in addition to assisting them with research, it was my job to ensure that all of the students contributed as much as they could and to ensure that, regardless of skill level, they all benefited from the experience. In our group we researched randomized methods for matrix multiplication. During this program, we developed two new algorithms for matrix generation. The first algorithm was a variation of an algorithm described by Drineas et. al. where the product AB is viewed as a sum of outer products and AB is estimated by taking a smaller sum of weighted outer products. We modified this algorithm by multiplying A and B by a Hadamard matrix and a random diagonal matrix in order to spread out the data contained in A and B. This resulted in a moderate speed increase over the method by Drineas et. al. In the second algorithm, we instead orthogonalized some of the outer products using a procedure similar to Gram-Schmidt with a special randomized inner product. We then project our orthogonalized outer products onto AB. Initial testing showed that we can obtain a 10 to 100 fold increase in accuracy without a significant change in work when compared to the method by Drineas et. al. It is our intention to continue our analysis and publish our results.

North Carolina State University

Grader, Lecture Assistant and Recitation Leader

North Carolina State University

LinkedIn
2008-8 - 2011-5 · 2 yrs 10 mos

I have graded and held office hours for various undergraduate level math courses up through Calculus III. I have also given weekly recitation lectures for Calculus I as well as for Calculus I for life-science majors.

Rensselaer Polytechnic Institute

REU (NSF Grant) Research under Peter Kramer

Rensselaer Polytechnic Institute

LinkedIn
2007-6 - 2008-5 · 1 yr

Troy, New York

Research was done on the effects of species size on the network properties of a food web model. A food web model examines the interactions (who eats whom) between a set of different species and tracks their population. We tried to improve upon existing models by incorporating the effects of species size. In most cases a creature can only consume other creatures that are within a certain size range, and a larger creature should require more energy to survive.

Saint-Gobain

Summer Intern

Saint-Gobain

LinkedIn
2006-4 - 2006-8 · 5 mos

Northborough, MA

Strain testing of novel composite materials For this position I preformed break tests of new materials where we measure stretching and force applied to the test material until it breaks. In addition, I photographed the broken ends with a microscope for later analysis. I also did other various lab work and computations.

Massachusetts Institute of Technology

Summer Intern

Massachusetts Institute of Technology

LinkedIn
2005-5 - 2005-7 · 3 mos

Cambridge, MA

Mathematical modeling of blood flow and drug deposition around a stent. MIT Biomedical Engineering Center For this position I assisted in modeling blood flow and drug deposition around a stent using proprietary software. I also added a text based user interface to other modeling programs written in Matlab.

Education

North Carolina State University

North Carolina State University

LinkedIn

Applied Mathematics

2008 - 2014 · 6 yrs

Summer 2014 is my expected graduation date.

North Carolina State University

North Carolina State University

LinkedIn

Mathematics

2008 - 2012 · 4 yrs

Received Masters in 2012 while working on my PhD.

Rensselaer Polytechnic Institute

Rensselaer Polytechnic Institute

LinkedIn

Physics and Mathematics

2004 - 2008 · 4 yrs

Thomas Wentworth's Contact Information

Email

******@***.com

Phone

(**) *** ****

Find the Right Leads
Find Verified Contact Data

Try with: Jensen Huang @ nvidia.com Click to autofill
LeadContact awards, five-star ratings, and GDPR compliance badges

What LeadContact does well

Find verified emails, phone numbers, and decision-makers with 98% accuracy.

Find Leads

Find Leads

Find the right people by company, role, industry, location, and more.

925M+ professional profiles

Find Leads
Find Emails

Find Emails

Access verified email addresses for your target contacts.

657M+ emails

Find Emails
Find Phone Numbers

Find Phone Numbers

Get cross-validated phone data from multiple top sources.

239M+ phone numbers

Find Phone Numbers

More Accurate. Lower Cost.

Find contact data in 1 tool with 98% accuracy

LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.

LeadContact Logo
Competitor Tools

All these = $289 per month

Great conversations start with the right contact.

It’s time to find yours.