Thomas Wentworth
Principal AI/ML Manager @ Aetna, a CVS Health Company
About
Experienced data scientist and manager with a passion for solving problems.
United States
Greater Boston
Internet
Technical Leadership, Written Communication, NoSQL, Research and Development (R&D), Pattern Recognition, Team Building, Neural Networks, Predictive Modeling, Solution Architecture, Cloud Infrastructure, Real World Evidence (RWE), Real World Data (RWD), Model Development, Data Management, Product Launch, Business Insights, Cross-functional Team Leadership, Generative AI, Data Strategies, GPT-4
Experience

Associate Director of Data Science
Cambridge, Massachusetts, United States
I lead a cross-functional team of data scientists and data analysts focused on improving data quality, extracting insights, and delivering scalable solutions across the Panjiva platform. Our work spans entity resolution, data imputation, and error detection/correction, powering core features of the product. Our work spans critical areas including classification, data value correction, and scalable data infrastructure. We apply advanced techniques to classify maritime shipment content using free-text fields, leading to a patented HS-Code imputation system that significantly outperformed prior neural network models. We correct unlabeled data errors by triangulating across multiple data sources, statistical signals, and qualitative inputs to drive measurable improvements in data reliability. We build and maintain reusable libraries to support model development, entity matching, and high-performance querying. To support entity resolution at scale, we develop tools that map noisy, unstructured entity mentions to standardized company profiles. This includes innovations like Tomenshtein (a Smith-Waterman-inspired similarity metric robust to truncation and translocation) and Min_Hashr (a min-hash-based fuzzy text indexer). Several of these tools are patent-pending and serve as foundational components of our platform’s data infrastructure. In addition to team leadership, I remain a hands-on contributor—personally driving multiple key projects such as our HS-Code parser and classifier, as well as being heavily involved in most other projects on the team.

Manager III (Data Science)
Boston, Massachusetts
I led a team of up to 6 data scientists focused on estimating shipping costs for Pricing. Our main project was PriceShip, an algorithm that leverages simulation, multiple standard and home-made machine learning algorithms, and internal tools to produce shipping estimates that are more accurate, more granular, aware of business initiatives (such as our growing logistics network), and up-to-date. PriceShip was also built for horizontal scalability, running in docker on a Kubernetes cluster with ZeroMQ connecting the various components.

Data Science Manager II
Greater Boston Area
I managed a small team responsible for estimating cost and revenue inputs for pricing at Wayfair. The team owns 4 main projects: 1) PriceShip: Described above. 2) LineShip: This is a historical-based shipping cost algorithm to estimate the average shipping cost for all of our products at Wayfair. The model was designed to handle a wide variety of product types, shipping carriers, sparse data and an ever evolving shipping network. 3) Cost and Revenue Adjustments: This is an algorithm used by pricing to align all non-shipping cost and revenue estimates with actuals. 4) ORCA: This project provides a clean dataset of all of our historical orders that is catered to the needs of pricing and is now used as a "source of truth" for over a dozen of projects across Pricing. By using ORCA, Data Scientists and business users across pricing have access to a dataset that is accurate and easy to understand without needing to understand, and write code to correct for the various nuances of its source tables.

Data Science Engineer II
Greater Boston Area
I continued developing and expanding LineShip (described above) to all of Wayfair US and CA pricing and developed the Cost and Revenue Adjustment model (described above). I also created and maintained a repository of shared python functions and classes for the team.

Postdoc
Koch Institute of Integrative Cancer Research
Cambridge, MA
I wrote algorithms to measure how RNA localization is affected by muscleblind.

PhD Graduate
Raleigh-Durham, North Carolina Area
Research under the advisement of Dr. Ilse Ipsen We are researching randomized sampling from matrices with orthonormal columns and associated matrix properties. Our results include bounds on the condition number of a sampled matrix, perturbation bounds on an associated matrix property called “leverage scores,” and an algorithm to generate test matrices with specified leverage scores. Our work has applications to a randomized least squares solver called Blendenpik. Blendenpik (which was originally described by Rokhlin and Tygert and then further examined by Avron et. al) can be much faster than more traditional deterministic methods for solving a least squares problem but comes with the trade-off of having a small failure probability and is most useful for problems involving massive amounts of data. In this application, our condition number bounds would give the probability that the solution is computed correctly and quickly. Our perturbation bound on leverage scores gives the necessary conditions for the leverage scores to be stable with respect to small matrix perturbations. Finally, our matrix generation algorithm provides researchers with a method to generate test matrices.

Instructor for MA 116: Introduction to Matlab
Raleigh-Durham, North Carolina Area
I taught two semesters of an introductory programing course in Matlab to approximately 70 students, most of who have had no prior programing experience. I focused on basic programing concepts and techniques with the goal of getting my students to write complete programs and scripts to accomplish a task as opposed to short, single-line code segments. With the exception of the textbook, I created all of the class materials for this course, including lesson plans, class lectures, practice problems, projects and tests.

REU Graduate Mentor
I worked with and lead a group of four students under the group’s faculty advisor Dr. Ilse Ipsen. The main goal of the program is to introduce undergraduates to research with the hope of publishing a paper. The students in our group had a particularly large variation in their abilities and, in addition to assisting them with research, it was my job to ensure that all of the students contributed as much as they could and to ensure that, regardless of skill level, they all benefited from the experience. In our group we researched randomized methods for matrix multiplication. During this program, we developed two new algorithms for matrix generation. The first algorithm was a variation of an algorithm described by Drineas et. al. where the product AB is viewed as a sum of outer products and AB is estimated by taking a smaller sum of weighted outer products. We modified this algorithm by multiplying A and B by a Hadamard matrix and a random diagonal matrix in order to spread out the data contained in A and B. This resulted in a moderate speed increase over the method by Drineas et. al. In the second algorithm, we instead orthogonalized some of the outer products using a procedure similar to Gram-Schmidt with a special randomized inner product. We then project our orthogonalized outer products onto AB. Initial testing showed that we can obtain a 10 to 100 fold increase in accuracy without a significant change in work when compared to the method by Drineas et. al. It is our intention to continue our analysis and publish our results.

Grader, Lecture Assistant and Recitation Leader
I have graded and held office hours for various undergraduate level math courses up through Calculus III. I have also given weekly recitation lectures for Calculus I as well as for Calculus I for life-science majors.

REU (NSF Grant) Research under Peter Kramer
Troy, New York
Research was done on the effects of species size on the network properties of a food web model. A food web model examines the interactions (who eats whom) between a set of different species and tracks their population. We tried to improve upon existing models by incorporating the effects of species size. In most cases a creature can only consume other creatures that are within a certain size range, and a larger creature should require more energy to survive.

Summer Intern
Northborough, MA
Strain testing of novel composite materials For this position I preformed break tests of new materials where we measure stretching and force applied to the test material until it breaks. In addition, I photographed the broken ends with a microscope for later analysis. I also did other various lab work and computations.

Summer Intern
Cambridge, MA
Mathematical modeling of blood flow and drug deposition around a stent. MIT Biomedical Engineering Center For this position I assisted in modeling blood flow and drug deposition around a stent using proprietary software. I also added a text based user interface to other modeling programs written in Matlab.
Education
Thomas Wentworth's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.





