Lilia Urmazova
Senior Staff ML Evaluation Engineer
About
With over 25 years of experience architecting quality strategies from the ground up, I specialize in the rigorous evaluation, benchmarking, and quality assurance of Machine Learning models and Generative AI systems. My expertise bridges traditional engineering disciplines with advanced ML concepts, focusing on LLM performance, alignment, and behavioral evaluation. My first and second educations, applied linguistics and software engineering, merge ideally, providing a comprehensive view of AI systems.
-
Armenia
Information Technology & Services
Vector Calculus, Matrices, Gradient Boosting, Linear Regression, Multivariable Calculus, Langfuse, Amazon Web Services (AWS), Git, MongoDB, GraphQL, Usability Testing, Prompt testing, Python (Programming Language), Java, Spring Framework, Selenium, Selenium WebDriver, TestNG, JUnit, REST APIs
Experience

Senior Staff ML Evaluation Engineer
NDA (Global developer-centric software company)
- Collaborating in a cross-functional ML team on cutting-edge R&D initiatives, focusing on scalable AI architecture. - Leading LLM evaluation to assess reasoning capabilities, tool-use proficiency, and output quality, selecting and validating evaluation metrics (e.g., G-Eval, RAGas, BLEU/ROUGE, CLIP family, or custom heuristics) - Evaluating AI agent capabilities, including visual, code, and text artifacts, using mutation testing and other reasonable approaches. - Developing end-to-end evaluation pipelines to automate performance tracking and regression testing for machine learning models.

AI QA Lead-Mentor
Armenia
#1 in the CIS for student IT employment. - Leveraged experience in AI testing to contribute to the hiring and onboarding of junior testers, including specialists in AI testing methodologies. - Authored and delivered onboarding educational materials for junior QAs and a textbook covering foundational QA skills (SQL, API, Web, test design, Agile), demonstrating expertise in knowledge transfer and process education.

AI QA Lead-Mentor
United States
- Developed and implemented comprehensive testing strategies for AI Vertex search and AI Visual search for merchandisers, including model performance testing, prompt testing, ethics testing, data validation, and bias assessment. - Covered LLM metrics with Python automation and frameworks for AI testing, using LLM-as-a-Judge approach, and model cost analysis. - Prepared and managed large-scale datasets in cloud storage for AI Vertex Search, optimizing data quality and accessibility for model training. - Led a team of 10 QA engineers, ensuring compliance with best testing practices and project deadlines in an Agile environment. - Collaborated with development teams and data specialists to define testing requirements and ensure end-to-end AI functionality quality.

AQA Lead
- Led and actively contributed to automation efforts to achieve a best-in-class test pass rate of 100% for Java TestNG API, SQL, and 95% for UI Java Selenium automated tests with BDD (Cucumber), demonstrating a commitment to robust and reliable testing. - Collaborated with the customers’ team to align the comprehensive manual and automation testing strategy with dynamic requirements in an Agile environment. - As a Lead, conducted 1:1, Performance reviews, controlled and motivated people, and set SMART goals, helped several subordinates to grow up to seniors and leads. - Received the “QA Engineer of the Year” accolade several times. - Leading a successful company-level QA onboarding project, which reduced hiring efforts by 50%.

QA Engineer
- Proactively minimized production defects through meticulous testing practices, resulting in a demonstrably positive impact on product quality and customer experience. - Led multiple projects as a solo QA, both manual and automation on TestComplete and SOAP UI. - Was selected as the best QA for a challenging project among 100+ QA in the company.

Manual QA Engineer
BMicro
- Being the only QA in the company, I pioneered the QA processes and practices - Negotiated and started the usage of testing documentation and bug tracking system - Performed a full-fledged testing of a box B2B product focusing on quality and customer needs
Lilia Urmazova's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.



