We are looking for a hands-on Senior Applied AI Evaluation Engineer to improve how our RAN R&D organization analyzes engineering tickets, supports Root-Cause Analysis, recommends ownership, and learns from resolved cases.
You will build rigorous evaluation datasets, establish meaningful baselines, compare internal and approved external tools, analyze failure modes, and prototype improvements across retrieval, classification, prompting, agent workflows, and model selection.
This role is focused on measurable, evidence-based improvement rather than AI demonstrations. You will assess whether AI-generated conclusions are accurate, grounded in evidence, appropriately calibrated, and useful to engineering teams.
As a secondary area of focus, you will analyze engineering workflows at case and team level to identify bottlenecks, handoffs, dependencies, and opportunities for process improvement.
You will build rigorous evaluation datasets, establish meaningful baselines, compare internal and approved external tools, analyze failure modes, and prototype improvements across retrieval, classification, prompting, agent workflows, and model selection.
This role is focused on measurable, evidence-based improvement rather than AI demonstrations. You will assess whether AI-generated conclusions are accurate, grounded in evidence, appropriately calibrated, and useful to engineering teams.
As a secondary area of focus, you will analyze engineering workflows at case and team level to identify bottlenecks, handoffs, dependencies, and opportunities for process improvement.
Requirements:
BSc or MSc in Computer Science, Data Science, Machine Learning, Statistics, Electrical Engineering, or a related field, or equivalent practical experience.
5+ years of hands-on experience in applied machine learning, data science, search, natural-language processing, analytics engineering, or AI-enabled software systems.
Recent experience evaluating LLM, RAG, search, or agent systems using representative datasets, task-specific metrics, human review, failure analysis, and regression testing.
Strong Python and SQL skills, with experience building maintainable data pipelines, experiment workflows, services, or analytical tools.
Practical experience with several areas such as information retrieval, embeddings, hybrid search, reranking, classification, structured outputs, tool calling, or common LLM failure modes.
Strong statistical judgment, including sampling, leakage prevention, uncertainty, calibration, precision and recall, temporal drift, and controlled comparison of competing approaches.
Ability to work with semi-structured engineering data from issue-tracking systems, source control, code reviews, continuous integration, logs, dashboards, and test systems.
Clear communication skills and the ability to explain results, limitations, and tradeoffs to technical and business stakeholders.
BSc or MSc in Computer Science, Data Science, Machine Learning, Statistics, Electrical Engineering, or a related field, or equivalent practical experience.
5+ years of hands-on experience in applied machine learning, data science, search, natural-language processing, analytics engineering, or AI-enabled software systems.
Recent experience evaluating LLM, RAG, search, or agent systems using representative datasets, task-specific metrics, human review, failure analysis, and regression testing.
Strong Python and SQL skills, with experience building maintainable data pipelines, experiment workflows, services, or analytical tools.
Practical experience with several areas such as information retrieval, embeddings, hybrid search, reranking, classification, structured outputs, tool calling, or common LLM failure modes.
Strong statistical judgment, including sampling, leakage prevention, uncertainty, calibration, precision and recall, temporal drift, and controlled comparison of competing approaches.
Ability to work with semi-structured engineering data from issue-tracking systems, source control, code reviews, continuous integration, logs, dashboards, and test systems.
Clear communication skills and the ability to explain results, limitations, and tradeoffs to technical and business stakeholders.
This position is open to all candidates.


















