Machine Learning Engineer, Model Evaluation

Hace 18 horas

san martin cp hualtaco i, piura, Perú XenonStack Moments Jornada completa S/ 444,000 Por obra

About Xenonstack XenonStack is a Data and AI Foundry for Agentic Systems, enabling enterprises to design, deploy, operate, and scale intelligent agents across digital and physical environments.

About Xenonstack XenonStack is a Data and AI Foundry for Agentic Systems, enabling enterprises to design, deploy, operate, and scale intelligent agents across digital and physical environments.

We Build Enterprise-grade Platforms Across The Agentic Stack

  • Akira AI — Reasoning and agent orchestration. Turn models into collaborative, policy-governed agents that learn and act together.
  • ElixirData — Agentic analytics intelligence. Explainable, decision-centric analytics for measurable business outcomes.
  • NexaStack — Agentic infrastructure automation. Secure, compliant AI deployment across cloud, edge, and on-prem.
  • MetaSecure — Trust, compliance and defense. Continuous assurance with AI-BOMs, risk scoring, and agentic security.

Our mission is to accelerate the world’s transition to AI + Human Intelligence by making agentic systems reliable, responsible, and enterprise-ready.

THE OPPORTUNITY

We are seeking an Machine Learning Engineer, Model Evaluation to ensure that large language models (LLMs) and agentic AI systems meet enterprise-grade standards of accuracy, safety, and trustworthiness .

This role focuses on evaluating, benchmarking, and stress-testing LLMs in real-world workflows, building frameworks for reliability, robustness, and continuous improvement . If you thrive at the intersection of AI research, applied testing, and responsible deployment , this is the role for you.

Key Responsibilities

  • Evaluation Frameworks
    • Design and implement LLM evaluation pipelines covering accuracy, robustness, safety, and bias.
    • Develop automated systems for benchmarking models on enterprise-relevant tasks.
  • Reliability Engineering
    • Conduct stress tests, adversarial testing, and edge-case evaluations.
    • Build tools to measure latency, consistency, and error recovery in multi-turn interactions.
  • Metrics & Monitoring
    • Define KPIs such as factual accuracy, hallucination rate, toxicity, and compliance alignment.
    • Establish real-time monitoring for drift, anomalies, and performance regressions.
  • Collaboration & Alignment
    • Partner with ML engineers, product managers, and domain experts to align evaluation with business objectives.
    • Work with Responsible AI teams to implement ethical, explainable, and compliant evaluation practices.
  • Continuous Improvement
    • Feed insights from evaluation into fine-tuning, RLHF/RLAIF pipelines, and model selection.
    • Maintain a central repository of test cases, benchmarks, and evaluation results.
  • Research & Innovation
    • Stay current with state-of-the-art LLM evaluation techniques, from academic benchmarks to applied enterprise metrics.
    • Explore automated evaluation using agentic test harnesses and synthetic data generation.

Skills & Qualifications

Must-Have

  • 3––6 years in AI/ML, NLP, or applied model evaluation.
  • Strong understanding of LLM architectures, prompt engineering, and failure modes.
  • Hands-on with evaluation frameworks (Eval harnesses, Ragas, OpenAI Evals, DeepEval).
  • Proficiency in Python and libraries like LangChain, LangGraph, LlamaIndex, Hugging Face.
  • Experience with vector databases, RAG pipelines, and knowledge graph integration.
  • Familiarity with bias/fairness testing and Responsible AI frameworks.

Good-to-Have

  • Experience with reinforcement learning (RLHF, RLAIF) and reward modeling.
  • Exposure to agentic evaluation frameworks (multi-agent stress testing, synthetic user simulators).
  • Knowledge of compliance and safety requirements for BFSI, GRC, or SOC use cases.
  • Contributions to open-source evaluation libraries or research papers.

WHY SHOULD YOU JOIN US?

  • Agentic AI Product Company

Ensure reliability in cutting-edge AI platforms that are redefining enterprise adoption.

  • A Fast‑Growing Category Leader

Be part of one of the fastest-growing AI Foundries , powering Fortune 500 enterprises with trustworthy AI.

  • Career Mobility & Growth

Grow into roles such as AI Systems Architect, Responsible AI Engineer, or Reliability Engineering Lead .

  • Global Exposure

Work on enterprise-scale evaluation challenges across BFSI, Healthcare, Telecom, and GRC.

  • Create Real Impact