|
We are seeking a Principal Agentic AI Operations Engineer to provide hands-on technical leadership for the operationalization, evaluation, and reliable production performance of the advanced agentic AI systems powering Claritev's next generation of healthcare products. This role is for an experienced engineer who excels at making agentic AI systems trustworthy at scale - measurable, observable, benchmarked, secure, and cost-efficient in production. Where AI engineering teams build agents, you will own how those agents are evaluated, deployed, monitored, governed, and continuously improved. You will define what "good" looks like for agent behavior in high-stakes healthcare workflows and build the evaluation harnesses, benchmarks, and operational tooling that prove it. You will partner closely with Product, Engineering, AI Science, Security, and business leaders to establish the operational backbone of Claritev's agentic AI platform, set engineering and evaluation standards, make key architectural decisions, and mentor engineers and scientists across the organization. Job Roles and Responsibilities
- Own the end-to-end operational lifecycle of production agentic AI systems: deployment, versioning, rollout strategies, monitoring, incident response, rollback, and continuous improvement.
- Design and build agent evaluation and benchmarking frameworks, including offline evaluation suites, golden datasets, simulation environments, scenario-based and multi-turn benchmarks, regression testing, and A/B and shadow-mode online evaluation.
- Define and track agent quality metrics across task completion, tool-call accuracy, trajectory efficiency, hallucination and grounding rates, latency, cost per task, escalation/human-handoff rates, and safety violations.
- Build observability and distributed tracing capabilities for agentic workflows - spanning LLM calls, tool invocations, retrieval steps, memory operations, and orchestration graphs - using modern APM, log analytics, and OpenTelemetry-based instrumentation.
- Establish LLMOps/AgentOps best practices, including CI/CD for agents, prompt and model versioning, evaluation gates in deployment pipelines, experiment tracking, drift detection, and automated incident troubleshooting.
- Operate and optimize agentic and RAG workloads across managed generative AI services, container orchestration platforms, vector search, and enterprise data platforms in a cloud environment.
- Implement guardrails, policy enforcement, and human-in-the-loop controls for agent actions, with full auditability of agent decisions and tool executions.
- Manage capacity, throughput, and cost governance for LLM and agent workloads, including model[SC1] [SC2] routing, caching, batching, and GPU/compute utilization.
- Ensure secure and responsible operation of AI, including privacy, PHI/PII protection, secrets management, identity and access controls, explainability, auditability, and compliance with HIPAA and applicable data-governance requirements.
- Establish reusable operational frameworks, evaluation harnesses, dashboards, and runbooks that enable teams to ship and operate agents efficiently and consistently.
- Provide technical leadership across complex, cross-functional initiatives; influence architecture, reliability, and evaluation decisions beyond an individual project.
- Mentor engineers and data scientists and promote a culture of measurement-driven development, operational excellence, and pragmatic innovation.
Job Requirements Education
- Bachelor's degree in Computer Science, Engineering, Data Science, a quantitative discipline, or a related field required.
- Master's degree or PhD preferred.
Experience
- 10+ years of hands-on experience in software engineering, ML engineering, platform/SRE engineering, or a related technical discipline.
- 5+ years of experience operating production-grade ML or AI systems (MLOps, LLMOps, or ML platform roles).
- 3+ years of experience with generative AI, LLMs, RAG, and/or agentic AI systems in production.
- Demonstrated experience building evaluation and benchmarking capabilities for LLM or agent systems - or contributing to public agent benchmarks (e.g., SWE-bench, GAIA, AgentBench, tau-bench, HELM-style harnesses) with measurable quality and reliability outcomes.
- Demonstrated experience leading complex technical initiatives from concept through production deployment and measurable business impact.
- Technical Skills
- Strong software engineering skills, including expert-level Python proficiency and experience designing scalable services, APIs, and distributed systems.
- Deep experience with agent and LLM evaluation: eval harness design, golden datasets, LLM-as-judge techniques, rubric-based scoring, statistical significance testing, and regression detection.
- Experience with evaluation and observability tooling such as LangSmith, Langfuse, Arize Phoenix, Ragas, DeepEval, promptfoo, OpenAI Evals, Braintrust, or equivalent - and building custom harnesses where off-the-shelf tools fall short.
- Experience with agentic AI frameworks and patterns, such as LangGraph, LangChain, AutoGen, CrewAI, or equivalent, including tool use, planning, orchestration, memory, and guardrails.
- Experience with cloud environments and large-scale data platforms, including managed generative AI/LLM services, compute and GPU infrastructure, and enterprise databases.
- Experience with vector databases and embedding-based retrieval, and building/operating RAG pipelines.
- Experience with MLOps/LLMOps practices: CI/CD, model and prompt versioning, canary and shadow deployments, monitoring, experimentation, and incident troubleshooting.
- Experience with infrastructure-as-code (e.g., Terraform), containerization, and Kubernetes-based deployment.
- Working knowledge of deep-learning frameworks such as PyTorch or TensorFlow.
Other Skills
- Strong problem-solving, critical-thinking, communication, and organizational skills.
- Ability to communicate agent quality, reliability, and cost trade-offs clearly to technical and non-technical stakeholders.
- Ability to operate effectively in a fast-moving, cross-functional environment.
- Preferred Qualifications
- Experience with Oracle Cloud Infrastructure (OCI), including its generative AI, data science, database, and observability capabilities.
- Experience in healthcare, health technology, insurance, claims, payment integrity, or other regulated industries.
- Experience operating AI systems that process sensitive data, including PHI or PII, in HIPAA-regulated environments.
- Experience with red-teaming, adversarial testing, or safety evaluation of LLM/agent systems.
- Experience with process automation and integration across enterprise workflows and systems.
- SRE background: SLOs/SLIs, error budgets, and reliability engineering applied to AI systems.
Compensation The salary range for this position is $165-185K. Specific offers take into account a candidate's education, experience and skills, as well as the candidate's work location and internal equity. This position is also eligible for health insurance, 401k and bonus opportunity.
Why Claritev? Healthcare is complex. We help make it clearer. At Claritev, you'll do work that matters. Together, we're helping make healthcare more transparent and affordable for all through the power of data, technology, and expertise. We offer meaningful opportunities to grow your career, collaborate with talented colleagues, and make an impact on the clients and communities we serve. If you're looking for purpose, growth, and a team that succeeds together, you'll find it here. What Guides Us At Claritev, innovation, agility, and a focus on results drive our success. We embrace bold thinking, work as one team, take ownership, and strive for excellence in everything we do - creating meaningful impact for our clients, communities, and each other.
|