Assure AI Engineer
2 positions
Keywords: AssureAI · Evaluation Datasets · Red-Teaming · Bias & Explainability (SHAP/LIME) · Threshold Gates · CI/CD · Audit Evidence · Python
You are one of two hands-on operators of AssureAI's Trustworthiness pillar inside the Governance Control Tower. Where the AI Trust & Compliance Engineer sets the strategy — which regulations map to which checks, what a passing threshold means — you build and run the evaluation suites that prove it, day in and day out, across a growing roster of Gemini-based agents. You report to the AI Trust & Compliance Engineer and work closely with the AI Engineers building the agents you test.
-
Build and maintain evaluation datasets and scenarios for each of AssureAI's 19 Trustworthiness checks (explainability, audit provenance, regulation coverage, safety guardrails) as new agents come online.
-
Run scheduled and on-commit bias, toxicity, red-team, and explainability (SHAP/LIME) suites; triage failures and route them to the right owner (AI Engineer, Integration Specialist, or Governance Engineer).
-
Maintain the CI/CD threshold-gate configuration so a failing Trustworthiness check blocks release rather than just flagging it.
-
Package audit evidence — run provenance, evidence exportability, regulation-coverage reports — for the AI Trust & Compliance Engineer's CSG review-board submissions.
-
Track dataset coverage and flag gaps as agent scope expands into new domains, data types, or regulatory contexts.
-
Split coverage with the second AssureAI Engineer across agent domains or pipeline stages so evaluation throughput scales with agent count.
-
5-7 years in QA/test engineering for ML or GenAI systems, ideally with a dedicated eval framework (DeepEval, Ragas, Promptfoo, or comparable).
-
Hands-on experience with AssureAI or a directly comparable AI-assurance/evaluation platform.
-
Working knowledge of bias/fairness testing, explainability techniques (SHAP/LIME), and red-teaming/adversarial-prompt methodology.
-
Comfortable reading AI regulatory requirements (EU AI Act, NIST AI RMF, ISO/IEC 42001) well enough to translate them into test coverage.
-
Proficient in Python; comfortable wiring test suites into CI/CD (Cloud Build, GitHub Actions).
-
Detail-oriented and comfortable owning a queue of failing checks across multiple agents at once.
-
Experience with Gemini Enterprise / ADK agents specifically, or another enterprise agent platform.
-
Familiarity with RAG evaluation (faithfulness, hallucination, context precision/recall).
-
Prior audit or compliance-adjacent work (SOC 2, ISO 27001, or similar) that makes evidence packaging second nature.
By month two: every live agent has an active Trustworthiness evaluation suite running on every commit, with clear pass/fail thresholds.
By month four: audit-evidence packages are produced on a standing cadence without ad hoc requests, and dataset coverage gaps are tracked and closed as new agents onboard.