Assure AI Engineer

Gemini – AssureAI Engineer

2 positions

Keywords: AssureAI · Evaluation Datasets · Red-Teaming · Bias & Explainability (SHAP/LIME) · Threshold Gates · CI/CD · Audit Evidence · Python

About the Role

You are one of two hands-on operators of AssureAI's Trustworthiness pillar inside the Governance Control Tower. Where the AI Trust & Compliance Engineer sets the strategy — which regulations map to which checks, what a passing threshold means — you build and run the evaluation suites that prove it, day in and day out, across a growing roster of Gemini-based agents. You report to the AI Trust & Compliance Engineer and work closely with the AI Engineers building the agents you test.

What You'll Own

  • Build and maintain evaluation datasets and scenarios for each of AssureAI's 19 Trustworthiness checks (explainability, audit provenance, regulation coverage, safety guardrails) as new agents come online.

  • Run scheduled and on-commit bias, toxicity, red-team, and explainability (SHAP/LIME) suites; triage failures and route them to the right owner (AI Engineer, Integration Specialist, or Governance Engineer).

  • Maintain the CI/CD threshold-gate configuration so a failing Trustworthiness check blocks release rather than just flagging it.

  • Package audit evidence — run provenance, evidence exportability, regulation-coverage reports — for the AI Trust & Compliance Engineer's CSG review-board submissions.

  • Track dataset coverage and flag gaps as agent scope expands into new domains, data types, or regulatory contexts.

  • Split coverage with the second AssureAI Engineer across agent domains or pipeline stages so evaluation throughput scales with agent count.

What We're Looking For

  • 5-7 years in QA/test engineering for ML or GenAI systems, ideally with a dedicated eval framework (DeepEval, Ragas, Promptfoo, or comparable).

  • Hands-on experience with AssureAI or a directly comparable AI-assurance/evaluation platform.

  • Working knowledge of bias/fairness testing, explainability techniques (SHAP/LIME), and red-teaming/adversarial-prompt methodology.

  • Comfortable reading AI regulatory requirements (EU AI Act, NIST AI RMF, ISO/IEC 42001) well enough to translate them into test coverage.

  • Proficient in Python; comfortable wiring test suites into CI/CD (Cloud Build, GitHub Actions).

  • Detail-oriented and comfortable owning a queue of failing checks across multiple agents at once.

Nice to Have

  • Experience with Gemini Enterprise / ADK agents specifically, or another enterprise agent platform.

  • Familiarity with RAG evaluation (faithfulness, hallucination, context precision/recall).

  • Prior audit or compliance-adjacent work (SOC 2, ISO 27001, or similar) that makes evidence packaging second nature.

What Success Looks Like

By month two: every live agent has an active Trustworthiness evaluation suite running on every commit, with clear pass/fail thresholds.

By month four: audit-evidence packages are produced on a standing cadence without ad hoc requests, and dataset coverage gaps are tracked and closed as new agents onboard.


Gemini – AssureAI Engineer

2 positions

Keywords: AssureAI · Evaluation Datasets · Red-Teaming · Bias & Explainability (SHAP/LIME) · Threshold Gates · CI/CD · Audit Evidence · Python

About the Role

You are one of two hands-on operators of AssureAI's Trustworthiness pillar inside the Governance Control Tower. Where the AI Trust & Compliance Engineer sets the strategy — which regulations map to which checks, what a passing threshold means — you build and run the evaluation suites that prove it, day in and day out, across a growing roster of Gemini-based agents. You report to the AI Trust & Compliance Engineer and work closely with the AI Engineers building the agents you test.

What You'll Own

  • Build and maintain evaluation datasets and scenarios for each of AssureAI's 19 Trustworthiness checks (explainability, audit provenance, regulation coverage, safety guardrails) as new agents come online.

  • Run scheduled and on-commit bias, toxicity, red-team, and explainability (SHAP/LIME) suites; triage failures and route them to the right owner (AI Engineer, Integration Specialist, or Governance Engineer).

  • Maintain the CI/CD threshold-gate configuration so a failing Trustworthiness check blocks release rather than just flagging it.

  • Package audit evidence — run provenance, evidence exportability, regulation-coverage reports — for the AI Trust & Compliance Engineer's CSG review-board submissions.

  • Track dataset coverage and flag gaps as agent scope expands into new domains, data types, or regulatory contexts.

  • Split coverage with the second AssureAI Engineer across agent domains or pipeline stages so evaluation throughput scales with agent count.

What We're Looking For

  • 5-7 years in QA/test engineering for ML or GenAI systems, ideally with a dedicated eval framework (DeepEval, Ragas, Promptfoo, or comparable).

  • Hands-on experience with AssureAI or a directly comparable AI-assurance/evaluation platform.

  • Working knowledge of bias/fairness testing, explainability techniques (SHAP/LIME), and red-teaming/adversarial-prompt methodology.

  • Comfortable reading AI regulatory requirements (EU AI Act, NIST AI RMF, ISO/IEC 42001) well enough to translate them into test coverage.

  • Proficient in Python; comfortable wiring test suites into CI/CD (Cloud Build, GitHub Actions).

  • Detail-oriented and comfortable owning a queue of failing checks across multiple agents at once.

Nice to Have

  • Experience with Gemini Enterprise / ADK agents specifically, or another enterprise agent platform.

  • Familiarity with RAG evaluation (faithfulness, hallucination, context precision/recall).

  • Prior audit or compliance-adjacent work (SOC 2, ISO 27001, or similar) that makes evidence packaging second nature.

What Success Looks Like

By month two: every live agent has an active Trustworthiness evaluation suite running on every commit, with clear pass/fail thresholds.

By month four: audit-evidence packages are produced on a standing cadence without ad hoc requests, and dataset coverage gaps are tracked and closed as new agents onboard.


Gemini – AssureAI Engineer

2 positions

Keywords: AssureAI · Evaluation Datasets · Red-Teaming · Bias & Explainability (SHAP/LIME) · Threshold Gates · CI/CD · Audit Evidence · Python

About the Role

You are one of two hands-on operators of AssureAI's Trustworthiness pillar inside the Governance Control Tower. Where the AI Trust & Compliance Engineer sets the strategy — which regulations map to which checks, what a passing threshold means — you build and run the evaluation suites that prove it, day in and day out, across a growing roster of Gemini-based agents. You report to the AI Trust & Compliance Engineer and work closely with the AI Engineers building the agents you test.

What You'll Own

  • Build and maintain evaluation datasets and scenarios for each of AssureAI's 19 Trustworthiness checks (explainability, audit provenance, regulation coverage, safety guardrails) as new agents come online.

  • Run scheduled and on-commit bias, toxicity, red-team, and explainability (SHAP/LIME) suites; triage failures and route them to the right owner (AI Engineer, Integration Specialist, or Governance Engineer).

  • Maintain the CI/CD threshold-gate configuration so a failing Trustworthiness check blocks release rather than just flagging it.

  • Package audit evidence — run provenance, evidence exportability, regulation-coverage reports — for the AI Trust & Compliance Engineer's CSG review-board submissions.

  • Track dataset coverage and flag gaps as agent scope expands into new domains, data types, or regulatory contexts.

  • Split coverage with the second AssureAI Engineer across agent domains or pipeline stages so evaluation throughput scales with agent count.

What We're Looking For

  • 5-7 years in QA/test engineering for ML or GenAI systems, ideally with a dedicated eval framework (DeepEval, Ragas, Promptfoo, or comparable).

  • Hands-on experience with AssureAI or a directly comparable AI-assurance/evaluation platform.

  • Working knowledge of bias/fairness testing, explainability techniques (SHAP/LIME), and red-teaming/adversarial-prompt methodology.

  • Comfortable reading AI regulatory requirements (EU AI Act, NIST AI RMF, ISO/IEC 42001) well enough to translate them into test coverage.

  • Proficient in Python; comfortable wiring test suites into CI/CD (Cloud Build, GitHub Actions).

  • Detail-oriented and comfortable owning a queue of failing checks across multiple agents at once.

Nice to Have

  • Experience with Gemini Enterprise / ADK agents specifically, or another enterprise agent platform.

  • Familiarity with RAG evaluation (faithfulness, hallucination, context precision/recall).

  • Prior audit or compliance-adjacent work (SOC 2, ISO 27001, or similar) that makes evidence packaging second nature.

What Success Looks Like

By month two: every live agent has an active Trustworthiness evaluation suite running on every commit, with clear pass/fail thresholds.

By month four: audit-evidence packages are produced on a standing cadence without ad hoc requests, and dataset coverage gaps are tracked and closed as new agents onboard.


See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available