#EG AI / LLM Specialist

This role is an AI/LLM Engineer focused on prompt engineering, model integration, evaluation, and production quality within NCS AI Central’s Forward Deployed Engineering model. This role helps take GenAI solutions from POC/POV to production by building prompts, integrating foundation models, creating automated evaluation and benchmarking frameworks, and continuously monitoring quality, safety, hallucination, bias, drift, latency, and cost.

You will work closely with AI Architects, Solution Architects, AI Engineers, and Testers to provide evidence-based model recommendations, support production readiness, and maintain reusable internal assets such as prompt libraries, evaluation templates, and benchmark datasets.

What will you do?

1. Model Integration & Prompt Engineering

  • Design, test, and optimise prompts and prompt chains for production use cases — balancing accuracy, latency, and cost.
  • Integrate foundation models into applications via APIs and gateways; advise on model/version selection for a given use case alongside AI Architects.
  • Support light fine-tuning and instruction-tuning work (LoRA/PEFT and similar techniques) where a use case calls for it, in partnership with AI Engineers.

2. Evaluation Framework & Benchmarking

  • Design and maintain evaluation harnesses (golden datasets, benchmark suites) that measure LLM/agentic system accuracy, consistency, and safety.
  • Run comparative model benchmarking (accuracy, latency, cost-per-query) to feed technical evidence into model-selection decisions led by AI/Solution Architects.
  • Build repeatable, automated regression suites that run on every prompt, model, or pipeline change, wired into CI.

3. Quality & Adversarial Testing

  • Proactively red-team AI systems — adversarial prompting, edge-case and jailbreak testing — to surface failure modes before clients do.
  • Detect and quantify hallucination, bias, and drift in production and pre-production systems, producing clear, defensible metrics (not qualitative impressions).

4. Governance & Reporting

  • Feed evaluation evidence into the PRR (Production Readiness Review) Evaluation & Quality pillar, supporting engagement teams at the Scale gate.
  • Maintain evaluation and prompt-version documentation and reporting standards that align with client compliance and audit needs (e.g., government AI governance requirements).

5. FDE & Development/Maintenance Coverage

  • During FDE engagements: rapidly prototype prompts and model integrations, and stand up lightweight evaluation harnesses to compare candidate models/approaches during POC/POV, giving the team fast, evidence-based go/no-go signals.
  • During system development & maintenance engagements: own ongoing prompt/model tuning and run continuous evaluation and regression monitoring on live production systems, flagging quality degradation over time.
  • Contribute reusable prompt libraries, evaluation templates, and benchmark datasets back into the shared internal asset library for reuse across engagements.

6. Collaboration

  • Provide technical benchmark evidence and integration recommendations to AI Architects and Solution Architects, who own the final client-facing model recommendation and proposal.
  • Partner with AI Engineers and Testers to distinguish functional QA (does it work) from output-quality evaluation (is it right), and to hand off tuned prompts/models cleanly into production builds.

The ideal candidate should possess:

  • 3+ years working hands-on with LLMs across prompt engineering, model integration, and evaluation — not evaluation alone.
  • Practical experience designing and optimising prompts and prompt chains for production applications, and integrating models via APIs/gateways.
  • Strong grasp of evaluation methodologies — accuracy/hallucination/toxicity metrics, human-in-the-loop evaluation, A/B testing.
  • Hands-on scripting ability (Python) to build and automate evaluation harnesses and integration/testing pipelines.
  • Statistical literacy — able to design a representative test/benchmark set and interpret results rigorously, not anecdotally.
  • Clear, structured written communication — able to translate evaluation results and model recommendations into a defensible report for both engineering and client audiences.
  • Working knowledge of the China AI model/tech stack (e.g., DeepSeek, Qwen, GLM, Kimi, MiniMax) — deployment patterns, licensing, and self-hosting requirements.

Preferred Qualifications

  • Hands-on fine-tuning/instruction-tuning experience (LoRA/PEFT or similar) on open-weight models.
  • Experience with LLM evaluation tooling (RAGAS, DeepEval, TruLens, promptfoo) or building custom eval frameworks.
  • Exposure to red-teaming/adversarial testing practices for generative AI systems.
  • Familiarity with regulated-sector AI governance expectations (Healthcare, Government, Financial Services).
  • Prior experience supporting client-facing presales or solutioning conversations with technical evidence (without owning the proposal).
  • Hands-on benchmarking or integration experience with Chinese open-weight models (DeepSeek, Qwen, GLM) alongside Western models.

Tech Stack (Illustrative)

  • Languages: Python (primary), SQL
  • Prompt & Integration: LangChain/LlamaIndex, model gateways (LiteLLM, Bedrock, Azure OpenAI), prompt-versioning tools
  • Fine-Tuning: LoRA/PEFT, Hugging Face Transformers (where applicable)
  • Eval Tooling: RAGAS, DeepEval, TruLens, promptfoo, custom harnesses
  • LLM Runtime: OpenAI/Azure OpenAI/Bedrock/Vertex APIs; DeepSeek/Qwen/GLM (China stack)
  • Data & Reporting: Pandas, Jupyter, BI/reporting tools for evaluation dashboards
  • CI Integration: GitHub Actions/GitLab CI for automated regression evaluation

Why Join NCS?

Grow with Us

  • Work on cutting-edge AI products that shape the future of technology
  • Collaborate with talented, passionate teams across research, engineering, and design
  • Access continuous learning opportunities and career development pathways

Make an Impact

  • Transform AI research into products that solve real problems for clients and users
  • Drive innovation in a leading Technology Services Firm with regional presence
  • Contribute to building a better future through responsible, human-centred AI

Thrive in Our Culture

  • Experience a human-to-human approach where relationships and collaboration matter
  • Be part of Team NCS, where bold ideas meet practical execution
  • Enjoy a supportive environment that values diversity, inclusion, and respect

We are driven by our AEIOU beliefs—Adventure, Excellence, Integrity, Ownership, and Unity—and we seek individuals who embody these values in both their professional and personal lives. We are committed to our Impact: Valuing our clients, Growing our people, and Creating our future.

Together, we make the extraordinary happen.

Learn more about us at ncs.co and visit our LinkedIn career site.

Scam Alert

We are aware of fraudulent job offers and impersonations of NCS recruiters. Phishing emails using convincing-looking but fake addresses are also commonly used to trick you into thinking that they come from official NCS sources.

Please note that all official communications from NCS Group will only be sent from verified corporate email addresses. Always check that the sender’s email address ends with the genuine NCS domain, @ncs.com.sg and beware of extra letters, symbols or misspellings. When in doubt, verify the sender’s identity by contacting us at [email protected].

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available