Software Engineer

Summary

Designs and builds AI-powered data ingestion and processing pipelines for a regulatory intelligence SaaS platform. Day-to-day work centers on Python (FastAPI/Django), web crawling (Scrapy/Playwright), LLM integration (LangChain/LlamaIndex), and deploying AI/ML services into production.

CUBE are a global RegTech business defining and implementing the gold standard of regulatory intelligence for the financial services industry. We deliver our services through intuitive SaaS solutions, powered by AI, to simplify the complex and everchanging world of compliance for our clients.

Why us?

🌍 CUBE is a globally recognized brand at the forefront of Regulatory Technology. Our industry-leading SaaS solutions are trusted by the world’s top financial institutions globally.

🚀 In 2024, we achieved over 50% growth, both organically and through two strategic acquisitions. We’re a fast-paced, high-performing team that thrives on pushing boundaries—continuously evolving our products, services, and operations. At CUBE, we don’t just keep up we stay ahead.

🌱 We believe our future is built by bold, ambitious individuals who are driven to make a real difference. Our “make it happen” culture empowers you to take ownership of your career and accelerate your personal and professional development from day one.

🌐 With over 700 CUBERs across 19 countries spanning EMEA, the Americas, and APAC, we operate as one team with a shared mission to transform regulatory compliance. Diversity, collaboration, and purpose are the heartbeat of our success.

💡 We were among the first to harness the power of AI in regulatory intelligence, and we continue to lead with our cutting-edge technology. At CUBE, You will work alongside some of the brightest minds in AI research and engineering in developing impactful solutions that are reshaping the world of regulatory compliance.

Key Deliverables — What Success Looks Like

Deliver scalable AI-enabled data ingestion and processing services.

Build reliable pipelines for extracting, transforming, and processing large volumes of structured and unstructured data.

Successfully deploy AI/ML models or LLM-powered services into production.

Ensure high system availability, performance, and operational visibility.

Deliver clean, maintainable, well-tested, and production-ready code.

About the Role:

Design, develop, and maintain scalable AI-powered data ingestion and processing platforms using Python. Build intelligent data pipelines, web crawling solutions, and AI/ML services that process structured and unstructured data while ensuring scalability, performance, and reliability.

Key Responsibilities:

  • Design and develop web crawling, data ingestion, and AI-powered processing solutions using Python.

  • Build scalable data pipelines for processing HTML, PDF, and other structured/unstructured data sources.

  • Develop REST APIs to support AI/ML services and data processing workflows.

  • Implement AI/ML models or integrate LLMs for document parsing, classification, summarization, and information extraction.

  • Design and optimize background processing, task scheduling, and distributed workloads.

  • Build monitoring, logging, and observability for AI pipelines and backend services.

  • Collaborate with Product, Data Science, DevOps, and Engineering teams to deliver production-ready AI solutions.

First 90 Days — Objectives

Day 30: Understand the AI platform, data processing architecture, and development standards.

Deliver enhancements to backend services or AI processing pipelines.

Day 60: Independently build and deploy AI-powered data processing components and APIs.

Integrate AI/ML models into production workflows with monitoring and logging.

Day 90: Own end-to-end delivery of AI-enabled backend services and data pipelines.

Optimize model inference, pipeline performance, and system reliability.

Required Skills & Experience:

  • 2-3 years of Python development experience with FastAPI, Django, or similar frameworks.

  • Strong experience in AI/ML, Generative AI, or LLM application development.

  • Experience with web scraping frameworks such as Scrapy or Playwright.

  • Hands-on experience with NLP, document processing, or information extraction techniques.

  • Experience integrating LLMs (OpenAI, Claude, Gemini, Llama, or similar) using frameworks like LangChain or LlamaIndex.

  • Strong knowledge of REST APIs, asynchronous programming, and distributed task processing (Celery, RabbitMQ, Kafka).

  • Experience parsing HTML, XML, and PDF documents using BeautifulSoup, lxml, or PyMuPDF.

  • Strong understanding of data structures, algorithms, and scalable data processing.

  • Experience with PostgreSQL, vector databases, or search platforms such as OpenSearch/Elasticsearch.

  • Experience with Docker, CI/CD, cloud platforms (AWS/Azure/GCP), and application monitoring.

Nice to Have:

Experience with RAG (Retrieval-Augmented Generation) architectures.

Knowledge of vector databases such as Pinecone, Weaviate, Milvus, or pgvector.

Experience with Hugging Face Transformers, PyTorch, or TensorFlow.

Exposure to MLOps tools such as MLflow, Kubeflow, or model serving frameworks.

Experience with Kubernetes and cloud-native deployments.

Knowledge of OCR, computer vision, or multimodal AI models.

Interested?

If you are passionate about leveraging technology to transform regulatory compliance and meet the qualifications outlined above, we invite you to apply. Please submit your resume detailing your relevant experience and interest in CUBE.​

CUBE is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

What this application asks

ashby

Name, Email, Resume

  • Current Location
  • Current Salary
  • Expected Salary
  • Notice Period
  • Total years of Experience

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available