Software Engineer, Kernel Programming

About FuriosaAI

FuriosaAI builds high-performance, high-efficiency AI compute for the Inference Era. Founded in 2017 by veteran semiconductor and AI algorithm engineers, Furiosa operates globally with offices in Korea and Silicon Valley, along with a compiler-focused R&D lab in Lisbon.

Our vision is to make AI computing sustainable, enabling access to powerful AI for everyone on Earth. We solve the AI hardware energy and operational cost crisis at the architectural level, rather than through brute force, building the world's first truly AI-native compute platform to unlock the full potential of artificial intelligence for every enterprise.

About the job

Lead the integration of diverse AI models including VLA, Vision, and Multimodal architectures by utilizing our kernel programming language to ensure both accuracy and performance while keeping the stack ready for developers to use.


Responsibilities

  • Design and implement efficient kernels on FuriosaAI’s kernel programming stack (including vISA, TCL), targeting Tensor Contract Processor (TCP) architectures.

  • Diagnose and optimize kernel performance with profiling tools and roofline analysis for each RNGD-accelerated AI model.

  • Develop and apply automated kernel generation and optimization for AI workloads.

  • Build diagnostic tools or testbeds for robust and reliable kernel validation.

  • Drive end-to-end programming enablement on RNGDs, creating reproducible guides and reference implementations.

Minimum Qualifications

  • BS in Computer Science, Artificial Intelligence, Electrical Engineering, or a related field.

  • Experience in low-level systems programming targeting XPU (e.g., NPU, GPU) architectures.

  • Experience collaborating across engineering, research, and product teams to align software development with product requirements.

Preferred Qualifications

  • MS or PhD in Computer Science, Artificial Intelligence, Electrical Engineering, or a related field.

  • Experience in optimizing high-performance kernels on AI accelerators (e.g., GPU, TPU) for AI products.

  • Understanding of XPU architecture (computation patterns, data movement) and software-hardware co-optimization strategies.

  • Experience in open-source or research projects on AI model architectures such as Diffusion, Mamba, and VLA.

  • Experience in designing efficient deep learning architectures and developing algorithms for AI applications.

Contact

What this application asks

greenhouse

First Name, Last Name, Email, Phone, Resume/CV, Cover Letter, Location

  • Preferred First Name optional
  • Desired Job Type choose one
  • Career Summary written answer
  • What is your expected annual salary for this role?
  • Do you have authorization to work in the country and/or state where the job is located? choose one
  • Do you require sponsorship for employment visa status in the country in the present or in the future? choose one
  • LinkedIn Profile optional
  • Website optional
  • What is your current annual salary (base + fixed bonus)?
  • When Would You Be Available to Start?

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available