Multimodal AI Systems Architect (AI Engineering)
We are seeking a talented Multimodal AI Systems Architect to develop and optimize AI systems that seamlessly integrate vision and audio models. This role focuses on enhancing our voice-to-voice interactions and multimodal retrieval capabilities, ensuring our systems are efficient and innovative.
Responsibilities:
- Integrate vision encoders and audio-native models into core agent reasoning loops.
- Optimize streaming latency for voice-to-voice AI interactions.
- Architect multimodal RAG systems capable of retrieving insights from videos and PDFs.
Qualifications:
- Experience with Whisper, CLIP, and multimodal LLM integration.
- Knowledge of streaming architectures and WebRTC.
- Expertise in cross-modal alignment.
As published by greenhouse
First Name, Last Name, Email, Phone, Resume/CV, Cover Letter
- Preferred First Name optional
- Website optional
- LinkedIn Profile optional
- Notice Period
- Current Annual Salary (with Currency)
- Expected Annual Salary (with Currency)
- Working Location choose any
- Do you have any Web3 experience? choose one
- Web3 Vertical Experience choose any
- Any personal experience in Web3 (e.g. side project, personal investment) if no professional experience. written answer