The opportunity
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.
What you'll do
Design and develop cutting-edge multimodal AI systems, integrating various: modalities such as text, speech, and vision.
Conduct research and experiments on our advanced compute infrastructure,: exploring novel ideas in multimodal representation learning, transfer learning, and more.
Collaborate closely with our world-class teams, learning from and: contributing to their expertise in the field.
Possess exceptional software engineering skills, with a proven track record: of building robust and scalable systems.
Have a strong command of Python and are well-versed in popular deep learning: frameworks like JAX, PyTorch, and TensorFlow, with an understanding of their multimodal capabilities.
Knowledge of distributed training strategies, especially for large-scale multimodal models.
What they're looking for
- Familiarity with autoregressive models, particularly their application in: multimodal tasks such as image or video captioning, speech-to-text generation.
- Bonus: Publications in top-tier venues demonstrating your expertise in multimodal AI research.
- Bonus: Experience in writing efficient GPU kernels using CUDA, optimising performance for multimodal tasks.
- Have a deep passion for machine learning and its potential to impact various: industries through multimodal applications.