The opportunity
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.
What you'll do
Maintain large-scale pipelines for processing web corpora.
Work on filtering and quality-scoring systems to identify high-value web documents.
Analyze web data composition across domains, languages and time periods.
Develop and maintain highly-performant deduplication pipelines.
Collaborate with cross-functional teams, including researchers and engineers,: to ensure data pipelines meet the demands of cutting-edge language models.
Strong software engineering skills, with proficiency in Python and experience building data pipelines.
What they're looking for
- Familiarity with data processing frameworks such as Apache Spark, Apache Beam, Pandas, or similar tools.
- Experience working with large-scale web datasets.
- Knowledge of data quality assessment techniques and experimentation with data mixtures.
- A passion for bridging research and engineering to solve complex data-related challenges in AI model training.