Member of Technical Staff (Software Engineer, Data Platform)Active$220K–$405K

The opportunity

The Data Platform team owns the end-to-end data lifecycle at Perplexity, from ingestion through processing, storage, and serving, powering product features, analytics, experimentation, AI workloads, and the company’s data lake.

What you'll do

  • Design and operate large-scale batch and streaming data pipelines that: directly power Perplexity product features, AI training and evaluation workflows, analytics, and experimentation.

  • Build event-driven and streaming systems (Kafka, Kinesis, PubSub, or similar): for real-time ingestion, transformation, and delivery, alongside batch frameworks for backfills, aggregations, and offline computation.

  • Lead the architecture of data orchestration using tools like Airflow or: Dagster, owning scheduling, dependency management, retries, SLAs, and end-to-end observability for critical data flows.

  • Set and enforce guarantees for data correctness, freshness, lineage, and: recoverability, designing systems that handle rapid scale growth, partial failures, and evolving schemas without disrupting AI workloads or product experiences.

  • Build self-serve data platforms that let engineers, data scientists, and: analysts safely discover data, define contracts, and create and operate their own pipelines with minimal friction.

  • Improve developer experience through better abstractions, opinionated paved: paths, and standards for data modeling, testing, validation, and deployment, treating the data platform as a product used by many teams.

What they're looking for

  • + years (Senior) or 8+ years (Staff) of software engineering experience.
  • Strong experience building production data infrastructure systems.
  • Hands-on experience with batch and/or streaming data processing at scale.
  • Deep familiarity with data orchestration systems (Airflow, Dagster, or similar).
  • Proficiency in Python and at least one additional backend language (Go, TypeScript, etc.).
  • Strong systems thinking around reliability, latency, cost, and complexity tradeoffs.
  • Experience supporting ML/AI workflows, training pipelines, or evaluation systems.
  • Familiarity with data quality, lineage, observability, and governance tooling.