Sr. Software Engineer - Ingestion Core teamNew$166K

The opportunity

Deeply understanding what’s in the enterprise data has been a challenge that Databricks has been addressing by providing analytics and machine learning tools. From data warehousing with Databricks SQL to large-scale distributed processing with Spark and advanced ML tools for…

What you'll do

  • Build distributed infrastructure to ingest data from diverse sources and: support streaming ingestion, incremental processing, and replication. This isn’t just about building plugin connectors.

  • Reduce end-to-end latency, increase throughput, and reduce costs from the: time data appears in source systems to when it is available in Delta Lake.

  • Design and optimize streaming and distributed workloads for throughput, cost, latency, reliability, and scale.

  • Optimize streaming workloads by exploring and applying ML techniques.

  • Build monitoring and observability capabilities (customer-facing and: internal) that provide visibility into ingestion workflows and the systems running them.

  • Collaborate with partner teams to enable use cases like RAG and AI agents.

What they're looking for

  • Build Agentic SDKs with strong evaluations.
  • + years of experience writing production code in one of: Java, Scala, Go, C++, or Python.
  • Experience architecting, developing, and deploying large-scale distributed and asynchronous systems.
  • Experience with distributed systems, streaming, Spark, databases, data processing, or CDC.