The opportunity
At Chime, we believe that everyone can achieve financial progress. We created Chime—a financial technology company, not a bank*—on the premise that core banking services should be helpful, easy, and free.
What you'll do
Design, build, and operate scalable ML and AI infrastructure on AWS.
Design and operate shared platform capabilities for LLM and agentic: workloads, including model access, prompt and configuration lifecycle, retrieval, tool integration, state management, and workflow orchestration.
Build evaluation frameworks for non-deterministic AI systems, including: offline benchmarks, regression testing, online quality signals, human feedback, and failure analysis.
Establish observability, reliability, and governance for models and agents,: covering traces, model and prompt versions, tool calls, latency, token usage, quality, safety, privacy, and cost.
Help teams make principled architecture decisions across traditional ML,: LLM-powered applications, and agentic workflows, and contribute to the platform’s technical roadmap.
Build distributed training, batch inference, and large-scale processing: systems using frameworks such as Ray or Spark.
What they're looking for
- Build and maintain infrastructure as code using Terraform.
- Support and evolve the feature store and feature pipelines.
- Develop data ingestion and streaming systems using technologies such as Kinesis, Kafka, Flink, or Spark.
- Improve CI/CD workflows for ML models, AI applications, and platform components.