The opportunity
This is a backend-leaning role focused on scaling systems. Billing and fraud will be a focus, but your remit spans every high-throughput system at Railway — workers, queues, event pipelines, and the databases underneath them.
What you'll do
Architect and scale the pipelines that turn raw usage into accurate,: real-time billing — metering, aggregation, rating, and invoicing across millions of events, from ingestion in ClickHouse to the rating engine.
Build payment flows that are correct under concurrency and partial failure:: idempotent charges, retries, reconciliation, and clean handling of provider edge cases (Stripe and beyond).
Develop fraud and abuse detection: signal collection, real-time scoring, automated mitigation — that protects platform margin without getting in legitimate users' way.
Scale the systems everything else depends on: Postgres under heavy write load, Node.js services under pressure, and long-running workflows orchestrated with Temporal where exactly-once semantics and durability actually matter.
Build TypeScript + GraphQL APIs where correctness and auditability are non-negotiable.
Write Engineering Requirement Documents to take something from idea, to: defined tasks, to implementation, to monitoring its success and scaling it further.
What they're looking for
- Contribute to our open-source repositories (CLI, Typescript SDK, Railpack,: etc.) — Rust experience, or the desire to learn it, helps here.
- Be oncall from time to time.
- Re-architect billing end-to-end: per-second usage metering at platform scale, idempotent payment processing that survives provider outages without double-charging, and credit, prepayment, and enterprise-invoicing models that hold up under audit.
- Stand up a fraud-detection service that scores signups and deployments in: real time and automatically throttles abuse (crypto mining, free-tier farming, stolen cards).