The opportunity
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more.
What you'll do
Continuously invest in documentation, metrics, monitors and other troubleshooting tools
Participate in on-call rotations during business hours and occasional: weekends. This is a challenging yet rewarding opportunity to help remediate the most pressing issues across the Palantir fleet.
Diagnose, resolve, and prevent issues encountered in the field. Deliver: end-to-end improvements to core products based on these issues you encounter in the field.
Improve observability by refactoring codepaths and introducing telemetry
Identify and implement data-driven opportunities for improved service resilience
Develop strategic opinions on stability investments and inform the vision for long-term product stability
What they're looking for
- Comfortable with and curious about large scale production systems and: technologies. For example, load balancing, monitoring, distributed systems, and configuration management.
- Confidence in troubleshooting complex issues independently using observability tools and stack traces
- Familiarity with monitoring tools such as Prometheus and health checks
- Experience coding with Java, Go and/or web technologies (e.g. HTML, CSS,: JavaScript, Python/Ruby, Django/Flask/Ruby on Rails, etc.) is a plus