The opportunity
Binance is building one of the largest internal AI agent fleets in the industry — hundreds of sandboxed agents powering automation across trading, compliance, customer service, risk, and beyond. This role sits at the core of that platform: you'll build the infrastructure and…
What you'll do
Build, publish, and maintain OpenClaw skills: modular capability units used by hundreds of agents across the org to automate repetitive work and unlock new business capabilities
Develop CLI tooling for agent operations: deployment, diagnostics, session management, skill registry, and developer workflows
Own end-to-end AI agent harness engineering: lifecycle management, tool execution, context/session tuning, compaction strategies, model routing
Instrument the agent fleet with data pipelines and dashboards; apply data: science techniques to understand token efficiency, failure modes, latency distribution, and business outcome correlation
Identify bottlenecks across the platform and drive measurable improvements in: agent throughput, response quality, and cost efficiency — directly supporting user growth and retention
Track the research frontier: papers, open-source releases, community developments — and rapidly prototype integrations (new model capabilities, reasoning techniques, agentic frameworks)
What they're looking for
- + years in software/platform engineering; 2+ year hands-on with LLM or AI agent systems in production
- AI Native mindset: you default to AI-assisted development, think natively in agent/tool/context primitives, and are allergic to doing manually what an agent could do
- Skill & CLI development: experience building modular, composable tools or CLI utilities for developer platforms; TypeScript and/or Python fluency
- Agent harness engineering: practical experience with OpenClaw, LangGraph, AutoGen, CrewAI, or equivalent orchestration runtimes
- LLM infrastructure: token management, model routing, context compaction, cost optimization at scale
- Data science capability: comfortable with log analysis, statistical profiling, SQL/Python for usage data; can translate raw telemetry into actionable insights that drive platform decisions
- Research awareness: follows model releases, agent framework updates, and relevant literature; can quickly assess what's worth integrating and what isn't
- Vibe coding: ships fast using AI-assisted workflows; iterative, pragmatic, high output-to-noise ratio with strong engineering fundamentals