Software Engineer- AI Datacenter Orchestration Platform

  • R&D
  • Tel Aviv, Israel
  • Full-time

Description

The Role

We're building a new orchestration platform for AI datacenters—software that turns a complex, multi-vendor stack of compute, networking, and accelerators into something customers can deploy, operate, and optimize with confidence.


You'll join a small, growing engineering team where every engineer ships. In this hands-on development role, you'll build backend services, contribute to system design, and own features from implementation through production, working closely with experienced engineers.


If you enjoy solving challenging engineering problems, building new products, and seeing your work run at scale in real AI datacenters, this role is for you.


What You'll Do


- Build and maintain core services and features for an orchestration platform spanning compute, networking, and accelerators.

- Translate customer requirements into APIs, workflows, and data models in collaboration with the team.

- Deliver features end-to-end, from prototyping and implementation to testing, deployment, and production support.

- Write reliable, maintainable code and contribute to shared libraries, automated tests, and CI/CD pipelines.

- Contribute to design discussions and make practical implementation tradeoffs around performance, reliability, and simplicity.

- Participate in code reviews, share knowledge, and help improve engineering practices.

- Troubleshoot issues and improve service performance, observability, and operational readiness.



Why This Role Is Different


You'll help build a new platform from its early stages, with meaningful ownership of features and a direct impact on the product. Working alongside experienced engineers, you'll contribute to the platform's foundation, deepen your distributed systems expertise, and ship software built for real-world scale.


Requirements

-5+ years of professional software development experience**, including building and maintaining production backend services.

- Strong backend development skills, preferably in Python, and experience building REST and/or GraphQL APIs.

- Practical understanding of distributed systems concepts, including scalability, resilience, and failure handling.

- Experience with SQL and/or NoSQL databases, data modeling, and query performance.

- Experience writing automated tests and working with Git and CI/CD workflows.

- Familiarity with multithreading, locks and synchronization mechanisms, performance analysis and optimization.

- Ability to take ownership of features, communicate clearly, and collaborate effectively within a small team.


Nice to Have


- Familiarity with TypeScript, React, and modern web application architecture.

- Experience with cloud platforms (AWS, GCP, or Azure), containers, and Kubernetes.

- Experience with observability tools such as OpenTelemetry, Prometheus, Grafana, or Datadog.

- Familiarity with secure coding practices and common application security risks.

- Experience with orchestration systems, provisioning pipelines, cluster management, or schedulers.

- Exposure to AI/GPU infrastructure or high-performance networking.

Apply to this job