Description
Hybrid | Israel
About the Company
DriveNets is a leader in large-scale networking solutions for AI infrastructure and service providers. The company's disaggregated networking architecture transforms the economics of large-scale infrastructures while maximizing performance, utilization, and operational efficiency. Its high-performance AI fabric maximizes GPU utilization and accelerates deployments by optimizing the AI stack end-to-end, resulting in higher tokens-per-second and lower cost-per-token. DriveNets' solutions power production networks for global tier-1 operators like AT&T and Comcast, and scale multi-vendor AI infrastructures at foundation model labs, NeoClouds, and enterprises.
Responsibilities
- Deploy and tune vLLM inference servers across customer environments on diverse GPU hardware (H100, A100, L40S, MI300X), including fully air-gapped, offline deployments.
- Optimize inference performance end-to-end - batching, tensor/pipeline parallelism, KV-cache and prefix caching, quantization (FP16/INT8/AWQ/GPTQ), and hardware-specific tuning across NVIDIA and AMD architectures.
- Maintain the LiteLLM gateway and Helm chart variations across Kubernetes flavors (OpenShift, EKS, AKS, GKE, bare metal).
- Own the model lifecycle - registry, canary rollouts, upgrade/rollback playbooks, and benchmarking of new open-weight models against our security-reasoning workloads.
- Build evaluation and regression-detection infrastructure to track finding quality, precision/recall, and cost-per-finding over time.
- Own observability for the AI stack - self-hosted Langfuse tracing, Grafana/OTel dashboards, and alerting on latency, error rates, and GPU health.
- Define AI infrastructure readiness for new customer deployments - GPU/driver validation, capacity planning, and standardized onboarding runbooks.
Requirements
Technical Skills
- Hands-on experience running LLM inference at scale (vLLM or similar) in production.
- Deep GPU knowledge - CUDA, NCCL, multi-GPU topologies, and attention-backend/quantization tuning (FlashAttention, PagedAttention, AWQ/GPTQ).
- Strong Kubernetes and Helm experience across managed and self-hosted clusters (OpenShift, EKS, AKS, GKE, bare metal).
- Experience with LLM gateways/routing (LiteLLM or equivalent) and observability tooling (Langfuse, Grafana, OTel, Prometheus).
- Solid scripting/automation skills (Python) for building eval harnesses and benchmarking pipelines.
- Familiarity with model lifecycle management - versioning, canary rollouts, and rollback strategies.
Soft Skills
- Strong ownership mentality - comfortable being the sole owner of a critical, customer-facing system.
- Able to work independently in constrained, air-gapped, or highly regulated environments.
- Clear communicator who can translate infrastructure decisions into customer-facing runbooks and dashboards.
- Structured, benchmark-driven approach to performance and cost optimization.
Nice to Have / Advantage
- Experience with AMD GPU architectures (ROCm/MI300X) alongside NVIDIA.
- Background in speculative decoding or advanced inference-serving research.
- Prior experience in security-sensitive or telecom/service-provider environments.
- Experience standing up BYOC (Bring Your Own Cloud) deployment models for enterprise customers.
If your experience is close but doesn't fulfil all requirements, please submit your application. DriveNets is on a mission to build a special company comprised of individuals with different backgrounds, perspectives, and experiences.
DriveNets is an equal opportunity employer. We do not discriminate based on upon race, religion, national origin, sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with disability, or other applicable legally protected characteristics.
More About DriveNets
Based in Israel with extended teams located in the US, Japan, and Romania, DriveNets operations cover more than twelve countries globally. Powering production networks for global tier-1 operators, DriveNets is a leader in large-scale networking solutions for AI infrastructure and service providers. Visit our website to learn more:
https://drivenets.com/company/