Senior AI Ops Engineer
Description
Hybrid | Israel
About the Company
DriveNets is a leader in large-scale networking solutions for AI infrastructure and service providers. The company's disaggregated networking architecture transforms the economics of large-scale infrastructures while maximizing performance, utilization, and operational efficiency. Its high-performance AI fabric maximizes GPU utilization and accelerates deployments by optimizing the AI stack end-to-end, resulting in higher tokens-per-second and lower cost-per-token. DriveNets' solutions power production networks for global tier-1 operators like AT&T and Comcast, and scale multi-vendor AI infrastructures at foundation model labs, NeoClouds, and enterprises.
Responsibilities
- Deploy and tune vLLM inference servers across customer environments on diverse GPU hardware (H100, A100, L40S, MI300X), including fully air-gapped, offline deployments.
- Optimize inference performance end-to-end - batching, tensor/pipeline parallelism, KV-cache and prefix caching, quantization (FP16/INT8/AWQ/GPTQ), and hardware-specific tuning across NVIDIA and AMD architectures.
- Maintain the LiteLLM gateway and Helm chart variations across Kubernetes flavors (OpenShift, EKS, AKS, GKE, bare metal).
- Own the model lifecycle - registry, canary rollouts, upgrade/rollback playbooks, and benchmarking of new open-weight models against our security-reasoning workloads.
- Build evaluation and regression-detection infrastructure to track finding quality, precision/recall, and cost-per-finding over time.
- Own observability for the AI stack - self-hosted Langfuse tracing, Grafana/OTel dashboards, and alerting on latency, error rates, and GPU health.
- Define AI infrastructure readiness for new customer deployments - GPU/driver validation, capacity planning, and standardized onboarding runbooks.
Requirements
- Technical Skills
- Hands-on experience running LLM inference at scale (vLLM or similar) in production.
- Deep GPU knowledge - CUDA, NCCL, multi-GPU topologies, and attention-backend/quantization tuning (FlashAttention, PagedAttention, AWQ/GPTQ).
- Strong Kubernetes and Helm experience across managed and self-hosted clusters (OpenShift, EKS, AKS, GKE, bare metal).
- Experience with LLM gateways/routing (LiteLLM or equivalent) and observability tooling (Langfuse, Grafana, OTel, Prometheus).
- Solid scripting/automation skills (Python) for building eval harnesses and benchmarking pipelines.
- Familiarity with model lifecycle management - versioning, canary rollouts, and rollback strategies.
- Soft Skills
- Strong ownership mentality - comfortable being the sole owner of a critical, customer-facing system.
- Able to work independently in constrained, air-gapped, or highly regulated environments.
- Clear communicator who can translate infrastructure decisions into customer-facing runbooks and dashboards.
- Structured, benchmark-driven approach to performance and cost optimization.
- Nice to Have / Advantage
- Experience with AMD GPU architectures (ROCm/MI300X) alongside NVIDIA.
- Background in speculative decoding or advanced inference-serving research.
- Prior experience in security-sensitive or telecom/service-provider environments.
- Experience standing up BYOC (Bring Your Own Cloud) deployment models for enterprise customers.
- If your experience is close but doesn't fulfil all requirements, please submit your application. DriveNets is on a mission to build a special company comprised of individuals with different backgrounds, perspectives, and experiences.
- DriveNets is an equal opportunity employer. We do not discriminate based on upon race, religion, national origin, sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with disability, or other applicable legally protected characteristics.
- More About DriveNets
- Based in Israel with extended teams located in the US, Japan, and Romania, DriveNets operations cover more than twelve countries globally. Powering production networks for global tier-1 operators, DriveNets is a leader in large-scale networking solutions for AI infrastructure and service providers. Visit our website to learn more:
- https://drivenets.com/company/
Every role in the library, ranked against your CV.
Rewritten from your real, matching experience for the job you pick.
Your whole pipeline in one place, with a fresh move every morning.