Skip to content

ML Infrastructure Engineer vs MLOps Engineer: Key Differences, Pay and When to Hire Each

Matt Li By Matt Li Co-Founder and Director 12 min read
TL;DR: An ML infrastructure engineer builds and runs the compute layer that models train and serve on: GPU clusters, schedulers, storage and serving engines. An MLOps engineer builds the model lifecycle on top of that layer: pipelines, registries, evaluation gates, deployment and monitoring. US postings on October 5, 2026 put ML infrastructure base pay at $160,360 to $345,040 and MLOps base pay at $127,500 to $350,000.

Apptronik, the humanoid robot maker in Austin, has both jobs open at once. Its Staff MLOps Engineer sets "the architecture for the platform layer above the training cluster". A Training Infrastructure engineer "owns the cluster layer beneath the platform". I read postings at six US employers to test that line.

Key takeaways
  1. 1Nuro's ML infrastructure team runs infrastructure-as-code pipelines for thousands of GPU and CPU nodes; its postings never mention model monitoring.
  2. 2PathAI's senior MLOps posting lists CI/CD for ML models, feature pipelines and model monitoring as the core of the job.
  3. 3Shield AI's "ML Ops" role is GPU scheduling and air-gapped platform work, so read the duties before the title.
  4. 4The BLS has no wage category for either job; its closest, software developers, had a median of $135,980 in May 2025.

What is the difference between an ML infrastructure engineer and an MLOps engineer?

An ML infrastructure engineer builds the compute platform that machine learning runs on. That means GPU clusters, job schedulers, storage and serving engines. An MLOps engineer builds the path a model takes through that platform: training pipelines, experiment tracking, model registries, evaluation gates, deployment and monitoring.

The first owns the machines. The second owns the model's journey across them.

The two teams serve different customers. The ML infrastructure team serves researchers and ML engineers who need compute. The MLOps team serves the same engineers once they have a model to ship, and the product teams that depend on it. MLOps Now describes MLOps work as "the operational aspects" of running models in production.

Small companies merge the two, and the cloud vendors' managed ML services cover much of the infrastructure half. Our ML infrastructure engineer role guide and MLOps engineer role guide cover each job on its own. If you are weighing MLOps against model building instead, read MLOps engineer vs machine learning engineer.

Responsibilities: who owns which layer

Nuro's ML infrastructure posting and PathAI's senior MLOps posting split the work along the same line. Nuro, a self-driving company, wants an engineer to provision compute and schedule jobs. PathAI, which builds AI for pathology labs, wants one to move models into production and keep them healthy.

ML infrastructure engineer (Nuro posting)
  • Scale infrastructure-as-code pipelines across thousands of GPU/CPU nodes
  • Optimize workload scheduling to cut job wait times and raise hardware use
  • Build pipelines for petabyte-scale sensor and telemetry data
  • Run feature caching and storage
  • Abstract cloud infrastructure behind one ML platform
MLOps engineer (PathAI posting)
  • Automate CI/CD for ML models and feature pipelines
  • Deploy and monitor models in a reproducible way
  • Evaluate and adopt new MLOps tools
  • Bridge research and production with ML and data science teams
  • Lead design of the MLOps suite for security and compliance

Both lists include infrastructure. PathAI's MLOps engineer still builds "infrastructure and automation, in AWS and on-premises". The difference is the unit of work. Nuro measures its engineer on nodes, queues and data throughput. PathAI measures its engineer on models: how they ship, how the team watches them, and whether anyone can repeat a deployment.

Apptronik spells out the same layers for a robot company. Its Staff MLOps posting lists the platform pieces that sit above the cluster:

Cluster layerInfrastructure
Training Infrastructure
Owns "the cluster layer beneath the platform"
Platform layerMLOps
Datasets to registry
Dataset lifecycle, experiment tracking, model registry, evaluation harnesses
RobotMLOps
Trained to deployed
Promotion path from "trained" to "qualified" to "deployed to robot", with rollback

The MLOps role at Apptronik also owns reproducibility. It must trace each trained model "back to the exact data and code that produced it." No ML infrastructure posting I read carried that duty. Lineage belongs to whoever owns the model's path.

Why the job title is a weak guide

Shield AI, a defense autonomy company, advertises a Senior Staff Engineer, ML Ops. Most of its duties are infrastructure. The role builds a Kubernetes-native platform, runs shared GPUs across cloud and on-premises sites, and schedules GPU jobs with KAI. It also deploys to "fully air-gapped systems". It asks for experience operating GPU-accelerated infrastructure and packaging it with Terraform and Helm.

The overlap runs the other way too. Apptronik's Senior Software Engineer, ML Infrastructure builds the model store and the promotion path to the robot. Researchers and engineers on MLOps, Autonomy, Data Platform and TeleOp teams use its platform daily. That is lifecycle work under an infrastructure title.

How the postings were picked. I searched public Greenhouse, Lever and Ashby job boards on October 5, 2026 for titles containing ML infrastructure or MLOps, and kept US roles with a posted base range. Postings open and close weekly, so treat the figures as a snapshot.

Titles keep splitting. In its Platform Engineering predictions for 2026, Platform Engineering expects the platform engineer title to break into specialisms, "AI-focused platform engineers" among them. For a job description, start from the layer you need covered and pick the title second.

Skills and tooling

Both roles share a base: Kubernetes, infrastructure as code and Python. PathAI's MLOps engineer needs Kubernetes, Helm and Terraform, while Nuro's ML infrastructure engineer can bring Terraform, Pulumi or Crossplane. The split shows in what each posting adds on top. On the MLOps side, Neptune.ai describes the role as a blend of software engineering and ML knowledge: the engineer has to understand how a model works to operate it well.

Skill areaML infrastructure engineerMLOps engineer
ComputeGPU cluster management, schedulers such as Ray, KubeRay, Slurm or Volcano (Nuro)Runs workloads on the cluster through Kubernetes and Airflow (PathAI)
Kubernetes depthAdministration with custom resource definitions and operators (Abridge)Service design on Kubernetes and the contract with compute below (Apptronik)
DataSpark or Beam for petabyte-scale extraction; feature stores such as Feast (Nuro)Dataset versioning with DVC, LakeFS or Delta (Apptronik)
Model toolingServing engines such as Triton, vLLM and TRT-LLM (Abridge)Experiment tracking with MLflow or W&B, model registry (Apptronik)
PerformanceKV-cache design, quantization, custom GPU kernels (Roblox)Packaging with ONNX, TensorRT or torch.compile for deployment (Apptronik)
MonitoringCluster health, GPU use, storage bottlenecksPrometheus, Grafana or Datadog plus model monitoring (PathAI)
ML knowledgeEnough to tune training and inference systemsPyTorch or scikit-learn preferred (PathAI)

The ML infrastructure side is drifting toward systems programming. Roblox's Principal ML Infrastructure Engineer works with FSDP, vLLM, SGLang and CUDA, and develops custom kernels where needed. Abridge's ML infrastructure posting centres on model serving and GPU use for its clinical documentation models.

The MLOps side draws on DevOps habits. People in AI, a recruiter, lists the skills employers ask for. They include cloud platforms, Kubernetes, Kubeflow, MLflow and CI/CD tools such as Jenkins and GitHub Actions. If your team runs hosted language models, the related LLMOps engineer covers prompts, evals and provider costs. To test Kubernetes depth for either hire, start with our Kubernetes interview questions.

A typical day for each role

Neither company publishes a diary, so I built the day below from the duty lists in the Nuro and PathAI postings. Treat it as a sketch of where the hours go.

Morning, ML infrastructure
Queue and cluster check: why training jobs waited overnight, which GPU nodes failed, what the scheduler did with them
Afternoon, ML infrastructure
Infrastructure as code: a Terraform change for a new node pool, a storage fix so data loading keeps GPUs busy
Morning, MLOps
Model health: monitoring dashboards for drift and latency, then a failed pipeline run for a retrained model
Afternoon, MLOps
Release work: packaging a data scientist's model and adding an evaluation gate before it ships

The ML infrastructure engineer's customers are inside the company, and an outage shows up as idle researchers and a growing job queue. The MLOps engineer's failures reach users: a model that degrades in production, or a release the team cannot roll back. So PathAI asks its MLOps hire to own monitoring. Nuro asks its infrastructure hire about storage bottlenecks.

Pay: what the postings say

Here are the posted base ranges from the seven postings, with the minimum experience each one asks for:

ML infrastructure engineer
Nuro, 3+ yrs$160K to $241K
Nuro senior, 4+ yrs$194K to $291K
Abridge, 5+ yrs$221K to $260K
Roblox principal, 5+ yrs$295K to $345K
MLOps engineer
PathAI senior, 5+ yrs$128K to $196K
PathAI assoc. director, 8+ yrs$182K to $278K
Shield AI senior staff$233K to $350K
Posted annual base pay, scale $0 to $360,000. Locations: Nuro Mountain View, Abridge San Francisco, Roblox San Mateo, PathAI Boston, Shield AI San Diego. Source: company postings read October 5, 2026.

At matched seniority, the ML infrastructure postings pay more. Nuro's senior range starts at $193,930 for four years of experience, while PathAI's senior MLOps range tops out at $195,500 for five. Location explains part of that gap: Nuro and Abridge hire in the Bay Area, PathAI in Boston.

The MLOps range climbs once the role takes on a team. PathAI's associate director posting pays up to $278,300 to lead a team of 6 to 7+ engineers.

Shield AI posts a second, higher range for the same role in San Mateo: $280,000 to $420,000. All of these figures are base pay. Nuro and Shield AI add a bonus and equity, and Roblox adds equity. Salary aggregators such as Glassdoor publish self-reported MLOps pay, but neither title has a federal wage category. The Bureau of Labor Statistics puts the median software developer wage at $135,980 in May 2025, below the bottom of most ranges above. For costs outside the US, see our MLOps engineer cost guide and ML engineer cost guide.

Career paths

ML infrastructure engineers often arrive from backend platform work or distributed systems. Nuro accepts three years in ML infrastructure, backend platform engineering or distributed systems. The ladder runs toward systems depth. Roblox's principal role asks for 3+ years tech-leading ML infrastructure engineers and owns the architecture for training, serving and features. Our platform engineer and site reliability engineer guides cover the neighbouring rungs.

MLOps engineers come from software engineering, DevOps and data engineering. PathAI's senior posting asks for 5+ years of software engineering and lists ML frameworks only as a preference. The ladder runs toward ownership of the model lifecycle and then a team. PathAI's associate director role needs 8 to 10+ years with 4+ managing engineers.

People in AI reports that titles such as "MLOps Lead" and "Head of MLOps" are becoming more common. It also cites LinkedIn's Emerging Jobs report: MLOps grew 9.8 times in five years. The DevOps engineer and ML model deployment engineer guides map the entry routes.

Seniority bands follow other engineering ladders. Our comparison of a lead engineer vs senior engineer explains where the team-leading step sits.

When to hire which

Hire the MLOps engineer first if your models run on a managed cloud service. Your problem then is getting models to production and keeping them there. The pipelines, registry and monitoring pay off from the first model in production, and the cloud provider runs the cluster.

Hire the ML infrastructure engineer first if you own or rent GPU capacity yourself, train large models, or serve them under tight latency targets. Abridge, Roblox and Nuro post for this layer to run large-scale training and low-latency serving. Signs you need one: long job queues, idle GPUs, storage bottlenecks and GPU bills you cannot attribute to a team.

6 to 7+
Engineers on PathAI's MLOps team under one associate director
5 to 10x
Growth in AI runs PathAI wants its inference stack to support
4+ yrs
Owning an MLOps platform end to end, Apptronik's alternative to 8+ years of general ML platform work
Source: PathAI and Apptronik postings, read October 5, 2026.

Most teams end up with both once training and serving grow. PathAI's MLOps lead must "modernize the AI Product inference stack to support 5-10x growth of AI runs". The lead also works with site reliability engineers on cost metrics. At that point the work no longer fits one person. Platform Engineering expects the pipelines to meet. Its Platform Engineering predictions say that mature platforms will offer "a single delivery pipeline serving app developers, ML engineers, and data scientists" by the end of 2026.

Whoever you hire will lean on AI coding tools; PathAI's posting asks for experience with Copilot or Cursor. Our guide to AI-native vs traditional engineers covers how to test for that. For another pair that splits on build versus outcome, see deployment strategist vs forward deployed engineer.

Hire ML infrastructure and MLOps engineers from Asia

We match you with vetted machine learning engineers, DevOps and MLOps engineers and Kubernetes engineers at 50-70% below US cost. Each hire carries a 90-day, one-time replacement guarantee. We employ them through our own EOR in 9 Asian markets, including Vietnam, the Philippines and Malaysia. The Second Talent Monthly plan is $4,999 a month for a senior engineer, all-inclusive, with $2,999 for the first month. Tell us which layer you need covered and we will send a shortlist.

Frequently Asked Questions

Is an ML platform engineer the same as an ML infrastructure engineer?

Often, but check the duties. Nuro's ML infrastructure team builds "a unified ML platform", while Apptronik's MLOps role owns its platform layer. If the posting centres on clusters and scheduling, it is infrastructure. If it centres on registries and releases, it is MLOps.

Can a DevOps engineer move into either role?

Yes. Both sets of postings ask for Kubernetes and Terraform, and PathAI adds Prometheus, Grafana or Datadog. MLOps adds model lifecycle knowledge; ML infrastructure adds GPU scheduling and distributed systems.

Can one engineer cover both at a startup?

Yes, while you run a few models on managed cloud services. Split the roles once you operate your own GPU capacity or ship models on a regular cadence.

Hire machine learning engineers.

Pre-vetted senior engineers from Asia at 50 to 70% below US hiring costs, with first profiles in 24 hours.

Hire ML engineers Apply as talent →
Matt Li

Written by

Matt Li is a tech-driven entrepreneur with deep expertise in global talent strategy, digital experience optimization, e-commerce, and Web3 innovation. He is the Co-Founder of Second Talent, a US-based company that connects businesses with top-tier tech professionals worldwide. Since launching the company in 2024, Matt has led its growth by leveraging technology to streamline remote hiring and scale distributed teams. With a background spanning product, operations, and innovation, Matt brings a cross-disciplinary perspective to the evolving digital economy. His work sits at the intersection of global talent, emerging technology, and scalable digital transformation.

More posts by Matt Li →

Keep Reading

Loading available times…